Natural language processing method, natural language processing system, and natural language processing program
The system addresses the challenge of unclear grammatical definitions in Japanese by using a word dictionary and templates to enhance semantic recognition, improving the naturalness of human-computer dialogue.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SEAMAN ARTIFICIAL INTELLIGENCE LAB CO LTD
- Filing Date
- 2024-07-02
- Publication Date
- 2026-07-30
AI Technical Summary
Japanese natural language processing systems struggle with accurately recognizing dependency relationships and informal expressions due to the lack of clear grammatical definitions, particularly in colloquial contexts, leading to unnatural interactions in human-computer dialogue.
A natural language processing system that utilizes a word dictionary with attribute information and predefined templates to analyze input sentences, allowing for the classification of elements and their combinations, enabling accurate semantic recognition and generation of natural responses.
This approach enhances the naturalness of human-computer dialogue by accurately recognizing dependency relationships and informal expressions, resulting in more appropriate and contextually relevant responses.
Smart Images

Figure 0007897610000001 
Figure 0007897610000002 
Figure 0007897610000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method for natural language processing of Japanese, a natural language processing system, and a natural language processing program.
Background Art
[0002] Conventionally, a technology for a computer that performs natural language processing to interact with a person is known. By performing natural language processing suitable for each language, it is possible to realize a more natural interaction between a computer and a person, and in particular, a natural language processing technology suitable for Japanese is known.
[0003] Patent Document 1 discloses a technique for more appropriately complementing functional words that need to be supplemented in an intermediate predicate part. Patent Document 1 assigns semantic labels to the functional expressions of the predicate part and further classifies them into three types: Mod, Foc, and T. A tense determination unit 71 determines whether it is necessary to complement any functional expression for an intermediate predicate part whose meaning, type, and the conjunction following the intermediate predicate part are the complementation target. A complementation processing unit 72 compares the semantic label of the functional expression of the intermediate predicate part to be complemented with the semantic label of the functional expression of the complementation source predicate part immediately after the intermediate predicate part. From the type (Mod, Foc, T) of the functional expression of the intermediate predicate part, the "missing" functional expression of the intermediate predicate part is determined. It is disclosed that only what the intermediate predicate part lacks is complemented from the functional expression of the subsequent predicate part.
[0004] Patent Document 2 discloses a technology that can generate appropriate question sentences at low cost. Patent Document 2 discloses that a question sentence candidate generation unit 29 generates question sentence candidates by taking at least one label as input to a template created from a question sentence, in which the words contained in the question sentence are left blank and the parts of speech and semantic attributes of the words are assigned to the blanks, and replacing the blanks with words characteristic of the input label that correspond to the parts of speech and semantic attributes assigned to the blanks. A question sentence evaluation unit 30 calculates a score representing plausibility for each question sentence candidate using a language model corresponding to the input label, and outputs the question sentence candidate with the highest plausibility as the question sentence. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2011-164678 [Patent Document 2] Japanese Patent Publication No. 2017-27233 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] The problem that this invention aims to solve is to provide a novel technology for more natural dialogue in Japanese natural language processing. [Means for solving the problem]
[0007] To solve the problems described above, the present invention provides a natural language processing method in which a computer stores in a memory a word dictionary in which attribute information is assigned to words, which are elements that have meaning in themselves, and a template in which combinations of multiple elements are defined, and selects a template corresponding to the input sentence based on the combination of elements that constitute the input sentence.
[0008] Furthermore, the present invention relates to a natural language processing system comprising: a storage unit that stores a word dictionary in which attribute information is assigned to words, which are elements that have meaning in themselves; and templates in which combinations of multiple elements are defined; and a template processing unit that selects a template corresponding to the input sentence based on the combination of elements that constitute the input sentence.
[0009] Furthermore, the present invention relates to a natural language processing program in which a computer storing a word dictionary in which attribute information is assigned to words, which are elements that have meaning in themselves, and templates in which combinations of multiple elements are defined, functions as a template processing unit that selects a template corresponding to an input sentence based on the combination of elements that constitute the input sentence.
[0010] In a preferred embodiment of the present invention, the elements include elements that have meaning in combination with words, the template has one or more groups of the words and groups of 0, 1 or more of the elements as comparison targets and positional information indicating the order in which the comparison targets are arranged in the input sentence, and the template is selected by comparing the groups to which the elements constituting the input sentence belong according to the order in which they are arranged.
[0011] Japanese lacks clear grammatical definitions regarding word order, making it difficult to recognize dependency relationships between words when word order changes. Furthermore, Japanese lacks clear grammatical definitions for colloquial, informal expressions. While some informal expressions contain important meanings that significantly impact the overall meaning of a sentence, natural language processing has historically struggled to accurately recognize their meanings. The aforementioned structure allows for the classification of input sentences as combinations of attributes, enabling abstract recognition of dependency relationships and informal expressions.
[0012] In a preferred embodiment of the present invention, the template includes a dialogue template for recognizing an input sentence from a combination of elements contained in the input sentence, and a word template for recognizing a plurality of consecutive elements contained in the input sentence as a single unit element. The plurality of consecutive elements constituting the input sentence are grouped into a single unit element by applying the word template, and the dialogue template is selected using the element grouped by the word template.
[0013] This configuration allows for the distinction between words and elements that modify other words and the words being modified, enabling more accurate semantic recognition of the input sentence.
[0014] In a preferred embodiment of the present invention, instruction information for a response to an input sentence is further stored in association with the template, and a response to an input sentence is generated according to the instruction information associated with the selected template. In a preferred embodiment of the present invention, the instruction information includes any of the following processes using attribute information of specific elements contained in the input sentence: substitution with words having common attributes, conversion of element conjugations, conversion of words to higher-level concepts, or conversion of words. In a preferred embodiment of the present invention, the instruction information is an instruction to generate a response using a specific template. In a preferred embodiment of the present invention, the attribute information includes a first attribute that defines the meaning of a word by a combination of classifications having a hierarchical structure of two or more layers, and the instruction information is an instruction to generate a response sentence using a predetermined word that is identified by manipulating the positional relationship of the first attribute of the word contained in the input sentence in the hierarchical structure. In a preferred embodiment of the present invention, the elements include elements that have meaning when combined with words, an element dictionary is stored in the memory unit in which attribute information indicating the meaning of the elements is assigned to the elements, and the instruction information is an instruction to generate a response sentence using a predetermined element that is identified by converting the attribute information of the elements included in the input sentence into predetermined attribute information. In a preferred embodiment of the present invention, the template includes one or more words and positional information of those words within a sentence, and the instruction information is an instruction to generate a response sentence that converts the positional information of a predetermined word included in the input sentence into another positional information.
[0015] Even when discussing a common subject, the way it is expressed varies depending on the speaker. In human-computer interaction, if the computer's expression of the subject is inappropriate, the user will perceive the interaction as unnatural. By adopting this structure, the meaning of the input sentence is abstracted by recognizing it as a combination of attribute information, a favorable response sentence for that meaning is generated, and a natural response sentence that reflects the changes in the expression of the subject is produced.
[0016] In a preferred embodiment of the present invention, groups of elements are determined for the elements that constitute the input sentence, and a new template is generated based on the groups of elements. In a preferred embodiment of the present invention, the word dictionary includes, as attribute information, a first attribute that defines the meaning of a word by a combination of classifications having two or more hierarchical structures, and generates a new template based on a group of words defined by the first attribute of the words constituting the input sentence. In a preferred embodiment of the present invention, an element dictionary, in which attribute information is assigned to elements that give meaning when combined with other words, is stored in a memory unit, and a new template is generated based on a group of elements defined by the attribute information of the elements that constitute the input sentence.
[0017] Furthermore, the present invention relates to a computer that stores element data representing elements constituting a document, element groups representing a collection of the element data, and a dialogue template in which one or more template elements including the element data and / or the element groups are arranged, as well as attributes and attribute values for each of the element data, in a storage unit. Determining the question template, which is a dialogue template corresponding to the question, The generation of estimated response sentences based on the aforementioned question template is performed, In the generation of the estimated response sentence, a process of generating an estimated response sentence by converting the template elements arranged on the response template for the question template based on the attribute and / or attribute value, a process of generating a response template by performing stretching to replace some of the template elements arranged on the question template with other template elements, and generating an estimated response sentence based on the template elements arranged on the generated response template, or, a process of searching for another response template in which the question template and some of the template elements are common, and generating an estimated response sentence based on the template elements arranged on the searched response template, is a natural language processing method for generating the estimated response sentence by one or more of the above.
[0018] With such a configuration, an estimated response sentence for the question sentence can be generated.
[0019] In a preferred form of the present invention, further, when the computer storing the generated estimated response sentence receives an input response sentence from a user, searches for the estimated response sentence based on the input response sentence, and obtains the attribute and / or attribute value corresponding to the input response sentence.
[0020] With such a configuration, the attribute and / or attribute value of the input response sentence can be obtained by searching for the estimated response sentence.
[0021] In a preferred form of the present invention, the attribute and / or attribute value corresponding to the estimated response sentence is output in association with the estimated response sentence, and a registration screen for an action to be referred to and executed is displayed by referring to these, stores the conditions related to the input attribute and / or attribute value and the action to be executed, and when receiving an input response sentence from a user, obtains the attribute and / or attribute value corresponding to the input response sentence and executes the action.
[0022] This configuration allows you to set the actions to be performed based on the input response, using the estimated response as a reference.
[0023] Furthermore, the present invention is a computer program that causes the computer to execute the natural language processing method described above.
[0024] Furthermore, the present invention includes a storage unit that stores element data representing elements constituting a document, element groups representing a collection of the element data, and a dialogue template in which one or more template elements including the element data and / or the element groups are arranged, and attributes and attribute values for each of the element data. A template processing unit that determines a question template, which is a dialogue template corresponding to the question text, The system includes a response text estimation unit that generates estimated response texts based on the aforementioned question template, The aforementioned response text estimation unit, A process of generating an estimated answer sentence by transforming the template elements arranged on the answer template for the question template based on the attributes and / or attribute values, A process of generating an answer template by performing stretching, which replaces some of the template elements arranged on the question template with other template elements, and then generating an estimated answer sentence based on the template elements arranged in the generated answer template, or A process that searches for other answer templates that share some template elements with the question template, and generates an estimated answer sentence based on the template elements arranged in the searched answer template. This is a natural language processing system that generates the estimated response sentence using one or more of the following: [Effects of the Invention]
[0025] According to the present invention, it is possible to provide a novel technology for more natural dialogue in Japanese natural language processing. [Brief explanation of the drawing]
[0026] [Figure 1] Configuration diagram of the natural language processing system of Embodiment 1. [Figure 2] Hardware configuration diagram of Embodiment 1. [Figure 3] Conceptual diagram of words, elements, etc., in Embodiment 1. [Figure 4] Data structure diagram of the word dictionary in Embodiment 1. [Figure 5] Data structure diagram of the word dictionary in Embodiment 1. [Figure 6] Data structure diagram of the usage dictionary for Embodiment 1. [Figure 7] A schematic diagram of the components of the dialogue data in Embodiment 1. [Figure 8] Schematic diagram of the Word template for Embodiment 1. [Figure 9] A schematic diagram of the dialogue template for Embodiment 1. [Figure 10] A flowchart illustrating the process for registering a dialogue template in Embodiment 1. [Figure 11] A flowchart illustrating the process for registering a Word template in Embodiment 1. [Figure 12] A flowchart illustrating the process of interacting with the user in Embodiment 1. [Figure 13] A flowchart illustrating the process from the input statement of Embodiment 1 to the identification of the dialogue template. [Figure 14] Configuration diagram of the natural language processing system of Embodiment 2. [Figure 15] A processing flowchart relating to the estimation process of the answer statement in Embodiment 2. [Modes for carrying out the invention]
[0027] <Embodiment 1> The following description of a natural language processing system according to Embodiment 1 of the present invention will be explained with reference to the drawings. Note that the embodiments shown below are examples of the present invention, and the present invention is not limited to these embodiments; various configurations can be adopted.
[0028] This embodiment describes the configuration and operation of a natural language processing system, but similar configurations, devices, computer programs, and program recording media storing such programs can achieve similar effects. The series of processes according to this embodiment, described below, are provided as a computer-executable program and can be provided via non-transient computer-readable recording media such as CD-ROMs and flexible disks, as well as via communication lines. In this embodiment, the program may also be implemented in a so-called cloud computing manner, where the program is launched on an external computer to realize its functions on a client terminal.
[0029] A natural language processing system consists of a computer device. This computer device includes an arithmetic unit such as a CPU (Central Processing Unit) and a memory device. The computer device can function as a natural language processing unit by executing a natural language processing program stored in the memory device using its arithmetic unit. The natural language processing method is implemented through the processing of the computer device, including the natural language processing unit.
[0030] <1.1. System Configuration> Figure 1 shows an example of a system configuration diagram of natural language processing system 1. As shown in Figure 1, natural language processing system 1 comprises an input sentence acquisition unit 10, a natural language processing unit 20, a response sentence output unit 30, and a database DB.
[0031] The input sentence acquisition unit 10 acquires text data as input sentences from the input data. The input sentence acquisition unit 10 outputs the input sentences to the natural language processing unit 20. Input data is input from the interlocutor, as well as from another natural language processing system, such as natural language processing system 1' or language generation AI systems such as LLM (Large Language Models). Input data may be audio data, image data, or text data. If the input data is audio data, the input sentence acquisition unit 10 performs transcription (text recognition) on the audio data to acquire the input sentences. If the input data is image data, the input sentence acquisition unit 10 performs OCR processing on the image data to acquire the input sentences.
[0032] The natural language processing unit 20 obtains the input text from the input text acquisition unit 10 and generates a response text by processing the input text in natural language. The natural language processing unit 20 outputs the response text corresponding to the input text to the response text output unit 30. In this embodiment, the natural language processing unit 20 includes a registration unit 21, a splitting unit 22, a word processing unit 23, a template processing unit 24, and a response text generation unit 25. In this embodiment, the registration unit 21, the splitting unit 22, and the word processing unit 23 are used to register a new template.
[0033] The response output unit 30 obtains the response from the natural language processing unit 20 and outputs output data generated from the obtained response. The output data is either audio data or text data.
[0034] The database DB is configured to communicate with the natural language processing unit 20.
[0035] <1.2. Hardware Configuration> In one embodiment, the natural language processing system 1 is configured as a dialogue device 100 and a natural language processing device 200. The dialogue device 100 includes a dialogue robot or a terminal device, and comprises an input sentence acquisition unit 10 and a response sentence output unit 30. The natural language processing device 200 includes a server device, and comprises a natural language processing unit 20. The dialogue device 100 and the natural language processing device 200 are connected to a communication network and configured to enable data communication.
[0036] Figure 2(a) shows an example of a hardware configuration diagram of the dialogue device 100. The dialogue device 100 comprises a communication unit 101, a control unit 102, a storage unit 103, an input unit 104, and an output unit 105 as its hardware configuration.
[0037] The communication unit 101 controls communication with the communication network and provides the necessary inputs and output data related to the operation results for the dialogue device 100.
[0038] The control unit 102 includes one or more processors such as a CPU, and controls the entire operation process of the interactive device 100 by executing the interactive program, OS, and other applications. The control unit 102 makes the computer device function as the interactive device 100 by executing processing based on the interactive program stored in the memory unit 103, thereby realizing the functional components described later.
[0039] The storage unit 103 is an HDD (Hard disk drive), ROM (Read Only Memory), RAM (Random Access Memory), etc., and stores the interactive program and data used by the control unit 102 when it executes processing based on the program.
[0040] The input unit 104 inputs input data, including voice data or text data, to the control unit 102. The input unit 104 consists of a microphone for acquiring voice data, a keyboard for acquiring text data, a touch panel, and the like.
[0041] The output unit 105 outputs output data that includes audio data or text data. The output unit 105 consists of a speaker that outputs audio data, a display that displays text data, and the like.
[0042] Figure 2(b) shows an example of the hardware configuration of the natural language processing unit 200. The natural language processing unit 200 comprises a communication unit 201, a control unit 202, and a storage unit 203 as its hardware configuration.
[0043] The communication unit 201 controls communication with the communication network and provides the necessary inputs and output data related to the operation results for the natural language processing unit 200.
[0044] The control unit 202 includes one or more processors such as a CPU, and controls the entire operation of the natural language processing unit 200 by executing natural language processing programs, an operating system, and other applications. The control unit 202 executes processing based on the natural language processing programs stored in the memory unit 203, thereby enabling the computer device to function as a natural language processing unit 200, and realizing the functional components described later.
[0045] The storage unit 203 is an HDD (Hard disk drive), ROM (Read Only Memory), RAM (Random Access Memory), etc., and stores the natural language processing program and data used by the control unit 202 when it executes processing based on the program.
[0046] <1.3. Database> The database DB comprises a word DB1 for storing word dictionaries, an element DB for storing element dictionaries (including a conjugation DB2 for storing conjugation dictionaries and a suffix DB3 for storing suffix dictionaries), a word template DB4 for storing word templates, a dialogue template DB5 for storing dialogue templates, a response template DB6 for storing response templates, a dialogue scenario DB7 for storing dialogue scenarios, and a word group DB8. The data of elements such as word dictionaries and element dictionaries are referred to as element data.
[0047] <2.1. Words, Elements> Figure 3 is a conceptual diagram showing the definitions of words, elements, etc., in this embodiment. In this embodiment, elements that have meaning in themselves are referred to as "words." Words include, for example, "I," "eat," "today," etc. Words include, in conventional parts of speech, nouns, verbs, adjectives, adjectival nouns, adverbs, attributive adjectives, conjunctions, and interjections. Words are stored in the word dictionary of word DB1.
[0048] In this embodiment, elements that do not have meaning on their own but are included in the input sentence in combination with words are referred to as "elements." Elements include word conjugations such as "watashi" (I am), "no" (my), "taberu" (to eat), and "tabitai" (want to eat). Elements also include particles and auxiliary verbs as conventional parts of speech. Conjugations are stored in the conjugation dictionary of the conjugation DB2.
[0049] Furthermore, elements may also include suffixes that connect to words or other elements (conjugations). Suffixes include, for example, "desu" (I want to eat) and "wa" (I want to eat). Suffixes are stored in the suffix dictionary of Suffix DB3. Conjugation DB2 and Suffix DB3 constitute the element DB.
[0050] <2.2. Words> In this embodiment, a combination of elements that includes at least one or more elements and represents a semantic unit contained in the input sentence is referred to as a "word." A word may be "one or more words only," or it may be "one or more words + one or more elements (inflection and / or suffixes)." Specific examples of words include, for example, "yakiniku bento: word + word," "I: word + inflection," "I want to eat: word + inflection," and "I want to eat: word + inflection + suffix." In addition, there are cases where multiple words and one or more elements are combined to form a word, such as "fingernail bones." Note that "I: word" may also be treated as a word.
[0051] <2.3. Word Groups, Element Groups> Here, a group defining a set of words is called a word group (dotted line in Figure 3), and a group defining a set of elements is called an element group (double-dotted line in Figure 3). In this embodiment, the element group includes a word group, which is a set of one or more words, and an element group, which is a set of one or more elements. Words stored in the word DB1 are assigned one or more word groups. Word groups can be defined by information set on a word, such as attribute groups defined by one or more attribute information (described later) set on the word, or word modification groups to which words that share the same associated element group belong, or by adding words to a group list, etc.
[0052] Furthermore, elements stored in the element DB belong to one or more element groups. Here, each word in the word DB1 is associated with an element group that it references (indicating elements that can be linked). An element group is a group that defines a set of elements that can be linked to the same element. Words that can be combined with elements are what are known as nominative cases and stems. Note that an element group includes an inflection group, which is a set of inflections, and an ending group, which is a set of endings. An inflection group is, for example, a set of multiple inflections that can be applied to a common stem. An ending group is, for example, a set of multiple endings that can be substituted without sounding unnatural in Japanese expression.
[0053] Information for defining word groups, such as a list of words belonging to a word group, may be stored in the word group DB8, or word groups may be defined based on attribute information of words. In this embodiment, a set of words to which the same element group is associated is defined as a word group called a common usage group, and a list of multiple words belonging to the group is stored in the word group DB8 using the group ID, which is the identifier of the common usage group, as the key. By defining common usage groups and usage groups by lists, words and elements are linked in the database DB.
[0054] <2.4.Attribute information> In this embodiment, the natural language processing unit 20 refers to the words and elements stored in the word DB1 and the element DB, and performs natural language processing based on the attribute information of the elements contained in the input sentence. The attribute information is information that indicates the meaning of an element, and includes word attributes related to words and element attributes related to elements.
[0055] Word attributes are information that indicates the meaning or quality of the word itself, and in this embodiment, they include a first attribute and a second attribute. The first attribute is defined as a word category that indicates the semantic content or quality of a word. In this embodiment, the word category is a classification that indicates the meaning of a word and has a hierarchical structure of two or more layers. The meaning of a word is represented by a combination of classifications. The second attribute is defined as a modifying attribute that indicates the semantic content or quality of a word. Modifying attributes are set for each word group called a word modification group, and each group has a modifying attribute item and a modifying attribute value that expresses the modifying meaning of the word using text, numbers, or logical values (truth values). A word modification group includes one or more arbitrarily designated words and / or one or more words determined by one or more arbitrarily designated word categories, and the modifying attributes of a word are determined according to the word modification group set for that word.
[0056] Element attributes are information about elements that complement a word, indicating the speaker's emotions, intentions, nuances, or situations that arise from the element, and in this embodiment, they include the third and fourth attributes. The third attribute is defined as a usage attribute that indicates the meaning of usage. A usage attribute has a usage attribute item and a usage attribute value that expresses the meaning of usage as text, a number, or a logical value (truth value). The fourth attribute is defined as a suffix attribute that indicates the meaning of a word's ending. A suffix attribute is defined as an item of the suffix attribute and a suffix attribute value that expresses the meaning of the ending attached to the word as text, a number, or a logical value (truth value).
[0057] <2.5. Vocabulary> Word DB1 stores a word dictionary that defines multiple words. Words are classified by word categories based on their meanings. In this embodiment, the word categories have a hierarchical structure corresponding to the degree of abstraction of their meanings. The hierarchical structure has two or more layers, and there is no limit to the number of layers.
[0058] <2.6. Word Categories> Figure 4(a) shows an example of the structure of word categories in a word dictionary. A word category has a word category ID, which is a unique identifier for the word category, a related word category ID that indicates related word categories, and a word category name. The related word category includes a word category ID that indicates a parent-child relationship in a hierarchical structure, and in this embodiment, it refers to the parent word category ID. For example, the word category name "brand" has a related word category ID for the word category name "distribution system" which is its parent, and the word category name "distribution system" has a related word category ID for the word category name "food" which is its parent. Note that related word categories are not limited to parent-child relationships, as long as they belong to the same word group, such as grandchild, sibling, or cousin relationships.
[0059] Figure 4(b) is an example of a data structure showing the first attribute of a word in a word dictionary. The example structure is shown. Figure 4(c) is a conceptual diagram showing the first attribute of a word as defined in Figures 4(a) and 4(b). The word categories in Figure 4(c) are defined hierarchically by the related word category IDs shown in Figure 4(a), with the left side showing higher-level word categories and the right side showing lower-level word categories. Lower-level word categories belong to higher-level word categories and are used in combination with higher-level classifications to indicate their meaning. In Figure 4(c), the word "Brand A" (word name) is defined, for example, by a hierarchical structure of word categories, starting from the highest level, such as food, distribution, and brand. For example, in this embodiment, the data shown in Figure 4(b) associates the word "Brand A" with the associated word category ID (10 in Figure 4(a)), which is the identification information of the lowest-level word category to which the word belongs (Brand in Figure 4(c)), and also associates with a word identifier (word ID) that is unique when used alone or in combination with the word category identifier.
[0060] The word categories and the parts of speech classified by the word categories are not limited to the examples shown in Figure 4. In this embodiment, word category names and word names are treated separately by the word category ID and word ID, and words with the same word name may exist under different word categories. Furthermore, each word indicating a word category name (the word category name in Figure 4(a)) may also be treated as a word by assigning a word identifier instead of a word category identifier.
[0061] Figure 5(a) shows an example of a data structure in a word dictionary, illustrating word data and its secondary attributes. Word data is data about a word, identified by a word ID, which is the identifier of the word. The word data, as shown in Figure 5 and the data stored in the word group DB8, has a word category ID and a word group ID, as illustrated in Figure 4. As shown in Figure 5(a), the word data includes a word ID, a word name, the reading of the word name (kana and kana), the part of speech, the word stem, the reading of the stem (kana and kana), the stem conjugation type, and modification attributes.
[0062] The conjugation type indicates the classification of conjugated words, such as verbs, and in this embodiment, it is information for identifying the conjugation table corresponding to the word. It may also be a conjugation table ID, etc.
[0063] Modifier attributes are data associated with words that define additional meanings for those words. Modifier attributes may be defined for each word group; for example, in this embodiment, common modifier attribute items are set for word groups defined by word categories. A word group to which common modifier attributes are assigned is called a word modification group.
[0064] In the illustrated example, when a word modification group is defined by word category ID=1 (word category: food), the word may have modification attributes such as "deliciousness," "junk food level," and "luxury level." Each item of the modification attribute is assigned a modification attribute value; for example, a word for fast food would be set to have a high junk food level, and a word for French food would be set to have a high luxury level. The modification attribute value can be set as text, a number, or a logical value. In this embodiment, a word modification group is set for each word category, but any other set of words may also be used as a word modification group. In the example above, a word modification group is defined by a single word category ID, but a single word modification group may be defined based on the AND or OR conditions of multiple word categories.
[0065] The word group DB8 stores the word group table that defines word groups. The word group table has a word group name and identifiers (word IDs) of one or more words belonging to the word group, with the word group ID as the key, and uses a word group identifier such as the word group ID as the key. In addition, if information indicating a word modification group is stored, the name or identifier of one or more modification attributes assigned to the word modification group, the data type of the modification attribute value, the number of characters, etc. may be stored. A word group may be defined by one or more word categories, and if it is defined by one word category, it does not have to be stored in the word group DB8.
[0066] <2.7. Conjugation Dictionary> The utilization DB2 (element DB) stores a utilization dictionary (utilization table) that defines multiple utilization data and groups thereof. In this embodiment, the utilization table is assigned a utilization table ID and is defined as a utilization group to which utilizations that can be combined with a common word belong. Figure 6 shows an example of the data structure of the utilization dictionary. As shown in Figure 6, the utilization data has a utilization name, utilization attributes, and question phrase markers for each utilization belonging to the group, with the utilization ID, which is a unique identifier for the utilization that is unique either on its own or in combination with the utilization table ID, etc., as the key.
[0067] The conjugation attribute items indicate the meanings that the conjugation encompasses, and are classified into basic, negative, question, future intention, past perfect, completion, desire, regret, antonym, and possibility. Each conjugation attribute has a conjugation attribute value that indicates the strength of its meaning. For example, in conjugations that mean future intention, such as "I intend to eat" and "I decided to eat," "I decided to eat" is defined as encompassing a stronger meaning of "future intention" than "I intend to eat," and is defined as having a future intention attribute value of "1" and an attribute value of "2," respectively.
[0068] The conjugation table is set up according to the conjugation type of the word. The conjugation type indicates the general conjugation form of the word, such as five-stage conjugation or lower-one-stage conjugation. In the illustrated example, the conjugation table is linked to words whose conjugation type is lower-one-stage conjugation, such as "tabe" (to eat) and "ne" (to sleep). In this embodiment, conjugation includes phrases that are conventionally defined as conjugation forms and phrases that are not conventionally defined as conjugation forms, and is defined as a phrase used as part of a conjugation expression of a verb, etc. Taking the verb "taberu" (to eat) as an example, "tabe" is the stem that does not conjugate, and "ru" is the conjugation that can be conjugated. Conjugation is set up according to the conjugation type of the stem. When the conjugation type is upper-one-stage conjugation or lower-one-stage conjugation of a verb, examples of conjugations include "ru", "masu", "nai", "masen", "ru?", "masu?", "masu ka?", "ru ki ga shinai", "ru ki ga shimasen", "you kana", "youtto", "temiru ka", "ru tsumori". Note that the conjugations listed here are just examples, and a variety of conjugations are defined.
[0069] Verb conjugations include colloquial and idiomatic expressions, each with its own meaning. Traditional Japanese verb conjugations defined meanings according to their form, such as the imperfective, conjunctive, terminal, attributive, conditional, and imperative forms. However, traditional conjugation systems did not classify colloquial conjugations, making it impossible to process the meanings encompassed by conjugations as natural language. A conjugation dictionary is a dictionary created with the purpose of processing colloquial and idiomatic expressions as natural language by defining conjugations and conjugation attributes.
[0070] Conjugation is not limited to particles and auxiliary verbs, but can include combinations of traditional parts of speech such as nouns, verbs, adjectives, adjectival nouns, pronouns, adverbs, attributive adjectives, conjunctions, and interjections, or parts of speech that have not been traditionally defined.
[0071] A conjugation dictionary may include dictionaries that define conjugations that assign meaning to parts of speech other than verbs, using them as stems. For example, particles such as "wa," "no," and "ga" that conjugate "person" can be grouped together as a conjugation group. In this case, the conjugation group may be defined, for example, by the word category associated with the word. In a word dictionary, words that have a conjugation type are associated with the corresponding conjugation dictionary.
[0072] <2.8. Word Ending Dictionary> The suffix DB3 stores a suffix dictionary that defines multiple suffix data. Suffix data is data that defines the meaning of a suffix. Suffix data has a suffix ID, which is a unique identifier for the suffix, a suffix name, and suffix attributes.
[0073] The suffix attribute items indicate the meaning that the suffix encompasses, and are classified into basic, negative, question, future intention, past perfect, completion, desire, regret, antonym, and possibility. Each suffix attribute has a suffix attribute value that indicates the strength of its meaning. For example, suffixes that mean "is that so?" and "right?" are defined as having a stronger meaning of "question" when expressed as "right?" than when expressed as "is that so?", and are defined as having question attribute values of "1" and "2" respectively.
[0074] <2.9. Examples of Dialogue Data Components> Figure 7 is a schematic diagram of the components of the dialogue data in this embodiment. In Figure 7(a), code A represents dialogue data (input sentence), code B represents a word, and code C represents an element. Dialogue data A1 is divided into words B1-B3 and elements C1-C3 by referring to the word dictionary and element dictionary. Since each of the words B1-B3 is linked to a common inflection group (element in the element DB) in the word DB1, it is possible to determine word D from the beginning of the input sentence based on combinations of elements such as "word only," "word + element (inflection)," and "word + element (inflection) + element (ending)." By referring to the word dictionary, the first to fourth attributes set for words D1-D3 are determined. Based on the words and elements contained in the word, each attribute is assigned a unique identifier and attribute value (such as word category ID, attribute name, and attribute value) as the first to fourth attributes. For example, for word D1, the attributes can be determined for each word, such as: 1st attribute: word category ID = 7 (person > first person), 2nd attribute: attribute name = attribute value...
[0075] <3.1. Template> A template is information that defines a combination of groups of multiple elements and their order. In this embodiment, a template has one or more thrones. A throne is a comparison target that is compared with one or more words contained in the input sentence, and one or more word groups and / or element groups are defined for each throne. This word group may be a set of words that share common attribute information (e.g., word category), or it may be a set of arbitrary words collected for a specific purpose in the word group DB8. Similarly, an element group may be a set of inflections (elements) defined by an inflection table, or it may be a set of arbitrary elements. The template also has positional information that indicates the order of words in the thrones / input sentence. By comparing the groups set in the thrones and their order with the words in the input sentence and their order, template matching can be performed for the input sentence. The positional information indicates the order of words in the thrones, i.e., the words contained in the sentence, and template matching is performed by comparing the words that make up the input sentence in the order of appearance with the comparison target whose order corresponds to the word.
[0076] In this embodiment, the template includes a word template, a dialogue template, and a response template. The word template is a template for determining the type, grouping, and structure of a word, and is stored in the word template DB4. The dialogue template is a template for understanding the content and type of the input text, and is stored in the dialogue template DB5. The response template is a template for creating a response text, and is stored in the response template DB6.
[0077] <3.2. Word Template> A word template is a template for recognizing multiple elements, such as words modified by other words, as a single word. A word template allows for the determination of combinations of consecutive words, consecutive words, consecutive one or more words, and combinations of one or more elements as words. Figure 8 shows the data structure diagram of the word template data. In this embodiment, a word template includes a word template ID, which is a unique identifier for the word template, a sample sentence, a word template name, and one or more thrones E, which are used for comparison with combinations of consecutive elements in the input sentence, etc. Each throne E includes location information, one or more words and / or elements assigned to the throne. In this embodiment, word groups and element groups are assigned to the thrones, and by referring to these groups, multiple types of words consisting of combinations of a first group of elements arranged consecutively and second and subsequent groups of elements can be recognized. A group is either a word group or an element group, and in this embodiment, a pair of a word group ID and a utilization table ID is specified for each throne.
[0078] Word template ID=1 is a word template that is matched when input text containing words such as "of the hand" or "of the foot" is entered. The first throne E1 corresponds to the word "hand" and the accompanying phrase "of". The first throne E1 has position information of 1 and has a word group ID (=11) that contains "hand" or "foot", and a conjugation table ID (=2) related to "of". In this embodiment, the first throne E1 contains a pair of word group and element group, but the word group and element group may be assigned to the throne separately, as shown for example as thrones E11 and E12. In the illustrated example, the first throne E11 has position information of 1 and has a word group ID (=11) that includes "hand" and "foot", and the second throne E12 has position information of 2 and has a conjugation table ID (=2) related to "no". In this embodiment, groups are assigned to each throne, but as shown for throne E21, in a throne included in a word template, the word itself or the element itself may be assigned instead of some or all of the groups. For example, the first throne E21 has position information of 1 and has the word ID (=111) for "hand" and the conjugation ID (=222) for "no". Furthermore, a group containing only one element may be defined and assigned to a throne.
[0079] Word template ID=2 is a word template that is matched when input text containing words such as "fingers," "back of hand," or "toes" is entered. The first throne E2 corresponds to the word "of hand," and the second throne E3 corresponds to the word "fingers." The first throne E2 has position information of 1 and has a group ID (=11) related to "hand" and a conjugation table ID (=2) related to "of." The second throne E3 has position information of 2 and has a word group ID (=12) for "fingers" and a conjugation table ID (=1) indicating that "fingers" is in its base form (without elements). In the illustrated example, a conjugation table ID=1 is assigned to thrones that do not contain conjugation data. Note that if the target throne does not contain conjugation data, the conjugation table ID may be omitted.
[0080] Word template ID=3 is a word template that is matched when input text containing words such as "fingernail bones," "finger skin," or "toe bones" is entered. The first throne E4 has position information of 1 and has a word group ID (=11) related to "hand" and a conjugation table ID (=2) related to "no." The second throne E5 has position information of 2 and has a word group ID (=12) related to "finger" and a conjugation table ID (=2) related to "no." The third throne E6 has position information of 3 and has a word group ID (=13) related to "bone" and a conjugation table ID (=1) indicating that "bone" is the base form. There is no limit to the number of thrones that make up a word template.
[0081] The splitting unit 22 splits the input sentence acquired by the input sentence acquisition unit 10 into elements. Specifically, the splitting unit 22 performs the splitting process according to the splitting rules shown below. The splitting unit 22 refers to the word dictionary in the word DB1 and identifies the words contained in the input sentence. The splitting unit 22 further refers to the element DB (inflection DB2, suffix DB3) and identifies the elements (inflection, suffix) contained in the input sentence.
[0082] The word processing unit 23 determines a word from the elements contained in the input sentence. The word processing unit 23 refers to the word template DB4 and determines whether the combination of elements arranged sequentially from the first word of the input sentence corresponds to a registered word template. If there is a combination that corresponds to a word template, the word processing unit 23 confirms it as a word based on the word template. Then, the word processing unit 23 searches for words to which a word template is not applied, starting from the beginning of the input sentence, and checks if there is a word or element corresponding to the next word. If a word or element cannot be identified, it is treated as a word consisting only of a single word. If a word or element can be identified, it checks for the presence of further consecutive words or elements. If further consecutive words or elements cannot be identified, for example, it is treated as a word consisting of a word and an element (inflection). If it can be identified, for example, it is treated as a word consisting of a word, an element (inflection), and an element (ending). The splitting unit 22 divides the input sentence into one or more words by repeatedly referring to words and elements sequentially from the beginning to the end of the input sentence. Furthermore, the word processing unit 23 also treats words and / or elements included in the input text to which no word template was applied as words, and proceeds to matching them with the dialogue template.
[0083] <3.3. Dialogue Template> The dialogue template contains required and optional comparison targets, which in this embodiment are distinguished by a required flag. Optional comparison targets indicate words that can be omitted in the input sentence. For example, in response to the question "What did you eat for dinner yesterday?", the input may include dialogue data such as "I ate curry for dinner" or "I ate curry." In this way, words indicating time or object, subjects, etc. (in this example, "yesterday," "I," "for dinner," etc.) may be omitted, and even when they are omitted, the same template can still be used for matching.
[0084] Figure 9 shows the data structure diagram of the dialogue template data. In Figure 8, the dialogue template has a dialogue template ID, which is a unique identifier for the dialogue template, a sample sentence, a dialogue template name, and one or more thrones (W1-4) to be compared. In this embodiment, a throne includes location information, a required flag, and a group of elements to be applied to the throne. The group is a word group or a word group and an element group, and in this embodiment, it is a pair of a word group ID and an incorporation table ID. The required flag is a logical value indicating whether the word corresponding to the throne is required in the dialogue data. For words determined by word template matching, the last word or word and element of the word template is used when matching with the throne in the dialogue template.
[0085] Dialogue template ID=1 is a greeting template. Word group ID=1 contains words such as "good morning" and "hello," and when an input sentence consisting of these words is entered, it is matched (Throne W1). Note that if the target throne does not contain inflection data, the inflection table ID may be omitted, or an ID indicating the base form may be defined. In the illustrated example, thrones that do not contain inflection are assigned inflection table ID=1. For example, as shown for thrones E11 and E12 in Figure 8, groups of elements may be assigned to each throne. Alternatively, as shown for throne E21 in Figure 8, elements themselves may be assigned to a throne included in a dialogue template, instead of some or all of the groups.
[0086] Dialogue template CT2 is a weather question template, and is a template that matches when dialogue data such as "Tell me today's weather" is entered. The first throne W2 has location information of 1, a required flag of false, and has a word group ID related to "today" included in "today's", and a conjugation table ID related to "no". The second throne W3 has location information of 2, a required flag of true, and has a word group ID related to "weather" and a conjugation table ID indicating that "weather" is in its base form. The third throne W4 has location information of 3, a required flag of false, and has a word group ID related to "teach" included in "teach", and a conjugation table ID related to "te".
[0087] The template processing unit 24 compares the words of the input sentence determined by the word processing unit 23 with the dialogue template to determine the dialogue template corresponding to the input sentence. In this embodiment, the template processing unit 24 refers to the dialogue templates in the dialogue template DB5 and selects the template corresponding to the input sentence based on the degree of agreement of the combination of groups to which each element constituting the input sentence belongs. Note that the combination of thrones does not necessarily have to be a perfect match; the dialogue template with the highest agreement rate may be determined as corresponding to the input sentence. The agreement rate may be determined by the agreement ratio of thrones, etc.
[0088] <3.4. Dialogue Scenarios> In this embodiment, the dialogue between a person and a dialogue device proceeds based on a dialogue scenario. The dialogue scenario is data that sets the general flow of the dialogue. Based on the dialogue template identified by the input sentence, the dialogue scenario sets branches and other parameters according to the subsequent dialogue (input sentence and its response sentence). The dialogue scenario is stored in the dialogue scenario DB7.
[0089] <3.5. Instruction information> Instruction information is set in conjunction with a dialogue template and indicates an action to be performed in response to an input sentence. Instruction information may also be linked to a dialogue template within the dialogue scenario. Actions include outputting a standard response sentence, as well as generating and outputting a response sentence. The response sentence generation unit 25 executes an action based on the instruction information linked to the dialogue template.
[0090] <3.6. Generating and Outputting Response Text> The natural language processing unit 20 acquires the input sentence and classifies it into a dialogue template by referring to the dialogue template DB5. The response sentence generation unit 25 then acquires predefined instruction information for the acquired dialogue template and generates a response sentence based on this instruction information. If the action includes generating a response sentence, the instruction information includes a response template that has conversion rules for the words and / or elements placed in the one or more thrones that are referenced. In this case, the response sentence generation unit 25 acquires a response template by referring to the response template DB6.
[0091] <3.7. Answer Template> An answer template is a template for responding to an input sentence and is mapped to a dialogue template by instruction information. An answer template includes references to some or all of the thrones in the mapped dialogue template. For example, when the input sentence "I ate cup noodles" is recognized as a dialogue template "Throne 1: Cup noodles" and "Throne 2: I ate", the answer sentence "Cup noodles" can be generated by directly referencing the word "cup noodles" in Throne 1. An answer template may also include conversion rules for the referenced thrones or dialogue templates (if they include references to all thrones). For example, the word category of the word "cup noodles" in Throne 1 can be made into a higher-level concept to generate the answer sentence "Ramen". It may also include fixed answer sentences in combination with references to the thrones in the dialogue template. For example, the fixed phrase "I ate" can be combined with the word "cup noodles" in Throne 1 to generate the answer sentence "I ate cup noodles".
[0092] <3.8. Conversion Rules> The conversion rules include subject conversion, interrogative conversion, inversion conversion, ellipsis conversion, first attribute conversion, second attribute conversion, third attribute conversion, fourth attribute conversion, etc. However, the conversion rules are not limited to these, and any rules can be set. Conversion rules can be set by specifying the attribute information of elements included in a given throne referenced in the dialogue template. A response template may contain a combination of multiple conversion rules.
[0093] Nominative case conversion is a conversion rule that changes the first, second, or third person of words related to "person" in the input sentence. This allows for conversions that reflect the speaker's and listener's perspectives. Question conversion is a set of conversion rules that transform an input sentence into a question. For example, it is a set of conversion rules that transform or add elements of an input sentence into elements that form a question. Inversion is a rule that moves the positional information of a specific word (e.g., "throne") in the input sentence to the end of the sentence, resulting in an inverted expression. Abbreviation conversion is a rule that omits specific words from the input text.
[0094] The first attribute conversion is a rule that converts the word categories of words contained in the input sentence. The first attribute conversion sets conversion conditions related to the first attribute, such as whether the word categories belong to the same word group or different word groups, and to what level of hierarchy to convert them to. This makes it possible to retrieve words of higher and lower concepts from the word DB1. Second attribute conversion is a rule that converts the modifier attributes of words contained in the input sentence. Second attribute conversion sets conditions for changing the modifier attributes or modifier attribute values of specific words. For example, a value is set to be added to or subtracted from the modifier attribute value of a specific word contained in the input sentence, and another word with the modified modifier attribute value can be determined. Third-party attribute conversion is a rule that converts the utilization attributes of utilization data included in the input sentence. Third-party attribute conversion sets conditions for changing the utilization attribute or utilization attribute value. For example, a value is set to be added or subtracted from the utilization attribute value of a specific utilization included in the input sentence, and another utilization with the modified utilization attribute value can be retrieved from the element database. The fourth attribute conversion is a rule that converts the suffix attribute or suffix attribute value of words contained in the input sentence. The fourth attribute conversion sets conditions for changing the suffix attribute or suffix attribute value. For example, a value is set to be added to or subtracted from the suffix attribute value of a specific word contained in the input sentence, and another word with the modified suffix attribute value can be retrieved from the element database.
[0095] <3.9. Control Commands> Actions based on instruction information may further include control commands to arbitrary devices. Control commands include outputting operation information to operate arbitrary devices, and include, for example, instructions to a predetermined external device (e.g., turning off the lights, opening the curtains, sending an email), instructions to a search device to retrieve information, and control of actuators, displays, and speakers installed in the dialogue device. Operation information may include arbitrary words or phrases included in the input sentence, or it may include the generated response sentence. Control commands to arbitrary devices include, for example, requests to a language generation AI system such as LLM, and the input sentence acquisition unit 10 performs further actions (e.g., generating and outputting text to the dialoguer based on this response) based on the response received from the language generation AI system.
[0096] <4.1. Registration process for new dialogue templates> Next, we will explain the process for registering a new dialogue template based on the input text. Figure 10 shows a flowchart of the process related to dialogue template registration.
[0097] The input text acquisition unit 10 acquires the input text entered from the terminal device (not shown) of the user who is trying to register a new dialogue template (step S11). The splitting unit 22 splits the input text into elements (step S12). The word processing unit 23 refers to the word template DB4 and determines the words included in the input text (step S13).
[0098] The registration unit 21 registers a new dialogue template from the input text. In step S14, the registration unit 21 refers to the word groups of words and element groups of elements contained in the words obtained from the input text in step S13, and associates them with the order in which the words appear in the input text. The order in which the words appear in the input text becomes the position information of the throne in the new dialogue template, and the referred group is linked to the throne.
[0099] In step S15, the dialogue template is stretched as needed. Stretching is a process that changes the content of the template generated according to the input text, such as expanding, contracting, or changing the range of attributes in each throne of the new dialogue template to be registered. In this embodiment, the registration unit 21 presents each group of words included in the input text to the terminal device of the user who intends to register a new dialogue template, enabling the registration of a dialogue template with a changed scope of application. For example, a group of words (e.g., a word category) included in the words in the input text can be used as the word group of the throne in the new dialogue template. Furthermore, any word group can be specified, such as a higher-level word category, a lower-level word category, a word modification group for that word, or any other group containing any word, and a dialogue template linked to the specified word group can be generated. In addition, the stretching process may be configured to search a word dictionary and allow the specification of a group for another word not included in the input text. Furthermore, for elements, the element group corresponding to the specified word may be selected in conjunction with the specified word group, or any element group may be specified. Furthermore, it may be possible to delete thrones set by words in the input text, or to change their location information.
[0100] In step S16, the registration unit 21 uses the order of the words in the input sentence as throne position information, associates the selected group with the throne, and registers the dialogue template in the dialogue template DB5. If multiple word groups are specified for the same throne through the stretching process, the registration unit 21 may register a dialogue template for each word group.
[0101] <4.2. Registering a New Word Template> Figure 11 shows a flowchart of the process related to Word template registration. In this embodiment, a Word template is configured by multiple groups, one or more word groups, and one or more element groups. The registration unit 21 receives the specified word groups and element groups and registers a new Word template.
[0102] First, the registration unit 21 receives a registration instruction request regarding a word template from the terminal device of a user who intends to register a new word template (step S21). The registration unit 21 accepts the selection of a first word group and a first element group to be assigned to the first throne (step S22). Here, the registration unit 21 may be configured to accept the selection of a word category or arbitrary word data, and to allow the specification of a word group. For example, when registering a word template to combine the word "A: finger" and the element "B: of", the word group that corresponds to A is defined by a word group based on a word category or by a word group to which an arbitrarily selected word belongs. For example, first, a word group containing "finger", "hand", "foot", "head", etc. is specified as A. Then, an element (conjugation) group containing "of" is specified as B. The specified groups A and B are set to the throne of the new word template (position information 1 in Figure 8), respectively. For thrones that do not contain elements, a system is created to indicate that they do not contain elements (for example, by setting them to null data or assigning a specific string).
[0103] The registration unit 21 accepts a selection of whether or not to end the combination of words and elements (conjugations) to be defined as a word template (step S23). If the combination is to continue (NO in step S23), the registration unit 21 proceeds to the following step S24 and executes the process. If the combination is to end (YES in step S23), the registration unit 21 proceeds to step S26 and completes the word template setting process.
[0104] In step S24, the registration unit 21 accepts the selection of a word group and an element group to be associated with the second throne. Here, the registration unit 21 may be configured to accept the selection of any word or element and to allow the specification of the word group to which the selected word belongs and the element group to which the element belongs. For thrones that do not contain a word or element, an indication that they do not contain a word or element is registered (for example, by setting them as null data or assigning a predetermined string). In step S25, the registration unit 21 accepts the selection of whether or not to finish adding combinations of word groups and element groups. If the combination is to be completed (YES in step S25), the registration unit 21 proceeds to the following step S26 and completes the word template setting process. If the combination is to be continued (NO in step S25), that is, to set up a third or subsequent throne, the registration unit 21 returns to step S24, accepts the selection of further word groups and element groups to be combined, and associates them with the thrones.
[0105] The registration unit 21 stores the word template in the word template DB4 based on the throne location information, word groups, and element groups set in steps S21 to S25 (step S27). For example, one configuration of a word template is that it consists only of multiple word groups. Examples of word templates in this configuration include text data such as "Tokyo Station" and "bookshelf." A word template consisting only of word groups includes so-called idioms and the like. The registration unit 21 can accept the registration of a word template consisting only of word groups by accepting the option to omit the element group in steps S22 and S24. If multiple word groups are specified for the same throne, the registration unit 21 may register a word template for each word group.
[0106] <5.1. Natural Language Processing of Input Text> Next, we will explain the process of interacting with the user using dialogue templates and the like. Figure 12 shows a flowchart of the process of interacting with the user according to this embodiment.
[0107] <5.2. Identifying Dialogue Templates> First, in step S31, the dialogue template corresponding to the input sentence is identified (step S31). Figure 13 shows a flowchart of the process from input sentence to identifying the dialogue template. The input sentence acquisition unit 10 acquires the input sentence and outputs it to the natural language processing unit 20 (step S41). The splitting unit 22 splits the input sentence into elements (step S42). The word processing unit 23 refers to the attribute information of the elements contained in the input sentence (step S43). The word processing unit 23 refers to a word dictionary to identify the word category (first attribute) and modification attribute (second attribute) that the word has. If the input sentence includes inflection, the word processing unit 23 refers to an inflection dictionary to identify the inflection attribute (third attribute) that the inflection has. If the input sentence includes a word ending, the word processing unit 23 refers to a word ending dictionary to identify the word ending attribute (fourth attribute) that the word ending has.
[0108] <5.3. Word Template Confirmation Process> The word processing unit 23 performs word determination processing based on the elements that have been divided (steps S44 to S49). First, the word processing unit 23 generates a temporary word by sequentially combining the divided elements from the beginning of the input sentence (step S45). In step S46, the word processing unit 23 determines whether the combination of the consecutively placed temporary words matches a word template stored in the word template DB4. If the combination of temporary words does not match a word template stored in the word template DB4 (NO in step S46), the word processing unit 23 returns to step S45 and determines whether there is a matching word template for another combination. If the combination matches a word template stored in the word template DB4 (YES in step S46), the word processing unit 23 determines the word based on the word template (step S47).
[0109] The word processing unit 23 repeatedly executes the processes in steps S44 to S49 until a word template to be applied to all words in the input sentence is determined (step S48). In step S47, elements for which there is no corresponding word template are determined as individual words. Once the word processing unit 23 has completed the determination process for all words in the input sentence, it proceeds to step S49 and executes the subsequent processes. The word processing unit 23 may also be configured to apply a word template to words that have already been determined, for example, by determining a word consisting of two elements and then determining a word consisting of that word and another element.
[0110] The template processing unit 24 refers to the dialogue template DB5 and extracts the dialogue template with the highest match rate to the combination of words arranged in a predetermined order determined in step S49 (step S50). For words determined by the word template, the last word or group of words and elements is compared with the group set as the throne of the dialogue template. If there are multiple dialogue templates with the highest match rate (YES in step S51), the template processing unit 24 proceeds to step S52 and performs a dialogue template filtering process (step S52). If there is only one dialogue template with the highest match rate (NO in step S51) or if the filtering process is completed, the process proceeds to step S53. If the match rates of multiple dialogue templates are the same, the filtering process refers to the dialogue scenario DB7 and determines the dialogue template according to the priority of the dialogue templates set in the scenario. For example, the history of the determination of each template in the dialogue scenario may be recorded, and the dialogue scenario with the highest percentage may be applied. In step S53, the template processing unit 24 identifies the dialogue template as the dialogue template corresponding to the input sentence. The template processing unit 24 associates the identified dialogue template with the input text and stores it in the database DB.
[0111] <5.4. Determination of Instruction Information> Once the dialogue template is determined, in step S32 of Figure 11, the response generation unit 25 refers to the dialogue scenario stored in the scenario DB and determines whether instruction information containing a response corresponding to the dialogue template is associated with it (step S32). If instruction information containing a standard response is set on the dialogue scenario (YES in step S32), the response generation unit 25 outputs the response included in the instruction information to the response output unit 30 (step S33).
[0112] If the instruction information does not include a standard response (NO in step S32), the response generation unit 25 proceeds to step S34. In step S34, the response generation unit 25 refers to the dialogue scenario and obtains instruction information including a response template. The response generation unit 25 converts words and elements in the input sentence according to the conversion rules, applies the converted sentence to the response template to generate a response, and passes it to the response output unit 30 (step S35). The response output unit 30 obtains the response from the natural language processing unit 20 and outputs the response (step S36). The response output unit 30 outputs the response by playing the audio data or displaying it as text data.
[0113] <6.1. Sample> This example shows the process from obtaining a dialogue template corresponding to the input text to outputting a response.
[0114] <6.2. Sample 1> The input text X1 is the text data "Thank you". The natural language processing unit 20 obtains the input text X1 and classifies it into a dialogue template X2 related to "thanks". When the response text generation unit 25 obtains the dialogue template X2, it refers to a dialogue scenario X3 related to "thanks" and obtains instruction information X4. Here, if the instruction information X4 is instruction information that includes a fixed-length response text X5 such as "You're welcome", the response text generation unit 25 generates the response text X5 "You're welcome".
[0115] <6.3. Sample 2> The input sentence Z1 is the text data "I am 18 years old today". The natural language processing unit 20 obtains the input sentence Z1 and classifies it into a dialogue template Z2 related to "age". The input sentence is a combination of the words "I", "today", and "I am 18 years old". The word "I" is word Z3, which corresponds to throne 1 in dialogue template Z2, and has the word group (word category) "person / 1st person" as its first attribute. The word "today" Z4 is word, which corresponds to throne 2, and has the word group "tense / today" as its first attribute. The word "I am 18 years old" is word Z5, which corresponds to throne 3, and has the word attribute "age". Dialogue scenario Z7A is associated with dialogue scenarios Z7A~C, and is written to refer to dialogue scenario Z7B related to "age" when input sentence Z1, which corresponds to dialogue template Z2, is input. When the dialogue template Z2 is determined, the response sentence generation unit 25 refers to dialogue scenario Z7B related to "age" and obtains instruction information Z8. Here, instruction information Z8 is instruction information that includes the specification for generating a response sentence using response template Z9. Response template Z9 refers to all throne 1-3 of dialogue template Z2, and has the following conversion rules: Z10: a conversion rule that refers to throne 1 and omits (deletes) it; Z11: a conversion rule that refers to throne 2 and performs inversion; and Z12: a conversion rule that refers to throne 3 and converts it into a question. The response sentence generation unit 25 omits word Z3 according to Z10, converts word Z5 into a question according to Z12, and inverts word Z4 according to Z11. As a result, the response sentence generation unit 25 generates the response sentence "Are you 18 today?" from the conversion rules of the response template.
[0116] In one embodiment, the natural language processing system 1 is configured as a dialogue device 100. The dialogue device 100 is embodied as a dialogue robot or a terminal device having dialogue means. The terminal device includes smartphones, tablet terminals, smart speakers, etc. In this case, the dialogue device 100 includes an input sentence acquisition unit 10, a natural language processing unit 20, and a response sentence output unit 30.
[0117] <Embodiment 2> Next, Embodiment 2 will be described with reference to Figures 14-15. Components similar to those in Embodiment 1 are denoted by the same reference numerals and their descriptions are omitted. In this embodiment, we will describe (1) a process for estimating (automatically generating) an answer to an arbitrary question, and (2) a process for asking additional questions during a conversation with the user if there are deficiencies in the user's response.
[0118] The input text to the system or the output text from the system regarding a question is called the question text, the output text from the system in response to the question text or the input text to the system is called the answer text, the dialogue template matched to the question text is called the question template, and the dialogue template used to generate the answer text is called the answer template. Furthermore, usage data and word ending data are collectively referred to as element data. Additionally, word data constituting the word dictionary stored in Word DB1, and element data constituting the element dictionary stored in Element DB, are collectively referred to as element data. Finally, a collection of element data, i.e., a word group and an element group, are collectively referred to as an element group. Furthermore, one or more template elements are arranged on the dialogue template. Template elements are element data and / or element groups (word data or word groups, word data and element data, word data and element groups, word groups and element data, word groups and element groups) that are set on one or more thrones provided in the dialogue template.
[0119] <7.1. System Configuration> Figure 14 shows an example of a system configuration diagram of the natural language processing system 1. In this embodiment, the natural language processing system 1 is configured as a dialogue device 100, a natural language processing device 200, and a user terminal device 300. Additional questions will be asked via the dialogue device 100 and the natural language processing device 200, and the estimation of the answer sentence will be done via the user terminal device 300 and the natural language processing device 200. The user terminal device 300 includes a communication unit 301, a control unit 302, a storage unit 303, an input unit 304, and an output unit 305, similar to the hardware configuration of the dialogue device 100 shown in Figure 2(a). The user terminal device 300 and the natural language processing device 200 are connected by a communication network and configured to enable data communication.
[0120] In this embodiment, the natural language processing unit 20 includes a registration unit 21, a division unit 22, a word processing unit 23, a template processing unit 24, a conversion processing unit 26, a response text estimation unit 27, and a dialogue processing unit 28. The conversion processing unit 26 performs text conversion according to predefined conversion rules. The response sentence estimation unit 27 generates one or more response sentences (estimated response sentences) for the input sentence (estimated response sentence creation process). The dialogue processing unit 28 executes processes related to the dialogue with the user. In this embodiment, it performs execution control of actions in response to the user's response (input response) according to the dialogue template, additional question processing, and dialogue scenario registration processing. In action execution control, it executes actions according to the user's input response according to a predefined combination of conditions and actions. In additional question processing, the dialogue processing unit 28 asks additional questions during the dialogue with the user to supplement elements missing from the actual response as answers to the questions. In dialogue scenario registration processing, the dialogue processing unit 28 receives data input related to the dialogue scenario from the user who designs the dialogue scenario and registers it in the dialogue scenario DB7.
[0121] <7.2.(1) Generation of Estimated Response Text> Next, (1) the generation of estimated answer sentences will be explained. In this embodiment, the answer sentence estimation unit 27 generates estimated answer sentences using answer templates corresponding to the question sentences.
[0122] Here, the answer estimation unit 27, (A) Any dialogue template predefined by the user; (B) One or more dialogue templates generated by stretching the question template based on some template elements included in the question template; and, (C) Based on some template elements included in the question template, one or more of the one or more dialogue templates obtained by searching the question template will be used as the answer template corresponding to the question.
[0123] The response estimation unit 27, for the response template, (D) Select any element data (word data or element data) from within each element group (word group or element group) arranged in the answer template, determine the combination, and generate one or more estimated answer sentences depending on the number of combinations; or, (E) From within each element group (word group or element group) arranged in the response template, select any element data (word data or element data) that matches the attribute and / or attribute value that constitutes the conversion condition, determine the combination, and generate one or more estimated response sentences depending on the number of combinations.
[0124] In other words, in generating estimated answer sentences, the answer sentence estimation unit 27 generates estimated answer sentences by one or more of the following processes: via the conversion processing unit 26, converting template elements arranged on the answer template for the question template based on attributes and / or attribute values that are conditions for conversion to generate estimated answer sentences; via the registration unit 21, performing stretching to replace some template elements arranged on the question template with other template elements to generate an answer template, and generating estimated answer sentences based on the template elements arranged in the generated answer template; or searching for other answer templates that share some template elements with the question template, and generating estimated answer sentences based on the template elements arranged in the searched answer template.
[0125] Multiple attributes and / or attribute values may be specified as conversion conditions. In this case, the conversion processing unit 26 generates one or more estimated response sentences corresponding to the conversion conditions for each response template, and then generates one or more estimated response sentences for yet another conversion condition.
[0126] <7.3. Determining the Answer Template> (A) A predefined dialogue template by the user may be a question template. The answer template for generating estimated answer sentences may be the question template itself, or in place of or in addition to the question template, one or more of (A) any pre-specified dialogue template, (B) a dialogue template generated by stretching, and (C) a retrieved dialogue template.
[0127] In generating an answer template, the answer sentence estimation unit 27 passes the question template to the registration unit 21, and the registration unit 21 performs a stretching process based on replacing some of the template elements arranged on the question template with other template elements to generate an answer template for generating an estimated answer sentence. If an answer template is registered in advance, the answer text estimation unit 27 searches the dialogue template DB5 for other dialogue templates that share some template elements with the question template, and uses the retrieved dialogue template as the answer template for generating the estimated answer text.
[0128] <7.4. Marker Elements> A marker element is a template element in the array of question templates that can serve as the subject of an answer, either as a single template element or a combination of multiple template elements. A template element that can serve as the subject of an answer is, for example, a template element related to the 5W1H (Who, What, When, Where, Why, How) arranged in a dialogue template, and refers to, for example, word data such as the following, or a group of words as a collection of word data such as the following. When: "When," "What time," Where: "Where," "At (a place)," Who: "With whom," "Who," What: "What," Why: "Why," How: "How"
[0129] If a marker element is set for a question template, (B) stretching may replace the marker element with another template element. Also, (C) search may search the dialogue template DB5 for other dialogue templates that share elements other than the marker element.
[0130] A correspondence between element data or element groups and marker elements may be defined. The correspondence may be defined by one or more of the following: combinations of word data, combinations of word groups, combinations of attributes of word data, combinations of word data and word groups, combinations of word data and the attributes of word data, or combinations of word groups and the attributes of word data. For example, a correspondence can be defined by associating a word group containing "who" and "who" with word data whose word category (first attribute) is "person". If a correspondence is defined, (B) stretching may replace the marker element with another template element for which the correspondence is defined. Also, (C) search may search the dialogue template DB5 for dialogue templates that contain other template elements for which the marker element and the correspondence are defined.
[0131] In this embodiment, the marker elements are set by the user who inputs the question. The user instructs the answer estimation unit 27 on the position information (the position of the throne) of the template element to be used as the marker element within the question template. If the question is input via voice, voice analysis may be performed, and the marker elements on the question template may be determined according to the parts that are emphasized by speaking loudly or slowly.
[0132] A single marker element may be a combination of multiple template elements, and may also include a combination of multiple template elements and their positional relationships. The positional relationship includes the front-to-back relationship or whether the template elements are consecutive or not. Marker elements may be defined for master data; for example, a marker element may be defined by specifying any template element arranged in a dialogue template stored in the dialogue template DB5, or it may be defined for word data stored in the word DB or word groups stored in the word group DB that can be arranged in the dialogue template as template elements. In addition, among the template elements arranged on the dialogue template, required template elements (required comparison targets) may be treated as marker elements.
[0133] <7.5. Examples of using estimated response text> For example, by providing estimated answer sentences, users can obtain data on combinations of question sentences and answer sentences.
[0134] Furthermore, by outputting a combination of the estimated response sentence, the attributes and / or attribute values that serve as conditions for the conversion, or the classification results based thereon, it is possible to understand what attributes, attribute values, or context the estimated response sentence was created for, and the design of the dialogue scenario via the dialogue processing unit 28 can be made more efficient.
[0135] <7.6. Dialogue Scenarios> A dialogue scenario may include at least one of the following when a dialogue takes place with a person (dialoguer) via the dialogue device 100: a step IN in which the dialoguer inputs an input sentence, and a step AC in which the dialogue device 100 performs an action. When multiple steps are arranged in a dialogue scenario, or when dialogue scenarios are connected to each other, the dialogue scenarios, the connections between dialogue scenarios, or both are designed so that the dialogue flow branches out according to the input sentence by defining combinations of step IN and step AC.
[0136] <7.6. Action Execution Process> The user can set instruction information for executing actions in a dialogue scenario via the dialogue processing unit 28. The instruction information indicates an action to be performed in response to an input sentence (answer sentence) from the interlocutor, and may include one or more of the following: outputting a standard answer sentence or standard question sentence, outputting a generated answer sentence or question sentence, or controlling equipment. The instruction information may also include an action that is performed definitively regardless of the input sentence from the interlocutor (for example, the question sentence that the dialogue device 100 initially outputs to the interlocutor).
[0137] The dialogue processing unit 28 displays a registration screen to the user for registering the conditions related to the input attributes and / or attribute values, and the actions to be performed. On the registration screen, the dialogue processing unit 28 outputs the attributes and / or attribute values corresponding to the estimated response sentences, associating them with the estimated response sentences. The user can then register the actions to be performed while understanding what attributes and / or attribute values will be input when certain actual response sentences are entered. The registration screen allows users to register dialogue scenarios, defining branching conditions for each step, actions to take if the conditions are met, and the connections between steps. Branching conditions can be set based on the matching template of the actual response text, the attributes and / or attribute values of the actual response text, whether or not the actual response text is missing essential response elements, etc. When interacting with a user, the dialogue processing unit 28 determines which branch the user's input response text belongs to by performing template matching, identifying word data in the input response text, and searching the database, and then executes an action.
[0138] <7.7.(2) Processing of additional questions> Next, (2) Additional Question Processing will be explained. In this embodiment, an example is given in which, while the user and the dialogue device 100 are having a dialogue using a natural language processing method, the dialogue device 100 asks the user a question using voice output data, and the user answers using voice input data.
[0139] Figure 15 is a processing flowchart related to the processing of additional questions. As shown in Figure 15, first, the natural language processing unit 20 outputs a question sentence set in the dialogue scenario (step S61). Then, it receives input data of the actual answer sentence from the user (step S62).
[0140] Next, the template processing unit 24 performs template matching on the actual answer sentence to check whether the actual answer sentence contains the required answer elements (step S63). Required answer elements are, for example, element data or element groups that have a defined correspondence with the marker elements of the question sentence, and the unit determines whether these template elements are included in the answer template. Alternatively, it determines whether the words included in the input sentence are word data that have a correspondence with the marker elements, and determines whether the required answer elements are included in the answer sentence. Note that the marker elements are assumed to be attached to the dialogue scenario or the question template for the question sentence.
[0141] If an element is missing (YES in step S64), the system uses the provided information to initiate an additional questioning process in step S65, prompting the user to provide the missing element, and then accepts the user's input of the additional actual answer.
[0142] Additional questions are asked using plain text or dialogue templates and are defined within the dialogue scenario for each marker element that was not answered (a missing required answer element). When setting additional question sentences using dialogue templates, the additional questions may be generated by referencing any throne value (e.g., word ID) included in the actual answer sentence.
[0143] For example, in the following case, the required response element "age" has been answered, but the "user" remains unknown. Therefore, an additional question is asked using the "additional question template for when 'user' is missing," which is linked to the marker element: "user." In this case, the value "16 years old" included in the response is directly substituted into the additional question template, making the intent of the additional question clearer.
[0144] Question: "Please tell us the names and ages of the service users." Marker elements: "User", "Age" Answer template: "I am (who) and (X years old)." Required answer elements: "User", "Age"
[0145] Additional question template when "users" are insufficient: "Are you (X years old)?" Additional question template when "age" is missing: "Please tell me (whose) age." User response when "users" are insufficient: "I am 16 years old." Additional question if there is a shortage of "users": "Are you 16 years old?" User's additional response: "It's me."
[0146] If the missing elements are resolved (NO in step S64), the process ends and proceeds to the next flow of the dialogue scenario.
[0147] As described above, according to the present invention, more natural dialogue can be performed in natural language processing of Japanese. [Explanation of Symbols]
[0148] 1. Natural Language Processing Systems 10 Input text acquisition unit 20 Natural Language Processing Unit 21 Registration Department 22 Division 23 Word Processing Unit 24 Template Processing Unit 25 Answer sentence generation section 26 Conversion Processing Unit 27 Answer sentence estimation part 28 Interactive Processing Unit 30 Answer text output section DB Database 100 Dialogue device 200 Natural Language Processing Devices
Claims
1. A computer that stores element data representing elements that constitute a document, element groups representing a collection of such element data, and a dialogue template in which one or more template elements including the element data and / or the element groups are arranged in a storage unit, Determining the question template, which is a dialogue template corresponding to the question, A dialogue template for generating estimated response sentences, which is based on a response template that includes at least the element group of template elements, and which generates estimated response sentences that are given by combinations of element data determined for each of the template elements, arranged according to the arrangement of the template elements in the response template, and In generating the estimated response sentence, A process to generate an answer template by performing stretching, which replaces some of the template elements arranged on the question template with other groups of elements, or A process of searching for other dialogue templates that share some template elements with the question template, and determining the answer template. A natural language processing method that generates the estimated response sentence based on the response template provided by one or more of the above.
2. The natural language processing method according to claim 1, wherein the computer generates a plurality of estimated response sentences based on any combination of element data selected from a plurality of element data included in the element group arranged in the response template.
3. The storage unit stores attributes and attribute values for each element data, The natural language processing method according to claim 1, wherein the computer generates a plurality of estimated response sentences based on a combination of the plurality of element data included in the element group arranged in the response template, determined based on the attribute and attribute value that serve as a conversion condition.
4. The storage unit stores attributes and attribute values for each element data, The system outputs the attributes and / or attribute values based on the element data included in the estimated response text, associates them with the estimated response text, and displays a registration screen for registering the actions to be performed by referring to these. The natural language processing method according to claim 1, which stores a dialogue scenario by associating the input attributes and / or attribute values with the actions to be performed.
5. The natural language processing method according to claim 4, wherein, in a scene in which a dialogue is performed with a user using the dialogue scenario, when an input response sentence from the user to a certain question is received, the attributes and / or attribute values corresponding to the input response sentence are obtained, and if the conditions are met, the corresponding action is executed.
6. A computer program that causes the computer to execute the natural language processing method described in any one of claims 1 to 5.
7. A storage unit that stores element data indicating elements that constitute a document, element groups indicating a collection of the element data, and a dialogue template in which one or more template elements including the element data and / or the element groups are arranged, A template processing unit that determines a question template, which is a dialogue template corresponding to the question text, A dialogue template for generating estimated response sentences, comprising: a response sentence estimation unit that generates estimated response sentences based on a response template that includes at least the element group of template elements, arranged according to the arrangement of the template elements in the response template, and given by a combination of element data determined for each of the template elements; The aforementioned response text estimation unit, A process to generate an answer template by performing stretching, which replaces some of the template elements arranged on the question template with other groups of elements, or A process of searching for other dialogue templates that share some template elements with the question template, and determining the answer template. A natural language processing system that generates the estimated response sentence based on the response template provided by one or more of the above.