Search device, search method, and recording medium

By using the control unit and grammar analysis technology in the search device, comparing the grammar of words and articles with the grammar of multiple use cases, the problem of difficulty in determining the meaning of polysemous entries in the prior art is solved, and a fast and simple word meaning query is achieved.

CN112528635BActive Publication Date: 2025-06-10CASIO COMPUTER CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010492240.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-17
Filing Date
2020-06-02
Publication Date
2025-06-10
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to simply determine the specific word meaning information required by the user from multiple word meaning information, especially when there are many word meanings corresponding to the entry and the use cases are complex.

Method used

By introducing a control unit into the search device, using dictionary data and grammar analysis results, the grammar of words and articles is compared with the grammar of multiple use cases, and the output is controlled to determine the required word meaning information.

Benefits of technology

It realizes the quick and simple determination of the word meaning information required by the user from multiple word meaning information, and improves the user's query efficiency in the case of multiple word entries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112528635B_ABST
    Figure CN112528635B_ABST
Patent Text Reader

Abstract

The electronic dictionary (10) has a control unit that, when a word and article data including the word have been specified, determines a plurality of usage examples including the word based on dictionary data, compares the grammar of each of the determined plurality of usage examples with the grammar of the article data, and controls the output of information related to the plurality of usage examples based on the result of the grammar comparison.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a retrieval device, a retrieval method, and a recording medium. Background Art

[0002] Generally, in an electronic dictionary, a plurality of entries are registered, and for each of the entries, semantic information including the meaning of a word and the like is associated and stored. Usually, at least one piece of semantic information is stored in one entry. Further, depending on the entry, a plurality of pieces of semantic information may be stored. For example, in a common English-Japanese dictionary, for the entry "run", semantic information corresponding to more than 100 meanings is associated and stored.

[0003] When there are a plurality of meanings corresponding to an entry, the user needs to refer to a legend of the meaning included in the semantic information or an example sentence (example text) of the word including the entry, and thus determine the meaning of the word that the user wants to know from among the plurality of meanings displayed as retrieval results. However, as described above, when there is a large amount of semantic information, it becomes difficult to determine the meaning of the word that the user wants to know. In particular, in order to determine the meaning of a word based on an example sentence, it is necessary to read the example sentences corresponding to the plurality of meanings respectively, and it is not possible to simply determine the meaning of the word that the user wants to know.

[0004] In the prior art, a natural language processing device capable of determining the meaning of a word used in a sentence is known (for example, refer to Patent Document 1). When the natural language processing device determines the meaning of a word included in a sentence from among a plurality of meanings, it displays a plurality of words or meanings representing the plurality of meanings that the word has, and designates the meaning that most conforms to the sentence through a dialog process with the user. In the natural language processing device, although the meaning can be designated with reference to the plurality of words or meanings displayed, it is necessary to designate the most conforming meaning after confirming the plurality of words or meanings.

[0005] [Patent Document 1] Japanese Patent Laid-Open No. 4-130577

[0006] Thus, in the prior art, it is not possible to simply determine the meaning of the word that the user wants to know from among the plurality of meanings corresponding to an entry. Summary of the Invention

[0007] The present invention has been made in view of the above problems, and an object thereof is to provide a retrieval device and a retrieval method capable of simply determining specific semantic information required from among a plurality of semantic information corresponding to one entry.

[0008] One aspect of the present invention relates to a retrieval device having a control unit. When a word and article data including the word are specified, the control unit determines a plurality of usage examples including the word based on dictionary data, compares the grammar of each of the determined plurality of usage examples with the grammar of the article data, and controls the output of information related to the plurality of usage examples based on the comparison result of the grammar.

[0009] Another aspect of the present invention relates to a retrieval method. When a word and article data including the word are specified, a retrieval device determines a plurality of usage examples including the word based on dictionary data, compares the grammar of each of the determined plurality of usage examples with the grammar of the article data, and controls the output of information related to the plurality of usage examples based on the comparison result of the grammar.

[0010] Still another aspect of the present invention relates to a recording medium that records a program for causing a computer to function as a control unit. When a word and article data including the word are specified, the control unit determines a plurality of usage examples including the word based on dictionary data, compares the grammar of each of the determined plurality of usage examples with the grammar of the article data, and controls the output of information related to the plurality of usage examples based on the comparison result of the grammar. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a functional block diagram showing the structure of an electronic circuit of a retrieval device according to an embodiment of the present invention.

[0012] Figure 2 It is a front view showing the external structure of an electronic dictionary in the present embodiment.

[0013] Figure 3 It is a diagram showing an example of information registered in dictionary data in the present embodiment.

[0014] Figure 4 It is a flowchart showing a dictionary control process performed by the electronic dictionary in the present embodiment.

[0015] Figure 5 It is a flowchart showing a dictionary control process performed by the electronic dictionary in the present embodiment.

[0016] Figure 6 It is a diagram showing an example of a retrieval word input screen.

[0017] Figure 7 It is a diagram showing an example of an article display screen.

[0018] Figure 8 It is a diagram showing an example of a modification relationship detected by grammatical analysis.

[0019] Figure 9 This is a diagram showing an example of a syntax tree generated through syntax analysis processing.

[0020] Figure 10 This is a diagram showing an example of a set of distances of modification relation tags.

[0021] Figure 11 This is a diagram showing an example of a sharing relation tag.

[0022] Figure 12 This is a diagram showing an example of the total of sharing relation tags.

[0023] Figure 13 This is a diagram showing the distances corresponding to the modification target relation tags and modification source relation tags corresponding to the input text.

[0024] Figure 14 This is a diagram showing the distances corresponding to the modification target relation tags and modification source relation tags corresponding to the use cases.

[0025] Figure 15 This is a diagram showing an example of the total of the distances corresponding to the sharing relation tags.

[0026] Figure 16 This is a flowchart showing a modified example of the dictionary control process performed by the electronic dictionary in this embodiment.

[0027] Figure 17 This is a diagram showing an example of a modification tag table.

[0028] Figure 18 This is a diagram showing an example of a sharing relation tag. Detailed implementation manners

[0029] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0030] Figure 1 This is a functional block diagram showing the structure of the electronic circuit of the retrieval device according to the embodiment of the present invention.

[0031] In this embodiment, an example is shown in which the retrieval device is configured as an electronic dictionary 10 for example. It should be noted that the retrieval device can be implemented by various electronic devices such as a personal computer, a smart phone, and a tablet PC in addition to the electronic dictionary 10.

[0032] The electronic dictionary 10 records, as dictionary data, information related to at least one meaning corresponding to each of the words set as multiple entries. The dictionary data includes usage examples (sample sentences) that include the words of the entries corresponding to the meanings. The electronic dictionary 10 has a retrieval function for retrieving information including the meanings corresponding to the entries by inputting a string (word) specifying an entry. In the retrieval function of the electronic dictionary 10, not only the string (word) of the retrieval word set as the specified entry, but also a retrieval can be executed by inputting an article including the string (word). The electronic dictionary 10 determines the meaning corresponding to the retrieval word (word) based on the similarity between the grammatical analysis result (first grammatical analysis information) of the article including the word to be retrieved as the retrieval object and the grammatical analysis result (second grammatical analysis information) of the usage examples corresponding to each of the multiple meanings registered in advance.

[0033] The electronic dictionary 10 has a structure of a computer that reads programs recorded on various recording media or programs transmitted, and controls operations according to the read programs. In its electronic circuit, a CPU (central processing unit) 11 is provided.

[0034] The CPU 11 functions as a control unit that controls the entire electronic dictionary 10. The CPU 11 controls the operations of each part of the circuit according to the control program pre-stored in the memory 12, or the control program read into the memory 12 from a recording medium 13 such as a ROM card via the recording medium reading unit 14, or the control program downloaded via the network N including the Internet and read into the memory 12 from the server 20 via the communication unit 15.

[0035] The control program stored in the memory 12 is started according to the input signal corresponding to the user operation from the key input unit 16, the input signal corresponding to the user operation from the touch panel display unit 17, the communication signal from the server 20 on the externally connected network N, or the connection communication signal with the external recording medium 13 such as an EEPROM (registered trademark), RAM, and ROM connected via the recording medium reading unit 14.

[0036] The memory 12, the recording medium reading unit 14, the communication unit 15, the key input unit 16, the touch panel display unit 17, etc. are connected to the CPU 11.

[0037] As a control program stored in the memory 12, a system program for managing the overall operation of the electronic dictionary 10 and a communication program for data communication with other electronic devices such as a server 20 and a personal computer connected to an external network N are stored. Furthermore, a dictionary control program 12a is stored in the memory 12, which executes a retrieval function of retrieving information corresponding to a dictionary entry based on an input string and outputting it. The dictionary control program 12a includes a syntax analysis program 12b for performing syntax analysis on article data.

[0038] Furthermore, in the memory 12, dictionary data 12c, a modification relation label distance table 12d, etc. are stored.

[0039] The dictionary data 12c includes, for example, a database that collects multiple dictionaries such as an English-Japanese dictionary, a Japanese-English dictionary, an English-English dictionary, and a Japanese dictionary. In the dictionary data 12c, semantic information that explains the meaning (semantic meaning) corresponding to each of the entries for each dictionary is included for each dictionary. There are cases where multiple semantic information are stored for one entry. Also, in the semantic information, an example of use (refer to Figure 3 ) indicating the usage example in the article corresponding to the semantic meaning of the entry (word) is set. It should be noted that the dictionary data 12c may not be built into the main body of the electronic dictionary 10, but may be obtained from an accessible dictionary database (for example, the server 20) via the network N.

[0040] The modification relation label distance table 12d stores syntax analysis information (second syntax analysis information) indicating the results of syntax analysis of usage examples corresponding to multiple semantic meanings of one entry for each entry registered in the dictionary data 12c. The syntax analysis information in the modification relation label distance table 12d is used to determine the similarity with the syntax analysis result (first syntax analysis information) of the article input together with the search word as the search object.

[0041] As a result of syntax analysis, for example, it includes: a modification type (modification relation label) indicating the modification relation between the word of the entry used in the usage example corresponding to the semantic meaning and other words outside the usage example; the distance between the word of the entry in the usage example and other words in the modification relation. Details of the modification type (modification relation label) indicating the modification relation and the distance between words will be described later (refer to Figures 8 - 10 )

[0042] The similarity of the syntactic analysis information in this embodiment is set, for example, based on the number of matches of the modification types (modification relationship labels) extracted by syntactic analysis of the article. That is, it is set that the more the same modification types (modification relationship labels) there are, the higher the similarity. Further, in the case of multiple syntactic analysis information (usage examples) with the same number of matches of the modification types (modification relationship labels), the case where the distance between the entries in the modification relationship with the same modification types (modification relationship labels) and other words is smaller is determined to have a high similarity (it should be noted that in the case where there are multiple other words in the modification relationship with the same modification types (modification relationship labels), the determination is made based on the sum of the distances corresponding to each of the multiple other words).

[0043] In the modification relationship label distance table 12d, for example, for each usage example corresponding to each word meaning of all entries, modification relationship data indicating the modification type (modification relationship label) and the distance between the words in the modification relationship is generated as a "modification relationship label distance set" and registered. One set is generated for one usage example corresponding to the word meaning. Further, it is also possible to generate a modification relationship label distance set not only for the usage examples registered in the dictionary data 12c, but also based on other articles, and register it in the modification relationship label distance table 12d.

[0044] It should be noted that the modification relationship label distance table 12d is not generated separately from the dictionary data 12c, and can also be registered as a part of the dictionary data 12c. Further, in addition to being pre-registered in the dictionary data 12c, the "modification relationship label distance set" can also be generated by performing a structure analysis process on an article different from the usage examples registered in the dictionary data 12c, and additionally registered in the modification relationship label distance table 12d. In this case, it is also possible to use the "modification relationship label distance set" registered in the dictionary data 12c and the "modification relationship label distance set" registered in the modification relationship label distance table 12d together to perform the dictionary control process described later.

[0045] In addition, the modification relationship label distance table 12d may not be built into the main body of the electronic dictionary 10, but obtained from an accessible dictionary database (for example, the server 20) via the network N.

[0046] Figure 2 It is the front view showing the external structure of the electronic dictionary 10.

[0047] In Figure 2 In the case of the electronic dictionary 10, a CPU 11, a memory 12, a recording medium reading unit 14, and a communication unit 15 are built in the lower side of the device main body that is opened and closed, and a key input unit 16 is provided, and a touch panel type display unit 17 is provided on the upper side.

[0048] The key input unit 16 includes a character input key 16a, a dictionary selection key 16b for selecting various dictionaries or various functions, a [Translate / Determine] key 16c, a [Return] key 16d, cursor keys (up, down, left, and right keys) 16e, a power button, and various other function keys. On the touch panel display unit 17, various menus, buttons 17a, etc. are displayed according to the execution of various functions.

[0049] The electronic dictionary 10 can input user instructions based on the user's operations on the key input unit 16 or touch operations (using a pen tip or fingertip) on the menus and buttons displayed on the display unit 17.

[0050] For the thus configured electronic dictionary 10, the CPU 11 controls the operations of each part of the circuit according to the commands described in the dictionary control program 12a, and the software and hardware cooperate to operate, thereby implementing the functions described in the following operation description.

[0051] Figure 3 It is a diagram showing an example of the information registered in the dictionary data 12c.

[0052] Figure 3 It shows the semantic information corresponding to the word "catch" set as an entry. For the word "catch" set as an entry, semantic information corresponding to multiple meanings 1, 2,... is registered. For example, in Figure 3 it shows an example in the case where the word "catch" is used as a verb with meaning 1 and as a noun with meaning 2. In addition, in Figure 3 it shows that usage examples 1 and 2 that establish correspondence with each of meanings 1 and 2 and use the meanings of the words are registered in the dictionary data 12c.

[0053] For example, as a usage example of meaning 1 when the word "catch" is used as a verb, for example, "He caught it all on video tape" is registered; as a usage example of meaning 2 when the word "catch" is used as a noun, for example, "There must be a catch somewhere" is registered.

[0054] Next, the operation of the electronic dictionary 10 in this embodiment will be described.

[0055] Figure 4 and Figure 5 It is a flowchart showing the dictionary control process performed by the electronic dictionary 10 in this embodiment.

[0056] If the power supply is turned on, the CPU 11 starts the dictionary control program 12a and begins dictionary control processing. The CPU 11 causes the touch panel display unit 17 to display a home screen as an initial screen (step S1). The home screen includes a menu for selecting a dictionary to be set as a retrieval target. In the menu, it is possible to select a dictionary to be set as a retrieval target. For example, as a dictionary to be set as a retrieval target, it is possible to select a dictionary that includes all dictionaries as a retrieval target, a specific range of dictionaries (e.g., English-language dictionaries, etc.), or a specific dictionary (e.g., ○○ English-Japanese dictionary, etc.).

[0057] If the CPU 11 selects a dictionary to be set as a retrieval target in the menu (step S2, YES), the CPU 11 causes the touch panel display unit 17 to display a retrieval word input screen in which an input area for inputting a retrieval word (word) is set (step S3).

[0058] Figure 6 FIG. is an example of a retrieval word input screen D1. As Figure 6 shown, in the retrieval word input screen D1, an input area AR11 for inputting a string to be set as a retrieval word, an article input area AR12 for inputting an article, and a retrieval start button B1 for instructing the execution of retrieval processing are provided. In the electronic dictionary 10 of the present embodiment, in order to retrieve entries (lexical meaning information) registered in the dictionary, not only a retrieval word (word) can be input, but also an article using the retrieval word can be input.

[0059] In the electronic dictionary 10 of the present embodiment, by not only inputting a retrieval word but also an article, it is possible to retrieve the lexical meaning corresponding to the usage example in which the retrieval word (word) is used in the same way as the input article. Therefore, even when there are a large number of lexical meanings set for one entry, the user can easily extract the required lexical meaning based on the usage example. It should be noted that when performing a dictionary search for a retrieval word that does not have a large number of lexical meanings, the search can also be performed by inputting only the retrieval word in the same way as a normal dictionary search.

[0060] If a string (word) to be set as a retrieval word is input by operating the character input key 16a and the execution of the search is instructed by operating the [Translate / Determine] key 16c or the search start button B1 (step S4, YES), the CPU 11 determines whether there is an article input to the article input area AR12. Here, when an article is not input together with the retrieval word (step S5, NO), the CPU 11 performs a search process on the dictionary data of the dictionary to be set as a retrieval target based on the retrieval word (step S6). That is, the CPU 11 retrieves an entry corresponding to the retrieval word from the dictionary data 12c, reads out the lexical meaning information corresponding to the retrieved entry from the dictionary data 12c, and displays it on the touch panel display unit 17.

[0061] On the other hand, when an article is input together with a search term (step S5, yes), the CPU 11 determines whether the input article has completed syntactic analysis. For example, in the electronic dictionary 10, when the following described modification relation analysis process (syntactic analysis) is performed on the article input to the article input area AR12, the memory 12 stores the processed article and the syntactic analysis result. The CPU 11 determines whether the input article exists in the processed articles. If it is determined that it does not exist (step S16, no), the modification relation analysis process (syntactic analysis) for the input article is executed (step S17). In addition, when the syntactic analysis for the input article is completed, the CPU 11 executes dictionary retrieval using the syntactic analysis result stored by the completed syntactic analysis (steps S18 and subsequent steps).

[0062] It should be noted that in the above description, in the search term input screen D1, a search term and an article are input, but in other methods, a search term and an article that are the objects of dictionary retrieval can be input. For example, when the CPU 11 instructs article display on the home screen (step S11, yes), for example, the touch panel type display unit 17 is caused to display an article display screen including the article corresponding to the text data stored in the memory 12.

[0063] Figure 7 It is a diagram showing an example of the article display screen D2. As Figure 7 shown, in the article display screen D2, in addition to the article display area for displaying the article, a search start button B2 for instructing the execution of the search process is also provided.

[0064] For example, when a touch operation (using a pen tip or a fingertip) is detected at a position where an article is displayed in the article display area, the CPU 11 determines the word displayed at the touch position (step S13), and discriminates the text data of an article including the word (step S14). For example, in Figure 7 it is assumed that the position corresponding to the word "caught" W1 has been touched. The CPU 11 detects the word "caught" based on the touch position, and extracts the text data of an article including "caught", which is "I caught the boy stealing fruit from our orchard.". The CPU 11 uses the word "caught" specified by the touch operation as the search term, and uses the text data "I caught the boy stealing fruit from our orchard." including the word "caught" as the input article. Thus, by simply performing a touch operation on the displayed article, a search term and an article can be easily input and dictionary retrieval can be executed.

[0065] Here, if the execution of the search is instructed by the operation of the [Translation / Decision] button 16c or the search start button B2 (step S15, Yes), the CPU 11 determines in the same manner as described above whether the input article exists in the processed articles. If it is determined that it does not exist (step S16, No), the CPU 11 executes the modification relation analysis process (syntactic analysis) for the input article (step S17). On the other hand, when the syntactic analysis for the input article is completed, the CPU 11 executes a dictionary search using the syntactic analysis result stored by the completed syntactic analysis (steps S18 and subsequent).

[0066] Next, Figure 4 the modification relation analysis process (syntactic analysis) in step S17 shown below will be described.

[0067] Figure 8 (A) of is a diagram showing an example of the modification relation (modification relation label) detected by syntactic analysis for the input article. Figure 8 (B) and (C) of are diagrams showing an example of the modification relation (modification relation label) detected by syntactic analysis for the use case (shown in ). Note that in the syntactic analysis process, the detailed description of the process using the existing method is omitted. Figure 3 In the syntactic analysis process, the modification relations between the word corresponding to the search word in the article (input word) and multiple other words are detected. In the modification relations, there are relations with other words (modified target words) before the input word in the document and relations with other words (modifying source words) after it, and the relation labels representing each relation are obtained.

[0068] For example, in the article shown in (A) of, with respect to the input word "caught", the other word "I" becomes the modified target word, and the relation label "nsubj" (indicating the subject noun) is obtained. Furthermore, with respect to the input word "caught", the other word "stealing" becomes the modifying source word, and the relation label "xcomp" (indicating the complement) is obtained.

[0069] In addition, in the modification relation analysis process, the distance in the article from the input word to other words is determined. In an article, when the distance between words is short, the relevance between the words can be regarded as high. For example, Figure 8 in (A) of, the distance from the input word "caught" to the word "I" is "-1", and the distance to the word "stealing" is "3".

[0070] Figure 8

[0071] ​​Note that, as described above, it is also possible to simply count the number of words from the input word to other words and set it as the distance, and it is also possible to use the syntactic analysis result to determine the distance. For example, perform syntactic analysis processing to generate a syntax tree representing the sentence structure of the article, and set the number of branches of the syntax tree as the distance between words. Thus, regardless of the presence or absence of words such as articles that are irrelevant to the modification relationship between words, although there will be changes in the distance in the count value of the simple number of words, by setting the number of branches of the syntax tree as the distance, the distance corresponding to the modification relationship between words can be determined.

[0072] For example, in the article "I have a pen.", the number of words from the input word "have" to the other word "pen" becomes "2". On the other hand, in the text "I have pens.", the number of words from the input word "have" to the other word "pens" is "1". That is, regardless of whether the articles are the same in terms of word usage, due to the difference in whether the relevant word candidate is in the singular form "pen" or the plural form "pens", there will be a difference in the presence or absence of articles. Therefore, when simply setting the number of words up to a word as the distance, even if the modification relationship between words is the same, the distance will change.

[0073] Figure 9 It is a diagram showing an example of a syntax tree generated through syntactic analysis processing. Figure 9 (A) of which represents the syntax tree corresponding to the aforementioned article "I have a pen.", Figure 9 and (B) of which represents the syntax tree of the aforementioned article "I have pens.".

[0074] As Figure 9 shown in (A) of which, the number of branches between the input word "have" K2 and the other word "pen" T2 in the article "I have a pen." is "5". Furthermore, as Figure 9 shown in (B) of which, the number of branches between the input word "have" K3 and the other word "pens" T3 in the article "I have pens." is "5". That is, regardless of the presence or absence of articles in the article, for the input word and the other word with the same modification relationship between the input word and the other word in the article with the same structure, the same distance can be determined.

[0075] In this way, perform syntactic analysis on the article including the input word and set the number of branches of the syntax tree as the distance between words. Thus, even if there are changes in the article due to differences in the presence or absence of articles, etc., the positional relationship (distance) between words can be correctly determined.

[0076] Next, the dictionary retrieval using the syntactic analysis result will be described.

[0077] The CPU 11 performs a retrieval process on the dictionary data of the dictionary set as the retrieval object based on the retrieval word (step S18). That is, the CPU 11 retrieves an entry corresponding to the lemma of the word of the retrieval word from the dictionary data 12c, and extracts the usage examples of all semantic information corresponding to the retrieved entry (step S19).

[0078] The CPU 11 selects the usage example corresponding to the entry that is the object of the process for determining the similarity to the input article (step S20). Here, when there is a usage example that is the object of the process (step S21, yes), since the process for all usage examples has not been completed, the CPU 11 moves to the process for the selected usage example. The CPU 11 determines whether the syntactic analysis process for the selected usage example is completed. That is, it determines whether the "modifier relationship label distance set" corresponding to the usage example is registered in the modifier relationship label distance table 12d.

[0079] When the syntactic analysis process for the usage example has not been completed (step S22, no), the CPU 11 extracts the text data (article) of the usage example from the dictionary data 12c and performs a modifier relationship analysis process (syntactic analysis) (step S24).

[0080] The modifier relationship analysis process (syntactic analysis) is executed in the same manner as the modifier relationship analysis process (syntactic analysis) for the previously input article (step S17). For example, in the case of the usage example 1 "He caught it all on video tape" shown in Figure 3 , as shown in (B) of Figure 8 , the modifier relationship and distance between the word corresponding to the entry "catch" and other words are determined. Similarly, in the case of the usage example 2 "There must be a catch somewhere" shown in Figure 3 , as shown in (C) of Figure 8 , the modifier relationship and distance are determined.

[0081] Here, for the result of the syntactic analysis that has been executed ("modifier relationship label distance set"), a correspondence is established with the semantic information (usage example) of the entry, and it is additionally stored in the modifier relationship label distance table 12d. Thus, when the same usage example becomes the object of the process, the syntactic analysis process can be omitted using the completed "modifier relationship label distance set".

[0082] In this way, when the syntactic analysis process for the usage example has not been completed, since the modifier relationship analysis process can be executed at that time, for example, in the case of a structure in which usage examples corresponding to the semantics can be additionally added to the dictionary data 12c, the newly added usage examples can also be set as the object of the process.

[0083] On the other hand, when the syntactic analysis processing for the use case is completed (step S22, yes), the CPU 11 reads out the modification relation data ("modification relation label distance set") representing the syntactic analysis result (second syntactic analysis information) corresponding to the use case from the modification relation label distance table 12d, and discriminates the similarity with the syntactic analysis result (first syntactic analysis information) of the input article.

[0084] Figure 10 It is a diagram showing an example of the "modification relation label distance set" registered in the modification relation label distance table 12d. In Figure 10 In it, for each of the multiple semantic meanings 1, 2, 3... corresponding to one entry registered in the dictionary data 12c, the use cases represent the "modification relation label distance set". In Figure 10 In it, it shows the cases where multiple use cases are set for one semantic meaning regarding semantic meanings 1 and 3, and one use case is set regarding semantic meaning 2. Therefore, regarding semantic meanings 1 and 3, multiple "modification relation label distance sets" corresponding to each of the multiple use cases are registered.

[0085] For example, in semantic meaning 1, there are multiple use cases 1, 2, 3..., and the "modification relation label distance sets" are respectively stored in correspondence with the multiple use cases 1, 2, 3....

[0086] In the "modification relation label distance set" corresponding to use case 1 in semantic meaning 1, it includes: the distances "-3", "-2", "-1", "1" corresponding to the relation labels "advmod", "aux", "nsubj", "dobj" respectively representing the modification relations between the word of the entry in use case 1 and the modified target word; and the distance "0" corresponding to the relation label "root" representing the modification relation with the modified source word.

[0087] In this way, if the syntactic analysis result ("modification relation label distance set") corresponding to the use case is registered in the modification relation label distance table 12d in advance, it is not necessary to perform the modification relation analysis processing for the use case every time a search term and article data are input, so the search time can be shortened and the accuracy can be improved.

[0088] When the CPU 11 obtains the "modification relation label distance set" for the use case to be processed, for each of the modification target relation label and the modification source relation label of the input article and the use case, it discriminates the common relation label (common relation label), and obtains the total of the common relation labels (step S25).

[0089] In Figure 11 It shows an example of the common relation label. Figure 11(A) represents the modification target relationship label and the modification source relationship label corresponding to the input article. Figure 11 (B1) of represents Figure 3 the modification target relationship label and the modification source relationship label corresponding to Use Case 1 shown. Figure 11 (C1) of represents Figure 3 the modification target relationship label and the modification source relationship label corresponding to Use Case 2 shown.

[0090] As Figure 11 shown by (B2), for the modification target relationship label of the input article and Use Case 1, two relationship labels "nsubj" and "dobj" are shared, and for the modification source relationship label of the input article and Use Case 1, 1 "root" is shared, and they are respectively judged as shared relationship labels. Therefore, regarding Use Case 1, as Figure 12 shown, the total of the shared relationship labels is obtained as "3".

[0091] On the other hand, as Figure 11 shown by (C2), there are no shared relationship labels in the modification target relationship label and the modification source relationship label of the input article and Use Case 2. That is, for the input article in which the word of the entry is used as a verb, there are shared relationship labels in Use Case 1 corresponding to the meaning of the verb, but there are no shared relationship labels in Use Case 2 corresponding to the meaning of the noun. Thus, according to the meaning of the entry, using the situation where the grammatical structures of the use cases are different, based on the shared relationship labels, it is possible to increase the priority of Use Case 1 with a high similarity and decrease the priority of Use Case 2 (or remove it from the search target).

[0092] Next, the CPU 11 respectively obtains the total of the differences between the distances corresponding to the shared relationship labels of the modification target relationship label and the distances corresponding to the shared relationship labels of the modification source relationship label, and sums the total values corresponding to the modification target relationship label and the modification source relationship label respectively (step S26).

[0093] Figure 13 represents the distance corresponding to each of the modification target relationship label and the modification source relationship label corresponding to the input article. Figure 14 represents Figure 3 the distance corresponding to each of the modification target relationship label and the modification source relationship label corresponding to Use Case 1 shown. The shared relationship labels of the input article and Use Case 1 are, as described above, the relationship labels "nsubj" and "dobj" are shared for the modification target relationship label, and "root" is shared for the modification source relationship label.

[0094] The distance corresponding to the common relation label "nsubj" of the input article is "-1", and the distance corresponding to the common relation label "nsubj" of Use Case 1 is "-1". Therefore, regarding the difference in distance for the common relation label "nsubj", it is "-1 - (-1) = 0". Similarly, regarding the difference in distance for the common relation label "dobj", it is "2 - 1 = 1". Similarly, the distance corresponding to the common relation label "root" is "2 - 2 = 0". Therefore, as Figure 15 shown, the total value corresponding to the modified target relation label and the modified source relation label for Use Case 1 is "1".

[0095] CPU 11 establishes a correspondence between the total number of common relation labels, the total distance, and the use cases to be processed and stores them in the memory 12 (step S27).

[0096] Next, similarly, CPU 11 selects 1 use case to be the next processing object corresponding to the entry (step S20), executes the aforementioned processing, obtains the total number of common relation labels and the total distance, and stores them in the memory 12 after establishing a correspondence with the use case (steps S21 - S27).

[0097] If the processing for all use cases is completed (step S21, no), then CPU 11 executes a process to determine the similarity between the result of the syntactic analysis of the input article (first syntactic analysis information) and the result of the syntactic analysis (second syntactic analysis information) including the total number of common relation labels and the total distance stored through the syntactic analysis for each use case (step S40).

[0098] First, CPU 11 selects the use case (lexical meaning) with the largest total number of common relation labels (step S28). That is, it determines the use case with the highest similarity of syntactic analysis information and the most consistent syntactic structure with the input article.

[0099] It should be noted that when there are multiple use cases with the same total number of common relation labels (step S29, yes), CPU 11 determines and selects the use case with the smallest total distance of the common relation labels as the use case with high similarity (step S30). The higher the relevance between other words in the modified relation corresponding to the common relation label and the words corresponding to the entry, the smaller the total distance of the common relation labels becomes. Thus, by selecting the use case with the smallest total distance of the common relation labels, it is easier to determine the use case where the usage of the words of the entry in the modified relation and other words is closer to the input article.

[0100] If the similarity between the input article and each use case is determined, the CPU 11 determines priorities for each use case or the word meanings corresponding to each use case based on the similarity determination result, and controls the output. That is, the CPU 11 gives the highest priority (upper position) to the use case with the highest similarity. For the other use cases, the use cases (word meanings) are sorted in descending order based on the total number of common relation tags (step S31). That is, multiple use cases are sorted in the order of higher similarity to the input article to determine the priorities.

[0101] Furthermore, when there are multiple use cases with the same total number of common relation tags, the CPU 11, in the same manner as described above, separately obtains the total of the distances corresponding to the common relation tags of each use case, and sorts them in ascending order based on the total of the distances (step S32). Thereby, it is possible to determine the priorities based on the relevance (distance) between the words in the entries in the use cases and other words.

[0102] The CPU 11 arranges the word meaning information (use cases, or word meanings corresponding to use cases) corresponding to multiple use cases whose priorities are determined based on the common relation tags according to the priorities, and causes the touch panel display unit 17 to perform display (step S33). It should be noted that the word meanings corresponding to the use cases without common relation tags may also be excluded from the display targets.

[0103] In this way, in the electronic dictionary 10 in the present embodiment, not only the string (word) of the search word of the designated entry is set, but also an article including the string (word) is input, whereby it is possible to give priority to the word meaning information of the use case close to the input article and display it as a search result. That is, even if there are multiple word meanings for one entry, the priorities can be determined based on the similarity of the grammatical analysis results of the input article and the use cases corresponding to the word meanings, so that it is possible to simply and effectively obtain the word meaning that the user wants to know.

[0104] It should be noted that in the foregoing description, grammatical analysis (modifier relation analysis processing) is performed on the article input by the user and the use cases in the electronic dictionary 10, but the data to be processed can also be sent to the server 20 (cloud) connected via the network N and processed.

[0105] (Variant example)

[0106] Next, a variant example of the process for determining the similarity between the result of the grammatical analysis of the input article (first grammatical analysis information) and the results of the grammatical analysis of each use case (second grammatical analysis information) will be described. In the foregoing description, the similarity is determined based on the total number of common relation tags and the total of the distances, but in the variant example, a weight value is obtained for each common relation tag (modifier relation tag), and the similarity is determined based on the total of the weight values of the common relation tags.

[0107] In the method of summing the number of shared relationship tags described above, all types of modification relationship tags are equivalently treated, and the number of one modification relationship tag is simply set to 1 for summation. However, depending on the entry and the meaning of the word contained in the entry, the tendency of the appearance frequency of the modification relationship tags used in the usage examples also varies. That is, in the article using the entry, depending on the entry, there are modification relationship tags that are likely to occur and those that are difficult to occur. Therefore, for each entry and based on the appearance frequency of the modification relationship tags, a weight value is used such that the more likely a modification relationship tag is to occur, the larger the value. Thus, even if the number of modification relationship tags that are likely to occur and those that are difficult to occur in the input article is the same, the meaning of the side using the modification relationship tag that is more likely to occur can be prioritized for display, achieving an improvement in accuracy.

[0108] Hereinafter, a dictionary control process that uses summation based on weight values to determine similarity will be described. It should be noted that in this dictionary control process, the processes of steps S1 to S19 shown in Figure 4 and the processes corresponding to steps S20 to S33 shown in Figure 5 are executed. In the flowchart shown in Figure 16 , the parts that execute the same processes as the flowchart shown in Figure 16 are given the same reference numerals. Regarding the parts shared with the descriptions using Figure 5 , the description will be omitted. Figure 4 and Figure 5

[0109] In the case of determining similarity based on the summation of weight values of shared relationship tags, for each of all the entries in the modification relationship tag distance table 12d, a weight value is set for the modification relationship tags used in the usage examples of each meaning corresponding to the entry. The weight value of the modification relationship tag is calculated as follows.

[0110] Figure 17 An example of a modification tag table that registers the modification relationship tags used in the usage examples of each meaning corresponding to the entry "catch" is shown.

[0111] In the modification tag table, all the modification relationship tags used in all the usage examples corresponding to the entry "catch" are set, and the frequency of each modification relationship tag is calculated. In the present embodiment, the highest frequency among the modification relationship tags used in all the usage examples within the same entry is set as fmax. In Figure 17In the example shown, since the frequency "17" of the modification relation label "dobj" is the highest, the frequency "17" of the modification relation label "dobj" is set as fmax. Moreover, the weight value of each modification relation label is the value obtained by dividing each frequency by fmax ("17").

[0112] In Figure 16 In step S20 shown, the CPU 11 selects one use case corresponding to a word entry that is set as the object of the process for determining the similarity with the input article. Here, when there is a use case set as the object of the process (step S21, yes), since the process for all use cases has not ended, the CPU 11 moves to the process for the selected use case.

[0113] In the explanation using Figure 5 although the modification relation analysis process (syntactic analysis) is executed individually when the syntactic analysis process for the selected use case is not completed, here the modification relation analysis process is executed for all use cases for which the syntactic analysis process is not completed, and the processing result is reflected in the modification label table, and the frequencies and weight values of all modification relation labels are calculated and set.

[0114] If the "modification relation label distance set" for the use case set as the object of the process is obtained, the CPU 11 determines the common relation labels (common relation labels) for each of the modification target relation label and the modification source relation label between the input article and the use case, and obtains the sum of the weight values of the common relation labels (step S45).

[0115] Figure 18 represents the common relation labels, the weight values of each common relation label, and the total value in the example shown above Figure 11 That is, it represents that there are two common relation labels "nsubj" and "dobj" between the input article and use case 1, the weight value of the common relation label "nsubj" is "0.941176", and the weight value of the common relation label "dobj" is "1". Therefore, for use case 1, the sum of the weight values of the common relation labels is obtained as "1.941176".

[0116] Next, the CPU 11 obtains the sum of the differences between the distances corresponding to the common relation labels of the modification target relation label and the distances corresponding to the common relation labels of the modification source relation label respectively, and sums the total values corresponding to the modification target relation label and the modification source relation label respectively (step S26).

[0117] The CPU 11 associates the sum of the weight values of the common relation labels and the sum of the distances with the use case set as the object of the process and stores them in the memory 12 (step S47).

[0118] Next, similarly, the CPU 11 selects one use case (step S20) that is to be the next processing target corresponding to the entry, executes the above-described processing, obtains the sum of the weight values of the common relationship labels and the sum of the distances, associates them with the use case, and stores them in the memory 12 (steps S21 to S47).

[0119] If the processing for all use cases is completed (step S21, no), the CPU 11 executes a process of determining the similarity between the result of the syntactic analysis of the input article (first syntactic analysis information) and the result of the syntactic analysis including the sum of the weight values of the common relationship labels and the sum of the distances stored by the syntactic analysis for each use case (second syntactic analysis information) (step S40).

[0120] First, the CPU 11 selects the use case (lexical meaning) with the largest sum of the weight values of the common relationship labels (step S48). That is, the use case with the highest similarity of the syntactic analysis information and the most consistent syntactic structure with the input article is determined.

[0121] Note that, in the case where there are multiple use cases with the same sum of the weight values of the common relationship labels (step S29, yes), the CPU 11 selects the use case with the smallest sum of the distances of the common relationship labels as the use case with high similarity (step S30).

[0122] If the similarity between the input article and each use case is determined, the CPU 11 determines the priority for each use case or the lexical meaning corresponding to each use case based on the similarity determination result and controls the output. That is, the CPU 11 gives the highest priority (upper position) to the use case with the highest similarity, and for the other use cases, sorts the use cases (lexical meanings) in descending order based on the sum of the weight values of the common relationship labels (step S51). That is, multiple use cases are sorted in the order of high similarity to the input article to determine the priority.

[0123] In addition, in the case where there are multiple use cases with the same sum of the weight values of the common relationship labels, the CPU 11, in the same manner as described above, separately obtains the sum of the distances corresponding to the common relationship labels of each use case, and sorts them in ascending order based on the sum of the distances (step S52). Thereby, the priority can be determined based on the relevance (distance) between the words of the entries in the use case and other words.

[0124] The CPU 11 arranges the semantic information (use case, or the lexical meaning corresponding to the use case) corresponding to the multiple use cases whose priorities are determined based on the common relationship labels according to the priority, and causes the touch panel display unit 17 to display it (step S33). Note that the lexical meaning corresponding to the use case without the common relationship label may also be excluded from the display target.

[0125] In this way, by using the weight value of the modification relationship tag calculated for each entry based on the occurrence frequency of the modification relationship tag, it is possible to preferentially display the word meaning using the modification relationship tag that is easy to generate.

[0126] It should be noted that in the foregoing description, the weight value of each modification relationship tag is set to the value obtained by dividing the frequency of each modification relationship tag by the highest frequency fmax among all the use cases of the modification relationship tags used within the same entry. As long as the modification relationship tag that is easier to generate becomes larger, the weight value can also be calculated by other methods.

[0127] Furthermore, in the description of using Figure 16 when there are multiple use cases (word meanings) with the same total weight value of the sharing relationship tags, the priority is determined based on the distance of the sharing relationship tags. However, it is also possible to arbitrarily implement the discrimination of the number, distance, and weight value of the foregoing sharing relationship tags used.

[0128] Also, for example, when the total number of sharing relationship tags is the same, the user can be allowed to select which one of the distance or the weight value to further use for discrimination of the priority.

[0129] It should be noted that in the foregoing embodiment, an English-language dictionary is taken as an example for description, but it is also possible to implement other language dictionaries as the object.

[0130] In addition, the methods described in the embodiment, that is, each method such as the processing shown in the flowchart, can be stored as a program executable by a computer in a recording medium such as a memory card (ROM card, RAM card, etc.), a magnetic disk (floppy disk, hard disk, etc.), an optical disk (CD-ROM, DVD, etc.), a semiconductor memory, etc., and distributed. Moreover, the computer reads the program recorded in the external recording medium and controls its operation using this program, thereby enabling the same processing as the functions described in the embodiment.

[0131] In addition, the data of the program for implementing each method can be transmitted over a network (Internet) in the form of program code, and it is also possible to obtain the program data from a computer (server device, etc.) connected to the network (Internet) to achieve the same functions as the foregoing embodiment.

[0132] It should be noted that the invention of the present application is not limited to the embodiments, and can be variously modified within the scope without departing from its gist during the implementation stage. Furthermore, the embodiments include inventions at various stages, and various inventions can be extracted through appropriate combinations of a plurality of structural elements disclosed. For example, when even if several structural elements are deleted from all the structural elements shown in the embodiments, or several structural elements are combined, the problems described in the "Problems to be Solved by the Invention" column can be solved and the effects described in the "Effects of the Invention" column can be obtained, the structure in which the structural elements are deleted or combined can be extracted as an invention.

Claims

1. A retrieval device having a control unit, wherein when a word and article data including the word have been specified, the control unit determines a plurality of usage examples including the word based on dictionary data, for each text of the determined plurality of usage examples and the text of the article data, determines a plurality of modification types representing the modification relationships of other multiple words with respect to the word, calculates the number of common modification types among the determined plurality of modification types, and controls the output of information related to the plurality of usage examples based on the calculation result, wherein in the calculation of the number of common modification types, the control unit performs a calculation of adding weights determined for each modification type.

2. The retrieval device according to claim 1, wherein, in the dictionary data, a plurality of semantic meanings corresponding to one word and usage examples including the one word corresponding to each of the plurality of semantic meanings are stored, the control unit controls the output of each of the determined usage examples or the semantic meanings corresponding to each of the determined usage examples based on the number of common modification types obtained by the calculation of adding the weights, as the information related to the plurality of usage examples.

3. The retrieval device according to claim 1, wherein, the control unit controls the output of information related to the plurality of usage examples based on the number of common modification types and the distance between words in the common modification types.

4. The retrieval device according to any one of claims 1 to 3, wherein, based on the occurrence frequencies of a plurality of modification types for each word, weights determined for each modification type are determined such that the higher the occurrence frequency, the greater the weight.

5. The retrieval device according to claim 4, wherein, the weight determined for each modification type sets the ratio of the occurrence frequency of each modification type to the highest occurrence frequency among the plurality of modification types as the weight.

6. A retrieval method, wherein when a word and article data including the word have been specified in a retrieval device, a plurality of usage examples including the word are determined based on dictionary data, for each text of the determined plurality of usage examples and the text of the article data, a plurality of modification types representing the modification relationships of other multiple words with respect to the word are determined, the number of common modification types among the determined plurality of modification types is calculated, and the output of information related to the plurality of usage examples is controlled based on the calculation result, wherein in the calculation of the number of common modification types, a calculation of adding weights determined for each modification type is performed.

7. A recording medium recording a program for causing a computer to function as a control unit, wherein when a word and article data including the word have been specified, the control unit determines a plurality of usage examples including the word based on dictionary data, for each text of the determined plurality of usage examples and the text of the article data, determines a plurality of modification types representing the modification relationships of other multiple words with respect to the word, Calculate the number of common modification types among the determined multiple modification types, and based on the calculation result, control the output of the information involved in the multiple use cases. In calculating the number of the common modification types, the control unit performs a calculation of adding the weight determined for each modification type.

Citation Information

Patent Citations

  • Natural language processor

    JP1992130577A

  • On-line dictionary and read understanding support system utilizing same

    JP1996235181A

  • Simple sentence similarity computer

    JP1997212509A

  • Text data similarity calculation method, text data similarity calculation apparatus, and text data similarity calculation program

    JP2006139708A