Compound word detection device and compound word detection method

The compound word detection device addresses the limitations of existing methods by using a time-series model to learn and detect compound words in technical fields, ensuring accurate identification of unknown words without manual intervention or web search reliance.

JP2025077540APending Publication Date: 2025-05-19HITACHI GE NUCLEAR ENERGY LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023189809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

Existing methods for detecting compound words in technical fields are inadequate, particularly for unknown words, as they rely on manual definition of part-of-speech sequences and web search data, which is limited for domain-specific terms.

Method used

A compound word detection device that learns a model using feature amounts of consecutive words from a compound word dictionary, employing a time-series model to automatically detect compound words without relying on manual pattern definition or web search data.

Benefits of technology

Accurately detects compound words in specialized fields, including unknown words, by comprehensively learning patterns and eliminating the need for time-consuming manual definition and web search dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025077540000001_ABST
    Figure 2025077540000001_ABST
Patent Text Reader

Abstract

To accurately detect a compound word of a technical field even when it is an unknown word.SOLUTION: A compound word detection device includes a compound word pattern learning unit and a compound word detection unit. The compound word pattern learning unit learns a model for detecting whether or not a plurality of consecutive words are a compound word using a feature amount of the plurality of consecutive words that are a result of dividing the compound word registered in a compound word dictionary. The compound word detection unit inputs the feature amount of the plurality of consecutive words to be detected relative to a learned model and receives an output of the learned model indicating whether or not the plurality of consecutive words to be detected are a compound word.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a compound word detection device and a compound word detection method.

Background Art

[0002] Words and expressions formed by combining two or more words with new meanings are called compound words. Detecting compound words from a target sentence is very important in natural language processing. In particular, compound word detection in technical term recognition is essential for effectively understanding and handling sentences in a technical field (a field specialized in the target application domain). That is, technical terms have unique meanings and usage examples that are not found in general expressions, and by recognizing technical terms, the context specific to that field can be correctly understood.

[0003] To detect compound words, it is effective to formulate the composition rules of the words that make up the compound words and extract combinations of words that conform to the rules from within the sentence. For example, the electronic device of Patent Document 1 extracts a group of consecutive words of a predetermined part of speech from a word sequence obtained by morphological analysis of text, and sets the combination pattern of the extracted words and the part of speech of each word as a candidate for the word composition rule. Then, the electronic device examines the appearance ratio of the pattern and the number of times the compound word that matches the pattern is searched on the web in a document set included in the learning sentence corpus, and determines candidates with high values as formal rules.

[0004] Also, a method of dividing text into compound words based on the probability that a word boundary exists between characters is effective. Non-Patent Document 1 performs Markov approximation on character classes obtained by clustering similar characters in a document set included in a learning sentence corpus using n-gram, and determines whether there is a boundary between consecutive character classes. Non-Patent Document 2 sets a word boundary at a position that maximizes the entropy of the word boundary probability learned using an n-gram model.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] However, when generating candidates for word composition rules, the electronic device of Patent Document 1 extracts words based on the part-of-speech sequence (for example, words with two or more consecutive nouns) defined manually in advance. Defining these manually is time-consuming and there may be omissions in the definition. Furthermore, when determining the formal rules, the electronic device examines the number of times a compound word has been web-searched. Technical terms that depend on a specific domain are basically rarely published on web pages, and among them, there are many technical terms that are not at all publicly available outside, such as company-confidential information. The electronic device of Patent Document 1 is not suitable for detecting compound words of technical terms.

[0008] In addition, since Non-Patent Document 1 and Non-Patent Document 2 generate a language model using the words themselves included in the learning text corpus, the behavior for unknown words that do not exist in the text corpus is not guaranteed.

[0009] Therefore, an object of the present invention is to accurately detect compound words in a specialized field even when they are unknown words.

Means for Solving the Problem

[0010] The compound word detection device of the present invention learns a model for detecting whether a plurality of consecutive words are compound words by using the feature amounts of a plurality of consecutive words that are the result of splitting the compound words registered in the compound word dictionary. The compound word detection device includes a compound word pattern learning unit and a compound word detection unit. The compound word detection unit receives the input of the feature amounts of a plurality of consecutive words to be detected with respect to the learned model, and receives the output of whether the plurality of consecutive words to be detected are compound words from the learned model. Other means will be described in the mode for carrying out the invention.

Effect of the Invention

[0011] According to the present invention, compound words in a specialized field can be accurately detected even when they are unknown words.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present invention (hereinafter referred to as "the present embodiment") will be described in detail with reference to the drawings. The present embodiment is an example of detecting Japanese compound words. However, the present invention is generally applicable to examples of determining compound word boundary lines in languages that do not have the habit of inserting blanks between words, such as Chinese and Korean. Furthermore, even in languages that insert blanks between words, such as English, it is applicable to fields where compound words such as "Central Processing Unit" are frequently used.

[0014] FIG. 1 is a diagram showing the configuration of a compound word detection device 1. The compound word detection device 1 is a general computer. The compound word detection device 1 includes a central control device 11, an input device 12 such as a keyboard and a mouse, an output device 13 such as a display, a main storage device 14, an auxiliary storage device 15, and a communication device 16.

[0015] The sub-word splitting unit 21, the morphological analysis unit 22, the compound word selection unit 23, the non-compound word generation unit 24, the compound word pattern learning unit 25, the word selection unit 26, the compound word candidate detection unit 27, and the compound word detection unit 28 are programs. In the following description, when the subject is described as "○○ unit is", it means that the central control device 11 reads each program from the auxiliary storage device 15 into the main storage device 14 and realizes the functions of each program (details will be described later). Note that these programs may be implemented as "○○ devices" which are hardware.

[0016] The auxiliary storage device 15 stores a text corpus 31, a morphological analysis dictionary 32, a compound word dictionary (positive examples) 33, a non-compound word dictionary (negative examples) 34, and a time series model 35 (details will be described later). The auxiliary storage device 15 may have a configuration independent of the compound word detection device 1.

[0017] FIG. 2 is a diagram showing an example of a compound word pattern. There is a regularity in the pattern of the arrangement of words constituting a compound word (hereinafter also referred to as a "compound word pattern" or simply a "pattern").

[0018] First, as a pattern 101 that does not depend on a domain, a compound word including a prefix 102 and a suffix 103 can be cited. In this case, for example, in the pattern of "prefix + verb", compound words such as "cancel", "approach", "handle", etc. apply, and in the pattern of "noun + suffix", compound words such as "tinged with black", "greasy", "two-person job", etc. apply.

[0019] Patterns 104 that depend on a specific domain, that is, technical terms, often form regular patterns. For example, the model numbers of equipment in a factory often adopt highly regular character strings 105 for each site. In this example, the model number is expressed by a combination of a specific character string and alphanumeric characters. For example, in the pattern where a numeric sequence follows the word "TS", and then a single alphabetic character comes and then a numeric sequence follows, "TS121T02", "TS014D33", "TS131D01", etc. apply.

[0020] When the co-occurrence of words is high in case 106, regular patterns are often formed. For example, when compound words such as "terminal mar 1", "terminal mar 2", "terminal mar 3" frequently appear, it can be seen that a symbol often comes after the word "terminal", and a pattern of "“terminal” + symbol" is derived. Note that "mar 1" means a symbol in which "1" is written in "○". The same applies to "mar 2" and "mar 3".

[0021] Also, when there is a connection by highly relevant word 107, a regular pattern is formed. For example, in the case of compound words such as "cell motor cover", "discharge resistor cover", and "insulated transformer cover", it can be seen that a series of nouns related to electrical machinery such as "cell motor", "discharge resistor", and "insulated transformer" tend to come before the word "cover". In this embodiment, these regular patterns 108 are comprehensively learned, and compound words that match these patterns are extracted from the text, enabling detection of even unknown compound words.

[0022] FIG. 3 is a diagram for explaining a pattern learning process according to Patent Document 1. The electronic device of Patent Document 1 executes the following processes.

[0023] <Process 1> Morphological analysis is performed on the learning example sentence "Go to Ueno Park Hall" 201, and the sentence is divided into a word segmentation result 202. <Process 2> Specify the part of speech of each word as a part-of-speech specification result 203. <Process 3> Extract, as a word group extraction result 204, the arrangement of parts of speech defined in advance manually from the part-of-speech specification result 203. In this example, two or more consecutive nouns are defined in advance as candidates for compound words, and a word sequence such as "Ueno", "Park", and "Hall" is extracted.

[0024] <Process 4> Generate, as a pattern candidate 205, candidates for the pattern of the word arrangement from these three words. Here, "all combinations of words and parts of speech" are generated. For example, a pattern in which two common nouns come after the word "Ueno", a pattern in which the word "Park" comes after a place name and then a common noun comes, and a pattern in which the word "Hall" comes after a place name and then a common noun comes are generated.

[0025] <Process 5> From among these pattern candidates 205, determine, as a formal pattern 206, those with a high appearance ratio of the corresponding pattern and a large number of times the compound word matching the corresponding pattern has been searched on the web. In this example, the pattern 206 configured in the order of "place name → common noun → "Hall"" is selected.

[0026] FIG. 4 is a diagram for explaining the pattern learning process according to the present embodiment. The compound word detection device 1 of the present embodiment executes the following processes.

[0027] <Process 11> According to the word appearance frequency in the text corpus 31, the learning example sentences are divided into a plurality of units called sub-words, and the sub-word segmentation result 301 is obtained. For example, since "Ueno Park Hall" frequently appears in the text corpus 31 in the same order, it is not further divided and becomes one sub-word. <Process 12> Automatically classify the positive examples (i.e., compound words) 302 from the sub-word segmentation result 301, and automatically generate negative examples (sequences of words that are not compound words) 303. Here, the criterion for distinguishing positive and negative examples is whether the sub-word can be further divided into words by morphological analysis. The fact that a certain sub-word can be divided into words means that a word sequence that can be grammatically divided into several parts is used as an integrated compound word within the text corpus 31 (details will be described later). <Process 13> Divide the positive example 302 into words.

[0028] <Process 14> For each of the positive example 304 and the negative example 305, calculate a feature vector (i.e., feature quantity) in units of word sequences, and perform learning of the time series model 35 (i.e., model). In this case, the pattern 306 of the word sequence is expressed as a feature vector by the time series model 35. Here, the feature vector is a vector having elements such as the part of speech, inflection form, word vector of the word, and whether the word contains an alphabet. However, the elements are not limited to these. Also, the word vector is a unique vector assigned to each word sequence unit, and is characterized in that words with high relevance have similar vectors. The word vector may be calculated in any manner, but for example, it is effective to calculate it using Word2Vec.

[0029] The time-series model 35 is a model that takes the feature vector of a certain word sequence as input and outputs whether the word sequence is a compound word. The compound word detection device 1 uses a large number of combinations of the feature vectors of word sequences as samples and the labels indicating whether the word sequences are compound words as "supervised learning data" to perform machine learning on the time-series model 35.

[0030] The advantages of this embodiment compared with Patent Document 1 are the following two points.

[0031] First, this embodiment detects compound words based on the learned time-series model 35 (i.e., the learned model), while Patent Document 1 detects compound words by extracting a word sequence that matches the learned pattern (in the example of FIG. 3, a compound word composed of a place name → a common noun → "hall" in this order) from the target sentence. That is, in Patent Document 1, since it is necessary to manually define the pattern of the part-of-speech sequence in advance, learning is time-consuming and there may be omissions in the pattern. On the other hand, in this embodiment, the pattern is comprehensively and automatically learned by the time-series model 35, so learning is not time-consuming and there are no omissions in the pattern.

[0032] Second, patterns that depend on the domain, such as equipment model numbers, often do not exist on the web, and Patent Document 1 that uses web search results cannot handle them. On the other hand, this embodiment does not require web search.

[0033] FIG. 5 is a diagram for explaining the pattern detection process according to Patent Document 1. The electronic device of Patent Document 1 executes the following processing.

[0034] <Process 21> Perform morphological analysis on the sentence "Going to Osaka Castle Hall" to be detected, split the sentence into words, and obtain the word segmentation result 401. <Process 22> Specify the part-of-speech of the words in the word segmentation result 401 and obtain the part-of-speech specification result 402. <Process 23> Extract the word sequence "Osaka Castle Hall" that matches the learned pattern as the compound word detection result 403.

[0035] FIG. 6 is a diagram for explaining the pattern detection process according to the present embodiment. The compound word detection device 1 of the present embodiment executes the following processes.

[0036] <Process 31> Perform morphological analysis on the sentence "Going to Osaka Castle Hall" which is the detection target, decompose the sentence into words, and obtain the word segmentation result 501. <Process 32> Calculate a feature vector for a word sequence that is part of the word segmentation result 501, and input the calculated feature vector into the learned time series model 35, thereby detecting "Osaka Castle Hall" as the compound word detection result 502.

[0037] FIG. 7 is a block diagram of the learning process. The learning sentence corpus 31 includes the example sentence "I had a stomachache and made a dash to the convenience store" 601. Here, it is assumed that the learning of subword segmentation has been completed using the same sentence corpus 31. Subword segmentation is a method of dividing a sentence based on the statistical information of the sentence corpus 31. It treats words with high occurrence frequencies as a single vocabulary without dividing them, and divides low-frequency words into shorter vocabularies. In this case, the learning in subword segmentation is to calculate frequency information using the sentences in the sentence corpus 31.

[0038] As another example of subword segmentation, there are methods such as treating a plurality of consecutive nouns as one subword, treating a plurality of consecutive verbs as one subword, and treating a plurality of consecutive Chinese characters as one subword, regardless of the occurrence frequency. In any case, the resulting vocabulary is called a subword. As the subword segmentation unit 21, for example, a library such as SentencePiece can be used.

[0039] As the learning process according to this embodiment, first, the sub-word splitting unit 21 splits the example sentence 601 into sub-words. As a result, the example sentence 601 is split into sub-words 602 of "stomachache", "te", "dash", "de", "convenience store", "to", "rush into", and "da". The slash " / " in FIG. 7 indicates the boundary of the sub-words.

[0040] Subsequently, the morphological analysis unit 22 performs morphological analysis on the split sub-words 602 and splits the sub-words 602 into words. As the morphological analysis unit 22, libraries such as ChaSen (tea whisk) and Mecab can be used. At this time, the sub-word "stomachache" is further split into words "stomach" and "ache". Similarly, the sub-word "dash" is split into words "dash" and "rush". The sub-word "convenience store" is decomposed into words "convenience" and "store". The sub-word "rush into" is split into words "rush" and "into" (reference numeral 603). The double slash " / / " in FIG. 7 indicates the boundary when the sub-word is further split into words by morphological analysis.

[0041] The compound word selection unit 23 regards the sub-words 603 that are further split into a plurality of words by morphological analysis as compound words and stores them in the compound word dictionary (positive examples) 33. In this example, "stomachache", "dash", "convenience store", and "rush into" are determined as compound words and registered in the compound word dictionary (positive examples) 33. Compound words can be further added manually to the compound word dictionary (positive examples) 33. It should be noted that the compound word selection unit 23 does not register "stomachache" by splitting it into "stomach" and "ache" and registering each of them in the compound word dictionary (positive example 33), but registers "stomachache" as it is in the compound word dictionary (positive example 33).

[0042] On the other hand, the non-compound word generation unit 24 concatenates the sub-words 603 that did not become compound words, or a part of the words constituting the compound word, to generate a non-compound word consisting of a word order that cannot become a compound word, and registers it in the non-compound word dictionary (negative example) 34. In this example, vocabulary such as "it hurts", "dash convenience store", and "run to the store" are generated as non-compound words and registered in the non-compound word dictionary (negative example) 34.

[0043] Then, the compound word pattern learning unit 25 uses these two dictionaries as teacher-assisted learning data to train the time series model 35. That is, the compound word pattern learning unit 25 uses the feature vectors calculated in word sequence units for the compound words and non-compound words registered in each dictionary as the input to the time series model 35. Moreover, the compound word pattern learning unit 25 trains the time series model 35 using the compound words registered in the compound word dictionary (positive example) 33 as positive examples and the non-compound words registered in the non-compound word dictionary (negative example) 34 as negative examples.

[0044] In this way, by automatically creating teacher-assisted learning data by the compound word selection unit 23 and the non-compound word generation unit 24, it becomes unnecessary to generate learning labels manually, and a trained time series model 35 is automatically generated just by preparing the text corpus 31. Also, by inputting the feature vector of the compound word instead of the word itself that constitutes the compound word into the time series model 35, it is possible to detect a compound word (i.e., an unknown word) having a similar feature even if it does not exist in the text corpus 31.

[0045] Here, the sub-word splitting in FIG. 7 may be any method as long as it can split the text into a plurality of word sequences. For example, the methods of Non-Patent Document 1 or Non-Patent Document 2 may be used. Also, these methods may be combined.

[0046] For example, the sub-word splitting unit 21 may split the sentences in the sentence corpus 31 into separate sub-words using three methods: sub-word splitting here, the method of Non-Patent Document 1, and the method of Non-Patent Document 2. Then, for all the sub-words resulting from the splitting, processing below morphological analysis may be executed. In this case, the variations of compound words registered in the compound word dictionary (positive examples) 33 increase, and the types of compound words that the time-series model 35 can represent become abundant.

[0047] Note that the following relationships hold between terms. · A sentence contains a plurality of word sequences. · One word sequence contains a plurality of consecutive words. · A sub-word is either one word sequence or one word (such as a particle). · A feature vector is defined for one word sequence (= a plurality of consecutive words). Note that other examples will be described later. · For sub-words (excluding those consisting of one word), selection between compound words and non-compound words is made.

[0048] Figure 8 is a flowchart of the learning process. In step S701, the sub-word splitting unit 21 executes learning for performing sub-word splitting using all the sentences in the sentence corpus 31. In step S702, the sub-word splitting unit 21 determines whether the processing of step S701 has been performed for all the texts in the sentence corpus 31. If the sub-word splitting unit 21 has performed the processing of step S701 for all the texts in the sentence corpus 31 (step S702 “Yes”), it proceeds to step S703; otherwise (step S702 “No”), it returns to step S701.

[0049] In step S703, the sub-word splitting unit 21 performs sub-word splitting on the sentences (example sentences) in the sentence corpus 31. The sub-word splitting unit 21 may use the above-mentioned multiple methods in combination. If different sub-words are obtained for each method, the sub-word splitting unit 21 may use the union of them as the sub-words. In step S704, the morphological analysis unit 22 divides the sub-word into words by performing morphological analysis on the sub-word. In step S705, the compound word selection unit 23 determines that the sub-word divided into words by morphological analysis is a compound word and stores it in the compound word dictionary (positive example) 33.

[0050] In step S706, the non-compound word generation unit 24 generates a non-compound word consisting of a word order that cannot be a compound word and stores it in the non-compound word dictionary (negative example) 34. In step S707, the compound word pattern learning unit 25 determines whether the processes of steps S703 to S706 have been performed on all the texts in the text corpus 31. If the processes of steps S703 to S706 have been performed on all the texts in the text corpus 31 (step S707 “Yes”), the compound word pattern learning unit 25 proceeds to step S708; otherwise (step S707 “No”), it returns to step S703.

[0051] In step S708, the compound word pattern learning unit 25 calculates a feature vector for each word sequence unit for the compound words stored in the compound word dictionary (positive example) 33 and the non-compound words stored in the non-compound word dictionary (negative example) 34. In step S709, the compound word pattern learning unit 25 uses the compound words stored in the compound word dictionary (positive example) 33 as positive examples and the non-compound words stored in the non-compound word dictionary (negative example) 34 as negative examples to train the time series model 35. Then, the training process ends.

[0052] Figure 9 is a diagram for explaining an example of the process by the time series model 35. The time series model 35 is typically a neural network that takes a feature vector of a word sequence as input. Examples of such a time series model 35 include RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc.

[0053] Now, assume that example sentence v is composed of words w1, w2, w3, w4, w5, …, wn. And assume that the feature vector of the word sequence “w1+w2” is x2, the feature vector of the word sequence “w1+w2+w3” is x3, the feature vector of the word sequence “w1+w2+w3+w4” is x4, …, and so on. The compound word pattern learning unit 25 calculates the feature vectors x2, x3, x4, … in order. Next, the compound word pattern learning unit 25 separates the first word w1 and repeats the same process. Then, it separates the next word w2 and repeats the same process. Note that these premises and the flow of the process are the same in FIG. 10 as well.

[0054] More specifically, assume that the example sentence v is “going to Osaka Castle Hall”. The compound word pattern learning unit 25 first calculates the following first group of feature vectors, then calculates the following second group of feature vectors, and then generates the following third group of feature vectors.

[0055] 〈First Group〉 Feature vector of “Osaka / Castle” Feature vector of “Osaka / Castle / Hall” Feature vector of “Osaka / Castle / Hall / to” Feature vector of “Osaka / Castle / Hall / to / go”

[0056] 〈Second Group〉 Feature vector of “Castle / Hall” Feature vector of “Castle / Hall / Hall” Feature vector of “Castle / Hall / to / go”

[0057] 〈Third Group〉 Feature vector of “Hall / to” Feature vector of “Hall / to / go”

[0058] Now, for the sake of convenience of explanation, pay attention to the above-mentioned <First Group>. At time point h1, the compound word pattern learning unit 25 inputs the feature vector x2 of "Osaka / Castle" to the time series model 35. At time point h2, the compound word pattern learning unit 25 inputs the feature vector x3 of "Osaka / Castle / Hall" to the time series model 35. At time point h3, the compound word pattern learning unit 25 inputs the feature vector x4 of "Osaka / Castle / Hall / to" to the time series model 35. At time point h4, the compound word pattern learning unit 25 inputs the feature vector x5 of "Osaka / Castle / Hall / to / go" to the time series model 35 (it may continue further in the case of other examples).

[0059] As a result of the above, the time series model 35 outputs the estimated result vector y2 at time point h1, outputs the estimated result vector y3 at time point h2, outputs the estimated result vector y4 at time point h3, and outputs the estimated result vector y5 at time point h4.

[0060] The compound word pattern learning unit 25 receives the estimated result vector y2 from the time series model 35 at time point h1, receives the estimated result vector y3 at time point h2, receives the estimated result vector y4 at time point h3, and receives the result estimation vector y5 at time point h4. The estimated result vector yn is the estimated result of whether the word sequence (a continuous part in the example v) corresponding to the feature vector xn is a compound word or not.

[0061] The result estimation vector 801 may be the two-dimensional vector "(1,0)" when the word sequence corresponding to the feature vector xn is included in the compound word dictionary (positive example) 33. The result estimation vector 801 may be the two-dimensional vector "(0,1)" when the word sequence corresponding to the feature vector xn is included in the non-compound word dictionary (negative example) 34.

[0062] (Variant example of the feature vector) In the above, for example, an example in which one feature vector is defined for the entire word sequence of "Osaka / Castle / Hall" was described. However, for each of "Osaka", "Castle", and "Hall", one feature quantity vector may be defined. In such an example, the feature quantity vector can easily represent the part-of-speech of each individual word.

[0063] FIG. 10 is a diagram for explaining another example of the processing by the time series model 35. Compared with FIG. 9, in FIG. 10, the content of the estimated result vector output by the time series model 35 is slightly different. In FIG. 10, the time series model 35 outputs, for example, the estimated result vector 901 as "yn=(a,b)". However, "a + b = 1, 0 ≦ a < 1, 0 ≦ b ≦ 1" holds.

[0064] At this time, using a predetermined constant T, if a ≧ T, the word sequence is a compound word (positive example), and its likelihood (confidence of being a compound word) is "a". If b < T, the word sequence is a non-compound word (negative example), and its likelihood (confidence of being a non-compound word) is "b". Here, T is generally a numerical value near 0.5, but this is not the limit.

[0065] Note that the time series model 35 used for learning may be any of the following neural networks, or any of the following improved systems, or a combination of these, not limited to just RNN and LSTM. ·GRU (Gated Recurrent Unit) ·BERT (Bidirectional Encoder Representations from Transformers) ·Bi-directional RNN ·Bi-directional LSTM ·Bi-directional GRU

[0066] The composite word pattern learning unit 25 may prepare, for example, two types of discriminators, namely, a discriminator using an RNN and a discriminator using an LSTM, and use the average value of the outputs calculated by each discriminator as the final result. In addition to the time series model 35, the composite word pattern learning unit 25 may use, for example, a model capable of two-class classification (two-class output) as follows. · CNN (Convolutional Neural Network) · SVM (Support Vector Machine) · RF (Random Forest) · DT (Decision Tree)

[0067] The reason why the time series model 35 is named "time series" is that, as is clear from FIGS. 9 and 10, for the time series model 35, the feature vectors of two words starting with a certain word, the feature vectors of three words, the feature vectors of four words,... are input in order in time series.

[0068] FIG. 11 is a block diagram of a process for detecting a composite word from a text to be detected. The text to be detected here is "Going to the Hall of Osaka Castle" 1001.

[0069] First, the morphological analysis unit 22 performs morphological analysis on the input text and splits the text into words.

[0070] Next, the word selection unit 26 selects combinations of a plurality of consecutive words from the text split into words. For example, from the text split into "Osaka / Castle / Hall / to / go", combinations such as "Osaka / Castle", "Osaka / Castle / Hall", "Osaka / Castle / Hall / to", "Osaka / Castle / Hall / to / go",..., "Hall / to / go", "to / go" are selected as word sequences. Note that " / " indicates the boundary of a word.

[0071] Furthermore, the compound word candidate detection unit 27 calculates a feature vector for each word sequence and inputs it into the learned time-series model 35. Then, it extracts the word sequences determined to be compound words as compound word candidates and calculates their likelihoods.

[0072] Finally, the compound word detection unit 28 determines the compound word 1002, which is the final detection result, from among the extracted compound word candidates. Here, "Osaka Castle Hall" is detected as a compound word because the time-series model 35 has been machine-learned using, for example, "Osaka Castle Hall" as positive example supervised learning data. Even if the time-series model 35 had been machine-learned using other word sequences such as "Himeji Castle Hall" and "Hikone Castle Hall" as positive example supervised learning data, the result would be the same.

[0073] FIG. 12 is a diagram for explaining the process of determining the compound word, which is the final result, from among the compound word candidates. For example, when the compound word candidate detection unit 27 extracts "Osaka / Castle / Hall", "Osaka / Castle", "Castle / Hall", "Osaka / Castle", and "Hall / Hall" as compound word candidates, the compound word detection unit 28 constructs a graph structure as shown in FIG. 12 using these word sequences.

[0074] The compound word detection unit 28 sets a directed graph between the word sequences so that the same word is not encountered on any route from the start point to the end point. Also, the compound word detection unit 28 shows, below the word sequence, the likelihood output as a result of inputting the feature vector of the word sequence into the learned time-series model 35. At this time, the problem of determining a compound word from these word sequences can be defined as the problem of finding the optimal route from these multiple routes. The compound word detection unit 28 calculates the evaluation value E(r) defined as "Equation 1" below, determines the route r that maximizes that value, and extracts the word sequence existing on that route as the compound word.

[0075]

Equation

[0076] Here, r is the root number. n(r) is the number of words existing on root r. L(k) is the likelihood of the word sequence to which the k-th word of root r belongs. Then, in the example of FIG. 12, E(1), E(2), E(3), and E(4) have the following values.

[0077] E(1) = 0.6 + 0.6 + 0.6 + 0.6 = 2.4 E(2) = 0.9 + 0.9 + 0.9 = 2.7 E(3) = 0.5 + 0.5 = 1.0 E(4) = 0.7 + 0.7 + 0.6 + 0.6 = 2.6

[0078] The compound word detection unit 28 determines root 2 that maximizes the evaluation value as the optimal root. The compound word detection unit 28 finally detects the word sequence "Osaka Castle Hall" existing on root 2 as a compound word. Thus, the evaluation value E(r) may be calculated for all roots. However, when the length of the word sequence that is a candidate for a compound word increases, the number of roots also increases explosively accordingly. Therefore, the compound word detection unit 28 may reduce the calculation amount by using, for example, dynamic programming. Further, the definition of the evaluation formula E(r) and the method for determining a compound word according to FIG. 12 are examples, and the compound word detection unit 28 may more simply detect, for example, a word sequence whose likelihood is high enough to satisfy a predetermined criterion as a compound word.

[0079] In the above, for example, an example in which one feature vector is defined for the entire word sequence "Osaka / Castle / Hall", and an example in which one feature amount vector is defined for each of "Osaka", "Castle", and "Hall" have been described. The information processing of FIG. 12 is particularly suitable for the latter example.

[0080] FIG. 13 is a flowchart of a process for detecting a compound word from a sentence to be detected. In step S1201, the morphological analysis unit 22 performs morphological analysis on the sentence to be detected and divides it into words. In step S1202, the word selection unit 26 selects a combination of a plurality of consecutive words as a word sequence. In step S1203, the compound word candidate detection unit 27 calculates a feature vector for each word sequence unit.

[0081] In step S1204, the compound word candidate detection unit 27 inputs the feature vector into the learned time series model 35. The time series model 35 outputs a result (estimated result vector) indicating whether the word sequence is a compound word. In step S1205, the compound word candidate detection unit 27 determines whether the word sequence is a compound word. If the output of the learned time series model 35 is "is a compound word" (step S1205 "Yes"), the process proceeds to step S1206; otherwise (step S1205 "No"), the process proceeds to step S1207.

[0082] In step S1206, the compound word candidate detection unit 27 extracts the word sequence as a compound word candidate and stores the likelihood in the main storage device 14. In step S1207, the compound word candidate detection unit 27 determines whether the processes of steps S1203 to S1206 have been performed for all word sequences. If it is determined that the processes have been performed (step S1207 "Yes"), the process proceeds to step S1208; otherwise (step S1207 "No"), the process returns to step S1204.

[0083] In step S1208, the compound word detection unit 28 reads the likelihood from the main storage device 14 for the character string extracted as a compound word candidate, and determines the compound word in the manner shown in FIG. 12. Then, the process of detecting the compound word ends.

[0084] FIG. 14 is an example of a display screen for interactively learning the time series model 35. Just by the user preparing the learning text corpus 31, the compound word dictionary (positive example) 33 and the non-compound word dictionary (negative example) 34 are automatically generated. Further, the compound word detection unit 28 displays the display screen of FIG. 14 on the output device 13. Then, the user can learn the time series model 35 by pressing the learning execution button 1301. On the display screen of FIG. 14, there are a column 1302 for displaying the compound words registered in the compound word dictionary (positive example) 33, a column 1303 for displaying the non-compound words registered in the non-compound word dictionary (negative example) 34, buttons 1304 and 1305 for selecting compound words and non-compound words, buttons 1306 and 1307 for deleting compound words and non-compound words, text boxes 1308 and 1309 for inputting an arbitrary word sequence, and buttons 1310 and 1311 for additionally registering these in each dictionary.

[0085] (Effect of the Embodiment) (1) The compound word detection device can detect compound words from among the unknown words in a specialized field using a time series model. (2) The compound word detection device can accurately create sub-words corresponding to positive examples of compound words. (3) The compound word detection device can use many word segmentation methods for sub-word segmentation. (4) The compound word detection device can surely create positive examples that can be compound words. (5) The compound word detection device can use word vectors or the like as feature amounts of a word sequence. (6) The compound word detection device can use a model with achievements in the field of natural language processing.

[0086] Note that the present invention is not limited to the above-described embodiments, and includes various modifications. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Further, it is possible to add, delete, or replace other configurations for a part of the configuration of each embodiment.

Explanation of Signs

[0087] 1 Compound word detection device 11 Central control device 12 Input device 13 Output device 14 Main memory device 15 Auxiliary storage device 16 Communication device 21 Sub-word division unit 22 Morphological analysis unit 23 Compound word selection unit 24 Non-compound word generation unit 25 Compound word pattern learning unit 26 Word selection unit 27 Compound word candidate detection unit 28 Compound word detection unit 31 Text corpus 32 Morphological analysis dictionary 33 Compound word dictionary (positive example) 34 Non-compound word dictionary (negative example) 35 Time series model

Claims

1. a compound word pattern learning unit that learns a model for detecting whether a plurality of consecutive words is a compound word or not, using feature amounts of a plurality of consecutive words that are the result of dividing a compound word registered in a compound word dictionary; Inputting features of a plurality of consecutive words to be detected into the trained model; a compound word detection unit that receives an output from the trained model indicating whether the plurality of consecutive words to be detected are a compound word; A compound word detection device comprising:

2. The compound word dictionary includes: storing compound words selected from word strings generated by applying a method for dividing a sentence into multiple word strings to a sentence corpus; 2. The compound word detection device according to claim 1,

3. The method of dividing the sentence into a plurality of word strings includes the steps of: Being plural, 3. The compound word detection device according to claim 2, wherein

4. The compound word dictionary includes: storing, as a compound word, a word string that can be further divided by morphological analysis from among the word strings generated by applying the method for dividing a sentence into a plurality of word strings to a sentence corpus; 4. The compound word detection device according to claim 2, further comprising:

5. The feature amount of the plurality of consecutive words is The word string includes at least one of the part of speech, conjugation form, and word vector of each word in the word string, and the presence or absence of alphabets in each word in the word string; 2. The compound word detection device according to claim 1,

6. The model is A neural network that uses time-series features of multiple consecutive words as input, and / or a model that outputs two classes, The neural network comprises: At least one of RNN, LSTM, Bi-directional LSTM, GRU, Bi-directional GRU, and BERT; The model that outputs the two classes is Including at least one of CNN, SVM, RF, and DT; 2. The compound word detection device according to claim 1,

7. A model is trained so as to be able to detect whether a plurality of consecutive words is a compound word or not by using the feature values ​​of the plurality of consecutive words that are the result of dividing a compound word registered in a compound word dictionary, and the feature values ​​of the plurality of consecutive words to be detected are inputted; a compound word detection unit that receives an output from the trained model indicating whether the plurality of consecutive words to be detected are a compound word; A compound word detection device comprising:

8. The feature amount of the plurality of consecutive words is The word string includes at least one of the part of speech, conjugation form, and word vector of each word in the word string, and the presence or absence of alphabets in each word in the word string; 8. The compound word detection device according to claim 7,

9. The model is A neural network that uses time-series features of multiple consecutive words as input, and / or a model that outputs two classes, The neural network comprises: At least one of RNN, LSTM, Bi-directional LSTM, GRU, Bi-directional GRU, and BERT; The model that outputs the two classes is Including at least one of CNN, SVM, RF, and DT; 8. The compound word detection device according to claim 7,

10. The compound word pattern learning unit of the compound word detection device includes: A model is trained to detect whether multiple consecutive words are a compound word or not, using features of multiple consecutive words that are the result of dividing a compound word registered in a compound word dictionary; The compound word detection unit of the compound word detection device includes: Inputting features of a plurality of consecutive words to be detected into the trained model; receiving an output from the trained model indicating whether the plurality of consecutive words to be detected are a compound word; A compound word detection method characterized by:

Citation Information

Patent Citations

  • Electronic device, morphological element compounding method, and its program

    JP2010009355A