Document processing apparatus, document processing method, and document processing program
The document processing apparatus addresses the challenge of representing low-frequency and unknown words by dividing them into sub-words and synthesizing with class vectors, effectively embedding meaning without additional learning, improving semantic accuracy.
Patent Information
- Application Number
- JP2021144245
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-03
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2041-09-03
AI Technical Summary
Existing natural language processing technologies struggle to generate distributed representations for words with low appearance frequency or those not included in the learned space, such as sub-words and unknown words, leading to insufficient meaning representation.
A document processing apparatus that utilizes morphological analysis to divide low-frequency words into sub-words, synthesizes position vectors with class vectors representing word types, and corrects distributed representations to embed meaning without additional learning.
Appropriately embeds meaning for low-frequency and unknown words, generating accurate distributed representations without additional training, enhancing semantic accuracy.
Smart Images

Figure 0007705314000001 
Figure 0007705314000002 
Figure 0007705314000003
Abstract
Description
Technical Field
[0001] The present invention relates to a document processing apparatus, a document processing method, and a document processing program, and is suitable for application to a document processing apparatus, a document processing method, and a document processing program that generate a distributed representation of words.
Background Art
[0002] Conventionally, in the field of natural language processing, a technique of distributed representation that represents a word as a vector composed of a plurality of continuous values and can represent features between words is known.
[0003] As a document related to this technical field, for example, Patent Document 1 can be cited. Patent Document 1 has an issue of "correcting the values of word vectors that represent the features of words with a multi-dimensional vector in a simple manner", and includes "a thesaurus storage unit 166 that stores the categories of words, and a first word vector storage unit 164 that stores a distributed representation model including word vectors that represent the features of words with a multi-dimensional vector. A category vector calculation unit 144 calculates a category vector that represents the word vectors of words belonging to the same category, and a word vector correction unit 146 corrects the word vectors to be close to the category vector of the category to which the word belongs." A natural language device is disclosed.
[0004] Here, in the distributed representation, by learning features from the context before and after each word using a large corpus, a vector representing the features of the word and the relevance to another word can be obtained. However, when the word has a low appearance frequency in the corpus to be learned or the context in which the word appears is biased, there is a problem that the exact features of the word cannot be obtained.
[0005] Regarding the above problem, Patent Document 1 describes a mechanism in which a thesaurus is used as external knowledge to correct the distributed representation of each word to be close to the category vector representing the category according to the category to which the word belongs.
Prior Art Documents
Patent Document
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, the technology disclosed in Patent Document 1 only corrects words represented in the learned distributed representation space, and does not disclose how to handle words not included in the distributed representation space (for example, sub-words split into partial strings, unknown words, etc., words that do not represent sufficient meaning by themselves). Therefore, the natural language device disclosed in Patent Document 1 has a problem that even for words included in the thesaurus, it is impossible to create a distributed representation for words not included in the learned distributed representation space.
[0008] The present invention has been made in consideration of the above points, and it is intended to propose a document processing apparatus, a document processing method, and a document processing program that can appropriately embed the meaning of words without performing additional learning for words that cannot sufficiently represent their meaning in a pre-learned distributed representation space.
Means for Solving the Problems
[0009] In order to solve such problems, the present invention provides a document processing apparatus that generates a distributed representation of words included in text input from the outside, the apparatus comprising: a class vector storage unit that stores class vectors, which are reference vectors for each type of word; a morphological analysis unit that divides the text into a word sequence by morphological analysis and attach information indicating the type to which each word sequence belongs ; a word distributed representation generation unit that converts the word sequence divided by the morphological analysis unit into a distributed representation for each word included in the word sequence; before a word distributed representation correction unit that corrects the distributed representation by the word distributed representation generation unit thus generated distributed representation to correct ; and is provided with, when the text includes low-frequency words whose appearance frequency in pre-training is lower than a predetermined level, after the morphological analysis unit divides the text into word sequences by morphological analysis, the morphological analysis unit further divides the low-frequency words included in the word sequences into sub-words of partial character strings, the word distribution representation generation unit converts the low-frequency words into a plurality of distribution representations each consisting of the distribution representation of each sub-word included in the low-frequency word, the word distribution representation correction unit synthesizes a position vector that fixes the order of the sub-words in the low-frequency word with the plurality of distribution representations of the low-frequency word generated by the word distribution representation generation unit, and then synthesizes the class vector of the type to which the low-frequency word belongs, and replaces the synthesized distribution representation with the distribution representation of the low-frequency word A document processing apparatus is provided.
[0010] Also, in order to solve such a problem, in the present invention, there is provided a document control method by a document processing apparatus that generates a distributed representation of words included in text input from the outside. The document processing apparatus stores class vectors that are reference vectors for each type of word. The document processing apparatus divides the text into a word sequence by morphological analysis and attach information indicating the type to which each word sequence belongs in a morphological analysis step, the document processing apparatus converts the word sequence divided in the morphological analysis step into a distributed representation for each word included in the word sequence in a word distributed representation generation step, and the document processing apparatus , before in the word distributed representation generation step generated distributed representation to correct in a word distributed representation correction step, and includes , when the text includes low-frequency words whose appearance frequency in pre-training is lower than a predetermined level, in the morphological analysis step, after dividing the text into word sequences by morphological analysis, the morphological analysis unit further divides the low-frequency words included in the word sequences into sub-words of partial character strings, in the word distribution representation generation step, the low-frequency words are converted into a plurality of distribution representations each consisting of the distribution representation of each sub-word included in the low-frequency word, in the word distribution representation correction step, a position vector that fixes the order of the sub-words in the low-frequency word is synthesized with the plurality of distribution representations of the low-frequency word generated in the word distribution representation generation step, and then the class vector of the type to which the low-frequency word belongs is synthesized, and the synthesized distribution representation is replaced with the distribution representation of the low-frequency word A document processing method is provided.
[0011] Also, in order to solve such a problem, in the present invention, there is provided a document control program for causing a computer that constitutes a document processing apparatus that generates a distributed representation of words included in text input from the outside to execute. The document processing apparatus stores class vectors that are reference vectors for each type of word , before divides the text into a word sequence by morphological analysis and attach information indicating the type to which each word sequence belongs in a morphological analysis process, and , before converts the word sequence divided in the morphological analysis process into a distributed representation for each word included in the word sequence in a word distributed representation generation process, and before in a word distributed representation correction process that corrects the distributed representation generated in the word distributed representation generation process already distributed representation to correct causes the computer to execute, and When the text includes low-frequency words whose appearance frequency in pre-training is lower than a predetermined level, in the morphological analysis process, after splitting the text into a word sequence by morphological analysis, further split the low-frequency words included in the word sequence into sub-words of partial character strings. In the word distribution representation generation process, convert the low-frequency words into a plurality of distribution representations composed of the distribution representations for each sub-word included in the low-frequency words. In the word distribution representation correction process, after synthesizing a position vector that fixes the order of the sub-words in the low-frequency words to the plurality of distribution representations of the low-frequency words generated in the word distribution representation generation process, synthesize it with the class vector of the type to which the low-frequency word belongs, and replace the synthesized distribution representation with the distribution representation of the low-frequency word. A document processing program is provided.
Advantages of the Invention
[0012] According to the present invention, for words that cannot fully represent meaning in a pre-trained distributed representation space, the meaning of the words can be appropriately embedded without additional learning. Other problems, configurations, and effects than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0015] (1) First Embodiment In the first embodiment, regarding the distributed representation of proper nouns, a reference vector (class vector) representing the type is generated for each type of proper noun, and the distributed representation of the proper noun with the meaning embedded is generated by synthesizing the distributed representation of the words constituting the proper noun and the class vector. An example of a document processing apparatus will be described. The proper nouns handled in this embodiment are underlearned words that cannot sufficiently represent the meaning in the pre-trained distributed representation space, and correspond to an example of low-frequency words with a low appearance frequency in previous learning and insufficient accuracy.
[0016] (1-1) Configuration FIG. 1 is a block diagram showing an example of the functional configuration of a document processing apparatus 100 according to the first embodiment of the present invention. The document processing apparatus 100 is a computer (information processing apparatus) in which a document processing program is installed. As shown in FIG. 1, it includes an input unit 101, a morphological analysis unit 102, a morphological analysis dictionary 103, a pre-trained neural language model 104 (word distributed representation generation unit 105, sentence distributed representation generation unit 106), an output unit 107, a target word extraction unit 108, a word distributed representation correction unit 109, a class vector confirmation unit 110, a position vector storage unit 111, a class vector storage unit 112, and a class vector generation unit 113. Each of the above units will be described in detail after FIGS. 2 and 3 are described. Note that the connection relationship between the configurations shown in FIG. 1 represents only a typical connection example and does not represent all connection relationships.
[0017] FIG. 2 is a diagram showing a configuration example of a document processing system 200 including a document processing apparatus 100. When the document processing system 200 is realized as, for example, a client-server system, as shown in FIG. 2, a server 201 which is a client server and a plurality of user terminals 203 used by a user are communicably connected via a network 202. The user terminal 203 is an information processing apparatus such as a personal computer (PC) or a smartphone. The network 202 may be any network such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network).
[0018] When the document processing system 200 is realized as a client-server system, the document processing program executed by the document processing apparatus 100 is installed in the server 201. Therefore, the server 201 executes morphological analysis, generation of distributed representation, generation of class vectors, generation of extended word vectors, etc. as the document processing apparatus 100. In this case, the user terminal 203 serves as an interface for executing input of a document to be processed, transmission of the document to be processed to the server 201, and reception of a processing result from the server 201.
[0019] Note that in the present embodiment, the document processing system 200 can also be realized in a stand-alone type. In this case, the document processing program executed by the document processing apparatus 100 is installed in the user terminal 203, and the server 201 and the network 202 are not essential. Therefore, the user terminal 203 executes input of a document to be processed, data conversion of the input document, output of a processing result, etc. as the document processing apparatus 100.
[0020] FIG. 3 is a block diagram showing a hardware configuration example of the document processing apparatus 100. The information processing apparatus 300 shown in FIG. 3 is an information processing apparatus that realizes the document processing apparatus 100 and corresponds to the server 201 and the user terminal 203 shown in FIG. 2.
[0021] As shown in FIG. 3, the information processing apparatus 300 includes a processor 301, a storage device 302, an input device 303, an output device 304, and a communication interface (communication IF) 305. The processor 301, the storage device 302, the input device 303, the output device 304, and the communication IF 305 are connected by a bus 306 which is an internal communication line.
[0022] The processor 301 is a processor that controls the information processing apparatus 300, and is, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The storage device 302 is a device that provides a working area for the processor 301, and is a non-temporary or temporary recording medium that stores various programs and data. Specifically, the storage device 302 is, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), or a flash memory. Note that the storage device 302 does not necessarily have to be a recording medium mounted inside the information processing apparatus 300, and a configuration may be adopted in which part or all of various programs and data are stored externally (such as an external storage medium or the cloud) connected to the information processing apparatus 300. Each processing unit of the document processing apparatus 100 shown in FIG. 1 is mainly realized by the processor 301 reading and executing various predetermined programs. Also, each storage unit (including the morphological analysis dictionary 103) of the document processing apparatus 100 shown in FIG. 1 is realized by the storage device 302. The input device 303 is a device for inputting data, and specifically, for example, a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, or a biosensor. The output device 304 is a device for outputting data, and specifically, for example, a display, a printer, or a speaker. The communication IF 305 is an interface that connects to the network 202 and transmits and receives data to and from the outside of the information processing apparatus 300, and specifically, for example, a NIC (Network Interface Card).
[0023] The functional configuration of the document processing apparatus 100 shown in FIG. 1 will be described.
[0024] The input unit 101 is a module that reads text data, and all processing and registration of text data are performed via the input unit 101. In the following description, the text included in the text data read by the input unit 101 may sometimes be simply referred to as text.
[0025] The morphological analysis unit 102 performs morphological analysis of the text input from the input unit 101, and outputs the results to the word distribution representation generation unit 105 of the pre-trained neural language model 104 and the target word extraction unit 108. Note that, as preprocessing for morphological analysis, the morphological analysis unit 102 performs processing such as unifying alphanumeric symbols and full-width katakana characters, and converting different font types such as old-style Chinese characters to common Chinese characters for Chinese characters, and performs normalization possible at the character unit.
[0026] In morphological analysis, the text is split into a word sequence using the morphological analysis dictionary 103, and reading and part-of-speech information are added. The morphological analysis dictionary 103 is a database having reading / part-of-speech information for each word, and the occurrence cost of words and the connection cost between words learned from a large amount of text corpora. Therefore, in the morphological analysis process by the morphological analysis unit 102, based on the information in the morphological analysis dictionary 103, the text of the text data is appropriately delimited at appropriate positions as Japanese and split into a word sequence.
[0027] The morphological analysis dictionary 103 is composed of a system dictionary which is a read-only dictionary prepared in advance on the system side, and a user dictionary which is a readable and writable dictionary that the user can arbitrarily register and modify. By editing the user dictionary, the user can newly register proper nouns in the morphological analysis dictionary 103. Note that some proper nouns may be registered in the system dictionary in advance. Also, the system dictionary of the morphological analysis dictionary 103 is the same dictionary as the morphological analysis dictionary used when creating the pre-trained neural language model 104 described later.
[0028] The morpheme analysis unit 102 further obtains a vocabulary list from the pre-trained neural language model 104 described later and performs word processing according to the vocabulary list. The vocabulary list is a list of the vocabulary included in the pre-trained neural language model 104. Since the neural language model needs to limit the number of vocabulary due to its mechanism, it divides the words segmented by normal morpheme analysis into sub-words of smaller partial character strings, i.e., sub-word segmentation, and performs unknown word processing to replace low-frequency words (words with low appearance frequencies in the previous learning) with special tokens indicating unknown words (words not treated as sufficiently learned words), thereby reducing the overall number of vocabulary.
[0029] Therefore, the morpheme analysis unit 102 re-divides the result of morpheme analysis using the morpheme analysis dictionary 103 into sub-word units defined in the vocabulary list, and replaces it with a special token indicating an unknown word if it does not exist in the vocabulary list.
[0030] The pre-trained neural language model 104 has a neural network and is a language model trained with a large corpus, which converts the word sequence output by the morpheme analysis unit 102 into a distributed representation that is a multi-dimensional real vector. The pre-trained neural language model 104 has a word distributed representation generation unit 105 and a sentence distributed representation generation unit 106.
[0031] As the pre-trained neural language model 104 having the word distributed representation generation unit 105 and the sentence distributed representation generation unit 106, for example, there is BERT (Bidirectional Encoder Representations from Transformers). At this time, the embedding layer in BERT corresponds to the word distributed representation generation unit 105, and the encoder layer corresponds to the sentence distributed representation generation unit 106.
[0032] The word distribution representation generation unit 105 converts the word sequence output by the morphological analysis unit 102 into a distribution representation (word distribution representation), which is a multi-dimensional real-valued vector for each word, through a neural network. More specifically, the word distribution representation generation unit 105 converts a word into a word ID, further converts this into a one-hot vector, and then inputs the one-hot vector into the neural network to obtain a distribution representation (word distribution representation) for each word. The word ID is a continuous integer uniquely assigned to the words in the vocabulary list.
[0033] A one-hot vector is a vector in which one component is "1" and the remaining components are all "0". A one-hot vector is a vector with the same dimension as the number of word IDs, and a word with a word ID of "n" is converted into a one-hot vector in which the nth element is "1". The word distribution representation generation unit 105 inputs the one-hot vector converted as described above into the neural network and converts it into a word distribution representation, which is a real-valued vector.
[0034] The sentence distribution representation generation unit 106 synthesizes the word distribution representations output by the word distribution representation generation unit 105 (or replaced by the word distribution representation correction unit 109 described later) to generate a distribution representation for the entire sentence (or word sequence). Note that the pre-trained neural language model 104 is not limited to BERT, and it only needs to be able to convert a word sequence into a distribution representation for each word and a distribution representation for the entire sentence. A language model with an architecture similar to BERT may also be used. Also, it may be configured such that Word2Vec is used in the word distribution representation generation unit 105, and the average of the word distribution representations output by the word distribution representation generation unit 105 is taken as the sentence distribution representation in the sentence distribution representation generation unit 106.
[0035] The output unit 107 outputs the distribution representation for the entire sentence (or word sequence) generated by the sentence distribution representation generation unit 106 by a predetermined output means using the output device 304 (such as display output to a display or data output to an output file). At this time, the output unit 107 may output a distribution representation (distribution vector) for the entire sentence, or may output a distribution vector for each word.
[0036] The target word extraction unit 108 extracts words of a predetermined part-of-speech type (e.g., proper nouns such as place names, personal names, restaurant names, etc.) corresponding to the class vectors stored in the class vector storage unit 112 from the word sequence output by the morphological analysis unit 102. Note that the words extracted here are the words before performing sub-word segmentation and out-of-vocabulary processing according to the vocabulary of the aforementioned pre-trained neural language model 104. In addition, the target word extraction unit 108 also extracts the word dispersion representations of the word sequences corresponding to the extracted proper nouns from the word dispersion representations output by the word dispersion representation generation unit 105.
[0037] The word dispersion representation correction unit 109 synthesizes the word dispersion representations of a plurality of words constituting the proper nouns output by the target word extraction unit 108, the class vectors stored in the class vector storage unit 112, and the position vectors stored in the position vector storage unit 111 to generate the word dispersion representation of the proper nouns. Then, the word dispersion representation correction unit 109 outputs the generated word dispersion representation of the proper nouns to the word dispersion representation generation unit 105, replaces the word dispersion representation group of the proper nouns in the word dispersion representation generation unit 105 with the generated word dispersion representation, and inputs the replaced word dispersion representation to the sentence dispersion representation generation unit 106.
[0038] The class vector confirmation unit 110 has a function of outputting, as a class vector candidate, the class vector generated by the class vector generation unit 113 described later to the output device 304 and receiving manual correction from the user to the class vector candidate by the input device 303. When the class vector confirmation unit 110 receives manual correction to the class vector candidate, it transfers the manually corrected class vector candidate to the class vector generation unit 113 and registers the class vector candidate as a class vector in the class vector storage unit 112 (alternatively, it may be directly registered in the class vector storage unit 112 from the class vector confirmation unit 110).
[0039] The position vector storage unit 111 stores the position vectors that the word distributed representation correction unit 109 synthesizes into the word distributed representation. The position vectors serve to provide position information by adding absolute values corresponding to their positions to the distributed representations of each word that constitutes a proper noun. Technologies for realizing position vectors include, for example, the position encoding of Transformer.
[0040] The class vector storage unit 112 stores the class vectors that the word distributed representation correction unit 109 synthesizes into the word distributed representation. The class vector storage unit 112 stores the class vectors created for each type of proper noun by the class vector creation process described later by the class vector generation unit 113.
[0041] The class vector generation unit 113 generates the class vectors that the word distributed representation correction unit 109 synthesizes into the word distributed representation and stores them in the class vector storage unit 112. The generation of the class vectors is executed before the document processing device 100 performs document processing.
[0042] (1-2) Class Vector Creation Process FIG. 4 is a flowchart showing an example of the processing procedure of the class vector creation process. The class vector creation process shown in FIG. 4 is a process executed at an arbitrary timing before the distributed representation generation process described later, and is a process in which the document processing device 100 shown in FIG. 1 creates class vectors and registers them in the class vector storage unit 112.
[0043] According to FIG. 4, first, the class vector generation unit 113 reads the vocabulary list of the pre-trained neural language model 104 via the input unit 101 (step S401). The vocabulary list is a list of all the vocabularies (all the vocabularies held by the word dispersion representation generation unit 105) included in the pre-trained neural language model 104 and can be obtained from the pre-trained neural language model 104. Further, when the vocabulary list read in step S401 includes sub-words, the target word extraction unit 108 excludes the sub-words and extracts only the words that are not divided into sub-words (step S402). The vocabulary list includes information on whether a word is a sub-word or not, and the target word extraction unit 108 excludes the sub-words based on this information. When the vocabulary list does not include sub-words, the process of step S402 is omitted.
[0044] Next, the morphological analysis unit 102 refers to the morphological analysis dictionary 103 to examine the part of speech of the words included in the vocabulary list read in step S401, other than the sub-words extracted in step S402, and extracts proper nouns (step S403). Then, the class vector generation unit 113 converts the words of the proper nouns extracted by the morphological analysis unit 102 in step S403 into word dispersion representations via the word dispersion representation generation unit 105 (step S404).
[0045] When all the proper nouns extracted by the morphological analysis unit 102 are converted into word dispersion representations in step S404, the class vector generation unit 113 calculates the average of the dispersion representations for each proper noun type in the word dispersion representation generation unit 105 to create a dispersion representation, and uses this as the class vector for each proper noun type (step S405). Note that the class vector created by the class vector generation unit 113 in step S405 can be corrected by the user in step S407 described later (that is, it is not necessarily the class vector finally registered in the class vector storage unit 112), so it can also be called a class vector candidate.
[0046] Furthermore, in order to visualize the word distribution representations of proper nouns and class vectors created in step S405, the class vector generation unit 113 performs dimensionality reduction on the word distribution representations of the words and creates a two-dimensional scatter plot (step S406). For dimensionality reduction, for example, principal component analysis can be used, but any other method may be used as long as it can aggregate features representing the differences in the distribution representations of each proper noun.
[0047] Next, the class vector generation unit 113 transmits the scatter plot created in step S406 to the class vector confirmation unit 110, and the class vector confirmation unit 110 outputs it using the output device 304 (a display device such as a display) (step S407).
[0048] FIG. 5 is a diagram showing an example of a screen display of a scatter plot of word distribution representations of proper nouns. The scatter plot 500 shown in FIG. 5 is a specific example of the scatter plot output in step S407 of FIG. 4. In the scatter plot 500, a group of distribution representations 501 of place names such as "Tokyo" and "Osaka" and a group of distribution representations 502 of restaurant names such as "○× Cafe" and "○△ Dining Hall" are shown. Also, the × marks 503 and 504 shown in FIG. 5 represent the distribution representations of class vector candidates generated from the distribution representations of the respective proper nouns of the place names and restaurant names (restaurants in the figure).
[0049] Here, if a distribution representation having a value significantly deviated from the majority of the distribution representations is included in the distribution representation group of each proper noun, the average of the distribution representation group may not be appropriate as the class vector. Therefore, in step S407, the class vector confirmation unit 110 may accept manual correction by the user using the input device 303 (such as a mouse or a keyboard) for the class vector candidates shown in the output scatter plot (for example, the × marks 503 and 504 in the scatter plot 500 of FIG. 5).
[0050] FIG. 6 is a diagram showing an example of a screen display during correction of a class vector candidate. The scatter diagram 600 shown in FIG. 6 is mostly the same as the scatter diagram 500 in FIG. 5, and a cross mark 601 is added. The cross mark 601 is the position of the class vector of the place name considered appropriate by the user. By the user performing a predetermined operation to select the cross mark 601 on the scatter diagram 600, it is determined that the class vector candidate of the place name is corrected from the cross mark 503 to the cross mark 601. Note that the appropriate position of the class vector candidate can be arbitrarily determined by the user. Specifically, for example, it may be determined based on the average of the vectors in the dispersion expression group of each proper noun, or it may be determined based on any one of the dispersion expressions, or it may be by other determination methods. When the correction of the class vector candidate is determined, the class vector confirmation unit 110 (or the class vector generation unit 113 that has received information from the class vector confirmation unit 110) newly creates a class vector candidate according to the above determination. Also, when the user determines that there is no need to correct the class vector candidate for each dispersion expression group, by performing a predetermined operation indicating confirmation, the confirmation of the class vector candidate displayed in the scatter diagram 500 (or scatter diagram 600) in step S407 is completed.
[0051] Then, according to the result of the correction or confirmation of the class vector candidate in step S407, the class vector generation unit 113 finally registers each confirmed class vector candidate as a class vector in the class vector storage unit 112 (step S408), and completes the class vector generation process. Note that in step S408, the class vector is registered in the class vector storage unit 112 as data paired with the proper noun type. As a result, the word dispersion expression correction unit 109 can obtain the class vector of the proper noun type by referring to the class vector storage unit 112 using the proper noun type such as "restaurant" or "place name" as a key.
[0052] (1-3) Dispersion Expression Correction Process FIG. 7 is a flowchart showing an example of the processing procedure of the distributed representation generation process. The distributed representation generation process shown in FIG. 7 is a process in which the document processing apparatus 100 shown in FIG. 1 converts the text included in the text data into a sentence distributed representation. The distributed representation generation process is executed, for example, when an input sentence (text data) in text is input to the input unit 101 during the operation of the document processing apparatus 100 after the class vector creation process shown in FIG. 4 is completed.
[0053] FIG. 8 is a flowchart showing an example of the processing procedure of the word distributed representation correction process. The word distributed representation correction process shown in FIG. 8 corresponds to the process of step S705 of the distributed representation generation process shown in FIG. 7 and is executed by the target word extraction unit 108 and the word distributed representation correction unit 109. Further, FIG. 9 is a diagram for explaining an example of the processing process of the input sentence by the distributed representation generation process.
[0054] Hereinafter, the distributed representation generation process in the present embodiment will be described in detail with reference to the specific example of FIG. 9 as appropriate along the flowcharts of FIGS. 7 and 8.
[0055] According to FIG. 7, first, the input unit 101 reads an input sentence 901 that is text data (step S701). FIG. 9 shows that the input sentence 901 is text "○×△ Café where".
[0056] Next, the morphological analysis unit 102 performs morphological analysis of the input sentence 901 input in step S701 (step S702). According to FIG. 9, the input sentence 901 is divided into three morphemes (words) 902 to 904, namely, "○×△ Café", "is", and "where" by morphological analysis. In addition, part-of-speech information 905 to 907 is assigned to each morpheme 902 to 904. Specifically, the morpheme 902 of "○×△ Café" is assigned part-of-speech information 905 of "proper noun - restaurant", the morpheme 903 of "is" is assigned part-of-speech information 906 of "particle", and the morpheme 904 of "where" is assigned part-of-speech information 907 of "pronoun".
[0057] Next, the morphological analysis unit 102 refers to the vocabulary list of the pre-trained neural language model 104 and divides the morphemes 902 to 904 into sub-words (step S703). In the case of FIG. 9, specifically, the proper noun "〇×△ Café" (morpheme 902) is sub-word divided into four tokens 908 to 911 of "〇", "×", "△", and "Café". Also, if the particle "wa" (morpheme 903) and the pronoun "doko" (morpheme 904) are pre-trained words registered in the vocabulary list of the pre-trained neural language model 104 and do not need to be divided more finely than the current situation, no sub-word division is performed and two tokens 912 and 913 are generated.
[0058] Note that the vocabulary list of the pre-trained neural language model 104 is learned from the appearance frequency of the learning corpus during the pre-training of the pre-trained neural language model 104. Therefore, in step S703, there may be cases where proper nouns are not divided into sub-words, and there may also be cases where common nouns are divided into sub-words. That is, if a proper noun appears frequently in pre-training, sub-word division is not required because sufficient accuracy has been obtained. On the other hand, if a common noun has a low appearance frequency in pre-training, sub-word division is performed because sufficient accuracy has not been obtained.
[0059] Next, the word distributed representation generation unit 105 converts each of the tokens (word sequences) 908 to 913 after the process of step S703 into word distributed representations 914 to 919 (step S704).
[0060] Next, the target word extraction unit 108 and the word distributed representation correction unit 109 execute a word distributed representation correction process (step S705). The word distributed representation correction process executed in step S705 corrects the distributed representation of the word (in this example, the proper noun word 902) that is split into sub-words and has no meaning in step S703 among the morphemes (words) 902 to 904 output in step S702, and generates a distributed representation by one word. As described above, a detailed example of the processing procedure is shown in FIG. 8. Hereinafter, the word distributed representation correction process will be described in detail with reference to FIG. 8.
[0061] According to FIG. 8, first, the target word extraction unit 108 focuses on the first word 902 among the words 902 to 904 output by the morphological analysis unit 102 in step S702 (step S801).
[0062] Next, the target word extraction unit 108 refers to the part-of-speech information 905 of the focused word 902 and determines whether it is a proper noun to be processed (step S802). The process of step S802 determines whether the focused word 902 is a word to be subject to the conversion process of the distributed representation of the word performed in steps S803 to S807 described later. In this case, the proper noun word 902 split into sub-words in step S703 of FIG. 7 corresponds to the word to be processed. In step S802, if it is a proper noun to be processed (YES in step S802), the process proceeds to step S803. If it is not a proper noun to be processed (NO in step S802), the process proceeds to step S808 described later.
[0063] Referring to FIG. 9, since the part-of-speech information 905 of "proper noun - restaurant" is a proper noun, it is determined as "YES" in step S802, and the process proceeds to step S803. In step S803, the word distributed representation correction unit 109 acquires the word distributed representations 914 to 917 of the words constituting the focused word 902 from the word distributed representation generation unit 105. As described above, the word distributed representations 914 to 917 are generated by the word distributed representation generation unit 105 in step S704 of FIG. 7.
[0064] Next, the word distribution representation correction unit 109 synthesizes position vectors for the word distribution representations 914 to 917 obtained in step S803 (step S804).
[0065] The process of step S804 will be described in detail. First, the word distribution representation correction unit 109 acquires from the position vector storage unit 111 the position vectors (position vectors 920 to 923 shown in FIG. 9) to be synthesized for each word of the word distribution representations 914 to 917. According to FIG. 9, the position vectors 920 to 923 are denoted as "P1", "P2", "P3", "P4" in order, but these represent the order of positions. That is, the position vector 920 "P1" is the vector indicating the first position, and the position vectors 921 to 923 are the vectors indicating the second, third, and fourth positions, respectively.
[0066] Furthermore, in step S804, the word distribution representation correction unit 109 adds the acquired position vectors 920 to 923 to the word distribution representations 914 to 917. Specifically, the position vector 920 of "P1" indicating the first position is synthesized into the word distribution representation 914 of the first word "○", the position vector 921 of "P2" indicating the second position is synthesized into the word distribution representation 915 of the second word "×", and the third and fourth words are synthesized in the same manner. The above synthesis process is a process necessary to generate a unique vector. By synthesizing the word distribution representations 914 to 917 and the position vectors 920 to 923, the word order of the sub-words in the word 902 can be maintained.
[0067] And after synthesizing the position vectors 920 to 923 into the word distribution representations 914 to 917 of all words in step S804, the word distribution representation correction unit 109 takes the arithmetic mean of these word distribution representations with the position vectors synthesized and synthesizes them into one distribution representation (step S805).
[0068] Next, the word distribution representation correction unit 109 synthesizes the one distribution representation synthesized in step S805 with the class vector (class vector 924 shown in FIG. 9) corresponding to the distribution representation (step S806).
[0069]
[0070]
[0069]
[0071]
[0072]
[0073] Next, the word distribution representation correction unit 109 (or the target word extraction unit 108) checks whether there is an unprocessed word (i.e., a word not focused on in step S801 or step S809) among the words 902 to 904 output by the morphological analysis unit 102 in step S702 of FIG. 7 (step S808). If there is an unprocessed word (YES in step S808), it focuses on one of the words (step S809) and repeats the processing after step S802. On the other hand, if there is no unprocessed word (NO in step S808), the word distribution representation correction process ends.
[0074] Specifically, for example, in step S808 after the processing of steps S802 to S807 for word 902, since words 903 and 904 are unprocessed (not focused on), in step S809, for example, it focuses on word 903 and the processing of step S802 is performed. However, since the part-of-speech information 906 of word 903 is "particle", it is determined in step S802 that it is not a proper noun to be processed (NO in step S802), and without any special conversion processing, it proceeds to step S808. As a result, in the example of FIG. 9, the word distribution representation 918 converted from word 903 in step S704 becomes the word distribution representation 926. In the next step S808, since word 904 is unprocessed, in step S809, it focuses on word 904 and the processing of step S802 is performed again. However, since word 904 is also not a proper noun to be processed, it is determined as NO in step S802 and proceeds to step S808 without any special conversion processing. As a result, in the example of FIG. 9, the word distribution representation 919 converted from word 904 in step S704 becomes the word distribution representation 927. In step S808 after the above processing, there is no unprocessed word (i.e., the processing of all words 902 to 904 is completed), so it is determined as NO in step S808 and the word distribution representation correction process ends.
[0075] Return to the description of FIG. 7. After the word distribution representation correction process in step S705 is performed, the sentence distribution representation generation unit 106 synthesizes the word distribution representations (for example, the word distribution representations 925 to 927 shown in FIG. 9) after being corrected by the word distribution representation correction unit 109 in step S705 to generate the sentence distribution representation of the entire input sentence 901 (step S706), and ends the distribution representation generation process.
[0076] As described above, according to the document processing apparatus 100 according to the present embodiment, for the distribution representation of low-frequency words (for example, proper nouns) for which sufficient semantic accuracy is not obtained in pre-training by the class vector creation process in FIG. 4, a class vector representing the type can be generated for each type of low-frequency word. Then, by the word distribution correction process in FIG. 8 performed in the distribution representation generation process in FIG. 7, for the low-frequency words included in the input sentence, after synthesizing the position vectors that fix the order of sub-words to the distribution representation of the low-frequency words and then synthesizing with the above class vectors, sufficient meaning can be imparted to the low-frequency words to correct (replace) the word distribution representation of the low-frequency words. Thus, the document processing apparatus 100 can generate a unique sentence distribution representation with the meaning of low-frequency words appropriately embedded without performing additional learning on a sentence including low-frequency words (for example, proper nouns) for which sufficient accuracy is not obtained in pre-training, by executing the distribution representation generation process in FIG. 7 after the class vector creation process in FIG. 4.
[0077] (2) Second Embodiment In the second embodiment, an example of a document processing apparatus that generates a reference vector (class vector) representing a symbol for the distribution representation of an emoji symbol, and synthesizes the distribution representation of the explanatory text of the emoji and the class vector to generate a distribution representation of the symbol with meaning embedded will be described. The emojis handled in this embodiment are under-learned words that cannot sufficiently represent meaning in the pre-trained distribution representation space, and correspond to an example of unknown words that have not been pre-trained.
[0078] (2-1) Configuration FIG. 10 is a block diagram showing a functional configuration example of a document processing apparatus 1000 according to a second embodiment of the present invention. The document processing apparatus 1000 is configured by adding an extended word vector generation unit 1003 and an extended word vector storage unit 1004 to the configuration of the document processing apparatus 100 according to the first embodiment. Note that the target word extraction unit 1001 and the word distribution expression correction unit 1002 in the document processing apparatus 1000 are processing units corresponding to the target word extraction unit 108 and the word distribution expression correction unit 109 in the document processing apparatus 100, but they perform slightly different processes, so new reference numerals are assigned.
[0079] Similar to the document processing apparatus 100 according to the first embodiment, the document processing apparatus 1000 can be included in the document processing system 200 shown in FIG. 2. Also, similar to the document processing apparatus 100 according to the first embodiment, the document processing apparatus 1000 is realized by an information processing apparatus 300 having the hardware configuration example shown in FIG. 3. Specifically, the extended word vector generation unit 1003 is realized by the processor 301 reading and executing various predetermined programs, and the extended word vector storage unit 1004 is realized by the storage device 302.
[0080] The extended word vector generation unit 1003 obtains the sentence distribution expression of the description sentence of the word (unknown word, which represents an emoji in this example) that is the basis of the extended word vector via the sentence distribution expression generation unit 106, generates a distribution expression (extended word vector) with the meaning as a symbol, and registers it in the extended word vector storage unit 1004. The extended word vector storage unit 1004 stores the extended word vector generated by the extended word vector generation unit 1003. The generation of the extended word vector will be described in detail in the extended word vector creation process shown in FIG. 11 described later.
[0081] FIG. 14 is a diagram showing an example of the explanatory text of an emoji. In FIG. 14, a plurality of emojis 1401 and the explanatory text 1402 for each emoji 1401 are shown in association with each other. When the emoji 1401 is a symbol of an emoji defined in UNICODE, the explanatory text 1402 describes the UNICODE explanatory text of the corresponding symbol. It is assumed that the symbols and explanatory texts of emojis defined in UNICODE are registered in the system dictionary of the morphological analysis dictionary 103. However, the description of the explanatory text 1402 is not limited to this, and it may be created independently by the user. Also, symbols other than those defined in UNICODE can be made into targets for generating extended word vectors by being registered in the user dictionary of the morphological analysis dictionary 103 or the like.
[0082] As described above, in the present embodiment, the symbol of the emoji is used as an example of an extended word (unknown word), and the extended word vector generation unit 1003 generates the extended word vectors of the respective emojis 1401 by executing the extended word vector creation process in the preprocessing, and registers them in the extended word vector storage unit 1004.
[0083] (2-2) Extended Word Vector Creation Process FIG. 11 is a flowchart showing an example of the processing procedure of the extended word vector creation process. The extended word vector creation process shown in FIG. 11 is a process executed at an arbitrary timing before the distributed representation generation process is executed, similar to the class vector creation process shown in FIG. 4 in the first embodiment, and is a process in which the document processing device 1000 creates an extended word vector and registers it in the extended word vector storage unit 1004.
[0084] According to FIG. 11, first, the class vector generation unit 113 generates a class vector representing a plurality of symbols (emojis 1401 in FIG. 14) (step S1101). The procedure for generating the class vector in step S1101 is almost the same as the procedure for the class vector creation process shown in FIG. 4, and in step S403 of FIG. 4, symbols are extracted instead of proper nouns. The class vector of the symbols generated in step S1101 is registered in the class vector storage unit 112.
[0085] Next, the class vector generation unit 113 calculates the Euclidean distance between the class vector of the symbol generated in step S1101 and the dispersion representation of each symbol, and obtains the maximum value of the Euclidean distance (step S1102).
[0086] Next, the extended word vector generation unit 1003 converts all the explanatory texts 1402 of the emoji into sentence dispersion representations via the input unit 101, the morphological analysis unit 102, the word dispersion representation generation unit 105, and the sentence dispersion representation generation unit 106 (step S1103). In step S1103, for all the explanatory texts 1402 of the target emojis 1401, conversion into sentence dispersion representations is performed. The generation procedure of the sentence dispersion representation in step S1103 is the same as the dispersion representation generation process shown in FIG. 7 (however, the correction of the word dispersion representation in step S705 is not performed).
[0087] Next, the extended word vector generation unit 1003 calculates the average of the sentence dispersion representations of all the explanatory texts created in step S1103 (step S1104). Then, the Euclidean distance between the average of the sentence dispersion representations of all the calculated explanatory texts and the dispersion representation of each explanatory text is calculated, and the maximum value of the Euclidean distance is obtained (step S1105).
[0088] Next, the extended word vector generation unit 1003 uses the maximum value of the Euclidean distance of the symbol dispersion representation calculated in step S1102 and the maximum value of the Euclidean distance of the sentence dispersion representation of the explanatory text calculated in step S1105 to obtain the ratio of the latter maximum value to the former maximum value (step S1106).
[0089] Next, the extended word vector generation unit 1003 subtracts the average of the dispersion representations of all the explanatory texts from the sentence dispersion representation of the explanatory text, and integrates the ratio of the Euclidean distance obtained in step S1106 to generate the dispersion representation (extended word vector) of the emoji corresponding to each explanatory text (step S1107).
[0090] Finally, the extended word vector generation unit 1003 registers the extended word vector generated in step S1107 in the extended word vector storage unit 1004 (step S1108), and ends the extended word vector creation process.
[0091] By executing the extended word vector creation process as described above, the extended word vector generation unit 1003 can generate a unique distributed representation (extended word vector) for each description (in other words, for each emoji corresponding to the description) while including the nature as a symbol.
[0092] (2-3) Word Distributed Representation Generation Process The distributed representation generation process executed in the second embodiment will be described. The distributed representation generation process in this embodiment is a process in which the document processing device 1000 shown in FIG. 10 converts the text included in the text data into a document distributed representation. The distributed representation generation process in this embodiment is executed, for example, when an input sentence (text data) in text is input to the input unit 101 during the operation of the document processing device 1000 after the class vector creation process shown in FIG. 4 and the extended word vector creation process shown in FIG. 11 are completed. Note that most of the distributed representation generation process in the second embodiment is executed in the same processing procedure as the distributed representation generation process shown in FIG. 7 in the first embodiment, and the word distributed representation correction process in step S705 is executed in a processing procedure specific to the second embodiment.
[0093] FIG. 12 is a flowchart showing an example of the processing procedure of the word distributed representation correction process in the second embodiment. The word distributed representation correction process shown in FIG. 12 corresponds to the process of step S705 of the distributed representation generation process shown in FIG. 7 in the second embodiment, and is executed by the target word extraction unit 1001 and the word distributed representation correction unit 1002. FIG. 13 is a diagram for explaining an example of the processing process of an input sentence by the distributed representation generation process in the second embodiment.
[0094] Hereinafter, the distributed representation generation process in this embodiment will be described in detail with reference to the specific example in FIG. 13 as appropriate according to the flowcharts in FIGS. 7 and 12. Note that the description of the process in FIG. 7 that is common to the first embodiment may be omitted.
[0095] In the distributed representation generation process, first, the input unit 101 reads the input sentence 1301 which is text data (step S701). In FIG. 13, an input sentence 1301 of text including an emoji 1302 is illustrated. Note that the emoji 1302 is an emoji with an explanatory text 1402 of "smiling face" in FIG. 14.
[0096] Next, the morphological analysis unit 102 performs morphological analysis on the input sentence 1301 input in step S701 (step S702). According to FIG. 13, the input sentence 1301 is divided into two morphemes (words) 1303 and 1304 by morphological analysis. Also, part-of-speech information 1305 and 1306 is assigned to each morpheme 1303 and 1304. Specifically, part-of-speech information 1305 of "interjection" is assigned to the morpheme 1303, and part-of-speech information 1306 of "symbol-emoji" is assigned to the morpheme 1304.
[0097] Next, the morphological analysis unit 102 refers to the vocabulary list of the pre-trained neural language model 104 and divides the morphemes 1303 and 1304 into sub-words (step S703). Specifically, since the morpheme 1303 which is an interjection exists in the vocabulary list, it is not divided into sub-words and is converted into a token 1307 of "thank you". On the other hand, since the morpheme 1304 which is an emoji symbol does not exist in the vocabulary list, it is converted into a special token "UNK" (token 1308) representing an unknown word in the pre-trained neural language model 104.
[0098] Next, the word distributed representation generation unit 105 converts the tokens (word sequences) 1307 and 1308 after the processing in step S703 into word distributed representations 1309 and 1310 respectively (step S704).
[0099] Next, the word distributed representation correction process shown in FIG. 12 is executed (step S705). The word distributed representation correction process executed in step S705 of the present embodiment corrects the distributed representation of a word (in this example, the emoji word 1304) that is divided into sub-words that do not make sense among the morphemes (words) 1303 and 1304 output in step S702, and generates a distributed representation by one word. As described above, a detailed processing procedure example is shown in FIG. 12. Hereinafter, the word distributed representation correction process will be described in detail with reference to FIG. 12.
[0100] According to FIG. 12, first, the target word extraction unit 1001 focuses on the first word 1303 among the words 1303 and 1304 output in step S702 (step 1201). Then, based on the part-of-speech information 1305 of the focused word 1303, the target word extraction unit 1001 refers to the extended word vector storage unit 1004 and determines whether it is a word to be processed (step S1202). In step S1202, when the extended word vector of the focused word 1303 is registered in the extended word vector storage unit 1004, it can be determined that it is a word to be processed. In step S1202, if it is an emoji to be processed (YES in step S1202), the process proceeds to step S1203, and if it is not an emoji to be processed (NO in step S1202), the process proceeds to step S1207.
[0101] Specifically, in the case of the word 1303, since the part-of-speech information 1305 is "interjection", it is determined in step S1202 that it is not a word to be processed (NO in step S1202), and the process proceeds to step S1207.
[0102] In step S1207, the word distribution representation correction unit 1002 (or the target word extraction unit 1001) checks whether there are any unprocessed words among the words 1303 and 1304 output in step S702 (that is, words not focused on in step S1201 or step S1208). If there are unprocessed words (YES in step S1207), one of them is focused on (step S1208), and the processing from step S1202 onwards is repeated. On the other hand, if there are no unprocessed words (NO in step S1207), the word distribution representation correction process ends.
[0103] Specifically, when word 1303 is being focused on, since word 1304 is unprocessed, it is determined as "NO" in step S1207, and word 1304 is focused on in step S1208, then the process proceeds to step S1202.
[0104] And in step S1202 when word 1304 is focused on, the part-of-speech information 1306 is "symbol-emoji", and since its extended word vector is registered in the extended word vector storage unit 1004, it is determined as the processing target (YES in step S1202), and the process proceeds to step S1203.
[0105] In step S1203, the word distribution representation correction unit 1002 refers to the class vector storage unit 112 to obtain the class vector 1312 of the symbol of the word 1304 focused on in step S1202. Also, the word distribution representation correction unit 1002 refers to the extended word vector storage unit 1004 to obtain the extended word vector 1311 corresponding to the word 1304 of the emoji (symbol) (step S1204).
[0106] Next, the word distribution representation correction unit 1002 adds the extended word vector 1311 obtained in step S1204 and the class vector 1312 obtained in step S1203 to generate the word distribution representation 1314 of the word 1304 of the emoji (step S1205).
[0107] Next, the word distribution representation correction unit 1002 replaces the word distribution representation 1310 of the word distribution representation generation unit 105 with the newly generated word distribution representation 1314 in step S1205 (step S1206).
[0108] When the process of step S1206 ends, the process proceeds to step S1207. In this case, since the processes after step S1202 have been completed for all words 1303, 1304, it is determined as "NO", and the word distribution representation correction process ends.
[0109] Return to the description of FIG. 7. After the word distribution representation correction process of step S705 is performed, the sentence distribution representation generation unit 106 synthesizes the word distribution representations (for example, the word distribution representations 1313, 1314 shown in FIG. 13) after being corrected by the word distribution representation correction unit 1002 in step S705 to generate the sentence distribution representation of the entire input sentence 1301 (step S706), and the distribution representation generation process ends.
[0110] As described above, according to the document processing apparatus 1000 according to the present embodiment, in the extended word vector creation process of FIG. 11, for the distributed representation of unknown words (e.g., emoji symbols) for which sufficient semantic accuracy has not been obtained in pre-training, a class vector representing a plurality of unknown words is generated, and a unique distributed representation (extended word vector) for each unknown word is generated using the descriptions of the plurality of unknown words. Then, in the word distribution correction process of FIG. 12 performed in the distributed representation generation process of FIG. 7, for the unknown words included in the input sentence, by synthesizing the distributed representation (extended word vector) of the unknown word and the class vector, sufficient meaning can be imparted to the unknown word to correct (replace) the word distributed representation of the unknown word. Thus, the document processing apparatus 1000 can generate a unique sentence distributed representation with the meaning of the unknown word appropriately embedded therein without performing additional learning on a sentence including unknown words (e.g., emoji symbols) for which sufficient accuracy has not been obtained in pre-training, by executing the distributed representation generation process of FIG. 7 after the class vector creation process of FIG. 11. In the above description, emoji symbols are used as an example, but the document processing apparatus 1000 according to the present embodiment can perform the same processing on unknown words such as kaomoji composed of symbols as long as there is a description corresponding to the symbol as shown in FIG. 14. In this case, the kaomoji shall be registered in the morphological analysis dictionary 103.
[0111] Note that the present invention is not limited to the above-described embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Further, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations is possible.
[0112] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be implemented in hardware by designing part or all of them, for example, using an integrated circuit. Further, each of the above-described configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that realizes each function. Information such as programs, tables, files, etc. that realize each function can be stored in a memory, a recording device such as a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC card, an SD card, or a DVD.
[0113] Also, in the drawings, control lines and information lines show those considered necessary for explanation, and not necessarily all control lines and information lines are shown on the product. In reality, it may be considered that almost all components are interconnected.
Explanation of Signs
[0114] 100, 1000 Document processing device 101 Input unit 102 Morphological analysis unit 103 Morphological analysis dictionary 104 Pre-trained neural language model 105 Word distribution representation generation unit 106 Sentence distribution representation generation unit 107 Output unit 108, 1001 Target word extraction unit 109, 1002 Word distribution representation correction unit 110 Class vector confirmation unit 111 Position vector storage unit 112 Class vector storage unit 113 Class vector generation unit 200 Document processing system 201 Server 202 Network 203 User terminal 300 Information processing device 301 Processor 302 Memory device 303 Input device 304 Output device 305 Communication interface 306 Bus 1003 Extended word vector generation unit 1004 Extended word vector storage unit
Claims
1. A document processing apparatus that generates a distributed representation of words included in text input from the outside, comprising: a class vector storage unit that stores class vectors, which are reference vectors for each type of word; a morphological analysis unit that divides the text into a word sequence by morphological analysis and assigns information indicating the type to which each word sequence belongs; a word distributed representation generation unit that converts the word sequence divided by the morphological analysis unit into a distributed representation for each word included in the word sequence; a word distributed representation correction unit that corrects the distributed representation generated by the word distributed representation generation unit; and when the text includes low-frequency words whose appearance frequency in pre-learning is lower than a predetermined level, after the morphological analysis unit divides the text into a word sequence by morphological analysis, the morphological analysis unit further divides the low-frequency words included in the word sequence into sub-words of partial character strings, the word distributed representation generation unit converts the low-frequency words into a plurality of distributed representations each composed of a distributed representation for each sub-word included in the low-frequency words, the word distributed representation correction unit synthesizes a position vector that fixes the order of the sub-words in the low-frequency words with the plurality of distributed representations of the low-frequency words generated by the word distributed representation generation unit, and then synthesizes the class vector of the type to which the low-frequency words belong, and replaces the synthesized distributed representation with the distributed representation of the low-frequency words. A document processing apparatus characterized by the above.
2. The document processing apparatus according to claim 1, further comprising a class vector generation unit that generates class vectors for each type based on the distributed representations of a plurality of the low-frequency words belonging to the same type and registers the generated class vectors in the class vector storage unit.
3. The class vector generation unit generates, as the class vector of the type to which the low-frequency words belong, the distributed representation generated by the word distributed representation generation unit for one word among the plurality of low-frequency words belonging to the same type, or the average of the distributed representations respectively generated by the word distributed representation generation unit for the plurality of low-frequency words belonging to the same type. The document processing apparatus according to claim 2, characterized by the above.
4. The word distributed representation correction unit... The document processing apparatus according to claim 3, characterized by the above.
5. The word distributed representation correction unit... After integrating a predetermined coefficient for adjusting the distributed representation of the low-frequency word generated by the word distributed representation generation unit so that the synthesized distributed representation is in the vicinity of the class vector, the distributed representation is synthesized with the class vector of the type to which the low-frequency word belongs, and the synthesized distributed representation is replaced with the distributed representation of the low-frequency word. The document processing apparatus according to claim 3, characterized in that.
5. Output the class vector of the type to which the low-frequency word belongs, which is generated by the class vector generation unit, as a class vector candidate and request confirmation by the user. If a correction request is made by the user, further include a class vector confirmation unit that corrects the class vector candidate according to the request. The class vector generation unit registers the class vector candidate corrected or confirmed by the class vector confirmation unit in the class vector storage unit as the class vector of the type to which the low-frequency word belongs. The document processing apparatus according to claim 2, characterized in that.
6. A sentence distributed representation generation unit that generates a distributed representation of the entire text by synthesizing the distributed representations generated by the word distributed representation generation unit or the word distributed representation correction unit for each word included in the word sequence constituting the text. An output unit that outputs the distributed representation of the entire text generated by the sentence distributed representation generation unit. The document processing apparatus according to claim 1, characterized in that.
7. For each type to which the low-frequency word belongs, generate class vectors of each type based on the distributed representations of a plurality of low-frequency words belonging to the same type, and register the generated class vectors in the class vector storage unit. Further include a class vector generation unit. Explanations for a plurality of unknown words that have not been pre-learned are retained. The class vector generation unit generates class vectors of each type based on the distributed representations of the unknown words belonging to the same type for each type to which the unknown words belong, and registers the generated class vectors in the class vector storage unit. The sentence distributed representation generation unit converts a plurality of explanations for the plurality of unknown words into distributed representations. An extended word vector generation unit that generates a distributed representation of each unknown word based on a comparison between the distributed representations of the plurality of explanations and the distributed representations of the explanations for each unknown word. An extended word vector storage unit that stores the distributed representations of each unknown word generated by the extended word vector generation unit. further comprising when any of the unknown words is included in the text, after the morphological analysis unit divides the text into a word sequence, the morphological analysis unit converts the unknown word included in the word sequence into a predetermined token, the word distributed representation generation unit generates a distributed representation of the predetermined token, the word distributed representation correction unit synthesizes the distributed representation of the unknown word stored in the extended word vector storage unit and the class vector of the type to which the unknown word belongs stored in the class vector storage unit, and sets the synthesized distributed representation as the distributed representation of the unknown word The document processing apparatus according to claim 6, characterized in that.
8. A document control method by a document processing apparatus that generates a distributed representation of words included in a text input from the outside, the document processing apparatus stores a class vector that is a reference vector for each type of word, a morphological analysis step in which the document processing apparatus divides the text into a word sequence by morphological analysis and assigns information indicating the type to which each word sequence belongs; a word distributed representation generation step in which the document processing apparatus converts the word sequence divided in the morphological analysis step into a distributed representation for each word included in the word sequence; a word distributed representation correction step in which the document processing apparatus corrects the distributed representation generated in the word distributed representation generation step; comprising when the text includes a low-frequency word whose appearance frequency in pre-learning is lower than a predetermined level, in the morphological analysis step, after dividing the text into a word sequence by morphological analysis, further dividing the low-frequency word included in the word sequence into sub-words of partial character strings, in the word distributed representation generation step, converting the low-frequency word into a plurality of distributed representations each consisting of distributed representations for each sub-word included in the low-frequency word, in the word distributed representation correction step, after synthesizing a position vector that fixes the order of the sub-words in the low-frequency word with the plurality of distributed representations of the low-frequency word generated in the word distributed representation generation step, synthesizing with the class vector of the type to which the low-frequency word belongs, and replacing the synthesized distributed representation with the distributed representation of the low-frequency word A document processing method characterized by the above.
9. A document control program for causing a computer constituting a document processing apparatus that generates a distributed representation of words included in a text input from the outside to execute, The document processing apparatus stores class vectors, which are reference vectors for each word type. A morphological analysis process that divides the text into a word sequence by morphological analysis and assigns information indicating the type to which each word sequence belongs. A word dispersion representation generation process that converts the word sequence divided by the morphological analysis process into a dispersion representation for each word included in the word sequence. A word dispersion representation correction process that corrects the dispersion representation generated by the word dispersion representation generation process. The computer is caused to execute When the text includes low-frequency words whose appearance frequency in pre-training is lower than a predetermined level, In the morphological analysis process, after dividing the text into a word sequence by morphological analysis, the low-frequency words included in the word sequence are further divided into sub-words of partial character strings. In the word dispersion representation generation process, the low-frequency words are converted into a plurality of dispersion representations each consisting of a dispersion representation for each sub-word included in the low-frequency words. In the word dispersion representation correction process, a position vector that fixes the order of the sub-words in the low-frequency words is synthesized with the plurality of dispersion representations of the low-frequency words generated by the word dispersion representation generation process, and then synthesized with the class vector of the type to which the low-frequency words belong, and the synthesized dispersion representation is replaced as the dispersion representation of the low-frequency words. A document processing program characterized by the above.
Citation Information
Patent Citations
Natural language processing device and natural language processing program
JP2021009538A