Information identification device, information identification method, and program

The information identification device splits and classifies input information using distributed representations and SVMs to automatically identify same-type information across varying formats, enhancing interoperability in business collaborations.

JP7710144B2Active Publication Date: 2025-07-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023568771
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-07-18
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

Existing systems fail to automatically identify information in different formats as being of the same type, necessitating manual examination and format conversion in information circulation among collaborating business operators.

Method used

An information identification device that splits input information into words, generates distributed representations for each word, combines these representations, and uses machine learning to classify the information into types, employing a support-vector machine (SVM) to determine confidence levels and handle new types through re-learning.

Benefits of technology

Enables automatic identification of information in different formats as being of the same type, improving classification accuracy and facilitating seamless information flow among diverse systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710144000006
    Figure 0007710144000006
  • Figure 0007710144000007
    Figure 0007710144000007
  • Figure 0007710144000008
    Figure 0007710144000008
Patent Text Reader

Abstract

An identification device 1 is equipped with: a character string division unit 11 for dividing information for training into words; a distributed representation generation unit 12 for generating a distributed representation for each word; a distributed representation-linking unit 13 for generating a distributed representation combination by combining the distributed representations of each word; a training data generation unit 14 for generating training data which includes the distributed representation combination and the group number of the group which corresponds to the information for training; and a learning unit for generating a classifier for identifying to which of the groups input information belongs by subjecting the training data to machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information identification device, an information identification method, and a program.

Background Art

[0002] B2B2X services in which a plurality of business operators in different industries cooperate through B2B2X are increasing. In providing such services, it is necessary to circulate information such as customer information, contract information, and invoice information among the business operators who cooperate.

[0003] In the circulation of information, identification of information is necessary, and as technologies related to the identification of information, there are morphological analysis (Non-Patent Document 1), distributed representation of characters (Non-Patent Document 2), and classifier confidence calculation technology (Non-Patent Document 3).

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] Currently, in order to circulate information among collaborating operators, an operator examines the content of the input information to determine its type, and manually determines the distribution destination and performs format conversion of the information.

[0006] In the future, in order to cope with the increasing number of services that involve collaboration among multiple companies in different industries, it is necessary to automate the information flow among collaborating operators. In automating the information flow, when information in different formats is input, it is required to identify that they are of the same type. Specifically, in order to distribute information to appropriate destinations for each type, even when existing systems used by each company are different and information of the same type is input in different formats, it is necessary to automatically identify that they are of the same type.

[0007] Non-Patent Documents 1-3 do not consider identifying that information in different formats is of the same type. Therefore, using these non-patent documents, it is not possible to identify that information in different formats is of the same type.

[0008] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technique capable of identifying that information in different formats is of the same type when the information is input.

Means for Solving the Problems

[0009] To achieve the above object, one aspect of the present invention includes a splitting unit that splits learning information into words, a distributed representation generation unit that generates a distributed representation for each word, a distributed representation concatenation unit that combines the distributed representations of each word to generate a combination of distributed representations, a learning data generation unit that generates learning data including the combination of distributed representations and a class number of a class corresponding to the learning information, and a learning unit that performs machine learning on the learning data to generate a classifier that identifies input information into any one of the classes.

[0010] One aspect of the present invention is an information identification method performed by an information identification device, the method including steps of splitting learning information into words, generating a distributed representation for each word, combining the distributed representations of each word to generate a combination of distributed representations, generating learning data including the combination of distributed representations and a class number of a class corresponding to the learning information, and performing machine learning on the learning data to generate a classifier that identifies input information into any one of the classes.

[0011] One aspect of the present invention is a program that causes a computer to function as the above information identification device.

Effects of the Invention

[0012] According to the present invention, it is possible to provide a technique that can identify that different forms of information are of the same class when the information is input.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

Figure 13

Figure 14A

Figure 14B

Figure 15

Figure 16

Mode for Carrying Out the Invention

[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0015] <Configuration of Information Identification Device> Fig. 1 shows a configuration example of the information identification device 1 of the present embodiment. The information identification device 1 is a device that identifies the type of input information in information circulation among multiple companies in different industries. The illustrated information identification device 1 includes a character string splitting unit 11, a distributed representation generation unit 12, a distributed representation concatenation unit 13, a learning data generation unit 14, a type classification unit 15, a type determination unit 16, a similar word extraction unit 17, a learning data update unit 18, a type holding unit 19, and a regular expression determination unit 20.

[0016] The character string splitting unit 11 (splitting unit) splits the input learning information into words. Specifically, the character string splitting unit 11 performs morphological analysis on the character string of the information and splits it into the smallest unit words that have meaning by themselves (see Non-Patent Document 1). The information of the present embodiment includes a key and a value.

[0017] The distributed representation generation unit 12 generates a distributed representation of each word split by the character string splitting unit 11. The distributed representation is one of natural language processing and is a technique for representing a word as a high-dimensional real vector (see Non-Patent Document 2). By mathematically expressing the meaning of a word, arithmetic processing using the meaning of the word becomes possible.

[0018] The distributed representation concatenation unit 13 combines the distributed representations of each word to generate a combination of distributed representations. In the present embodiment, the distributed representation concatenation unit 13 calculates the sum of the distributed representations of each word as a combination of distributed representations.

[0019] The learning data generation unit 14 generates learning data including a combination of distributed representations and a type number of the type corresponding to the learning information.

[0020] The type classification unit 15 includes a learning unit and a classifier. The learning unit performs machine learning on the learning data to generate a classifier. The classifier is a learned model for identifying the input information into any type.

[0021] The classifier may use an SVM (support-vector machine) (see Non-Patent Document 3). The SVM is a method that focuses on the confidence level of how confident a multi-class classifier is in its recognition and aims for more accurate estimation. When classifying into K classes, the SVM uses the concept of One vs All SVM to create K decision boundaries and determines the classification type using the confidence level calculated from the distance between the object to be classified and each decision boundary. As the method for calculating the confidence level of the disperser in the present embodiment, the distance from the decision boundary of the SVM is utilized. The farther away from the decision boundary in the positive direction when viewed from the decision boundary, the greater the confidence level. The calculation of the confidence level will be described later.

[0022] The type determination unit 16 (determination unit) determines whether the type of the input information is any of the existing types or a new type using the confidence level of the highest value (maximum value) output by the classifier. Specifically, when the confidence level of the highest value is positive, the type determination unit 16 determines that the type of the input information is the type of the highest value, and when the confidence level of the highest value is negative, the type determination unit 16 determines that the type of the input information is a new type. When the type determination unit 16 determines that the type of the input information is a new type, the type determination unit 16 generates a type number of the new type (new type number) and stores the type number together with the new type in the type holding unit 19.

[0023] When the type determination unit 16 determines that the type is a new type, the similar word extraction unit 17 extracts the similar words of the key of the input information (first similar words) and the similar words of the value of the input information (second similar words), respectively. Then, the distributed representation generation unit 12 generates the distributed representation of the similar words of the key and the distributed representation of the similar words of the value.

[0024] The distributed representation concatenation unit 13 combines the distributed representation of the similar words of the key and the distributed representation of the similar words of the value to generate a combined additional distributed representation. Also, the distributed representation concatenation unit 13 may combine the distributed representation of the key and the distributed representation of the similar words of the value to generate a combined additional distributed representation. Further, the distributed representation concatenation unit 13 may combine the distributed representation of the similar words of the key and the distributed representation of the value to generate a combined additional distributed representation.

[0025] The learning data update unit 18 generates learning data (additional learning data) including a new type number. Specifically, the learning data update unit 18 generates learning data including a combination of additional distributed representations and a new type number. The learning unit of the type determination unit 16 retrains the classifier using the additional learning data generated by the learning data update unit 18.

[0026] In the type holding unit 19, a type and a type number are stored in association with each other. When the value of the input information cannot be converted into a distributed representation (in the case of a regular expression), the regular expression determination unit 20 determines the type of the input information.

[0027] The Middle B system 3 is a system of a service provider. The information identification device 1 of the present embodiment is a device operated by Middle B. The operator terminal 7 is a terminal used by an operator of Middle B. The First B system 5 is a system of a cooperating business operator related to the service provided by Middle B. In the example shown in FIG. 1, the information identification device 1 identifies the type of information input from the Middle B system 3 and outputs the identification result to the First B system 5.

[0028] <Learning phase> FIG. 2 is an explanatory diagram for explaining the operation of the information identification device 1 in the learning phase. The operator terminal 7 transmits learning information (character string) input by the operator to the character string splitting unit 11 of the information identification device 1 (step S11). The character string splitting unit 11 splits the input information into words by morphological analysis (step S12).

[0029] In the illustrated example, the character string splitting unit 11 splits "Name: Yamada Taro" into "Name", "Yamada", and "Taro" and outputs them to the distributed representation generation unit 12. "Name" is the key and "Yamada Taro" is the value.

[0030] The distributed representation generation unit 12 generates the distributed representation of each word segmented by the string segmentation unit 11 and outputs it to the distributed representation concatenation unit 13 (step S13). That is, the distributed representation generation unit 12 converts each word into a high-dimensional real vector.

[0031] The distributed representation concatenation unit 13 combines the distributed representations of each word to generate a combination of distributed representations and outputs it to the learning data generation unit 14 (step S14). By using the combination of distributed representations, it is possible to prevent multiple categories from being assigned to a word with multiple meanings and improve the classification accuracy. In the present embodiment, the distributed representation concatenation unit 13 calculates the sum of the distributed representations of each word as the combination of distributed representations.

[0032] The learning data generation unit 14 receives the type number of the learning information input in step S11 from the operator terminal 7 (step S15), generates learning data including the combination of distributed representations and the type number, and outputs it to the type classification unit 15 (step S16). The illustrated learning data includes the sum of the distributed representations of "name", "Yamada", and "Taro" and the type number: 0. The learning data is data for training the classifier of the type classification unit 15. The learning unit of the type classification unit 15 generates a classifier by machine learning using the learning data.

[0033] As described above, in the learning phase of the present embodiment, the information identification device 1 divides the information including key and value into words, calculates the distributed representation of each word, generates a combination of the distributed representations of each word, and generates learning data including the combination of distributed representations and the type number. Thereby, in the identification phase described later, even if a word with multiple meanings or a synonym is input, it can be identified as an appropriate type.

[0034] Figure 3 shows the distributed representations (vectors) of each word ("surname", "name", "address", "Yamaguchi") generated by the distributed representation generation unit 12, and the sum of the distributed representations generated by the distributed representation concatenation unit 13 ("surname: Yamaguchi", "name: Yamaguchi", "address: Yamaguchi"). As shown in the figure, it shows that the sums of distributed representations of the same type are mapped (i.e., clustered) to nearby positions.

[0035] As can be seen from Figure 3, by using the sum of distributed representations as a combination of distributed representations, information of the same type is mapped to nearby positions. Therefore, by using the sum of the distributed representation of the Key and the distributed representation of the Value to generate learning data in multi-class classification, it becomes possible to identify the meaning of words with multiple meanings and to identify that different words are of the same type.

[0036] Figure 4 is a sequence diagram showing the operation of the information identification device 1 in the learning phase.

[0037] The operator terminal 7 receives an operator's instruction and inputs learning information to the string splitting unit 11 of the information identification device 1 (step S21). Here, a plurality of pieces of information of the type "name" are input. The string splitting unit 11 splits the information into words (step S22). The distributed representation generation unit 12 generates distributed representations of each word split for each piece of information and outputs them to the distributed representation concatenation unit 13 (step S23).

[0038] The distributed representation concatenation unit 13 combines the distributed representations of each word for each piece of information to generate a combination of distributed representations and outputs it to the learning data generation unit 14 (step S24). Here, the sum of the distributed representations of each word is used as the combination of distributed representations.

[0039] The operator terminal 7 receives an operator's instruction and requests the type number of the information transmitted in S21 from the type holding unit 19 (step S25). Here, the operator terminal 7 requests the type number of "name". The type holding unit 19 transmits the type number corresponding to "name" to the operator terminal 7 (step S26). When the operator terminal 7 acquires the type number, it transmits the type number to the learning data generation unit 14 (step S27).

[0040] The learning data generation unit 14 generates learning data including the sum of the distributed representations received in step S24 and the type number received in S27, and sends it to the type classification unit 15 (step S28). The learning unit of the type classification unit 15 generates a classifier by machine learning the learning data (step S29).

[0041] In addition, in FIG. 4, the operator terminal 7 requests the type number from the type holding unit 19 in step S25, but the learning data generation unit 14 may request the type number from the type holding unit 19. In this case, in step S26, the type holding unit 19 sends the type number to the learning data generation unit 14, and the learning data generation unit 14 generates learning data using the sum of the distributed representations and the type number obtained from the type holding unit 19 (step S28).

[0042] Next, the operation of registering information that cannot be represented in a distributed manner, such as a postal code or a telephone number, in the type holding unit 19 will be described (steps S30, S31). For information such as a postal code or a telephone number that cannot be represented in a distributed manner, in the identification phase described later, the regular expression determination unit 20 of the information identification device 1 determines what type it is from the pattern of numbers or character strings using a regular expression. Note that steps S30 and S31 are performed asynchronously with steps S21 to S28.

[0043] The operator terminal 7 receives the pattern of the regular expression corresponding to the type from the partner of the First B System 5 (step S30), and transmits the type and the pattern of the regular expression to the type holding unit 21 of the information identification device 1 (step S31). The type holding unit 19 stores the transmitted type and the pattern of the regular expression in its own storage unit.

[0044] In FIG. 4, an example of a regular expression pattern of a postal code type and an example of a regular expression pattern of a telephone number type are illustrated. In the regular expression of the postal code, it starts with the "〒" mark and represents that it ends with three digits, a "-" (hyphen), and four digits.

[0045] <Identification phase> In the identification phase, the information identification device 1 determines whether the type of the input information can be identified as any of the types held in the type holding unit 19 by using the confidence level calculated by the classifier of the type classification unit 15. In the present embodiment, the information identification device 1 determines that identification is possible when the highest value of the confidence level output by the classifier is positive, and determines that identification is impossible and that it is a new type not held in the type holding unit 19 when the highest value of the confidence level is negative.

[0046] FIGS. 5 and 6 are explanatory diagrams for explaining the outline of the operation of the information identification device 1 in the identification phase. FIG. 5 is an explanatory diagram when the type can be identified, and FIG. 6 is an explanatory diagram when the type cannot be identified.

[0047] In FIG. 5, the distributed representation concatenation unit 13 outputs the sum (combination of distributed representations) of the distributed representations of the information input from Middle B via the character string division unit 11 and the distributed representation generation unit 12 to the type classification unit 15. The classifier of the type classification unit 15 outputs the confidence level for each type. In the illustrated example, it is assumed that the type classification unit 15 can identify three types: "name", "address", and "company name".

[0048] When the highest value of the confidence level for each type is positive, the type determination unit 16 assigns the type with the highest value to the input information. Here, since the confidence level of the "name" with the highest value is positive, the type determination unit 16 determines that the type of the input information is "name", and outputs the input information and the determined type to the First B System 5.

[0049] In FIG. 6, similar to FIG. 5, the distributed representation concatenation unit 13 outputs the sum of the distributed representations to the type classification unit 15, and the type classification unit 15 outputs the confidence level for each type. In the illustrated example, the confidence level of "name" with the highest value output by the type classification unit 15 is negative. Since the highest value of the confidence level is negative, the type determination unit 16 determines that the input information does not match any existing type, and outputs the distributed representations of the key and value of the input information to the similar word extraction unit 17. Each functional unit 11-18 outputs the data processed by the functional unit and also outputs the data input to the functional unit. Therefore, the type determination unit 16 obtains the distributed representations of the key and value generated by the distributed representation generation unit 12 via the distributed representation concatenation unit 13 and the type classification unit 15. The subsequent processing will be described later.

[0050] FIG. 7 is a diagram showing the overall flow of the information identification device 1 when identifiable information in FIG. 5 is input.

[0051] The Middle B system 3 transmits arbitrary information to the character string splitting unit 11 of the information identification device 1 (step S41). The character string splitting unit 11 splits the input information into words (step S42). The distributed representation generation unit 12 generates the distributed representation of each split word (step S43). The distributed representation concatenation unit 13 combines the distributed representations of each word (step S44). Here, as identifiable information, "name: Taro Watanabe" is input in S41, and the distributed representation concatenation unit 13 calculates the sum of the distributed representations of "name", "Watanabe", and "Taro".

[0052] The type classification unit 15 outputs the confidence level for each type for the input combination of distributed representations (step S45).

[0053] The type determination unit 16 receives the confidence level output from the type classification unit 15 (step S46), and determines whether the confidence level of the highest value is positive (step S47). Here, since the confidence level of the highest value is positive (step S47: YES), the type determination unit 16 determines that the information input in S41 is the type with the highest confidence level. (Step S48). That is, the type determination unit 16 assigns the type with the highest confidence level to the input information. Here, the type of "name" is assigned to "Name: Taro Watanabe".

[0054] Then, the type determination unit 16 adds the determined type (or type number) to the input information and outputs it to the first B system 5. Each functional unit 11-18 outputs the data processed by the functional unit and also outputs the data input to the functional unit. Therefore, the type determination unit 16 acquires the information input to the character string splitting unit 11 via the distributed representation generation unit 12, the distributed representation concatenation unit 13, and the type classification unit 15.

[0055] FIG. 8 and FIG. 9 are diagrams showing the overall flow of the information identification device 1 when unidentifiable information in FIG. 6 is input. Steps S41 to S45 shown in FIG. 8 are the same as steps S41 to S45 described in FIG. 7. However, in FIG. 8, the "Person in Charge: Taro Watanabe" input in step S41 is assumed to be unidentifiable information that does not correspond to any existing type.

[0056] The type determination unit 16 receives the confidence level output from the type classification unit 15 (step S46). In FIG. 8, since the confidence level of the highest value of "name" is negative (step S47: NO), the type determination unit 16 determines that the type of the information input in S41 is unidentifiable by the current type classification unit 15. In this case, the type determination unit 16 adds a new type number to the type holding unit 19 (step S49). Then, the type determination unit 16 outputs the distributed representation of the key (person in charge) and the distributed representation of the value (Taro Watanabe) of the information input in S41 to the similar word extraction unit 17 (step S50).

[0057] Proceeding to FIG. 9, as described above, in step S49, the type determination unit 16 adds a new type and type number ("person in charge" - 3) to the type holding unit 19. Also, in step S50, the distributed representation of the key ("person in charge") of the information input in S41 and the distributed representation of the value ("Taro Watanabe") are output to the similar word extraction unit 17.

[0058] The similar word extraction unit 17 extracts similar words of the key and similar words of the value using the distributed representation of the key and the distributed representation of the value, and outputs them to the distributed representation generation unit 12 (step S51). For example, the similar word extraction unit 17 extracts similar words of the key and value using the FastText model of the non-patent document 2 and the cosine similarity. As similar words of "person in charge" (key), "engaged in", "employed in", "director", etc. are extracted, and as similar words of "Taro Watanabe" (value), "Watanabe", "Ito", "Sato", etc. are extracted.

[0059] The distributed representation generation unit 12 generates distributed representations of each similar word of the key and each similar word of the value (step S52). The distributed representation connection unit 13 combines the distributed representations of the key and the similar words of the key and the distributed representations of the value and the similar words of the value respectively to generate additional combinations of distributed representations (step S53). For example, the following combinations of distributed representations are generated.

[0060] · The sum of the distributed representation of (key) and the distributed representation of (value) · The sum of the distributed representation of (key) and the distributed representation of (similar word of value) · The sum of the distributed representation of (similar word of key) and the distributed representation of (similar word of value) · The sum of the distributed representation of (similar word of key) and the distributed representation of (value) The learning data generation unit 14 generates additional learning data including the combinations of distributed representations generated by the distributed representation connection unit 13 and the new type number issued by the type determination unit 16 in step S49, and outputs it to the type classification unit 15 (step S54). The learning unit of the type classification unit 15 retrains the classifier using the additional learning data.

[0061] FIG. 10 is obtained by adding to the explanatory diagram shown in FIG. 3 the scattered expressions (vectors) of a new type (person in charge) and similar words of the new type (involved), and the sum of the scattered expressions ("person in charge: Yamaguchi", "involved: Yamaguchi"). As shown in the figure, it shows that the sum of the scattered expressions of "person in charge" and the sum of the scattered expressions of "involved" are mapped to nearby positions.

[0062] As can be seen from FIG. 10, by taking the sum of the scattered expressions as a combination of the scattered expressions, information of the same type is mapped to nearby positions. Therefore, by using the sum of the scattered expressions of the key and value of the new type and the sum of the scattered expressions of the similar words of the key and value of the new type to generate additional learning data for the new type and training the classifier, it becomes possible to identify the new type.

[0063] That is, when generating a new type, as additional learning data, a combination of the scattered expression of the key that is the new type and the scattered expression of the value is used. At that time, by extracting the key and value similar words and training the learning data that utilizes the sum of the scattered expressions of the extracted similar words, it becomes possible to generate a new type of class in the classifier and improve the identification accuracy of the new type.

[0064] FIGS. 11A and 12 are sequence diagrams showing the operation of the information identification device 1 in the identification phase. FIG. 11A shows the operation when the value of the input information can be expressed in a scattered form, and FIG. 11B shows an example of the scattered expressions of similar words and additional learning data for re-training. FIG. 12 shows the operation when the value of the input information cannot be expressed in a scattered form.

[0065] In FIG. 11A, the Middle B system inputs information including a key and a value to the character string splitting unit 11 of the information identification device 1 (step S61). Here, "person in charge: Taro Watanabe" (key: value) is input.

[0066] The character string splitting unit 11 splits the information into words (character strings) (step S62). Here, it is split into "person in charge", "Watanabe", and "Taro". The distributed representation generation unit 12 generates the distributed representation of each split word and outputs it to the distributed representation concatenation unit 13 (step S63).

[0067] The distributed representation concatenation unit 13 combines the distributed representations of each word to generate a combination of distributed representations and outputs it to the type classification unit 15 (step S65). Here, the sum of the distributed representation of "person in charge", the distributed representation of "Watanabe", and the distributed representation of "Taro" is generated. The classifier of the type classification unit 15 outputs the confidence level for each type for the input combination of distributed representations (step S65). When there are K types classified by the classifier, the classifier calculates K confidence levels.

[0068] When the maximum confidence level output by the classifier is a positive value, the type determination unit 16 determines that the type of the information input in S61 is the type with the maximum confidence level. Then, the type determination unit 16 transmits an identification result including the information input in S61 and the type with the maximum confidence level (type name and / or type number) to the first B system 5 (step S66). Here, [(person in charge: Watanabe Taro), 3] is transmitted. 3 is the type number of "person in charge".

[0069] On the other hand, when the maximum confidence level output by the classifier is a negative value, the type determination unit 16 registers the key of the information input in S61 in the type holding unit 19 (step S67). The type holding unit 19 pays out the type number of the registered key (step S68) and outputs a completion notification to the type determination unit 16 (step S69). Here, the type holding unit 19 pays out "3" as the type number of "person in charge" and holds it in association with "person in charge". Note that the type determination unit 16 may pay out the type number of the key and register the key and the type number in the type holding unit 19 in step S67.

[0070] When the type determination unit 16 receives the completion notification, it outputs the distributed representation of the key and the distributed representation of the value to the similar word extraction unit 17 (step S70). Here, the type determination unit 16 outputs the distributed representation of "person in charge" (key) and the distributed representation of "Taro Watanabe" (value).

[0071] The similar word extraction unit 17 extracts similar words similar to the distributed representation of the key and similar words similar to the distributed representation of the value, and outputs them to the distributed representation generation unit 12 (step S71). At this time, the similar word extraction unit 17 also outputs the distributed representation of the key and the distributed representation of the value obtained in step S70 to the distributed representation generation unit 12 together with the similar words.

[0072] The distributed representation generation unit 12 generates the distributed representation of each input similar word (step S72). The distributed representation concatenation unit 13 generates combinations of additional distributed representations by combining the distributed representation of the key and the distributed representations of its similar words output in step S70, and the distributed representation of the value and the distributed representations of its similar words output in step S70, and sends them to the learning data update unit 18 (step S73).

[0073] The learning data update unit 18 generates additional learning data including the combinations of additional distributed representations and the new type numbers paid out in step S68, and outputs it to the type classification unit 15 (step S74). The learning unit of the type classification unit 15 uses the additional learning data to re-learn and update the classifier, and notifies the distributed representation concatenation unit 13 of the completion of re-learning (step S75).

[0074] The distributed representation concatenation unit 13 returns to step S64, and re-inputs the combination of the distributed representations of the information input in step S61 to the classifier after re-learning of the type classification unit 15. The classifier outputs the confidence levels of each type (step S65). At this time, the confidence level of the "person in charge" type becomes the maximum, and the type determination unit 16 determines that the type of the information input in S61 is "person in charge".

[0075] FIG. 11B shows an example of the distributed representation of similar words and the learning data. In the list 131 of the distributed representations of the illustrated similar words, as similar words of the key "person in charge", distributed representations such as "be involved in", "engage in", and "director" are shown, and as similar words of the value "Taro Watanabe", distributed representations such as "Okamoto", "Yamamoto", and "Watanabe" are shown.

[0076] In the list 132 of the learning data, a plurality of learning data including combinations of distributed representations and type numbers are shown. As additional learning data 133 for relearning a new type of "person in charge", for example, the sum of the distributed representation of "person in charge" (key) and the distributed representation of "Okamoto" (similar word of value) and the learning data of type number "3" are exemplified.

[0077] Next, referring to FIG. 12, the operation when the value of the input information cannot be represented as a distributed representation will be described.

[0078] The Middle B system 3 inputs information including a key and a value to the string splitting unit 11 of the information identification device 1 (step S61). The string splitting unit 11 splits the information into words (step S62). Here, "phone number: 090 - 1234 - 5678" (key: value) is input and split into "phone", "number", and "090 - 1234 - 5678".

[0079] The distributed representation generation unit 12 attempts to generate distributed representations for each split word, but if there is a word that cannot be represented as a distributed representation, an error occurs. Here, the distributed representation generation unit 12 outputs the "090 - 1234 - 5678" (value) that cannot be represented as a distributed representation to the regular expression determination unit 20 (step S81).

[0080] When the value is input, the regular expression determination unit 20 requests and acquires all pairs of the regular expressions and type numbers registered in the type holding unit 19 from the type holding unit 19 (steps S82, S83). Then, the regular expression determination unit 20 determines which pattern of the acquired regular expressions the pattern of the value string input in step S82 matches.

[0081] When it matches the pattern of any regular expression, the regular expression determination unit 20 determines that the type of the information input in S61 is the type of the regular expression of the matched pattern. Then, the regular expression determination unit 20 transmits an identification result including the information input in S61 and the type of the matched pattern to the first B system 5 (step S84).

[0082] On the other hand, when it does not match the pattern of any regular expression, the regular expression determination unit 20 transmits an error indicating that the type cannot be specified to the middle B system 3 (step S85).

[0083] <Confidence level of type classification degree> Next, the confidence level for each type output by the classifier (SVM) of the type classification unit 15 will be described. The confidence level is calculated using the boundary surface in each class of the classifier that performs multi-class classification, the distance between the input information and the boundary surface, and on which side (positive or negative) of the boundary surface the input information exists. That is, the classifier calculates a positive or negative confidence level from the distance between the boundary surface and the input information.

[0084] As a prerequisite, in order to determine whether there is a type to be assigned using the confidence level, it is necessary that there are three or more types that the classifier can identify. When there is only one type that can be identified, a boundary cannot be drawn, so the confidence level cannot be calculated. Also, when there are two types that can be identified, since the confidence level of one of the types is always positive, it cannot be determined that it is unidentifiable.

[0085] FIG. 13 is an explanatory diagram for explaining a method of calculating the distance between a boundary surface (hyperplane) and a point. Let the foot of the perpendicular drawn from the point X(x - ) to the boundary surface be H(h). Then, the following vector is parallel to the normal vector w of the boundary surface because it is perpendicular to the boundary surface.

[0086]

Equation

[0087] Therefore, it can be expressed as follows using the real number k.

[0088]

Number

[0089] Here, H(h) is a point on the boundary surface w T Since it is a point on x + b = 0, the following equation holds.

[0090]

Number

[0091] Therefore, the distance d between the desired point and the boundary surface is calculated by the following equation.

[0092]

Number

[0093] Therefore, the following equation holds. That is, the boundary surface w in each class of the classifier performing multi-class classification T x + b = 0 and the distance from the input information X(x - ) is expressed by the following equation.

[0094]

Number

[0095] <Implementation Example of the Distributed Representation Connection Part and the Type Judgment Part> Figure 14A shows an implementation example of the distributed representation connection part 13 in Python, and Figure 14B shows an implementation example of the type judgment part 16 in Python.

[0096] The implementation example of the distributed representation connection part 13 shown in Figure 14A shows that the combination of distributed representations is calculated by taking the sum of the distributed representation of key (name) and the distributed representation of value (the distributed representation of Watanabe + the distributed representation of Taro).

[0097] In the implementation example of the type determination unit 16 shown in FIG. 14B, when the confidence level of the highest value output by the type classification unit 15 is greater than 0 (positive value), "true" (identifiable) is output, and when the confidence level of the highest value is 0 or less (0 or negative value), "false" (unidentifiable) is output.

[0098] The information identification device 1 of the present embodiment described above includes a character string division unit 11 that divides learning information into words, a distributed representation generation unit 12 that generates a distributed representation of each word, a distributed representation connection unit 13 that combines the distributed representations of each word to generate a combination of distributed representations, a learning data generation unit 14 that generates learning data including the combination of distributed representations and the type number of the type corresponding to the learning information, and a learning unit that performs machine learning on the learning data to generate a classifier that identifies input information into any type.

[0099] Further, the information identification device 1 of the present embodiment includes a type determination unit 16 that determines whether the type of the input information is any of the existing types or a new type using the confidence level of the highest value output by the classifier.

[0100] Further, the type determination unit 16 of the information identification device 1 of the present embodiment includes a learning data update unit 18 that generates a new type number for a new type and generates additional learning data including the new type number, and the learning unit re-learns the classifier using the additional learning data.

[0101] Thereby, as shown in FIG. 15, in the present embodiment, even when existing systems of each company input information of the same type in different formats, it is possible to identify that the information is of the same type.

[0102] Further, when unknown information that cannot be identified by the classifier is input, it is correctly identified as information of a new type, learning data for the new type is generated, and the classifier is re-learned. As a result, the re-learned classifier can classify a new type, and thereby automatic information identification in the information flow becomes possible.

[0103] The information identification device 1 described above can use, for example, a general-purpose computer system as shown in FIG. 16. The illustrated computer system includes a CPU (Central Processing Unit), a processor 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906. The memory 902 and the storage 903 are storage devices. In this computer system, the functions of the information identification device 1 are realized by the CPU 901 executing a predetermined program loaded onto the memory 902.

[0104] The information identification device 1 may be implemented on one computer or may be implemented on a plurality of computers. Further, the information identification device 1 may be a virtual machine implemented on a computer. The program of the information identification device 1 can be stored in a computer-readable recording medium such as an HDD, an SSD, a USB (Universal Serial Bus) memory, a CD (Compact Disc), or a DVD (Digital Versatile Disc), or can be distributed via a network.

[0105] Note that the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist thereof.

Explanation of Reference Numerals

[0106] 1: Information identification device 11: Character string splitting unit (splitting unit) 12: Dispersion expression generation unit 13: Dispersion expression connection unit 14: Learning data generation unit 15: Type classification unit 16: Type determination unit (determination unit) 17: Similar word extraction unit 18: Learning data update unit 19: Type holding unit 20: Regular Expression Judgment Unit 3: Middle B System 5: First B System 7: Operator Terminal

Claims

1. A splitting unit that splits learning information into words; A distributed representation generation unit that generates a distributed representation of each word; A distributed representation concatenation unit that combines the distributed representations of each word to generate a combination of distributed representations; A learning data generation unit that generates learning data including the combination of the distributed representations and the type number of the type corresponding to the learning information; A learning unit that performs machine learning on the learning data to generate a classifier that classifies input information into any one of the types; A determination unit that, when the confidence level of the highest value output by the classifier is positive, determines that the type of the input information is the type of the highest value, and when the confidence level of the highest value is negative, determines that the type of the input information is a new type. An information identification device.

2. The determination unit determines whether the type of the input information is any one of the existing types or the new type by using the confidence level of the highest value. The information identification device according to claim 1.

3. The determination unit generates a new type number for the new type, and includes a learning data update unit that generates additional learning data including the new type number, and the learning unit relearns the classifier using the additional learning data. The information identification device according to claim 1 or 2.

4. When the determination unit determines a new type, it includes a similar word extraction unit that extracts a first similar word of the key of the input information and a second similar word of the value of the input information respectively, the distributed representation generation unit generates a distributed representation of the first similar word and a distributed representation of the second similar word, the distributed representation concatenation unit combines the distributed representation of the first similar word and the distributed representation of the second similar word to generate a combination of additional distributed representations, and the learning data update unit generates the additional learning data including the combination of the additional distributed representations and the new type number. The information identification device according to claim 3.

5. The distributed representation concatenation unit calculates the sum of the distributed representations of each word as the combination of the distributed representations. The information identification device according to any one of claims 1 to 4.

6. An information identification method performed by an information identification device, including the steps of splitting learning information into words, generating a distributed representation of each word, combining the distributed representations of each word to generate a combination of distributed representations, generating learning data including the combination of the distributed representations and the type number of the type corresponding to the learning information, A step of generating a classifier that identifies input information into any type by machine learning of the learning data; When the confidence of the highest value output by the classifier is positive, determining the type of the input information as the type of the highest value, and when the confidence of the highest value is negative, determining the type of the input information as a new type; An information identification method.

7. A program for causing a computer to function as the information identification device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text classification processing method, text classification processing device and text classification processing program

    JP2008084064A

  • Information determination model learning device and program thereof

    JP2019215705A