Knowledge Graph Generation Device, Computer Program, and Method
The knowledge graph generation device automatically generates knowledge graphs from frame images by detecting objects and representing their relationships, addressing the time-consuming issue of user interaction in existing techniques.
Patent Information
- Application Number
- JP2021099556
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-15
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Existing knowledge graph completion techniques require user interaction to complete knowledge graphs with unknown entities, leading to time-consuming processes.
A knowledge graph generation device that acquires frame images, performs image recognition to detect objects, and generates text representing relationships between objects, allowing for the automatic generation of knowledge graphs without user input.
Enables the efficient generation of knowledge graphs consisting of two entities and their relationships without requiring user answers, thereby reducing processing time and improving automation.
Smart Images

Figure 0007687071000002 
Figure 0007687071000003 
Figure 0007687071000004
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for generating a knowledge graph consisting of two entities and a relationship (relation) between the two entities.
Background Art
[0002] A knowledge graph is an effective description method for intuitively expressing the relationship between two entities.
[0003] According to Patent Document 1, a knowledge graph completion device capable of improving the accuracy of knowledge graph completion, when an unknown entity exists in a knowledge graph storage unit that stores a knowledge graph, based on the pattern of the known knowledge graph stored in the knowledge graph storage unit, generates triple candidates (referring to two entities and the relationship between them) for the unknown entity, calculates the confidence of each triple candidate, generates a question sentence for the triple candidate with the highest confidence, and asks the user a question. Using the user's answer to the question, the knowledge graph for the unknown entity is completed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, according to the technology disclosed in Patent Document 1, in order to complement a knowledge graph including unknown entities, interaction with the user is necessary. For this reason, there is a problem that it takes time for the user to answer.
[0007] In order to solve the above problems, an aspect according to the present disclosure provides a knowledge graph generation device, a computer program, and a method capable of generating a knowledge graph including two entities and a relation between the two entities without asking the user for an answer.
Means for Solving the Problems
[0008] To achieve this object, an aspect according to the present disclosure is a knowledge graph generation device, including an acquisition unit that acquires a frame image including a plurality of objects, and among the plurality of objects, Two objects and the the relationship between two objects and to represent text generate text a generation unit, and the relationship between two entities corresponding to the two objects Two entities and the to extract from the and and text extract from the extracted the two entities and between the two entities relationship and It is characterized by comprising a knowledge graph generation means for generating a knowledge graph consisting of
[0009] Here, further, it may be configured to include a storage means for storing the knowledge graph, and a writing means for writing the knowledge graph generated by the knowledge graph generation means into the storage means.
[0010] Here, further, image recognition is performed on the acquired frame image to detect a plurality of objects, and a feature vector corresponding to the object is output to the text generation means, and it may be configured to include an image recognition means.
[0012] Here, the text generation means Furthermore, similar texts similar to the said text is generated, and the knowledge graph generation means Using the said text and the similar texts, consisting of two entities and the relationship between the two entities a plurality of knowledge graphs to generate may be formed.
[0013] Here, the knowledge graph generation means generates a plurality of knowledge graph candidates consisting of a plurality of entities corresponding to the plurality of objects By generating all combinations of two of the entities and the relationships between entities , a relationship between two entities and the two entities, and for each of the generated knowledge graph candidates the knowledge graph candidates a sentence is created, and from the plurality of knowledge graph candidates, among the sentences thus created, the one texts generated from similar to the text corresponding to the The candidates of the text as a knowledge graph knowledge graph may be selected.
[0014] Here, the knowledge graph generation means calculates the similarity to a sentence for each of the generated plurality of knowledge graph candidates texts created from , and Generated by the said text generation means selects, as a knowledge graph, a knowledge graph candidate whose calculated similarity is equal to or greater than a predetermined threshold. corresponding to the text
[0016] Also, one aspect of the present disclosure is an information processing apparatus that performs information processing using a knowledge graph, and is the knowledge graph generation device described aboveand , Generated by the said knowledge graph generation device a storage means for storing a plurality of generated knowledge graphs; image a reception means for receiving input of data, and means for generating a knowledge graph from the data input by the reception means; and The frame image of the said image data from the execution means for determining whether there is a knowledge graph that matches the generated knowledge graph among the plurality of knowledge graphs stored in the storage means, and if there is a knowledge graph determined to match, executing a predetermined process that is to be executed when such a determination is made for the knowledge graph. From the frame image of the said image data by the knowledge graph generation device It is characterized by comprising: Here, the image data is photographed image data, the storage means stores in advance the correspondence between each of the plurality of knowledge graphs and the risk level of the photographed situation, and the execution means outputs an alarm sound when the risk level corresponding to the determined matching knowledge graph is equal to or higher than a threshold value as the predetermined process. Also, the image data is photographed image data, the storage means stores in advance the correspondence between each of the plurality of knowledge graphs and the scene in the photographed image data, and the execution means may display identification information for identifying the scene corresponding to the determined matching knowledge graph as the predetermined process.
[0017] Also, one aspect of the present disclosure is a control computer program used in a knowledge graph generation device that is a computer, and the computer includes an acquisition step of acquiring a frame image including a plurality of objects, and among the plurality of objects, Two objects and the the relationship between two objects and representing text generating text a generation step, and extracting the relationship between two entities corresponding to the two objects from the Two entities and the relationship between two entities and from the text and extracting extracted the two entitiesand between the two entities relationship and characterized by executing a knowledge graph generation step of generating a knowledge graph consisting of
[0018] Here, the knowledge graph generation device includes storage means for storing a knowledge graph, and the computer program Cause the computer to further includes a writing step of writing the knowledge graph generated by the knowledge graph generation step into the storage means execute may be used as well
[0019] Here, the computer program Cause the computer to further includes an image recognition step of performing image recognition on the acquired frame image to detect a plurality of objects, and outputting a feature vector corresponding to the object to the text generation step execute may be used as well
[0021] Here, the text generation step Furthermore, similar texts similar to the said text generates, and the knowledge graph generation step Using the said text and the similar texts, consisting of two entities and the relationship between the two entities a plurality of knowledge graphs to generate may be formed
[0022] Here, the knowledge graph generation step includes a plurality of entities corresponding to the plurality of objects By generating all combinations of two of the entities and the relationships between entities , generates a plurality of knowledge graph candidates consisting of two entities and the relationship between the two entities, and for each of the generated knowledge graph candidates the knowledge graph candidates creates a sentence, and from the plurality of knowledge graph candidates, among the sentences created, the one that is texts generated from similar to the text corresponding to the sentence generation step The candidates of the text as a knowledge graph knowledge graph may be selected
[0023] Here, for each of the generated plurality of knowledge graph candidates, the knowledge graph generation step texts created from for each Generated by the said text generation step Calculate the similarity with the article, and if the calculated similarity is equal to or greater than a predetermined threshold corresponding to the text A candidate for the knowledge graph may be selected as the knowledge graph.
[0025] In addition, one aspect of the present disclosure is a knowledge graph generation device to execute A method including: an acquisition step of acquiring a frame image including a plurality of objects; among the plurality of objects, Two objects and the The relationship between two objects and To represent text Generate text A generation step; and the relationship between two entities corresponding to the two objects Two entities and the The relationship between two entities and From the text Extract, extracted The two entities and between the two entities Relationship and A knowledge graph generation step of generating a knowledge graph consisting of is characterized by including.
[0026] Here, the knowledge graph generation device includes storage means for storing the knowledge graph, and the method may further include a writing step of writing the knowledge graph generated by the knowledge graph generation step into the storage means.
[0027] Here, the method may further include an image recognition step of performing image recognition on the acquired frame image to detect a plurality of objects and outputting a feature vector corresponding to the objects to the text Generation step.
[0029] Here, the text Generation step is Furthermore, similar texts similar to the said text Generate, and the knowledge graph generation step is Using the said text and the similar texts, consisting of two entities and the relationship between the two entities A plurality of knowledge graphs to generate May be formed.
[0030] Here, the knowledge graph generation step includes a plurality of entities corresponding to the plurality of objects By generating all combinations of two of the entities and the relationships between entities A plurality of knowledge graph candidates each consisting of two entities and a relationship between the two entities are generated, and for each of the generated knowledge graph candidates, the Knowledge Graph candidates and generating a sentence from the plurality of knowledge graph candidates, the sentences generated are selected from the plurality of knowledge graph candidates. texts generated from Similar to text Knowledge graph corresponding to The candidates of the text as a knowledge graph You may choose.
[0031] Here, the knowledge graph generating step includes: texts created from About Generated by the said text generation step The similarity with the text is calculated, and if the calculated similarity is equal to or greater than a predetermined threshold value, corresponding to the text A candidate knowledge graph may be selected as the knowledge graph. Effect of the Invention
[0033] According to this embodiment, it is possible to produce an excellent effect of generating a knowledge graph consisting of two entities and a relation between the two entities without requesting an answer from the user. [Brief description of the drawings]
[0034]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
MODE FOR CARRYING OUT THE INVENTION
[0035] 1 Embodiment A knowledge graph generation device 10 according to one embodiment will be described.
[0036] 1.1 Knowledge Graph Generation Device 10 As shown in FIG. 1, the knowledge graph generation device 10 includes a CPU (Central Processing Unit) 106, a ROM (Read Only Memory) 107, a RAM (Random access memory) 108, a bus 109, a storage unit 105, and an input / output unit 120. The CPU 106, the ROM 107, the RAM 108, the storage unit 105, and the input / output unit 120 are interconnected via the bus 109.
[0037] The RAM 108 is composed of a non-volatile semiconductor memory that can be read from and written to, and provides a work area during program execution by the CPU 106.
[0038] The ROM 107 is composed of a non-volatile semiconductor memory that can only be read from, and stores a control program and the like, which is a computer program for executing processing in the knowledge graph generation device 10.
[0039] The CPU 106 operates according to the control program stored in the ROM 107.
[0040] When the CPU 106 uses the RAM 108 as a work area and operates according to the control program stored in the ROM 107, the CPU 106, the ROM 107, and the RAM 108 functionally constitute an image recognition unit 101, a caption generation unit 102, a similar sentence generation unit 103, and a knowledge graph generation unit 104.
[0041] (1) Input / Output Unit 120 and Storage Unit 105 The input / output unit 120 (acquisition means, writing means) reads data stored in the storage unit 105 and writes data to the storage unit 105.
[0042] The memory unit 105 (memory means) is composed of, for example, a hard disk unit.
[0043] As shown in FIG. 1, the memory unit 105 stores moving image data 131 and a knowledge database 150.
[0044] (Moving image data 131) The moving image data 131 is, for example, moving image data having a format defined by MPEG (Moving Picture Experts Group). FIG. 2 shows an example of a moving image obtained by photographing a man with a smartphone descending a staircase. As shown in this figure, the moving image data 131 includes still images (frame images) 132, 133, 134, 135,... arranged in time series. The still images 132, 133, 134, 135,... are each identified by an identification number.
[0045] Each still image includes one or more object images. Each object image represents one or more objects. Here, an object is a person or an object.
[0046] For example, the still image 134 shown in FIG. 4 includes three object images 134a, 134b, and 134c. The object image 134a represents a person and a smartphone. The object image 134b represents a person's arm and a smartphone. The object image 134c represents a person's foot and a staircase. Here, the person, the person's arm, the person's foot, the smartphone, and the staircase are each an object.
[0047] (Knowledge database 150) The knowledge database 150 is a database for storing a knowledge graph.
[0048] As shown in FIG. 3, the knowledge database 150 has areas for storing a plurality of knowledge graphs 151, 151a, 151b,....
[0049] The knowledge graph 151 is composed of an entity 1 (152), an entity 2 (154), and a relation 153 that associates the entity 1 (152) with the entity 2 (154). The relation 153 is labeled. The knowledge graph 151 is identified by an identification number 155. Other knowledge graphs 151a, 151b, ··· also have the same structure as the knowledge graph 151.
[0050] For example, the knowledge graph 151a is composed of an entity 1 (152a) "person", an entity 2 (154a) "staircase", and a relation 153a that associates the entity 1 (152a) with the entity 2 (154a). The relation 153a is labeled with "walk". The knowledge graph 151a is identified by an identification number 155a "ID002".
[0051] Also, the knowledge graph 151b is composed of an entity 1 (152b) "person", an entity 2 (154b) "smartphone", and a relation 153b that associates the entity 1 (152b) with the entity 2 (154b). The relation 153b is labeled with "look at". The knowledge graph 151b is identified by an identification number 155b "ID003".
[0052] Here, an object and an entity are substantially the same concept. The entity included in the object image is called an object, and the entity that is a component of the knowledge graph is called an entity.
[0053] (2) Image recognition unit 101 The image recognition unit 101 (image recognition means) performs image recognition on the acquired frame image as follows to detect a plurality of objects. Note that the entity (object) is represented as a code as described below.
[0054] The image recognition unit 101 incorporates, as an example, a neural network. The neural network is publicly known and its description is omitted. This neural network has, for example, previously learned a large number of images representing animals (including humans) and objects, the actions of animals and objects, the states of animals and objects, the relationships between objects, etc.
[0055] Here, animals include, for example, humans, men, women, dogs, cats, etc. Objects include, for example, containers, smartphones, stairs, etc. Also, the actions of animals include walking, looking, standing, etc. Furthermore, the state of an object, for example, when the object is a container, is the state where the lid of the container is open or closed, and when the object is a smartphone, is the state where the LED lamp indicating an incoming call is blinking.
[0056] The image recognition unit 101 reads the moving image data 131 from the storage unit 105 via the input / output unit 120.
[0057] As an example, as shown in FIG. 4, the image recognition unit 101 uses the incorporated neural network to detect one or more objects from the still images 134 included in the moving image data 131 by image recognition (process P01).
[0058] Next, the image recognition unit 101 generates feature vectors for the regions (objects) corresponding to the animals and objects recognized and distinguished from the still images included in the moving image data 131, and outputs the generated feature vectors to the caption generation unit 102. Also, the image recognition unit 101 may generate codes (numbers) indicating the animals and objects, the actions of the animals, and the states of the objects, etc., recognized and distinguished from the still images included in the moving image data 131, and output the generated codes to the caption generation unit 102 and the knowledge graph generation unit 104.
[0059] (3) Caption generation unit 102 The caption generation unit 102 (character expression generation means) generates a caption (character expression, sentence) representing the relationship between two objects among a plurality of objects as follows.
[0060] The caption generation unit 102 incorporates, for example, a Seq2Seq model such as a Transformer model. The Transformer model is composed of an encoder and a decoder, and uses word distance (Positional Encoding), attention mechanism (Multi-head Attention), fully connected (Feed Forward), etc. Since the Transformer model and the Seq2Seq model are well-known, the description is omitted.
[0061] As an example, as shown in FIG. 4, the caption generation unit 102 receives a feature vector from the image recognition unit 101. When receiving the feature vector, the caption generation unit 102 uses a language model such as a Transformer model to combine animals and objects corresponding to the received feature vector, animal actions, and object states, etc., to generate a caption (process P02).
[0062] Examples of captions generated by the caption generation unit 102 are "A man is walking up the stairs", "A man is looking at a smartphone", etc. Thus, the caption is a sentence representing the relationship between two objects.
[0063] The caption generation unit 102 outputs the generated caption to the similar sentence generation unit 103.
[0064] As described above, the caption generation unit 102 generates one sentence indicating the relationship between two objects.
[0065] (4) Similar sentence generation unit 103 The similar sentence generation unit 103 (text expression generation means) incorporates, for example, the above-described Transformer model. Note that the similar sentence generation unit 103 may use the Transformer model incorporated in the caption generation unit 102.
[0066] The similar sentence generation unit 103 receives a caption from the caption generation unit 102. Next, the similar sentence generation unit 103 uses the Transformer model to generate a similar sentence similar to the received caption.
[0067] As an example, as shown in FIG. 5(a), the similar sentence generation unit 103 generates a similar sentence 141a "A man is walking up the stairs" and a similar sentence 141b "A person is walking up the stairs" from the caption 141 "A man is walking up the stairs". Also, as an example, as shown in FIG. 5(b), the similar sentence generation unit 103 generates a similar sentence 142a "A man is looking at a mobile phone" and a similar sentence 142b "A man is looking at a smartphone" from the caption 142 "A man is looking at a smartphone".
[0068] The similar sentence generation unit 103 outputs the received caption and the generated similar sentences to the knowledge graph generation unit 104.
[0069] Here, as an example, for words that are similar in meaning such as "man", "male", "person", etc., respective occurrence probabilities may be set in advance. Also, as an example, for words that are similar in meaning such as "smartphone", "mobile phone", "smartphone", etc., respective occurrence probabilities may be set in advance. In this case, when generating a similar sentence, the occurrence probability of the generated similar sentence may be set according to the probability set for each word. The similar sentence generation unit 103 may adopt the similar sentence when the set probability among a plurality of similar sentences is equal to or greater than a predetermined threshold.
[0070] In this way, the similar sentence generation unit 103 generates one similar sentence that is similar to the caption and shows the relationship between two objects.
[0071] In this way, the caption generation unit 102 and the similar sentence generation unit 103 generate a plurality of sentences indicating the relationship between two objects.
[0072] (5) Knowledge graph generation unit 104 The knowledge graph generation unit 104 (knowledge graph generation means) extracts the relationship between two entities corresponding to two objects from the sentences representing the relationship between the two objects as shown below, and generates a knowledge graph consisting of the two entities and the extracted relationship.
[0073] The knowledge graph generation unit 104 receives feature vectors indicating animals, objects, animal actions, object states, etc. from the image recognition unit 101. Entities representing animals and objects are generated from the received feature vectors.
[0074] The knowledge graph generation unit 104 receives captions and similar sentences from the similar sentence generation unit 103. Here, captions and similar sentences are referred to as sentences.
[0075] For all the received sentences, the knowledge graph generation unit 104 interprets each sentence and extracts the actor, object, and predicate.
[0076] For example, as shown in Fig. 6(a), the knowledge graph generation unit 104 interprets the sentence "A man is walking on the stairs" (process P06) and extracts the actor 145a "A man", the object 145b "on the stairs", and the predicate 145c "is walking".
[0077] Next, among the generated entities, the knowledge graph generation unit 104 designates the entity corresponding to the actor as entity 1. Also, among the generated entities, the knowledge graph generation unit 104 designates the entity corresponding to the object as entity 2. Furthermore, the knowledge graph generation unit 104 uses the predicate as the relation. In this way, the knowledge graph generation unit 104 generates a knowledge graph consisting of entity 1, entity 2, and the relation.
[0078] For example, as shown in FIG. 6(a), the knowledge graph generation unit 104 sets the actor 145a "a man" as the entity 1 (162) "man", the object 145b "the stairs" as the entity 2 (164) "stairs", and the predicate 145c "is walking" as the relation 163 "walk", and generates a knowledge graph 161 composed of the entity 1 (162), the entity 2 (164), and the relation 163.
[0079] In this way, as an example shown in FIG. 7, for each of the caption 141, the similar sentences 141a, 141b, the caption 142, the similar sentences 142a, 142b, the knowledge graph generation unit 104 generates knowledge graphs 171, 172, 173, 174, 175, 176.
[0080] Next, the knowledge graph generation unit 104 writes the generated knowledge graph into the knowledge database 150 of the storage unit 105 via the input / output unit 120.
[0081] Further, the knowledge graph generation unit 104 may extract a plurality of relationships between two entities from a plurality of sentences (captions and similar sentences), and generate the two entities and each of the plurality of extracted relationships as a plurality of knowledge graphs.
[0082] (Modification Example 1) Note that the knowledge graph generation unit 104 may perform Semantic Role Labeling using a Deep Learning-based method, and generate a knowledge graph using the result.
[0083] That is, the knowledge graph generation unit 104 pre-learns and stores a method for generating a knowledge graph from a sentence based on a known knowledge graph or a known sentence and a known knowledge graph, and may generate a knowledge graph from the sentence based on the stored method.
[0084] Here, examples of the known knowledge graph and the known sentence (original sentence) are as follows.
[0085] Entity 1: "carnival glass" Relation: "derived from" Entity 2: "glass" Original sentence: "The word "carnival glass" is derived from "glass"" Using these, the knowledge graph generation unit 104 learns, end-to-end, a model that estimates entities 1 and 2 and the relation of the knowledge graph from the sentence.
[0086] Using this model, the knowledge graph generation unit 104 splits the sentence into words, converts them into feature vectors described later, and inputs them to the encoder of the Transformer model. Next, the decoder of the Transformer model may be replaced with a fully connected layer or the like to estimate words corresponding to entities 1 and 2 and the relation.
[0087] (Modification example 2) Further, the knowledge graph generation unit 104 may perform semantic role assignment based on predetermined rules by performing syntactic analysis or the like, and generate a knowledge graph using the result. Semantic role assignment is performed in the following order to obtain a knowledge graph.
[0088] (a) Morphological analysis: The sentence is segmented into morphemes, and the part of speech is estimated.
[0089] (b) Syntactic analysis: Discover structures such as the main part, the predicate part, noun phrases, and verb phrases that exist above the part of speech.
[0090] (c) Semantic analysis: Estimate the meaning of what the subject and the object of the action are in the sentence.
[0091] Fig. 6(b) shows an example of generating a knowledge graph by performing morphological analysis, syntactic analysis, and semantic analysis on a sentence.
[0092] The knowledge graph generation unit 104 performs morphological analysis on the sentence 31 and decomposes it into morphemes 32a, 32b, 32c, 32d, and 32e (process P07). The sentence 31 "I ate onigiri" is decomposed into the morpheme 32a "I", the morpheme 32b "wa", the morpheme 32c "onigiri", the morpheme 32d "wo", and the morpheme 32e "ate (past tense)". The morphemes 32a, 32b, 32c, 32d, and 32e are a pronoun, a particle, a noun, a particle, and a verb (past tense), respectively.
[0093] Next, the knowledge graph generation unit 104 performs syntactic analysis on the morphemes 32a, 32b, 32c, 32d, and 32e, classifies them into the noun phrase 33a "I", the particle 33b "wa", and the verb phrase 33c "ate onigiri", and further sets the noun phrase 33a and the particle 33b as the main part 33d "I wa", and the verb phrase 33c as the predicate part 33e "ate onigiri" (process P08).
[0094] Next, the knowledge graph generation unit 104 performs semantic analysis on the main part 33d and the predicate part 33e to generate the subject 34a "I", the action 34b "eat", the object of the verb 34c "onigiri", and the tense 34d "past" (process P09).
[0095] Next, among the entities generated from the code received from the image recognition unit 101, the knowledge graph generation unit 104 sets the entity corresponding to the subject 34a "I" as the entity 1 (35a) "I", sets the action 34b "eat" as the relation 35b, and sets the entity corresponding to the object of the verb 34c "onigiri" among the entities generated from the code received from the image recognition unit 101 as the entity 2 (35c), and generates a knowledge graph 35 composed of the entity 1 (35a), the relation 35b, and the entity 2 (35c) (process P10).
[0096] 1.2 Summary As described above, the knowledge graph generation device 10 includes an image recognition unit 101 that performs image recognition on the acquired frame image to detect objects and generates the two entities from two of the detected objects, a caption generation unit 102 that generates an original sentence representing the relationship between the two detected objects from the feature vectors corresponding to the two detected objects, a similar sentence generation unit 103 that generates a plurality of similar sentences similar to the generated original sentence from the generated original sentence, and a knowledge graph generation unit 104 that generates a plurality of knowledge graphs composed of the two entities and the relationship between the two entities using the generated original sentence and the plurality of similar sentences.
[0097] In this way, the knowledge graph generation device 10 can generate a knowledge graph composed of two entities and the relation between the two entities from an image without asking the user for an answer.
[0098] Note that since the caption represents the relationship between people and objects in the frame image, it is useful to use the caption in order to generate a knowledge graph that reflects the meaning represented by the frame image.
[0099] (1) Note that in the above, the knowledge graph generation device 10 is described as including a caption generation unit 102 and a similar sentence generation unit 103, but it is not limited to this.
[0100] The knowledge graph generation device 10 may not include the similar sentence generation unit 103.
[0101] In this case, as an example, the caption generation unit 102 may generate a caption "A man is walking on the stairs" from two object images 134a, 134c (FIG. 4) and output the generated caption to the knowledge graph generation unit 104 without generating a similar sentence of the caption.
[0102] Here, the knowledge graph generation unit 104 may use the generated caption to extract the relationship between two entities, and generate a knowledge graph consisting of two entities and the relationship between the two entities.
[0103] (2) The knowledge graph generation unit 104 may generate a knowledge graph as follows.
[0104] Non-Patent Document 1 describes an end-to-end multi-task method that uses Semantic Role Labeling (SRL, which extracts information such as who did what to whom) as the main task and predicate identification, dependency parsing, and part-of-speech identification as sub-tasks. The knowledge graph generation unit 104 may use this method to generate a knowledge graph from the results of SRL, with the subject as entity 1, the action as the relation, and the object of the action as entity 2.
[0105] 1.3 Variation A variation of the knowledge graph generation device 10 will be described.
[0106] In this variation, as shown in FIG. 8 as an example, the knowledge graph generation unit 104 vectorizes captions 181a and 181b (which may be similar sentences) as the sentences received from the similar sentence generation unit 103, and generates a feature vector 182 for each sentence (process P11). Here, as an example, caption 181a is "A man is walking up the stairs", and caption 181b is "A man is looking at a smartphone".
[0107] The procedure for generating the feature vector for each sentence will be described.
[0108] The knowledge graph generation unit 104 obtains a fixed-dimensional distributed representation from the received text. For example, a model based on deep learning such as SWEM (Simple Word-Embedding-based Methods), LSTM (Long Short Term Memory), CNN (Convolutional Neural Network), or BERT (Bidirectional Encoder Representations from Transformers) may be used to convert the text into a feature vector. Since SWEM, LSTM, CNN, and BERT are well-known, the description is omitted here for simplicity. Here, when the text is "A man is walking up the stairs", the feature vector is, for example, (2.3159367E-1, 5.31529129E-1, -6.28219426E-1, -7.73212969E-1, ···) (see Figure 8).
[0109] In this way, the knowledge graph generation unit 104 generates the feature vector for each text.
[0110] The memory unit 105 stores the classification model 183 in advance.
[0111] As shown as an example in Figure 8, the classification model 183 is composed of a plurality of relations. For example, the classification model 183 includes relation 183a "walk", relation 183b "see", ···, etc.
[0112] Next, the knowledge graph generation unit 104 determines whether a relation (for example, "walk") exists in the classification model 183 among the feature vectors generated for each text. If it exists, the relation 186 "walk" is specified (process P12). If it does not exist, the knowledge graph generation unit 104 does not specify a relation from this feature vector. Therefore, the following processing is not executed, and the knowledge graph is not generated.
[0113] Next, the knowledge graph generation unit 104 generates all combinations of two entities among the entities generated from the codes received from the image recognition unit 101 and the identified relation (process P13).
[0114] Here, in the example shown in FIG. 8, entity 1 is "man", "staircase", "smartphone", entity 2 is also "man", "staircase", "smartphone", and the relations are "walk" and "look at".
[0115] Also, the combinations of entity 1, entity 2, and the identified relation are knowledge graphs 184a, 184b, 184c,... in the example shown in FIG. 8.
[0116] Next, the knowledge graph generation unit 104 creates sentences from the generated combinations (process P14). In the example shown in FIG. 8, sentences 185a, 185b, 185c,... are generated.
[0117] Next, the knowledge graph generation unit 104 determines which of the generated sentences 185a, 185b, 185c,... the received caption 181a is similar to (process P15). When making the similarity determination, for example, feature vectors may be generated from the sentences 185a, 185b, 185c,... and the generated feature vectors may be compared with the feature vector of the received caption 181a.
[0118] Next, the knowledge graph generation unit 104 selects the knowledge graph corresponding to the sentence similar to the received caption 181a from among the sentences 185a, 185b, 185c,... The knowledge graph generation unit 104 writes the selected knowledge graph into the knowledge database 150 (process P16).
[0119] Note that the knowledge graph generation unit 104 may select a knowledge graph corresponding to the sentence among sentences 185a, 185b, 185c, ··· that is most similar to the received caption 181a. Further, the knowledge graph generation unit 104 may select a knowledge graph corresponding to a sentence among sentences 185a, 185b, 185c, ··· whose similarity to the received caption 181a is equal to or greater than a predetermined threshold value.
[0120] As described above, the knowledge graph generation unit 104 selects a plurality of candidate pairs each consisting of two entities from a plurality of entities corresponding to a plurality of objects, extracts relationships from the generated sentences (captions or similar sentences), generates a plurality of knowledge graph candidates each consisting of the extracted relationships and each of the plurality of selected candidate pairs, and may select a knowledge graph from the plurality of generated knowledge graph candidates using the sentences.
[0121] Further, for each of the plurality of generated knowledge graph candidates, the knowledge graph generation unit 104 calculates the similarity to the generated sentence (caption or similar sentence), and may select, as a knowledge graph, a knowledge graph candidate whose calculated similarity is equal to or greater than a predetermined threshold value.
[0122] Thus, also in the modification example, it is possible to generate a knowledge graph consisting of two entities and a relation between the two entities from the image.
[0123] Further, by generating a plurality of similar sentences, a plurality of types of knowledge graphs can be temporarily generated, and among them, a knowledge graph with high accuracy can be added to the knowledge database.
[0124] Also, in a natural language model, by utilizing the property that words having similar meanings have similar feature vectors, similar sentences can be easily generated.
[0125] 1.4 Application Example (1) A risk prediction device 10a (information processing device) as Application Example (1) of the embodiment will be described.
[0126] The risk prediction device 10a estimates the risk level of the captured situation based on the captured moving image data.
[0127] (1) Risk prediction device 10a As shown in FIG. 9, the risk prediction device 10a includes a CPU 106, a ROM 107, a RAM 108, a storage unit 105 (storage means), a bus 109, an input unit 111, a camera 112, an output unit 113, a speaker 114, and an input / output unit 120. The CPU 106, ROM 107, RAM 108, storage unit 105, input unit 111, output unit 113, and input / output unit 120 are interconnected via the bus 109.
[0128] The CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the risk prediction device 10a have the same configurations as the CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the knowledge graph generation device 10 in the embodiment, respectively.
[0129] In addition to the functions of the knowledge graph generation device 10 in the embodiment, the risk prediction device 10a has unique functions.
[0130] Here, the description will focus on the differences from the knowledge graph generation device 10.
[0131] The ROM 107 stores a control program and the like, which are computer programs for executing the processing in the risk prediction device 10a.
[0132] When the CPU 106 operates according to the control program stored in the ROM 107 using the RAM 108 as a work area, the CPU 106, ROM 107, and RAM 108 functionally constitute an image recognition unit 101, a caption generation unit 102, a similar sentence generation unit 103, a knowledge graph generation unit 104, and a risk determination unit 110.
[0133] The image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 each have the same configuration as the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 10 of the knowledge graph generation device 10 in the embodiment.
[0134] (2) Memory unit 105 The memory unit 105 stores the moving image data 131, the knowledge database 150, and the rule table 165. The memory unit 105 also has an area for storing the moving image data 170.
[0135] The moving image data 131 and the knowledge database 150 are the same as the moving image data 131 and the knowledge database 150 in the embodiment, respectively.
[0136] It is assumed that the knowledge database 150 already stores the knowledge graph generated from the moving image data 131.
[0137] As an example, the rule table 165 is a data table that stores a plurality of rule data 166 in advance as shown in FIG. 10. Each rule data 166 indicates the risk level when using the knowledge graph included in the knowledge database 152.
[0138] Each rule data 166 includes an identification number 167, a condition 168, and a risk level 169 associated therewith.
[0139] The identification number 167 is identification information for uniquely identifying the corresponding rule data 166.
[0140] Condition 168 indicates the condition to which the corresponding risk level 169 is applied. Condition 168 consists of identification information that identifies the knowledge graph included in the knowledge database 150. Condition 168 may, for example, consist of one piece of identification information that identifies one knowledge graph. In this case, Condition 168 indicates that only one knowledge graph is used. Also, Condition 168 may, for example, consist of two pieces of identification information that identify two knowledge graphs. In this case, the two knowledge graphs are combined by an AND condition. Furthermore, Condition 168 may, for example, consist of three or more pieces of identification information that identify three or more knowledge graphs.
[0141] Risk level 169 indicates the risk level when the corresponding rule data 166 is applied. Risk level 169 takes, for example, any value from "0", "1", "3", ···, "9". "0" is the lowest risk level, and "9" is the highest risk level.
[0142] Risk level 169 is set manually according to the knowledge graph.
[0143] For example, as shown in FIG. 10, the rule data identified by the identification number "R001" includes the condition "ID002" and the risk level "1". The condition "ID002" indicates the knowledge graph 151a (FIG. 3) included in the knowledge database 150. Since the knowledge graph 151a indicates that a person is walking on the stairs, its risk level is relatively low (risk level "1").
[0144] Also, for example, as shown in FIG. 10, the rule data identified by the identification number "R002" includes the condition "ID003" and the risk level "0". The condition "ID003" indicates the knowledge graph 151b (FIG. 3) included in the knowledge database 150. Since the knowledge graph 151b indicates that a person is looking at a smartphone, its risk level is extremely low (risk level "0").
[0145] Also, for example, as shown in FIG. 10, the rule data identified by the identification number "R003" includes the condition "ID002 AND ID003" and the risk level "5". "ID002" included in the condition indicates the knowledge graph 151a included in the knowledge database 150, and "ID003" included in the condition indicates the knowledge graph 151b included in the knowledge database 150. That is, the condition "ID002 AND ID003" indicates the case where the knowledge graph 151a and the knowledge graph 151b are established. In this case, since it indicates that a person is looking at a smartphone while walking on the stairs, the risk level is relatively high (risk level "5").
[0146] (3) Input unit 111 and camera 112 The camera 112 (reception means) generates moving image data by shooting and outputs the generated moving image data to the input unit 111.
[0147] The input unit 111 (reception means) receives the moving image data from the camera 112 and writes the received moving image data as moving image data 170 to the storage unit 105 via the bus 109.
[0148] (4) Output unit 113 and speaker 114 The output unit 113 outputs an electrical signal indicating an alarm sound to the speaker 114 under the control of the risk determination unit 110.
[0149] When the speaker 114 receives an electrical signal indicating an alarm sound from the output unit 113, it converts the received electrical signal into sound and outputs it as an alarm sound.
[0150] (5) Image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 In addition to the functions described in the embodiments, when moving image data 170 is written into the storage unit 105, the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 generate one or more knowledge graphs from the moving image data 170 in the same manner as the method described in the embodiments, and output the generated one or more knowledge graphs to the risk determination unit 110.
[0151] (6) Risk determination unit 110 When moving image data 170 is written into the storage unit 105, the risk determination unit 110 (execution means) receives one or more knowledge graphs generated based on the moving image data 170 from the knowledge graph generation unit 104.
[0152] The risk determination unit 110 determines whether one or more knowledge graphs generated based on the moving image data 170 match any of the conditions in the rule table 160.
[0153] If one or more knowledge graphs generated based on the moving image data 170 match any of the conditions in the rule table 160, the risk determination unit 110 extracts the corresponding risk level from the rule table 160.
[0154] Next, the risk determination unit 110 compares the extracted risk level with a threshold value (for example, "4"). If the extracted risk level is lower than the threshold value, the risk determination unit 110 determines that the risk level is low and ends the process.
[0155] On the other hand, if the extracted risk level is higher than the threshold value or the extracted risk level matches the threshold value, the risk determination unit 110 determines that the risk level is high and controls the output unit 113 to output an electrical signal indicating an alarm sound to the speaker 114. In this case, the speaker 114 converts the electrical signal received from the output unit 113 into sound and outputs it as an alarm sound.
[0156] In this way, the risk determination unit 110 executes processing on the data input by the camera 112 and the input unit 111 by using the knowledge graph stored in the knowledge database 150 of the storage unit 105.
[0157] (7) Operations in the risk prediction device 10a The operations in the risk prediction device 10a will be described with reference to the flowchart shown in FIG. 11.
[0158] The camera 112 generates moving image data by shooting (step S201).
[0159] The image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 generate a knowledge graph from the moving image data generated by the camera 112 (step S202).
[0160] The risk determination unit 110 compares the generated knowledge graph with the rule table 160 (step S203).
[0161] When the generated knowledge graph matches the conditions in the rule table 160 (''YES'' in step S204), the risk determination unit 110 extracts the risk level from the rule data including the conditions that match the generated knowledge data (step S205).
[0162] The risk determination unit 110 compares the extracted risk level with the threshold value (step S206).
[0163] When the extracted risk level is higher than the threshold value or when the extracted risk level matches the threshold value (''≧'' in step S206), the risk determination unit 110 controls the output unit 113 to output an alarm sound. The speaker 114 outputs the alarm sound (step S207). Thus, a series of processing is terminated.
[0164] If the generated knowledge graph does not match the conditions in the rule table 160 ( "NO" in step S204), or if the extracted risk level is lower than the threshold value ( "<" in step S206), the process ends.
[0165] As described above, the operation of the risk prediction device 10a ends.
[0166] (8) Summary As described above, using the knowledge database and the rule table, it is possible to estimate the risk level of the situation where the moving image is taken from the moving image data generated by photographing with the camera.
[0167] 1.5 Application Example (2) The search device 10b (information processing device) as the application example (2) of the embodiment will be described.
[0168] (1) Search Device 10b The search device 10b searches for a scene that matches or is similar to the scene of the newly input moving image from among a plurality of scenes of the moving image data stored in advance.
[0169] Here, a scene has a meaning as a single block and is composed of a plurality of consecutive still images. For example, a scene where a man is eating, a scene where a man is walking up the stairs, a scene where a man is looking at a smartphone, etc. Also, for example, when a man is walking up the stairs, the scene of walking on the upper part of the stairs can be regarded as one scene, the scene of walking on the middle part of the stairs can be regarded as one scene, and the scene of walking on the lower part of the stairs can be regarded as one scene.
[0170] As shown in FIG. 12, the search device 10b is composed of a CPU 106, a ROM 107, a RAM 108, a storage unit 105 (storage means), a bus 109, an input unit 111, a camera 112, an output unit 116, and a monitor 117. The CPU 106, ROM 107, RAM 108, storage unit 105, input unit 111, output unit 116, and input / output unit 120 are interconnected via the bus 109.
[0171] The CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the search device 10b each have the same configuration as the CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the knowledge graph generation device 10 of the embodiment.
[0172] In addition to the functions of the knowledge graph generation device 10 of the embodiment, the search device 10b has unique functions.
[0173] Here, the description will focus on the differences from the knowledge graph generation device 10.
[0174] The ROM 107 stores a control program and the like, which are computer programs for executing the processing in the search device 10b.
[0175] The CPU 106 uses the RAM 108 as a work area and operates according to the control program stored in the ROM 107. As a result, the CPU 106, ROM 107, and RAM 108 functionally constitute an image recognition unit 101, a caption generation unit 102, a similar sentence generation unit 103, a knowledge graph generation unit 104, and a similarity determination unit 115.
[0176] The image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 each have the same configuration as the image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 of the knowledge graph generation device 10 of the embodiment.
[0177] (2) Storage Unit 105 The storage unit 105 stores moving image data 131, a knowledge database 150, and a scene table 190. The storage unit 105 also has an area for storing the moving image data 170.
[0178] The moving image data 131 and the knowledge database 150 are the same as the moving image data 131 and the knowledge database 150 in the embodiment, respectively. It is assumed that the knowledge database 150 of the search device 10b already stores knowledge graphs for all scenes in the moving image data 131, respectively.
[0179] As an example, as shown in FIG. 13, the scene table 190 is a data table that stores a plurality of scene data 191 in advance. Each scene data 191 shows the correspondence between the knowledge graph included in the knowledge database 152 and the scenes in the moving image data 131.
[0180] Each scene data 191 includes the identification number 192 of the scene and the identification number 193 of the knowledge graph in association with each other.
[0181] The identification number 192 of the scene is identification information for uniquely identifying the scenes in the moving image data 131.
[0182] The identification number 193 of the knowledge graph is identification information for uniquely identifying the knowledge graph included in the knowledge database 152.
[0183] If the knowledge graph included in the knowledge database 152 is specified, the scenes in the moving image data 131 corresponding to the specified knowledge graph can be specified using the scene table 190.
[0184] (3) Input unit 111 and camera 112 The camera 112 (receiving means) generates moving image data by shooting and outputs the generated moving image data to the input unit 111.
[0185] The input unit 111 (receiving means) receives the moving image data from the camera 112 and writes the received moving image data as moving image data 170 to the storage unit 105 via the bus 109.
[0186] (4) Output unit 116 and monitor 117 The output unit 116 outputs, to the monitor 117, an identification number for identifying the specified scene under the control of the similarity determination unit 115.
[0187] When the monitor 117 receives, from the output unit 116, an identification number for identifying the specified scene, the monitor 117 displays the received identification number.
[0188] (5) Image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 In addition to the functions described in the embodiments, when the moving image data 170 is written into the storage unit 105, the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 generate one or more knowledge graphs from the moving image data 170 in the same manner as the method described in the embodiments, and output the generated one or more knowledge graphs to the similarity determination unit 115.
[0189] (6) Similarity determination unit 115 When the moving image data 170 is written into the storage unit 105, the similarity determination unit 115 receives, from the knowledge graph generation unit 104, one or more knowledge graphs generated based on the moving image data 170.
[0190] The similarity determination unit 115 determines whether one or more knowledge graphs generated based on the moving image data 170 match any of the knowledge graphs in the knowledge database 150.
[0191] When one or more knowledge graphs generated based on the moving image data 170 match any of the knowledge graphs in the knowledge database 150, the similarity determination unit 115 extracts, from the scene table 190, the identification number of the scene corresponding to the identification number for identifying the matching knowledge graph.
[0192] Next, the similarity determination unit 115 outputs the extracted identification number of the scene to the output unit 116 and controls the output unit 116 to display the extracted identification number of the scene on the monitor 117.
[0193] On the other hand, when one or more knowledge graphs generated based on the moving image data 170 do not match any of the knowledge graphs in the knowledge database 150 of the storage unit 105, the similarity determination unit 115 outputs a message indicating that there is no matching scene to the output unit 116 and controls to display the message on the monitor 117. The monitor 117 displays the message.
[0194] In this way, the similarity determination unit 115 uses the knowledge graphs stored in the knowledge database 150 of the storage unit 105 to execute processing on the data input by the camera 112 and the input unit 111.
[0195] (7) Operations in the search device 10b The operations in the search device 10b will be described with reference to the flowchart shown in FIG. 14.
[0196] The camera 112 generates moving image data by shooting (step S231).
[0197] The image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 generate a knowledge graph from the moving image data generated by the camera 112 (step S232).
[0198] The similarity determination unit 115 compares the generated knowledge graph with the knowledge graphs in the knowledge database 150 (step S233).
[0199] When the generated knowledge graph matches any of the knowledge graphs in the knowledge database 150 (”YES” in step S234), the similarity determination unit 115 extracts, from the scene table 190, the identification number of the scene corresponding to the identification number that identifies the matching knowledge graph (step S235). Next, the similarity determination unit 115 outputs the extracted scene identification number to the output unit 116 and controls the output unit 116 to display the extracted scene identification number on the monitor 117. The monitor 117 displays the extracted scene identification number (step S236). Thus, the series of processes ends.
[0200] On the other hand, when none of the one or more knowledge graphs generated based on the moving image data 170 match any of the knowledge graphs in the knowledge database 150 (”NO” in step S234), the similarity determination unit 115 outputs, to the output unit 116, a message indicating that there is no matching scene and controls the output unit 116 to display the message on the monitor 117. The monitor 117 displays the message (step S237). Thus, the series of processes ends.
[0201] (8) Summary As described above, using the knowledge database and the scene table, it is possible to search for a scene that matches or is similar to the scene in the moving image data generated by shooting with a camera from among the scenes of the moving images stored in advance.
[0202] Note that in the above, when the generated knowledge graph matches any of the knowledge graphs in the knowledge database 150, the monitor 117 displays the identification number of the extracted scene. However, the present invention is not limited to this.
[0203] When the generated knowledge graph matches any of the knowledge graphs in the knowledge database 150, the similarity determination unit 115 extracts, from the scene table 190, the identification number of the scene corresponding to the identification number for identifying the matching knowledge graph, outputs the extracted identification number of the scene to the output unit 116, and the output unit 116 extracts, from the moving image data 131, the scene identified by the received identification number of the scene, and may output the identification number of the scene and the extracted scene to the monitor 117. The monitor 117 displays the identification number of the scene and the scene.
[0204] 1.6 Application Example (3) The VQA (Visual Question Answering) device 10c (information processing device) as the application example (3) of the embodiment will be described.
[0205] The VQA device 10c obtains an answer to a question from an image and question data using a knowledge database.
[0206] (1) VQA device 10c As shown in FIG. 15, the VQA device 10c includes a CPU 106, a ROM 107, a RAM 108, a storage unit 105 (storage means), a bus 109, an input unit 126, an output unit 116, a monitor 117, and an input / output unit 120. The CPU 106, ROM 107, RAM 108, storage unit 105, input unit 126, output unit 116, and input / output unit 120 are interconnected via the bus 109.
[0207] The CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the VQA device 10c each have the same configuration as the CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the knowledge graph generation device 10 of the embodiment.
[0208] In addition to the functions of the knowledge graph generation device 10 of the embodiment, the VQA device 10c has unique functions.
[0209] Here, the description will focus on the differences from the knowledge graph generation device 10.
[0210] The ROM 107 stores a control program and the like, which are computer programs for executing the processing in the VQA device 10c.
[0211] The CPU 106 uses the RAM 108 as a work area and operates according to the control program stored in the ROM 107. As a result, the CPU 106, ROM 107, and RAM 108 functionally constitute an image recognition unit 101, a caption generation unit 102, a similar sentence generation unit 103, a knowledge graph generation unit 104, and a VQA unit 140.
[0212] The image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 each have the same configuration as the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 of the knowledge graph generation device 10 in the embodiment.
[0213] (2) Storage unit 105 The storage unit 105 stores the moving image data 131 and the knowledge database 150.
[0214] The moving image data 131 and the knowledge database 150 are the same as the moving image data 131 and the knowledge database 150 in the embodiment, respectively.
[0215] It is assumed that the knowledge database 150 already stores the knowledge graph generated from the moving image data 131.
[0216] (3) Input unit 126 The input unit 126 (reception means) receives the input of the input data 127.
[0217] As shown as an example in FIG. 16, the input data 127 is composed of an image 127a and question data 127b. The image 127a is, for example, a still image. The question data 127b is, for example, text data representing a question such as "What is the person looking at?".
[0218] The input unit 126 outputs the received input data 127 to the VQA unit 140.
[0219] (4) Output unit 116 and monitor 117 The output unit 116 receives answer data for the question from the VQA unit 140. When receiving the answer data, the received answer data is output to the monitor 117.
[0220] When the monitor 117 receives the answer data from the output unit 116, the received answer data is displayed.
[0221] (5) Image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 In addition to the functions described in the embodiment, the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 receive the image 127a from the VQA unit 140. When receiving the image 127a, in the same manner as the method described in the embodiment, one knowledge graph 301 (FIG. 16) is generated from the image 127a, and the generated knowledge graph 301 is output to the VQA unit 140.
[0222] (6) VQA unit 140 The VQA unit 140 receives the input data 127 from the input unit 126. As described above, the input data 127 includes, as an example, the image 127a and the question data 127b.
[0223] When receiving the input data 127, the VQA unit 140 outputs the image 127a to the image recognition unit 101 and instructs the image recognition unit 101, the caption generation unit 102, the similar sentence generation unit 103, and the knowledge graph generation unit 104 to generate the knowledge graph of the image 127a. The VQA unit 140 receives the knowledge graph 301 (FIG. 16) from the knowledge graph generation unit 104 (process P51).
[0224] Also, when receiving the input data 127, the VQA unit 140 performs language analysis on the question data 127b to generate a question-type knowledge graph 302 (FIG. 16) (process P52).
[0225] The VQA unit 140 compares the knowledge graph 301 with the knowledge graph in the knowledge database 150 to determine whether there is a matching knowledge graph (process P53).
[0226] If there is a matching knowledge graph, the VQA unit 140 extracts the matching knowledge graph 151b (FIG. 16) from the knowledge database 150.
[0227] Next, the VQA unit 140 compares the extracted knowledge graph 151b with the generated question-type knowledge graph 302 (process P54) to identify the entity corresponding to "What?" in the question-type knowledge graph 302. Here, in the example shown in FIG. 16, the VQA unit 140 obtains "smartphone" as the answer data 303.
[0228] The VQA unit 140 outputs the answer data 303 to the output unit 116 and controls to output the answer data 303 to the monitor 117. The monitor 117 displays the answer data 303 (process P55).
[0229] In this way, the VQA unit 140 uses the knowledge graph of the knowledge database 150 stored in the storage unit 105 to execute processing on the data input by the input unit 126.
[0230] (7) Summary As described above, using the knowledge database, answers to questions can be obtained from images and questions.
[0231] 1.7 Application Example (4) The knowledge graph completion device 10d (information processing device) as the application example (4) of the embodiment will be described.
[0232] (1) Knowledge Graph Completion Device 10d As shown in FIG. 17, the knowledge graph completion device 10d includes a CPU 106, a ROM 107, a RAM 108, a storage unit 105 (storage means), a bus 109, an input unit 118, a microphone 119, an output unit 113, a speaker 114, and an input / output unit 120. The CPU 106, ROM 107, RAM 108, storage unit 105, input unit 118, output unit 113, and input / output unit 120 are interconnected via the bus 109.
[0233] The CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the knowledge graph completion device 10d have the same configurations as those of the CPU 106, ROM 107, RAM 108, storage unit 105, bus 109, and input / output unit 120 of the knowledge graph generation device 10 of the embodiment, respectively.
[0234] In addition to the functions of the knowledge graph generation device 10 of the embodiment, the knowledge graph completion device 10d has unique functions.
[0235] Here, the description will focus on the differences from the knowledge graph generation device 10.
[0236] The ROM 107 stores a control program and the like, which are computer programs for executing the processing in the knowledge graph completion device 10d.
[0237] The CPU 106 operates according to a control program stored in the ROM 107, using the RAM 108 as a work area. As a result, the CPU 106, ROM 107, and RAM 108 functionally constitute an image recognition unit 101, a caption generation unit 102, a similar sentence generation unit 103, a knowledge graph generation unit 104, a dialogue generation unit 121, a voice recognition unit 122, a language analysis unit 123, and a completion unit 124.
[0238] The image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 each have the same configuration as the image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 of the knowledge graph generation device 10 of the embodiment.
[0239] (2) Storage unit 105 The storage unit 105 stores moving image data 131 and a knowledge database 150.
[0240] As an example, the knowledge database 150 already includes a knowledge graph 151c including an unknown entity 154c, in addition to the knowledge graphs 151, 151a, 151b,... as shown in FIG. 18.
[0241] (3) Input unit 118 and microphone 119 The microphone 119 (reception means) receives an input of voice, converts the voice into an analog electrical signal, and outputs the analog voice signal to the input unit 118.
[0242] The input unit 118 (reception means) receives the analog voice signal from the microphone 119, converts the received analog into a digital voice signal, and outputs the digital voice signal to the voice recognition unit 122 via the bus 109.
[0243] (4) Output unit 113 and speaker 114 The output unit 113 receives digital message sound data from the dialogue generation unit 121, converts the received digital message sound data into an analog message sound signal, and outputs the analog message sound signal to the speaker 114.
[0244] The speaker 114 receives an analog message sound signal from the output unit 113, converts the received message sound signal into sound, and outputs it as sound.
[0245] (5) Image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 The image recognition unit 101, caption generation unit 102, similar sentence generation unit 103, and knowledge graph generation unit 104 have the functions described in the embodiments.
[0246] (6) Dialogue generation unit 121 The dialogue generation unit 121 extracts a knowledge graph including unknown entities from the knowledge database 150. Next, for the extracted knowledge graph, it generates a dialogue sentence asking what the unknown entity is (process P41 in FIG. 18). Also, the dialogue generation unit 121 outputs the extracted knowledge graph including unknown entities to the complement unit 124.
[0247] In the knowledge database 150 shown in FIG. 18, as an example, the knowledge graph 151c includes an unknown entity 154c. Since the unknown entity 154c is an object to be operated, the dialogue generation unit 121 generates the dialogue sentence 201 "What is the object that a person operates?"
[0248] The dialogue generation unit 121 converts the generated dialogue sentence into digital message sound data, outputs the message sound data to the output unit 113, and controls so that an analog message sound signal is output from the speaker 114.
[0249] (7) Speech recognition unit 122 The voice recognition unit 122 receives a digital voice signal from the input unit 118. The voice recognition unit 122 performs voice recognition on the digital voice signal to generate a set of phonemes (process P42 in FIG. 18).
[0250] For example, when the digital voice signal is "personal computer", the voice recognition unit 122 generates a set of phonemes "pa", "so", "ko", "n", "de", "su".
[0251] The voice recognition unit 122 outputs the generated set of phonemes to the language analysis unit 123.
[0252] (8) Language analysis unit 123 The language analysis unit 123 receives a set of phonemes from the voice recognition unit 122.
[0253] The language analysis unit 123 performs language analysis on the received set of phonemes to generate words (process P43 in FIG. 18).
[0254] As an example, when receiving a set of phonemes "pa", "so", "ko", "n", "de", "su", the language analysis unit 123 generates the words "personal computer" and "is" from the set of phonemes.
[0255] The language analysis unit 123 outputs the generated words to the complement unit 124.
[0256] (9) Complement unit 124 The complement unit 124 receives a knowledge graph including unknown entities from the dialogue generation unit 121. The complement unit 124 also receives words from the language analysis unit 123.
[0257] The complement unit 124 complements the knowledge graph including unknown entities with the words received from the language analysis unit 123 (process P44 in FIG. 18).
[0258] In the example shown in FIG. 18, the complementing unit 124 complements the unknown entity 154c of the knowledge graph 151c with the word "personal computer" to generate the knowledge graph 203.
[0259] The complementing unit 124 rewrites the knowledge graph including the unknown entity in the knowledge database 150 into the knowledge graph generated by complementing (process P45 in FIG. 18).
[0260] (10) Summary As described above, the dialogue generation unit 121, the speech recognition unit 122, the language analysis unit 123, and the complementing unit 124 use the knowledge graph of the knowledge database 150 stored in the storage unit 105 to execute processing on the data input by the microphone 119 and the input unit 118.
[0261] As described above, the knowledge graph including the unknown entity in the knowledge database can be rewritten into the knowledge graph generated by complementing.
[0262] 2 Other Modification Examples Aspects according to the present disclosure are not limited to the embodiments and modification examples described above. You may configure as shown below.
[0263] (1) The caption generation unit 102 may generate multiple types of sentences as follows.
[0264] In sentence generation, when a part of the sentence (x1:t-1) is given, the probability of the subsequent word (xt) is predicted. The subsequent word is selected according to the probability of the subsequent word (see FIG. 19).
[0265] As shown in FIG. 19, when the word 321 "The" is given, the next following word 322 "car" is selected. Here, assume that the probabilities of the words "nice", "dog", "car",... are predicted to be "0.5", "0.4", "0.2",... respectively. The probability of the word 322 "car" is the third highest "0.2" from the top. Thus, it is also possible to select words other than the most probable one. Next, as the word following the word 322 "car", the word 323 "drives" is selected. Here, assume that the probabilities of the words "drives", "is", "stops",... are predicted to be "0.6", "0.4", "0.1",... respectively. The word 323 "drives" has the highest probability among the multiple selectable words.
[0266] Each word xi in the sentence is selected from the vocabulary V (xi ∈ V). If the sentence is x = (x1,..., xn), the probability of generating the sentence x is as follows.
[0267] [Number] There are roughly two ways to select the following word, as follows.
[0268] (a) Deterministic decoding: Select the word with the highest probability (b) Probabilistic decoding: Sample from the predicted probability distribution P(xt|x1,..., xt-1).
[0269] By using probabilistic decoding, multiple types of sentences can be generated.
[0270] Note that in order to prevent the generation of words with low probabilities, top-K sampling (selecting from the top K) is effective.
[0271] (2) An example of a knowledge graph consisting of entity 1, relation, and entity 2 is shown in FIG. 20.
[0272] The knowledge graphs 351 and 352 shown in FIG. 20 are the same as the knowledge graphs 151a and 151b shown in FIG. 3, respectively. The knowledge graphs 351 and 352 each show human behavior.
[0273] Here, the knowledge graph can show not only human behavior but also the following relationships.
[0274] The knowledge graph 353 consists of entity 1 "smartphone", relation "display", and entity 2 "message", indicating that the smartphone displays the message. That is, the knowledge graph 353 shows the operation of the smartphone.
[0275] The knowledge graph 354 consists of entity 1 "Mr. A", relation "friend", and entity 2 "Mr. B", indicating that Mr. A is a friend of Mr. B. That is, the knowledge graph 354 shows the relationship between Mr. A and Mr. B.
[0276] The knowledge graph 355 consists of entity 1 "student", relation "go to", and entity 2 "school", indicating that the student goes to school. That is, the knowledge graph 355 shows the behavior of the student.
[0277] The knowledge graph 356 consists of entity 1 "student", relation "belong to", and entity 2 "school", indicating that the student belongs to the school. That is, the knowledge graph 356 shows the belonging relationship of the student.
[0278] The knowledge graph 357 consists of entity 1 "elephant", relation "bigger than", and entity 2 "mouse", indicating that the elephant is bigger than the mouse. That is, the knowledge graph 357 shows the size relationship between the elephant and the mouse.
[0279] The knowledge graph 358 consists of entity 1 "iron ball", relation "heavier than", and entity 2 "feather", indicating that the iron ball is heavier than the feather. That is, the knowledge graph 358 shows the weight relationship between the iron ball and the feather.
[0280] The knowledge graph 359 consists of entity 1 "Tamagoyaki", relation "material", and entity 2 "egg", indicating that the material of tamagoyaki is egg. That is, the knowledge graph 359 shows that tamagoyaki and egg are related by the material.
[0281] The knowledge graph 360 consists of entity 1 "Tamagoyaki", relation "cooking method", and entity 2 "bake", indicating that tamagoyaki is cooked by baking. That is, the knowledge graph 360 shows that tamagoyaki and the act of baking are related by the cooking method.
[0282] The knowledge graph 361 consists of entity 1 "Tamagoyaki", relation "cooking utensil", and entity 2 "frying pan", indicating that tamagoyaki is cooked using a frying pan. That is, the knowledge graph 361 shows that tamagoyaki and the act of baking are related by the cooking utensil.
[0283] The knowledge graph 362 consists of entity 1 "hen", relation "parent - child", and entity 2 "chick", indicating that the hen and the chick are in a parent - child relationship.
[0284] The knowledge graph 363 consists of entity 1 "egg", relation "hatch", and entity 2 "chick", indicating that a chick hatches from an egg. That is, the knowledge graph 363 shows the relationship of the change from an egg to a chick.
[0285] The knowledge graph 364 consists of entity 1 "sky", relation "color", and entity 2 "blue", indicating that the color of the sky is blue. That is, the knowledge graph 364 shows one attribute of the sky, which is color.
[0286] The knowledge graph 365 consists of entity 1 "typhoon", relation "generate", and entity 2 "storm", indicating that a storm is generated by a typhoon. That is, the knowledge graph 365 shows the causal relationship between a typhoon and a storm.
[0287] As described above, the knowledge graph can represent not only human actions but also various relationships between entity 1 and entity 2.
[0288] (3) In the above embodiment, the image recognition unit 101 identifies a plurality of objects from one still image included in the moving image data. The knowledge graph generation unit 104 generates a knowledge graph showing the relationship between two entities among the plurality of identified objects (a plurality of entities) from one still image.
[0289] However, this is not limiting.
[0290] The image recognition unit 101 identifies a first object from the first still image included in the moving image data. Further, the image recognition unit 101 may identify a second object included in the moving image data. The knowledge graph generation unit 104 may generate a knowledge graph showing the relationship between the first object (first entity) identified from the first still image and the second object (second entity) identified from the second still image.
[0291] (4) Instead of the method shown in the above embodiment, the following may be done.
[0292] The image recognition unit 101 may detect a plurality of objects from the frame image, recognize "staircase", "smartphone", etc. from the detected objects, and output a code indicating the recognition result to the caption generation unit 102. The caption generation unit 102 may receive the code and generate a caption (literal expression, text) representing the relationship between two objects using the received code.
[0293] (5) The above embodiment and modification examples may be combined.
Industrial Applicability
[0294] Aspects according to the present disclosure have an excellent effect of being able to generate a knowledge graph consisting of two entities and the relationship between the two entities without asking the user for an answer, and are suitable as a technology for generating a knowledge graph.
Explanation of Signs
[0295] 10 Knowledge graph generation device 10a Risk prediction device 10b Search device 10c VQA device 10d Knowledge graph completion device 101 Image recognition unit 102 Caption generation unit 103 Similar sentence generation unit 104 Knowledge graph generation unit 105 Storage unit 106 CPU 107 ROM 108 RAM 109 Bus 110 Risk determination unit 111 Input unit 112 Camera 113 Output unit 114 Speaker 115 Similarity determination unit 116 Output unit 117 Monitor 118 Input unit 119 Microphone 120 Input / output unit 121 Dialogue generation unit 122 Speech recognition unit 123 Language analysis unit 124 Completion unit 140 VQA unit
Claims
1. An acquisition means for acquiring a frame image including a plurality of objects; A sentence generation means for generating a sentence representing two objects among the plurality of objects and the relationship between the two objects; A knowledge graph generation means for extracting two entities corresponding to the two objects and the relationship between the two entities from the sentence, and generating a knowledge graph consisting of the extracted two entities and the relationship between the two entities A knowledge graph generation device characterized by comprising the above.
2. Furthermore, A storage means for storing a knowledge graph; A writing means for writing the knowledge graph generated by the knowledge graph generation means into the storage means The knowledge graph generation device according to claim 1, characterized by comprising the above.
3. Furthermore, An image recognition means for performing image recognition on the acquired frame image to detect a plurality of objects, and outputting a feature vector corresponding to the object to the sentence generation means The knowledge graph generation device according to claim 1, characterized by comprising the above.
4. The sentence generation means further generates a similar sentence similar to the sentence, The knowledge graph generation means generates a plurality of knowledge graphs consisting of two entities and the relationship between the two entities by using the sentence and the similar sentence The knowledge graph generation device according to claim 1, characterized by the above.
5. The knowledge graph generation means generates a plurality of combinations of two entities and the relationship between entities among the plurality of entities corresponding to the plurality of objects, thereby generating a plurality of knowledge graph candidates consisting of two entities and the relationship between the two entities. For each generated knowledge graph candidate, a sentence is created from the knowledge graph candidate, and among the plurality of knowledge graph candidates, the knowledge graph candidate corresponding to the sentence similar to the sentence generated by the sentence generation means is selected as the knowledge graph The knowledge graph generation device according to claim 1, characterized by the above.
6. The knowledge graph generation means calculates the similarity between the sentence created from each of the generated plurality of knowledge graph candidates and the sentence generated by the sentence generation means, and selects the knowledge graph candidate corresponding to the sentence whose calculated similarity is equal to or higher than a predetermined threshold as the knowledge graph The knowledge graph generation device according to claim 5, characterized in that...
7. An information processing apparatus that performs information processing using a knowledge graph, comprising: the knowledge graph generation device according to any one of claims 1 to 6; storage means for storing a plurality of knowledge graphs generated by the knowledge graph generation device; reception means for receiving input of image data; means for generating a knowledge graph from the frame image of the image data input by the reception means using the knowledge graph generation device; determining whether there is a knowledge graph that matches the knowledge graph generated by the knowledge graph generation device from the frame image of the image data among the plurality of knowledge graphs stored in the storage means, and if there is a knowledge graph determined to match, executing a predetermined process that is to be executed when such a determination is made for the knowledge graph Execution means; An information processing apparatus characterized by comprising:
8. The image data is captured image data, the storage means stores in advance the correspondence between each of the plurality of knowledge graphs and the risk level of the captured situation, the execution means outputs an alarm sound when the risk level corresponding to the knowledge graph determined to match is equal to or higher than a threshold value as the predetermined process The information processing apparatus according to claim 7, characterized in that...
9. The image data is captured image data, the storage means stores in advance the correspondence between each of the plurality of knowledge graphs and the scene in the captured image data, the execution means displays identification information for identifying the scene corresponding to the knowledge graph determined to match as the predetermined process The information processing apparatus according to claim 7, characterized in that...
10. A control computer program used in a knowledge graph generation device that is a computer, the program causing the computer to: acquire a frame image including a plurality of objects; generate a sentence representing two objects among the plurality of objects and the relationship between the two objects; extract from the sentence two entities corresponding to the two objects and the relationship between the two entities, and generate a knowledge graph consisting of the extracted two entities and the relationship between the two entities A computer program for causing [the computer] to execute
11. The knowledge graph generation device includes storage means for storing a knowledge graph, The computer program further causes the computer to execute a writing step of writing the knowledge graph generated in the knowledge graph generation step into the storage means The computer program according to claim 10, characterized in that
12. The computer program further causes the computer to execute an image recognition step of performing image recognition on the acquired frame image to detect a plurality of objects and outputting a feature vector corresponding to the object to the sentence generation step The computer program according to claim 10, characterized in that
13. The sentence generation step further generates a similar sentence similar to the sentence, The knowledge graph generation step generates a plurality of knowledge graphs each composed of two entities and the relationship between the two entities using the sentence and the similar sentence The computer program according to claim 10, characterized in that
14. The knowledge graph generation step generates all combinations of two entities among the plurality of entities corresponding to the plurality of objects and the relationship between the entities, thereby generating a plurality of candidates for knowledge graphs each composed of two entities and the relationship between the two entities, creating a sentence from each candidate for the generated knowledge graph, and selecting, as a knowledge graph, a candidate for the knowledge graph corresponding to a sentence similar to the sentence generated in the sentence generation step among the sentences created from the plurality of candidates for the knowledge graphs The computer program according to claim 10, characterized in that
15. The knowledge graph generation step calculates the similarity between the sentences created from each of the generated plurality of candidates for knowledge graphs and the sentence generated in the sentence generation step, and selects, as a knowledge graph, a candidate for the knowledge graph corresponding to a sentence whose calculated similarity is equal to or greater than a predetermined threshold The computer program according to claim 14, characterized in that
16. A method executed by a knowledge graph generation device, comprising: an acquisition step of acquiring a frame image including a plurality of objects; A sentence generation step of generating a sentence representing two objects among the plurality of objects and the relationship between the two objects; A knowledge graph generation step of extracting two entities corresponding to the two objects and the relationship between the two entities from the sentence, and generating a knowledge graph consisting of the two extracted entities and the relationship between the two entities A method characterized by including the above.
17. The knowledge graph generation device includes a storage means for storing a knowledge graph, The method further includes A writing step of writing the knowledge graph generated by the knowledge graph generation step into the storage means The method according to claim 16, characterized by the above.
18. The method further includes An image recognition step of performing image recognition on the acquired frame image to detect a plurality of objects, and outputting a feature vector corresponding to the object to the sentence generation step The method according to claim 16, characterized by including the above.
19. The sentence generation step further generates a similar sentence similar to the sentence, The knowledge graph generation step generates a plurality of knowledge graphs consisting of two entities and the relationship between the two entities using the sentence and the similar sentence The method according to claim 16, characterized by the above.
20. The knowledge graph generation step generates a plurality of candidates for knowledge graphs consisting of two entities and the relationship between the two entities by generating all combinations of two entities among the plurality of entities corresponding to the plurality of objects and the relationship between the entities. For each candidate knowledge graph, a sentence is created from the candidate knowledge graph, and from the plurality of candidate knowledge graphs, a candidate knowledge graph corresponding to a sentence similar to the sentence generated by the sentence generation step among the created sentences is selected as the knowledge graph The method according to claim 16, characterized by the above.
21. The knowledge graph generation step calculates the similarity between the sentences created from each of the generated plurality of candidate knowledge graphs and the sentence generated by the sentence generation step, and selects a candidate knowledge graph corresponding to a sentence whose calculated similarity is equal to or greater than a predetermined threshold as the knowledge graph The method according to claim 20, characterized by the above.
Citation Information
Patent Citations
Entity recognition using multiple data streams for supplementing missing information related to entity
JP2020064613A
Information processing device, information processing method, and information processing program
JP2020129193A
Image search device, image search method, and computer program
JP2020149337A
Knowledge graph complementing device and knowledge graph complementing method
JP2020191009A