Information processing apparatus, method, and program
The information processing apparatus simplifies real-space searches by using an object arrangement characteristic database to predict answers, addressing the complexity of traditional methods and enabling efficient and accurate responses.
Patent Information
- Application Number
- JP2023200462
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
AI Technical Summary
Conventional technologies face complexity in searching the real space due to intricate operations such as data collection, parameter adjustment, and response operation setup for recognized results.
An information processing apparatus that includes an object arrangement information acquisition unit, an inquiry information input unit, and a prediction unit, which uses an object arrangement characteristic database to predict answers based on object arrangement information and inquiry information, simplifying the search process.
Enables efficient searching of the real space without the complexity of traditional methods, allowing for accurate and streamlined searches based on object arrangement characteristics.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, method, and program.
Background Art
[0002] Recently, machines have been able to recognize object characteristics such as the type and position of objects included in image information or three-dimensional shape information by image recognition or three-dimensional shape recognition from image information or three-dimensional shape information measured by sensors such as cameras and LiDAR. Furthermore, search in the real space for what is included in the image information or three-dimensional shape information is becoming possible.
[0003] In Patent Document 1, a predetermined action of a person recognized from an image is identified to detect a suspicious person. In Patent Document 2, the joint positions of a person identified from an image are recognized to recognize the working situation.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
[0006] However, in the conventional technology, for each task of searching the real space, there is a problem that the operations such as data collection, adjustment of parameters and conditions, and setting of the response operation of the system for the recognized results are complicated.
[0007] In view of the above, an object of the present invention is to perform a search of the real space without difficulty.
Means for Solving the Problems
[0008] An information processing apparatus according to an embodiment of the present invention includes: an object arrangement information acquisition unit that acquires object arrangement information including an object type and an arrangement relationship of an object, which is generated based on measurement information obtained by measuring a real space; an inquiry information input unit that inputs inquiry information from a user; and a prediction unit that inputs the object arrangement information and the inquiry information, and predicts an answer using an object arrangement characteristic database that holds object arrangement characteristics representing a positional relationship of a plurality of objects.
Effects of the Invention
[0009] According to the present invention, a search of the real space can be performed without difficulty.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0011] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the invention according to the claims, and not all combinations of features described in the embodiments are essential for the solution means of the invention. In each figure, the same components may be denoted by the same reference numerals and the description thereof may be omitted.
[0012] <Embodiment 1> In this embodiment, a case where the present invention is applied to the use of an image captured by an imaging device such as a surveillance camera (surveillance system) will be described as an example. Image recognition and three-dimensional space recognition are one of the most basic problems in computer vision. Image recognition and three-dimensional space recognition are applied not only to uses such as recognition of the type of an object, grasping of the position, and counting of the number of objects, but also to various tasks such as recognition of a place, obstacle avoidance in automatic driving, and danger prediction. Object detection from an image or a three-dimensional shape model is realized, for example, by a neural network that selects a region of interest (rectangle) and discriminates the type of an object included therein.
[0013] Based on the results of such image recognition and three-dimensional space recognition, search tasks in the real space are carried out, such as monitoring like detecting suspicious persons, work analysis like grasping the work procedures of factory workers, and logistics management like detecting lost items in factory logistics. However, in order to realize the task of searching the real space, for each task, operations such as data collection for generating an identifier, adjustment of parameters and conditions, and setting of the system's response actions for the recognized results are required, and the preparation is complicated.
[0014] For example, as a specific example, the case of constructing an alert sending system that sends an alert when a suspicious person approaches a car in a monitoring task will be described. First, a large amount of data for detecting people, cars, and objects held by people from the images of surveillance cameras or cameras of mobile robots is collected. Then, ground truth data is manually given to the data, and the neural network is trained with the data and the ground truth data paired to construct a recognizer.
[0015] Next, parameters and conditions of the recognized object types and their relative positional relationships are registered. For example, in order to detect that a person is holding a hammer and approaching within 1m of a car, each object is detected from the image and a rectangle is drawn, and the distance of the centroid positions of each rectangle is registered as a parameter. Then, the operations of the system when the conditions are met (for example, sending an alert, sending an email to the user, etc.) are set. In this way, a system for monitoring suspicious persons is constructed.
[0016] On the one hand, to have a human security guard perform similar monitoring tasks, for example, it can be achieved by making an inquiry such as "There have been many crimes of cars being damaged recently. Please alert me if a suspicious person approaches the car." That is, even without a detailed explanation of data collection, parameter setting, and response actions, a person can achieve the task based on experience and common sense. In this embodiment, it is an object to realize such a search instruction to a human and the execution of the human search by an information processing apparatus. That is, for each task, it is an object to reduce the labor of operations such as data collection for generating an identifier, adjustment of parameters and conditions, and setting of the system's response actions for the recognized results.
[0017] In this embodiment, based on an object arrangement characteristic database that aggregates the "object arrangement relationship" in the real space, information related to the objects in the real space and the relationship between those objects (corresponding to human common sense) is used. By this, the search in the real space can be realized without explaining the conditions in detail as in the case of making an inquiry to a human security guard.
[0018] <Operation Outline> FIG. 1 is a diagram for explaining the usage scenario and operation concept of a monitoring system according to Embodiment 1 of the present invention. F011 is a monitoring camera. F012 is a mobile robot equipped with a camera. The mobile robot F012 acquires image information F021 which is information of an image obtained by photographing the real space F001, and three-dimensional shape information F022 which is a three-dimensional shape model.
[0019] Here, in the real space F001 of FIG. 1, as described above, it shows a situation where a suspicious person, that is, a person holding a hammer, approaches the car. Such a situation is held in F031 as text information (a story including a time series: hereinafter referred to as text) that describes "object arrangement information" which is the type of objects included in the image information and the relationship between the positions of those objects. F031 is an object arrangement information holding unit. Specifically, the object arrangement information holding unit F031 holds it as text such as "A person took out a hammer from a bag, is approaching 1 m beside the car while holding the hammer...". The generation of text from images and three-dimensional information will be described later.
[0020] F032 is inquiry information. The inquiry information F032 indicates a prompt for specifying a monitoring target. The inquiry information F032 is the text input by the user as "Please send an alert if a suspicious person approaches the car."
[0021] F041 is an example of splitting the text describing the object arrangement information in the real space held in the object arrangement information holding unit F031 into "tokens", which are the smallest units that a computer can interpret. F042 is an example of splitting the text that is the inquiry information F032 into "tokens", which are the smallest units that a computer can interpret. That is, each of F041 and F042 is a token. F043 is SEQ, which is a token representing the boundary between the text describing the object arrangement information and the text that is the inquiry information. By inputting these tokens F041 to F043 into the object arrangement characteristic database F051 that aggregates the "object arrangement relationships" in the real space, the response text F061 is obtained. As the response text F061, for example, a text such as "There is a person trying to hit the car with a hammer. An alert is issued,..." is obtained.
[0022] FIG. 2 is a diagram showing the functional module configuration of the information processing apparatus 1 according to Embodiment 1 of the present invention. The information processing apparatus 1 includes an object arrangement information acquisition unit 101, an inquiry information input unit 102, an object arrangement characteristic database 103, and a prediction unit 104. The object arrangement characteristic database is not necessarily always possessed by the information processing apparatus 1, and a database can also be held in another apparatus.
[0023] The object arrangement information acquisition unit 101 acquires, as first object arrangement information, the text describing the objects detected from the image captured by a camera, which is an example of a measurement unit, and their arrangement relationships, from an object arrangement information holding unit. The camera is, for example, the monitoring camera F011 in FIG. 1 or the camera provided in the mobile robot F012. The object arrangement information holding unit is, for example, the object arrangement information holding unit F031 in FIG. 1. The image captured by the camera is an example of the measurement information measured by the measurement unit.
[0024] Article generation from an image can be implemented by the method of Johnson et al. (see Non-Patent Document 1) in which the image is convolved by a neural network to extract an object region and the relationship between the objects is described by an LSTM. Thereby, the object arrangement information acquisition unit 101 acquires, for example, an article such as "A person is holding a hammer. The person and the car are located at a distance of 1 m" from the image shown in the image information F021. The object arrangement information acquisition unit 101 outputs the acquired object arrangement information to the prediction unit 104 as first object arrangement information.
[0025] The inquiry information input unit 102 receives an inquiry from the user by a keyboard and inputs it as inquiry information in an article. The inquiry information includes the arrangement relationship (i.e., second object arrangement information) of the real space that the user wants to inquire about. Specifically, the inquiry information is an article such as "Please send an alert if a suspicious person approaches the car". The article of the inquiry information includes the arrangement relationship of "a suspicious person and a car approaching each other". The inquiry information input unit 102 outputs such inquiry information to the prediction unit 104.
[0026] The object arrangement characteristic database 103 is a database that holds object arrangement characteristics representing the positional relationship of a plurality of objects. The object arrangement characteristic is knowledge data obtained by generalizing the three-dimensional positional relationship of objects in the real world. That is, the characteristics of the arrangement relationship of the objects are held inside the database in order to determine that "a person and a car are located at a distance of 1 m" and "a suspicious person approaches the car" are similar. Also, the characteristics of the arrangement relationship of the objects are held inside the database in order to determine that "a person is holding a hammer" and "a suspicious person" are similar. That is, according to the object arrangement characteristic database 103, when data including the types of two or more objects and their arrangement relationships is input, the similarity with another arrangement relationship can be predicted using the characteristics of the arrangement relationship of the objects held inside. The concept of similarity will be described later with reference to FIGS. 6 to 8.
[0027] The object arrangement characteristic database 103 in this embodiment is a pre-trained neural network that is learned to infer the similarity of the arrangement relationship between two objects. The object arrangement characteristic database 103 is a neural network that interprets natural language and is learned to output answers related to the arrangement relationship of objects. Specifically, it is a neural network with 24 layers of the Transformer by Ashish et al. (see Non-Patent Document 2). In this embodiment, it is assumed that the input dimension number and the output dimension number of the Transformer are 512 dimensions, that is, a configuration in which a maximum of 512 pieces of object characteristic information are input and an output of 512 dimensions of the same number is obtained. Specifically, the encoder network used in the method by Jacob et al. (see Non-Patent Document 3) is adopted. For the learning of this network, a sentence related to the arrangement relationship of objects is input, and learning is performed to predict the words included in the subsequent sentence (answer) in order.
[0028] The prediction unit 104 inputs the output of the object arrangement information acquisition unit 101 and the output of the inquiry information input unit 102, and uses the object arrangement characteristic database 103 to evaluate the similarity of the object arrangements included in the object arrangement information and the inquiry information. Based on the evaluation result, the prediction unit 104 predicts a response sentence for the inquiry information and outputs the response sentence to a display, which is an example of the output unit.
[0029] Figure 3 is a diagram showing the hardware configuration of the information processing apparatus 1. The information processing apparatus 1 includes a CPU H11, a system bus H21, a ROM H12, a RAM H13, an external memory H14, an input unit H15, a display unit H16, a communication interface H17, and an I / O H18. CPU is an abbreviation for Central Processing Unit. ROM is an abbreviation for Read Only Memory. RAM is an abbreviation for Random Access Memory. I / O is an abbreviation for Input / Output.
[0030] The CPU executes the processing of this embodiment by executing a program that describes the operations in this embodiment. Also, the CPU H11 controls various devices connected to the system bus H21. The ROM H12 stores the BIOS program and the boot program. The RAM H13 is used as the main memory device of the CPU H11. The external memory H14 stores the programs processed by the information processing apparatus 100. The input unit H15 performs processing to receive inputs such as information from a keyboard, mouse, etc. The display unit H16 outputs the calculation results of the information processing apparatus 100 to the display device according to instructions from the CPU H11. Note that the display device can be of any type, such as a liquid crystal display device, a projector, an LED indicator, etc. The communication interface H17 performs information communication via a network. The communication interface can be Ethernet, or of any type such as USB, serial communication, wireless communication, etc. USB is the abbreviation for Universal Serial Bus. The object characteristic information input unit 101 inputs object characteristic group information via the communication interface H17. The prediction unit 103 outputs prediction results via the communication interface H17. The I / O H18 performs other input and output operations.
[0031] Figure 4 is a flowchart for explaining the operations of the information processing apparatus 1. The processing described in Figure 4 is automatically started when the information processing apparatus 1 is activated as the computer that executes the information processing apparatus 1 is powered on.
[0032] In step S101, the information processing apparatus 100 initializes the system. That is, it reads a program from the external memory H14 to make the information processing apparatus 1 operable. Also, if necessary, it reads the weight parameters of the neural network, which is the object placement characteristic database 103, from the external memory H14 into the RAM H13. Once a series of initialization processes are completed, it proceeds to step S102.
[0033] In step S102, the object arrangement information acquisition unit 101 acquires, as object arrangement information, a sentence including the arrangement relationship of objects in the real space from the object arrangement information holding unit F031. In step S103, the inquiry information input unit 102 inputs, as inquiry information, an inquiry from the user in the form of a sentence.
[0034] In step S104, the prediction unit 104 inputs the object arrangement information and the inquiry information into the object arrangement characteristic database 103, which is a neural network, performs forward propagation, and obtains an answer to the inquiry.
[0035] In step S105, the information processing device 100 performs an end determination. If the inquiry has not ended, it returns to step S102, and if the inquiry has ended, the process ends.
[0036] FIG. 5 is a flowchart for explaining the details of the process of step S104, which is a prediction process. In step S1001 of FIG. 5, the prediction unit 104 converts a sentence describing the arrangement relationship of objects, which is object arrangement information, into a format interpretable by a neural network. Specifically, the prediction unit 104 performs lexical analysis on the sentence, divides the text into tokens such as words, sub-words, and symbols, and assigns an ID to each token. Specifically, the conversion (encoding) into tokens shall employ the method of Yonghui et al. (see Non-Patent Document 4). In step S1002, the prediction unit 104 converts the sentence, which is inquiry information, into tokens in the same manner as in step S1001.
[0037] In step S1003, the prediction unit 104 combines the two tokenized sentences. Specifically, the prediction unit 104 arranges the two token groups, inserts a special token representing the boundary between the object arrangement information and the inquiry information, and generates a combined token for input into the object arrangement characteristic database 103.
[0038] In step S1004, the prediction unit 104 inputs the combined token into the object arrangement characteristic database 103.
[0039] In step S1005, the prediction unit 104 propagates the calculation results forward through each layer of the neural network, which is the object arrangement characteristic database 103, and obtains an output vector as an output token. Specifically, the prediction unit 104 weights the input token with the Transformer block of the object arrangement characteristic database 103. The object arrangement characteristic database 103 calculates scores for all word candidates held by the object arrangement characteristic database 103. Subsequently, the object arrangement characteristic database 103 outputs the word with the highest score and repeats calculating scores for the word candidates again to obtain a group of output tokens. The concept of the object arrangement characteristic database 103 interpreting the arrangement information of the object and calculating scores for words will be described later.
[0040] In step S1006, the prediction unit 104 converts (decodes) the output token into a sentence. For decoding, the method of Yonghui et al. is adopted.
[0041] FIG. 6, FIG. 7, and FIG. 8 are diagrams showing an example of the processing of step S1005 in which the prediction unit 104 interprets the arrangement relationship of the object and predicts an answer using the object arrangement characteristic database 103.
[0042] D001 in FIG. 6(A) is a structural diagram representing the arrangement relationship of the objects included in the measurement information obtained by measuring the real space in a graph format. The structural diagram D001 shows that the hammer is connected to the person and the bonnet of the car and they are located at spatially close positions.
[0043] D002 in FIG. 6(B) is a structural diagram representing the arrangement relationship of the objects included in the inquiry information in a graph format. The structural diagram D002 shows that the car and the suspicious person included in the inquiry information are located at close positions. In the structural diagram D002, the arrangement information of the objects not included in the inquiry information is indicated by "?".
[0044] The object arrangement characteristic database 103 is a neural network that has learned the prior probability of the arrangement relationship of the objects indicated by "?" included in the structural diagram D002. That is, as shown by the analogy result D011 in the structural diagram D003 of FIG. 6(C), the object arrangement characteristic database 103 analogizes that the suspicious person is a human, and if it is a suspicious person, objects such as a hammer, a mallet, and a mask are spatially close to the human. In addition, from the fact that the suspicious person is close to the car, the object arrangement characteristic database 103 analogizes that the car, the bonnet, and the damage are located spatially close to each other.
[0045] D004 in FIG. 7 is a diagram conceptually showing how the object arrangement characteristic database 103 interprets the object arrangement information and the inquiry information. D021 is a token obtained by converting a sentence that is the object arrangement information in the real space. D022 is a token obtained by converting a sentence that is the inquiry information. D023 is a token indicating the delimiter of the sentence. FIG. D004 is a diagram shown as a two-dimensional array to show the mutual relationship of the words input to the object arrangement characteristic database 103. D024 is shown in a dark gray color, which is a color indicating that the tokens corresponding to "suspicious person" and the tokens that are spatially and semantically close have a high degree of relevance. Specifically, the "suspicious person" included in the inquiry information indicates that it is spatially and semantically close to the "human" and "hammer" included in the object arrangement information.
[0046] These relationships are pre-learned in the Transformer block in the object arrangement characteristic database 103. That is, the similarity in this embodiment is the attention value of the tokens representing two objects, which is output using the weights pre-learned in the Transformer block. The higher the degree of relevance of the positional relationship between the two objects, the larger the attention value.
[0047] D005 in Fig. 8 shows how the output layer of the object arrangement characteristic database 103 predicts a sentence. D031 is a predicted token. D032 is the token to be predicted subsequently. D033 is a candidate for the predicted word, and a prediction confidence D034 is assigned to each word. Based on the prior probability of the arrangement of the objects held by the object arrangement characteristic database 103 and the arrangement relationship between the words included in the input sentence indicated by D004, the next word to be output is selected.
[0048] Here, assume that a sentence such as "A person took out a hammer from a bag and is approaching 1 m beside the car while holding the hammer..." which is object arrangement information is input to the object arrangement characteristic database 103. Also assume that a sentence such as "Please send an alert if a suspicious person approaches the car" which is inquiry information is input to the object arrangement characteristic database 103. In this case, the object arrangement characteristic database 103 operates as described above and obtains an output (answer) such as "A suspicious person is trying to hit the car with a hammer. Issue an alert."
[0049] <Effect> As described above, in Embodiment 1, based on the object arrangement information generated based on the measurement information obtained by measuring the real space and the arrangement relationship of the objects included in the inquiry information, an answer regarding the similarity of the object arrangement relationship is predicted. By doing so, for each task of searching the real space, the real space can be queried by a sentence without operations such as data collection for generating a discriminator, adjustment of parameters and conditions, and selection of the system operation for the recognized result. Therefore, the construction of the search system and the complexity of the inquiry can be reduced.
[0050] <Modification Example 1-1> In Embodiment 1, a sentence describing the arrangement relationship of objects included in the measurement information obtained by measuring the real space was used as the object arrangement information. The object arrangement information of the present invention is not limited to a sentence, and may be any data structure that allows the object arrangement characteristic database 103 to determine the arrangement relationship of objects. The present invention may hold the object arrangement information in a form that summarizes the arrangement relationship, such as a list, or in a form that follows a specific rule, such as the yaml format. The present invention may hold the object arrangement information not as character data but as metadata, for example, in a configuration that holds a scene graph with object IDs as nodes and their relative position relationships as edges. Thus, the object arrangement information is information generated based on the measurement information obtained by measuring the real space, representing two or more object type information and at least one or more position relationship information between the objects. The position relationship information may be an adverb representing the relationship between the object type names and their positions in the sentence. Also, the position relationship information may be the distance between objects or the direction in which objects are located as relative position information of the objects.
[0051] In Embodiment 1, an answer was generated based on the first object arrangement information, which is a sentence describing the arrangement relationship of objects included in the measurement information obtained by measuring the real space, and the arrangement relationship of objects included in the inquiry sentence. That is, even if the second object arrangement information, which is the arrangement relationship of objects included in the inquiry sentence, was not clearly generated, it was a configuration interpreted inside the neural network. On the other hand, a configuration can also be realized in which the second object arrangement information is explicitly generated from the inquiry sentence, and an answer is generated based on the first object arrangement information and the second object arrangement information. For example, for the inquiry sentence, it is held in yaml format (this is the second object arrangement information) using a neural network trained to output the arrangement relationship of objects in yaml format. Subsequently, the first object arrangement information and the second object arrangement information are input to the object arrangement characteristic database 103 to obtain an answer. By explicitly generating the second object arrangement information from the inquiry information in this way, the user can determine whether the prediction unit appropriately recognizes the arrangement relationship of the objects included in the inquiry information. Also, when it is not recognized, by correcting the inquiry sentence into a recognizable form and re-inputting it, the inquiry can be carried out with higher accuracy.
[0052] In Embodiment 1, an answer was generated based on the similarity of the arrangement relationship of objects. If an answer can be generated based on the degree of relevance between the event in the real space and the event in the inquiry information, in addition to the arrangement relationship, time series information (first time series information) can also be taken into account to generate an answer. Specifically, the object arrangement information generated from the time series image information is given the shooting time information for each, and is made into a sentence. For example, in the case of Embodiment 1, it is a sentence such as "A person took out a hammer from a bag." By doing so, since the change in the arrangement of objects can be grasped with higher accuracy, the accuracy of the answer is improved.
[0053] Furthermore, if time-series information (second time-series information) is also attached to the inquiry information, an answer may be generated based on the similarity between the time-series information in the real space and the time-series information included in the inquiry information. That is, assume that an inquiry sentence such as "Please send an alert as soon as possible before the suspicious person damages the car" is input. In this case, when an event such as "a person took out a hammer from a bag" is obtained in the real space rather than an event such as "a person with a hammer approaches the car", an answer that sends an alert can be output. In this way, the accuracy of the answer can be improved.
[0054] In addition to the object type information, the characteristic information of the object can also be held in the object arrangement information. The characteristic information of the object is information for more specifically identifying the object, such as the size, color, orientation, speed, etc. of the object. By providing such information, the object included in the inquiry sentence can be more accurately identified. For example, assume that an inquiry sentence such as "Please send an alert so that the suspicious person does not damage my red car" is input. In this case, when an event such as "a person is approaching a blue car" is obtained in the real space, it is possible to suppress accidentally sending an alert. In this way, the accuracy of the answer can be improved.
[0055] In Embodiment 1, the object arrangement characteristic database 103 was a neural network model using a Transformer. The object arrangement characteristic database 103 is not limited to this, and it is sufficient if it can generate an answer based on the arrangement relationship of objects. It may be a convolutional network, a fully connected network, an RCN, etc., and there is no particular limitation. Further, the object arrangement characteristic database 103 may be a Bayesian network, not limited to a neural network model. Also, the object arrangement characteristic database 103 may be a database that holds object characteristic group information. When using a database, it may be configured to return an answer similar to the positional relationship of the objects included in the input real-space object arrangement information and inquiry information from the past object arrangement information collected and registered in the object arrangement characteristic database 103. Using such a configuration, it can be realized with a smaller amount of calculation compared to a neural network.
[0056] In Embodiment 1, the similarity of the object arrangements was the attention value of the tokens representing two objects calculated using the weights held inside the neural network held by the object arrangement characteristic database 103. The similarity is not limited to this as long as it can represent whether the arrangement relationships of the objects are similar. For example, when two arrangement relationships are input, the difference in the distance between the objects and the difference in the direction in which another object is located with respect to a certain object in the two arrangement relationships can be used as the similarity. If it is held in a graph structure, for example, the graph shape can be used as the graph similarity using the Graph Edit Distance algorithm.
[0057] In Embodiment 1, the answer to the inquiry information was a sentence. That is, when the inquiry sentence contained the phrase "send an alert", if the arrangement relationship of the object met the inquiry conditions, an answer "send an alert" could be obtained. When it matches the sentence pattern of sending an alert in such an answer sentence, the alert sending unit can be configured to send an alert. Also, even if it is not configured to perform a predetermined operation when it matches the sentence pattern, the object arrangement characteristic database 103 may be configured to directly output a signal of 1 when performing a predetermined operation and 0 otherwise. Specifically, for example, a fully connected layer of a neural network is connected to the output layer of the object arrangement characteristic database 103. And it can be realized by learning so that when the inquiry information contains an instruction for a predetermined operation and the arrangement relationship in the real space matches the arrangement relationship included in the inquiry information, it becomes 1, and 0 otherwise. By doing so, when the conditions included in the inquiry information are met, a predetermined operation can be instructed to the information processing device 1. Here, the predetermined operation is not limited to sending an alert, and as long as it is configured to execute the operation included in the inquiry information, a configuration that drives a specific device or software via I / O H18, such as turning on a lamp or sending an email, can be realized.
[0058] If the inquiry information input unit 102 in Embodiment 1 is a keyboard and the answer output by the prediction unit 104 is displayed on a display as a display unit, it can be configured as a chat system for searching the real space. The input unit is not limited to a keyboard and can be arbitrary as long as it can input an inquiry sentence such as a touch display or voice input. The display unit is also not limited to a display and can be arbitrary as long as it can output an answer such as a projector or voice output.
[0059] <Embodiment 2> In Embodiment 1, an answer was generated based on one piece of inquiry information input by the user. On the other hand, with one piece of inquiry information, there may be cases where it is impossible to determine whether it is similar to the object arrangement information. That is, in this case, it is a case where the information is insufficient, such as when the arrangement relationship of the objects included in the inquiry information is insufficient or ambiguous. In Embodiment 2, a configuration for supplementing such insufficient information will be described. Specifically, a reliability is assigned to the words output by the object arrangement characteristic database, and when the reliability is lower than a predetermined value, a configuration for requesting input of more detailed inquiry information will be described.
[0060] The configuration diagram and processing flow of the information processing apparatus according to Embodiment 2 are the same as those in Embodiment 1. What is different between Embodiment 2 and Embodiment 1 is that in step S1005, the object arrangement characteristic database 103 predicts the reliability of the answer, and if the reliability is equal to or lower than a predetermined value, an answer is generated that requests input of more detailed inquiry information. Specifically, when calculating the score of each word candidate in step S1005, the prediction unit 104 determines that the arrangement relationship of the objects included in the inquiry sentence is insufficient if the score is equal to or lower than a predetermined value. Also, in such a case, an output for obtaining a more detailed arrangement relationship is generated. For example, assume that inquiry information such as "If there is a suspicious person, issue an alert" is input. In this case, when predicting the "..." part of "suspicious person, is,..." in the output of the object arrangement characteristic database 103, if there are multiple word candidates and it is not possible to narrow it down to one word, that is, if the reliability is low. In this case, an answer such as "Please input where the suspicious person is for issuing an alert" is output, which requests the arrangement relationship regarding the location where the suspicious person is located. That is, the answer output by the information processing apparatus 1 is a question requesting additional information. The user inputs, as inquiry information, insufficient information (third object arrangement information) such as "If a suspicious person enters within a range of 3 m from my car, issue an alert" in response to this answer (question requesting additional information). The information processing apparatus 1 obtains the arrangement relationship, which is the insufficient information of "within a range of 3 m from the car".
[0061] <Effect> In this modified example, when the reliability of the output of the object arrangement characteristic database 103 decreases, an answer is generated that prompts the input of more detailed object arrangement relationships into the inquiry information. By doing so, based on the object arrangement relationships in the real space, it becomes possible to generate an answer that more accurately matches the inquiry information.
[0062] In this modified example, the score assigned to the word candidates is regarded as the reliability. Even without such a configuration, as long as it is a configuration that generates an answer that causes more detailed object arrangement information to be input as the inquiry information when the inquiry information is ambiguous. For example, a fully connected layer learned to output a binary value indicating whether the inquiry information is ambiguous is connected to the object arrangement characteristic database 103, and when an ambiguous output is obtained, it may be configured to output an answer such as "Please input more details." By doing so, when the inquiry information is ambiguous, incorrect answers can be suppressed, and answers can be generated with higher accuracy.
[0063] Note that even if the reliability is not directly calculated, when the prediction means determines that the inquiry information is ambiguous as an internal state, it is also possible to output an answer that directly obtains the insufficient information. The object arrangement specific database 103 may be learned to request more detailed input when the inquiry information is ambiguous. That is, an object arrangement information and an inquiry information that cannot be answered from the object arrangement information are input, and a dataset that answers to obtain the insufficient information is prepared, and the object arrangement characteristic database 103 is learned. By doing so, an answer can be generated without directly calculating the reliability.
[0064] Furthermore, in the case where the inquiry information is ambiguous, when multiple events among the arrangement relationships of the objects in the real space match, a configuration may be adopted to generate a response that prompts the user to input one of them as the inquiry information. Specifically, when the difference between the score values of multiple candidates is equal to or less than a predetermined value, a response such as "Please select" is output to prompt the user to select which one of the two candidates it is. By doing so, when the inquiry information is ambiguous, the user can be made to select the target condition, and responses can be generated with higher accuracy.
[0065] Furthermore, a configuration may be adopted to prompt the user to newly input additional inquiry information as to whether the arrangement relationship of the objects included in the inquiry information recognized by the object arrangement characteristic database 103 is correct. Specifically, the arrangement relationship of the objects included in the inquiry information is replaced with another expression according to the object arrangement characteristics included in the object arrangement characteristic database 103 and then answered. Also, in addition to the inquiry information, the arrangement relationship of the objects included in the response replaced with another expression is held in the inquiry information holding unit as supplementary information to the inquiry information. For example, for an inquiry such as "If a suspicious person approaches the car, issue an alert", a search condition such as "When a suspicious person approaches within 1 m of the car, I will send you a notification email. Is that okay?" is generated as a response. At this time, the prediction unit 104 uses the object arrangement characteristics included in the object arrangement characteristic database 103 to predict that "the distance between the suspicious person and the car is less than 1 m" and that "an alert means sending a notification email". Then, it answers the user as to whether the predicted content is as intended by the inquiry. In this way, by prompting the user to answer whether the conditions specified by the user are as intended, it is possible to prevent the specification of incorrect conditions, and the search conditions in the real space can be registered more easily.
[0066] Also, it is possible to adopt a configuration in which additional inquiry information is newly requested from the user only when the reliability of the answer replaced with another expression according to the object arrangement characteristics included in the object arrangement characteristic database 103 is greater than a predetermined value. By doing so, additional inquiry information is requested from the user only when uncertain inquiry information is input, and the labor of inputting the user's inquiry information can be reduced.
[0067] <Embodiment 3> In Embodiment 1, the prediction unit 104 predicted based on the object arrangement information, which is a sentence representing the arrangement relationship of objects in the real space based on the measurement information previously measured by the measuring device and held by the object arrangement information holding unit. In Embodiment 3, a configuration will be described in which object arrangement information is generated based on measurement information obtained by measuring the real space in real time, and an inquiry is made with respect to the object arrangement information that is sequentially updated.
[0068] FIG. 9 is a diagram showing a search system for a real space including the information processing apparatus 2 according to Embodiment 3 of the present invention. The information processing apparatus 2 includes a measurement unit 201, a measurement information holding unit 202, an object arrangement information generation unit 203, an object arrangement information holding unit 204, and an inquiry information holding unit 205 in addition to the configuration of the information processing apparatus 1. Hereinafter, in the information processing apparatus 2, the configurations added to the information processing apparatus 1 will be described in detail. The same configurations as those of the information processing apparatus 1 are denoted by the same reference numerals and the description thereof is omitted.
[0069] The measurement unit 201 is a depth camera that acquires an image and a depth image as measurement information obtained by measuring the real space. The image and the depth image input from the depth camera are held by the measurement information holding unit 202. The measurement information holding unit 202 holds the image and the depth image as the measurement information measured by the measurement unit 201.
[0070] The object arrangement information generation unit 203 generates, as object arrangement information, an image held by the measurement information holding unit 202 and a sentence describing the object types and their arrangement relationships included in the depth image. The object arrangement information holding unit 204 holds the object arrangement information, which is the sentence generated by the object arrangement information generation unit 203. The object arrangement information holding unit 204 also outputs the held object arrangement information to the object arrangement information acquisition unit 101.
[0071] The inquiry information holding unit 205 holds a history of the inquiry information input by the inquiry information input unit 102. The inquiry information holding unit 205 also outputs the held inquiry information to the prediction unit 104.
[0072] FIG. 10 is a flowchart for explaining the operation of the information processing apparatus 2 according to Embodiment 3 of the present invention. The information processing apparatus 2 according to Embodiment 3 executes, in addition to the processing of Embodiment 1 shown in FIG. 4, measurement information input processing (step S201), object arrangement information generation processing (step S202), and inquiry information input determination processing (step S203). Hereinafter, the processing added to the information processing apparatus 1 in the information processing apparatus 2 will be described in detail. The same processing as that of the information processing apparatus 1 will be denoted by the same reference numerals and the description thereof will be omitted.
[0073] In step S201 following step S101, the measurement unit 201, which is a depth camera, inputs an image and a depth image as measurement information. The measurement unit 201 outputs the input image and depth image as measurement information to the measurement information holding unit 202. The measurement information holding unit 202 holds the measurement information from the measurement unit 201.
[0074] In step S202, the object arrangement information generation unit 203 generates three-dimensional shape data by SLAM. SLAM is an abbreviation for Simultaneous Localization and Mapping. Also in step S202, the object arrangement information generation unit 203, together, performs object labeling of pixels by semantic segmentation based on the input image. Also in step S202, the object arrangement information generation unit 203 assigns the object label to the three-dimensional shape data based on the generated three-dimensional shape data (three-dimensional shape model) and the object label. A detailed description of these series of processes is given in Non-Patent Document 5, and this is incorporated by reference. Subsequently, in step S202, the object arrangement information generation unit 203 verbalizes the generated three-dimensional shape data based on the relative positional relationship of the objects in the three-dimensional space. For verbalization, the method described in Non-Patent Document 6, which is a captioning method of a Transformer-based three-dimensional model that inputs the three-dimensional shape data and learns to generate a caption focusing on the relative positional relationship of the objects, is incorporated by reference. The object arrangement information generation unit 203 outputs, in this way, object arrangement information, which is a sentence describing the arrangement relationship of the objects generated based on the measurement information, to the object arrangement information holding unit 204. The object arrangement information holding unit 204 holds the object arrangement information from the object arrangement information generation unit 203. Following step S202, the process of step S102 is executed.
[0075] In step S203 following step S102, the inquiry information input unit 102 determines whether new inquiry information is input. If the inquiry information input unit 102 determines that new inquiry information is input, the process of step S103 is executed. If the inquiry information input unit 102 determines that no new inquiry information is input, the process of step S104 is executed.
[0076] In step S105, the information processing apparatus 100 performs an end determination. If the inquiry has not ended, it returns to step S201, and if the inquiry has ended, the process ends.
[0077] <Effect> As described above, based on the object arrangement information generated by the measurement unit in real time and the arrangement relationship of the objects included in the inquiry information, an answer regarding the similarity of the arrangement relationship of the objects is predicted. By doing so, it is possible to inquire about the ever-changing real space in the text without much effort, and the construction of the search system and the complexity of the inquiry can be reduced.
[0078] <Modification Example 3-1> In Embodiment 3, the measurement unit was a depth camera. However, in the present invention, the measurement unit is arbitrary as long as it is a sensor capable of acquiring objects in the real space and their positional relationships. The measurement result of the sensor is sensor information. For example, the measurement unit may be a stereo camera or a multi-camera. Furthermore, the measurement unit may be a 3D LiDAR that acquires a three-dimensional point cloud. To grasp the object types and their positional relationships when using 3D LiDAR, for example, the method of Charles et al., which is a neural network for recognizing object types from a three-dimensional point cloud (see Non-Patent Document 7), is used to assign object type labels to the point cloud. By doing so, it becomes possible to use the object arrangement information generated based on the measurement information measured by a plurality of measurement methods, and inquiries closer to the real space become possible.
[0079] When the output reliability in Embodiment 2 is low, assuming that there are deficiencies or errors in the object arrangement information generated based on the measurement information, the text generation from the measurement information according to Embodiment 3 may be redone. Specifically, when the reliability predicted by the prediction unit 104 is lower than a predetermined value, parameters related to object detection are adjusted, such as lowering the threshold so that more objects can be detected from the image information. Next, the detected objects with the changed parameters are input and converted into text. By doing so, even when there are excesses or deficiencies in the once-generated object arrangement information, the object arrangement information can be regenerated again, and it becomes possible to answer the inquiry information with higher accuracy.
[0080] In Embodiment 3, a sentence describing the arrangement relationship of objects included in image information or three-dimensional shape information is generated as object arrangement information, and the prediction unit 104 predicts an answer based on the generated sentence (object arrangement information) and the query sentence. For example, if a neural network is trained to input three-dimensional shape data and a query sentence and generate an answer, the prediction unit 104 can predict the answer without generating a sentence as object arrangement information. For such learning, the method of Shuquan et al. (see Non-Patent Document 8) can be used, and thus detailed description is omitted.
[0081] In Embodiment 3, it was a configuration to update both the object arrangement information based on the measurement information of the real space and the query information. These updates may be only one of them. That is, a configuration can also be realized in which query information is registered in advance and an answer is generated each time the object arrangement information is updated. Conversely, a configuration can also be realized in which a user inputs a plurality of query information for the once-generated object arrangement information and updates the query information.
[0082] <Modification Example 3-2> In Embodiment 1, a method of applying this embodiment to a monitoring system that queries the monitoring conditions of the real space in a sentence based on the object arrangement information of the real space was described. If a user can search for desired conditions based on the object arrangement information of the real space, the task of querying the real space is not limited to the monitoring task.
[0083] For example, Embodiment 1 can also be applied to the work analysis of the device assembly process in a factory. That is, the time-series positional relationship of the hands, tools, and parts shown in the video of the operator's hand acquired by the measurement unit is retained as object arrangement information and used for the analysis of work procedures and work differences for each operator. Object detection is performed on the image obtained from the camera that captures the operator's hand to obtain the relative position and orientation of the object. By associating this relative position and orientation with the shooting time when the camera captures the image, the object and their time-series positional relationship are obtained. This is put into words and retained. Subsequently, an inquiry sentence is input, and the prediction unit 104 generates an answer using the object arrangement characteristic database 103. Specifically, assume that a sentence is input as inquiry information. The input sentence is "The correct work procedure is to first hold the tweezers in the left hand and hold the part in the right hand. Subsequently, hold the part with the tip of the tweezers and attach it to the device in front. If there is a variation in the work time of the operator in this process, please tell the reason." In this case, the prediction unit 104 gives an answer such as "The work time of operator B is slow. The reason for the slowness is that the part is pinched against the belly of the tweezers, so the part is likely to fall when attaching it to the device." That is, for the "object arrangement relationship in which the part is located at the tip of the tweezers" included in the inquiry information, the "object arrangement relationship in which the part is located at the belly of the tweezers" measured in the real space is recognized, and the difference in the work procedure is extracted.
[0084] In this way, by comparing the object arrangement relationship in the real space with the object arrangement relationship included in the inquiry information, this information processing apparatus can also be applied to the inquiry task in work analysis.
[0085] Furthermore, the present invention can also be applied to the task of searching for the location of an article. For example, in a logistics warehouse, the incoming articles are photographed (measured by the measuring unit) using a surveillance camera or a camera arranged on a moving body such as a robot. The positional relationship of the articles included in the measurement information is associated with the time as object arrangement information and written and held. Subsequently, the user inputs an inquiry sentence, and the prediction unit 104 generates an answer using the object arrangement characteristic database 103. For example, it is assumed that a sentence such as "There should be 10 articles A on shelf B, but there are only 9. Where is the remaining 1 article?" is input as the inquiry information. The prediction unit 104 gives an answer such as "Article A fell during transportation from point C to shelf B. Article A is on the passage from point C to shelf B." That is, for the "time-series position information of article A in the real space", the "position information of article A not on shelf B" included in the inquiry information is searched. Thus, the present invention can also be applied to management tasks in logistics.
[0086] In this modification example, an example of using the present invention for work analysis and logistics management has been described. Thus, as long as it is a task of associating and searching for the object arrangement relationship in the real space and the object arrangement relationship included in the inquiry information, the present invention can be applied not only to surveillance, work analysis, and logistics management tasks but also to any task.
[0087] The inquiry information holding unit 205 can be configured to hold inquiry information of a plurality of tasks instead of holding inquiry information of one task. Specifically, as the inquiry information, there are inquiries such as "Inquiry 1: Issue an alert if a suspicious person approaches the vehicle. Inquiry 2: Send an email to the security guard if there is a lost item in the parking lot." If the inquiry information holding unit 205 holds the inquiry information of a plurality of tasks in this way, the prediction unit 104 predicts an answer for each inquiry, so that a plurality of search tasks can be simultaneously executed by one information processing device.
[0088] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0089] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.
[0090] The disclosure of this embodiment includes the following configurations. (Configuration 1) Object arrangement information acquisition means for acquiring object arrangement information including the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space, Inquiry information input means for inputting inquiry information from a user, Prediction means for inputting the object arrangement information and the inquiry information, and predicting an answer to the inquiry information using an object arrangement characteristic database that holds object arrangement characteristics representing the positional relationship of a plurality of objects. An information processing apparatus characterized by comprising: (Configuration 2) The object arrangement information acquired by the object arrangement information acquisition means is first object arrangement information, The arrangement relationship of the object included in the inquiry information is second object arrangement information, The prediction means predicts an answer to the inquiry information based on the first object arrangement information and the second object arrangement information using the object arrangement characteristic database. The information processing apparatus according to Configuration 1, characterized by: (Configuration 3) The prediction means predicts an answer regarding the similarity between the first object arrangement information and the second object arrangement information. The information processing apparatus according to Configuration 2, characterized by: (Configuration 4) The prediction means generates third object arrangement information related to the inquiry information that is not included in the second object arrangement information, using the object arrangement characteristic database, predicts an answer to the inquiry information based on the similarity of the first object arrangement information, the second object arrangement information, and the third object arrangement information The information processing apparatus according to Configuration 2 or Configuration 3, characterized in that (Configuration 5) The first object arrangement information further includes first time-series information consisting of time-series positional relationship information of the object, The prediction means generates, using the object arrangement characteristic database, second time-series information consisting of time-series positional relationship information of the object included in the inquiry information, in association with the second object arrangement information, predicts an answer to the inquiry information based on the similarity between the first object arrangement information and the second object arrangement information and the similarity between the first and second time-series information The information processing apparatus according to any one of Configurations 2 to 4, characterized in that (Configuration 6) The information processing apparatus further includes inquiry information holding means for holding a plurality of pieces of inquiry information, The second object arrangement information is the arrangement relationship of the objects included in the plurality of pieces of inquiry information held by the inquiry information holding means The information processing apparatus according to any one of Configurations 2 to 5, characterized in that (Configuration 7) The prediction means extracts deficiency information not included in the inquiry information using the object arrangement characteristic database, and predicts an answer that requests the inquiry information for supplementing the deficiency information The information processing apparatus according to any one of Configurations 2 to 6, characterized in that (Configuration 8) The prediction means generates, in association with the answer, a reliability degree of the answer to the inquiry information, When the reliability is lower than a predetermined value, predict an answer that requests inquiry information for supplementing the second object arrangement information so as to supplement the missing information not included in the inquiry information. The information processing apparatus according to Configuration 7, characterized by the above. (Configuration 9) Measurement information holding means for holding measurement information measured by a sensor, Object arrangement information generation means for generating the object arrangement information as a sentence based on the measurement information, Object arrangement information holding means for holding the object arrangement information generated by the object arrangement information generation means, further comprising The object arrangement information acquisition means acquires the object arrangement information held by the object arrangement information holding means acquires The information processing apparatus according to any one of Configurations 1 to 8, characterized by the above. (Configuration 10) The object arrangement information is text information representing, in text, the object type and arrangement relationship of an object generated based on measurement information obtained by measuring the real space. The information processing apparatus according to any one of Configurations 1 to 9, characterized by the above. (Configuration 11) The object arrangement characteristic database is text information representing, in text, the object type and arrangement relationship of an object generated based on measurement information obtained by measuring the real space, inquiry text information obtained by inputting an inquiry from a user as inquiry information in text, and is a neural network that interprets natural language, which is learned to input the above and output an answer related to the arrangement relationship. The information processing apparatus according to any one of Configurations 1 to 10, characterized by the above. (Method 1) An object arrangement information acquisition step of acquiring object arrangement information including the object type and arrangement relationship of an object generated based on measurement information obtained by measuring the real space, An inquiry information input step of inputting inquiry information from a user, A prediction step of inputting the object arrangement information and the inquiry information and predicting an answer to the inquiry information using an object arrangement characteristic database that holds object arrangement characteristics representing the positional relationships of a plurality of objects; A method characterized by comprising the above. (Program 1) A computer is An object arrangement information acquisition means for acquiring object arrangement information including the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space; An inquiry information input means for inputting inquiry information from a user; An object arrangement characteristic database that holds object arrangement characteristics representing the positional relationships of a plurality of objects, and A prediction means for inputting the object arrangement information and the inquiry information and predicting an answer to the inquiry information using the object arrangement characteristic database; A program characterized by causing the computer to function as the above.
Explanation of Signs
[0091] 101: Object arrangement information acquisition unit, 102: Inquiry information input unit, 103: Object arrangement characteristic database, 104: Prediction means
Claims
1. An object arrangement information acquisition means for acquiring object arrangement information including the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space; An inquiry information input means for inputting inquiry information from a user; A prediction means for inputting the object arrangement information and the inquiry information and predicting an answer to the inquiry information using an object arrangement characteristic database that holds object arrangement characteristics representing the positional relationship of a plurality of objects; An information processing apparatus characterized by comprising the above.
2. The object arrangement information acquired by the object arrangement information acquisition means is first object arrangement information, The arrangement relationship of the object included in the inquiry information is second object arrangement information, The prediction means predicts an answer to the inquiry information based on the first object arrangement information and the second object arrangement information using the object arrangement characteristic database. The information processing apparatus according to claim 1, characterized by the above.
3. The prediction means predicts an answer regarding the similarity between the first object arrangement information and the second object arrangement information. The information processing apparatus according to claim 2, characterized by the above.
4. The prediction means generates third object arrangement information related to the inquiry information that is not included in the second object arrangement information using the object arrangement characteristic database, and predicts an answer to the inquiry information based on the similarity between the first object arrangement information, the second object arrangement information, and the third object arrangement information. The information processing apparatus according to claim 2, characterized by the above.
5. The first object arrangement information further includes first time-series information consisting of time-series positional relationship information of the object, The prediction means generates second time-series information consisting of time-series positional relationship information of the object included in the inquiry information in association with the second object arrangement information using the object arrangement characteristic database, and predicts an answer to the inquiry information based on the similarity between the first object arrangement information and the second object arrangement information and the similarity between the first and second time-series information. The information processing apparatus according to claim 2, characterized by the above.
6. Further comprising an inquiry information holding means for holding a plurality of pieces of inquiry information, The second object arrangement information is the arrangement relationship of the objects included in the plurality of pieces of inquiry information held by the inquiry information holding means. The information processing apparatus according to claim 2, characterized by the above.
7. The prediction means extracts shortage information not included in the inquiry information using the object arrangement characteristic database, and predicts an answer that requests the inquiry information for supplementing the shortage information. The information processing apparatus according to claim 2, characterized in that.
8. The prediction means generates and associates a reliability degree of an answer to the inquiry information with the answer, and when the reliability degree is lower than a predetermined value, predicts an answer that requests the inquiry information for supplementing the second object arrangement information so as to supplement the shortage information not included in the inquiry information. The information processing apparatus according to claim 7, characterized in that.
9. Measurement information holding means for holding measurement information measured by a sensor; Object arrangement information generation means for generating the object arrangement information as a sentence based on the measurement information; Object arrangement information holding means for holding the object arrangement information generated by the object arrangement information generation means; further comprising: The object arrangement information acquisition means acquires the object arrangement information held by the object arrangement information holding means acquires The information processing apparatus according to claim 1, characterized in that.
10. The object arrangement information is text information representing the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space, in a sentence. The information processing apparatus according to claim 1, characterized in that.
11. The object arrangement characteristic database is text information representing the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space, in a sentence; inquiry text information obtained by inputting an inquiry from a user as inquiry information in a sentence; a neural network that interprets natural language, which is input with the above and is learned to output an answer related to the arrangement relationship. The information processing apparatus according to claim 1, characterized in that.
12. An object arrangement information acquisition step of acquiring object arrangement information including the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space; An inquiry information input step of inputting inquiry information from a user; A prediction step of inputting the object arrangement information and the inquiry information, and predicting an answer to the inquiry information using an object arrangement characteristic database that holds object arrangement characteristics representing the positional relationship of a plurality of objects. A method characterized by comprising.
13. A computer An object arrangement information acquisition means for acquiring object arrangement information including the object type and arrangement relationship of an object, generated based on measurement information obtained by measuring the real space, An inquiry information input means for inputting inquiry information from a user, An object arrangement characteristic database that holds object arrangement characteristics representing the positional relationships of a plurality of objects, and A prediction means that inputs the object arrangement information and the inquiry information and predicts an answer to the inquiry information using the object arrangement characteristic database, A program characterized by being caused to function as such.
Citation Information
Patent Citations
Work analysis device and work analysis method
JP2023041969A
Monitoring system and monitoring method
JP7111422B2