Artificial intelligence systems, artificial intelligence programs, and natural language processing systems
By constructing an artificial intelligence system with a first platform and a second platform, the problem of not being able to understand the meaning of sentences in existing technologies has been solved, enabling natural communication and emotion simulation, and possessing the same psychological and social integration capabilities as humans.
Patent Information
- Application Number
- CN202180007651.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-14
- Filing Date
- 2021-02-09
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2041-02-09
AI Technical Summary
Existing artificial intelligence and natural language processing systems cannot understand the meaning of sentences, resulting in an inability to engage in natural conversations and exchanges, to simulate the thought process of others, and to integrate into human society.
By constructing an artificial intelligence system with a first platform and a second platform, using data models to generate and simulate human objects, and combining this with a natural language processing system to parse the meaning of sentences, the system can achieve the recognition and output of the external world.
It achieves the understanding of the meaning of natural language, can communicate naturally with people, simulate the other person's thought process, integrate into human society, and possess the same psychological and emotional judgment abilities as humans.
Smart Images

Figure CN114902236B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an artificial intelligence system, an artificial intelligence program, and a natural language processing system.
[0002] This application claims priority based on International Application PCT / JP2020 / 005696 filed on February 14, 2020, and incorporates all the recitations described in said International Application. BACKGROUND
[0003] Nowadays, artificial intelligence has an intelligence exceeding that of humans if limited to specific fields such as Go and chess. However, even if limited to Go and chess, it is not very practical in real life. The artificial intelligence we expect is not specialized in limited fields, but an artificial intelligence that can do anything like a human being. Such an artificial intelligence is called general artificial intelligence.
[0004] As an artificial intelligence like a human being, one of them is an artificial intelligence that can naturally communicate with humans, that is, can naturally converse.
[0005] PRIOR ART DOCUMENTS
[0006] PATENT DOCUMENTS
[0007] Patent Document 1: Japanese Patent Application Publication No. 2006-178063 SUMMARY OF THE INVENTION
[0008] (I) PROBLEMS TO BE SOLVED BY THE INVENTION
[0009] Patent Document 1 discloses a conversation system that makes a corresponding answer if an emotion appears. However, this system only responds to words, and only that cannot be called a conversation.
[0010] In addition, there are AI speakers, applications of smartphones, and the like that can have a conversation, but all of them can only reproduce a scenario prepared in advance, and cannot have a natural conversation or chat. Humans can have a natural conversation because they can perceive, empathize with, and understand the intentions of the other party, that is, understand the unexpressed meaning, and have many considerations in their heads such as if they say this, the other party will think this, and have a conversation. Only by saying a scenario prepared in advance, a natural conversation cannot be made.
[0011] An object of the present application is to provide an artificial intelligence system that can understand the feelings of the other party and can behave as having the same psychology as a person integrated into society.
[0012] (II) TECHNICAL SOLUTION
[0013] An artificial intelligence system of the present application determines an output to an outside based on information inputted from the outside, the artificial intelligence system having: a storage section that pre-stores a data model that imitates a person and a person's thinking; a generation section that extracts the data model from the storage section and generates a person object that can reproduce a person's action and thinking; a world building section that has a first platform and a second platform that arrange the person object and builds a world in which the person object's action and thinking are developed; an outside world reproducing section that arranges the person object on the first platform based on information inputted from the outside and reproduces an outside world; and an output determining section that grasps a situation of the outside by recognizing the outside world reproduced on the first platform and determines an output to the outside by arranging the person object on the second platform and operating it.
[0014] Thus, the output determining section recognizes the outside world from the first platform. That is, since a person in the outside is recognized as the person object, the thinking content of the person that cannot be acquired by a camera or the like can be grasped. In addition, using the second platform, it is possible to simulate what the other party would think if such an action is taken, and it is possible to make a natural response.
[0015] In the above artificial intelligence system, it can be that person objects of oneself and the other party are arranged on the first platform and the second platform. It can be that the output determining section determines the output in such a manner that the thinking of the person object of the other party is felt to be satisfactory. Thus, it is possible to make a response that takes the other party into consideration, and it is possible to realize an artificial intelligence that behaves as a person having the same psychology as a person.
[0016] In the above artificial intelligence system, it can be that the person object arranged on the first platform has a first platform on a lower layer side of the first platform of the person. It can be that the world building section reproduces the outside world on the first platform on the lower layer side based on information inputted to the person. Thus, it is possible to think from the standpoint of the other party, and it is possible to have a more natural exchange with a person.
[0017] In the above artificial intelligence system, it can be that the data model of the person has two kinds of desires, i.e., a low-level desire that is generated by a body and a high-level desire that pursues something with a high social value. It can be that the output determining section determines the output in such a manner that the low-level desire is suppressed and the high-level desire is satisfied. Thus, it is possible to judge "should ~", right and wrong, it is possible to have a natural conversation with a person, and it is possible to realize an artificial intelligence that is integrated into a human society.
[0018] In addition, the artificial intelligence program of the present application determines an output to the outside based on information input from the outside, so that the computer functions as a generation section that generates a person object capable of reproducing the actions and thoughts of a person in accordance with a data model that imitates a person and the thoughts of the person, a world building section that has a first platform and a second platform on which the person object is arranged and builds a world in which the actions and thoughts of the person object are developed, an external world reproduction section that arranges the person object on the first platform and reproduces the external world based on information input from the outside, and an output determination section that grasps the external world by recognizing the external world reproduced on the first platform and determines an output to the outside by arranging the person object on the second platform and operating it.
[0019] Thus, the output determination section recognizes the external world in accordance with the first platform. That is, since a person in the outside is recognized as a person object, the thoughts of the person, which cannot be acquired by a camera or the like, can also be recognized. In addition, using the second platform, it is possible to simulate what the other person would think if the person acted like that and the like, and a natural response can be made.
[0020] Here, the words of a person are referred to as natural language, and the field of artificial intelligence that handles natural language is referred to as natural language processing. The biggest problem of natural language processing is that the meaning of a sentence cannot be understood. Therefore, in a conversational AI, only a scenario prepared in advance is spoken, and a natural conversation cannot be made.
[0021] Therefore, another object of the present application is to provide a method of understanding the meaning of natural language.
[0022] The natural language processing system of the present application has an input device that inputs a sentence of natural language, a storage device that stores an object that represents a person or thing, and a control device that decomposes an input sentence from the input device into words and analyzes the meaning, characterized in that the storage device stores the object and the name of the object in association with each other, and the control device generates the object based on the words of the input sentence and the storage device, and changes the object based on the words of the input sentence, thereby analyzing the meaning.
[0023] By extracting words from a string of natural language and changing the attributes of the object, it is possible to reproduce a state close to an actual person or thing using a computer. That is, the object can realize the same attributes and activities as an actual person or thing within the computer. This can be said to be a meaning understanding of a sentence of natural language.
[0024] (Three) Advantages
[0025] Such an artificial intelligence system can understand the feelings of the other person, can behave as having the same psychology as a person, and can be integrated into human society. Attached Figure Description
[0026] Figure 1 This is a block diagram showing the structure of the artificial intelligence system in Implementation 1.
[0027] Figure 2 This is a conceptual diagram showing the structure of the first platform and the second platform of the artificial intelligence system according to Embodiment 1 of this disclosure.
[0028] Figure 3 This is a flowchart illustrating typical steps when using the artificial intelligence system disclosed herein for processing.
[0029] Figure 4 This is a conceptual diagram representing a second platform where a robot is deployed in a room.
[0030] Figure 5 This is a conceptual diagram representing a second platform where a robot is deployed in a room.
[0031] Figure 6 This is a conceptual diagram representing a second platform where a robot is deployed in a room.
[0032] Figure 7 This is a flowchart illustrating typical steps when using the artificial intelligence system disclosed herein for processing.
[0033] Figure 8 It is the table used in the good and evil judgment procedure.
[0034] Figure 9 It is a table used in the social judgment of good and evil in the ought-to-be determination procedure.
[0035] Figure 10 It is a table used in the determination of personal conduct in the ought-to-be determination procedure.
[0036] Figure 11 This is a conceptual diagram illustrating the structure of an artificial intelligence system according to another embodiment of the present disclosure. Detailed Implementation
[0037] (Details of the embodiments of the invention in this application)
[0038] Next, an embodiment of the artificial intelligence system of this disclosure will be described with reference to the accompanying drawings. In the following drawings, the same or equivalent parts will be given the same reference numerals and will not be described repeatedly.
[0039] (Implementation Method 1)
[0040] The structure of the artificial intelligence system according to Embodiment 1 of this disclosure will be described. Figure 1 This is a block diagram showing the structure of the artificial intelligence system in Implementation 1.
[0041] The artificial intelligence system 11 of Embodiment 1 is applied to a robot having a control device 12. As a structure of the robot, there is a body identical to that of a human. That is, there are portions corresponding to hands and feet of a human, and there are motors 13 that drive them. In addition, there is a camera 14 that corresponds to eyes of a human, and video data is acquired by taking the outside with the camera 14. In addition, there is a microphone (a member having a sound collecting function) 15 that corresponds to ears of a human, and sounds from the outside are heard with the microphone 15. In addition, there is a speaker 16 that corresponds to a mouth of a human, and it is possible to make a conversation by emitting sounds with the speaker 16. An artificial intelligence program is executed by the control device 12 that controls each of the above-mentioned portions. The artificial intelligence program controls the motors 13 that move the hands and feet, and the speaker 16 that speaks, on the basis of external information acquired from the camera 14 and the microphone 15. That is, the artificial intelligence program determines the action of the robot body on the basis of information from the outside. Furthermore, the control device 12 has a CPU (Central Processing Unit) and a main memory in which the artificial intelligence program of the present disclosure is loaded.
[0042] The artificial intelligence system 11 is an artificial intelligence system that determines an output to the outside on the basis of information input from the outside, and has a control device 12 including a control section 18, and a database 17 that functions as a storage section that stores data models of a human, an object, and the like in advance. The control section 18 determines an output to the outside, such as the motors 13 and the speaker 16, on the basis of information input from the camera 14 and the microphone 15.
[0043] The control section 18 has a generation section 24 that generates an object in accordance with a data model, a world building section 25, an external world reproduction section 26, and an output determination section 27 that determines an output.
[0044] The data model represents a human, an object, a concept, and the like, and is stored in the database 17. When the control section 18 recognizes a human or an object on the basis of external information from the camera 14 and the microphone 15, the corresponding data model is extracted with the external world reproduction section 26, an object is generated with the generation section 24, and a world is built by arranging the object in the world building section 25. The object is the same as an object of an object-oriented language, and can be freely manipulated, and is generated in a memory.
[0045] The artificial intelligence system 11 is a system capable of having a natural conversation, dialogue with a person. Therefore, it has a function of inputting a voice of a person who is conversing from the microphone 15 and converting the voice data into a natural language string (text data) in real time. Further, the input is not limited to the conversational voice from the microphone 15, but can be a natural language sentence acquired by the camera 14. In this case, it is converted into a natural language string by image character recognition. Also, the input is not limited to a sentence, but there is a case where a situation is read and understood from a scene in front of the camera 14, and this case will be described later.
[0046] In addition, the artificial intelligence system 11 as a natural language processing system is not limited to a conversation, dialogue with a person, but can be widely applied to a case where the meaning of a natural language is understood such as summarization of a sentence, machine translation.
[0047] Here, a method of understanding the meaning of a natural language according to the present application is described. The greatest feature of the meaning understanding according to the present application is that a person or thing existing in the real world is expressed as an object. The object is an object of an object-oriented language, and can be said to be a model simulating a person or thing existing in the real world.
[0048] The object has a property, an attribute expressing a property of the thing, and a method expressing an activity of the thing. For example, in the case of an apple object, the shape attribute is round, spherical, and the color attribute is red. In addition, the thing also has a position attribute. In the case of a person object, for example, there are a name attribute, a gender, and a residence attribute. In addition, as the method, there are walking, running, eating, and the like. The object is written as a class in a program and generated as an object in a heap area of a memory. These objects are managed in a word dictionary corresponding to a word of a natural language. The word dictionary is stored in the database 17 or a storage device such as a hard disk drive (HDD) or a main memory. For example, the apple class is managed in the word dictionary corresponding to the word "apple".
[0049] The object expresses a person or thing, and is arranged in a world or a space corresponding to the person or thing. If one easy-to-understand example is given, it is a three-dimensional space. The three-dimensional space is a space generated by virtually simulating a three-dimensional space of the real world in a computer, for example, by three-dimensional computer graphics (3DCG). Further, it can be only a wire frame.
[0050] The understanding of the meaning of natural language will be described from here. Assume that a natural language sentence such as "There is an apple on the table" is input. First, the input sentence is broken down into words by morphological analysis. Thus, it is broken down into "table / on / there / is / apple / ." Here, "table" is taken out, and the word dictionary is searched to find, generating a table object. Similarly, an apple object is also generated. This is the understanding of the meaning of natural language words such as "table" and "apple".
[0051] Next, the verb "there is" of "There is a table" and "There is an apple" is the meaning of existence. Therefore, they are arranged in a three-dimensional space. This is the understanding of the meaning of the natural language verb "there is".
[0052] The word "on" refers to the direction opposite to the direction of gravity. By 3DCG, gravity can also be simulated, and its direction can be set, so the apple object is arranged on the table object arranged in the three-dimensional space. This is the understanding of the meaning of "on" in natural language. In this way, the meaning of natural language is understood.
[0053] In this way, in the three-dimensional space of 3DCG, the state of the apple being placed on the table is reproduced. This can be said to be the same as the image that appears in the mind when a person reads a sentence such as "There is an apple on the table." Next, a sentence such as "Pick up the apple upward" is input. When the meaning of this sentence is understood, first, it is determined that "apple" is the apple on the table in the current state. Then, since it is to be "picked up upward", the apple is moved upward in the three-dimensional space of 3DCG. This is the understanding of the meaning of "pick up".
[0054] If this is an instruction to the robot "Pick up the apple upward", the robot understands the meaning and picks up the apple on the table in front of it upward. This means that a robot that understands the meaning of natural language can communicate naturally with a person.
[0055] Next, the above is the understanding of the meaning of a sentence. A sentence is a series of multiple sentences. In the first sentence, a scene in which there is an apple on the table is generated. First, the objects and characters that appear in the scene are generated to generate the scene. In the subsequent sentences, the objects that appear are operated to move them. In this way, it can be said that the understanding of the meaning of the article can be expressed by changing the scene by setting the scene and the objects that appear in the scene. Then, by storing the scene before the change and the scene after the change, it is possible to remember the history so far. This corresponds to the episode memory in a person's memory. The scenes are saved in the order in which the events occurred. The function of maintaining the order of saving the scene corresponds to time.
[0056] Here, an attempt is made to consider the relationship of the three-dimensional space of 3DCG and the apple. This can be said to be the relationship of the world and the object. The object configuration in the three-dimensional space is the position characteristic of the object. It can be said that the relationship of the object and the world is determined by the position characteristic. Also, the method of "pick up" indicates moving upward along the Z coordinate of the representation of the three-dimensional space. That is, assuming that the position characteristic of the object is set using the coordinate axes of X, Y, and Z, it can be said that pick up means changing the Z coordinate. Further, it can be said that the verb of move means changing the position characteristic of the object. In this way, it can be said that the method is to change the characteristic, and it can be said that the verb of the natural language corresponds to the method. That is, the meaning understanding of the natural language is to generate an object as an object, and the verb is to change the characteristic of the object.
[0057] In addition, the relationship of the objects to each other can also be defined using the characteristic, and the relationship of the table and the apple can be said to be separated by the distance calculated from the difference in the respective position characteristics. In this way, the relative relationship between the objects can be determined using the characteristic.
[0058] Next, the meaning of "get the apple" is considered. In this case, not the three-dimensional world but the possession space is considered. The possession space is a space that represents the meaning of having an object. The object configured in this world has a characteristic called possession. For example, in the case of "A got the apple from B," A and B are configured in the possession space, and initially there is an apple in the possession characteristic of B. Next, if the verb "get" is made to act, the apple existing in the possession characteristic of B is transferred to the possession characteristic of A. This is the meaning of "get." From the perspective of A, the event is "give." In this way, the meanings of "give" and "get" are realized by the human object and the possession space.
[0059] In this way, the meaning of the sentence of the natural language can be expressed by the world, or the space, the object configured in the space, and the change in the characteristic of the object. Also, it can be said that this is the meaning understanding of the sentence of the natural language. The space referred to here is not limited to the three-dimensional space. The three-dimensional space is defined by the position coordinates, and the possession space is defined by who has what. The space referred to here is a space that manages the attributes of each object configured. The three-dimensional space is a space that manages objects using absolute positions, and the possession space is a space that manages the possession relationship between objects.
[0060] Also, as another space, an attempt is made to consider the family relationship space. In this case, the relative relationship of the characters is expressed. Is the other person a father or a son with respect to oneself? In this case, the relative relationship of the other person with respect to oneself is set in the character relationship characteristic of the human object, and the relationship of oneself changes depending on the other person. This is the family relationship space. According to this, the meaning of "become a parent" can be understood as the meaning of having a child.
[0061] Thus, the meaning of so-called natural language can be said to be that an object is generated, and an appropriate property of the object is set, or an operation such as a change is performed. Also, the property can be said to belong to a certain space or world.
[0062] As shown in FIG. 1, the world building unit 25 has a first platform 21 and a second platform 22. On the first platform 21, the world building unit 25 reproduces the outside world itself as a virtual world. For example, if the outside world is set to be a room, a three-dimensional world is set on the first platform 21. Also, if it is recognized as a table by acquisition with the camera 14, a data model of the table is taken out from the database 17, the table object is generated with the generation unit 24, and the table object is arranged in the three-dimensional world, that is, the first platform 21. Figure 2 Thus, the room acquired with the camera 14 is reproduced on the first platform 21. The table object is a three-dimensional object, and thus can move freely in the room. This can be said to be the same as a table in the real world. That is, a situation in which operations can be performed freely as a human imagines in the mind is built. Also, the output determination unit 27 recognizes the real world through the outside world built on the first platform 21. That is, the output determination unit 27 recognizes the outside world built on the first platform 21 as the real world itself.
[0063] The second platform 22 of the world building unit 25 is built and operated by the output determination unit 27. Operations that can be performed on the object arranged on the second platform 22 can be said to be simulation using the second platform 22. That is, the output determination unit 27 can most appropriately determine the output by simulating using the second platform 22. In other words, this can be said that the output determination unit 27 has a function as "consciousness" that recognizes the world developed on the first platform 21 as the real world, and determines the action by trial and error using the second platform 22.
[0064] The object arranged on the platforms 21, 22 exists as a subject such as a person, and also as a thing other than the subject. The difference is whether it has so-called "psychology". The psychology refers to having emotions such as joy and sadness, and having a function equivalent to the artificial intelligence program of the present disclosure. The subject has psychology, and is not limited to an actual person, but also includes characters appearing in movies and novels, roles that do not exist in reality, gods, demons, and the like. The world developed in the first platform 21 is not limited to the real world, but also a world seen in a movie, a world imagined by reading a novel. As a characteristic, it is generated based on information from the outside, and the output determination unit 27 cannot directly change it.
[0065]
[0066] Next, a processing method in an artificial intelligence program using a humanoid robot will be described. The robot is a humanoid robot having hands and feet, and is capable of walking and holding objects. In addition, the robot has a camera, a microphone, and a speaker, and is capable of communicating with a person through voice. The above components are controlled by a control device 12 that controls the robot.
[0067] Figure 3 is a flowchart showing typical procedures when processing is performed by the artificial intelligence system 11 of the present disclosure. Figures 4 to 6 is a conceptual diagram showing that the robot 31 is disposed on a second floor of the room 30.
[0068] As shown in Figure 4 , in the room 30 in which the robot 31 is disposed, a shelf 33 is installed on a wall 32. Also, it is assumed that a battery 34 is placed on the shelf 33. In addition, it is assumed that a chair 35 is disposed in the room 30.
[0069] Next, referring to Figure 3 , how the robot 31 recognizes the situation will be described. The robot 31 acquires the situation of the room 30 using the camera, and analyzes the image to recognize the shelf 33, the battery 34, the chair 35, and the like. That is, the situation outside is obtained using a sensor (S11). A plurality of data models are stored in a database, and the recognized object is taken out from the database 17, and an object of the object is generated (S12). The object has data such as shape and size as three-dimensional data in the case of an object, for example. Also, information such as attributes and functions of the object, such as color and weight, is associated. That is, it can be said that the artificial intelligence can understand the meaning. By understanding the meaning, it means that, for example, when "the height of the table" is specified, it is understood that it corresponds to the height data of the table object.
[0070] The data model is realized by a class in an object-oriented language. The camera 14, which corresponds to the eyes of the robot 31, photographs the real world in front of the eyes, and converts it into three-dimensional data in real time through image analysis, and recognizes the object based on its shape. For example, when it is determined to be a "chair", a chair class is called up, and a chair object is generated. The chair object is directly recognized by the robot 31. As to the chair class, it has legs, a seat surface, and the like as components of the chair 35, and also has three-dimensional data thereof. Also, the components such as the legs and the seat surface are set in a manner consistent with the three-dimensional data of the recognized chair 35. That is, the outside world is reproduced based on the acquired data (S13).
[0071] The object has attributes such as material, color, weight, hardness, and the like. For the color of the chair 35 acquired by image analysis, a color attribute is set in the chair object. Similarly, when the material of the chair 35 is judged to be wood from the data of image analysis, the material is set to wood in the material attribute of the chair object. In the data model of wood, data such as weight, hardness, and the like are recorded, and thus the weight, hardness, and the like of the chair object are set from them. Only the image can be acquired directly with the camera 14, but it is also possible to recognize data such as weight, hardness, and the like that cannot be measured directly.
[0072] The object recognized with the camera 14 exists in the three-dimensional space of the real world. Therefore, the object generated from the data model also needs to be arranged in the three-dimensional space. Here, the three-dimensional space in which the object is arranged is referred to as a three-dimensional virtual world with respect to the three-dimensional world of the outside reality. The three-dimensional virtual world is constructed on the first platform 21 and the second platform 22 of the world construction unit 25, and can have a plurality of objects. The three-dimensional virtual world also has functions of operating the object. For example, as an arrangement function, there is a function of having a position and an object as parameters, and when the object and the position are passed, the specified object is arranged at the specified position in the three-dimensional virtual world. In addition, as a movement function, there is a function of having an object and a movement destination position as parameters, and moving the specified object to the specified position.
[0073] The three-dimensional virtual world, the data model are generated so as to imitate the real world as much as possible. In the real world, objects as solids do not overlap, and thus, for example, when two balls collide, overlap does not occur, but the balls bounce and collide to make a sound. In the three-dimensional virtual world, it is also programmed so that the objects do not overlap with each other and bounce and make a sound when colliding. In addition, gravity is also set as in the real world. That is, gravity always acts on the object toward the lower side of the vertical direction. Thus, the up-down direction can also be set in the three-dimensional virtual world.
[0074] Further, since time flows in the real world, this is also implemented by the program. The so-called time can be expressed by a one-dimensional time axis that flows from the past to the present and to the future. Further, the situation that is cut at the instant of the present is the situation that is unfolded in front of the eyes. That is, in the first platform 21, the situation that is unfolding at the present is the present recognized by the robot 31.
[0075] The situation of the three-dimensional virtual world at a certain instant is saved as an event, and the events are stored along the time axis, which is a story (narrative). The story sets the flow of time in the direction from the past to the present. That is, in the story, the scenes and events are managed in the order of time, which corresponds to the concept of "time". In the program, it can be implemented by using data structures such as arrays, lists, and the like.
[0076] Further, the world construction section 25 reproduces the real world, the story as a virtual world. The outside world reproduction section 26 reproduces the outside real world, and the output determination section 27 performs the operation.
[0077] The outside virtual world which faithfully reproduces the outside real world is developed on the first platform 21. Based on the information of the outside situation from the camera 14, the microphone 15, the world construction section 25 generates the outside virtual world on the first platform 21.
[0078] The physical object is constituted by the object of the three-dimensional data, and thus is logically able to move in the virtual world. However, the outside virtual world developed on the first platform 21 faithfully reproduces the outside real world, and thus, if the chair 35 is moved only in the outside virtual world, it will lead to the divergence from the real world. Therefore, the object of the first platform 21 is made not to be freely moved in the output determination section 27. Therefore, the second platform 22 exists as a platform which is able to develop a virtual world different from the outside virtual world. That is, the second platform 22 is able to freely perform the operation with the output determination section 27 (consciousness).
[0079] Here, it is assumed that the robot 31 is at work in a company, and the boss now enters the room 30, and the boss says "bring the battery on the shelf" to the robot 31. The robot 31 judges that the person who speaks is the boss by the face authentication, and converts the word to the text data by the voice recognition, and understands the meaning as the content said by the boss. In this case, it is understood that the meaning that the task is given is to take the battery 34 loaded on the shelf 33 and give it to the boss. That is, here, it is judged that the timing of the determination output has come (Yes in S14).
[0080] The output determination section 27 explores how to act to take the battery 34. For this, the second platform 22 is used. That is, here, the simulation is performed in the second platform 22 (S15).
[0081] First, the output determination section 27 constructs the three-dimensional virtual world same as the first platform 21 on the second platform 22. At this time, the robot 31 as itself is also arranged in the three-dimensional virtual world. By this, the output determination section 27 simulates the own action in the second platform 22, and is able to determine the best action. That is, the output is determined (Yes in S16).
[0082] Since the purpose is to take the battery 34, first, it is moved to the vicinity of the shelf 33 in a manner to approach the battery 34 as much as possible, and then, in order to take the battery 34, the hand is stretched to the upper side of the shelf 33 (refer to Figure 5 ). It is thus known that even if the hand is stretched, the hand cannot reach the battery 34. That is, the height is not enough.
[0083] Therefore, a method of supplementing the height is explored. The robot 31 has knowledge in the database 17 and explores a method of raising the height from it. In this way, the method of "standing on a chair" is found. Therefore, the chair is explored next, and the chair 35 is found in the room 30. Then, an attempt is made to simulate moving the chair 35 to under the shelf 33 on the second platform 22 and standing on it. In this way, it is possible to simulate reaching the shelf 33, and it is possible to obtain the battery 34 (refer to Figure 6 ). Then, it is simulated to take the battery 34 down from the chair 35 and take the battery 34 to the boss. If there is no problem, the series of actions of the self simulated is recorded. Then, the boss is acted upon in accordance with the determined output content (S17).
[0084] The self is configured on the second platform 22, and it is possible to objectively recognize the self, and the self is configured on the first platform 21, and it is subjective. The subjective self is the state of the self acquired by the sensor of the present self. For example, in the case where the camera 14 corresponding to the eyes acquires the hands and legs of the self, and there are temperature sensors and tactile sensors on the hands and legs, it is the data being detected by the self from those sensors.
[0085] It is explained that the present world in which the first platform 21 cannot be operated in the output determination section 27, but the body of the self such as the hands and legs of the self, the speaker 16 corresponding to the mouth, and the like can be operated. The reason is that if the output determination section 27 wants to raise the hand, the motor 13 of the hand of the robot 31 is driven, and the hand of the robot 31 in the outside real world is raised. In addition, this situation is acquired by the camera 14, and the hand of the self in the outside virtual world of the first platform 21 is also raised.
[0086] Therefore, by performing the action determined by the simulation in the first platform 21, it is possible to actually act in the real world and complete the task of obtaining the battery 34 and taking it to the boss.
[0087] In this way, the artificial intelligence system 11 of the present disclosure is able to determine the best action while simulating in the second platform 22. It can be said that this is roughly the same as the thinking of a person.
[0088] The world recognized by the output determination section 27 as the object acquired by the camera 14, the data model (object) is explained, but the data model is not limited to the physically existing, and anything that a person can recognize can become a data model. For example, try to consider the company organization. The positions of the section chief, the department head, and the like are not physically existing, but are concepts existing in the head of a person. Even such a concept, it is possible to recognize in the output determination section 27 using the data model (object) and the world in which the data model (object) is configured.
[0089] For example, when a company organization virtual world that imitates a company organization of a real world is generated, the company organization virtual world has a structure in which a plurality of positions are arranged in order of the positions. A data model of a real world employee, i.e., an employee object, is arranged in a corresponding position of the company organization virtual world. The company organization virtual world has a promotion function as a function of operating the employee object, which promotes an employee object of a section chief by one rank to become a division chief.
[0090] For example, assume that an acquaintance says "I became a division chief." If it is known that the person was a section chief before, he was promoted, which means that the value of the person was raised, so if an answer is given as "Congratulations," a meaningful conversation is established. This is to understand the meaning of words. By being able to understand the meaning, it is possible to have a natural conversation with a person.
[0091] In addition, the external world is not limited to a world that exists in reality, but can be a world in a movie, a novel. In that case, a virtual world is constructed in the first platform 21 based on video, text data acquired with the camera 14.
[0092] Next, a method of the robot of another embodiment using the artificial intelligence system 11 to determine an action is described. The robot has sensors that detect not only the external environment with a camera, a microphone, etc., but also its own internal state. As one of them, a battery level detection sensor is provided.
[0093] In addition, the artificial intelligence system 11 has positive and negative psychological states. The positive psychological state is a happy, full stomach, and the like, which is a psychological state that is good for oneself, and the negative psychological state is a sad, dangerous, and the like, which is a psychological state that is not good for oneself.
[0094] The output determination section 27 has a psychological state determination program, which inputs information from various sensors that detect the external environment, the internal state, to determine the psychological state. For example, in a case where the battery level from the battery level detection sensor decreases below a lower limit value, it is determined to be a negative psychological state of a full stomach, and in a case where it rises above an upper limit value, it is determined to be a positive psychological state of a full stomach.
[0095] The output determination section 27 explores an action for eliminating the psychological state when it feels a negative psychological state such as a full stomach. The database has a cause-effect dictionary in which causes and effects are paired.
[0096] The cause-effect dictionary records a rule such as "if A, then B" by managing pairs of a cause and an effect and recording them in a database. The cause-effect dictionary is one kind of long-term storage. As examples of the cause-effect dictionary, there are "if study, then brain will become good" and "the more you practice, the faster you run". The contents of the cause-effect dictionary are added through experience and learning. Also, here, data such as "if charge, then battery will recover" and "if replace battery, then battery will recover" is stored.
[0097] Next, the flowchart of FIG. 8 will be used to explain the method by which the output determination section 27 determines an action. When the empty stomach state included in the psychological state determination program of the output determination section 27 is detected in step SI, the output determination section 27 determines a negative psychological state. Figure 7
[0098] When the output determination section 27 detects a negative psychological state, it explores an action to eliminate the psychological state in the next step S2. By exploring the cause-effect dictionary of the database 17, for example, two actions of "if charge, then battery will recover" and "if replace battery, then battery will recover" are obtained.
[0099] In the following step S3, it is determined which of the obtained actions is selected. The output determination section 27 configures its own object on the second platform 22 and simulates the action obtained in step 2. In the action, a cost such as movement, operation, and the like, and a cost such as expense are consumed. By simulation, the cost at the time of charging is set to +5 and the cost at the time of replacing the battery is set to +10. It is assumed that data necessary for these cost calculations is stored in the database 17. The output determination section 27 selects the action with the lowest cost among the actions, and in this case, charging is selected.
[0100] When the action is determined, the action like simulation is performed in the next step S4. The first platform 21 is directly constructed as a model of the real world and is able to manipulate its own object. The own object configured on the first platform 21 is linked to the motor 13 and the speaker 16 which are output to the outside, and the output determination section 27 is constructed so that when the own object on the first platform 21 is manipulated, the own body in the outside world actually moves. This can be said to correspond to the primary motor cortex in the human brain.
[0101] Therefore, when the output determination section 27 applies the action selected in step S3 to the own object of the first platform 21, the own body is actually able to move, that is, charge, in the outside real world.
[0102] When the actual charging is performed and the battery level detection sensor exceeds the upper limit value, the psychological state determination program is the satiated psychological state in step S5. In this way, the output determination section 27 determines that the empty stomach psychological state is eliminated, and ends the action exploration to eliminate the empty stomach.
[0103] Thus, the artificial intelligence system 11 does not directly respond to the outside world using sensors, but is able to simulate and act while constructing a virtual world inside. This is the greatest advantage of the present application.
[0104] In order to easily explain this situation, a frog that directly responds to the outside world is considered. The frog is set to recognize a fly as bait by the activity of a black dot, and when the fly is recognized when hungry, the frog extends its tongue to prey on the fly to eat. If the environment does not change at all, the frog is able to survive in this program, but if the environment changes to a red fly instead of a black fly living in the environment, the frog is not able to prey on the red fly in the program that can only react to black dots and will starve to death.
[0105] On the other hand, if there is a platform that can simulate, even if the environment changes to a new environment, a new action corresponding to the environment can be generated to simulate and actually try the new environment, so even if the environment changes, the plan can be flexibly responded to and the action can be taken. This function is possessed by the human brain, and is also the artificial intelligence system 11 of the present application.
[0106] Here, the database 17 is explained. In the database 17, the shape, name, color, etc. of an object are recorded. This corresponds to the semantic memory of a human being's storage. Semantic memory is knowledge such as "apples are red" and "there are twelve months in a year".
[0107] In addition, on the first platform 21, an event that is currently occurring in the real world is unfolded. Then, when a certain emotion (psychological state) such as joy or surprise occurs, this is taken as an opportunity to save the situation unfolded at that time on the first platform 21 as a story in the database 17. Then, the artificial intelligence system 11 later unfolds the saved event on the second platform 22 and is able to recognize it again. This can be said to be a human being's "memory". "Memory" is a type of storage in which a scene appears as a video in the mind, and is called episodic memory. In this way, the database 17 not only records semantic memory, but also records episodic memory. The hippocampus in the human brain is equivalent to storing episodic memory. In this way, the database 17 functions as long-term memory that records semantic memory and episodic memory.
[0108] Furthermore, episodic memory is not only able to store the present world unfolded on the first platform 21, but also able to store a world created (imagined) by the output determination unit 27 and unfolded on the second platform 22.
[0109] In addition, the second platform 22 is not only able to be used for past events, but also able to be used in the case of imagining future events. In the second platform 22, "time" can be set to "yesterday" and "tomorrow".
[0110] The artificial intelligence system 11 adds a common part of the scenario memory, a repeatedly occurring event, and the like to the cause-effect dictionary, the semantic memory at the time when the output determination section 27 is not operating, that is, the so-called sleeping time for a human. If the output determination section 27 is always activated, this processing cannot be performed. Therefore, when the output determination section 27 is in the activated state for a long time, there is a desire (psychological state) to desire to end the activation of the output determination section 27. This corresponds to the sleepiness of a human who wants to sleep.
[0111] The explanation about the platform is also supplemented. It is about the advantage of reconstructing the world in the first platform 21.
[0112] For example, there are a table and a chair in front of the eyes, and they are sequentially subjected to image recognition by the camera 14, and the table is first recognized, and then the chair is recognized. That is, the table and the chair which should exist at the same time have the order of first discovering the table and then discovering the chair, and the world is not acquired as it is. It is constructed as a virtual world on the platform, and the output determination section 27 can feel the world as it is existing now by recognizing it. That is, it can be said that the world developed in the first platform 21 is the instant of "now".
[0113] Then, the world constructed on the first platform 21 or the second platform 22 and recognizable by the output determination section 27 corresponds to the short-term memory or the working memory of a human. Furthermore, not only the objects that can be seen with the eyes but also the position, money, and the like that cannot be detected with a sensor are constructed on the platform.
[0114] When the past memory developed in the second platform 22 and the "now" developed in the first platform 21 are maintained in a data structure capable of maintaining a plurality of events and scenes in order, the data structure is "time". The so-called "time" is not capable of being detected with a sensor or the like, but by providing it as an object that can be arranged on the platform, it can be recognized with the output determination section 27. The characteristic of "time" is that it flows in only one direction from the past to the present and from the present to the future in the real world. In addition, the cause and the effect registered in the cause-effect dictionary are managed in the order of time in which the cause is the previous event and the effect is the subsequent event.
[0115] Next, a method in which the output determination section 27 understands the rules of the society default of good and evil is described. For example, consider the following situation.
[0116] When I was walking in the park, a small boy was crying. Then, I went and asked "What's wrong, little boy?", and it turned out that the ball was caught by a branch and could not be taken out. Therefore, I removed the ball for him.
[0117] This is an extremely common action to help a child in trouble. However, on closer examination, it is incredible that it is taken for granted that such an action is taken. It is possible to pass by without feeling anything, but when a child is crying in front of one's eyes, no one can pass by without feeling anything. For example, even if the wind blows and the leaves shake, it is normal to pass by without any concern. What is the difference in the mind? It can be said that there is a social default rule of "should take good actions" in the background of the human mind.
[0118] To make robots realize the social default rule, they must be able to understand the meaning of good and evil. However, this is surprisingly difficult.
[0119] Robots can be programmed to take all "good" actions such as "help a child in trouble" and "pick up litter on the road" because they take actions determined by programming, but there is no end to it. Humans can understand what is good and what is bad in the course of social life even if they are not taught all of these actions. It is not that they act after being taught all good and evil actions.
[0120] Moreover, humans can also not do what they know is good and do what they know is evil. Good and evil are not simply related to human actions. That is, it is not possible to learn good and evil from human actions through mechanical learning.
[0121] Therefore, regarding good and evil, rather than focusing on actions alone, the psychological state in the background thereof is attempted to be focused on. In this way, it is known that good actions are the psychological state of "should ~". That is, good actions can be said to be the rule set by society to individuals by default. Therefore, the output determination section 27 has a "should determination program".
[0122] Here, the "should determination program" is considered. First, assume that the subject objects of oneself and the other party are arranged on the platforms 21, 22, and the output determination section 27 determines the positive and negative emotions of oneself and the other party. The output determination section 27 determines the action in principle in such a way that oneself is positive. However, in the case of acting in society, the "should determination program" works.
[0123] The output determination section 27 arranges the objects of oneself and the boy as subjects on the second platform 22 and simulates the actions that oneself can take. One is a case where one simply leaves without doing anything, and the other is a case where one helps the boy. The output determination section 27 calculates the positive and negative emotions of oneself in numerical values, for example, according to the situation. In the case of leaving, one has been walking in the park so far, and will simply continue walking next, and there is no change, so it is 0.
[0124] As Figure 8In the case of helping the boy, as a result of greeting the boy and performing some operation, it is expected that labor and time will be spent, and the output determination unit 27 expects positive and negative emotions of -5 (negative 5) at this time.
[0125] The output determination unit 27 determines an action in a manner that the positive and negative emotions of the self are the largest, and thus, when judged based on these conditions, an action of directly leaving is taken. However, in this case, the propriety determination program performs the operation.
[0126] The propriety determination program considers the positive and negative emotions of the other person in the case where the other person exists other than the self. The positive and negative emotions of the other person are estimated based on various conditions, as the positive and negative emotions of the other person cannot be detected with a sensor or the like. In this case, the other person is in a "crying" condition. The crying condition can be discriminated based on the camera 14, the sound. Then, the state of "crying" is registered as a negative emotion, for example, -10, in the database 17. In addition, if the boy is helped, the positive and negative emotions of the boy are estimated to become +5 (positive 5).
[0127] Then, the positive and negative emotions of the other person are set, and the output determination unit 27 considers the positive and negative emotions of the other person on the positive and negative emotions of the self, for example, determines an action by addition. In this way, in the case of leaving, the positive and negative emotions of the other person are -10, and thus, the total is -10. On the other hand, in the case of helping, the positive and negative emotions of the other person are +5 (positive 5), and thus, the total is ±0 (positive and negative 0), and since it is larger in the case of helping, the action of helping is selected.
[0128] In the simulation, in the case of leaving, the output determination unit 27 feels a negative emotion of -10, which indicates that just imagining that a person in difficulty is present in front of the self and directly leaving will result in an unpleasant feeling, a feeling that this is a bad thing. Of course, not only in the simulation, but also in the case of actually leaving, an unpleasant feeling will be generated.
[0129] In this way, when the output determination unit 27 determines an action, a bias can be applied so that the action of the ego is suppressed, and an action of altruism is taken. Thereby, it is possible to realize an output determination unit 27 that performs a good deed, that is, helps a person in difficulty, sympathizes with the other person. That is, the output determination unit 27 determines an output based on the total of the quantified positive and negative emotions of the person object disposed on the second platform 22. In this way, it is possible to perform a more sympathetic response to the other person, and it is possible to become an artificial intelligence that exhibits a psychology similar to that of a human being.
[0130] It is also possible to have a good deed in the case where the other person does not exist. For example, if an empty can is dropped in a park, picking it up and putting it in a garbage can is a good deed. This case is also considered.
[0131] As Figure 9As shown, the own positive and negative emotions in the case of leaving without doing anything are 0 because of not doing anything, and the own positive and negative emotions in the case of picking up the empty can are -5 because of the need to expend labor.
[0132] Next, assume that the empty can is recognized to have fallen. The empty can is one kind of garbage, and is judged to be of negative value according to semantic memory. When it is recognized that something has positive and negative value, next, the social value in that case is considered. The society in that case is, for example, the residents who are using the park, and the like. Then, when the person's positive and negative emotions are estimated, it is set to -10 in the case of dropping garbage, and 0 in the case of no garbage.
[0133] Thus, the total in the case of picking up garbage is -5, and in the case of not picking up garbage is -10, and the value in the case of picking up garbage becomes large, and the action of picking up garbage is taken. Thus, even in the absence of the other party, assuming the regional residents, society, and the like, the choice of good action is decided.
[0134] Thus, by correcting the action decision of the self by estimating the positive and negative emotions of the subject other than the self, without storing the infinite existence of good and evil actions, the ethics of good and evil can be realized.
[0135] Further, the value of the positive and negative emotions set by the output determination section 27, the shouldness determination program is not fixed, and differs according to the robot, and this becomes the personality of the robot. For example, if there is a tendency to set the positive and negative emotions of the other party higher, it is an altruistic, excellent personality, and if the own positive and negative emotions are set higher, it is a robot of a selfish, selfish personality.
[0136] Next, "should" in the case of not good and evil is explained. For example, when parents say to a child "don't just play, learn", the background is the existence of the idea that "learning should be done more than playing". This "should learn" is not social good and evil. It is not related to society, but to the individual.
[0137] Here, the positive and negative emotions of the individual are divided into two, low-level desires and high-level desires. The low-level desire is something that is obtained immediately, obtained at low cost, and in short, is the body, the flesh, and is typically a desire from instinct. That is, it is a desire that is the motive force of action based on the instinct of an animal that avoids unhappiness and pursues happiness. It is a desire for food, sex, sleep, comfort, safety, and the pursuit of pleasure. In addition, the desire that is obtained immediately without deep thought is also included in the low-level desire. For example, it is a game, gambling, or a hobby such as alcohol, cigarettes, coffee, and the like, or a drug.
[0138] The high-level desire is a desire to compare with others and to have a higher value in society than others, and is a thing that cannot be obtained without spending a long time and a high cost. For example, it is a desire to have a position such as a general manager, a doctor, a politician, a professor, a professional athlete, a singer, an actor, a celebrity, a high income, a high education, and the like. The high-level desire is not limited to such a general society, but also includes a victory in a sports meeting, a rise in value and position in a school, and the like.
[0139] Therefore, if the positive and negative emotions of a person are divided into the low-level desire and the high-level desire, it can be said that the low-level desire is suppressed and the high-level desire should be selected. In short, it can be said that the present pleasure should be suppressed and the long-term self-growth should be selected.
[0140] The case where the high-level desire is satisfied if a good university is entered is considered. This can be inferred using the cause-effect dictionary. For example, "If you study, your head will become good" and "If your head is good, you can enter a good university" are chained with the cause-effect according to the cause-effect dictionary. In this way, "study" is obtained as one of the actions that can be taken by the person himself or herself now in order to enter a good university.
[0141] The low-level desire exists in two types, a type that can be detected using a sensor such as hunger and battery level, and a type that arises from the inside such as a desire to play and a desire to drink. The world construction unit 25 generates a so-called desire to play that arises from the inside as the low-level desire of the person.
[0142] The output determination unit 27 attempts to take the action of "study" in order to enter a good university, but on the contrary, feels the low-level desire of "desire to play". The output determination unit 27 determines the action from the two options of "play" and "study".
[0143] If "play" is selected, the pleasure is high, the low-level desire is, for example, +10, but the study cannot be done, the test score decreases, and the high-level desire is, for example, -20, and the total is -10. On the other hand, if "study" is selected, the play cannot be done immediately, the low-level desire is not satisfied, and is, for example, -5, but the test score is improved, and a good school can be entered, so the high-level desire is, for example, +20, and the total is +15. Therefore, it is decided to select the action of "study". As a result, the action for suppressing the low-level desire and achieving the high-level desire can be promoted. This is the "shouldness program". The parameters set differ depending on the personality and experience, and this is the individuality. That is, it is the reason why the robot becomes a robot that easily falls into pleasure or a robot that is patient.
[0144] To give advice for the other's trouble, it is necessary to understand the meaning of the trouble, and if the meaning of the default rule of "should" is not understood, the trouble of "want to play but must study" is not understood, and even the conversation does not exist. In the artificial intelligence system of the present invention, since the meaning of the so-called "should" can be understood, a natural conversation with a person can be performed.
[0145] The feeling of "should" is a feeling that is commonly possessed by the same person living in society. In other words, it can be said that if the common feeling is possessed, it can be accepted as a member of the society. One of them is the rule of good and evil.
[0146] AI described in movies, novels only performs actions that are logically correct, but does not feel the human psychology of feeling the other's feelings. According to the artificial intelligence system of the present invention, it is possible to sympathize with the other's feelings, take the action that should be done, and naturally be accepted by society.
[0147] In addition, the lower order desire and the higher order desire can be considered as follows. The person object imitates a person, and as a program, has a desire and an emotion that a person has. For example, as an attribute, it has a degree of hunger as a numerical value in a hunger degree characteristic. In addition, as a method, it has an action of "eating". The association is performed in such a way that the desire to eat food increases if the hunger degree characteristic increases. This is appetite. The appetite is the motive force of the action of wanting to eat, and the hunger degree and the desire are associated in such a way that the size of the desire to want to eat is proportional to the hunger degree characteristic. In this way, it is possible to reproduce appetite using a computer program. The desire of a human being is generated based on the body, the instinct such as sleep desire and sexual desire. The desire required to maintain the state of the object itself, for example, in the case where the subject is a robot, is called a lower order desire.
[0148] The desire of a human being is generated based on the body, the instinct such as sleep desire and sexual desire. The desire required to maintain the state of the object itself, for example, in the case where the subject is a robot, is called a lower order desire.
[0149] Next, a slightly complicated psychological state will be explained. People have various psychological states, from emotions such as "happy" and "sad" to complex emotions such as "regret" and "jealousy," to ethical views such as good and evil, and a conversation is established by being able to understand the psychological state of the other party. If the other party is happy, the answer is "good," and if the other party is sad, the response is "that's really sad," and the other party feels that the other person understands their feelings. And, this is the daily conversation.
[0150] That is, in a daily conversation, the most important thing is to be able to understand the psychological state of the other party. Furthermore, making a response corresponding to that psychological state is a daily conversation.
[0151] Next, the method of analyzing the psychological state as a psychological model will be explained.
[0152] Then, the first psychological model of "restraint" will be explained. It was explained earlier that in play and learning, learning is chosen. The lower desire is suppressed, and the higher desire should be chosen, and the force comes from the "should program," and the state of suppressing the lower desire at this time can be defined as "restraint." It refers to a psychological model that suppresses the desire that mainly comes from the body, such as difficulty and pain.
[0153] When the lower desire is suppressed and the higher desire is chosen, the psychological model that focuses on the lower desire is "restraint," and the psychological model that focuses on the "higher desire" is "encouragement" and "effort." For example, regarding the behavior of "learning," since the lower desire is negative, it is an action that is not desired, but the higher desire at the time of realization is high. This can be said to be aiming at the goal, suppressing the lower desire, and taking action that becomes a higher desire. This is the psychological model of "encouragement" and "effort." The condition of running a marathon is aiming at the finish line as a goal, suppressing the lower desire to rest, and taking action to continue running as a higher desire, so it is a psychological model of "encouragement."
[0154] Then, the person in this condition is given "cheering!" This is a psychological model of "cheering." Cheering is a call word that affirms and encourages the action of the higher desire chosen by the other party in the case where the psychological model of the other party is "encouragement."
[0155] Next, the psychological model of "regret" will be explained. To "regret," a goal is first needed. The goal is a higher desire like getting into a university that one wants to be in the future. Then, it is produced when focusing on the past action that the higher desire is not realized and it is possible to realize the goal if the past action is changed. In order to be able to understand this concept, it is necessary to understand the "if" assumption.
[0156] The second platform 22 is used in order to understand the assumption. For example, consider the case of regret about failing the university entrance examination, if only I had studied harder. First, configure the past self on the second platform 22, simulate the action that should have been taken among the actions that could have been chosen at that time, and the case where learning was performed. Register the information that the head will become good if learning, and the head is good if the university is entered in the cause-effect dictionary, and it is possible to simulate that the university can be entered by learning. The output determination unit 27 compares the university entry based on the simulation result with the examination failure that actually occurred, and identifies that the reason is the action of "not studying" being chosen. That is, the artificial intelligence system 11 compares the reality with the simulation, and when it understands that the past action of oneself was the reason for not achieving the goal, a psychological model of "if only I had studied harder" is generated.
[0157] Not only one's own case, but also in imagining the psychological model of the other party, if the other party is judged to be regretting, the conversation will be established if "it is really too bad" or "if only I had studied harder" is said to the other party.
[0158] Next, the psychological model of "excuse" is explained. "Excuse" is one of the psychological models generated in the case where the advanced desire as a goal is not achieved. "Regret" is a psychological model in which the reason for not achieving the purpose is attributed to oneself, but "excuse" is a psychological model in which the reason for not achieving the goal is attributed to something other than oneself. For example, it is a case where the reason for the examination failure is said to be "the neighbor was too noisy to concentrate on studying".
[0159] In the case of "regret", the reason is oneself, and therefore negative emotion is generated, but in the case of "excuse", the reason is something other than oneself, and therefore negative emotion is not generated. In this way, even if the same result occurs, various psychological models are generated depending on the personality, and the answer differs depending on the personality. This is the reason why even if a large amount of conversation data is concentrated and learning is performed by mechanical learning, a correct answer cannot be obtained.
[0160] Next, the psychological model of "pride" is explained. "Pride" is a psychological model in which oneself satisfies the advanced desire and intentionally shows it to others. The so-called advanced of the advanced desire is to be valuable in society, and is generated by comparison with others. That is, it can be said that oneself is valuable to society. Intentionally showing this to others who have not satisfied the advanced desire is a behavior that makes oneself feel more satisfied and makes positive emotion rise. This is the psychological model of "pride".
[0161] "Pride", "envy", and "jealousy" can be said to be psychological models opposite thereto. That is, it is negative emotion generated when it is known that others have obtained an advanced desire that oneself wants but oneself does not obtain.
[0162] Next, a psychological model of "shame" is explained. A higher-order desire is a desire to be a being above a given value in society. In other words, a being below the given value can be said to be an ordinary being. If the value further decreases to below a certain benchmark, it can be said to be a being worse than ordinary. "Shame" is a psychological model when it decreases below the certain benchmark. The benchmark is determined by the society to which the being belongs, for example, in a track team, it is generally 6 seconds to run 50 m, and if it takes more than 7 seconds, it is "shame". This is not explicitly stated, but exists as a default rule of the society, and there can be various benchmarks such as clothing, ability, and the like. Understanding the rule and maintaining the minimum benchmark can also be said to be the minimum condition for being accepted by the society. It can be said that it is necessary for an artificial intelligence that is accepted by society to understand the psychological model of shame.
[0163] Next, a psychological model of "win" and "lose" is explained. For this purpose, first, the concept of "opposition" is explained. "Opposition" is not a psychological model, but a concept, and is configured on the platform and output by the determination unit 27. The concept called opposition is a situation in which two subjects compete. The competition is an action to determine which is superior. The subjects of opposition are not limited to two, but can be more than two. In addition, the subjects are not limited to one person, but can be a combination of a group, a team, a country. Opposition is a concept established in situations such as sports, war, games, and the like.
[0164] The platform can set the concept of opposition, configure two subjects, compete according to the rule of determining who is superior, and the subject who becomes superior is "win" and the subject who becomes inferior is "lose". "Win" is a positive emotion and "lose" is a negative emotion.
[0165] Next, "metaphor" is explained. For example, try to consider the word "exam war". "Examination" is not a mutual killing, but is completely different from war, but both "war" and "examination" are competing with each other, which is the same, and can be said to belong to the concept of "opposition". In that case, "examination" is compared with "war" via the concept of opposition, and the intensity of the examination is emphasized as much as war. If "examination" and "war" are configured in the concept of opposition set on the second platform 22, the output determination unit 27 that recognizes it can be implemented by projecting the intense impression that "war" has to "examination".
[0166] In this way, the expression of metaphor can also be understood by using a psychological model. The understanding of metaphor is also the most difficult in the natural language processing of the past.
[0167] Next, another embodiment of the present disclosure is explained. The artificial intelligence system of the present disclosure is set to read the following story to answer the question. The story is as follows.
[0168] "A certain A and B are in a room. A has a basket and B has a box. A has marbles. A puts marbles in the basket. A goes for a walk. B takes A's marbles from the basket and puts them in the box. A comes back. A wants to play with the marbles. Then the question is, where does A find the marbles?"
[0169] The generating section 24 reads the story and generates a virtual world in the first platform 21. At first, as shown in Fig. 4, a data model of a room 41 is generated and arranged in the first platform 21. Then, objects of A 42 and B 43 are arranged in the room 41, and objects of a basket 44 and a box 45 are arranged in the room 41. This state is set as a first scene. Figure 11
[0170] Next, a scene in which A 42 puts marbles in the basket 44 is implemented in the virtual world. A scene having an activity or a change is set as an event. The next event is an event in which A 42 goes out of the room 41. After that, a scene in which B 43, the basket 44, and the box 45 are arranged in the room 41 is continued. In this way, the story is managed by a series of scenes and events.
[0171] When the story is read in this way, in the virtual world of the first platform 21, in the last scene, the marbles are put in the box 45. Here, it is assumed that the question "where does A find the marbles?" is asked. In the virtual world of the first platform 21, since the marbles are put in the box 45, it is directly answered that "find the box".
[0172] However, when the marbles are moved from the basket 44 to the box 45, A 42 is outside and does not know this, and therefore "find the basket" is the correct answer. Therefore, in order to be able to correctly answer the question, it is assumed that a person arranged in the first platform 21 has an artificial intelligence program 46. In this case, A 42 arranged in the first platform 21 has the artificial intelligence program 46, as shown in Fig. 5, and the artificial intelligence program 46 has a first platform, i.e., a lower side first platform 47, for A 42. Then A 42 constructs a world which A 42 considers to be real on the lower side first platform 47 based on information obtained from the outside by seeing or hearing. Figure 11
[0173] Here, A 42 himself put the glass marbles into the basket 44 and left the room 41, so he did not see B move the glass marbles to the box 45. That is, the glass marbles on the lower side first platform 47 of A 42 are still in the basket 48. The question is "where does A look for the glass marbles?" This is a question that cannot be answered from the standpoint of A 42. That is, it must be judged based on the situation that the artificial intelligence program 46 of A 42 considers to be real, not based on the actual real world situation. That is the virtual world constructed on the lower side first platform 47 of A 42, where the glass marbles are not put into the box 49 but into the basket 48, so A 42 should look for the "basket".
[0174] In this way, the artificial intelligence program is also provided on the main body (human object) of the first platform 21, that is, by configuring the artificial intelligence program in a nested structure, it is possible to have a more human-like mind based on the consideration from the other party's standpoint.
[0175] In this way, the structure that considers from the other party's standpoint is advocated as "mind logic",
[0176] "Mind logic" is said to be a capability that only humans have. By installing the lower side first platform, it is also possible to realize "mind logic".
[0177] In addition, the nested structure is not only double, but can also be deepened to triple, quadruple, or any number of times. However, since the amount of calculation becomes large, it is preferable to limit the structure to a nested structure of double or at most triple.
[0178] It should be understood that the embodiments disclosed herein are illustrative in all aspects, and are not limiting in any aspect. The scope of the present application is not defined by the above description, but by the claims, and is intended to include all modifications within the meaning and range of equivalents of the claims.
[0179] Explanation of Reference Signs
[0180] 11 - artificial intelligence system; 12 - control device; 13 - motor; 14 - camera; 15 - microphone; 16 - speaker; 17 - database; 18 - control section; 21 - first platform; 22 - second platform; 24 - generation section; 25 - world construction section; 26 - external situation reproduction section; 27 - output determination section; 30 - room; 31 - robot; 32 - wall; 33 - shelf; 34 - battery; 35 - chair; 41 - room; 42 - A; 43 - B; 44, 48 - basket; 45, 49 - box; 46 - artificial intelligence program; 47 - lower side first platform.
Claims
1. An artificial intelligence system that determines an output to an outside based on information inputted from the outside, the artificial intelligence system comprising: a storage section that stores in advance a data model that imitates a human and a human's thought; a generation section that extracts the data model from the storage section and generates a human object that can reproduce a human's action and thought; a world building section that has a first platform and a second platform on which the human object is arranged and builds a world in which the human object's action and thought are developed; an outside world reproducing section that arranges the human object on the first platform based on information inputted from the outside and reproduces an outside world; and an output determining section that grasps a situation of the outside by recognizing the outside world reproduced on the first platform and determines an output to the outside by arranging the human object on the second platform and operating it.
2. The artificial intelligence system according to claim 1, wherein the human object of a self and an opposite party is arranged on the first platform and the second platform, and the output determining section determines an output in such a manner that a thought of the opposite party's human object is felt to be satisfactory.
3. The artificial intelligence system according to claim 1 or 2, wherein the human's data model has two kinds of desires, a low-level desire that is generated by a body and a high-level desire that pursues something with a high social value, and the output determining section determines an output in such a manner that the low-level desire is suppressed and the high-level desire is satisfied.
Citation Information
Patent Citations
Interactive processing device
JP2006178063A
Artificial intelligence device for autonomously constructing knowledge system by language input
JP2016040730A