Information processor, information processing method and program
The information processing device generates feature maps from object position and behavior information to analyze the impact of environmental changes on inference results, addressing the limitations of conventional methods by providing insights into the physical factors influencing machine learning outcomes.
Patent Information
- Application Number
- JP2024023305
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2025-09-01
AI Technical Summary
Conventional methods for acquiring pixels that affect inference results in machine learning models fail to capture physical information in the environment that influences these results.
An information processing device that generates feature maps based on object position and behavior information, processes specific channels, and compares inference outputs before and after processing to determine the impact of environmental information on inference results.
Enables the acquisition of physical information affecting inference results by analyzing changes in inference outputs, thereby understanding the basis for determining environmental inferences more effectively.
Smart Images

Figure 2025126938000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In machine learning, it is difficult to grasp the basis for determining the inference results shown by a machine learning model. For example, Non-Patent Document 1 describes a technique for grasping the basis for determining the inference results shown by a machine learning model. The technique described in Non-Patent Document 1 randomly applies a mask to an image input to a machine learning model, and acquires important pixels in the image that affect the inference results based on the change in the inference results depending on whether the mask is present in the image. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] V.Petsiuk, A.Das, K.Saenko, “Rise: randomized input sampling for explanation of black-box models,” British Machine Vision Conference(BMVC), 2018. Summary of the Invention [Problem to be solved by the invention]
[0004] The conventional technology described in Non-Patent Document 1 acquires pixels that affect the inference results from the images input to the machine learning model, but has the problem of being unable to acquire physical information in the environment that affects the inference results.
[0005] The present disclosure is intended to solve the above-mentioned problems, and aims to provide an information processing device that is capable of acquiring physical information in an environment that affects the inference results of the environment. [Means for solving the problem]
[0006] The information processing device according to the present disclosure includes an environmental information acquisition unit that acquires environmental information including positional information of objects included in the environment or behavioral information of the objects; a map generation unit that generates a first feature map including a plurality of channels set based on the positional information of the objects or the behavioral information of the objects; a map processing unit that processes at least a first channel out of the plurality of channels included in the first feature map; and an inference unit that receives as input at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information that is information regarding the inference of the environment based on the second feature map. [Effects of the Invention]
[0007] According to the present disclosure, an information processing device according to the present disclosure generates a first feature map including multiple channels set based on object position information or object behavior information included in environmental information, processes at least a first channel of the multiple channels included in the first feature map, inputs at least a portion of the first feature map including the first channel and a second feature map including the processed first channel, and outputs first inference information that is information related to environmental inference based on at least a portion of the first feature map and second inference information that is information related to environmental inference based on the second feature map. Based on the first inference information and second inference information related to the environment to be inferred, the information processing device according to the present disclosure is able to obtain physical information in the environment that affects the environmental inference result. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an example of the configuration of an information processing device according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an overview of processing performed by an information processing device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram illustrating an outline of a feature amount map converted from environmental information. [Figure 4] 1 is a block diagram showing a hardware configuration for realizing the functions of an information processing device according to a first embodiment. [Figure 5] 5A, 5B and 5C illustrate a basketball scenario. [Figure 6] FIG. 2 is a schematic diagram showing an overview of basketball scenario identification by the information processing device according to the first embodiment. [Figure 7] FIG. 1 is a schematic diagram illustrating an overview of a feature map converted from a basketball scenario. [Figure 8] 8A and 8B are conceptual diagrams showing an outline of the process of applying a mask to a three-point shot channel included in a feature amount map. [Figure 9] 3 is a flowchart showing an information processing method according to the first embodiment. [Figure 10] FIG. 1 is a conceptual diagram showing an overview of processing using one scenario. [Figure 11] 11A and 11B are conceptual diagrams showing an outline of the process of masking the position information of an opponent who takes a shot after dribbling, in a channel included in the feature amount map. [Figure 12] 12A and 12B are conceptual diagrams showing an outline of the process of masking the position information of an enemy that moves little in a channel included in a feature amount map. [Figure 13] 13A, 13B, and 13C are diagrams showing feature maps in which changes are made to the information of various channels. [Figure 14] 14A and 14B are conceptual diagrams showing an outline of a battle game in which bees and hornets fight each other. [Figure 15] FIG. 1 is a schematic diagram illustrating an overview of a process for identifying multiple scenarios of a fighting game between bees and hornets using a machine learning model. [Figure 16]FIG. 10 is a schematic diagram showing an outline of a feature map converted from a scenario of a fighting game in which bees and hornets fight. [Figure 17] 17A and 17B are conceptual diagrams showing an outline of the process of applying a mask to the position information of a specific hornet in a channel included in a feature map. [Figure 18] 18A and 18B are conceptual diagrams showing an outline of the process of applying a mask to the position information of all hornets in the channels included in the feature amount map. [Figure 19] 19A and 19B are conceptual diagrams showing an outline of a process for masking attacks on bees in channels included in a feature map. [Figure 20] FIG. 10 is a schematic diagram showing an outline of a feature map converted from environmental information indicating the environment of a real space. [Figure 21] FIG. 10 is a block diagram showing an example of the configuration of an information processing device according to a second embodiment. [Figure 22] 10 is a flowchart showing an information processing method according to the second embodiment. [Figure 23] FIG. 11 is a block diagram showing an example of the configuration of an information processing device according to a third embodiment. [Figure 24] FIG. 11 is a conceptual diagram showing an overview of processing by an information processing device according to a third embodiment. [Figure 25] 11 is a flowchart showing an information processing method according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Embodiment 1 (Outline of information processing device) FIG. 1 is a block diagram showing an example configuration of an information processing device 1 according to a first embodiment. In FIG. 1, the information processing device 1 acquires environmental information including position information or behavioral information of objects included in the environment, generates a first feature map including multiple channels set based on the position information or behavioral information of the objects, processes a first channel among the multiple channels, and outputs first inference information regarding environmental inference based on the first feature map including the first channel and second inference information based on a second feature map including the processed first channel. By using the first inference information and the second inference information, the information processing device 1 can shorten the number of attempts required to understand the basis for determining an inference result regarding the environment.
[0010] The environment is the environment to be inferred, and may be a real-world environment or a virtual environment operated by a simulator. The object is a physical object contained in the environment. For example, objects include not only machine learning agents but also traffic lights and warehouses contained in the environment. The agent is the subject of decision-making and action, and the environment is the object that sends and receives information to and from the agent. The agent is realized as a program that constitutes a machine learning application, learns from its own experiences, and determines future actions based on the results of that learning.
[0011] The environmental information includes physical information such as the position or behavior of an object, which is an example of an object. The object's position information is, for example, two-dimensional coordinate information of the object on a captured image obtained by capturing an image of a real space, or two-dimensional coordinate information of an agent on an image showing a simulation environment. The object's behavior information may be, for example, information indicating a time-series change in the object's position. Note that the object's position information or behavior information may be two-dimensional information or multidimensional information of three or more dimensions.
[0012] The environmental information may be time-series data or data at a specific time. The first feature map includes a plurality of channels generated based on the environmental information. Each channel is set with location information indicating the location of an object or behavior information indicating the behavior of the object.
[0013] Of the multiple channels included in the first feature map, the channel that has been processed by the map processing unit 123 is the first channel. The feature map including the processed first channel is the second feature map. "Processing" includes not only applying a mask to an object included in the environment, but also making changes to the position or behavior of the object. The information processing device 1 processes the environmental information included as a channel in the first feature map, inputs the feature map before processing and the feature map after processing, performs inference regarding the environment, and compares the output first inference information with the output second inference information. The result of this comparison makes it possible to know how much the information set in the processed first channel, among the channels included in the first feature map, affected the inference result regarding the environment.
[0014] The information processing device 1 will be described in detail below. As shown in Fig. 1, the information processing device 1 includes a communication unit 11, a calculation unit 12, a storage unit 13, and a display unit 14. The communication unit 11 communicates with an external device via a network. For example, the communication unit 11 may be a communication device capable of mobile communication using a communication method such as LTE, 3G, 4G, or 5G. The communication unit 11 may also include a short-range wireless communication means such as Bluetooth (registered trademark). The information processing device 1 does not necessarily have to include the communication unit 11. For example, when environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1 does not include the communication unit 11, the environmental information acquisition unit 121 realized by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.
[0015] The calculation unit 12 controls the overall operation of the information processing device 1. The calculation unit 12 includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123, an inference unit 124, and an output unit 125. The calculation unit 12 executes an information processing application, thereby realizing the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, the inference unit 124, and the output unit 125.
[0016] The storage unit 13 stores, for example, an information processing application and information used in the arithmetic processing of the arithmetic unit 12. The storage unit 13 is a storage device provided in a computer that functions as the information processing device 1, and is a storage such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Note that the storage unit 13 may be provided outside the information processing device 1 as long as it is accessible by the information processing device 1.
[0017] The display unit 14 is a display device included in the information processing device 1. The display unit 14 is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display device. The display unit 14 may be provided outside the information processing device 1 as long as the information processing device 1 can display information. For example, the display unit 14 may be a display device included in a terminal device that can communicate with the information processing device 1 via a network.
[0018] (Environmental Information Acquisition Department) The environment information acquisition unit 121 acquires environment information including position information or behavior information of objects included in the environment. For example, when the inference target is an environment modeled after the Gym format in Open AI, which is a learning environment often used in reinforcement learning, the environment information acquisition unit 121 acquires time-series environment information from the simulator by using the environment observation function of the simulator. For example, specifically, the environmental information acquisition unit 121 communicates with an external device that functions as a simulator via a communication network using the communication unit 11, or connects to a simulator provided in the information processing device 1 via a signal line, and acquires time-series environmental information for the target period from the simulator.
[0019] The time-series environmental information is information indicating the positions or actions of objects included in the simulation environment. Fig. 2 is a conceptual diagram showing an overview of processing by the information processing device 1. For example, the environmental information indicating the simulation environment of a fighting game is time-series information made up of environmental information for each time, as shown in Fig. 2, and includes, for example, position information indicating the position of an agent at each time.
[0020] When the real-space environment (hereinafter referred to as the real environment) is the inference target, the environment information acquisition unit 121 acquires time-series environment information from, for example, a sensor arranged in the real space. The sensor may be a visible light camera, a GPS (Global Poisoning System), a ToF (Time of Flight) sensor, or an infrared camera. For example, specifically, the environment information acquisition unit 121 connects to a sensor via a communication network or via a signal line via the communication unit 11, and acquires time-series environment information indicating the real environment from the sensor. The environment information indicating the real environment is, for example, a captured image of the real environment captured by a camera.
[0021] (Map generation section) The map generation unit 122 generates a feature amount map based on the environmental information acquired by the environmental information acquisition unit 121. The feature amount map generated by the map generation unit 122 is a first feature amount map including a plurality of channels set based on the position information of the object or the behavior information of the object. For example, the map generation unit 122 generates a channel for each time based on the environmental information for each time, thereby generating a feature map made up of time-series channels as shown in Fig. 2. Numerical values indicating the position information or behavior information of an object are set in the channel, and the environment in which the object exists can be expressed in a binary image format or a multi-value image format.
[0022] The map generation unit 122 assigns a numerical value indicating a color gradation, such as a value from 1 to 255, to each object identified from the environmental information, or assigns a numerical value according to the role of the object. The map information in binary or multi-value image format is information indicating an image having pixel values according to the numerical values assigned to the object and other parts. For example, when representing an environment in which an object exists as a binary image, the map generation unit 122 assigns the value "1" to the object out of the two values "1" and "0" and assigns the value "0" to areas where no object exists, so that in the binary image, the object is represented as a white area and areas other than the object are represented as black areas. In addition, the map generation unit 122 may assign a value of "1" to objects performing the target behavior and represent them as white areas in the binary image, and assign a value of "0" to areas where no objects exist and objects not performing the target behavior and represent them as black areas.
[0023] FIG. 3 is a conceptual diagram showing an overview of a feature map converted from environmental information, and shows a feature map converted by the map generation unit 122 from environmental information indicating the simulation environment of a competitive game. In FIG. 3, it is assumed that objects represented by circles and squares exist in the simulation environment indicated by the environmental information at a certain time, with allies being white objects and opponents being gray objects. Based on the environmental information, the map generation unit 122 determines, for example, that each type of object exists in the simulation environment, assigns a value of "1" to the object, and assigns a value of "0" to areas where no object exists. The numerical values set for the objects, etc. are parameter values corresponding to the colors on the image.
[0024] For example, the feature map shown in Fig. 3 includes channels ch1, ch2, and ch3, etc. These channels include binary images representing the simulation environment. As shown in Figure 3, channel ch1 indicates the position of the opponent's object. In the binary image contained in channel ch1, the image area corresponding to the position of the opponent's object is expressed in white, and the image area where this object does not exist is expressed in black. Channel ch2 indicates the position of the ally's object. In the binary image contained in channel ch2, the image area corresponding to the position of the ally's object is expressed in white, and the image area where this object does not exist is expressed in black. Channel ch3 indicates the opponent's object that is performing an "attack," which is the target behavior for judging the degree of contribution to environmental inference. In the binary image contained in channel ch3, the image area corresponding to the position of the opponent's object performing the "attack" behavior is expressed in white, and the other image areas are expressed in black.
[0025] Note that the map generating unit 122 may assign the value "0" to objects out of the two values "1" and "0" and assign the value "1" to other objects. As a result, in the binary image, image areas where objects exist are displayed in black, and image areas where no objects exist are displayed in white.
[0026] Furthermore, when an object is performing multiple actions, the map generating unit 122 may assign a digital value to the object by calculating the logical product (AND) or logical sum (OR) of the numerical values assigned to each of the multiple actions. For example, a value of "1" is assigned to the action of "advancing," and a value of "0" is assigned to the action of "attacking." When identifying an object that is "attacking" while "advancing" as an object that is "advancing," the map generation unit 122 assigns the object a value of "1," which is the logical sum of the value "1" assigned to "advance" and the value "0" assigned to "attack." When identifying an object that is "attacking" while "advancing" as an object that is "attacking," the map generation unit 122 assigns the object a value of "0," which is the logical product of the value "1" assigned to "advance" and the value "0" assigned to "attack."
[0027] The feature map may also include multiple channels including a third channel. The third channel is a channel generated by the map generating unit 122 by processing an area of the environment indicated by the environmental information that does not correspond to the position information of the object. For example, the map generation unit 122 assigns a common numerical value different from the numerical value assigned to the object to areas that do not correspond to the position information of an object in an image that shows the environment in a multi-valued image format, and sets the parts other than the object to the same color, etc. This makes it possible to generate a third channel that includes an image that clearly shows the presence of the object. Note that the values assigned to areas that do not correspond to the position information of the object do not necessarily have to all be the same value, as long as they are close to each other. As a non-limiting example, any value between 1 and 10 may be set.
[0028] Furthermore, the feature map may include multiple channels including a fourth channel. The fourth channel is a channel generated by the map generating unit 122 by processing an area of the environment indicated by the environmental information that does not correspond to the behavior information of the object. For example, the map generation unit 122 assigns a common numerical value different from the numerical value assigned to the object to areas that do not correspond to the behavioral information of an object in an image that shows the environment in a multi-valued image format, and sets the parts other than the object to the same color, etc. This makes it possible to generate a fourth channel that includes an image that clearly shows the presence of an object performing a specific behavior. Note that the values assigned to areas that do not correspond to the behavioral information of an object do not necessarily have to all be the same value, as long as they are close to each other. As a non-limiting example, any value between 1 and 10 may be set.
[0029] Furthermore, the numerical value assigned to an object may be a value that corresponds to the meaning of the object's action. For example, if the object represents a ball, a value corresponding to the strength or angle at which the ball was thrown may be assigned to the object. In this way, the feature map generated by the map generating unit 122 includes feature amounts having information for each channel.
[0030] (Map processing section) The map processing unit 123 processes a part of the feature map. "Processing" also includes making changes to the information set in a channel in the feature map. That is, since each channel included in the feature map has information such as the position or behavior of an object, it is possible to make changes to the information included in a specific channel. The map processing unit 123 processes at least the first channel out of the multiple channels included in the feature map. The first channel is a channel in which information for which the degree of contribution to environmental inference is to be determined is set, and is one channel included in one feature map. Note that the channels to be processed may be multiple channels including the first channel.
[0031] The change may be made by masking the object. Masking is a process of eliminating the presence of the object, for example, by changing the numerical value assigned to the object to the same value as the numerical value assigned to everything else in the environment other than that object. For example, when masking an object in a binary image in which an object is assigned a value of "1" and other objects are assigned a value of "0," thereby rendering the image area where the object exists in white and the image area where the object does not exist in black, the map processing unit 123 assigns the value "0" to the object. As a result, as shown in Fig. 2, the image area in the binary image where the object exists also becomes black, and the object is masked. Furthermore, for example, in a binary image in which an object is assigned a value of "0" and everything else is assigned a value of "1," so that image areas where objects exist are represented in black and image areas where objects do not exist are represented in white, the object may be masked by assigning the value "1" to the object. The map processing unit 123 may assign 0.5, which is the intermediate value in the range of 0 to 1, to the object or other parts in the environment to be inferred, or may assign other values. Note that masking is a process for eliminating the presence of an object, and while it has been described as, for example, making the numerical value assigned to the object the same as the numerical value assigned to everything other than that object in the environment, this is not limited to this. As a non-limiting example, in the case of a multi-value image, the numerical value assigned to the object may be a value close to the numerical value assigned to everything other than that object in the environment. For example, a value within a range of plus or minus 5 of the value assigned to everything other than that object in the environment may be set.
[0032] (Inference part) The inference unit 124 receives as input at least a portion of a first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information and second inference information. The first inference information is information related to the inference of the environment based on at least a portion (one channel) of the first feature map. The second inference information is information related to the inference of the environment based on the second feature map. For example, when the simulation environment of a competitive game is the inference target, the inference information may include the probability that an ally will avoid an attack from the opponent, the probability that the ally will directly attack the opponent, and the probability that the ally will attack the opponent's base, as shown in FIG. 2. Since the feature map includes observation data of the real environment indicated by the environmental information or data indicating the simulated environment, the first inference information and second inference information inferred by the inference unit 124 based on this feature map are physical information that contribute to the inference results of the environment. Therefore, by comparing the first inference information and the second inference information, it is possible to grasp the extent to which the physical information that contributes to the inference result of the environment contributes to the inference result.
[0033] The inference unit 124 may include a trained model. The trained model outputs an inference result of the environment indicated by the environmental information when information included in the feature map is input. This trained model is generated by supervised learning using the information included in the feature map as training data, and is a machine learning model from which interpretability is to be obtained. "Interpretability" indicates the degree of influence that information input to a machine learning model has on the inference results of the machine learning model. The trained model is stored in the memory unit 13 shown in FIG. 1, or in an external storage device that the inference unit 124 can access via the communication unit 11 over a communication network. The inference unit 124 acquires the trained model from the memory unit 13 or the external storage device, and performs inference on the environment based on the acquired trained model.
[0034] As shown in Figure 2, the inference unit 124 compares first inference information, which is the inference result output from a trained model to which a first feature map that has not been processed by the map processing unit 123 has been input, with second inference information, which is the inference result output from the trained model to which a second feature map that has been processed by the map processing unit 123 has been input. Based on the result of comparing the first inference information with the second inference information, the inference unit 124 determines the degree of influence (hereinafter referred to as the degree of contribution) that the information of the processed channel has on the inference result by the trained model. The inference unit 124 is capable of interpreting how the elements included in the environmental information influence the inference result. In FIG. 2, the trained model is referred to as a "model." It will also be referred to as a "model" in the following figures.
[0035] The inference unit 124 performs inference using, of the feature amount maps generated by the map generation unit 122, a feature amount map including a first channel that has not been processed by the map processing unit 123, and a feature amount map including a first channel that has been processed by the map processing unit 123. The inference unit 124 may also perform inference using all feature amount maps generated by the map generation unit 122 and all feature amount maps including feature amount maps whose channels have been processed by the map processing unit 123.
[0036] The inference information may be a prediction probability output from a trained model, or may be an index for measuring prediction performance. For example, when the environment includes multiple classes, the inference unit 124 may output first index information, which is an index for measuring the inference performance of each of the multiple classes, based on the first inference information and information on the correct class of the environment included in the environment information, and may output second index information, which is an index for measuring the inference performance of each of the multiple classes, based on the second inference information and information on the correct class of the environment included in the environment information. The first index information and the second index information are indexes for measuring prediction performance, and may be, for example, the accuracy rate of each class, the accuracy rate of all classes, the recall rate, the precision rate, or the F1 value, etc. Furthermore, the inference unit 124 may acquire the accuracy rates output from the trained model when a feature map with a change made to the channel and a feature map with no change made to the channel are input, based on the inference result output from the trained model and the correct label of the dataset input to the trained model, and calculate the difference between these accuracy rates as the contribution.
[0037] (output section) The output unit 125 outputs information to the display unit 14. For example, the output unit 125 acquires first inference information and second inference information from the inference unit 124, and outputs the acquired information to the display unit 14. The display unit 14 displays the information output from the output unit 125. The information displayed on the display unit 14 may display the outputs of the inference unit 124 according to whether or not there is a change in the feature map, or may display the difference between these outputs. The information processing device 1 is required to have at least the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, and the inference unit 124, and the output unit 125 is an optional function. For example, the function of the output unit 125 may be a function provided in the inference unit 124.
[0038] (program) 4 is a block diagram showing a hardware configuration that realizes the functions of the information processing device 1. For example, the information processing device 1 has, as its hardware configuration, a communication interface 100, an input / output interface 101, a processor 102, and a memory 103. The communication unit 11 shown in FIG. 1 is a communication device that communicates with an external device of the information processing device 1 via the communication interface 100. The calculation unit 12 shown in FIG. 1 outputs display information to the display unit 14 shown in FIG. 1 via the input / output interface 101. The calculation unit 12 is the processor 102. The memory 103 is the storage unit 13 shown in FIG. 1.
[0039] The programs constituting the information processing application are stored in the memory 103. The processor 102 reads and executes the programs stored in the memory 103, thereby realizing the functions of the environmental information acquisition unit 121, map generation unit 122, map processing unit 123, inference unit 124, and output unit 125 provided in the calculation unit 12. The memory 103 is, for example, an HDD or SSD. The memory 103 may also be provided in an external device with which data can be exchanged via the communication interface 100.
[0040] For example, the environmental information acquisition unit 121 acquires environmental information received by the communication unit 11 via the communication interface 100. The output unit 125 outputs the inference information as display information to the display unit 14 via the input / output interface 101. The output unit 125 may also transmit the inference information to an external device via the communication interface 100 by the communication unit 11.
[0041] Next, a specific example of the operation of the information processing device 1 will be described. 5A, 5B, and 5C are diagrams showing basketball scenarios, each showing a scenario of teammate and opponent movements on a basketball court. For example, scenarios that the opponent can take in a basketball simulation include "emphasis on one-on-one" as shown in FIG. 5A, "emphasis on basket" as shown in FIG. 5B, and "emphasis on outside attacks" as shown in FIG. 5C. The inference unit 124 distinguishes between these scenarios.
[0042] Figures 5A, 5B, and 5C show the movement of players and the ball on a basketball court. In Figures 5A, 5B, and 5C, the thick circle represents the ball. The white circles represent teammates and are assigned consecutive numbers from 1 to 5. The black circles represent opponents and are also assigned consecutive numbers from 1 to 5. The "one-on-one focused" scenario shown in Figure 5A shows a strategy in which an opposing player marks a teammate one-on-one, and the opposing player number 3 dribbles past the teammate number 3 and takes a two-point shot from under the basket. The "Emphasis on the basket" scenario shown in Figure 5B is similar to Figure 5A in that the opposing player marks a teammate one-on-one, but the opposing player number 4 passes the ball to the opposing player number 5 who is under the basket, and the opposing player number 5 receives the pass and takes a two-point shot from under the basket. The "outside attack emphasis" scenario shown in Figure 5C depicts a strategy in which an opposing player marks a teammate one-on-one, and the opposing player number 4 attempts a three-point shot from outside the three-point line.
[0043] FIG. 6 is a schematic diagram illustrating an overview of basketball scenario identification by the information processing device 1. The basketball simulation environment uses a learning environment that is often used in reinforcement learning. For example, the basketball simulation environment is realized following the Open AI Gym format. In the reinforcement learning environment, the environment information acquisition unit 121 acquires, by an observation function, time-series environment information during the simulation, information indicating the positions or actions of opposing players and friendly players. A scenario is information in which a data set provided for executing a basketball simulation is classified into classes, and is environmental information that indicates one scene to be simulated. Classes include, for example, "emphasis on one-on-one play," "emphasis on basket play," and "emphasis on outside attacks." A scenario may also include time-series information.
[0044] The "actions" in the basketball simulation include at least an action of a player making a two-point shot, an action of a player making a three-point shot, an action of a player dribbling, an action of a player passing, and the like. That is, the actions performed by players on both sides within the environment representing the basketball court are the objects of simulation.
[0045] 6, the environmental information acquired by the environmental information acquisition unit 121 is a plurality of time-series scenarios classified into a plurality of classes. Based on this environmental information, the map generation unit 122 generates a time-series feature map for each scenario, as shown in Fig. 6. The map processing unit 123 adds changes to part of the feature map generated by the map generation unit 122. Also, in Figure 6, the softmax function is a function used when predicting multiple classes. It normalizes the input value and converts it into a range from 0 to 1, and has the characteristic that the sum of all the values becomes 1. By using the softmax function, the probability of belonging to each class can be calculated, and the class with the highest probability becomes the predicted class. With argmax, the class with the highest probability is selected from the output value of the softmax function, and this class can be assigned as the identification label.
[0046] When an unaltered feature map and an altered feature map are input, the trained model included in the inference unit 124 outputs, for example, using a softmax function, a predicted probability as first inference information based on the unaltered feature map, and a predicted probability as second inference information based on the altered feature map. The predicted probability is the probability that an environment will be inferred to be a specific class, for example, the probability that a scenario will be inferred to be a specific class.
[0047] The inference unit 124 assigns an identification label indicating the class of the scenario based on the predicted probability using the argmax model, and calculates the accuracy rate based on the identification label and the correct class of that scenario.The inference unit 124 then compares the accuracy rate based on the feature map without any changes with the accuracy rate based on the feature map with changes, and based on this comparison, it is possible to grasp the degree of contribution of the changed information to the inference result.
[0048] The environmental information may be set by a user using an input device (not shown). In this case, the environmental information may include information other than the positions and actions of the players and the ball. Examples of the information include the remaining game time, the size of the court, the distance between the players and the goal, and the remaining physical strength of each player.
[0049] Furthermore, the dataset of environmental information may be provided with information indicating the enemy's strategy (correct class) in the environment (scenario) indicated by this environmental information. By comparing the output (discrimination class) of a machine learning model to which a feature map converted from certain environmental information is input, with the correct class provided for that environmental information, it is possible to calculate the accuracy rate for each class.
[0050] FIG. 7 is a schematic diagram illustrating an outline of a feature map converted from a basketball scenario. For example, the map generation unit 122 converts all or part of the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in FIG. 7. If the environmental information includes teammate and opponent players and the ball position, a pixel value of 1 to 255 representing 256 gradations of color is assigned to the environmental information. If the teammate and opponent players are not present, a pixel value of 0 (black) representing 256 gradations is assigned to the environmental information. Furthermore, teammate and opponent players may be assigned numbers according to their roles. For example, the point guard may be assigned the number 1, the shooting guard may be assigned the number 2, the small forward may be assigned the number 3, the power forward may be assigned the number 4, and the center may be assigned the number 5.
[0051] 7 is information consisting of three channels (ch1, ch2, ch3) in which object position information is set, and four channels (ch4, ch5, ch6, ch7) in which object behavior information is set. The map generation unit 122 sets "1" to the position information set in channels ch1, ch2, and ch3 if an object exists, and sets "0" if the object does not exist.
[0052] The map generation unit 122 sets action information indicating an action of a two-point shot by an opposing player for channel ch4, and sets action information indicating an action of a three-point shot by an opposing player for channel ch5. The map generation unit 122 sets action information indicating an action of a dribble by an opposing player for channel ch6, and sets action information indicating an action of a pass by an opposing player for channel ch7. Furthermore, if the behavior indicated by the behavior information set on channels ch4, ch5, ch6, and ch7 is being performed, a numerical value such as the above-mentioned 1 to 255 is set, or a numerical value corresponding to the strength of the behavior is set, and if the behavior is not being performed, a value of 0 is set.
[0053] The map processing unit 123 performs processing such as changing at least a specific feature amount map among the feature amount maps generated by the map generating unit 122. In the feature amount map, position information or behavior information of an object is set for each channel. The map processing unit 123 can change the specific information set for the channel.
[0054] FIG. 8A shows a feature map converted from environmental information indicating a basketball simulation environment, and FIG. 8B is a conceptual diagram showing an outline of the process of masking channel ch5 for "three-point shot" included in the feature map shown in FIG. 8A. Here, the feature map shown in FIG. 8A is a time-series feature map converted from multiple scenarios shown in FIG. 8B. For example, it is expected that the identification of the scenario "emphasis on outside attacks" shown in FIG. 5C will be influenced by players who perform "three-point shots" in this scenario. Therefore, the map processing unit 123 masks the behavior information for "three-point shot," which is set to channel ch5 among the channels included in the feature map.
[0055] For example, a player taking a three-point shot set in channel ch5 is assigned a value of "1," and other parts of the basketball court are assigned a value of "0." As a result, in the binary image of the basketball court included in channel ch5, the image area where the player taking a three-point shot is located is displayed in white, and the other areas are displayed in black, as shown in Fig. 8B. The map processing unit 123 assigns the same value "0" to this player as to the other parts, thereby masking this player as shown in FIG. 8B.
[0056] The inference unit 124 outputs first inference information based on the feature map to which no changes have been made by the map processing unit 123, and outputs second inference information based on the feature map to which changes have been made by the map processing unit 123. For example, if the inference unit 124 includes a trained model that identifies basketball tactics (scenarios), when a feature map including channel ch5 in which a player taking a three-point shot is not masked is input, the trained model outputs first inference information, which is the accuracy rate of each tactic, as shown in FIG. 8B. Furthermore, when a feature map including channel ch5 in which a player taking a three-point shot is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the accuracy rate of each tactic, as shown in FIG. 8B. The tactics include "emphasis on one-on-one," "emphasis on basket," and "emphasis on outside attacks," and the inference information is the accuracy rate of each tactic.
[0057] For example, suppose that the first inference information based on a feature map including channel ch5, where the player taking a three-point shot is not masked, has an accuracy rate of 0.92 for "one-on-one focus," 0.94 for "focus on basket," and 0.85 for "focus on outside attack." Also, suppose that the second inference information based on a feature map including channel ch5, where the player taking a three-point shot is masked, has an accuracy rate of 0.88 for "one-on-one focus," 0.94 for "focus on basket," and 0.32 for "focus on outside attack."
[0058] In this case, by masking players who are taking three-point shots, the accuracy rate of "emphasis on outside attack" changes (decreases) significantly from 0.85 to 0.32, so it can be determined that players who are taking three-point shots have a significant impact on the prediction of "emphasis on outside attack." In this way, the information processing device 1 can acquire channel information that is important for each class. The trained model was generated by converting the training dataset into a feature map in advance and using the feature map as input through supervised learning. For example, 3DCNN may be used to generate the trained model.
[0059] In addition, the map processing unit 123 masks various channel information by, for example, changing the channel information to be masked, and by understanding the degree of contribution to each class, it is possible to obtain channel information that contributes greatly to the inference result, and it is also possible to obtain channel information that contributes little to the inference result. The position information or behavior information of an object included in a channel that has a high contribution to the inference result is physical information about the environment in which the object exists. Based on this information, it is possible to understand the class of the inference target and its contribution to the inference result.
[0060] (Output information) The output unit 125 outputs the first inference information and the second inference information output from the inference unit 124 to the display unit 14. For example, the output unit 125 outputs both the accuracy rate of the strategy identification based on the feature map without any changes to channel ch5 and the accuracy rate of the strategy identification based on the feature map with changes to channel ch5 to the display unit 14. The display unit 14 displays each accuracy rate as shown in FIG. 8B. By visually checking the accuracy rates displayed on the display unit 14, the impact of the information with the changes made to channel ch5 on the strategy identification can be recognized. Furthermore, the output unit 125 may output either the accuracy rate of the tactic identification based on the feature map with no change made to channel ch5 or the accuracy rate of the tactic identification based on the feature map with a change made to channel ch5 to the display unit 14. The display unit 14 displays the information output from the output unit 125. Furthermore, the output unit 125 may output information combining the accuracy rate of tactical identification based on the feature map with no change made to channel ch5 and the accuracy rate of tactical identification based on the feature map with a change made to channel ch5 to the display unit 14. The display unit 14 displays the information output from the output unit 125.
[0061] The output unit 125 may output information based on the first inference information and the second inference information output from the inference unit 124 to the display unit 14. For example, the output unit 125 calculates the difference between the accuracy rate of tactic identification based on a feature map with no change made to channel ch5 and the accuracy rate of tactic identification based on a feature map with a change made to channel ch5, and outputs information about the calculated difference to the display unit 14. The display unit 14 displays the information about the difference. By visually checking the difference information displayed on the display unit 14, it is possible to grasp the extent to which the information set on channel ch5 affects the inference. In addition, the information based on the first inference information and the second inference information is not limited to differential information, and the value calculated by adding, multiplying, or averaging the two may be used as an index value for the degree of contribution.
[0062] The output unit 125 may output information of the first channel to the display unit 14. The first channel is a channel that has been subjected to processing such as change by the map processing unit 123. For example, the output unit 125 outputs information of channel ch5 before the change is applied to the display unit 14. The display unit 14 displays information on channel ch5. By visually checking the information of channel ch5 displayed on display unit 14, it can be recognized that the information that influences the prediction of "emphasis on outside attacks" is channel ch5.
[0063] (Information processing method) FIG. 9 is a flowchart showing an information processing method according to the first embodiment. The environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment (step ST1). The map generation unit 122 generates a first feature map including multiple channels set based on the position information or behavior information of the objects (step ST2). The map processing unit 123 processes at least a first channel out of the multiple channels included in the first feature map (step ST3). The inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information that is information related to the inference of the environment based on at least a portion of the first feature map and second inference information that is information related to the inference of the environment based on the second feature map (step ST4). By having the information processing device 1 execute the information processing method shown in FIG. 20, it is possible to acquire physical information in the environment that affects the inference result of the environment.
[0064] (Variation 1) In the explanation so far, we have used a dataset of environmental information showing multiple scenarios to evaluate the degree to which channel information in the feature map contributes to the inference results. Next, a case will be described in which environmental information indicating one scenario is used to evaluate the degree of contribution of channel information in a feature map to an inference result. Specifically, it calculates the predicted probability of which class a scenario belongs to. The contribution of this channel information to the inference result is evaluated based on the difference between the predicted probability based on a feature map with no channel changes and the predicted probability based on a feature map with channel changes.
[0065] 10 is a conceptual diagram showing an outline of processing using one scenario, in which a basketball simulation environment is the inference target. In FIG. 10, the environment information acquisition unit 121 acquires a data set of one scenario as environment information. For example, the scenario is "emphasis on one-to-one". Based on this environment information, the map generation unit 122 generates a time-series feature amount map for the "emphasis on one-to-one" scenario, as shown in FIG. 10. This feature amount map is a feature amount map including seven channels, similar to FIG. 8A. The map processing unit 123 adds changes to a part of the feature amount map generated by the map generation unit 122.
[0066] When an unaltered feature map and a varied feature map are input, the trained model included in the inference unit 124 uses, for example, a softmax function to output a predicted probability as first inference information based on the unaltered feature map, and a predicted probability as second inference information based on the varied feature map. Based on the result of comparing the predicted probability based on the unaltered feature map with the predicted probability based on the varied feature map, the inference unit 124 can grasp the degree of contribution of the varied information to the inference result.
[0067] Fig. 11A shows a feature map converted from environmental information showing a basketball simulation environment. Fig. 11B is a conceptual diagram showing an outline of a process for masking channel ch2, which is set to "enemy position" and is included in the feature map shown in Fig. 11A. The feature map shown in Fig. 11A is a time-series feature map converted from time-series environmental information of one scenario (one-to-one focused) shown in Fig. 11B. Channel ch2 included in the feature map for each time is the first channel in which position information indicating the positions of five enemy players (1) to (5), numbered 1 to 5 as "enemy positions," is set.
[0068] The map processing unit 123 processes the part of the object included in the first channel, thereby making it possible to evaluate the influence of the part of the object included in the first channel on the inference result. For example, it is expected that the prediction for the scenario "Focus on one-on-one" will be influenced by the position information of the opponent (3) player who shoots after dribbling in the scenario. Therefore, in the example of Fig. 11B, the map processing unit 123 applies a mask to the position of the opponent (3) player among the five opponent players (1) to (5). The position information of the opponent players (1) to (5) included in channel ch2 is assigned a value of "1," and the other parts of the basketball court are assigned a value of "0." As a result, as shown in Fig. 11B, in the binary image of the basketball court included in channel ch2, the image areas where the opponent players (1) to (5) are present are displayed in white, and the other areas are displayed in black. When masking the position information of the opponent (3) player, the map processing unit 123 masks this player by assigning the same value "0" as the other parts to the position of the opponent (3) player who is shooting from a dribble among the opponent (1) to (5) players, as shown in FIG. 11B.
[0069] The inference unit 124 outputs first inference information based on the feature map to which no changes have been made by the map processing unit 123, and outputs second inference information based on the feature map to which changes have been made by the map processing unit 123. For example, if the inference unit 124 is a trained model that predicts basketball strategies (scenarios), when a feature map including at least channel ch2 in which the position information of the opponent (3) player is not masked is input, the trained model outputs first inference information, which is the predicted probability of each scenario shown in Fig. 11B. Also, when a feature map including channel ch2 in which the position information of the opponent (3) player is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the predicted probability of each scenario shown in Fig. 11B.
[0070] The inference unit 124 compares the predicted probability based on a feature map (first feature map) including channel ch2 in which the position information of the enemy (3) player is not masked with the predicted probability based on a feature map (second feature map) including channel ch2 in which the position information of the enemy (3) player is masked. The predicted probability is the probability of inferring that the environment (scenario) is a specific class. The influence of the channel information included in the feature map on the predicted probability can be evaluated.
[0071] In Figure 11B, the predicted probability of "emphasis on one-on-one" based on the feature map including channel ch2 where the position information of the enemy (3) player is not masked is 0.82, while the predicted probability of "emphasis on one-on-one" based on the feature map including channel ch2 where the position information of the enemy (3) player is masked is 0.56, showing a large change in predicted probability. In this way, the inference unit 124 receives the first feature amount map and the second feature amount map as input, and outputs first inference information and second inference information, which are the probabilities (predicted probabilities) of inferring that the environment belongs to a specific class.
[0072] The output unit 125 calculates the difference between the predicted probability based on the feature map including channel ch2 in which the position information of the enemy (3) player is not masked and the predicted probability based on the feature map including channel ch2 in which the position information of the enemy (3) player is masked, and outputs information about the calculated difference to the display unit 14. The display unit 14 displays the information about the difference. By visually checking the difference information displayed on the display unit 14 and masking the position information of the opponent (3) player, it can be seen that the difference in predicted probability in the scenario "Focus on one-on-one" decreased by 0.26, the difference in predicted probability in the scenario "Focus on basket" increased by 0.22, and the difference in predicted probability in the scenario "Focus on outside attack" increased by 0.04. In other words, the predicted probability in the scenario "Focus on one-on-one" changed (decreased) the most, and it can be seen that the opponent (3) player taking a shot from a dribble contributed greatly to the prediction in "Focus on one-on-one".
[0073] The output unit 125 may output the information set in the changed channel ch2 and information indicating that the masked object is the position information of the enemy (3) player to the display unit 14. The display unit 14 displays the information output from the output unit 125. By visually checking the information displayed on the display unit 14, it can be recognized that the information that influences the "one-on-one focused" prediction is the position information of the enemy (3) player included in channel ch2. The output unit 125 may also output either the information set in the changed channel ch2 or the information indicating that the masked object is the position information of the enemy (3) player to the display unit 14. The display unit 14 displays the information output from the output unit 125. Furthermore, the output unit 125 may output information that combines the information set in the changed channel ch2 and information indicating that the masked object is the position information of the enemy (3) player to the display unit 14. The display unit 14 displays the information output from the output unit 125.
[0074] (Variation 2) Furthermore, the map processing unit 123 may process the part of the object that is performing dynamic actions among the objects included in the first channel. For example, as described above, the map processing unit 123 identifies the opponent (3) player, who is the object that is performing a shot from a dribble, among the position information of the opponent (1) to (5) players set in channel ch2, and masks only the position information of the opponent (3) player in channel ch2. In other words, the position information of this object is the part of the object that is performing dynamic actions. Even if the specific behavior is not clear, by masking only the location information of objects that are performing dynamic behavior, it is possible to evaluate the impact of this object's location information on the inference results.
[0075] Fig. 12A shows a feature map converted from environmental information indicating a basketball simulation environment, and is the same feature map as Fig. 11A. Fig. 12B is a conceptual diagram showing an outline of a process of applying a mask to channel ch2, to which the "enemy position" included in the feature map shown in Fig. 12A is set. The feature map shown in Fig. 12A is a time-series feature map converted from the time-series environmental information of one scenario (emphasis on one-to-one) shown in Fig. 12B. Channel ch2 included in each feature map at each time is a first channel in which position information indicating the positions of five enemy players (1) to (5) assigned numbers 1 to 5 as "enemy positions" is set.
[0076] For example, in the scenario "Focus on one-on-one", the position information of the enemy (4) player who does not take any specific action and moves little is masked. In this case, the map processing unit 123 masks the position information of the enemy (4) player among the five enemy players (1) to (5) whose position information is set to channel ch2. The position information of the opponent players (1) to (5) included in channel ch2 is assigned a value of "1," and the other parts of the basketball court are assigned a value of "0." As a result, as shown in Fig. 12B, in the binary image of the basketball court included in channel ch2, the image areas where the opponent players (1) to (5) are present are displayed in white, and the other areas are displayed in black. When masking the position information of the enemy (4) player, the map processing unit 123 masks this player by assigning the same value "0" to the position information of the enemy (4) player among the enemy (1) to (5) players, as is the case with the other players, as shown in FIG. 12B.
[0077] The inference unit 124 outputs first inference information based on the feature map to which no changes have been made by the map processing unit 123, and outputs second inference information based on the feature map to which changes have been made by the map processing unit 123. For example, if the inference unit 124 is a trained model that predicts basketball strategies (scenarios), when a feature map including at least channel ch2 in which the position information of the opponent (4) player is not masked is input, the trained model outputs first inference information, which is the predicted probability of each scenario shown in Fig. 12B. When a feature map including channel ch2 in which the position information of the opponent (4) player is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the predicted probability of each scenario shown in Fig. 12B.
[0078] The inference unit 124 compares the predicted probability based on a feature map (first feature map) including channel ch2 in which the position information of the enemy (4) player is not masked with the predicted probability based on a feature map (second feature map) including channel ch2 in which the position information of the enemy (4) player is masked. The predicted probability is the probability of inferring that the environment (scenario) is of a specific class. The influence of the channel information included in the feature map on the predicted probability can be evaluated.
[0079] 12B, the predicted probability of "emphasis on one-on-one play" based on a feature map including channel ch2 in which the position information of the opponent (4) player is not masked is 0.82, the predicted probability of "emphasis on under the basket" is 0.10, and the predicted probability of "emphasis on outside attacks" is 0.10. Also, the predicted probability of "emphasis on one-on-one play" based on a feature map including channel ch2 in which the position information of the opponent (4) player is masked is 0.80, the predicted probability of "emphasis on under the basket" is 0.12, and the predicted probability of "emphasis on outside attacks" is 0.08. In this way, even if the position information of the enemy (4) player, who does not take any specific action and moves little, is masked, the change in the prediction probability is small, and it can be seen that the contribution of the position information of the enemy (4) player to the prediction of the "one-on-one emphasis" scenario using the discriminative model is small.
[0080] In addition, the map processing unit 123 masks various pieces of information, for example by changing the information to be masked among the pieces of information set for one channel, and by grasping the degree of contribution to each class, it is possible to obtain channel information that contributes greatly to the inference result, and as described above, it is also possible to obtain channel information that contributes greatly to the inference result, and it is also possible to obtain channel information that contributes little.
[0081] In this way, the information processing device 1 can acquire channel information or information on objects included in channels that significantly contributed to changes in predicted probability based on the results of comparing predicted probabilities based on feature maps including channels that have not been changed with predicted probabilities based on feature maps including channels that have been changed. Furthermore, the extent to which the channel information or information on objects included in channels affects the inference results can be determined based on the results of the comparison of predicted probabilities.
[0082] The information processing device 1 applies processing (changes) to at least the channel portion included in the feature map converted from a single scenario, and obtains information that contributed to the change in predicted probability based on the result of comparing the predicted probability based on the feature map that has not been processed with the predicted probability based on the feature map that has been processed. In addition, the information processing device 1 may apply processing (changes) to at least the channel portion included in the feature maps converted from multiple scenarios, and calculate, for each scenario, the difference between the predicted probability based on the feature map that has not been processed and the predicted probability based on the feature map that has been processed as the contribution for each scenario. Furthermore, the information processing device 1 may calculate the contributions to all scenarios assumed in the simulation based on the sum, product, or average of the prediction probabilities calculated for each scenario, etc. This allows the information processing device 1 to grasp the channels that contribute to predictions in multiple scenarios.
[0083] As a method of adding a change to the feature amount map, for example, a change may be added to specific information in the time-series feature amount map, or a change may be added to information at a specific time, as shown in Figures 13A, 13B, and 13C. This allows the information processing device 1 to evaluate the influence of specific information in the time-series feature amount map on inference, and also to evaluate the influence of information at a specific time on inference. 13A shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all of the information relating to actions has been changed. For example, the map processing unit 123 applies changes to each of the channels in which information relating to actions has been set, namely, channel ch4 representing "two-point shot," channel ch5 representing "three-point shot," channel ch6 representing "dribble," and channel ch7 representing "pass," in the channel information included in the feature map shown in FIG. 13A.
[0084] 13B shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all information at a specific time is changed. For example, the map processing unit 123 changes all of the channel information for channels ch1 to ch7 at the specific time included in the feature map shown in FIG. 13B.
[0085] 13C shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all information relating to a specific opponent has been changed. For example, the map processing unit 123 changes each of the channel information relating to the specific opponent included in the feature map shown in FIG. 13C, including channel ch2 indicating "opponent position," channel ch3 indicating "ball position," channel ch4 indicating "two-point shot," channel ch5 indicating "three-point shot," channel ch6 indicating "dribble," and channel ch7 indicating "pass."
[0086] Changes may be made to the time-series feature map shown in Figures 13A, 13B, and 13C, or these may be combined. Furthermore, the methods of making changes shown in Figures 8A, 8B, and 8C, the methods of making changes shown in Figures 11A, 11B, and 11C, and the methods of making changes shown in Figures 12A, 12B, and 12C may be combined, or these may be combined with the methods of making changes shown in Figures 13A, 13B, and 13C. This enables the information processing device 1 to evaluate the influence of specific information in the time-series feature map on inference, and to evaluate the influence of information at a specific time on inference.
[0087] (Variation 2) 14A and 14B are diagrams showing an overview of a fighting game between bees and hornets. FIG. 14A shows a simulation environment in which bees and hornets fight, and FIG. 14B shows rules for the actions of each agent in the simulation environment shown in FIG. 14A. A simulator for a fighting game between bees and hornets operates in a reinforcement learning environment, similar to the simulation environment for basketball described above, for example. In this game, each object behaves based on a rule base.
[0088] As shown in FIG. 14A, this battle environment includes competing bees and hornets, flowers from which the bees gather resources, a hornet's nest, and the beehive itself as agents (objects). As shown in FIG. 14B, hornets have three behavioral patterns: an "avoidance type" that prioritizes fleeing from bees; a "direct attack type" that prioritizes approaching and attacking bees; and a "base attack type" that prioritizes approaching and attacking the beehive. The hornets behave based on one of these behavioral patterns. The bees acquire positional information and behavioral information of the objects from the simulator and, based on the acquired information, identify which behavioral pattern the hornets' behavior is. The information acquired from the simulator may be all information in the battle environment, or it may be limited information, such as only information identified within the bees' field of vision.
[0089] FIG. 15 is a schematic diagram illustrating an overview of a process for identifying multiple scenarios of a fighting game between bees and hornets using a machine learning model. In FIG. 15, environmental information indicating multiple scenarios as the environment of the fighting game is time-series information consisting of environmental information for each time, and includes, for example, position history information indicating the position of an object at each time or behavior history information indicating the behavior of an object at each time. The environmental information acquisition unit 121 acquires the time-series environmental information from a simulator. The map generation unit 122 generates channels for each time based on the environmental information for each time, thereby generating a feature map consisting of time-series channels, as shown in FIG. 15.
[0090] The map processing unit 123 processes at least one channel out of multiple channels included in the feature map. For example, the map processing unit 123 applies a change (e.g., a mask) to an object included in a channel. The inference unit 124 receives as input at least a portion of the feature map including channels to which no change has been applied and a feature map including channels to which changes have been applied by the map processing unit 123, and outputs first inference information and second inference information. For example, as shown in FIG. 15 , the first inference information and the second inference information include the accuracy rate of hornets avoiding attacks from honeybees, the accuracy rate of hornets approaching honeybees and attacking them directly, and the accuracy rate of hornets approaching beehives and attacking them. The inference unit 124 finds the degree of contribution that the information of the changed channel makes to the inference result based on the result of comparing the accuracy rate, which is the first inference information, with the accuracy rate, which is the second inference information.
[0091] FIG. 16 is a schematic diagram showing an outline of a feature map converted from a scenario of a fighting game in which bees and hornets fight. The map generation unit 122 converts all or part of the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in FIG. 16. The environmental information is assigned a "1" to the location of the hornet, the location of the bee, the location of the flower, and the location of the beehive if each of these exists, and a "0" if each does not exist. A "1" is assigned to each of the actions of attacking the bee, attacking the beehive, moving left, moving up, moving right, and moving down, and a "0" is assigned if each of these actions is not performed. Note that the feature map may have information indicating a composite element of movement direction and action content, such as "attack right," set to a channel.
[0092] 16 is information consisting of four channels (ch1, ch2, ch3, ch4) in which object position information is set, and six channels (ch5, ch6, ch7, ch8, ch9, ch10) in which object behavior information is set. That is, the map generating unit 122 sets "1" to the position information set in channels ch1, ch2, ch3, and ch4 if an object exists, and sets "0" if an object does not exist.
[0093] (Variation 3) FIG. 17A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and FIG. 17B is a conceptual diagram showing an outline of the process of applying a mask to channel ch2 for the "position of hornet A" included in the feature map shown in FIG. 17A. Here, the feature map shown in FIG. 17A is a feature map converted from the time-series environmental information for each of the multiple scenarios shown in FIG. 17B. For example, it is expected that the "position of hornet A" will have an impact on identifying the hornet's strategy (scenario). Therefore, the map processing unit 123 applies a mask to the position information "position of hornet A" set to channel ch2 among the channels included in the feature map.
[0094] For example, the position information of the hornet set in channel ch2 is assigned a value of "1," and the rest of the information is assigned a value of "0." As a result, as shown in Fig. 17B, in the binary image of the battle environment included in channel ch2, the image area where the hornet exists is displayed in white, and the rest of the area is displayed in black. The map processing unit 123 assigns the same value "0" to "Wasp A" as to the other parts, thereby masking the position information of "Wasp A" as shown in FIG. 17B.
[0095] If the inference unit 124 includes a trained model that identifies the hornet's tactics, when a feature map including channel ch2 in which the location information of "Hornet A" is not masked is input, the trained model outputs first inference information, which is the accuracy rate of each tactic, as shown in FIG. 17B. Furthermore, when a feature map including channel ch2 in which the location information of "Hornet A" is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the accuracy rate of each tactic, as shown in FIG. 17B. The tactics include "evasion," "direct attack," and "base attack," and the inference information is the accuracy rate of each tactic.
[0096] For example, the first inference information based on a feature map including channel ch2 where the location information of "Hornet A" is not masked is assumed to have an accuracy rate of 97% for "avoidance," 82% for "direct attack," and 92% for "base attack." Also, the second inference information based on a feature map including channel ch2 where the location information of "Hornet A" is masked is assumed to have an accuracy rate of 42% for "avoidance," 80% for "direct attack," and 92% for "base attack."
[0097] In this case, by masking the location information of "Hornet A," the accuracy rate of "Avoid" changes (decreases) significantly from 97% to 42%, so it can be determined that "Hornet A" has a significant influence on the prediction of the hornet's strategy. In this way, the information processing device 1 can acquire channel information that is important for each class. The trained model was generated by converting the training dataset into a feature map in advance and using the feature map as input through supervised learning. For example, 3DCNN may be used to generate the trained model.
[0098] The output unit 125 outputs both the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is not masked and the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is masked to the display unit 14. The display unit 14 displays each accuracy rate as shown in FIG. 17B. By visually checking the accuracy rates displayed on the display unit 14, it is possible to interpret that the trained model has identified the tactic based on the position information of hornet A. In addition, the output unit 125 may output to the display unit 14 either the accuracy rate of the strategy identification based on the feature map in which the "position of hornet A" in channel ch2 is not masked or the accuracy rate of the strategy identification based on the feature map in which the "position of hornet A" in channel ch2 is masked, and the display unit 14 may display the accuracy rate output from the output unit 125. In addition, the output unit 125 may output to the display unit 14 information that combines the accuracy rate of strategy identification based on a feature map in which the "position of hornet A" in channel ch2 is not masked and the accuracy rate of strategy identification based on a feature map in which the "position of hornet A" in channel ch2 is masked, and the display unit 14 may display the information output from the output unit 125.
[0099] (Variation 4) Fig. 18A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and Fig. 18B is a conceptual diagram showing an outline of the process of masking channel ch2 for the positions of all hornets included in the feature map shown in Fig. 18A. Here, the feature map shown in Fig. 18A is a time-series feature map converted from the time-series environmental information for each of the multiple scenarios shown in Fig. 18B. For example, if the accuracy rate of a specific strategy changes when the position information of all hornets is masked, it is possible to interpret that the trained model identifies the hornets' strategies based on the position information of all hornets.
[0100] For example, the position information of the hornet set in channel ch2 is assigned a value of "1," and the rest of the information is assigned a value of "0." As a result, as shown in Fig. 18B, in the binary image of the battle environment included in channel ch2, the image area where the hornet exists is displayed in white, and the rest of the area is displayed in black. The map processing unit 123 assigns the same value "0" to all the hornets as to the other parts, thereby masking the position information of all the hornets as shown in FIG. 18B.
[0101] When a feature map including channel ch2 in which the location information of all hornets has not been masked is input, the trained model outputs first inference information, which is the accuracy rate of each tactic shown in FIG. 18B. Furthermore, when a feature map including channel ch2 in which the location information of all hornets has been masked by map processing unit 123 is input, the trained model outputs second inference information, which is the accuracy rate of each tactic shown in FIG. 18B. The inference information is the accuracy rate of each tactic: "evasion," "direct attack," and "base attack."
[0102] For example, the first inference information based on a feature map including channel ch2 in which the location information of all hornets is not masked shows that the accuracy rate for "avoidance" is 97%, the accuracy rate for "direct attack" is 82%, and the accuracy rate for "base attack" is 92%. Also, the second inference information based on a feature map including channel ch2 in which the location information of all hornets is masked shows that the accuracy rate for "avoidance" is 20%, the accuracy rate for "direct attack" is 30%, and the accuracy rate for "base attack" is 88%.
[0103] In this case, by masking the location information of all hornets, the accuracy rate for "evasion" dropped significantly from 97% to 20%, and the accuracy rate for "direct attack" dropped significantly from 82% to 30%. On the other hand, the accuracy rate for "base attack" only changed from 92% to 88%, a small decrease. In this way, by masking the location information of all hornets contained in channel ch2 multiple times and calculating the contribution to each class, it is possible to grasp information that contributes significantly to predicting hornet strategies, as well as information that contributes little to predicting hornet strategies.
[0104] (Variation 5) FIG. 19A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and FIG. 19B is a conceptual diagram showing an outline of the process of applying a mask to channel ch5 for bee attacks included in the feature map shown in FIG. 19A. Here, the feature amount map shown in Fig. 19A is a time-series feature amount map converted from the multiple scenarios shown in Fig. 19B. For example, the environmental information acquisition unit 121 acquires environmental information indicating one scenario (strategy), and the map generation unit 122 generates a 10-channel feature amount map as shown in Fig. 16 from the environmental information, similar to the multiple scenarios.
[0105] The map processing unit 123 masks all of the behavioral information "bee attack" set in channel ch5. When a feature map including at least channel ch5 in which all of the behavioral information "bee attack" is not masked is input, the trained model included in the inference unit 124 outputs first inference information, which is the predicted probability of each scenario shown in FIG. 19B. When a feature map including channel ch5 in which all of the behavioral information "bee attack" is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the predicted probability of each scenario shown in FIG. 19B.
[0106] 19B, the predicted probability of the scenario "Avoid" based on the feature map including channel ch5 where "Bee Attack" is not masked is 2%, the predicted probability of "Direct Attack" is 92%, and the predicted probability of "Base Attack" is 6%. Also, the predicted probability of "Avoid" based on the feature map including channel ch5 where "Bee Attack" is masked is 42%, the predicted probability of "Direct Attack" is 38%, and the predicted probability of "Base Attack" is 20%.
[0107] In this case, by masking all of the behavioral information of "bee attacks," the prediction rate of "avoidance" changes (increases) significantly from 2% to 42%, the prediction rate of "direct attacks" changes (decreases) significantly from 92% to 38%, and the prediction rate of "base attacks" changes (increases) significantly from 6% to 20%. This allows us to determine that "bee attacks" contribute greatly to the tactical prediction of "direct attacks." In this way, by masking various pieces of information and performing inference, the information processing device 1 can determine channel information that contributes greatly to the prediction of each class, and can also determine channel information that contributes little to the prediction of each class.
[0108] Although the supervised learning has been described in particular as classification (identification), if correct answer data is available, a regression model may be used as the trained model. For example, in regression, the emphasis rate on 3-point shots and the emphasis rate on 2-point shots are expressed by weighting. Let the emphasis rate on 3-point shots be x and let it be [x, 1-x]. Then, x is regressed using machine learning on ([0.9, 0.1]). x was perfectly correct and was ([0.9,0.1]), but when the action information of a three-point shot is masked, it changes to [0.5,0.5], and it can be seen that the channel set with the action information of a three-point shot makes a significant contribution when x = 0.9.
[0109] Furthermore, when environmental information is considered in terms of a data set, important channels can be identified by examining which combinations of channels included in the feature map should be masked to increase the RMSE (root mean square error). The same can be done for clustering, which is unsupervised learning. However, since it is not possible to compare the results with the correct answer, it is not possible to evaluate the accuracy rate. However, the information processing device 1 can output the contribution of each channel by focusing on changes in the predicted probability, thereby making it possible to present the channel that serves as the basis for determining which cluster it belongs to. Furthermore, if the current cluster is assumed to be the correct answer, the contribution of each channel can be output, making it possible to present the channels that provide the basis for why they belong to that cluster.
[0110] (Variation 6) FIG. 20 is a schematic diagram showing an outline of a feature map converted from environmental information indicating the environment of the real space, and shows a view ahead of the vehicle as seen from an autonomous vehicle. The environmental information acquisition unit 121 acquires, as environmental information, a captured image of the area ahead of the vehicle captured by an on-board camera, object class information (e.g., person or car) acquired by object detection ahead of the vehicle, which is one type of machine learning, a parallax image of the area ahead of the vehicle, the distance to an object ahead of the vehicle measured by a sensor mounted on the vehicle, or speed information of an oncoming vehicle. The map generation unit 122 converts the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in FIG. 20.
[0111] 20, the environmental information indicates that at a certain time, there are objects at an intersection ahead of the vehicle, including a person crossing a crosswalk, a vehicle traveling ahead in the same lane, and a vehicle approaching in the oncoming lane. The map generation unit 122 assigns, for example, a value of "1" to the objects and a value of "0" to areas where no objects exist.
[0112] For example, the feature map shown in FIG. 20 includes channels ch1, ch2, and ch3. These channels include binary images that represent the environment of the real space. Channel ch1 indicates the position of a car. In the binary image included in channel ch1, image areas corresponding to the position of the car are represented in white, and image areas where no car is present are represented in black. Channel ch2 indicates the position of a person. In the binary image included in channel ch2, image areas corresponding to the position of the person are represented in white, and image areas where no person is present are represented in black. Channel ch3 indicates an object that is "approaching" the vehicle. In the binary image included in channel ch3, image areas corresponding to the car that is "approaching" among the objects are represented in white, and the other image areas are represented in black.
[0113] When the inference unit 124 includes a trained model, by inputting an unaltered feature map and an altered feature map, it is possible to present information that the machine learning model is focusing on during autonomous driving. Furthermore, when an operation is performed on the vehicle, the information processing device 1 can present the reason for the operation.
[0114] The environment to be inferred may also be a surveillance area monitored by a surveillance camera. The environmental information acquisition unit 121 acquires, as environmental information, captured images of the surveillance area captured by the surveillance camera, object class information (e.g., people or ornaments) acquired by object detection in the surveillance area, which is one of machine learning methods, or human behavior information. The map generation unit 122 converts the environmental information acquired by the environmental information acquisition unit 121 into a feature map. When the inference unit 124 includes a trained model, by inputting an unaltered feature map and an altered feature map, when a surveillance system equipped with a surveillance camera detects a suspicious person, it can present the reason why the trained model detected the suspicious person as being based on what human behavior.
[0115] The object's position information or object's behavior information may be obtained by object recognition from image information showing the environment. The information processing device 1 can be applied to any environment in the real world as long as it is possible to acquire environmental information by using sensor information such as image recognition technology, GPS (Global Poisoning System), ToF (Time of Flight) sensor, or infrared camera. When image information is input as environmental information, the environmental information acquisition unit 121 recognizes objects included in the image information and detects the objects in the environment and their position information. The inference unit 124 predicts the behavior of the objects by image recognition and acquires behavior information indicating the predicted behavior. The environment information acquisition unit 121 acquires image information as environment information, but the position information of the object may be detected using a sensor such as a GPS.
[0116] As described above, the information processing device 1 according to the first embodiment includes: an environment information acquisition unit 121 that acquires environment information including position information or behavior information of objects included in the environment; a map generation unit 122 that generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a map processing unit 123 that processes at least a first channel among the plurality of channels included in the first feature map; and an inference unit 124 that receives as input at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information that is information regarding the inference of the environment based on at least a portion of the first feature map and second inference information that is information regarding the inference of the environment based on the second feature map. Based on the first inference information and the second inference information regarding the environment to be inferred, the information processing device 1 can acquire physical information in the environment that affects the inference result of the environment. The conventional technology described in Non-Patent Document 1 randomly masks an image and searches for pixels that affect the image classification results. Furthermore, pixels have a relative meaning but no spatial or physical meaning. In other words, conventional technology is unable to extract the physical information that contributes to the inference result, and is therefore unable to present what information has an overall impact on the accuracy rate of the inference result, and to what extent. Furthermore, with conventional techniques, it takes a lot of time and requires many trials before it is possible to interpret how pixels affect the inference results. On the other hand, the information processing device 1 can acquire physical information about the environment, and therefore can also solve the above problem.
[0117] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs the output first inference information and second inference information to the display unit 14. By visually checking the first inference information and second inference information displayed on the display unit 14, it is possible to recognize the influence that the information processed in the first feature map has on the inference.
[0118] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs information based on the output first inference information and second inference information to the display unit 14. By visually checking the information displayed on the display unit 14, it is possible to grasp the extent to which the information processed in the first feature map affects the inference.
[0119] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs information of the first channel to the display unit 14. By visually checking the first channel displayed on the display unit 14, it is possible to recognize that the information that influences the inference is the first channel.
[0120] In the information processing device 1 according to the first embodiment, the map processing unit 123 performs processing on the part of the object included in the first channel, thereby making it possible to evaluate the influence of the part of the object included in the first channel on the inference result.
[0121] In the information processing device 1 according to the first embodiment, the map processing unit 123 processes the part of the objects that are performing dynamic actions among the objects included in the first channel, thereby making it possible to evaluate the influence that the part of the objects that are performing dynamic actions among the objects included in the first channel has on the inference result.
[0122] In the information processing device 1 according to the first embodiment, the multiple channels include a third channel. The map generation unit 122 processes an area of the environmental information that does not correspond to the position information, and generates the third channel. This enables the information processing device 1 to generate the third channel including an image that clearly shows the presence of an object.
[0123] In the information processing device 1 according to embodiment 1, the multiple channels include a fourth channel. The map generation unit 122 processes an area of the environmental information that does not correspond to the behavioral information, and generates the fourth channel. This enables the information processing device 1 to generate the fourth channel including an image that clearly shows the presence of an object performing a specific behavior.
[0124] In the information processing device 1 according to the first embodiment, the environment is an environment operated by a simulator. The object position information or object behavior information is information generated by the simulator. This allows the information processing device 1 to acquire physical information in the environment that affects the inference result of the simulator about the environment.
[0125] In the information processing device 1 according to the first embodiment, the position information or behavior information of the object is acquired by object recognition from image information showing the environment. This allows the information processing device 1 to acquire the information obtained by object recognition using the image information as environmental information.
[0126] In the information processing device 1 according to the first embodiment, the inference unit 124 receives the first feature amount map and the second feature amount map, and outputs the first inference information and the second inference information by regression, thereby enabling the information processing device 1 to perform inference by regression.
[0127] In the information processing device 1 according to the first embodiment, the inference unit 124 receives the first feature map and the second feature map, and outputs first inference information and second inference information, which are probabilities of inferring that an environment belongs to a specific class. This makes it possible to evaluate the influence of the channel information included in the feature map on the probability of inferring that an environment belongs to a specific class.
[0128] In the information processing device 1 according to the first embodiment, the environment includes a plurality of classes. The inference unit 124 outputs first index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the first inference information and information on the correct class of the environment included in the environment information, and outputs second index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the second inference information and information on the correct class of the environment included in the environment information. This allows the information processing device 1 to grasp the contribution of information to an inference result based on the first inference information and the second inference information.
[0129] In the information processing device 1 according to the first embodiment, the inference unit 124 includes a trained model, which enables the information processing device 1 to acquire physical information about the environment that affects the inference result about the environment by the trained model.
[0130] The information processing method according to the first embodiment includes a step (ST1) in which an environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment; a step (ST2) in which a map generation unit 122 generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a step (ST3) in which a map processing unit 123 processes at least a first channel out of the plurality of channels included in the first feature map; and a step (ST4) in which an inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information which is information relating to the inference of the environment based on at least a portion of the first feature map, and second inference information which is information relating to the inference of the environment based on the second feature map. By the information processing device 1 executing the information processing method according to the first embodiment, it is possible to acquire physical information in the environment that affects the inference result of the environment.
[0131] A computer that executes a program according to the first embodiment executes the following steps: (ST1) an environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment; (ST2) a map generation unit 122 generates a first feature map including multiple channels set based on the position information or behavior information of the objects; (ST3) a map processing unit 123 processes at least a first channel among the multiple channels included in the first feature map; and (ST4) an inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information that is information regarding environmental inference based on at least a portion of the first feature map and second inference information that is information regarding environmental inference based on the second feature map. By executing this program, the computer functions as an information processing device 1, and is therefore able to acquire physical information in the environment that affects the results of environmental inference.
[0132] Embodiment 2 Fig. 21 is a block diagram showing an example configuration of an information processing device 1A according to embodiment 2. In Fig. 21, the information processing device 1A includes a communication unit 11, a calculation unit 12A, a storage unit 13, and a display unit 14. The communication unit 11 communicates with an external device via a network. The calculation unit 12A controls the overall operation of the information processing device 1A. It should be noted that the information processing device 1A does not necessarily have to include the communication unit 11. For example, when environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1A does not include the communication unit 11, the environmental information acquisition unit 121 realized by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.
[0133] The calculation unit 12A includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123A, an inference unit 124, an output unit 125, and a map processing control unit 126. When the calculation unit 12A executes an information processing application, the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123A, the inference unit 124, the output unit 125, and the map processing control unit 126 are realized. The storage unit 13 stores, for example, an information processing application and information used in the arithmetic processing of the arithmetic unit 12A. The storage unit 13 is a storage device provided in a computer that functions as the information processing device 1A, and is a storage such as an HDD or SSD. Note that the storage unit 13 may be provided outside the information processing device 1A as long as it is accessible by the information processing device 1A. The storage unit 13 is also the memory 103 shown in FIG. 4.
[0134] The map processing control unit 126 performs control to generate a third feature map including at least the second channel, which has been processed by the map processing unit 123A based on the first inference information and the second inference information, among the multiple channels. The first inference information is an inference result output from a trained model to which a first feature map that has not been processed by the map processing unit 123A has been input. The second inference information is an inference result output from the trained model to which a second feature map that has been processed by the map processing unit 123A has been input. The inference unit 124 receives the third feature map and outputs the third inference information, which is information regarding the inference of environmental information based on the third feature map.
[0135] FIG. 22 is a flowchart showing an information processing method according to the second embodiment. The map processing control unit 126 specifies to the map processing unit 123A, among the feature amount maps generated by the map generating unit 122, the channels to be changed (step ST1A). For example, by comprehensively changing the channels included in the feature amount map converted from the environmental information indicating the simulation environment or the real environment in advance and performing the environmental inference shown in the first embodiment, the channel that most contributed to the inference result or the combination of channels that contributed most to the inference result may be searched for. The searched channel or combination of channels is set in the map processing control unit 126. When performing new inference in a simulation environment or a real environment, the map processing control unit 126 designates the previously set channel or combination of channels to the map processing unit 123A as the channel to be changed. Alternatively, the channels or combinations of channels that may be related to the simulation environment or the real environment may be narrowed down in advance, and the map processing control unit 126 may specify the channels to be changed from the previously narrowed down channels or combinations. Note that the channel narrowing down may be specified by the user using an input device (not shown), for example.
[0136] The map processing unit 123A performs processing on a channel designated by the map processing control unit 126 from among a plurality of channels included in the feature map (step ST2A). For example, the map processing unit 123A makes a change to the designated channel. The inference unit 124 receives a feature map including channels that have not been processed and a feature map including channels that have been processed, and outputs first inference information, which is information regarding the inference of the environment based on the feature map including channels that have not been processed, and second inference information, which is information regarding the inference of the environment based on the feature map including channels that have been processed (step ST3A). Next, the inference unit 124 searches for a channel that contributes to the inference result based on the result of comparing the first inference information and the second inference information (step ST4A). For example, if a channel that contributes highly to the inference result is found (step ST4A; YES), the processing of FIG. 22 ends. On the other hand, if the inference unit 124 does not find a channel that contributes highly to the inference result (step ST4A; NO), the processing returns to step ST1A and repeats the series of processing shown in FIG. 22. This makes it possible to find a channel that influences changes in the accuracy rate or prediction probability, which are the inference results by the inference unit 124.
[0137] As described above, the information processing device 1A according to the second embodiment includes the map processing control unit 126 that performs control to generate a third feature map including at least the second channel processed by the map processing unit 123A based on the first inference information and the second inference information among the multiple channels. The inference unit 124 receives the third feature map and outputs the third inference information, which is information related to the inference of environmental information based on the third feature map. This enables the information processing device 1A to find channels that affect changes in the accuracy rate or prediction probability, which are the inference results.
[0138] Embodiment 3 Fig. 23 is a block diagram showing an example configuration of an information processing device 1B according to embodiment 3. In Fig. 23, the information processing device 1B includes a communication unit 11, a calculation unit 12B, a storage unit 13, and a display unit 14. The communication unit 11 communicates with an external device via a network. The calculation unit 12B controls the overall operation of the information processing device 1B. It should be noted that the information processing device 1B does not necessarily have to include the communication unit 11. For example, when environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1B does not include the communication unit 11, the environmental information acquisition unit 121 realized by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.
[0139] The calculation unit 12B includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123, an inference unit 124A, and an output unit 125A. The calculation unit 12B executes an information processing application, thereby realizing the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, the inference unit 124A, and the output unit 125A. The storage unit 13 stores, for example, an information processing application, information used in the calculation processing of the calculation unit 12B, and a large-scale language model 13a (hereinafter referred to as LLM 13a). The storage unit 13 is a storage device provided in a computer that functions as the information processing device 1B, and is a storage such as an HDD or SSD. Note that the storage unit 13 may be provided outside the information processing device 1B as long as it is accessible by the information processing device 1B.
[0140] LLM13a is a language model for natural language processing that is trained using a large dataset. An example of an LLM13a is Open AI's GPT (Generative Pre-trained Transformer). By processing vast amounts of text data, LLM13a learns the associations and contexts of words and phrases, enabling it to generate context-appropriate sentences.
[0141] The inference unit 124A outputs the first inference information and the second inference information to the LLM 13a. The first inference information is an inference result output from a trained model to which a first feature amount map that has not been processed by the map processing unit 123 has been input. The second inference information is an inference result output from the trained model to which a second feature amount map that has been processed by the map processing unit 123 has been input. The LLM 13a outputs supplemental information based on the first inference information and the second inference information. Here, the supplemental information is information for explaining which information contributed to the inference result by the inference unit 124A. The output unit 125A outputs the first inference information, the second inference information, and the supplemental information to the display unit 14. As a result, the display unit 14 displays the first inference information, the second inference information, and the supplemental information.
[0142] Fig. 24 is a conceptual diagram showing an overview of processing by information processing device 1 B. In Fig. 24, environmental information indicating a plurality of scenarios as the environment of a fighting game is time-series information consisting of environmental information for each time, and includes, for example, position history information indicating the position of an object at each time or behavior history information indicating the behavior of an object at each time. FIG. 25 is a flowchart showing an information processing method according to the third embodiment. The environmental information acquisition unit 121 acquires, for example, environmental information as shown in Fig. 24 from the simulator (step ST1B). The map generation unit 122 generates a feature map including the channels shown in Fig. 24 based on the environmental information at each time (step ST2B).
[0143] The map processing unit 123 processes at least one channel out of the multiple channels included in the feature map (step ST3B). For example, the map processing unit 123 applies a change (e.g., a mask) to an object included in the channel. The inference unit 124A receives as input at least a portion of the feature map including channels to which no change has been applied and a feature map including channels to which changes have been applied by the map processing unit 123, and outputs first inference information and second inference information (step ST4B). For example, as shown in FIG. 24, the first inference information and the second inference information include the accuracy rate of hornets avoiding attacks from honeybees, the accuracy rate of hornets approaching honeybees and attacking them directly, and the accuracy rate of hornets approaching honeybees and attacking them. The inference unit 124A finds the degree of contribution that the information of the changed channel makes to the inference result based on the result of comparing the accuracy rate, which is the first inference information, with the accuracy rate, which is the second inference information.
[0144] The LLM 13a receives input from the inference unit 124A the feature map, information contained in the channels of the feature map, information obtained by processing (changing) this information by the map processing unit 123, the first inference information, and the second inference information (step ST5B). When this information is input, the LLM 13a outputs supplemental information explaining which information contributes to the inference result by the machine learning model (trained model) included in the inference unit 124A (step ST6B).
[0145] For example, suppose that the first inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is not masked is that the accuracy rate for "Avoid" is 97%, the accuracy rate for "Direct Attack" is 82%, and the accuracy rate for "Attack Base" is 92%. Furthermore, suppose that the second inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is masked is that the accuracy rate for "Avoid" is 42%, the accuracy rate for "Direct Attack" is 80%, and the accuracy rate for "Attack Base" is 92%. In this case, because masking the location information of "Hornet A" significantly changes (decreases) the accuracy rate for "Avoid" from 97% to 42%, the LLM13a outputs supplementary information such as, "Hornet A has a significant influence on the prediction of the hornet's strategy."
[0146] When environmental information is input, the LLM 13a may divide the input environmental information into various channels and output multiple feature map candidates, each representing a corresponding channel as text data. For example, the LLM 13a outputs feature map candidates including text data such as "location of the hornet," "location of the bee," "location of the flower," "beehive," "attack the bee," "attack the beehive," "move left," "move up," "move right," and "move down." The feature map candidates are output from the LLM 13a to the output unit 125A and then to the map generation unit 122. The output unit 125A outputs the feature map candidates to the display unit 14, which displays the feature map candidates. The map generation unit 122 generates a feature map corresponding to a feature map candidate selected using an input device (not shown) or automatically selected.
[0147] The information processing device 1B may also include a map processing control unit 126. As in the second embodiment, the map processing control unit 126 specifies, in the feature amount map, a channel to which a change is to be added, to the map processing unit 123. For example, the map processing control unit 126 may automatically determine candidates for the channel to be specified to the map processing unit 123 based on supplemental information output from the LLM 13a.
[0148] As described above, in the information processing device 1B according to the third embodiment, the inference unit 124A outputs the first inference information and the second inference information to the LLM 13a. The LLM 13a outputs supplemental information based on the first inference information and the second inference information. The output unit 125A outputs the first inference information, the second inference information, and the supplemental information to the display unit 14. By referring to the supplemental information, it is possible to improve the interpretability of the explanation of which input information contributes to the inference made by the inference unit 124A.
[0149] Various aspects of the present disclosure are summarized below as appendices.
[0150] (Appendix 1) an environment information acquisition unit that acquires environment information including position information of objects included in the environment or behavior information of the objects; a map generation unit that generates a first feature amount map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit that processes at least a first channel among the plurality of channels included in the first feature amount map; an inference unit that receives as input at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding inference of the environment based on at least a portion of the first feature map, and second inference information that is information regarding inference of the environment based on the second feature map. 1. An information processing device comprising: (Appendix 2) an output unit that outputs the output first inference information and the output second inference information to a display device; 2. The information processing device according to claim 1, (Appendix 3) and an output unit that outputs information based on the output first inference information and the output second inference information to a display device. 2. The information processing device according to claim 1, (Appendix 4) an output unit that outputs the information of the first channel to a display device; 2. The information processing device according to claim 1, (Appendix 5) The map processing unit performs the processing on the part of the object included in the first channel. 5. The information processing device according to claim 1, wherein: (Appendix 6) The map processing unit performs the processing on a portion of the object that is performing a dynamic action among the objects included in the first channel. 6. The information processing device according to claim 5, (Appendix 7) a map processing control unit that performs control to generate a third feature amount map including at least a second channel, among the plurality of channels, on which the processing has been performed by the map processing unit based on the first inference information and the second inference information; The inference unit receives the third feature map and outputs third inference information that is information regarding the inference of the environmental information based on the third feature map. 7. The information processing device according to claim 6, (Appendix 8) the plurality of channels includes a third channel; The map generating unit performs the processing on an area of the environmental information that does not correspond to the position information, and generates the third channel. 8. The information processing device according to claim 1, wherein: (Appendix 9) the plurality of channels includes a fourth channel; The map generating unit performs the processing on an area of the environmental information that does not correspond to the behavioral information, and generates the fourth channel. 9. The information processing device according to any one of Supplementary Note 1 to Supplementary Note 8, (Appendix 10) the environment is a simulator-driven environment; The position information of the object or the behavior information of the object is information generated by the simulator. 10. The information processing device according to any one of Supplementary Note 1 to Supplementary Note 9. (Appendix 11) The position information of the object or the behavior information of the object is acquired by object recognition from image information showing the environment. 11. The information processing device according to any one of Supplementary Note 1 to Supplementary Note 10. (Appendix 12) The inference unit receives the first feature map and the second feature map, and outputs the first inference information and the second inference information by regression. 12. The information processing device according to any one of claims 1 to 11. (Appendix 13) The inference unit receives the first feature map and the second feature map, and outputs the first inference information and the second inference information, which are probabilities of inferring that the environment belongs to a specific class. 13. The information processing device according to any one of Supplementary Note 1 to Supplementary Note 12. (Appendix 14) the environment includes a plurality of classes; The inference unit outputs first index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the first inference information and information on the correct class of the environment included in the environmental information, and outputs second index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the second inference information and information on the correct class of the environment included in the environmental information. 14. The information processing device according to any one of claims 1 to 13. (Appendix 15) The inference unit includes a trained model. 15. The information processing device according to any one of Supplementary Note 12 to Supplementary Note 14. (Appendix 16) the inference unit outputs the first inference information and the second inference information to a language model; the language model outputs supplemental information based on the first inference information and the second inference information; an output unit that outputs the first inference information, the second inference information, and the supplemental information to a display device; 16. The information processing device according to any one of Supplementary Note 12 to Supplementary Note 15. (Appendix 17) An information processing method executed by an information processing device, an environment information acquisition unit acquiring environment information including position information of an object included in the environment or behavior information of the object; a map generation unit generating a first feature map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit processing at least a first channel among the plurality of channels included in the first feature amount map; an inference unit receiving input of at least a part of the first feature amount map including the first channel and a second feature amount map including the first channel on which the processing has been performed by the map processing unit, and outputting first inference information that is information related to the inference of the environment based on at least a part of the first feature amount map and second inference information that is information related to the inference of the environment based on the second feature amount map; 1. An information processing method comprising: (Appendix 18) On the computer, an environment information acquisition unit acquiring environment information including position information of an object included in the environment or behavior information of the object; a map generation unit generating a first feature map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit processing at least a first channel among the plurality of channels included in the first feature amount map; a program for causing an inference unit to execute a step in which at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit are input, and the inference unit outputs first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information which is information regarding the inference of the environment based on the second feature map.
[0151] It is possible to combine the embodiments, modify any of the components of the embodiments, or omit any of the components of the embodiments. [Explanation of symbols]
[0152] 1A, 1B Information processing device, 11 Communication unit, 12, 12A, 12B Calculation unit, 13 Memory unit, 13a Large-scale language model, 14 Display unit, 100 Communication interface, 101 Input / output interface, 102 Processor, 103 Memory, 121 Environmental information acquisition unit, 122 Map generation unit, 123, 123A Map processing unit, 124, 124A Inference unit, 125, 125A Output unit, 126 Map processing control unit.
Claims
1. an environment information acquisition unit that acquires environment information including position information of objects included in the environment or behavior information of the objects; a map generation unit that generates a first feature amount map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit that processes at least a first channel among the plurality of channels included in the first feature amount map; an inference unit that receives as input at least a portion of the first feature amount map including the first channel and a second feature amount map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding inference of the environment based on at least a portion of the first feature amount map, and second inference information that is information regarding inference of the environment based on the second feature amount map.
1. An information processing device comprising:
2. an output unit that outputs the output first inference information and the output second inference information to a display device; 2. The information processing apparatus according to claim 1, wherein:
3. an output unit that outputs information based on the output first inference information and the output second inference information to a display device; 2. The information processing apparatus according to claim 1, wherein:
4. an output unit that outputs the information of the first channel to a display device; 2. The information processing apparatus according to claim 1, wherein:
5. The map processing unit performs the process on the part of the object included in the first channel.
5. The information processing apparatus according to claim 4,
6. The map processing unit performs the processing on a portion of the object that is performing a dynamic action among the objects included in the first channel.
6. The information processing apparatus according to claim 5,
7. a map processing control unit that performs control to generate a third feature amount map including at least a second channel, among the plurality of channels, on which the processing has been performed by the map processing unit based on the first inference information and the second inference information; The inference unit receives the third feature map and outputs third inference information that is information regarding the inference of the environmental information based on the third feature map.
7. The information processing apparatus according to claim 6,
8. the plurality of channels includes a third channel; The map generating unit performs the processing on an area of the environmental information that does not correspond to the position information, and generates the third channel.
2. The information processing apparatus according to claim 1, wherein:
9. the plurality of channels includes a fourth channel; The map generating unit performs the processing on an area of the environmental information that does not correspond to the behavioral information, and generates the fourth channel.
2. The information processing apparatus according to claim 1, wherein:
10. the environment is a simulator-driven environment; The position information of the object or the behavior information of the object is information generated by the simulator.
2. The information processing apparatus according to claim 1, wherein:
11. The position information of the object or the behavior information of the object is acquired by object recognition from image information showing the environment.
2. The information processing apparatus according to claim 1, wherein:
12. The inference unit receives the first feature map and the second feature map, and outputs the first inference information and the second inference information by regression.
2. The information processing apparatus according to claim 1, wherein:
13. The inference unit receives the first feature map and the second feature map, and outputs the first inference information and the second inference information, which are probabilities of inferring that the environment belongs to a specific class.
2. The information processing apparatus according to claim 1, wherein:
14. the environment includes a plurality of classes; The inference unit outputs first index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the first inference information and information on the correct class of the environment included in the environmental information, and outputs second index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the second inference information and information on the correct class of the environment included in the environmental information.
2. The information processing apparatus according to claim 1, wherein:
15. The inference unit includes a trained model.
15. The information processing device according to claim 12, wherein the information processing device is a computer.
16. the inference unit outputs the first inference information and the second inference information to a language model; the language model outputs supplemental information based on the first inference information and the second inference information; an output unit that outputs the first inference information, the second inference information, and the supplemental information to a display device; 2. The information processing apparatus according to claim 1, wherein:
17. An information processing method executed by an information processing device, an environment information acquisition unit acquiring environment information including position information of an object included in the environment or behavior information of the object; a map generating unit generating a first feature map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit processing at least a first channel among the plurality of channels included in the first feature amount map; an inference unit receiving input of at least a part of the first feature amount map including the first channel and a second feature amount map including the first channel processed by the map processing unit, and outputting first inference information that is information regarding inference of the environment based on at least a part of the first feature amount map and second inference information that is information regarding inference of the environment based on the second feature amount map; 1. An information processing method comprising:
18. On the computer, an environment information acquisition unit acquiring environment information including position information of an object included in the environment or behavior information of the object; a map generating unit generating a first feature map including a plurality of channels set based on the position information of the object or the behavior information of the object; a map processing unit processing at least a first channel among the plurality of channels included in the first feature amount map; A program for causing an inference unit to execute a step in which at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit are input, and the inference unit outputs first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information which is information regarding the inference of the environment based on the second feature map.