Information processing device, information processing method, and program

The information processing device generates feature maps from environmental information to identify the impact of object position and behavior on inference outcomes, addressing the lack of physical information acquisition in conventional methods.

WO2025177593A1PCT designated stage Publication Date: 2025-08-28MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/028165
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2024-08-07
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Conventional methods for determining inference results in machine learning models fail to acquire physical information in the environment that affects these results.

Method used

An information processing device that generates feature maps based on object position and behavior information, processes specific channels, and compares inference results before and after processing to identify the impact of environmental information on inference outcomes.

Benefits of technology

Enables the acquisition of physical information affecting inference results by analyzing changes in feature maps, thereby understanding the basis for determining environmental inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024028165_28082025_PF_FP_ABST
    Figure JP2024028165_28082025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device (1) comprises an environment information acquisition unit (121) that acquires environment information that includes position information for an object included in an environment or behavior information for the object, a map generation unit (122) that generates a first feature quantity map that includes a plurality of channels set on the basis of the position information for the object or the behavior information for the object, a map processing unit (123) that processes at least a first channel from among the plurality of channels included in the first feature quantity map, and an inference unit (124) that receives as input at least a portion of the first feature quantity map that includes the first channel and a second feature quantity map that includes the first channel as processed by the map processing unit (123) and outputs first inference information that is information about an inference for the environment based on the at least one portion of the first feature quantity map and second inference information that is information about an inference for the environment based on the second feature quantity map.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] In machine learning, it is difficult to grasp the basis for determining the inference results provided by a machine learning model. For example, Non-Patent Document 1 describes a technique for grasping the basis for determining the inference results provided by a machine learning model. The technique described in Non-Patent Document 1 randomly applies a mask to an image input to a machine learning model, and acquires important pixels in the image that affect the inference results based on changes in the inference results depending on whether or not the mask is present in the image.

[0003] V. Petsiuk, A. Das, K. Saenko, “Rise: randomized input sampling for explanation of black-box models,” British Machine Vision Conference (BMVC), 2018.

[0004] The conventional technology described in Non-Patent Document 1 acquires pixels that affect the inference results from the images input to the machine learning model, but has the problem of being unable to acquire physical information in the environment that affects the inference results of the environment.

[0005] The present disclosure is intended to solve the above-mentioned problems, and aims to provide an information processing device that is capable of acquiring physical information in an environment that affects the inference results of the environment.

[0006] The information processing device according to the present disclosure includes an environmental information acquisition unit that acquires environmental information including positional information of objects included in the environment or behavioral information of the objects; a map generation unit that generates a first feature map including a plurality of channels set based on the positional information of the objects or the behavioral information of the objects; a map processing unit that processes at least a first channel out of the plurality of channels included in the first feature map; and an inference unit that receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information that is information regarding the inference of the environment based on the second feature map.

[0007] According to the present disclosure, an information processing device according to the present disclosure generates a first feature map including multiple channels set based on object position information or object behavior information included in environmental information, processes at least a first channel of the multiple channels included in the first feature map, inputs at least a portion of the first feature map including the first channel and a second feature map including the processed first channel, and outputs first inference information that is information regarding environmental inference based on at least a portion of the first feature map and second inference information that is information regarding environmental inference based on the second feature map. Based on the first inference information and the second inference information regarding the environment to be inferred, an information processing device according to the present disclosure can obtain physical information in the environment that affects the environmental inference result.

[0008] 1 is a block diagram showing an example configuration of an information processing device according to embodiment 1. FIG. 1 is a conceptual diagram showing an overview of processing by the information processing device according to embodiment 1. FIG. 2 is a conceptual diagram showing an overview of a feature map converted from environmental information. FIG. 3 is a block diagram showing a hardware configuration for realizing the functions of the information processing device according to embodiment 1. FIG. 5A, FIG. 5B, and FIG. 5C are diagrams showing basketball scenarios. FIG. 6 is a schematic diagram showing an overview of basketball scenario identification by the information processing device according to embodiment 1. FIG. 7 is a schematic diagram showing an overview of a feature map converted from a basketball scenario. FIG. 8A and FIG. 8B are conceptual diagrams showing an overview of processing for applying a mask to a channel for a three-point shot included in a feature map. FIG. 9 is a flowchart showing an information processing method according to embodiment 1. FIG. 10 is a conceptual diagram showing an overview of processing using one scenario. FIG. 11A and FIG. 11B are conceptual diagrams showing an overview of processing for applying a mask to position information of an opponent who takes a shot from a dribble, in a channel included in a feature map. FIG. 12A and FIG. 12B are conceptual diagrams showing an overview of processing for applying a mask to position information of an opponent who moves little, in a channel included in a feature map. 13A, 13B, and 13C are diagrams showing feature maps in which changes have been made to the information of various channels. FIGS. 14A and 14B are conceptual diagrams showing an overview of a fighting game between bees and hornets. This is a schematic diagram showing an overview of a process for identifying multiple scenarios of a fighting game between bees and hornets using a machine learning model. This is a schematic diagram showing an overview of a feature map converted from a scenario of a fighting game between bees and hornets. FIGS. 17A and 17B are conceptual diagrams showing an overview of a process for masking the position information of a specific hornet in a channel included in the feature map. FIGS. 18A and 18B are conceptual diagrams showing an overview of a process for masking the position information of all hornets in a channel included in the feature map. FIGS. 19A and 19B are conceptual diagrams showing an overview of a process for masking attacks on bees in a channel included in the feature map. This is a schematic diagram showing an overview of a feature map converted from environmental information indicating the environment of a real space. FIG. 10 is a block diagram showing a configuration example of an information processing device according to a second embodiment.Fig. 1 is a flowchart showing an information processing method according to embodiment 2. Fig. 2 is a block diagram showing an example of the configuration of an information processing device according to embodiment 3. Fig. 3 is a conceptual diagram showing an overview of processing by the information processing device according to embodiment 3. Fig. 4 is a flowchart showing an information processing method according to embodiment 3.

[0009] Embodiment 1. (Overview of Information Processing Device) Fig. 1 is a block diagram showing an example configuration of an information processing device 1 according to embodiment 1. In Fig. 1, the information processing device 1 acquires environmental information including position information or behavioral information of objects included in the environment, generates a first feature map including multiple channels set based on the position information or behavioral information of the objects, processes a first channel among the multiple channels, and outputs first inference information regarding environmental inference based on the first feature map including the first channel and second inference information based on a second feature map including the processed first channel. By using the first inference information and the second inference information, the information processing device 1 can shorten the number of attempts required to understand the basis for determining an inference result regarding the environment.

[0010] The environment is the environment to be inferred, and may be a real-world environment or a virtual environment operated by a simulator. The object is a physical object contained in the environment. For example, objects include not only machine learning agents but also traffic lights and warehouses contained in the environment. The agent is the subject of decision-making and action, and the environment is the object that sends and receives information to and from the agent. The agent is realized as a program that constitutes a machine learning application, learns from its own experiences, and determines future actions based on the results of that learning.

[0011] The environmental information includes physical information such as the position or behavior of an object, which is an example of an object. The object's position information is, for example, two-dimensional coordinate information of the object on a captured image obtained by capturing an image of a real space, or two-dimensional coordinate information of an agent on an image showing a simulation environment. The object's behavior information may be, for example, information indicating a time-series change in the object's position. Note that the object's position information or behavior information may be two-dimensional information or three-dimensional information.

[0012] The environmental information may be time-series data or data at a specific time. The first feature map includes a plurality of channels generated based on the environmental information. Each channel is set with location information indicating the location of an object or behavior information indicating the behavior of the object.

[0013] Of the multiple channels included in the first feature map, the channel that has been processed by the map processing unit 123 is the first channel. The feature map including the processed first channel is the second feature map. "Processing" includes not only applying a mask to an object included in the environment, but also making changes to the position or behavior of the object. The information processing device 1 processes the environmental information included as a channel in the first feature map, inputs the feature map before processing and the feature map after processing, performs inference regarding the environment, and compares the output first inference information with the output second inference information. The result of this comparison makes it possible to know how much the information set in the processed first channel, among the channels included in the first feature map, affected the inference result regarding the environment.

[0014] The information processing device 1 will be described in detail below. As shown in FIG. 1 , the information processing device 1 includes a communication unit 11, a calculation unit 12, a storage unit 13, and a display unit 14. The communication unit 11 communicates with an external device via a network. For example, the communication unit 11 may be a communication device capable of mobile communication using a communication method such as LTE, 3G, 4G, or 5G. The communication unit 11 may also include a short-range wireless communication means such as Bluetooth (registered trademark). Note that the information processing device 1 does not necessarily include the communication unit 11. For example, if environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1 does not include the communication unit 11, the environmental information acquisition unit 121 implemented by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.

[0015] The calculation unit 12 controls the overall operation of the information processing device 1. The calculation unit 12 includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123, an inference unit 124, and an output unit 125. The calculation unit 12 executes an information processing application, thereby realizing the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, the inference unit 124, and the output unit 125.

[0016] The storage unit 13 stores, for example, an information processing application and information used in the arithmetic processing of the arithmetic unit 12. The storage unit 13 is a storage device provided in a computer that functions as the information processing device 1, and is a storage such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). Note that the storage unit 13 may be provided outside the information processing device 1 as long as it is accessible by the information processing device 1.

[0017] The display unit 14 is a display device included in the information processing device 1. The display unit 14 is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display device. The display unit 14 may be provided outside the information processing device 1 as long as the information processing device 1 can display information. For example, the display unit 14 may be a display device included in a terminal device that can communicate with the information processing device 1 via a network.

[0018] (Environmental Information Acquisition Unit) The environmental information acquisition unit 121 acquires environmental information including position information or behavioral information of objects included in the environment. For example, when the inference target is an environment modeled after the Gym format in Open AI, which is a learning environment often used in reinforcement learning, the environmental information acquisition unit 121 acquires time-series environmental information from the simulator by utilizing the environment observation function of the simulator. For example, specifically, the environmental information acquisition unit 121 communicates with an external device functioning as a simulator via a communication network using the communication unit 11, or connects to a simulator included in the information processing device 1 via a signal line, and acquires time-series environmental information for a target period from the simulator.

[0019] The time-series environmental information is information that indicates the positions or actions of objects included in the simulation environment. Fig. 2 is a conceptual diagram showing an overview of processing by the information processing device 1. For example, the environmental information that indicates the simulation environment of a fighting game is time-series information made up of environmental information for each time, as shown in Fig. 2, and includes, for example, position information that indicates the position of an agent at each time.

[0020] When a real-space environment (hereinafter referred to as the real environment) is the inference target, the environmental information acquisition unit 121 acquires time-series environmental information from, for example, a sensor arranged in the real space. The sensor may be a visible light camera, a GPS (Global Poisoning System), a ToF (Time of Flight) sensor, or an infrared camera. Specifically, for example, the environmental information acquisition unit 121 connects to the sensor via a communication network or a signal line via the communication unit 11, and acquires time-series environmental information indicating the real environment from the sensor. The environmental information indicating the real environment is, for example, an image of the real environment captured by a camera.

[0021] (Map Generation Unit) The map generation unit 122 generates a feature map based on the environmental information acquired by the environmental information acquisition unit 121. The feature map generated by the map generation unit 122 is a first feature map including multiple channels set based on the position information or behavior information of an object. For example, the map generation unit 122 generates a channel for each time based on the environmental information for each time, thereby generating a feature map consisting of time-series channels as shown in FIG. 2. Numerical values ​​indicating the position information or behavior information of an object are set in the channels, and the environment in which the object exists can be represented in a binary image format or a multi-value image format.

[0022] The map generation unit 122 assigns a numerical value indicating a color gradation, such as a range from 1 to 255, or a numerical value corresponding to the role of each object, to each object identified from the environmental information. Binary or multi-value image map information is information representing an image having pixel values ​​corresponding to the numerical values ​​assigned to objects and other parts. For example, when representing an environment in which an object exists as a binary image, the map generation unit 122 assigns the value "1" to the object and the value "0" to parts where no object exists, thereby representing the object as a white area in the binary image and parts other than the object as a black area. Alternatively, the map generation unit 122 may assign the value "1" to objects performing a target behavior and represent them as white areas in the binary image, and assign the value "0" to parts where no object exists and objects not performing the target behavior and represent them as black areas.

[0023] FIG. 3 is a conceptual diagram showing an overview of a feature map converted from environmental information, showing a feature map converted by the map generation unit 122 from environmental information indicating the simulation environment of a fighting game. In FIG. 3, it is assumed that objects represented by circles and squares exist in the simulation environment indicated by the environmental information at a certain time, with allies being white objects and opponents being gray objects. Based on the environmental information, the map generation unit 122 determines, for example, that each type of object exists in the simulation environment, assigns a value of "1" to the object, and assigns a value of "0" to areas where no object exists. The numerical values ​​assigned to the objects, etc., are parameter values ​​corresponding to the colors on the image.

[0024] For example, the feature map shown in FIG. 3 includes channels ch1, ch2, and ch3. These channels include binary images representing a simulation environment. As shown in FIG. 3, channel ch1 indicates the position of an opponent's object. In the binary image included in channel ch1, an image region corresponding to the position of the opponent's object is represented in white, and an image region in which this object is not present is represented in black. Channel ch2 indicates the position of an ally's object. In the binary image included in channel ch2, an image region corresponding to the position of the ally's object is represented in white, and an image region in which this object is not present is represented in black. Channel ch3 indicates an opponent's object that is performing an "attack" action, which is the target action for determining the degree of contribution to environmental inference. In the binary image included in channel ch3, an image region corresponding to the position of the opponent's object performing the "attack" action is represented in white, and the other image regions are represented in black.

[0025] Alternatively, the map generating unit 122 may assign the value "0" to an object out of the two values ​​"1" and "0" and assign the value "1" to other objects. As a result, in the binary image, image areas where objects exist are displayed in black, and image areas where no objects exist are displayed in white.

[0026] Furthermore, when an object is performing multiple actions, the map generation unit 122 may assign a digital value to the object by calculating the logical product (AND) or logical sum (OR) of the numerical values ​​assigned to each of the multiple actions. For example, the value "1" is assigned to the action of "advancing" and the value "0" is assigned to the action of "attacking." When identifying an object that is "attacking" while "advancing" as an object that is "advancing," the map generation unit 122 assigns the value "1," which is the logical sum of the value "1" assigned to "advance" and the value "0" assigned to "attack." When identifying an object that is "attacking" while "advancing" as an object that is "attacking," the map generation unit 122 assigns the value "0," which is the logical product of the value "1" assigned to "advance" and the value "0" assigned to "attack."

[0027] The feature map may also include multiple channels, including a third channel. The third channel is a channel generated by the map generation unit 122 by processing areas of the environment indicated by the environmental information that do not correspond to the object's position information. For example, the map generation unit 122 assigns a common numerical value different from the numerical value assigned to the object to areas that do not correspond to the object's position information in an image that indicates the environment in a multi-valued image format, and sets the parts other than the object to a uniform color, etc. This makes it possible to generate a channel that includes an image that clearly indicates the presence of the object as the third channel. Note that the values ​​assigned to the areas that do not correspond to the object's position information do not necessarily have to be the same, as long as they are similar values. As a non-limiting example, a value between 1 and 10 may be assigned.

[0028] Furthermore, the feature map may include multiple channels, including a fourth channel. The fourth channel is a channel generated by the map generation unit 122 by processing regions of the environment indicated by the environmental information that do not correspond to the behavioral information of objects. For example, the map generation unit 122 assigns a common numerical value different from the numerical value assigned to the object to regions that do not correspond to the behavioral information of objects in an image that indicates the environment in a multi-valued image format, and sets the parts other than the object to the same color, etc. This makes it possible to generate a channel as the fourth channel that includes an image that clearly shows the presence of an object performing a specific behavior. Note that the values ​​assigned to the regions that do not correspond to the behavioral information of objects do not necessarily have to be the same, as long as they are similar values. As a non-limiting example, a value between 1 and 10 may be assigned.

[0029] Furthermore, the numerical value assigned to an object may be a value that corresponds to the meaning of the object's action. For example, if the object represents a ball, a value corresponding to the strength or angle at which the ball was thrown may be assigned to the object. In this way, the feature map generated by the map generation unit 122 includes feature values ​​having information for each channel.

[0030] (Map Processing Unit) The map processing unit 123 processes a portion of the feature map. "Processing" also includes making changes to the information set in a channel in the feature map. That is, since each channel included in the feature map contains information such as the position or behavior of an object, it is possible to make changes to the information included in a specific channel. The map processing unit 123 processes at least the first channel out of the multiple channels included in the feature map. The first channel is a channel in which information for which the degree of contribution to environmental inference is to be determined is set, and is one channel included in one feature map. Note that the channels to be processed may be multiple channels including the first channel.

[0031] The change may be made by masking the object. Masking is a process of eliminating the presence of an object, for example, by changing a numerical value assigned to the object to the same value as a numerical value assigned to everything other than the object in the environment. For example, when masking an object in a binary image in which the value "1" is assigned to the object and the value "0" is assigned to everything else, thereby rendering the image area where the object exists white and the image area where the object does not exist black, the map processing unit 123 assigns the value "0" to the object. As a result, as shown in FIG. 2 , the image area where the object exists in the binary image also becomes black, thereby masking the object. Alternatively, for example, in a binary image in which the value "0" is assigned to the object and the value "1" is assigned to everything else, thereby rendering the image area where the object exists black and the image area where the object does not exist white, the map processing unit 123 may assign the value "1" to the object to mask the object. The map processing unit 123 may assign 0.5, which is the intermediate value in the numerical range from 0 to 1, or another value to the object or other parts in the environment to be inferred. It should be noted that masking is a process for eliminating the presence of an object, and while the masking has been described as, for example, making the numerical value assigned to an object the same as the numerical value assigned to everything else in the environment other than that object, it is not limited to this. As a non-limiting example, in the case of a multi-value image, the numerical value assigned to an object may be a value close to the numerical value assigned to everything else in the environment other than that object. For example, a value within a range of plus or minus 5 of the value assigned to everything else in the environment other than that object may be set.

[0032] (Inference Unit) The inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information and second inference information. The first inference information is information related to the inference of the environment based on at least a portion (one channel) of the first feature map. The second inference information is information related to the inference of the environment based on the second feature map. For example, when the inference target is a simulated environment of a competitive game, the inference information may include, as shown in FIG. 2 , the probability that an ally will avoid an attack from an opponent, the probability that the ally will directly attack the opponent, and the probability that the ally will attack the opponent's base. Because the feature map includes observation data of the real environment indicated by the environmental information or data indicating the simulated environment, the first inference information and the second inference information inferred by the inference unit 124 based on this feature map are physical information that contribute to the inference result of the environment. Therefore, by comparing the first inference information with the second inference information, it is possible to grasp the extent to which the physical information that contributes to the inference result of the environment contributes to the inference result.

[0033] The inference unit 124 may include a trained model. When information contained in the feature map is input, the trained model outputs an inference result of the environment indicated by the environmental information. This trained model is generated by supervised learning using the information contained in the feature map as training data, and is a machine learning model from which interpretability is to be acquired. "Interpretability" indicates the degree of influence that the information input to the machine learning model has on the inference result of the machine learning model. The trained model is stored in the memory unit 13 shown in FIG. 1 or in an external storage device accessible by the inference unit 124 via the communication unit 11 via a communication network. The inference unit 124 acquires the trained model from the memory unit 13 or the external storage device, and performs inference on the environment based on the acquired trained model.

[0034] As shown in FIG. 2 , the inference unit 124 compares first inference information, which is the inference result output from a trained model to which a first feature map not processed by the map processing unit 123 has been input, with second inference information, which is the inference result output from the trained model to which a second feature map processed by the map processing unit 123 has been input. Based on the result of comparing the first inference information with the second inference information, the inference unit 124 determines the magnitude of influence (hereinafter referred to as the contribution) that the information of the processed channel has on the inference result of the trained model. The inference unit 124 can interpret how elements included in the environmental information affect the inference result. Note that in FIG. 2 , the trained model is referred to as a "model." This will also be referred to as a "model" in the following figures.

[0035] The inference unit 124 performs inference using, of the feature amount maps generated by the map generation unit 122, a feature amount map including a first channel that has not been processed by the map processing unit 123, and a feature amount map including a first channel that has been processed by the map processing unit 123. Alternatively, the inference unit 124 may perform inference using all feature amount maps generated by the map generation unit 122 and all feature amount maps including feature amount maps whose channels have been processed by the map processing unit 123.

[0036] The inference information may be a prediction probability output from the trained model or an index measuring prediction performance. For example, when the environment includes multiple classes, the inference unit 124 may output first index information, which is an index for measuring the inference performance of each of the multiple classes, based on the first inference information and information about the correct class of the environment included in the environment information, and output second index information, which is an index for measuring the inference performance of each of the multiple classes, based on the second inference information and information about the correct class of the environment included in the environment information. The first index information and the second index information are indexes for measuring prediction performance, and may be, for example, the accuracy rate of each class, the accuracy rate of all classes, the recall rate, the precision rate, or the F1 score. Furthermore, the inference unit 124 may acquire the accuracy rates output from the trained model when a feature map with a change in the channel and a feature map with an unchanged channel are input, respectively, based on the inference results output from the trained model and the correct label of the dataset input to the trained model, and calculate the difference between these accuracy rates as the contribution.

[0037] (Output Unit) The output unit 125 outputs information to the display unit 14. For example, the output unit 125 acquires first inference information and second inference information from the inference unit 124 and outputs the acquired information to the display unit 14. The display unit 14 displays the information output from the output unit 125. The information displayed on the display unit 14 may display the output of the inference unit 124 depending on whether or not there is a change in the feature map, or may display the difference between these outputs. Note that the information processing device 1 is required to have at least the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, and the inference unit 124, and the output unit 125 is an optional function. For example, the function of the output unit 125 may be a function provided by the inference unit 124.

[0038] (Program) Fig. 4 is a block diagram showing a hardware configuration that realizes the functions of the information processing device 1. For example, the information processing device 1 has, as its hardware configuration, a communication interface 100, an input / output interface 101, a processor 102, and a memory 103. The communication unit 11 shown in Fig. 1 is a communication device that communicates with an external device of the information processing device 1 via the communication interface 100. The calculation unit 12 shown in Fig. 1 outputs display information to the display unit 14 shown in Fig. 1 via the input / output interface 101. The calculation unit 12 is the processor 102. The memory 103 is the storage unit 13 shown in Fig. 1.

[0039] The programs constituting the information processing application are stored in the memory 103. The processor 102 reads and executes the programs stored in the memory 103, thereby realizing the functions of the environmental information acquisition unit 121, map generation unit 122, map processing unit 123, inference unit 124, and output unit 125 provided in the calculation unit 12. The memory 103 is, for example, an HDD or SSD. The memory 103 may also be provided in an external device capable of exchanging data via the communication interface 100.

[0040] For example, the environmental information acquisition unit 121 acquires environmental information received by the communication unit 11 via the communication interface 100. The output unit 125 outputs the inference information as display information to the display unit 14 via the input / output interface 101. The output unit 125 may also transmit the inference information to an external device via the communication interface 100 via the communication unit 11.

[0041] Next, a specific example of the operation of the information processing device 1 will be described. Figures 5A, 5B, and 5C are diagrams showing basketball scenarios, illustrating scenarios of teammate and opponent movements on a basketball court. For example, scenarios that the opponent can take in a basketball simulation include "emphasis on one-on-one" as shown in Figure 5A, "emphasis on basket" as shown in Figure 5B, and "emphasis on outside attacks" as shown in Figure 5C. The inference unit 124 distinguishes between these scenarios.

[0042] Figures 5A, 5B, and 5C show the movement of players and a ball on a basketball court. In Figures 5A, 5B, and 5C, a thick circle represents the ball. Furthermore, white circles represent teammate players and are assigned consecutive numbers 1 through 5. Black circles represent opposing teammates and are similarly assigned consecutive numbers 1 through 5. The "one-on-one focused" scenario shown in Figure 5A illustrates a strategy in which an opposing teammate marks a teammate player one-on-one, and opposing teammate player number 3 dribbles past teammate player number 3 and takes a two-point shot from under the basket. The "under the basket focused" scenario shown in Figure 5B illustrates a strategy in which an opposing teammate marks a teammate player one-on-one, as in Figure 5A, but opposing teammate player number 4 passes the ball to opposing teammate player number 5, who is under the basket, and opposing teammate player number 5 receives the pass and takes a two-point shot from under the basket. The "outside attack emphasis" scenario shown in Figure 5C shows a strategy in which an opposing player marks a teammate one-on-one, and the opposing player number 4 takes a three-point shot from outside the three-point line.

[0043] FIG. 6 is a schematic diagram illustrating an overview of basketball scenario identification by the information processing device 1. The basketball simulation environment uses a learning environment commonly used in reinforcement learning. For example, the basketball simulation environment is realized following the Open AI Gym format. In the reinforcement learning environment, the environmental information acquisition unit 121 acquires, using an observation function, time-series environmental information during the simulation, such as information indicating the positions or actions of opposing and friendly players. Furthermore, a scenario is information obtained by classifying a data set provided for executing a basketball simulation, and is environmental information indicating one simulated scene. Classes include, for example, "emphasis on one-on-one play," "emphasis on basket play," and "emphasis on outside attacks." The scenario may also include time-series information.

[0044] The "actions" in a basketball simulation include at least an action of a player making a two-point shot, an action of a player making a three-point shot, an action of a player dribbling, an action of a player passing, etc. In other words, the actions performed by players on both sides present in the environment representing the basketball court are the targets of the simulation.

[0045] In FIG. 6 , the environmental information acquired by the environmental information acquisition unit 121 is a plurality of time-series scenarios classified into a plurality of classes. Based on this environmental information, the map generation unit 122 generates a time-series feature map for each scenario, as shown in FIG. 6 . The map processing unit 123 modifies a portion of the feature map generated by the map generation unit 122. Also in FIG. 6 , the softmax function is a function used to predict multiple classes. It normalizes input values ​​and converts them into a range from 0 to 1, and the sum of all the values ​​equals 1. Using the softmax function, the probability of belonging to each class can be calculated, and the class with the highest probability is the predicted class. In argmax, the class with the highest probability is selected from the output value of the softmax function, and this class can be assigned as an identification label.

[0046] When an unaltered feature map and a varied feature map are input, the trained model included in the inference unit 124 outputs, for example, a softmax function, a predicted probability as first inference information based on the unaltered feature map, and a predicted probability as second inference information based on the varied feature map. The predicted probability is the probability that an environment will be inferred to be a specific class, for example, the probability that a scenario will be inferred to be a specific class.

[0047] The inference unit 124 uses the argmax model to assign an identification label indicating the class of the scenario based on the predicted probability, and calculates the accuracy rate based on the identification label and the correct class of that scenario.The inference unit 124 then compares the accuracy rate based on the feature map without any changes with the accuracy rate based on the feature map with changes, and based on this comparison, it is possible to grasp the degree of contribution of the changed information to the inference result.

[0048] The environmental information may be set by a user using an input device (not shown). In this case, the environmental information may include information other than the positions and actions of the players and the ball. Examples of the information include the remaining game time, the size of the court, the distance between the players and the goal, and the remaining physical strength of each player.

[0049] The dataset of environmental information may also be provided with information indicating the enemy's strategy (correct class) in the environment (scenario) indicated by this environmental information. By comparing the output (discrimination class) of a machine learning model to which a feature map converted from certain environmental information is input with the correct class provided for that environmental information, it is possible to calculate the accuracy rate for each class.

[0050] FIG. 7 is a schematic diagram illustrating an overview of a feature map converted from a basketball scenario. For example, the map generation unit 122 converts all or part of the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in FIG. 7. If the environmental information includes teammate and opponent players and the ball position, a pixel value ranging from 1 to 255 representing 256 gradations of color is assigned to the environmental information. If the teammate and opponent players are not present, a pixel value of 0 (black) representing 256 gradations is assigned to the environmental information. Furthermore, teammate and opponent players may be assigned numbers according to their roles. For example, the point guard may be assigned the number 1, the shooting guard may be assigned the number 2, the small forward may be assigned the number 3, the power forward may be assigned the number 4, and the center may be assigned the number 5.

[0051] 7 is information consisting of three channels (ch1, ch2, ch3) in which object position information is set, and four channels (ch4, ch5, ch6, ch7) in which object behavior information is set. The map generation unit 122 sets "1" to the position information set in channels ch1, ch2, and ch3 if an object exists, and sets "0" if an object does not exist.

[0052] The map generation unit 122 sets action information indicating an action of a two-point shot by an opposing player for channel ch4, and sets action information indicating an action of a three-point shot by an opposing player for channel ch5. The map generation unit 122 sets action information indicating an action of a dribble by an opposing player for channel ch6, and sets action information indicating an action of a pass by an opposing player for channel ch7. Furthermore, if an action indicated by the action information set in channels ch4, ch5, ch6, and ch7 is being performed, a numerical value such as the above-mentioned 1 to 255, or a numerical value corresponding to the strength of the action, is set, and if the action is not being performed, a value of 0 is set.

[0053] The map processing unit 123 performs processing such as changing at least a specific feature amount map among the feature amount maps generated by the map generating unit 122. In the feature amount map, position information or behavior information of an object is set for each channel. The map processing unit 123 can change the specific information set for the channel.

[0054] FIG. 8A shows a feature map converted from environmental information representing a basketball simulation environment, and FIG. 8B is a conceptual diagram illustrating an outline of a process for masking channel ch5 for "three-point shot" included in the feature map shown in FIG. 8A. Here, the feature map shown in FIG. 8A is a time-series feature map converted from multiple scenarios shown in FIG. 8B. For example, it is expected that the identification of the "emphasis on outside attacks" scenario shown in FIG. 5C will be influenced by players who perform "three-point shots" in this scenario. Therefore, the map processing unit 123 masks the behavior information for "three-point shot," which is set to channel ch5 among the channels included in the feature map.

[0055] For example, a value of "1" is assigned to a player taking a three-point shot, which is set in channel ch5, and a value of "0" is assigned to other parts of the basketball court. As a result, as shown in FIG. 8B , in the binary image of the basketball court included in channel ch5, the image area where the player taking a three-point shot is located is displayed in white, and the other areas are displayed in black. The map processing unit 123 masks this player by assigning the same value of "0" to this player as the other areas, as shown in FIG. 8B .

[0056] The inference unit 124 outputs first inference information based on a feature map that has not been changed by the map processing unit 123, and outputs second inference information based on a feature map that has been changed by the map processing unit 123. For example, if the inference unit 124 includes a trained model that identifies basketball strategies (scenarios), when a feature map including channel ch5 in which a player taking a three-point shot is not masked is input, the trained model outputs first inference information, which is the accuracy rate of each strategy, as shown in FIG. 8B . Furthermore, when a feature map including channel ch5 in which a player taking a three-point shot is masked by the map processing unit 123, the trained model outputs second inference information, which is the accuracy rate of each strategy, as shown in FIG. 8B . The strategies include "emphasis on one-on-one play," "emphasis on basket play," and "emphasis on outside attacks," and the inference information is the accuracy rate of each strategy.

[0057] For example, suppose that the first inference information based on a feature map including channel ch5 in which the player taking the three-point shot is not masked has an accuracy rate of 0.92 for "one-on-one focus," 0.94 for "focus on basket," and 0.85 for "focus on outside attack." Also, suppose that the second inference information based on a feature map including channel ch5 in which the player taking the three-point shot is masked has an accuracy rate of 0.88 for "one-on-one focus," 0.94 for "focus on basket," and 0.32 for "focus on outside attack."

[0058] In this case, by masking players taking three-point shots, the accuracy rate for "emphasis on outside attack" significantly changes (decreases) from 0.85 to 0.32, which indicates that players taking three-point shots have a significant impact on the prediction of "emphasis on outside attack." In this way, the information processing device 1 can acquire important channel information for each class. The trained model is generated by supervised learning, where a training dataset is converted into a feature map in advance and the feature map is used as input. A 3D CNN, for example, may be used to generate the trained model.

[0059] Furthermore, the map processing unit 123 masks various channel information by, for example, changing the channel information to be masked, and by grasping the degree of contribution to each class, it is possible to acquire channel information that contributes greatly to the inference result, and also to acquire channel information that contributes little to the inference result. The position information or behavior information of an object included in a channel that contributes highly to the inference result is physical information about the environment in which this object exists. Based on this information, it is possible to grasp the class of the inference target and its contribution to the inference result.

[0060] (Output Information) The output unit 125 outputs the first inference information and the second inference information output from the inference unit 124 to the display unit 14. For example, the output unit 125 outputs both the accuracy rate of the strategy identification based on the feature map without any changes to channel ch5 and the accuracy rate of the strategy identification based on the feature map with any changes to channel ch5 to the display unit 14. The display unit 14 displays each accuracy rate as shown in FIG. 8B . By visually checking the accuracy rates displayed on the display unit 14, the impact of the information with the changes to channel ch5 on the strategy identification can be recognized. Furthermore, the output unit 125 may output to the display unit 14 either the accuracy rate of the strategy identification based on the feature map without any changes to channel ch5 or the accuracy rate of the strategy identification based on the feature map with any changes to channel ch5. The display unit 14 displays the information output from the output unit 125. Furthermore, the output unit 125 may output to the display unit 14 information that combines the accuracy rate of tactic identification based on the feature amount map with no change made to channel ch5 and the accuracy rate of tactic identification based on the feature amount map with a change made to channel ch5. The display unit 14 displays the information output from the output unit 125.

[0061] The output unit 125 may output information based on the first inference information and the second inference information output from the inference unit 124 to the display unit 14. For example, the output unit 125 calculates the difference between the accuracy rate of tactic identification based on a feature map with no changes made to channel ch5 and the accuracy rate of tactic identification based on a feature map with changes made to channel ch5, and outputs the calculated difference information to the display unit 14. The display unit 14 displays the difference information. By visually checking the difference information displayed on the display unit 14, it is possible to grasp the extent of the influence that the information set in channel ch5 has on the inference. Note that the information based on the first inference information and the second inference information is not limited to difference information, and a value calculated by adding, multiplying, averaging, or the like of both the information and the second inference information may be used as the index value of the degree of contribution.

[0062] The output unit 125 may output information of the first channel to the display unit 14. The first channel is a channel that has been subjected to processing such as changes by the map processing unit 123. For example, the output unit 125 outputs information of channel ch5 before the changes are made to the display unit 14. The display unit 14 displays the information of channel ch5. By visually checking the information of channel ch5 displayed on the display unit 14, it can be recognized that the information that influences the "emphasis on outside attacks" prediction is channel ch5.

[0063] (Information Processing Method) FIG. 9 is a flowchart showing an information processing method according to the first embodiment. The environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment (step ST1). The map generation unit 122 generates a first feature map including multiple channels set based on the position information or behavior information of the objects (step ST2). The map processing unit 123 processes at least the first channel among the multiple channels included in the first feature map (step ST3). The inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information, which is information related to the inference of the environment based on at least a portion of the first feature map, and second inference information, which is information related to the inference of the environment based on the second feature map (step ST4). By executing the information processing method shown in FIG. 20, the information processing device 1 can acquire physical information in the environment that affects the inference results of the environment.

[0064] (Variation 1) In the explanation so far, the degree of contribution of channel information in a feature map to an inference result has been evaluated using a dataset of environmental information indicating multiple scenarios. Next, a case will be explained in which the degree of contribution of channel information in a feature map to an inference result is evaluated using environmental information indicating a single scenario. Specifically, the prediction probability of which class a single scenario belongs to is calculated. The degree of contribution of this channel information to the inference result is evaluated based on the difference between the prediction probability based on a feature map with no changes made to the channels and the prediction probability based on a feature map with changes made to the channels.

[0065] FIG. 10 is a conceptual diagram showing an overview of processing using one scenario, illustrating a case where a basketball simulation environment is the inference target. In FIG. 10 , the environment information acquisition unit 121 acquires a data set of one scenario as environment information. For example, the scenario is "emphasis on one-to-one." Based on this environment information, the map generation unit 122 generates a time-series feature map for the "emphasis on one-to-one" scenario, as shown in FIG. 10 . This feature map is a feature map including seven channels, similar to FIG. 8A . The map processing unit 123 makes changes to a portion of the feature map generated by the map generation unit 122.

[0066] When an unaltered feature map and a varied feature map are input, the trained model included in the inference unit 124 uses, for example, a softmax function to output a predicted probability as first inference information based on the unaltered feature map, and a predicted probability as second inference information based on the varied feature map. Based on the result of comparing the predicted probability based on the unaltered feature map with the predicted probability based on the varied feature map, the inference unit 124 can grasp the degree of contribution of the varied information to the inference result.

[0067] FIG. 11A shows a feature map converted from environmental information indicating a basketball simulation environment. FIG. 11B is a conceptual diagram showing an outline of a process for masking channel ch2, which is set as "enemy position" and is included in the feature map shown in FIG. 11A . The feature map shown in FIG. 11A is a time-series feature map converted from time-series environmental information of one scenario (emphasis on one-on-one play) shown in FIG. 11B . Channel ch2 included in each feature map at each time is a first channel, in which position information indicating the positions of five enemy players (1) to (5), numbered 1 to 5 as "enemy positions," is set.

[0068] The map processing unit 123 processes the portion of the object included in the first channel. This allows the influence of the portion of the object included in the first channel on the inference results to be evaluated. For example, it is expected that the position information of the opponent (3) player, who shoots after dribbling in the scenario, will have an effect on the prediction of the "one-on-one emphasis" scenario. Therefore, in the example of FIG. 11B , the map processing unit 123 applies a mask to the position of the opponent (3) player among the five opponent players (1) to (5). The position information of the opponent players (1) to (5) included in channel ch2 is assigned a value of "1," and the value of "0" is assigned to other parts of the basketball court. As a result, as shown in FIG. 11B , in the binary image of the basketball court included in channel ch2, the image areas where the opponent players (1) to (5) are present are displayed in white, and the other areas are displayed in black. When masking the position information of the enemy (3) player, the map processing unit 123 masks this player by assigning the same value "0" to the position of the enemy (3) player who is shooting from a dribble among the enemy (1) to (5) players, as is the case with the other parts, as shown in FIG. 11B.

[0069] The inference unit 124 outputs first inference information based on a feature map that has not been changed by the map processing unit 123, and outputs second inference information based on a feature map that has been changed by the map processing unit 123. For example, if the inference unit 124 is a trained model that predicts basketball strategies (scenarios), when a feature map that includes at least channel ch2 in which the position information of the opponent (3) player is not masked is input, the trained model outputs first inference information that is the predicted probability of each scenario shown in FIG. 11B. Furthermore, when a feature map that includes channel ch2 in which the position information of the opponent (3) player is masked by the map processing unit 123 is input, the trained model outputs second inference information that is the predicted probability of each scenario shown in FIG. 11B.

[0070] The inference unit 124 compares the predicted probability based on a feature map (first feature map) including channel ch2 in which the position information of the enemy (3) player is not masked with the predicted probability based on a feature map (second feature map) including channel ch2 in which the position information of the enemy (3) player is masked. The predicted probability is the probability of inferring that the environment (scenario) is of a specific class. The influence of the channel information included in the feature map on the predicted probability can be evaluated.

[0071] 11B , the predicted probability of “one-on-one emphasis” based on the feature map including channel ch2 in which the position information of the enemy (3) player is not masked is 0.82, while the predicted probability of “one-on-one emphasis” based on the feature map including channel ch2 in which the position information of the enemy (3) player is masked is 0.56, showing a large change in predicted probability. In this way, the inference unit 124 receives the first feature map and the second feature map, and outputs first inference information and second inference information, which are the probabilities (predicted probabilities) of inferring that the environment is of a specific class.

[0072] The output unit 125 calculates the difference between the predicted probability based on the feature map including channel ch2 in which the position information of the opponent (3) player is not masked and the predicted probability based on the feature map including channel ch2 in which the position information of the opponent (3) player is masked, and outputs the calculated difference information to the display unit 14. The display unit 14 displays the difference information. By visually checking the difference information displayed on the display unit 14, it can be seen that by masking the position information of the opponent (3) player, the difference in predicted probability in the "one-on-one emphasis" scenario decreases by 0.26, the difference in predicted probability in the "under-the-basket emphasis" scenario increases by 0.22, and the difference in predicted probability in the "outside attack emphasis" scenario increases by 0.04. In other words, it can be seen that the predicted probability in the "one-on-one emphasis" scenario changes (decreases) the most, and that the opponent (3) player shooting from a dribble contributes significantly to the prediction in the "one-on-one emphasis" scenario.

[0073] The output unit 125 may output to the display unit 14 the information set in the channel ch2 to which the change has been made and information indicating that the masked object is the position information of the enemy (3) player. The display unit 14 displays the information output from the output unit 125. By visually checking the information displayed on the display unit 14, it is possible to recognize that the information affecting the "one-on-one focused" prediction is the position information of the enemy (3) player included in channel ch2. Furthermore, the output unit 125 may output to the display unit 14 either the information set in the channel ch2 to which the change has been made or the information indicating that the masked object is the position information of the enemy (3) player. The display unit 14 displays the information output from the output unit 125. Furthermore, the output unit 125 may output to the display unit 14 information that combines the information set in the channel ch2 to which the change has been made and the information indicating that the masked object is the position information of the enemy (3) player. The display unit 14 displays the information output from the output unit 125.

[0074] (Variation 2) Furthermore, the map processing unit 123 may perform processing on the portion of the objects included in the first channel that are performing dynamic actions. For example, as described above, the map processing unit 123 identifies the enemy (3) player, who is the object performing a shot after dribbling, from the position information of the enemy (1) to (5) players set in channel ch2, and masks only the position information of the enemy (3) player in channel ch2. In other words, the position information of this object is the portion of the object performing dynamic actions. In this way, even if the specific actions are not clear, by masking only the position information of the object performing dynamic actions, it is possible to evaluate the impact of the position information of this object on the inference results.

[0075] FIG. 12A shows a feature map converted from environmental information indicating a basketball simulation environment, and is the same feature map as FIG. 11A . FIG. 12B is a conceptual diagram outlining the process of masking channel ch2, to which the "enemy position" included in the feature map shown in FIG. 12A is set. The feature map shown in FIG. 12A is a time-series feature map converted from the time-series environmental information of one scenario (emphasis on one-on-one play) shown in FIG. 12B . Channel ch2 included in each feature map at each time is a first channel, to which position information indicating the positions of five enemy players (1) to (5), numbered 1 to 5 as "enemy positions," is set.

[0076] For example, in the scenario "Emphasis on One-on-One," it is assumed that the position information of the enemy (4) player, who does not take any specific action and moves little, is masked. In this case, the map processing unit 123 masks the position information of the enemy (4) player among the five enemy players (1) to (5) whose position information is set in channel ch2. The value "1" is assigned to the position information of the enemy players (1) to (5) included in channel ch2, and the value "0" is assigned to the other parts of the basketball court. As a result, as shown in FIG. 12B , in the binary image of the basketball court included in channel ch2, the image area where the enemy players (1) to (5) are present is displayed in white, and the other parts are displayed in black. When masking the position information of the enemy (4) player, the map processing unit 123 assigns the same value "0" to the position information of the enemy (4) player among the enemy players (1) to (5), as the other parts, thereby masking this player as shown in FIG. 12B .

[0077] The inference unit 124 outputs first inference information based on a feature map that has not been changed by the map processing unit 123, and outputs second inference information based on a feature map that has been changed by the map processing unit 123. For example, if the inference unit 124 is a trained model that predicts basketball strategies (scenarios), when a feature map that includes at least channel ch2 in which the position information of the opponent (4) player is not masked is input, the trained model outputs first inference information that is the predicted probability of each scenario shown in FIG. 12B. When a feature map that includes channel ch2 in which the position information of the opponent (4) player is masked by the map processing unit 123 is input, the trained model outputs second inference information that is the predicted probability of each scenario shown in FIG. 12B.

[0078] The inference unit 124 compares the predicted probability based on a feature map (first feature map) including channel ch2 in which the position information of the enemy (4) player is not masked with the predicted probability based on a feature map (second feature map) including channel ch2 in which the position information of the enemy (4) player is masked. The predicted probability is the probability of inferring that the environment (scenario) is of a specific class. The influence of the channel information included in the feature map on the predicted probability can be evaluated.

[0079] For example, as shown in FIG. 12B , the predicted probability of "emphasis on one-on-one play" based on a feature map including channel ch2 in which the position information of the opponent (4) player is not masked is 0.82, the predicted probability of "emphasis on basket," is 0.10, and the predicted probability of "emphasis on outside attack" is 0.10. Furthermore, the predicted probability of "emphasis on one-on-one play" based on a feature map including channel ch2 in which the position information of the opponent (4) player is masked is 0.80, the predicted probability of "emphasis on basket," is 0.12, and the predicted probability of "emphasis on outside attack" is 0.08. Thus, even if the position information of the opponent (4) player, who does not perform any specific actions and moves little, is masked, the change in predicted probability is small, and it can be seen that the position information of the opponent (4) player makes little contribution to the prediction of the "emphasis on one-on-one play" scenario using the discriminative model.

[0080] In addition, the map processing unit 123 masks various pieces of information, for example by changing the information to be masked among the pieces of information set for one channel, and by grasping the degree of contribution to each class, it is possible to obtain channel information that contributes greatly to the inference result, and as described above, it is also possible to obtain channel information that contributes greatly to the inference result, and it is also possible to obtain channel information that contributes little.

[0081] In this way, the information processing device 1 can acquire channel information or information on objects included in channels that significantly contributed to changes in the predicted probability based on the result of comparing the predicted probability based on a feature map including channels that have not been changed with the predicted probability based on a feature map including channels that have been changed. Furthermore, the extent to which the channel information or information on objects included in channels affects the inference result can be determined based on the result of the comparison of the predicted probabilities.

[0082] The information processing device 1 processes (changes) at least a portion of a channel included in a feature map converted from a single scenario, and acquires information on the contribution to changes in predicted probabilities based on a comparison between predicted probabilities based on an unprocessed feature map and predicted probabilities based on a processed feature map. The information processing device 1 may also process (change) at least a portion of a channel included in feature maps converted from multiple scenarios, and calculate, for each scenario, the difference between the predicted probabilities based on the unprocessed feature map and the predicted probabilities based on the processed feature map as the contribution. Furthermore, the information processing device 1 may calculate the contribution to all scenarios assumed in the simulation based on the sum, product, or average of the predicted probabilities calculated for each scenario. This allows the information processing device 1 to identify channels contributing to predictions in multiple scenarios.

[0083] As a method of applying changes to the feature map, for example, as shown in FIGS. 13A , 13B , and 13C , a change may be made to specific information in the time-series feature map, or a change may be made to information at a specific time. This allows the information processing device 1 to evaluate the influence of specific information in the time-series feature map on inference, and also to evaluate the influence of information at a specific time on inference. FIG. 13A shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all information related to actions has been changed. For example, the map processing unit 123 applies changes to each of the channels of channel information included in the feature map shown in FIG. 13A , which are channels set with information related to actions: channel ch4 indicating a “two-point shot,” channel ch5 indicating a “three-point shot,” channel ch6 indicating a “dribble,” and channel ch7 indicating a “pass.”

[0084] 13B shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all information at a specific time is changed. For example, the map processing unit 123 changes all of the channel information for channels ch1 to ch7 at the specific time included in the feature map shown in FIG.

[0085] 13C shows a time-series feature map converted from environmental information representing a basketball simulation environment, in which all information relating to a specific opponent has been changed. For example, the map processing unit 123 changes each of the channel information included in the feature map shown in FIG. 13C , which relates to the specific opponent, including channel ch2 indicating "opponent position," channel ch3 indicating "ball position," channel ch4 indicating "two-point shot," channel ch5 indicating "three-point shot," channel ch6 indicating "dribble," and channel ch7 indicating "pass."

[0086] Changes may be made to the time-series feature map shown in Figures 13A, 13B, and 13C, or these may be combined. Furthermore, the methods of making changes shown in Figures 8A, 8B, and 8C, the methods of making changes shown in Figures 11A, 11B, and 11C, and the methods of making changes shown in Figures 12A, 12B, and 12C may be combined, or these may be combined with the methods of making changes shown in Figures 13A, 13B, and 13C. This allows the information processing device 1 to evaluate the influence of specific information in the time-series feature map on inference, and to evaluate the influence of information at a specific time on inference.

[0087] (Variation 2) Figures 14A and 14B are diagrams showing an overview of a fighting game between bees and hornets. Figure 14A shows a simulation environment in which bees and hornets fight, and Figure 14B shows rules for the movements of each agent in the fighting environment shown in Figure 14A. A simulator for a fighting game between bees and hornets operates in a reinforcement learning environment, similar to the simulation environment for basketball described above, for example. In this game, each object behaves based on a rule base.

[0088] As shown in FIG. 14A , this battle environment includes competing bees and hornets, flowers from which the bees gather resources, a hornet's nest, and the beehive itself as agents (objects). As shown in FIG. 14B , hornets have three behavioral patterns: an "avoidance type" that prioritizes fleeing from the bees; a "direct attack type" that prioritizes approaching and attacking the bees; and a "base attack type" that prioritizes approaching and attacking the beehive. The hornets behave based on one of these behavioral patterns. The bees acquire positional information and behavioral information of the objects from the simulator and, based on the acquired information, identify which behavioral pattern the hornets' behavior belongs to. The information acquired from the simulator may be all information in the battle environment or may be limited information, such as only information identified within the bees' field of vision.

[0089] FIG. 15 is a schematic diagram illustrating an overview of a process for identifying multiple scenarios of a fighting game between bees and hornets using a machine learning model. In FIG. 15 , environmental information indicating multiple scenarios as the environment of the fighting game is time-series information consisting of environmental information for each time, and includes, for example, position history information indicating the position of an object at each time or behavior history information indicating the behavior of an object at each time. The environmental information acquisition unit 121 acquires the time-series environmental information from a simulator. The map generation unit 122 generates channels for each time based on the environmental information for each time, thereby generating a feature map consisting of time-series channels, as shown in FIG. 15 .

[0090] The map processing unit 123 processes at least one channel among the multiple channels included in the feature map. For example, the map processing unit 123 applies a change (e.g., a mask) to an object included in the channel. The inference unit 124 receives at least a portion of the feature map including channels to which no change has been applied and a feature map including channels to which changes have been applied by the map processing unit 123, and outputs first inference information and second inference information. For example, as shown in FIG. 15 , the first inference information and the second inference information include the accuracy rate of hornets avoiding attacks from honeybees, the accuracy rate of hornets approaching honeybees and attacking them directly, and the accuracy rate of hornets approaching beehives and attacking them. The inference unit 124 calculates the contribution of information from the changed channel to the inference result based on a comparison between the accuracy rate of the first inference information and the accuracy rate of the second inference information.

[0091] FIG. 16 is a schematic diagram showing an outline of a feature map converted from a scenario of a fighting game in which bees and hornets fight. The map generation unit 122 converts all or part of the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in FIG. 16. The environmental information is assigned a "1" to the location of the hornet, the location of the bee, the location of the flower, and the location of the beehive if each of these is present, and a "0" if each is not present. A "1" is assigned to each of the actions of attacking the bee, attacking the beehive, moving left, moving up, moving right, and moving down, and a "0" is assigned to each of these actions if each of these actions is not performed. Note that the feature map may have information indicating a combination of movement direction and action content, such as "attack right," set as a channel.

[0092] 16 is information consisting of four channels (ch1, ch2, ch3, ch4) in which object position information is set, and six channels (ch5, ch6, ch7, ch8, ch9, ch10) in which object behavior information is set. That is, for the position information set in channels ch1, ch2, ch3, and ch4, the map generation unit 122 sets "1" if an object exists, and sets "0" if an object does not exist.

[0093] (Variation 3) Fig. 17A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and Fig. 17B is a conceptual diagram showing an outline of the process of applying a mask to channel ch2 for the "position of hornet A" included in the feature map shown in Fig. 17A. Here, the feature map shown in Fig. 17A is a feature map converted from the time-series environmental information for each of the multiple scenarios shown in Fig. 17B. For example, it is expected that the "position of hornet A" will have an impact on identifying the hornet's strategy (scenario). Therefore, the map processing unit 123 applies a mask to the position information "position of hornet A" set to channel ch2 among the channels included in the feature map.

[0094] For example, the value "1" is assigned to the position information of the hornet set in channel ch2, and the value "0" is assigned to other parts. As a result, as shown in FIG. 17B , in the binary image of the battle environment included in channel ch2, the image area where the hornet exists is displayed in white, and the other parts are displayed in black. The map processing unit 123 assigns the same value "0" to "Hornet A" as the other parts, thereby masking the position information of "Hornet A" as shown in FIG. 17B .

[0095] If the inference unit 124 includes a trained model that identifies hornet tactics, when a feature map including channel ch2 in which the location information of "Hornet A" is not masked is input, the trained model outputs first inference information, which is the accuracy rate of each tactic, as shown in FIG. 17B. Furthermore, when a feature map including channel ch2 in which the location information of "Hornet A" is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the accuracy rate of each tactic, as shown in FIG. 17B. The tactics include "evasion," "direct attack," and "base attack," and the inference information is the accuracy rate of each tactic.

[0096] For example, suppose that the first inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is not masked shows that the accuracy rate for "avoidance" is 97%, the accuracy rate for "direct attack" is 82%, and the accuracy rate for "base attack" is 92%. Also, suppose that the second inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is masked shows that the accuracy rate for "avoidance" is 42%, the accuracy rate for "direct attack" is 80%, and the accuracy rate for "base attack" is 92%.

[0097] In this case, by masking the location information of "Hornet A," the accuracy rate of "Avoid" changed significantly (decreased) from 97% to 42%, so it can be determined that "Hornet A" has a significant impact on the prediction of the hornet's strategy. In this way, the information processing device 1 can acquire important channel information for each class. The trained model is generated by supervised learning, in which a training dataset is converted into a feature map in advance and the feature map is used as input. For example, a 3D CNN may be used to generate the trained model.

[0098] The output unit 125 outputs to the display unit 14 both the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is not masked and the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is masked. The display unit 14 displays each accuracy rate as shown in FIG. 17B . By visually checking the accuracy rates displayed on the display unit 14, it is possible to obtain an interpretation that the trained model has identified the tactic based on the position information of hornet A. Note that the output unit 125 may output to the display unit 14 either the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is not masked or the accuracy rate of the tactic identification based on the feature map in which the "position of hornet A" in channel ch2 is masked, and the display unit 14 may display the accuracy rates output from the output unit 125. In addition, the output unit 125 may output to the display unit 14 information that combines the accuracy rate of strategy identification based on a feature map in which the ``position of hornet A'' in channel ch2 is not masked and the accuracy rate of strategy identification based on a feature map in which the ``position of hornet A'' in channel ch2 is masked, and the display unit 14 may display the information output from the output unit 125.

[0099] (Variation 4) Fig. 18A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and Fig. 18B is a conceptual diagram showing an outline of a process of masking channel ch2 for the positions of all hornets included in the feature map shown in Fig. 18A. Here, the feature map shown in Fig. 18A is a time-series feature map converted from the time-series environmental information for each of the multiple scenarios shown in Fig. 18B. For example, if the accuracy rate of a specific strategy changes by masking the position information of all hornets, it is possible to interpret that the trained model identifies the hornets' strategies based on the position information of all hornets.

[0100] For example, the value "1" is assigned to the position information of the hornet set in channel ch2, and the value "0" is assigned to other parts. As a result, as shown in Fig. 18B, in the binary image of the battle environment included in channel ch2, the image area where the hornet exists is displayed in white, and the other parts are displayed in black. The map processing unit 123 assigns the same value "0" to all hornets as the other parts, thereby masking the position information of all hornets, as shown in Fig. 18B.

[0101] When a feature map including channel ch2 in which the location information of all hornets has not been masked is input, the trained model outputs first inference information, which is the accuracy rate of each tactic shown in Figure 18B. Furthermore, when a feature map including channel ch2 in which the location information of all hornets has been masked by map processing unit 123 is input, the trained model outputs second inference information, which is the accuracy rate of each tactic shown in Figure 18B. The inference information is the accuracy rate of each tactic: "evasion," "direct attack," and "base attack."

[0102] For example, suppose that the first inference information based on a feature map including channel ch2 in which the location information of all hornets is not masked has an accuracy rate of 97% for "avoidance," 82% for "direct attack," and 92% for "base attack." Also, suppose that the second inference information based on a feature map including channel ch2 in which the location information of all hornets is masked has an accuracy rate of 20% for "avoidance," 30% for "direct attack," and 88% for "base attack."

[0103] In this case, by masking the location information of all hornets, the accuracy rate for "evasion" drops significantly from 97% to 20%, and the accuracy rate for "direct attack" drops significantly from 82% to 30%. On the other hand, the accuracy rate for "base attack" only changes from 92% to 88%, a small decrease. In this way, by performing multiple masks on the location information of all hornets included in channel ch2 and determining the contribution to each class, it is possible to grasp information that contributes greatly to predicting hornet strategies, and also to grasp information that contributes little to predicting hornet strategies.

[0104] (Variation 5) Fig. 19A shows a feature map converted from a scenario of a fighting game in which bees and hornets fight, and Fig. 19B is a conceptual diagram showing an outline of a process for masking channel ch5 for bee attacks included in the feature map shown in Fig. 19A. Here, the feature map shown in Fig. 19A is a time-series feature map converted from multiple scenarios shown in Fig. 19B. For example, the environmental information acquisition unit 121 acquires environmental information indicating one scenario (tactic), and the map generation unit 122 generates a 10-channel feature map such as that shown in Fig. 16 from the environmental information, similar to multiple scenarios.

[0105] The map processing unit 123 masks all of the behavioral information "bee attack" set in channel ch5. When a feature map including at least channel ch5 in which all of the behavioral information "bee attack" is not masked is input, the trained model included in the inference unit 124 outputs first inference information, which is the predicted probability of each scenario shown in FIG. 19B. When a feature map including channel ch5 in which all of the behavioral information "bee attack" is masked by the map processing unit 123 is input, the trained model outputs second inference information, which is the predicted probability of each scenario shown in FIG. 19B.

[0106] 19B , based on a feature map including channel ch5 where "bee attacks" are not masked, the predicted probability of the scenario "avoidance" is 2%, the predicted probability of "direct attacks" is 92%, and the predicted probability of "base attacks" is 6%. Also, based on a feature map including channel ch5 where "bee attacks" are masked, the predicted probability of "avoidance" is 42%, the predicted probability of "direct attacks" is 38%, and the predicted probability of "base attacks" is 20%.

[0107] In this case, by masking all of the behavioral information of "bee attacks," the prediction rate of "evasion" changes (increases) significantly from 2% to 42%, the prediction rate of "direct attacks" changes (decreases) significantly from 92% to 38%, and the prediction rate of "base attacks" changes (increases) significantly from 6% to 20%. This makes it possible to determine that "bee attacks" contribute significantly to the tactical prediction of "direct attacks." In this way, by masking various pieces of information and performing inference, the information processing device 1 can determine channel information that contributes greatly to the prediction of each class, and can also determine channel information that contributes little to the prediction of each class.

[0108] Although supervised learning has been described, particularly with respect to classification (identification), a regression model may be used as the trained model if correct answer data is available. For example, in regression, the three-point shot emphasis rate and the two-point shot emphasis rate are expressed by weighting. Let x be the three-point shot emphasis rate and [x, 1-x]. x is regressed using machine learning of ([0.9, 0.1]). If x is perfectly correct and becomes ([0.9, 0.1]), but if the behavioral information of a three-point shot is masked and changes to [0.5, 0.5], it can be seen that the channel set with the behavioral information of a three-point shot makes a significant contribution when x = 0.9.

[0109] Furthermore, when environmental information is considered in terms of data sets, important channels can be identified by examining which combinations of channels included in the feature map, when masked, result in a larger root mean squared error (RMSE). A similar approach can be implemented for clustering, which is unsupervised learning. However, since matching with the correct answer is not possible, evaluation based on the accuracy rate is not possible. However, the information processing device 1 can output the contribution of each channel by focusing on changes in prediction probability. This makes it possible to present channels that serve as the basis for which cluster a device belongs. Furthermore, if the current cluster is assumed to be the correct answer, the contribution of each channel can be output. This makes it possible to present channels that serve as the basis for why a device belongs to that cluster.

[0110] (Variation 6) Fig. 20 is a schematic diagram showing an overview of a feature map converted from environmental information indicating the environment of real space, illustrating a view ahead of the vehicle as seen from an autonomous vehicle. The environmental information acquisition unit 121 acquires, as environmental information, an image captured from an onboard camera of the area ahead of the vehicle, object class information (e.g., a person or a car) acquired by object detection ahead of the vehicle, which is one type of machine learning, a parallax image of the area ahead of the vehicle, the distance to an object ahead of the vehicle measured by a sensor mounted on the vehicle, or speed information of an oncoming vehicle. The map generation unit 122 converts the environmental information acquired by the environmental information acquisition unit 121 into the feature map shown in Fig. 20.

[0111] 20, the environmental information indicates that at a certain time, there are objects at an intersection ahead of the vehicle, including a person crossing a crosswalk, a vehicle traveling ahead in the same lane, and a vehicle approaching in the oncoming lane. The map generation unit 122 assigns, for example, a value of "1" to the objects and a value of "0" to areas where no objects exist.

[0112] For example, the feature map shown in FIG. 20 includes channels ch1, ch2, and ch3. These channels include binary images showing the environment of the real space. Channel ch1 indicates the position of a car. In the binary image included in channel ch1, image areas corresponding to the position of the car are represented in white, and image areas where no car is present are represented in black. Channel ch2 indicates the position of a person. In the binary image included in channel ch2, image areas corresponding to the position of the person are represented in white, and image areas where no person is present are represented in black. Channel ch3 indicates an object that is "approaching" the vehicle. In the binary image included in channel ch3, image areas corresponding to the car that is "approaching" among the objects are represented in white, and the other image areas are represented in black.

[0113] When the inference unit 124 includes a trained model, the information processing device 1 can present information that the machine learning model focuses on during autonomous driving by inputting an unchanged feature map and an altered feature map. Furthermore, when an operation is performed on the vehicle, the information processing device 1 can present the reason for the operation.

[0114] The environment to be inferred may also be a surveillance area monitored by a surveillance camera. The environmental information acquisition unit 121 acquires, as environmental information, captured images of the surveillance area captured by the surveillance camera, object class information (e.g., people or ornaments) acquired by object detection in the surveillance area, which is one type of machine learning, or human behavior information. The map generation unit 122 converts the environmental information acquired by the environmental information acquisition unit 121 into a feature map. When the inference unit 124 includes a trained model, by inputting an unaltered feature map and an altered feature map, when a surveillance system equipped with a surveillance camera detects a suspicious person, it is possible to present the reason why the trained model detected the suspicious person as being based on what human behavior.

[0115] The object's location information or object's behavior information may be acquired by object recognition from image information indicating the environment. The information processing device 1 can be applied to any real-world environment as long as environmental information can be acquired using sensor information such as image recognition technology, a Global Poisoning System (GPS), a Time of Flight (ToF) sensor, or an infrared camera. When the environmental information acquisition unit 121 receives image information as environmental information, it recognizes objects included in the image information and detects the objects and their location information in the environment. The inference unit 124 predicts the object's behavior by image recognition and acquires behavior information indicating the predicted behavior. While the environmental information acquisition unit 121 acquires image information as environmental information, the object's location information may be detected using a sensor such as a GPS.

[0116] As described above, the information processing device 1 according to the first embodiment includes an environment information acquisition unit 121 that acquires environment information including position information or behavior information of objects included in the environment; a map generation unit 122 that generates a first feature map including multiple channels set based on the object position information or the object behavior information; a map processing unit 123 that processes at least a first channel among the multiple channels included in the first feature map; and an inference unit 124 that receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information, which is information related to inferring the environment based on at least a portion of the first feature map, and second inference information, which is information related to inferring the environment based on the second feature map. Based on the first inference information and second inference information related to the environment to be inferred, the information processing device 1 can acquire physical information in the environment that affects the result of inferring the environment. The conventional technology described in Non-Patent Document 1 randomly masks an image and searches for pixels that affect the image identification result. Furthermore, pixels have relative meaning but no spatial or physical meaning. That is, conventional technologies cannot extract physical information that contributes to an inference result, and are therefore unable to present what information has an overall impact on the accuracy rate of the inference result, etc. Furthermore, conventional technologies require numerous trials and are time-consuming to be able to interpret how pixels affect the inference result. In contrast, the information processing device 1 can acquire physical information in the environment, and therefore can also solve the above-mentioned problems.

[0117] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs the output first inference information and second inference information to the display unit 14. By visually checking the first inference information and the second inference information displayed on the display unit 14, it is possible to recognize the influence that the information processed in the first feature map has on the inference.

[0118] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs information based on the output first inference information and second inference information to the display unit 14. By visually checking the information displayed on the display unit 14, it is possible to grasp the magnitude of the influence that the information processed in the first feature map has on the inference.

[0119] The information processing device 1 according to the first embodiment includes an output unit 125 that outputs information of the first channel to the display unit 14. By visually checking the first channel displayed on the display unit 14, it is possible to recognize that the information that influences the inference is from the first channel.

[0120] In the information processing device 1 according to the first embodiment, the map processing unit 123 performs processing on the part of the object included in the first channel, thereby making it possible to evaluate the influence of the part of the object included in the first channel on the inference result.

[0121] In the information processing device 1 according to the first embodiment, the map processing unit 123 processes the part of the objects that are performing dynamic actions among the objects included in the first channel, thereby making it possible to evaluate the influence that the part of the objects that are performing dynamic actions among the objects included in the first channel has on the inference result.

[0122] In the information processing device 1 according to the first embodiment, the multiple channels include a third channel. The map generation unit 122 processes an area of ​​the environmental information that does not correspond to the position information to generate the third channel. This enables the information processing device 1 to generate the third channel including an image that clearly indicates the presence of an object.

[0123] In the information processing device 1 according to the first embodiment, the multiple channels include a fourth channel. The map generation unit 122 processes an area of ​​the environmental information that does not correspond to the behavioral information to generate the fourth channel. This enables the information processing device 1 to generate the fourth channel including an image that clearly shows the presence of an object performing a specific behavior.

[0124] In the information processing device 1 according to the first embodiment, the environment is an environment operated by a simulator. The object position information or object behavior information is information generated by the simulator. This allows the information processing device 1 to acquire physical information about the environment that affects the simulator's inference results about the environment.

[0125] In the information processing device 1 according to the first embodiment, the position information or behavior information of an object is acquired by object recognition from image information indicating an environment. This allows the information processing device 1 to acquire the information obtained by object recognition using image information as environmental information.

[0126] In the information processing device 1 according to the first embodiment, the inference unit 124 receives the first feature amount map and the second feature amount map, and outputs the first inference information and the second inference information by regression, thereby enabling the information processing device 1 to perform inference by regression.

[0127] In the information processing device 1 according to the first embodiment, the inference unit 124 receives the first feature map and the second feature map, and outputs first inference information and second inference information, which are probabilities of inferring that an environment belongs to a specific class. This makes it possible to evaluate the influence of the channel information included in the feature map on the probability of inferring that an environment belongs to a specific class.

[0128] In the information processing device 1 according to the first embodiment, the environment includes a plurality of classes. The inference unit 124 outputs first index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the first inference information and information on the correct class of the environment included in the environmental information, and outputs second index information, which is an index for measuring the inference performance of each of the plurality of classes, based on the second inference information and information on the correct class of the environment included in the environmental information. This allows the information processing device 1 to grasp the contribution of information to an inference result based on the first inference information and the second inference information.

[0129] In the information processing device 1 according to the first embodiment, the inference unit 124 includes a trained model, which enables the information processing device 1 to acquire physical information about the environment that affects the results of inferring the environment using the trained model.

[0130] The information processing method according to the first embodiment includes a step (ST1) in which an environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment, a step (ST2) in which a map generation unit 122 generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects, a step (ST3) in which a map processing unit 123 processes at least a first channel out of the plurality of channels included in the first feature map, and a step (ST4) in which an inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information that is information regarding environmental inference based on at least a portion of the first feature map and second inference information that is information regarding environmental inference based on the second feature map. By executing the information processing method according to the first embodiment, the information processing device 1 can acquire physical information in the environment that affects the environmental inference result.

[0131] A computer that executes a program according to the first embodiment executes the following steps: an environmental information acquisition unit 121 acquires environmental information including position information or behavior information of objects included in the environment (ST1); a map generation unit 122 generates a first feature map including multiple channels set based on the object position information or the object behavior information (ST2); a map processing unit 123 processes at least a first channel among the multiple channels included in the first feature map (ST3); and an inference unit 124 receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit 123, and outputs first inference information that is information regarding environmental inference based on at least a portion of the first feature map and second inference information that is information regarding environmental inference based on the second feature map (ST4). By executing this program, the computer functions as an information processing device 1, and is therefore able to acquire physical information in the environment that affects the results of environmental inference.

[0132] Embodiment 2. FIG. 21 is a block diagram showing an example configuration of an information processing device 1A according to embodiment 2. In FIG. 21, the information processing device 1A includes a communication unit 11, a calculation unit 12A, a storage unit 13, and a display unit 14. The communication unit 11 communicates with external devices via a network. The calculation unit 12A controls the overall operation of the information processing device 1A. Note that the information processing device 1A does not necessarily need to include the communication unit 11. For example, if environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1A does not include the communication unit 11, the environmental information acquisition unit 121 realized by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.

[0133] The calculation unit 12A includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123A, an inference unit 124, an output unit 125, and a map processing control unit 126. The calculation unit 12A executes an information processing application, thereby realizing the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123A, the inference unit 124, the output unit 125, and the map processing control unit 126. The storage unit 13 stores, for example, the information processing application and information used in the calculation processing of the calculation unit 12A. The storage unit 13 is a storage device included in a computer functioning as the information processing device 1A, and is a storage such as an HDD or SSD. Note that the storage unit 13 may be provided external to the information processing device 1A as long as it is accessible by the information processing device 1A. The storage unit 13 is also the memory 103 shown in FIG. 4.

[0134] The map processing control unit 126 performs control to generate a third feature map that includes at least the second channel, among the multiple channels, processed by the map processing unit 123A based on the first inference information and the second inference information. The first inference information is an inference result output from a trained model to which a first feature map that has not been processed by the map processing unit 123A has been input. The second inference information is an inference result output from the trained model to which a second feature map that has been processed by the map processing unit 123A has been input. The inference unit 124 receives the third feature map and outputs third inference information that is information regarding the inference of environmental information based on the third feature map.

[0135] FIG. 22 is a flowchart showing an information processing method according to the second embodiment. The map processing control unit 126 specifies, to the map processing unit 123A, the channels to which changes are to be applied from the feature map generated by the map generation unit 122 (step ST1A). For example, by comprehensively applying changes to the channels included in the feature map converted from environmental information indicating the simulation environment or the real environment in advance and performing the environmental inference described in the first embodiment, the channel that most contributed to the inference result or the combination of channels that contributed most to the inference result may be searched for. The searched channel or channel combination is set in the map processing control unit 126. When performing new inference in the simulation environment or the real environment, the map processing control unit 126 specifies a pre-set channel or combination of channels to the map processing unit 123A as the channel to which changes are to be applied. Alternatively, the channels or channel combinations that may be related to the simulation environment or the real environment may be narrowed down, and the map processing control unit 126 may specify the channel to which changes are to be applied from the pre-narrowed channels or combinations. Note that the channel narrowing down may be specified by the user, for example, using an input device (not shown).

[0136] The map processing unit 123A processes a channel specified by the map processing control unit 126 among the multiple channels included in the feature map (step ST2A). For example, the map processing unit 123A makes changes to the specified channel. The inference unit 124 receives a feature map including unprocessed channels and a feature map including processed channels, and outputs first inference information, which is information about the inference of the environment based on the feature map including the unprocessed channels, and second inference information, which is information about the inference of the environment based on the feature map including the processed channels (step ST3A). Next, the inference unit 124 searches for channels that contribute to the inference result based on the results of comparing the first inference information and the second inference information (step ST4A). For example, if a channel that contributes highly to the inference result is found (step ST4A; YES), the process of FIG. 22 ends. On the other hand, if the inference unit 124 does not find a channel that contributes highly to the inference result (step ST4A; NO), the process returns to step ST1A and repeats the series of processes shown in FIG. 22. This makes it possible to find channels that affect changes in the accuracy rate or prediction probability, which are the inference results of the inference unit 124.

[0137] As described above, the information processing device 1A according to the second embodiment includes a map processing control unit 126 that performs control to generate a third feature map including at least the second channel processed by the map processing unit 123A based on the first inference information and the second inference information among the multiple channels. The inference unit 124 receives the third feature map and outputs the third inference information, which is information related to the inference of environmental information based on the third feature map. This enables the information processing device 1A to find channels that affect changes in the accuracy rate or prediction probability, which are the inference results.

[0138] Embodiment 3. FIG. 23 is a block diagram showing an example configuration of an information processing device 1B according to embodiment 3. In FIG. 23, the information processing device 1B includes a communication unit 11, a calculation unit 12B, a storage unit 13, and a display unit 14. The communication unit 11 communicates with external devices via a network. The calculation unit 12B controls the overall operation of the information processing device 1B. Note that the information processing device 1B does not necessarily need to include the communication unit 11. For example, if environmental information generated by a simulation using a simulator is stored in the storage unit 13, even if the information processing device 1B does not include the communication unit 11, the environmental information acquisition unit 121 realized by the calculation unit 12 can acquire the environmental information directly from the storage unit 13.

[0139] The calculation unit 12B includes an environmental information acquisition unit 121, a map generation unit 122, a map processing unit 123, an inference unit 124A, and an output unit 125A. The calculation unit 12B executes an information processing application, thereby realizing the functions of the environmental information acquisition unit 121, the map generation unit 122, the map processing unit 123, the inference unit 124A, and the output unit 125A. The storage unit 13 stores, for example, the information processing application, information used in the calculation processing of the calculation unit 12B, and a large-scale language model 13a (hereinafter referred to as LLM 13a). The storage unit 13 is a storage device included in a computer functioning as the information processing device 1B, and is a storage such as an HDD or SSD. Note that the storage unit 13 may be provided external to the information processing device 1B as long as it is accessible by the information processing device 1B.

[0140] The LLM 13a is a language model that performs natural language processing and is trained using a large data set. An example of the LLM 13a is Open AI's Generative Pre-trained Transformer (GPT). By processing a huge amount of text data, the LLM 13a learns the associations and contexts of words and phrases, and is able to generate sentences that are appropriate to the context.

[0141] The inference unit 124A outputs first inference information and second inference information to the LLM 13a. The first inference information is an inference result output from a trained model to which a first feature map that has not been processed by the map processing unit 123 has been input. The second inference information is an inference result output from the trained model to which a second feature map that has been processed by the map processing unit 123 has been input. The LLM 13a outputs supplemental information based on the first inference information and the second inference information. Here, the supplemental information is information for explaining which information contributed to the inference result by the inference unit 124A. The output unit 125A outputs the first inference information, the second inference information, and the supplemental information to the display unit 14. As a result, the display unit 14 displays the first inference information, the second inference information, and the supplemental information.

[0142] FIG. 24 is a conceptual diagram showing an overview of processing by information processing device 1B. In FIG. 24, environmental information indicating multiple scenarios as the environment of a competitive game is time-series information consisting of environmental information for each time, and includes, for example, position history information indicating the position of an object at each time or behavior history information indicating the behavior of an object at each time. FIG. 25 is a flowchart showing an information processing method according to embodiment 3. For example, the environmental information acquisition unit 121 acquires environmental information such as that shown in FIG. 24 from a simulator (step ST1B). The map generation unit 122 generates a feature map including the channels shown in FIG. 24 based on the environmental information for each time (step ST2B).

[0143] The map processing unit 123 processes at least one channel among the multiple channels included in the feature map (step ST3B). For example, the map processing unit 123 applies a change (e.g., a mask) to an object included in the channel. The inference unit 124A receives at least a portion of the feature map including channels to which no changes have been applied and a feature map including channels to which changes have been applied by the map processing unit 123, and outputs first inference information and second inference information (step ST4B). For example, as shown in FIG. 24 , the first inference information and the second inference information include the accuracy rate of hornets avoiding attacks from honeybees, the accuracy rate of hornets approaching honeybees and attacking them directly, and the accuracy rate of hornets approaching honeybees and attacking them. The inference unit 124A calculates the contribution of information from the changed channel to the inference result based on a comparison between the accuracy rate of the first inference information and the accuracy rate of the second inference information.

[0144] The LLM 13a receives from the inference unit 124A the feature map, information contained in the channels of the feature map, information obtained by processing (changing) this information by the map processing unit 123, the first inference information, and the second inference information (step ST5B). When this information is input, the LLM 13a outputs supplemental information explaining which information contributes to the inference result by the machine learning model (trained model) included in the inference unit 124A (step ST6B).

[0145] For example, suppose that the first inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is not masked is 97% accurate, the "direct attack" is 82% accurate, and the "base attack" is 92% accurate. Furthermore, suppose that the second inference information based on a feature map including channel ch2 in which the location information of "Hornet A" is masked is 42% accurate, the "direct attack" is 80% accurate, and the "base attack" is 92% accurate. In this case, because masking the location information of "Hornet A" significantly changes (decreases) the accuracy rate of "Evasion" from 97% to 42%, the LLM 13a outputs supplemental information such as, for example, "Hornet A has a significant influence on the hornet's strategy predictions."

[0146] When environmental information is input, the LLM 13a may divide the input environmental information into various channels and output a plurality of feature map candidates, each representing a corresponding channel as text data. For example, the LLM 13a outputs feature map candidates including text data such as "location of hornet," "location of bee," "location of flower," "beehive," "attack bee," "attack beehive," "move left," "move up," "move right," and "move down." The feature map candidates are output from the LLM 13a to the output unit 125A and then to the map generation unit 122. The output unit 125A outputs the feature map candidates to the display unit 14, which displays the feature map candidates. The map generation unit 122 generates a feature map corresponding to a feature map candidate selected using an input device (not shown) or automatically selected.

[0147] The information processing device 1B may also include a map processing control unit 126. As in the second embodiment, the map processing control unit 126 specifies, in the feature amount map, a channel to which a change is to be added, to the map processing unit 123. For example, the map processing control unit 126 may automatically determine candidate channels to be specified to the map processing unit 123 based on supplemental information output from the LLM 13a.

[0148] As described above, in the information processing device 1B according to the third embodiment, the inference unit 124A outputs the first inference information and the second inference information to the LLM 13a. The LLM 13a outputs supplemental information based on the first inference information and the second inference information. The output unit 125A outputs the first inference information, the second inference information, and the supplemental information to the display unit 14. By referring to the supplemental information, it is possible to improve interpretability by explaining which input information contributes to the inference made by the inference unit 124A.

[0149] Various aspects of the present disclosure are summarized below as appendices.

[0150] (Supplementary Note 1) An information processing device comprising: an environmental information acquisition unit that acquires environmental information including position information of objects included in an environment or behavioral information of the objects; a map generation unit that generates a first feature map including a plurality of channels set based on the position information of the objects or the behavioral information of the objects; a map processing unit that processes at least a first channel of the plurality of channels included in the first feature map; and an inference unit that receives as input at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information that is information regarding the inference of the environment based on the second feature map. (Supplementary Note 2) The information processing device according to Supplementary Note 1, further comprising: an output unit that outputs the outputted first inference information and the second inference information to a display device. (Supplementary Note 3) The information processing device according to Supplementary Note 1, further comprising: an output unit that outputs information based on the outputted first inference information and the second inference information to a display device. (Supplementary Note 4) The information processing device according to Supplementary Note 1, further comprising an output unit that outputs information of the first channel to a display device. (Supplementary Note 5) The information processing device according to any one of Supplementary Notes 1 to 4, further comprising: the map processing unit performing the processing on a portion of the object included in the first channel. (Supplementary Note 6) The information processing device according to Supplementary Note 5, further comprising: the map processing unit performing the processing on a portion of the object that is performing a dynamic action, among the objects included in the first channel.(Supplementary Note 7) The information processing device according to Supplementary Note 6, further comprising: a map processing control unit that controls generation of a third feature amount map including at least a second channel among the plurality of channels, the second channel having undergone the processing by the map processing unit based on the first inference information and the second inference information, wherein the inference unit receives the third feature amount map as input and outputs third inference information that is information related to the inference of the environmental information based on the third feature amount map. (Supplementary Note 8) The information processing device according to any one of Supplementary Notes 1 to 7, further comprising: the plurality of channels including a third channel; and the map generation unit performs the processing on a region of the environmental information that does not correspond to the position information, thereby generating the third channel. (Supplementary Note 9) The information processing device according to any one of Supplementary Notes 1 to 8, further comprising: the plurality of channels including a fourth channel; and the map generation unit performs the processing on a region of the environmental information that does not correspond to the behavior information, thereby generating the fourth channel. (Supplementary Note 10) The information processing device according to any one of Supplementary Notes 1 to 9, wherein the environment is an environment operated by a simulator, and the position information of the object or the behavior information of the object is information generated by the simulator. (Supplementary Note 11) The information processing device according to any one of Supplementary Notes 1 to 10, wherein the position information of the object or the behavior information of the object is acquired by object recognition from image information indicating the environment. (Supplementary Note 12) The information processing device according to any one of Supplementary Notes 1 to 11, wherein the inference unit receives the first feature map and the second feature map as input, and outputs the first inference information and the second inference information by regression. (Supplementary Note 13) The information processing device according to any one of Supplementary Notes 1 to 12, wherein the inference unit receives the first feature map and the second feature map as input, and outputs the first inference information and the second inference information, which are probabilities of inferring the environment as being of a specific class.(Supplementary Note 14) The information processing device according to any one of Supplementary Notes 1 to 13, wherein the environment includes a plurality of classes, and the inference unit outputs first index information which is an index for measuring inference performance of each of the plurality of classes based on the first inference information and information on a correct class of the environment included in the environmental information, and outputs second index information which is an index for measuring inference performance of each of the plurality of classes based on the second inference information and information on a correct class of the environment included in the environmental information. (Supplementary Note 15) The information processing device according to any one of Supplementary Notes 12 to 14, wherein the inference unit includes a trained model. (Supplementary Note 16) The information processing device according to any one of Supplementary Notes 12 to 15, wherein the inference unit outputs the first inference information and the second inference information to a language model, and the language model outputs supplemental information based on the first inference information and the second inference information, and further comprises an output unit which outputs the first inference information, the second inference information, and the supplemental information to a display device. (Supplementary Note 17) An information processing method executed by an information processing device, comprising: an environmental information acquisition unit acquiring environmental information including position information of objects included in an environment or behavior information of the objects; a map generation unit generating a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a map processing unit processing at least a first channel out of the plurality of channels included in the first feature map; and an inference unit receiving input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputting first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map and second inference information which is information regarding the inference of the environment based on the second feature map.(Supplementary Note 18) A program for causing a computer to execute the following steps: an environmental information acquisition unit acquires environmental information including position information of objects included in an environment or behavior information of the objects; a map generation unit generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a map processing unit processes at least a first channel out of the plurality of channels included in the first feature map; and an inference unit receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information which is information regarding the inference of the environment based on the second feature map.

[0151] It is possible to combine the embodiments, modify any of the components of the embodiments, or omit any of the components of the embodiments.

[0152] The information processing device according to the present disclosure can be used for inference using a machine learning model.

[0153] 1, 1A, 1B Information processing device, 11 Communication unit, 12, 12A, 12B Calculation unit, 13 Storage unit, 13a Large-scale language model, 14 Display unit, 100 Communication interface, 101 Input / output interface, 102 Processor, 103 Memory, 121 Environmental information acquisition unit, 122 Map generation unit, 123, 123A Map processing unit, 124, 124A Inference unit, 125, 125A Output unit, 126 Map processing control unit.

Claims

1. An information processing device comprising: an environmental information acquisition unit that acquires environmental information including position information of objects included in an environment or behavioral information of the objects; a map generation unit that generates a first feature map including a plurality of channels set based on the position information of the objects or the behavioral information of the objects; a map processing unit that processes at least a first channel out of the plurality of channels included in the first feature map; and an inference unit that receives as input at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information that is information regarding inference of the environment based on at least a portion of the first feature map, and second inference information that is information regarding inference of the environment based on the second feature map.

2. The information processing device according to claim 1, further comprising an output unit that outputs the output first inference information and the output second inference information to a display device.

3. The information processing device according to claim 1, further comprising an output unit that outputs information based on the output first inference information and the output second inference information to a display device.

4. The information processing device according to claim 1, further comprising an output unit that outputs the information of the first channel to a display device.

5. An information processing device according to any one of claims 1 to 4, characterized in that the map processing unit performs the processing on the part of the object included in the first channel.

6. The information processing device according to claim 5, wherein the map processing unit performs the processing on a portion of the objects that are performing dynamic actions among the objects included in the first channel.

7. An information processing device as described in claim 6, further comprising a map processing control unit that controls the generation of a third feature map that includes at least a second channel among the plurality of channels, the second channel having undergone the processing by the map processing unit based on the first inference information and the second inference information, wherein the inference unit receives the third feature map as input and outputs third inference information that is information regarding the inference of the environmental information based on the third feature map.

8. An information processing device as described in any one of claims 1 to 7, characterized in that the multiple channels include a third channel, and the map generation unit performs the processing on an area of ​​the environmental information that does not correspond to the location information, and generates the third channel.

9. An information processing device as described in any one of claims 1 to 8, characterized in that the multiple channels include a fourth channel, and the map generation unit performs the processing on an area of ​​the environmental information that does not correspond to the behavioral information, and generates the fourth channel.

10. An information processing device according to any one of claims 1 to 9, characterized in that the environment is an environment operated by a simulator, and the position information of the object or the behavior information of the object is information generated by the simulator.

11. An information processing device according to any one of claims 1 to 10, characterized in that the position information of the object or the behavior information of the object is acquired by object recognition from image information showing the environment.

12. An information processing device described in any one of claims 1 to 11, characterized in that the inference unit receives the first feature map and the second feature map as input, and outputs the first inference information and the second inference information by regression.

13. An information processing device as described in any one of claims 1 to 12, characterized in that the inference unit receives the first feature map and the second feature map as input, and outputs the first inference information and the second inference information, which are probabilities of inferring that the environment is of a specific class.

14. An information processing device as described in any one of claims 1 to 13, characterized in that the environment includes a plurality of classes, and the inference unit outputs first index information which is an index for measuring the inference performance of each of the plurality of classes based on the first inference information and information on the correct class of the environment included in the environmental information, and outputs second index information which is an index for measuring the inference performance of each of the plurality of classes based on the second inference information and information on the correct class of the environment included in the environmental information.

15. An information processing device according to any one of claims 1 to 14, characterized in that the inference unit includes a trained model.

16. An information processing device as described in any one of claims 1 to 15, characterized in that the inference unit outputs the first inference information and the second inference information to a language model, the language model outputs supplementary information based on the first inference information and the second inference information, and the information processing device is further provided with an output unit that outputs the first inference information, the second inference information, and the supplementary information to a display device.

17. An information processing method executed by an information processing device, comprising: a step in which an environmental information acquisition unit acquires environmental information including position information of objects included in the environment or behavior information of the objects; a step in which a map generation unit generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a step in which a map processing unit processes at least a first channel out of the plurality of channels included in the first feature map; and a step in which an inference unit receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information which is information regarding the inference of the environment based on the second feature map.

18. A program for causing a computer to execute the following steps: an environmental information acquisition unit acquires environmental information including position information of objects included in the environment or behavior information of the objects; a map generation unit generates a first feature map including a plurality of channels set based on the position information of the objects or the behavior information of the objects; a map processing unit processes at least a first channel out of the plurality of channels included in the first feature map; and an inference unit receives input of at least a portion of the first feature map including the first channel and a second feature map including the first channel processed by the map processing unit, and outputs first inference information which is information regarding the inference of the environment based on at least a portion of the first feature map, and second inference information which is information regarding the inference of the environment based on the second feature map.