Information processing device, information processing method, and information processing program
Patent Information
- Application Number
- JP2026516901
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2024-07-05
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2044-07-05
Smart Images

Figure 0007927204000001 
Figure 0007927204000002 
Figure 0007927204000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to causality-based reasoning.
Background Art
[0002] Conventional models using an attention mechanism have long faced the challenge of balancing the increase in computational complexity against the improvement of prediction accuracy. In a conventional attention mechanism, since weighting is performed for each element of input sequence data, the computational complexity O(n 2 ) increases secondarily. "n" represents the length of the sequence data. Along with this increase in computational complexity, the problem that model inference time increases in large-scale data or real-time applications has arisen. Furthermore, for improving prediction accuracy, in order for the attention mechanism to function properly, it is necessary to accurately grasp the structural causal relationships hidden in sequence data and determine an appropriate computational range. However, with conventional methods, grasping of causal relationships and determination of the computational range are insufficient, which may restrict the improvement of prediction accuracy.
[0003] Patent Document 1 discloses a model acquisition method that improves model performance while reducing consumption of computational resources. This method trains a model in accordance with a learning objective corresponding to syntactic information for a self-attention module.
Prior Art Literature
Patent Literature
[0004]
Patent Document 1
Summary of the Invention
Problem to be Solved by the Invention
[0005] This disclosure aims to enable the calculation of the feature quantities of each of the multiple elements shown in sequential data while suppressing an increase in computational complexity. [Means for solving the problem]
[0006] The information processing device disclosed herein is A causal relationship estimation unit takes sequence data representing a sequence of multiple elements as input and uses a causal relationship estimation model to estimate the causal relationships between the multiple elements, A chunk determination unit determines, for each element, the chunk that constitutes the calculation range of the self-attention mechanism based on the estimated causal relationship, A self-attention calculation unit that calculates the self-attention value for each of the multiple elements by the self-attention mechanism, with each of the chunks of the multiple elements set to the calculation range, A feature calculation unit that takes the aforementioned self-attention values for each of the aforementioned elements as input and uses an inference model to calculate the feature quantities for each of the aforementioned elements, It is equipped with. [Effects of the Invention]
[0007] According to this disclosure, chunks are determined based on the causal relationships between multiple elements shown in the sequential data. This makes it possible to calculate the feature vectors of each of the multiple elements shown in the sequential data while suppressing the increase in computational complexity. [Brief explanation of the drawing]
[0008] [Figure 1] A diagram showing the configuration of the information processing device 100 in Embodiment 1. [Figure 2] Configuration diagram of the inference unit 110 in Embodiment 1. [Figure 3] Configuration diagram of the learning unit 120 in Embodiment 1. [Figure 4] Flowchart of the inference process in Embodiment 1. [Figure 5] This figure shows an example of a causal relationship graph 192 and a matrix 193 in Embodiment 1. [Figure 6]Figure showing an example of the matrix 193 in the first embodiment. [Figure 7] Figure showing an example of the parent-child node group 195 and the chunk 194 in the first embodiment. [Figure 8] Figure showing an example of the chunk 194 in the first embodiment. [Figure 9] Figure showing an example of the chunk 194 in the first embodiment. [Figure 10] Figure showing an example of the chunk 194 in the first embodiment. [Figure 11] Figure showing a flowchart of the learning process in the first embodiment. [Figure 12] Figure showing a flowchart of the learning process in the first embodiment. [Figure 13] Hardware configuration diagram of the information processing apparatus 100 in the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] In the embodiments and the drawings, the same or corresponding elements are assigned the same reference numerals. Descriptions of elements assigned the same reference numerals as already described elements will be omitted or simplified as appropriate. Arrows in the drawings mainly indicate the flow of data or the flow of processing.
[0010] First Embodiment An embodiment for calculating the feature amount of each of a plurality of elements included in sequence data while suppressing an increase in computational complexity will be described with reference to FIGS. 1 to 13.
[0011] ***Description of Configuration*** The configuration of the information processing apparatus 100 will be described based on FIG. 1. The information processing apparatus 100 is a computer including hardware such as a processor 101, a memory 102, an auxiliary storage device 103, a communication device 104, and an input / output interface 105. These pieces of hardware are connected to each other via signal lines.
[0012] The processor 101 is an IC that performs arithmetic operations and controls other hardware. For example, the processor 101 is a CPU, DSP, or GPU. IC is an abbreviation for Integrated Circuit. CPU is an abbreviation for Central Processing Unit. DSP is an abbreviation for Digital Signal Processor. GPU is an abbreviation for Graphics Processing Unit.
[0013] Memory 102 is a volatile or non-volatile storage device. Memory 102 is also called main memory. For example, memory 102 is RAM. Data stored in memory 102 is saved to auxiliary storage device 103 as needed. RAM is an abbreviation for Random Access Memory.
[0014] The auxiliary storage device 103 is a non-volatile storage device. For example, the auxiliary storage device 103 is a ROM, HDD, flash memory, or a combination thereof. Data stored in the auxiliary storage device 103 is loaded into memory 102 as needed. ROM is an abbreviation for Read Only Memory. HDD is an abbreviation for Hard Disk Drive.
[0015] The communication device 104 is a receiver and transmitter. For example, the communication device 104 is a communication chip or NIC. Communication of the information processing device 100 is performed using the communication device 104. NIC is an abbreviation for Network Interface Card.
[0016] The input / output interface 105 is a port to which input and output devices are connected. For example, the input / output interface 105 is a USB terminal, the input devices are a keyboard and mouse, and the output device is a display. Input and output of the information processing device 100 are performed using the input / output interface 105. USB is an abbreviation for Universal Serial Bus.
[0017] The information processing device 100 includes elements such as an inference unit 110 and a learning unit 120. These elements are implemented in software.
[0018] The auxiliary storage device 103 stores information processing programs that enable the computer to function as an inference unit 110 and a learning unit 120. The information processing programs are loaded into memory 102 and executed by the processor 101. The auxiliary storage device 103 also stores the operating system. At least a portion of the OS is loaded into memory 102 and executed by the processor 101. Processor 101 executes information processing programs while running the operating system. OS is an abbreviation for Operating System.
[0019] The data for the information processing program (input data, output data, etc.) is stored in the storage unit 190. Memory 102 functions as a storage unit 190. However, storage devices such as auxiliary storage device 103, registers in the processor 101, and cache memory in the processor 101 may function as a storage unit 190 instead of memory 102, or together with memory 102.
[0020] Information processing programs can be recorded (stored) in a computer-readable format on non-volatile recording media such as optical discs or flash memory.
[0021] The configuration of the inference unit 110 will be explained based on Figure 2. The inference unit 110 comprises a causal relationship estimation unit 111, a chunk determination unit 112, a self-attention calculation unit 113, and a feature calculation unit 114. Parameter data 191, sequence data 181, and inference data 182 are stored in the storage unit 190.
[0022] Based on Figure 3, the configuration of the learning unit 120 will be explained. The learning unit 120 includes a causal relationship estimation unit 121, a chunk determination unit 122, a self-attention calculation unit 123, a feature calculation unit 124, and a learning control unit 125. The training data 183 and the correct answer data 184 are stored in the memory unit 190.
[0023] ***Explanation of operation*** The operating procedure of the information processing device 100 corresponds to the information processing method. Furthermore, the operating procedure of the information processing device 100 corresponds to the processing procedure of the information processing program.
[0024] Based on Figure 4, the inference process of the information processing method will be explained. The trained parameters of the causal relationship estimation model and the inference model are stored in parameter data 191. For example, the parameters of the neural network (NN) used in each model are stored in parameter data 191. Causal relationship estimation models and inference models will be discussed later.
[0025] In step S111, the causal relationship estimation unit 111 obtains the parameters of the causal relationship estimation model from the parameter data 191 and sets them in the causal relationship estimation model. Furthermore, the feature calculation unit 114 obtains the parameters of the inference model from the parameter data 191 and sets them in the inference model.
[0026] In step S112, the causal relationship estimation unit 111 acquires the series data 181. For example, a user inputs sequence data 181 into the information processing device 100, and the causal relationship estimation unit 111 receives the input sequence data 181.
[0027] Series data is data that represents a sequence of multiple elements. For example, sequential data can represent multiple numbers arranged in a time series or multiple words arranged along a sentence. Examples of sequential data include natural language text data or FA sensor data. FA is an abbreviation for factory automation.
[0028] The causal relationship estimation unit 111 then takes the sequence data 181 as input and estimates the causal relationships between multiple elements using a causal relationship model. This generates a causal relationship graph 192, which shows the estimated causal relationship.
[0029] Let's explain the causal relationship model. The causal relationship model is a machine learning model that estimates the causal relationships between multiple elements shown in the input data. The causal relationship model is stored in the memory unit 190. Causal relationship models, for example, estimate causal relationships using convex optimization methods. The causal relationship model outputs a causal relationship graph 192 as data showing the estimated causal relationships.
[0030] Causal relationship graph 192 shows the linear causal relationships between multiple elements. Specifically, the causal relationship graph 192 shows multiple elements as nodes, and connects nodes that have a direct causal relationship with paths.
[0031] Causal relationship graph 192 includes path coefficients for each path. The path coefficient indicates the strength of the causal relationship between nodes connected by a path.
[0032] The causal relationship graph 192 can be represented by matrix 193.
[0033] Figure 5 shows examples of causal relationship graph 192 and matrix 193. Sequence data 181 represents a sequence of nine elements. These nine elements can be extracted, for example, using a neural network.
[0034] The causal relationship graph 192 has nine nodes (squares) representing nine elements. In the causal relationship graph 192, nodes that have a direct causal relationship are connected by paths (arrows). Each path is assigned a path coefficient that indicates the strength of the causal relationship between the nodes connected by that path. In a path connecting two nodes, one node (the starting point of the arrow) is called the parent node, and the other node (the ending point of the arrow) is called the child node. For example, the sixth element node has a direct causal relationship with the third, fifth, eighth, and ninth element nodes, respectively. The third element node is the parent node of the sixth element. The fifth, eighth, and ninth element nodes are child nodes of the sixth element.
[0035] Matrix 193 has nine queries (rows) and nine keys (columns) that are the same as the nine elements. The attention described later is calculated using the keys that are either the parent node or child node of the query. The shaded areas in matrix 193 indicate node pairs that have a direct causal relationship.
[0036] Figure 6 shows an example of matrix 193. Matrix 193 can be replaced by matrix 193I and matrix 193T. Matrix 193I is obtained by adding the identity matrix to matrix 193. Matrix 193T is obtained by adding the transpose and identity matrix to matrix 193.
[0037] Returning to Figure 4, we will continue the explanation from step S113. In step S113, the chunk determination unit 112 determines chunks 194 for each element based on the estimated causal relationships.
[0038] Chunk 194 is the computational range for the self-attention mechanism.
[0039] Specifically, the chunk determination unit 112 selects a group of parent-child nodes 195 from the causal relationship graph 192 for each element, and determines a chunk 194 for each element based on the group of parent-child nodes 195. The parent-child node group 195 consists of a target element node representing the target element, and nodes that have a parent-child relationship with the target element node.
[0040] For example, the chunk determination unit 112 determines the chunk 194 based on the parent-child node group 195 as follows. The chunk determination unit 112 determines the group of elements corresponding to the parent-child node group 195 as chunk 194. Figure 7 shows an example of a parent-child node group of 195 and a chunk of 194. The parent-child node group 195 of the sixth element consists of the third element node, the fifth element node, the sixth element node, the eighth element node, and the ninth element node. In this case, the third, fifth, sixth, eighth, and ninth elements, which correspond to the parent-child node group of the sixth element, become chunk 194 of the sixth element. The length of chunk 194 of the sixth element is 5.
[0041] Returning to Figure 4, we continue the explanation of step S113. For example, the chunk determination unit 112 determines the chunk 194 based on the parent-child node group 195 as follows. The chunk determination unit 112 compares the number of nodes in the parent-child node group 195 with a constant. The constant is a predetermined fixed value. If the number of nodes in the parent-child node group 195 is the same as a constant, the chunk determination unit 112 determines the group of elements corresponding to the parent-child node group 195 to be chunk 194. If the number of nodes in the parent-child node group 195 is less than a constant, the chunk determination unit 112 determines the group of elements corresponding to the parent-child node group 195 and one or more dummy elements to form chunk 194. If the number of nodes in the parent-child node group 195 exceeds a constant, the chunk determination unit 112 determines that the same number of elements as the constant from the element group corresponding to the parent-child node group 195 will be chunk 194. Specifically, the chunk determination unit 112 selects the same number of elements as the constant from the element group corresponding to the parent-child node group 195 as chunk 194 based on the path coefficient group for the parent-child node group 195.
[0042] Figure 8 shows an example of chunk 194. The parent-child node group 195 of the first element consists of the first element node, the seventh element node, the fourth element node, and the second element node. The number of nodes in the parent-child node group 195 of the first element is 4. When the constant is 4, the number of nodes in the parent-child node group 195 of the first element is the same as the constant. In this case, the 1st, 7th, 4th, and 2nd elements become chunk 194 of the 1st element. In chunk 194, the remaining elements (7th, 4th, and 2nd elements), excluding the target element (1st element), are sorted in descending order of their path coefficients.
[0043] Figure 9 shows an example of chunk 194. The parent-child node group 195 of the second element consists of the second element node, the third element node, and the first element node. The number of nodes in the parent-child node group 195 of the second element is 3. If the constant is 4, the number of nodes in the parent-child node group 195 of the second element is less than the constant. In this case, the second, third, first, and zeroth elements become chunk 194 of the second element. The zeroth element is a dummy element that is completed for chunk 194. In chunk 194, the remaining elements (element 1 and element 3), excluding the target element (element 2) and the dummy element (element 0), are sorted in descending order of their path coefficients.
[0044] Figure 10 shows an example of chunk 194. The parent-child node group 195 of the sixth element consists of the sixth element node, the ninth element node, the eighth element node, the third element node, and the fifth element node. The number of nodes in the parent-child node group 195 of the sixth element is 5. If the constant is 4, the number of nodes in the parent-child node group 195 of the sixth element exceeds the constant. In this case, the parent-child node group 195 of the sixth element is sorted in descending order of path coefficients. Then, a constant number of elements are selected in descending order of path coefficients, and these selected elements become chunk 194 of the sixth element. Chunk 194 of the sixth element consists of the sixth, ninth, eighth, and third elements. The fifth element is not included in chunk 194 of the sixth element. In other words, the fifth element is not used.
[0045] Returning to Figure 4, we will continue the explanation from step S114. In step S114, the self-attention calculation unit 113 uses each chunk 194 of the multiple elements as the calculation range and calculates the self-attention value for each of the multiple elements using the self-attention mechanism.
[0046] The self-attention mechanism is implemented using an algorithm to calculate the value of self-attention.
[0047] In step S115, the feature calculation unit 114 takes the self-attention values of each of the multiple elements as input and uses an inference model to calculate the features of each of the multiple elements. This creates inference data 182, which represents the features of each of the multiple elements.
[0048] Let's explain the inference model. An inference model is a machine learning model that calculates multiple features corresponding to multiple self-attention values shown in the input data, and then calculates an inference result based on these multiple features. These features are also called latent features. The inference result varies depending on the task and may show prediction results or classification results, among others. The inference model outputs inference data 182, which represents multiple features and the inference results.
[0049] Based on Figures 11 and 12, the learning process of the information processing method will be explained. In step S121, the causal relationship estimation unit 121 initializes the parameters of the causal relationship estimation model. Furthermore, the feature calculation unit 114 initializes the parameters of the inference model.
[0050] In step S122, the causal relationship estimation unit 121 acquires the training data 183. For example, a user inputs training data 183 into the information processing device 100, and the causal relationship estimation unit 111 receives the input training data 183.
[0051] Training data 183 is a sequence of data used for training.
[0052] The causal relationship estimation unit 111 then takes the sequence of multiple elements for training shown in the training data 183 as input and estimates the causal relationships between the multiple elements for training using a causal relationship model. This generates a causal relationship graph 192, which shows the estimated causal relationship.
[0053] In step S123, the chunk determination unit 122 determines chunks 194 for each element based on the estimated causal relationships. The determination method is the same as the method in step S113.
[0054] In step S124, the self-attention calculation unit 123 calculates the self-attention value for each of the multiple learning elements using the self-attention mechanism, with each chunk 194 of the multiple learning elements as the calculation range.
[0055] In step S125, the feature calculation unit 124 uses an inference model with the self-attention values of each of the multiple elements for training as input to calculate the features of each of the multiple elements for training. This creates the inference data 182.
[0056] In step S126, the learning control unit 125 determines whether the learning has converged based on the feature quantities of each of the multiple elements calculated in step S125.
[0057] The convergence of learning is determined as follows: First, the learning control unit 125 acquires the correct answer data 184. For example, a user inputs correct answer data 184 into the information processing device 100, and the learning control unit 125 receives the input correct answer data 184.
[0058] The 184 ground truth data points show the ground truth for multiple features and the ground truth for the inference results.
[0059] Next, the learning control unit 125 compares the inference data 182 with the correct answer data 184. Then, the learning control unit 125 determines whether the learning has converged based on the comparison results. For example, the learning control unit 125 calculates the feature difference. The feature difference indicates the magnitude of the difference between multiple features shown in the inference data 182 and multiple features shown in the ground truth data 184. The learning control unit 125 then compares the feature difference with a threshold and determines that the learning has converged if the feature difference is smaller than the threshold.
[0060] If the learning has converged, the process proceeds to step S127. If the learning has not converged, the learning control unit 125 updates the parameters of the causal relationship estimation model and the inference model based on the comparison results. The process then proceeds to step S122.
[0061] In step S127, the causal relationship estimation unit 121 generates sequence data that represents a sequence of multiple features corresponding to multiple elements for learning as a new sequence of multiple elements for learning. The causal relationship estimation unit 121 then uses the generated sequential data as input to estimate the causal relationships of multiple new elements for learning using a causal relationship estimation model.
[0062] In step S128, the learning control unit 125 determines whether the difference between the causal relationship estimated in step S127 and the causal relationship estimated in step S122 is large.
[0063] The magnitude of the difference between causal relationships is determined as follows: First, the learning control unit 125 updates the causal relationship graph 192 created in step S122 to the causal relationship graph 192 created in step S127 and calculates the graph update amount. The graph update amount indicates the amount of update from the causal relationship graph 192 created in step S122 to the causal relationship graph 192 created in step S127. The learning control unit 125 then compares the graph update amount with a threshold and determines that the difference between causal relationships is large if the graph update amount is greater than the threshold.
[0064] If the difference between causal relationships is large, the learning control unit 125 updates the parameters of the causal relationship estimation model and the inference model based on the comparison results. The process then proceeds to step S122. If the difference between the causal relationships is not large, the process proceeds to step S129.
[0065] In step S129, the causal relationship estimation unit 111 stores the parameters of the causal relationship estimation model in the parameter data 191. Furthermore, the feature calculation unit 114 stores the parameters of the inference model in the parameter data 191.
[0066] The learning process does not necessarily have to include steps S127 and S128. In this case, if the learning has converged in step S126, the process will proceed to step S129.
[0067] ***Effects of Embodiment 1*** Embodiment 1 aims to improve the accuracy of a predictive model by automatically determining the range of feature extraction based on causal relationships during the training and inference of a predictive model using an attention mechanism.
[0068] Embodiment 1 proposes a new method to simultaneously reduce the computational complexity of the attention mechanism and improve prediction accuracy. By combining causal relationship estimation, self-attention chunk determination, and self-attention value calculation, the system enables appropriate control of the calculation range, which is expected to improve the performance of the predictive model.
[0069] In models using attentional mechanisms, insufficient correlation between features can lead to decreased accuracy. Embodiment 1 improves the performance of the model during training and inference by estimating a causal graph from input data and appropriately saving and utilizing the trained parameters.
[0070] In Embodiment 1, the causal relationship estimation means and the latent feature calculation means are appropriately trained during the training phase, and during inference, these trained parameters are read from the DB to estimate causal relationships and calculate features. This improves the performance of the model.
[0071] Embodiment 1 enables the automatic determination of the feature extraction range based on causal relationships during model training and inference, thereby improving the model's prediction accuracy. Furthermore, the linear structure learning of the causal relationship estimation means and the latent feature calculation means ensures that the model's features are calculated appropriately, improving the performance of the predictive model.
[0072] ***Summary of Embodiment 1*** The information processing device 100 is a system that includes a causal relationship estimation unit 111, a chunk determination unit 112, a self-attention calculation unit 113, and a feature calculation unit 114. These components work together to automatically and effectively calculate features. The causal relationship estimation unit 111 estimates a causal graph from the input data. The chunk determination unit 112 identifies the range in which the direct causal relationships of the data can be elucidated, based on the obtained causal graph. The self-attention calculation unit 113 calculates the self-attention value from that range. The feature extraction unit 114 extracts latent features using machine learning. This integrated system automatically and effectively obtains useful information, including self-awareness, from input data, enabling the construction of advanced predictive models.
[0073] A convex optimization method is employed to determine the scope of the self-attention chunk, and the causal relationships of the input data are judged. When determining the chunk size, data with a direct causal relationship is selected, and the chunk size is determined based on that data. A constant is set to define the chunk length. If the chunk length is less than the constant, the causal graph is sorted in descending order of its path coefficients, and any missing nodes are filled with zero nodes. If the chunk length exceeds the constant, the causal graph is sorted in descending order of its path coefficients, and any excess nodes are not included in the chunk. By combining a convex optimization method with a chunking method that utilizes causal relationship data, the computational range can be appropriately controlled, resulting in efficient and effective calculation of self-attention.
[0074] ***Supplement to Embodiment 1*** Embodiment 1 realizes the following features and is expected to improve the performance of the prediction model. (1) Consideration of correlation and determination of direct causality To account for the correlation of features used in prediction, direct causal relationships between nodes in the input sequence data are determined. To calculate the features of self-attention, the calculation range (chunk) is automatically determined from the input data of length n (key, query). (2) Improvement of prediction accuracy The structural causal relationships between input sequence data (key, query) are estimated, and the chunk size is appropriately determined based on the parent-child relationships of these causal relationships. This allows the model to capture important correlations, leading to expectations of high prediction accuracy. (3) Reducing the computational complexity during inference By appropriately determining the computational scope (chunk) based on causal relationship estimations, the computational complexity during inference is significantly reduced. Specifically, the time complexity is O(n) 2The time complexity is reduced from ) to O(n·k), where "k" indicates the chunk size. This enables efficient inference, improves real-time performance, and allows for adaptation to large-scale data. These features allow for the construction of a predictive model that appropriately captures the characteristics of the input sequence data and takes into account the correlations between nodes, thereby improving prediction accuracy.
[0075] The hardware configuration of the information processing device 100 will be explained based on Figure 13. The information processing device 100 includes a processing circuit 109. The processing circuit 109 is hardware that implements the inference unit 110 and the learning unit 120. The processing circuit 109 may be dedicated hardware, or it may be a processor 101 that executes a program stored in memory 102.
[0076] If the processing circuit 109 is dedicated hardware, the processing circuit 109 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC, an FPGA, or a combination thereof. ASIC is an abbreviation for Application Specific Integrated Circuit. FPGA is an abbreviation for Field Programmable Gate Array.
[0077] The information processing device 100 may include multiple processing circuits that replace the processing circuit 109.
[0078] In the processing circuit 109, some functions may be implemented by dedicated hardware, while the remaining functions may be implemented by software or firmware.
[0079] Thus, the functions of the information processing device 100 can be realized through hardware, software, firmware, or a combination thereof.
[0080] Embodiment 1 is an example of a preferred form and is not intended to limit the technical scope of this disclosure. Embodiment 1 may be implemented in part or in combination with other forms. The procedure described using flowcharts, etc., may be modified as appropriate.
[0081] The "part" of each element of the information processing device 100 may be read as "processing," "process," "circuit," or "circuit."
[0082] The various aspects of this disclosure are described below as appendices. (Note 1) A causal relationship estimation unit takes sequence data representing a sequence of multiple elements as input and uses a causal relationship estimation model to estimate the causal relationships between the multiple elements, A chunk determination unit determines, for each element, the chunk that constitutes the calculation range of the self-attention mechanism based on the estimated causal relationship, A self-attention calculation unit that calculates the self-attention value for each of the multiple elements by the self-attention mechanism, with each of the chunks of the multiple elements set to the calculation range, A feature calculation unit that takes the aforementioned self-attention values for each of the aforementioned elements as input and uses an inference model to calculate the feature quantities for each of the aforementioned elements, An information processing device equipped with the following features.
[0083] (Note 2) The aforementioned causal relationship estimation model is a model that estimates causal relationships using a convex optimal method. The information processing device described in Appendix 1.
[0084] (Note 3) The causal relationship estimation unit creates a causal relationship graph as data showing the estimated causal relationships, by making each of the multiple elements a node and connecting the nodes that have a direct causal relationship with paths. The chunk determination unit selects, for each element, a group of parent-child nodes from the causal relationship graph, consisting of a node representing the element and nodes that have a parent-child relationship with the element's node, and determines the chunk for each element based on the group of parent-child nodes. The information processing device described in Appendix 1 or Appendix 2.
[0085] (Note 4) The chunk determination unit is an information processing device according to Appendix 3, which determines a group of elements corresponding to the parent-child node group into the chunk.
[0086] (Note 5) The chunk determination unit, If the number of nodes in the parent-child node group is the same as a constant, the group of elements corresponding to the parent-child node group is determined to be the chunk. If the number of nodes in the parent-child node group is less than the constant, the group of elements corresponding to the parent-child node group and one or more dummy elements are determined to make up the chunk. If the number of nodes in the parent-child node group exceeds the constant, the same number of elements as the constant are selected from the element group corresponding to the parent-child node group to form the chunk. The information processing device described in Appendix 3.
[0087] (Note 6) The causal relationship graph includes, for each path, a path coefficient indicating the magnitude of the causal relationship between the nodes connected by that path. The chunk determination unit, when the number of nodes in the parent-child node group exceeds the constant, selects the same number of elements as the constant from the element group corresponding to the parent-child node group as the chunk, based on the path coefficient group for the parent-child node group. The information processing device described in Appendix 5.
[0088] (Note 7) Using a causal relationship estimation model, the causal relationships between the multiple elements are estimated by taking sequential data representing a sequence of multiple elements as input. Based on the estimated causal relationships, for each element, the chunk that constitutes the calculation range of the self-attention mechanism is determined. The self-attention mechanism calculates the self-attention value for each of the multiple elements by setting each of the aforementioned chunks to the calculation range. Using the aforementioned self-attention values for each of the aforementioned elements as input, an inference model is used to calculate the feature quantities for each of the aforementioned elements. Information processing methods.
[0089] (Note 8) A causal relationship estimation process that takes sequence data representing a sequence of multiple elements as input and uses a causal relationship estimation model to estimate the causal relationships between the multiple elements, Based on the estimated causal relationships, a chunk determination process is performed to determine the chunks that constitute the calculation range of the self-attention mechanism for each element, A self-attention calculation process that uses the self-attention mechanism to calculate the self-attention value for each of the multiple elements, with each of the chunks of the multiple elements set as the calculation range, A feature calculation process that uses an inference model to calculate the feature quantities of each of the aforementioned elements, taking the aforementioned self-attention values of each of the aforementioned elements as input. An information processing program that causes a computer to execute something. [Explanation of Symbols]
[0090] 100 Information processing device, 101 Processor, 102 Memory, 103 Auxiliary storage device, 104 Communication device, 105 Input / Output interface, 109 Processing circuit, 110 Inference unit, 111 Causal relationship estimation unit, 112 Chunk determination unit, 113 Self-attention calculation unit, 114 Feature calculation unit, 120 Learning unit, 121 Causal relationship estimation unit, 122 Chunk determination unit, 123 Self-attention calculation unit, 124 Feature calculation unit, 181 Sequence data, 182 Inference data, 183 Training data, 184 Ground truth data, 190 Storage unit, 191 Parameter data, 192 Causal relationship graph, 193 Matrix, 194 Chunk, 195 Parent-child node group.
Claims
1. A causal relationship estimation unit takes sequence data representing a sequence of multiple elements as input and uses a causal relationship estimation model to estimate the causal relationships between the multiple elements, A chunk determination unit determines, for each element, the chunk that constitutes the calculation range of the self-attention mechanism based on the estimated causal relationship, A self-attention calculation unit that calculates the self-attention value for each of the multiple elements by the self-attention mechanism, with each of the chunks of the multiple elements set to the calculation range, A feature calculation unit that takes the aforementioned self-attention values for each of the aforementioned elements as input and uses an inference model to calculate the feature quantities for each of the aforementioned elements, An information processing device equipped with the following features.
2. The aforementioned causal relationship estimation model is a model that estimates causal relationships using a convex optimal method. The information processing apparatus according to claim 1.
3. The causal relationship estimation unit creates a causal relationship graph as data showing the estimated causal relationships, by making each of the multiple elements a node and connecting the nodes that have a direct causal relationship with paths. The chunk determination unit selects, for each element, a group of parent-child nodes from the causal relationship graph, consisting of a node representing the element and nodes that have a parent-child relationship with the element's node, and determines the chunk for each element based on the group of parent-child nodes. The information processing apparatus according to claim 1 or claim 2.
4. The information processing apparatus according to claim 3, wherein the chunk determination unit determines the group of elements corresponding to the parent-child node group as the chunk.
5. The chunk determination unit, If the number of nodes in the parent-child node group is the same as a constant, the group of elements corresponding to the parent-child node group is determined to be the chunk. If the number of nodes in the parent-child node group is less than the constant, the element group corresponding to the parent-child node group and one or more dummy elements are determined to form the chunk. If the number of nodes in the parent-child node group exceeds the constant, the same number of elements as the constant are selected from the element group corresponding to the parent-child node group to form the chunk. The information processing apparatus according to claim 3.
6. The causal relationship graph includes, for each path, a path coefficient indicating the magnitude of the causal relationship between the nodes connected by that path. The chunk determination unit, when the number of nodes in the parent-child node group exceeds the constant, selects the same number of elements as the constant from the element group corresponding to the parent-child node group as the chunk, based on the path coefficient group for the parent-child node group. The information processing apparatus according to claim 5.
7. The information processing device is Using a causal relationship estimation model, the causal relationships between the multiple elements are estimated by taking sequential data representing a sequence of multiple elements as input. Based on the estimated causal relationships, for each element, the chunk that constitutes the calculation range of the self-attention mechanism is determined. The self-attention mechanism calculates the self-attention value for each of the multiple elements by setting each of the aforementioned chunks to the calculation range. Using the aforementioned self-attention values for each of the aforementioned elements as input, an inference model is used to calculate the feature quantities for each of the aforementioned elements. Information processing methods.
8. A causal relationship estimation process that takes sequence data representing a sequence of multiple elements as input and uses a causal relationship estimation model to estimate the causal relationships between the multiple elements, Based on the estimated causal relationships, a chunk determination process is performed to determine the chunks that constitute the calculation range of the self-attention mechanism for each element, A self-attention calculation process that uses the self-attention mechanism to calculate the self-attention value for each of the multiple elements, with each of the chunks of the multiple elements set as the calculation range, A feature calculation process that uses an inference model to calculate the feature quantities of each of the aforementioned elements, taking the aforementioned self-attention values of each of the aforementioned elements as input. An information processing program that causes a computer to execute something.
Citation Information
Patent Citations
Industrial process space-time causal directed graph modeling method based on reinforcement learning
CN116880381A
Method, device and system for estimating causal relation between observation variables
JP2019207685A
Monitoring system and monitoring method
JP2020052714A
Estimation device, estimation method, and estimation program
JP2023013810A
Pre-trained model acquisition method, apparatus, electronic device, storage medium, and computer program
JP7379792B2