Method and apparatus for training artificial neural network

The artificial neural network learning method addresses the challenge of predicting arbitrary missing values in GNNs by preprocessing data, applying missing values, and using a graph autoencoder, resulting in improved prediction accuracy and robustness for industrial applications.

WO2025127531A1PCT designated stage expired Publication Date: 2025-06-19POSCO HLDG INC

Patent Information

Application Number
PCT/KR2024/019106
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-11-28
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Graph Neural Networks (GNNs) face challenges in predicting arbitrary missing values, which can lead to errors in correlation analysis and prediction tasks, particularly in industrial applications where system data may contain missing values.

Method used

An artificial neural network learning method and device that preprocesses node and edge data to generate embedded data, applies arbitrary missing values to create learning data, and uses a graph autoencoder to learn predictions for these missing values, with an evaluation unit terminating learning when prediction accuracy reaches a preset range.

Benefits of technology

This approach enhances the accuracy of prediction and correlation analysis in GNNs by effectively handling missing values, thereby expanding the usability of GNNs in industrial sites and improving their robustness in real-world applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019106_19062025_PF_FP_ABST
    Figure KR2024019106_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a technique for training an artificial neural network, and provides an apparatus and a method for training an artificial neural network, the apparatus comprising: a preprocessing unit for generating multiple pieces of embedding data by using data of multiple nodes and edges generated according to a state of system data converted into a graph structure, and generating training data by applying a random missing value to each of the multiple pieces of embedding data; a training unit for inputting training data into a graph auto-encoder, and training prediction for a missing value by using output data of the graph auto-encoder; and an evaluation unit for terminating training when the accuracy of prediction for the missing value is included in a preconfigured range.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial neural network learning method and device thereof

[0001] The present disclosure relates to techniques for learning artificial neural networks. In particular, the present disclosure relates to techniques for learning to predict arbitrary missing values ​​in a graph neural network.

[0002] With the advancement of artificial intelligence and deep learning technologies, various industries are seeking to incorporate artificial intelligence into their industrial settings.

[0003] Artificial intelligence technology encompasses a variety of neural network structures and technologies. Depending on the intended use, various neural network architectures, such as CNNs and RNNs, are utilized. For example, in image processing, CNNs, which extract features based on Euclidean space, are widely used.

[0004] Recently, efforts are being made to utilize artificial intelligence technology by considering the interrelationship between elements by incorporating GNN (Graph Neural Network) technology in the case of graph structure systems where data is formed through interrelationships between various elements.

[0005] GNN is a neural network technology that can be effectively applied when data is represented as a graph, a structure connected by numerous nodes and edges. While significant improvements are needed in terms of technological maturity, it is highly regarded for its potential.

[0006] Research is ongoing to predict and analyze user connectivity in social networks, predict molecular connectivity, predict molecular properties based on element type, analyze connections between cited documents, and analyze associations and correlations in knowledge graphs using recently developed GNNs.

[0007] However, in the case of GNN, although the vector value for each node is calculated including the node's state and edge information, if there is some loss in the vector value itself for each node, errors may occur in the correlation analysis and prediction described above.

[0008] Therefore, in order to improve the accuracy of prediction and correlation analysis using GNN, techniques for preventing and recovering the loss of node vector values ​​are required.

[0009] The present disclosure provides techniques for learning artificial neural networks. In particular, the disclosure provides techniques for learning to predict arbitrary missing values ​​in graph neural networks.

[0010] In one aspect, the present embodiments provide an artificial neural network learning device for learning an artificial neural network, comprising: a preprocessing unit for generating a plurality of embedding data using a plurality of node and edge data generated according to the state of system data converted into a graph structure, and for generating learning data by applying an arbitrary missing value to each of the plurality of embedding data; a learning unit for inputting the learning data into a graph autoencoder and learning a prediction for the missing value using output data of the graph autoencoder; and an evaluation unit for terminating learning when the prediction accuracy for the missing value falls within a preset range.

[0011] In another aspect, the present embodiments provide an artificial neural network learning method for learning an artificial neural network, comprising a preprocessing step of generating a plurality of embedding data by using a plurality of node and edge data generated according to the state of system data converted into a graph structure, and generating learning data by applying an arbitrary missing value to each of the plurality of embedding data, a learning step of inputting the learning data to a graph autoencoder and learning a prediction for the missing value by using output data of the graph autoencoder, and an evaluation step of terminating learning when the prediction accuracy for the missing value falls within a preset range.

[0012] According to the present disclosure, a technique for learning an artificial neural network can be provided. In particular, a technique for learning to predict arbitrary missing values ​​in a graph neural network can be provided.

[0013] FIG. 1 is a diagram illustrating the configuration of an artificial neural network learning device according to one embodiment.

[0014] Figure 2 is a diagram for explaining the conversion into a graph structure according to one embodiment.

[0015] FIG. 3 is a diagram illustrating an operation for constructing a data set including nodes and edges according to one embodiment.

[0016] FIG. 4 is a diagram for explaining an operation of generating embedded data according to one embodiment.

[0017] FIG. 5 is a drawing for explaining an example of applying a random loss value according to one embodiment.

[0018] FIG. 6 is a drawing for explaining another example of applying a random loss value according to one embodiment.

[0019] Fig. 7 is a diagram for explaining the configuration of an autoencoder according to one embodiment.

[0020] Fig. 8 is a drawing for explaining the overall operating structure of an artificial neural network learning device according to one embodiment.

[0021] Figure 9 is a diagram for explaining an artificial neural network learning method according to one embodiment.

[0022] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to exemplary drawings. When adding reference numerals to components in each drawing, identical components may have the same numerals as much as possible even if they are shown in different drawings. In addition, when describing the present embodiments, if it is determined that a detailed description of a related known configuration or function may obscure the gist of the technical idea of ​​the present invention, the detailed description may be omitted. When "includes," "has," "consists of," etc. are used in this specification, other parts may be added unless "only" is used. When a component is expressed in the singular, it may include a case in which the plural is included unless specifically stated otherwise.

[0023] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, sequence, or number of the components are not limited by the terms.

[0024] In a description of the positional relationship of components, when it is described that two or more components are "connected," "combined," or "connected," it should be understood that the two or more components may be directly "connected," "combined," or "connected," but that the two or more components may also be further "interposed" with another component to be "connected," "combined," or "connected." Here, the other component may be included in one or more of the two or more components that are "connected," "combined," or "connected" to each other.

[0025] In the description of the temporal flow relationship related to components, operation methods, or manufacturing methods, for example, when the temporal or flow relationship is described as “after”, “following”, “next to”, “before”, etc., it may also include cases where it is not continuous, unless “immediately” or “directly” is used.

[0026] Meanwhile, when numerical values ​​or corresponding information (e.g., levels, etc.) for components are mentioned, even without separate explicit description, the numerical values ​​or corresponding information may be interpreted as including an error range that may occur due to various factors (e.g., process factors, internal or external impact, noise, etc.).

[0027] Below, the devices and methods according to the present disclosure are described in more detail with reference to the drawings. Each algorithm described below is provided as an example, and various algorithms that perform the same purpose and function can be applied to the present disclosure. Furthermore, programs capable of performing the functions of the present embodiments are also included in the present disclosure, and recording media containing the programs are also construed as being included in the present disclosure.

[0028] FIG. 1 is a diagram illustrating the configuration of an artificial neural network learning device according to one embodiment.

[0029] Referring to FIG. 1, an artificial neural network learning device (100) for learning an artificial neural network may include a preprocessing unit (110) that generates a plurality of embedding data by using a plurality of nodes and edge data generated according to the state of system data converted into a graph structure, and generates learning data by applying an arbitrary missing value to each of the plurality of embedding data.

[0030] For example, the preprocessing unit (110) can generate a graph structure in which system data is converted into a graph structure including node and edge data, and edges include relationship information between nodes. For example, the preprocessing unit (110) can determine whether system data such as process data and molecular structure data can be converted into a graph structure. If the system data can be converted into a graph structure, the preprocessing unit (110) converts the system data into a graph structure including nodes and edges. For example, a node represents an element configuration including data information, and an edge can include connection or correlation information between nodes.

[0031] Graph structures are used to process data using GNNs and can be created by transforming social networks, text, images, and process data.

[0032] In addition, the preprocessing unit (110) can generate a plurality of node and edge data by generating corresponding node and edge data according to each state of the system data. For example, the preprocessing unit (110) can convert the system data into various nodes and edges according to its state. In particular, even if the system structure is the same, it can be displayed as various graphs depending on changes in the relationship between nodes or changes in the state of nodes. In order to learn an artificial neural network, a certain level or more of learning data is required, and in order to generate this, the preprocessing unit (110) can generate each node and edge data in response to changes in the state of the system data. Through this, the preprocessing unit (110) can configure a data set for the node and edge data.

[0033] Additionally, the preprocessing unit (110) may generate embedding data for each node using a preset embedding algorithm for each of the plurality of nodes and edge data. For example, the preprocessing unit (110) may generate node and edge data and apply a preset embedding algorithm to each node and edge data to generate embedding data for the node. The embedding data may be in vector format, but is not limited thereto. Furthermore, the embedding data may include one or more tokens.

[0034] Embedding data can be generated for each node. Furthermore, embedding data can be generated for each node and edge data set, as described above. Embedding data can be generated separately for each node and edge data set.

[0035] Additionally, the preprocessing unit (110) can generate training data by modifying the embedded data. In particular, training data can be generated by modifying specific tokens in the embedded data to perform predictive learning for arbitrary missing values. Various methods can be used to apply arbitrary missing values ​​to the embedded data.

[0036] For example, the preprocessing unit (110) can apply an arbitrary missing value by replacing some tokens randomly selected from among the tokens constituting the embedding data with preset mask tokens. For example, the preprocessing unit (110) can select a specific token from among the tokens constituting the embedding data through a random argument. The selected token can be replaced with a preset mask token. For example, the mask token can be set to various values ​​such as -999, "blank", "Null", "[Null]", etc., and can be set to a specific value according to user input. The portion replaced with the mask token becomes an arbitrary missing value in the embedding data.

[0037] As another example, the preprocessing unit (110) can apply the above random missing value by replacing some tokens sampled according to the span length among the tokens constituting the embedding data with preset mask tokens. For example, the preprocessing unit (110) can replace tokens of the embedding data extracted according to the span length with preset mask tokens. The span length can be extracted from a preset Poisson distribution. For example, if the Poisson distribution is set to 3, the probability that the span length will be extracted within 3 to 6 appears high. The preprocessing unit (110) can replace consecutive tokens corresponding to the span length with a preset mask token. As described above, the mask token can be set in various ways. The span length can be extracted as 0, in which case the mask token can be added between arbitrary tokens.

[0038] The preprocessing unit (110) can apply arbitrary missing values ​​to the embedded data to set some tokens as training data, with some tokens masked. A single embedded data may be generated into multiple training data sets depending on each masking process. Alternatively, a single embedded data set may be subject to only one masking process, resulting in only one corresponding training data set.

[0039] An artificial neural network learning device (100) may include a learning unit (120) that inputs learning data into a graph autoencoder and learns predictions for missing values ​​using output data of the graph autoencoder.

[0040] When the learning data is prepared, the learning unit (120) can set the learning data as an input value of the graph autoencoder and perform a learning operation to predict a missing value masked in the embedded data using the output value.

[0041] For example, a graph autoencoder may be configured as a graph convolutional neural network (GCN) in the form of an autoencoder, where the input and output are set to the same dimension. Furthermore, the output data may be a graph created by the preprocessing unit (110) as a snapshot of a specific state value of the system data, i.e., a data set of nodes and edges.

[0042] The learning unit (120) can perform a learning operation through supervised learning or unsupervised learning on the learning data and output data. For example, the learning unit (120) can perform supervised learning on a graph autoencoder so that the output data appears as the same value as the embedding data corresponding to the input learning data. Alternatively, when the learning data is output as output data of the graph autoencoder, the learning unit (120) can perform a learning operation through supervised learning to check whether the masked tokens in the learning data are output so as to match in the output data.

[0043] The artificial neural network learning device (100) may include an evaluation unit (130) that terminates learning when the prediction accuracy for the missing value falls within a preset range.

[0044] For example, the evaluation unit (130) measures the prediction accuracy for the missing gap, and terminates learning when the prediction accuracy falls within a preset range. The prediction accuracy can be preset by the user.

[0045] Alternatively, the evaluation unit (130) can evaluate the prediction accuracy with a validation data set while increasing the epoch and stop learning when the prediction accuracy exceeds a preset standard. This can improve accuracy while preventing problems caused by overfitting. To this end, the data set including node and edge data can be separated and stored as a learning data set, a validation data set, and a test data set. The learning data set refers to a data set used for learning, and the validation data set refers to a data set that is not used for learning but is set to be used for evaluation, such as when learning is completed. The test set refers to a separate data set used for testing after learning is completed.

[0046] Through the aforementioned operations, the artificial neural network learning device (100) can expand the usability of graph artificial neural networks, which have traditionally been utilized primarily for information about nodes and edges, such as prediction of node states, prediction of edges, graph classification, and node value clustering. Furthermore, it can improve accuracy and support the use of GNN even when some data has problems. That is, the artificial neural network learning device (100) can train the artificial neural network to predict missing values ​​with high accuracy even when some missing values ​​exist in the embedded data. Through this, it is possible to prepare for various problem situations that may arise in actual industrial sites.

[0047] Below, each operation of the aforementioned artificial neural network learning device is described in more detail with reference to the drawings. Each embodiment described below is provided for convenience of explanation and is not limited to the embodiments described below.

[0048] This embodiment describes a learning system for a device that performs simulations to predict arbitrary missing values ​​in system data, such as a process system, using a graph neural network. To achieve this, the present embodiment applies masking and text infilling to graph neural network learning, enabling prediction of arbitrary missing values, something previously impossible with graph neural networks.

[0049] In recent years, the application scope of graph neural networks (GNNs) has expanded, making them suitable for a variety of tasks. A graph neural network is a type of neural network implementation that takes as input a data structure called a graph, which consists of nodes and edges connecting them. Unlike conventional neural networks, a graph neural network is an algorithm that assumes that the information of each individual is influenced not only by its own state but also by the state and information values ​​of surrounding individuals.

[0050] Since the ultimate goal of graph neural networks is to learn the state embedding of each node (a value that represents information about each node and its surrounding nodes), it has the advantage of being able to learn not only the objects themselves but also the relationships between objects.

[0051] Recently developed graph-related tasks include node value classification, node value prediction, node value clustering, edge value classification, connection relationship prediction, and graph classification. These can be used to predict and analyze user connectivity in social networks, predict molecular connectivity, predict molecular properties based on element type, analyze connections between cited documents, and analyze associations and correlations in knowledge graphs.

[0052] However, development of tasks that can address cases where graphs contain arbitrary missing values ​​is not underway. When applying GNNs to industrial settings, diverse system data may exist, and missing values ​​may occur at specific nodes. Therefore, the artificial neural network learning technology of the present disclosure, which predicts missing values, is required to improve accuracy and enhance usability.

[0053] To achieve this, this disclosure utilizes an artificial intelligence model called a graph autoencoder to perform missing value prediction, a feat previously unachievable in graph neural networks. Furthermore, it provides a technique for generating learning data and performing learning using masking or text infilling techniques.

[0054] A graph is a data structure composed of a set of nodes and the edges connecting them. It is generally defined as G=(V,E), where V is the set of nodes constituting the graph, and E is the set of edges connecting two nodes. Edges can be structured in various ways, depending on whether they are directed or weighted.

[0055] Graphs are well-suited for handling abstract concepts like relationships and interactions. In particular, they offer the advantage of simplifying complex problems into simple representations. Furthermore, graphs enable representation and learning in non-Euclidean spaces. While images, text, and structured data commonly used in AI can be represented in grid formats, data like social media data and molecular data cannot be defined as structured data. Therefore, unstructured data must be represented in a non-Euclidean space, and in this case, graphs can be used for processing.

[0056] Figure 2 is a diagram for explaining the conversion into a graph structure according to one embodiment.

[0057] Referring to FIG. 2, the preprocessing unit can transform system data into a graph structure including node and edge data. Here, the graph structure can be created so that edges include relationship information between nodes.

[0058] For example, in the case of molecular structure (250), it cannot be expressed as formal data. Therefore, when data on molecular structure is input as system data, it is converted into a graph structure (290). A graph structure can be created by converting the atoms constituting each molecule into nodes (200) and configuring the connection relationships between atoms as edges (210).

[0059] Similarly, in a process system where each process is connected to form the entire process system, the process system can be converted into a graph structure with nodes and edges. The preprocessing unit can perform embedding using the node and edge data converted into a graph structure to generate embedded data.

[0060] FIG. 3 is a diagram illustrating an operation for constructing a data set including nodes and edges according to one embodiment.

[0061] Referring to Figure 3, the preprocessing unit can generate multiple node and edge data sets by generating corresponding node and edge data for each state of the system data. To generate a large number of training data sets and perform efficient training using information about various states, the preprocessing unit can generate node and edge data sets.

[0062] For example, when a graph structure is transformed, such as 300, embedded data can be generated using the corresponding graph structure. In addition, the preprocessing unit can generate node and edge data corresponding to each system data state by reflecting the state changes of various system data. For example, in a specific state, such as 310, a connection relationship may exist between nodes 2 and 3. Or, in a specific state, such as 320, a connection relationship may not exist between nodes 3 and 4. Or, as in 330, the state data of node 1 may be different from 300, and in this case, different state data, such as 1', may be used.

[0063] In this way, the preprocessing unit can construct a node and edge data set by generating various node and edge data according to each state of the system data. The preprocessing unit can then apply arbitrary missing values ​​to create a data subset, which can then be used as training data.

[0064] FIG. 4 is a diagram for explaining an operation of generating embedded data according to one embodiment.

[0065] Referring to Fig. 4, the preprocessing unit can convert a graph structure (400) including node and edge data into embedded data (410, 420) using an embedding algorithm. Each node of the graph structure (400) can be expressed as v1 to v6. The edge connecting each node is expressed as e, and the connected node number is assigned to e. 12 , e 26 It can be expressed as follows.

[0066] If the graph is expressed as a formula, it can be expressed as a set of G={V,E}, and accordingly, G=((v1, ..., v6), (e 12, ..., e 56)) is expressed as a matrix. In order to express the graph structure, an adjacency matrix or a degree matrix can be used to express it as a matrix.

[0067] For example, G can be expressed as an adjacency matrix A. The rows and columns each represent nodes, and it can be expressed as a matrix A that has a value of 1 if there is a connection between nodes and 0 if there is not.

[0068] The preprocessing unit can represent the matrix generated using node and edge data as vector values ​​using a preset embedding algorithm. For example, v1 can be vectorized as [123.11,1,92, "a",...,0] and generated as embedded data (410). Similarly, v2 through v6 can be converted into embedded data. For example, v6 can be vectorized as [532.06,26,1, "e",...,240] and generated as embedded data (420).

[0069] There are no restrictions on embedding algorithms, and any algorithm that can transform a graph into embedding data can be applied without any limitations. For example, either the transductive method or the inductive method can be used. The transductive method obtains an embedding vector for each node, and the node's embedding position can be obtained as a result of learning. The inductive method obtains the function (i.e., the encoder) that converts each node into an embedding. GNNs are a representative example of the inductive method. Thus, there are no restrictions on how embedding data can be obtained.

[0070] FIG. 5 is a drawing for explaining an example of applying a random loss value according to one embodiment.

[0071] Referring to Figure 5, the preprocessing unit can convert embedded data into training data. For example, the preprocessing unit can generate a training data set by applying arbitrary missing values ​​to the embedded data. Various methods can be used to apply arbitrary missing values ​​to the embedded data.

[0072] For example, the preprocessing unit can apply an arbitrary missing value by replacing some randomly selected tokens among the tokens that constitute the embedding data with preset mask tokens.

[0073] For example, the preprocessing unit can generate training data by inserting arbitrary missing values ​​into the embedding data v1 to v6 for the graph structure. Taking v1 as an example, the preprocessing unit can randomly select tokens 1 and 92 from among the tokens that compose v1. Once a token is selected, the preprocessing unit can replace the token with a preset mask token to generate training data containing the mask token, as in 510. That is, the 1 and 92 tokens can be replaced with a mask token replaced with [Blank].

[0074] Alternatively, the preprocessing unit can select only one token when selecting random tokens. If embedding data such as 420 is generated, the preprocessing unit can randomly select a single token. If token 240 is selected, the preprocessing unit can replace token 240 with a [Blank] mask token to generate training data such as 520.

[0075] Mask tokens can be applied in various ways depending on the settings. For example, mask tokens can be set to various values ​​such as -999, "Blank", "Null", and "[Null]".

[0076] In addition, the preprocessing unit can generate training data by applying various algorithms that can modify portions of the embedding data. For example, token masking, as described above, can be used. This involves extracting random tokens and replacing them with the [MASK] token. Alternatively, text infiltrating can be used. Furthermore, various natural language processing masking techniques, such as token deletion, sentence permutation, and document rotation, can be applied.

[0077] FIG. 6 is a drawing for explaining another example of applying a random loss value according to one embodiment.

[0078] Referring to Figure 6, the preprocessing unit can apply a random missing value by replacing some tokens sampled according to the span length among the tokens constituting the embedding data with preset mask tokens. For example, the span length can be extracted from a preset Poisson distribution, and the number of some tokens sampled according to the span length can be determined.

[0079] Unlike the method of replacing randomly extracted random tokens in Fig. 5 with mask tokens, the preprocessing unit can also use the text infilling method to generate learning data.

[0080] For example, a preprocessor could sample text spans of a length equal to the span length extracted from a Poisson distribution with lambda = 3 and replace them with a single [MASK] token. In other words, a span of a given length can be extracted and replaced with a single mask token.

[0081] When embedding data for each node is generated, such as 410 and 420, a text span of "1,92" can be sampled from a span with a length of 2. The sampled "1,92" can be replaced with a single token, [Blank], which is a mask token. Through this, training data (610) for v1 can be generated.

[0082] Additionally, the span length can be set to 0. In a case such as 420, if a span with a length of 0 is applied, a text span called 'empty' with a length of 0 can be sampled. The sampled 'empty' can be replaced with a single token, [Blank], which is a mask token. Through this, training data (620) for v6 can be generated.

[0083] A span can be thought of as a token in a text. When this span is set to follow a Poisson distribution with lambda = 3, a span length between 0 and 6 is likely to be selected.

[0084] Training data can be generated using the exemplary methods described above. In addition, the preprocessing unit can generate training data using various methods to introduce noise into the embedded data. Performance differences in learning results may occur depending on the noise application method. In this case, the preprocessing unit can generate training data by applying a preset, appropriate noise application method based on the type of embedded data or system data.

[0085] Fig. 7 is a diagram for explaining the configuration of an autoencoder according to one embodiment.

[0086] Referring to Figure 7, the learning unit performs learning by inputting learning data into a graph neural network in the form of an autoencoder, where the input and output dimensions are identical. For example, the autoencoder may be a denoising encoder.

[0087] Autoencoders are configured to have identical input and output dimensions. The hidden layer can be configured to reduce the dimensionality of the input data and then expand it to produce output data. Alternatively, the hidden layer can be configured to expand the dimensionality of the input data and then reduce it to produce output data.

[0088] There are various types of autoencoders, and the basic structure is shown in Figure 7. Input vector: x∈[0,1] D , and hidden representation: y∈[0,1] d is the following deterministic mapping: y=f θ (x)=s(Wx+b), where θ=W,b is the parameter, W is the weight matrix of dx D, and b is the bias vector. The latent representation calculated in the hidden layer: y value is the "Reconstructed" vector: x′∈[0,1] D It is re-mapped back to . The formula calculated at this time is , and θ′=W′,b′. At this time, the transpose matrix W of the weight matrix of the input layer before mapping T If W′ is the remapping weight matrix, then the autoencoder is said to have tied weights.

[0089] In addition to these, there are various types of autoencoders, such as stack Otter encoder, sparse Otter encoder, and denoising autoencoder.

[0090] In particular, denoising autoencoders perform learning by adding noise to input data and comparing output data derived from data learned based on the added noise with input data before noise insertion. In other words, learning can be performed by checking whether z1 is identical to x'1 before noise is inserted into x1.

[0091] The learning unit inputs training data into an autoencoder composed of graph convolutional networks, produces output data, and then compares the output data with the embedding data (pre-masking data) to perform learning. The learning unit can learn to predict arbitrary missing values ​​through supervised learning.

[0092] Graph convolutional networks are a type of spatial method. Like CNNs, they are convolutional networks that obtain information from neighboring nodes around a target vertex, aggregate it, and update it with information about the target node. As the term suggests, GCNs are models that learn the inherent information of a graph by performing convolution operations using kernels or filters, just like CNNs do. However, GCNs are not fully spatial GCNs. To be precise, GCNs can be seen as a method that bridges the spectral method to the spatial method. This is because they utilize convolution operations from adjacent nodes by taking advantage of the property that convolution operations in the spatial domain are equivalent to products in the Fourier domain.

[0093] From a spatial methodology perspective, GCN performs neighborhood aggregation, which means updating its own information using information from the neighboring nodes surrounding a target node. The amount of information to be transmitted from each neighboring node to the target node is called a message. Thus, graph convolution can be viewed as a weighted average or simple sum of the information from the neighboring nodes connected to the target node.

[0094] The goal of GCN is to find the optimal filter W that can create a hidden representation of the target's information using information from the target's neighboring nodes. This W filter is a weight and a learnable parameter that can be shared when using the hidden representation of other nodes.

[0095] Graph autoencoders can be constructed using these GCN properties.

[0096] Figure 8 is a diagram illustrating the overall operational structure of an artificial neural network learning device according to one embodiment. Referring to Figure 8, the learning operation described above will be described again in its entirety.

[0097] If system data can be converted into a graph structure, an artificial neural network learning device can use the system data to generate a graph containing nodes and edges in the form G=(V,E). Assuming it is expressed as an adjacency matrix, G can be expressed as A. In other words, the artificial neural network learning device defines the system as a graph and converts it by defining what the nodes and edges will be.

[0098] For example, in the case of a process system, a process graph is generated based on the structure and form of the process data, and this is embedded using an embedding algorithm. Once the embedded data is generated, the artificial neural network learning device sorts the data according to the form of the embedded graph and creates a training data set with portions of the embedded data masked.

[0099] The artificial neural network learning device performs learning by inputting learning data arranged in a graph format and with masking applied as input values ​​to a graph autoencoder (800). Learning is performed by judging the accuracy of the output data of the graph autoencoder (800) with the embedded data that embeds the graph, and the learning process is repeated until the accuracy of the masked portion reaches a certain level or higher.

[0100] Additionally, the artificial neural network learning device can increase the amount of the learning data set by repeatedly performing additional masking operations when learning data is insufficient.

[0101] If necessary, a process of retraining the simulation model using additional data, such as transfer learning, may be performed.

[0102] The aforementioned operations enable prediction of arbitrary missing values ​​in systems such as process systems. In particular, the present embodiment applies masking or text infilling techniques to graph neural network learning, providing a task previously impossible to predict arbitrary missing values ​​in graph neural networks.

[0103] Below, the artificial neural network learning device described above is again described from a methodological perspective, with reference to the drawings. Each step below is exemplary; the steps can be split or combined as needed. Furthermore, additional steps, such as transfer learning and additional training data generation, may be applied as needed.

[0104] Figure 9 is a diagram for explaining an artificial neural network learning method according to one embodiment.

[0105] Referring to FIG. 9, an artificial neural network learning method for learning an artificial neural network may include a preprocessing step of generating a plurality of embedding data by using a plurality of node and edge data generated according to the state of system data converted into a graph structure, and generating learning data by applying an arbitrary missing value to each of the plurality of embedding data (S900).

[0106] For example, the preprocessing step can create a graph structure in which system data is converted into a graph structure containing node and edge data, and edges contain relationship information between nodes. For example, the preprocessing step can determine whether system data such as process data or molecular structure data can be converted into a graph structure. If the system data can be converted into a graph structure, the preprocessing step converts the system data into a graph structure containing nodes and edges. The graph structure is for processing data using GNN, and can be created by converting social network data, text, images, process data, etc.

[0107] Additionally, the preprocessing step can generate multiple node and edge data by generating corresponding node and edge data according to each state of the system data. For example, the preprocessing step can transform system data into various nodes and edges depending on its state. In particular, even within the same system structure, various graphs can be displayed depending on changes in the relationship between nodes or changes in the state of nodes. A certain level of training data is required for artificial neural network training, and to generate this, the preprocessing step can generate each node and edge data in response to changes in the state of the system data. Through this, the preprocessing step can construct a data set for node and edge data.

[0108] Additionally, the preprocessing step can generate embedding data for each node using a preset embedding algorithm for each node and edge data. For example, the preprocessing step can generate node and edge data and apply a preset embedding algorithm to each node and edge data to generate embedding data for the node. The embedding data may be in vector format, but is not limited thereto. Furthermore, the embedding data may include one or more tokens.

[0109] Embedding data can be generated for each node. Furthermore, embedding data can be generated for each node and edge data set, as described above. Embedding data can be generated separately for each node and edge data set.

[0110] Additionally, the preprocessing step can generate training data by transforming the embedding data. Specifically, to facilitate prediction learning for arbitrary missing values, specific tokens in the embedding data can be transformed to generate training data. Various methods can be used to apply arbitrary missing values ​​to the embedding data.

[0111] For example, the preprocessing step can apply an arbitrary missing value by replacing some randomly selected tokens among the tokens constituting the embedding data with preset mask tokens. For example, the preprocessing step can select a specific token among the tokens constituting the embedding data through a random argument. The selected token can be replaced with a preset mask token. For example, the mask token can be set to various values ​​such as -999, "blank", "Null", "[Null]", etc., and can be set to a specific value based on user input. The portion replaced with the mask token becomes a random missing value in the embedding data.

[0112] As another example, the preprocessing step can apply the above random missing value by replacing some tokens sampled according to the span length among the tokens constituting the embedding data with preset mask tokens. For example, the preprocessing step can replace tokens of the embedding data extracted according to the span length with preset mask tokens. The span length can be extracted from a preset Poisson distribution. For example, if the Poisson distribution is set to 3, the span length is likely to be extracted as a value between 0 and 6 within the Poisson distribution. The preprocessing step can replace consecutive tokens corresponding to the span length with a preset mask token. As described above, the mask token can be set in various ways. The span length can be extracted as 0, in which case the mask token can be added between arbitrary tokens.

[0113] The preprocessing step can apply arbitrary missing values ​​to the embedding data, masking some tokens to set the data as training data. A single embedding data can be generated from multiple training data sets, depending on each masking operation. Alternatively, a single embedding data can be masked only once, resulting in only one corresponding training data set.

[0114] Meanwhile, the artificial neural network learning method may include a learning step of inputting learning data into a graph autoencoder and learning to predict missing values ​​using output data of the graph autoencoder (S910).

[0115] Once the training data is prepared, the training step can set the training data as the input value of the graph autoencoder and perform a training operation to predict the missing values ​​masked in the embedding data using the output value.

[0116] For example, a graph autoencoder can be configured as a graph convolutional neural network (GCN) in the form of an autoencoder, where the input and output are set to the same dimension. Furthermore, the output data can be a preprocessed data set of nodes and edges, i.e., a snapshot of a specific state value of the system data, created as a graph.

[0117] The learning stage can perform learning operations on the training data and output data through supervised or unsupervised learning. For example, the learning stage can supervise a graph autoencoder so that the output data appears as the same value as the corresponding embedding data for the input training data. Alternatively, if the training data is output as the graph autoencoder's output data, the learning stage can perform learning operations through supervised learning to check whether the masked tokens in the training data are output in a consistent manner in the output data.

[0118] The artificial neural network learning method may include an evaluation step for terminating learning when the prediction accuracy for the missing value falls within a preset range (S920).

[0119] For example, the evaluation step measures the prediction accuracy for the missing pair, and terminates training when the prediction accuracy falls within a preset range. The prediction accuracy can be preset by the user.

[0120] Alternatively, the evaluation step can evaluate the prediction accuracy using a validation data set while increasing the number of epochs, and stop learning when the accuracy exceeds a preset threshold. This can improve accuracy while preventing overfitting. To achieve this, the dataset containing node and edge data can be separated and stored as a training data set, a validation data set, and a test data set. The training data set refers to the data set used for learning, while the validation data set refers to the data set not used for learning but set to be used for evaluation, such as when learning is completed. The test set refers to a separate data set used for testing after learning is completed.

[0121] Through the aforementioned operations, the artificial neural network learning method can expand the usability of graph artificial neural networks, which have traditionally focused on node and edge information, such as node state prediction, edge prediction, graph classification, and node value clustering. Furthermore, it can improve accuracy and support the use of GNNs even when some data is problematic. In other words, the artificial neural network learning method can train the artificial neural network to predict missing values ​​with high accuracy even when the embedded data contains missing values. This allows for preparation for various problematic situations that may arise in real-world industrial settings.

[0122] The above description is merely an illustrative example of the technical idea of ​​the present disclosure, and those skilled in the art to which the present disclosure pertains will appreciate that various modifications and variations can be made without departing from the essential characteristics of the technical idea of ​​the present disclosure. In addition, the present embodiments are not intended to limit the technical idea of ​​the present disclosure but rather to explain it, and therefore the scope of the technical idea of ​​the present disclosure is not limited by these embodiments. The scope of protection of the present disclosure should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included within the scope of the rights of the present disclosure.

[0123]

[0124] CROSS-REFERENCE TO RELATED APPLICATION

[0125] This patent application claims priority under 35 USC § 119(a) to Korean Patent Application No. 10-2023-0180888, filed December 13, 2023, the entire contents of which are incorporated herein by reference. Furthermore, this patent application claims priority in countries other than the United States for the same reasons, the entire contents of which are incorporated herein by reference.

Claims

1. In an artificial neural network learning device that learns an artificial neural network, A preprocessing unit that generates multiple embedding data by using multiple node and edge data generated according to the status of system data converted into a graph structure, and generates learning data by applying an arbitrary missing value to each of the multiple embedding data; A learning unit that inputs the above learning data into a graph autoencoder and learns predictions for the missing values ​​using the output data of the graph autoencoder; and An artificial neural network learning device including an evaluation unit that terminates the learning when the prediction accuracy for the above-described missing value falls within a preset range.

2. In paragraph 1, The above preprocessing unit, An artificial neural network learning device that converts the above system data into the graph structure including the node and edge data, and generates the graph structure such that the edge includes relationship information between the nodes.

3. In paragraph 1, The above preprocessing unit, An artificial neural network learning device that generates the node and edge data corresponding to each state of the above system data, thereby generating the plurality of node and edge data.

4. In paragraph 1, The above preprocessing unit, An artificial neural network learning device that generates embedding data for each node using a preset embedding algorithm for each of the plurality of nodes and edge data.

5. In paragraph 1, The above preprocessing unit, An artificial neural network learning device that applies the arbitrary missing value by replacing some tokens randomly selected from among the tokens constituting the above-mentioned embedding data with preset mask tokens.

6. In paragraph 1, The above preprocessing unit, An artificial neural network learning device that applies the random missing value by replacing some tokens sampled according to the span length among the tokens constituting the above embedding data with preset mask tokens.

7. In paragraph 6, The above span length is, An artificial neural network learning device in which the number of sampled tokens is determined according to the span length, and the tokens are extracted from a preset Poisson distribution.

8. In paragraph 6, The above preprocessing unit An artificial neural network learning device that samples two or more tokens when the above span length is extracted as 2 or more, replaces the two or more tokens with one of the preset mask tokens, and applies the random missing value.

9. In paragraph 1, The above graph autoencoder, An artificial neural network learning device consisting of graph convolutional networks in which inputs and outputs are set to the same dimension.

10. In paragraph 9, The above learning department, An artificial neural network learning device that supervises the graph autoencoder so that the output data appears as the same value as the embedding data corresponding to the input learning data.

11. In an artificial neural network learning method for learning an artificial neural network, A preprocessing step of generating multiple embedding data by using multiple node and edge data generated according to the status of system data converted into a graph structure, and generating learning data by applying an arbitrary missing value to each of the multiple embedding data; A learning step of inputting the above learning data into a graph autoencoder and learning a prediction for the missing value using the output data of the graph autoencoder; and An artificial neural network learning method including an evaluation step of terminating the learning when the prediction accuracy for the above-mentioned missing value is within a preset range.

12. In paragraph 11, The above preprocessing step is, An artificial neural network learning method for generating a graph structure by converting the above system data into the graph structure including the node and edge data, wherein the edge includes relationship information between the nodes.

13. In paragraph 11, The above preprocessing step is, An artificial neural network learning method for generating a plurality of node and edge data by generating corresponding node and edge data according to each state of the above system data.

14. In paragraph 11, The above preprocessing step is, An artificial neural network learning method for generating embedding data for each node using a preset embedding algorithm for each of the above-described plurality of nodes and edge data.

15. In paragraph 11, The above preprocessing step is, An artificial neural network learning method for applying the arbitrary missing value by replacing some tokens randomly selected from among the tokens constituting the above embedding data with preset mask tokens.

16. In paragraph 11, The above preprocessing step is, An artificial neural network learning method for applying the above random loss value by replacing some tokens sampled according to the span length among the tokens constituting the above embedding data with preset mask tokens.

17. In paragraph 16, The above span length is, An artificial neural network learning method in which the number of sampled tokens is determined according to the span length, and the tokens are extracted from a preset Poisson distribution.

18. In paragraph 16, The above preprocessing step An artificial neural network learning method for applying the random missing value by sampling two or more tokens and replacing the two or more tokens with one of the preset mask tokens when the above span length is extracted as 2 or more.

19. In paragraph 11, The above graph autoencoder, An artificial neural network learning method consisting of graph convolutional networks in which input and output are set to the same dimension.

20. In paragraph 19, The above learning steps are: An artificial neural network learning method for supervising the graph autoencoder so that the output data appears as the same value as the embedding data corresponding to the input learning data.

Citation Information

Patent Citations

  • Electronic device arranging one-off social gathering based on location

    KR1020240138619A

  • Exclusive jig for 3D printing tire puzzle mold processing

    KR1020250024172A

  • Ice maker

    KR1020250064338A

  • Accident prevention monitoring method and system for tower crane

    KR102623060B1

  • Anomaly detecting method in sequence of control segment of automation equipment using graph autoencoder

    US20230027840A1

Cited By

  • Workshop production line whole-process data processing and intelligent monitoring method and system

    CN121069859A