Multitask processing model learning method and multitask processing execution method using machine learning model learned based on method
By using Riemannian geometry to perform geometric alignment in the latent space, the parameters of the multi-task processing model are optimized, which solves the limitation of transfer learning on small molecular structure datasets, achieves efficient prediction performance and stability, and expands the applicability of the model.
Patent Information
- Application Number
- CN202480041721.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-23
- Filing Date
- 2024-06-24
- Publication Date
- 2026-01-16
AI Technical Summary
Existing transfer learning techniques exhibit limitations on small and complex molecular structure datasets, especially in high-dimensional and structurally complex non-Euclidean spaces, resulting in insufficient applicability of models to new domains or tasks.
We employ Riemannian geometry to perform geometric alignment in the integrated latent space. By obtaining the geometric alignment vector and calculating the geometric alignment loss, we update the parameters of the multi-task processing model. By combining embedding vectors, perturbation vectors, original and perturbation latent vectors, we optimize the knowledge transfer between the source and target tasks.
It improves the model's predictive performance and stability on small datasets, expands the model's applicability, especially in complex regression problems, and maintains geometric consistency across tasks, thus enhancing the model's generalization performance.
Smart Images

Figure CN121359147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a multi-task processing model learning method and a multi-task processing execution method using a machine learning model learned based on the method. More specifically, it relates to a multi-task processing model learning method and a multi-task processing execution method using a machine learning model learned based on the method, which migrates and learns knowledge data scattered in a latent space per task to each other through geometrical alignment in one integrated latent space in order to process a multi-task based on outputs of multiple domains. BACKGROUND
[0002] Machine learning and artificial intelligence models require large data. However, it is actually limited to always have sufficient data. In particular, when trying to apply a model to a new domain or a new task, it further deepens its limitations. Representatively, a molecular structure dataset is a good example of such a situation. In the field of chemistry and pharmacy, predicting the characteristics of a new molecule requires data support, but it is difficult to obtain experimental data on each molecule, and it requires a lot of cost. Therefore, the necessity of transfer learning technology that applies the knowledge of a learned model to a new task increases.
[0003] However, the development of existing transfer learning has been mainly focused on classification problems of large data sets such as image or text data. Therefore, the existing transfer learning technology shows limitations when applied to regression problems or small and complex data sets such as molecular data sets. In particular, in the case of high-dimensional and complex structures in which the binding relationship between each structure is very important, such as a molecular structure data, the existing transfer learning technology based on Euclidean space cannot effectively handle the complex structure in a non-Euclidean space.
[0004] On the other hand, Riemannian geometry can more effectively represent and analyze the complex structure of data by implementing calculus computation in curved space. This Riemannian geometry method advantageously acts on geometrical alignment between a source task and a target task by assuming that latent vectors exist on a curved manifold.
[0005] Therefore, in the above background, there is an urgent need to introduce a new technology that exhibits high prediction performance and stability even in a small data set, implements more effective transfer learning, and improves model normalization performance, thereby improving its generalization performance. SUMMARY
[0006] PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] An embodiment of the present application aims to provide a multi-task processing model learning method that migrates and learns knowledge data scattered per task in a latent space by geometrical alignment in a comprehensive latent space for processing multi-tasks based on multiple domains, and a multi-task processing execution method using a machine learning model learned based on the method.
[0008] However, the technical problems to be solved by the present application and embodiments of the present application are not limited to the technical problems described above, and other technical problems can exist.
[0009] Means for solving the problems
[0010] The multi-task processing model learning method provided by the embodiments of the present application is a method for causing a multi-task processing model to learn by a computing system including a memory and a processor, the method including the steps of: initializing a multi-tasking model for processing multi-tasks with respect to multiple domains; acquiring prescribed experimental data; and training the multi-tasking model based on the acquired experimental data, the step of training the multi-tasking model including the steps of: acquiring a geometric alignment vector that supports geometrical alignment between data in a comprehensive latent space based on the experimental data; calculating a geometric alignment loss based on the acquired geometric alignment vector; and updating parameters of the multi-tasking model based on the calculated geometric alignment loss.
[0011] In another aspect, the step of acquiring the geometric alignment vector includes the step of acquiring an embedding vector that is obtained by projecting the experimental data to a prescribed embedding space and converting to a vector form by an embedding module included in the multi-tasking model.
[0012] In another aspect, the step of acquiring the geometric alignment vector further includes the step of acquiring a plurality of perturbation vectors that are obtained by moving the embedding vector in a prescribed direction by a perturbation module included in the multi-tasking model.
[0013] In another aspect, the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original latent vector, which is a vector obtained by projecting the embedding vector into a latent space of a first task (Taskl) according to an encoder module (Encoder Module) included in the multi-task processing model; and obtaining a perturbation latent vector, which is a vector obtained by projecting the perturbation vector into the latent space of the first task according to the encoder module included in the multi-task processing model.
[0014] In another aspect, the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original transfer vector, which is a vector obtained by mapping the original latent vector into a latent space of a second task (Task2) according to a transfer module (Transfer Module) included in the multi-task processing model; and obtaining a perturbation transfer vector, which is a vector obtained by mapping the perturbation latent vector into the latent space of the second task according to the transfer module included in the multi-task processing model.
[0015] In another aspect, the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original inverse vector, which is a vector obtained by remapping the original transfer vector into the latent space of the first task according to an inverse transfer module (Inverse Transfer Module) included in the multi-task processing model; and obtaining a perturbation inverse vector, which is a vector obtained by remapping the perturbation transfer vector into the latent space of the first task according to the inverse transfer module included in the multi-task processing model.
[0016] In another aspect, the step of calculating the geometric alignment loss includes the steps of: calculating a regression loss, an autoencoder loss, a consistency loss, a mapping loss, and a distance loss from the geometric alignment vectors; and calculating an integrated loss by weighted sum of the calculated regression loss, autoencoder loss, consistency loss, mapping loss, and distance loss.
[0017] In another aspect, the step of updating the parameters of the multi-task processing model includes the step of updating the parameters in a direction that minimizes the integrated loss.
[0018] In another aspect, the step of calculating the geometric alignment loss further includes the steps of: obtaining predicted values based on the latent vectors from a regressor module included in the multi-task processing model; calculating a mean squared error from the obtained predicted values and label values corresponding to the latent vectors; and calculating the regression loss from the calculated mean squared error.
[0019] In another aspect, the step of calculating the geometric alignment loss further includes the steps of: calculating a mean squared error from the original latent vectors and the original inverse vectors; and calculating the autoencoder loss from the calculated mean squared error.
[0020] In another aspect, the step of calculating the geometric alignment loss further includes the steps of: calculating a mean squared error from perturbation transfer vectors obtained by mapping from the latent space of the first task to the latent space of the second task and perturbation transfer vectors obtained by mapping from the latent space of the second task to the latent space of the first task; and calculating the consistency loss from the calculated mean squared error.
[0021] In another aspect, the step of calculating the geometric alignment loss further includes the steps of: calculating a mean squared error from label values for the first task and predicted values based on original inverse vectors for the second task; and calculating the mapping loss from the calculated mean squared error.
[0022] In another aspect, the step of calculating the geometric alignment loss further includes the steps of: calculating a first migration vector displacement, the first migration vector displacement being a distance between an original migration vector and a perturbed migration vector with respect to the first task; calculating a second migration vector displacement, the second migration vector displacement being a distance between an original migration vector and a perturbed migration vector with respect to the second task; calculating a mean square error based on the calculated first migration vector displacement and the second migration vector displacement; and calculating the distance loss based on the calculated mean square error.
[0023] In another aspect, the experimental data includes substance inherent property information and substance physical property specific information, the substance inherent property information being information that specifically specifies an inherent property held by a prescribed matter, and the substance physical property specific information being information that specifically specifies a data value possessed by a prescribed matter with respect to a prescribed physical property.
[0024] In another aspect, the multi-task processing model learning server according to an embodiment of the present disclosure includes at least one memory and at least one processor that reads at least one application program stored in the memory, and causes a multi-task processing model to learn, and the instructions of the processor include instructions to perform the steps of: initializing a multi-tasking model for processing a multi-task based on a plurality of domains; acquiring prescribed experimental data; acquiring a geometric alignment vector that supports geometric alignment between data in one comprehensive manifold based on the acquired experimental data; calculating a geometric alignment loss based on the acquired geometric alignment vector; and updating parameters of the multi-task processing model based on the calculated geometric alignment loss.
[0025] Effects of Invention
[0026] The multi-task processing model learning method according to an embodiment of the present disclosure and the multi-task processing execution method using a machine learning model learned based on the method have the following effects: the knowledge learned in a source task is transferred to a target task through transfer learning to eliminate a data deficiency problem, and thus a multi-task processing model that maintains high performance even in a small data set can be provided.
[0027] Accordingly, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned based on the method according to one embodiment of the present application have an effect of being able to expand the range of application to a field in which a machine learning model is difficult to be applied due to insufficient data or domain knowledge.
[0028] Also, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned based on the method according to one embodiment of the present application have an effect of being able to exert a high prediction performance even on a complex regression problem such as a molecular dataset by providing a dedicated transfer learning technique that can be effectively applied to a regression problem.
[0029] Also, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned based on the method according to one embodiment of the present application have an effect of being able to maintain the geometric consistency between tasks and improve the efficiency of transfer learning by optimizing the knowledge transfer between a source task and a target task in a Riemannian geometry manner.
[0030] Also, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned based on the method according to one embodiment of the present application have an effect of being able to further improve the generalization performance of the model by normalizing multiple aspects of the model by combining multiple loss functions.
[0031] Accordingly, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned based on the method according to one embodiment of the present application have an effect of being able to improve the quality of the entire related industry by providing a multi-task processing model that can be widely applied to various substances (raw materials).
[0032] However, the effects obtainable in the present application are not limited to the above-mentioned effects, and other effects not mentioned above can be clearly understood from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 An example of a block diagram of a computing system that implements a multi-task processing learning model providing service according to one embodiment of the present application is shown.
[0034] Figure 2 An example of a block diagram of a computing device that implements a multi-task processing learning model providing service according to one embodiment of the present application is shown.
[0035] Figure 3 An example of a block diagram of another aspect of a computing device that implements a multi-task processing learning model providing service according to one embodiment of the present application is shown.
[0036] Figure 4 and Figure 5 An example of a conceptual diagram for illustrating a multi-task processing learning model according to an embodiment of the present disclosure is shown.
[0037] Figure 6 An internal block diagram of a multi-task processing learning model according to an embodiment of the present disclosure is shown.
[0038] Figure 7 An example of a conceptual diagram for illustrating a multi-task processing model learning method according to an embodiment of the present disclosure is shown.
[0039] Figure 8 A block diagram type flowchart for illustrating a multi-task processing model learning method according to an embodiment of the present disclosure is shown.
[0040] Figure 9 A block diagram type flowchart for illustrating a multi-task processing learning model training method according to an embodiment of the present disclosure is shown.
[0041] Figure 10 An example of a conceptual diagram for illustrating a multi-task processing learning model training method according to an embodiment of the present disclosure is shown.
[0042] Figure 11 and Figure 12 An example of a diagram for illustrating a regression loss calculation method according to an embodiment of the present disclosure is shown.
[0043] Figure 13 An example of a diagram for illustrating a comprehensive latent space mapping method according to an embodiment of the present disclosure is shown.
[0044] Figure 14 and Figure 15 An example of a diagram for illustrating a consistency loss calculation method according to an embodiment of the present disclosure is shown.
[0045] Figure 16 and Figure 17 An example of a diagram for illustrating a mapping loss calculation method according to an embodiment of the present disclosure is shown.
[0046] Figure 18 An example of a diagram for illustrating a comprehensive loss calculation method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0047] This invention is capable of various modifications and embodiments, which will be described in detail by illustrating specific embodiments with reference to the accompanying drawings. The effects, features, and methods of implementing this invention will become clear from the detailed embodiments described below with reference to the accompanying drawings. However, this invention is not limited to the embodiments disclosed below and can be implemented in a variety of forms. In the following embodiments, terms such as "first," "second," etc., are not limiting in meaning but are used to distinguish one constituent element from other constituent elements. Furthermore, singular expressions include plural expressions unless the context clearly differs. Also, terms such as "including" or "having" indicate the presence of a feature or constituent element described in the specification, without pre-excluding the possibility of adding more than one other feature or constituent element. Furthermore, for ease of explanation, the size of constituent elements may be enlarged or reduced in the drawings. For example, the sizes and thicknesses of the various structures shown in the drawings are arbitrarily illustrated for ease of explanation, and therefore, this invention is not necessarily limited to the illustrated content.
[0048] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing with reference to the accompanying drawings, the same or corresponding constituent elements are given the same reference numerals, and repeated descriptions of them are omitted.
[0049] [An exemplary system for providing services using a multi-task processing learning model]
[0050] The following describes in detail, with reference to the accompanying drawings, an exemplary system that implements a multi-task processing learning model to provide services for processing multi-task outputs based on multiple domains. This model enables the transfer and learning of knowledge data scattered across a latent space for each task within a unified latent space through geometric alignment, and performs multi-task processing based thereon.
[0051] Figure 1 An example block diagram of a computing system that provides services for a multi-task processing learning model, according to an embodiment of the present invention, is shown.
[0052] Reference Figure 1 The computing system 1000 of the present invention, which provides services for a multi-task processing learning model, includes a user computing device 110, a server computing system 130, and a training computing system 150, which are capable of communicating via a network 170.
[0053] The multi-task processing model learning method provided by one embodiment of the present application and the multi-task processing execution method using the machine learning model learned based on the method can be 1) implemented and provided locally by the user computing device 110, 2) implemented and provided in the form of a web service by the server computing system 130 in communication with the user computing device 110, and 3) implemented and provided in cooperation with each other by the user computing device 110 and the server computing system 130.
[0054] At this time, in an embodiment, the user computing device 110 and / or the server computing system 130 can cause the machine learning model (120 and / or 140) to learn through interaction with a training computing system 150 connected in communication via the network 170. The training computing system 150 can be independent of the server computing system 130 or can be a part of the server computing system 130.
[0055] Also, at this time, regarding the artificial intelligence model, it can be 1) caused to learn locally by the user computing device 110 directly, 2) caused to learn by the server computing system 130 and the user computing device 110 interacting with each other via the network 170, and 3) caused to learn by the separate training computing system 150 using various training techniques and learning techniques. Also, it can be implemented in such a manner that the artificial intelligence model learned by the training computing system 150 is transmitted to the user computing device 110 and / or the server computing system 130 via the network 170 to be provided / updated.
[0056] In some embodiments, the training computing system 150 can be a part of the server computing system 130 or a part of the user computing device 110.
[0057] The user computing device 110 can include a smart phone, a mobile phone, a device for digital broadcasting, a PDA (personal digital assistants), a PMP (portable multimedia player), a desktop computer, a wearable device, an embedded computing device, a tablet PC, and all other types of computing devices.
[0058] Such a user computing device 110 includes at least one processor 111 and a memory 112. The processor 111 may consist of at least one processor or multiple processors electrically connected to each other, selected from a central processing unit (CPU), a graphics processing unit (GPU), application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions.
[0059] The memory 112 may include one or more non-transitory / transitory computer-readable storage media, such as RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM (Erasable Programmable Read Only Memory), flash memory devices, disks, and combinations thereof, and may include web storage of a server that performs the storage function of the memory on the Internet. Such a memory 112 may store data 113 and instructions 114 required by the at least one processor 111 to perform functions such as enabling the artificial intelligence model to learn or performing multi-task processing learning through the artificial intelligence model.
[0060] In one embodiment, the user computing device 110 may store at least one or more machine learning models 120.
[0061] In detail, the machine learning model 120 can be a wide variety of machine learning models such as multiple neural networks (e.g., deep neural networks) or other types of machine learning models including nonlinear models and / or linear models, and can be configured as a combination of them.
[0062] At this time, the neural network may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, and / or other types of neural networks.
[0063] In one embodiment, the user computing device 110 may receive at least one machine learning model 120 from the server computing system 130 via the network 170, store it in the memory 112, and then execute the stored machine learning model 120 through the processor 111 to perform multi-task processing learning, etc.
[0064] In another embodiment, the server computing system 130 may include at least one or more machine learning models 140, which perform actions and are linked with the user computing device 110 in a manner that communicates relevant data with the user computing device 110 to provide services to the user through a multi-task processing learning model.
[0065] For example, user computing device 110 can provide services through a multitasking learning model in such a way that server computing system 130 provides outputs via a webpage using machine learning model 140 in response to user input.
[0066] Furthermore, the artificial intelligence model can also be implemented in such a way that at least a portion of the machine learning model (120 and / or 140) is executed in the user computing device 110, and the remainder is executed in the server computing system 130.
[0067] Furthermore, the user computing device 110 may include at least one input component 121 for sensing user input. For example, the user input component 121 may include a touch sensor (e.g., a touchscreen and / or touchpad, etc.) for sensing the touch of a user's input medium (e.g., a finger or stylus), an image sensor for sensing the user's motion input, a microphone for sensing the user's voice input, a button, a mouse, and / or a keyboard, etc. Moreover, when the user input component 121 receives input from an external controller (e.g., a mouse and / or keyboard, etc.) via an interface, it may include both an interface and an external controller.
[0068] The server computing system 130 may include at least one processor 131 and a memory 132. The processor 131 may consist of at least one or more processors electrically connected to each other, including a central processing unit (CPU), a graphics processing unit (GPU), application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions.
[0069] Furthermore, the memory 132 may include one or more non-transitory / transitory computer-readable storage media, such as RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), flash memory devices, disks, etc., and combinations thereof. Such a memory 132 may store data 133 and instructions 134 required for the processor 131 to perform functional actions, such as enabling an artificial intelligence model to learn, or performing multi-task learning through the artificial intelligence model.
[0070] In one embodiment, the server computing system 130 may be implemented using at least one or more computing devices. For example, the server computing system 130 may be implemented such that multiple computing devices operate according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Furthermore, the server computing system 130 may include multiple computing devices connected via a network 170.
[0071] Furthermore, the server computing system 130 can store at least one machine learning model 140. For example, the server computing system 130 may include neural networks and / or other multi-layer nonlinear models as machine learning models 140. Exemplary neural networks may include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.
[0072] The training computing system 150 includes at least one processor 151 and a memory 152. The processor 151 may consist of at least one or more processors electrically connected to each other, including a central processing unit (CPU), a graphics processing unit (GPU), application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions.
[0073] Furthermore, the memory 152 may include one or more non-transitory / transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, and disks, as well as combinations thereof. Such a memory 152 can store data 153 and instructions 154 required by the processor 151 for tasks such as learning artificial intelligence models.
[0074] For example, training computation system 150 may include model trainer 160, which uses a variety of training or learning techniques, such as error backpropagation (according to...). Figure 3 The framework shown enables the machine learning models (120 and / or 140) stored in the user computing device 110 and / or server computing system 130 to learn (train).
[0075] For example, such a model trainer 160 can update more than one parameter of the machine learning model (120 and / or 140) in a backpropagation manner according to a defined loss function.
[0076] In some implementations, backpropagation of the error may include truncated backpropagation through time. The model trainer 160 may perform various generalization techniques (e.g., weight decay, drop-out, and / or knowledge distillation) to improve the generalization ability of the trained machine learning models (120 and / or 140).
[0077] In particular, the model trainer 160 can train machine learning models (120 and / or 140) based on a series of training data 161. The training data 161 can include data in various forms, such as images, audio samples, and / or text. Examples of usable image types can include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images, and / or various other image formats.
[0078] Such training data 161 can be provided by user computing device 110 and / or server computing system 130. In the case where the machine learning model (120 and / or 140) is trained by the computing device to learn from the specific data of user computing device 110, the machine learning model (120 and / or 140) can be characterized into a personalized model.
[0079] Furthermore, the model trainer 160 includes computer logic for providing the desired functionality.
[0080] Furthermore, the model trainer 160 can be implemented as hardware, firmware, and / or software controlling a general-purpose processor. In one implementation, the model trainer 160 may include a program file stored in a storage device, which may be loaded into memory 152 and executed by one or more processors 151. In another implementation, the model trainer 160 includes one or more sets of computer-executable data 153 and instructions 154 stored in a tangible computer-readable storage medium, such as a RAM hard disk or optical or magnetic medium.
[0081] Network 170 includes, but is not limited to, 3GPP (3rd Generation Partnership Project) networks, LTE (Long Term Evolution) networks, WIMAX (World Interoperability for Microwave Access) networks, the Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth networks, satellite broadcasting networks, analog broadcasting networks, and / or DMB (Digital Multimedia Broadcasting) networks.
[0082] Typically, communication via network 170 can be conducted using any type of wired and / or wireless connection, through a wide variety of communication protocols (e.g., TCP / IP (Transmission Control Protocol / Internet Protocol), HTTP (Hypertext Transfer Protocol), SMTP (Simple Mail Transfer Protocol), and / or FTP (File Transfer Protocol), etc.), encodings or formats (e.g., HTML (Hypertext Markup Language) and / or XML (Extensible Markup Language), etc.), and / or protection schemes (e.g., VPN (Virtual Private Network), Secure HTTP, and / or SSL (Secure Socket Layer), etc.).
[0083] Figure 2 An example block diagram of a computing device for providing services through a multi-task learning model, as provided in an embodiment of the present invention, is shown.
[0084] Reference Figure 2The computing device 100 included in the user computing device 110, server computing system 130, and training computing system 150 includes various applications (e.g., application 1 to application N). Each application may include a machine learning library and more than one machine learning model. For example, applications may include image processing (e.g., detection, classification, and / or segmentation) applications, text communication applications, email applications, dictation applications, virtual keyboard applications, browser applications, and / or chatbot applications, etc.
[0085] In an embodiment, the computing device 100 may include a model trainer 160 for enabling an artificial intelligence model to learn, and by storing the learned artificial intelligence model and enabling it to perform actions, it can provide output data corresponding to specified input data (as an example, information on the inherent properties of a substance and / or specific information on the physical properties of a substance, etc.).
[0086] Each application of computing device 100 can communicate with multiple other components of computing device 100, such as at least one sensor, context manager, device state component, and / or additional components. In one embodiment, each application can communicate with various device components using an API (Application Programming Interface) (e.g., a public API). In one embodiment, the API used by each application can be application-specific.
[0087] Figure 3 An example block diagram of another aspect of a computing device for providing services through a multi-task learning model, as provided in one embodiment of the present invention, is shown.
[0088] Reference Figure 3 The computing device 200 includes various applications (e.g., application 1 to application N). Each application can communicate with the central intelligence layer. For example, applications may include image processing applications, text messaging applications, email applications, dictation applications, virtual keyboard applications, and / or browser applications, etc. In one embodiment, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API applicable to all applications).
[0089] The central intelligence layer can include multiple machine learning models. For example, such as Figure 3As shown, at least a portion of each machine learning model can be provided for each application and can be managed by a central intelligence layer. In another implementation, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model relative to all applications. In some implementations, the central intelligence layer can be contained within the operating system of computing device 200, or implemented differently.
[0090] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for computing device 200. For example... Figure 3 As shown, the central device data layer can communicate with multiple other components of the computing device 200, such as more than one sensor, context manager, device state component, and / or additional components. In some implementations, the central device data layer can communicate with individual device components using APIs (e.g., private APIs).
[0091] The techniques described in this specification can be referenced not only to servers, databases, software applications, and other computer-based systems, but also to the actions taken and the information transmitted to or from said systems. It should be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible structures, combinations, divisions of operations, and functionalities formed between components. For example, the procedures described in this specification can be implemented using multiple devices or components operating as a single device or component or in combination. Databases and applications can be implemented in a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0092] [Multi-tasking Learning Model (MtLM)]
[0093] Figure 4 as well as Figure 5 An example of a concept diagram is shown to illustrate a multi-task processing learning model (MtLM) provided in one embodiment of the present invention.
[0094] Reference Figure 4 as well as Figure 5The multi-task processing learning model (MtLM) (Geometrically Aligned Transfer Encoder Model) provided by the embodiments of the present invention can be a machine learning model for processing multi-task outputs based on multiple domains, which aligns knowledge data (as an example, latent vectors, etc.) scattered in the latent space for each task in a comprehensive latent space (M: Manifold) through geometric transfer.
[0095] That is, the Multi-Task Processing Learning Model (MtLM) provided in the embodiment not only learns knowledge data about various domains simultaneously, but also efficiently learns the relationships between the domains, thereby expanding the learning domain and performing effective multi-task processing learning that can learn the local patterns of each domain and the common principles between multiple domains at the same time.
[0096] Therefore, the Multi-Task Processing Learning Model (MtLM) can directly improve the processing performance and accuracy of various tasks in multi-task processing based on the model learned as described above.
[0097] In this embodiment, the multi-task processing learning model (MtLM) is capable of performing pre-training based on specified experimental data.
[0098] The experimental data provided in the embodiments refers to the learning data used to train the multi-task processing learning model (MtLM), which may include specified information on the inherent properties of the material and specific information on the physical properties of the material.
[0099] In this case, the material inherent property information provided in the embodiment may refer to information that specifically specifies the inherent properties possessed by the specified material.
[0100] For example, information about the inherent properties of a substance may include the specified substance name, molecular structural formula and / or chemical formula, etc.
[0101] Furthermore, the specific information on the physical properties of a substance provided in the embodiments may refer to information that specifically specifies the data values that a given substance possesses regarding a given physical property.
[0102] For example, specific information about the physical properties of a substance may include specified physical property (i.e., domain) values such as boiling point, melting point, refractive index, solubility, viscosity, surface tension, density, strength and / or thermal conductivity.
[0103] On the other hand, in the embodiments, the multi-task processing learning model (MtLM) that performs pre-learning as described above can accept input of specified material inherent property information and / or material physical property specific information, and output data predicted based on the input information and the learned knowledge.
[0104] As an example, the Multi-Task Processing Learning Model (MtLM) can accept input of specified material inherent property information and output specific information of the material's physical properties predicted based on the input information and learned knowledge.
[0105] As another embodiment, the Multi-Task Processing Learning Model (MtLM) can accept specific information about the physical properties of a material as input and output information about the inherent properties of the material predicted based on the input information and the learned knowledge.
[0106] As another embodiment, the Multi-Task Processing Learning Model (MtLM) can accept inputs of specified material intrinsic property information and material physical property specific information, and output the best material intrinsic property information and material physical property specific information predicted based on the input information and learned knowledge.
[0107] Figure 6 An internal block diagram of a multi-task processing learning model (MtLM) provided by an embodiment of the present invention is shown.
[0108] Reference Figure 6 On the other hand, the multi-task processing learning model (MtLM) provided in the embodiment may include at least one embedding module (EBM), encoder module (ECM), regressor module (RGM), transfer module (TFM), inverse transfer module (ITM), perturbation module (PBM), and loss calculation module (LCM).
[0109] In detail, the embedding module (EBM) provided in the embodiments of the present invention can be a pre-encoder module that converts specified input data into an embedding vector.
[0110] That is, an embedding module (EBM) can be a module that projects specific input data into a specified embedding space and transforms it into vector form.
[0111] As an example, the embedding module (EBM) can provide embedding vectors about the input data according to the DMPNN (Directed Message Passing Neural Network) structure.
[0112] Furthermore, the encoder module (ECM) provided in the embodiments of the present invention can be a module that takes a specified embedding vector as input, projects the input embedding vector onto the latent space corresponding to the corresponding task, and transforms it into a latent vector.
[0113] That is, the encoder module (ECM) can be a module that extracts the main features of the input embedding vector and represents them on the corresponding latent space.
[0114] In an embodiment, such an encoder module (ECM) may include multiple encoder modules (ECMs) corresponding to multiple domains respectively.
[0115] As an example, the encoder module (ECM) may include a first encoder module (ECM) corresponding to a first domain (e.g., boiling point) and a second encoder module (ECM) corresponding to a second domain (e.g., melting point).
[0116] In this embodiment, any one of the multiple encoder modules (ECM) can be the encoder module (ECM) corresponding to the source task of transfer learning provided in the embodiment of the present invention, i.e., the source encoder module (ECM).
[0117] Furthermore, any of the remaining encoder modules (ECMs) other than the source encoder module (ECM) can be an encoder module (ECM) corresponding to the target task of transfer learning provided in the embodiments of the present invention, i.e., a target encoder module (ECM).
[0118] Furthermore, the Regressor Module (RGM) provided in the embodiments of the present invention can be a Head module that takes a specified latent vector as input and generates a final predicted value based on the input latent vector.
[0119] This regressor module (RGM) can directly participate in the generation of the final output to determine the model's predictive performance.
[0120] Furthermore, in an embodiment, the regressor module (RGM) may include multiple regressor modules (RGM) corresponding to multiple domains respectively.
[0121] As an example, the regressor module (RGM) may include a first regressor module (RGM) corresponding to a first domain (e.g., boiling point) and a second regressor module (RGM) corresponding to a second domain (e.g., melting point).
[0122] In this embodiment, any one of the multiple regressor modules (RGM) can be a regressor module (RGM) corresponding to the source task of transfer learning provided in the embodiment of the present invention, i.e., a source regressor module (RGM).
[0123] Furthermore, any of the remaining regressor modules (RGMs) other than the source regressor module (RGM) can be a regressor module (RGM) corresponding to the target task of transfer learning provided in the embodiments of the present invention, i.e., the target regressor module (RGM).
[0124] Furthermore, the transfer module (TFM) provided in the embodiments of the present invention can be a module that maps a specified latent vector to the latent space of other tasks and converts it into a transfer vector.
[0125] In detail, in an embodiment, the transfer module (TFM) can map a specific latent vector to the latent space of another task and transform it into a transfer vector according to the Riemannian geometry method.
[0126] In this process, the transfer module (TFM) can achieve geometric alignment between the mapped tasks according to embodiments of the present invention. This will be described in detail later in the section on multi-task processing model learning methods.
[0127] That is, in the embodiments, the transfer module (TFM) is able to efficiently perform the transfer of knowledge data between multiple tasks in a manner that maps the potential vectors about the first task to the potential space about the second task through the geometric alignment provided by the embodiments of the present invention.
[0128] In this embodiment, the transfer module (TFM) utilizes an autoencoder structure to support data processing that improves the accuracy and consistency of the transformed vector (i.e., the transfer vector).
[0129] Furthermore, in an embodiment, the migration module (TFM) may include multiple migration modules (TFMs) corresponding to multiple domains respectively.
[0130] As an example, the migration module (TFM) may include a first migration module (TFM) corresponding to a first domain (e.g., boiling point) and a second migration module (TFM) corresponding to a second domain (e.g., melting point).
[0131] In this embodiment, any one of the multiple transfer modules (TFMs) can be a transfer module (TFM) corresponding to the source task of transfer learning provided in the embodiment of the present invention, i.e., a source transfer module (TFM).
[0132] Furthermore, any of the remaining transfer modules (TFMs) other than the source transfer module (TFM) can be a transfer module (TFM) corresponding to the target task of transfer learning provided in the embodiments of the present invention, i.e., a target transfer module (TFM).
[0133] Furthermore, the inverse transfer module (ITM) provided in the embodiments of the present invention can be a module that reconstructs the transfer vectors that have been mapped to the potential space of other tasks by the transfer module (TFM) and transformed, and then remaps them to the original potential space.
[0134] In this way, in the embodiment, the inverse migration module (ITM) is able to generate a vector that reconstructs the migration vector and transforms it back into the original state (hereinafter referred to as the inverse vector).
[0135] In this embodiment, the inverse migration module (ITM) can improve the stability of the above reconstruction process and the accuracy and consistency of the corresponding migration vectors by utilizing an autoencoder structure.
[0136] In an embodiment, such an inversion migration module (ITM) may include multiple inversion migration modules (ITMs) corresponding to multiple domains respectively.
[0137] As an example, the reverse migration module (ITM) may include a first reverse migration module (ITM) corresponding to a first domain (e.g., boiling point) and a second reverse migration module (ITM) corresponding to a second domain (e.g., melting point).
[0138] In this embodiment, any one of the multiple inverse transfer modules (ITMs) can be an inverse transfer module (ITM) corresponding to the source task of transfer learning provided in the embodiment of the present invention, i.e., a source inverse transfer module (ITM).
[0139] Furthermore, any of the remaining inverse transfer modules (ITMs) other than the source inverse transfer module (ITM) can be an inverse transfer module (ITM) corresponding to the target task of transfer learning provided in the embodiments of the present invention, i.e., a target inverse transfer module (ITM).
[0140] Furthermore, the perturbation module (PBM) provided in the embodiments of the present invention can be a module that applies a specified change to a specified embedding vector to generate multiple perturbation vectors.
[0141] In detail, in an embodiment, a perturbation module (PBM) may be a module that generates multiple perturbation vectors (i.e., perturbation points) around a given embedding vector by applying a change that moves a particular embedding vector in a specified direction.
[0142] At this point, the generated perturbation vectors can be designed to maintain a relative distance from the corresponding embedding vectors, thereby effectively assisting in geometric alignment.
[0143] That is, the perturbation module (PBM) described above assists in the geometric alignment of the model by generating multiple perturbation vectors, thereby helping to keep the coordinate systems of the source and target tasks consistent.
[0144] Furthermore, in the embodiments, the perturbation module (PBM) is able to support the calculation of the distance between the specified embedding vector and the multiple perturbation vectors generated based thereon, and to keep the displacement between the source task and the target task consistent according to the calculated distance.
[0145] As a result, the perturbation module (PBM) can more easily maintain the consistency of the model in the latent space.
[0146] According to an embodiment, the perturbation module (PBM) can prevent model overfitting and improve generalization performance by forcibly maintaining a prescribed embedding vector and the relationship between multiple perturbation vectors generated therefrom.
[0147] Furthermore, the loss calculation module (LCM) provided in the embodiments of the present invention can be a module that calculates various loss functions based on various vectors obtained through the multi-task processing learning model (MtLM).
[0148] In this embodiment, the loss calculation module (LCM) can calculate the regression loss, autoencoder loss, consistency loss, mapping loss, distance loss, and / or integrated loss provided by the embodiments of the present invention. These will be described in detail later in the section on multi-task processing model learning methods.
[0149] Therefore, the Loss Calculation Module (LCM) can support the normalization and learning of different parts of the model, and can provide feedback for model learning to achieve model optimization.
[0150] On the other hand, in embodiments of the present invention, the multi-task processing learning model (MtLM) can perform model optimization and updates through various data processing procedures linked with the above-mentioned modules.
[0151] For example, the multi-task processing learning model (MtLM) can perform model optimization and parameter updates in conjunction with the above modules, based on optimization algorithms such as AdamW (Adam with WeightDecay; Adaptive Moment Estimator Optimizer).
[0152] In this way, in embodiments of the present invention, the multi-task processing learning model (MtLM) not only learns knowledge data about various domains simultaneously, but also efficiently learns the relationships between the domains, thereby expanding the learning domain and performing effective multi-task processing learning that can learn local patterns based on each domain as well as common principles between multiple domains.
[0153] Therefore, the Multi-Task Processing Learning Model (MtLM) can directly improve the processing performance and accuracy of various tasks in multi-task processing based on the model learned as described above.
[0154] [Service Implementation Methods for Multi-Task Processing Learning Models]
[0155] The following provides a detailed description of a method for implementing a multi-task processing learning model (MtLM) service in a computing system 1000 according to an embodiment of the present invention. The multi-task processing learning model (MtLM) provides services for processing multi-tasks for outputs based on multiple domains by using geometric alignment to transfer and learn knowledge data scattered in the latent space for each task to each other in a comprehensive latent space, and performing multi-task processing based thereon.
[0156] Generally, existing transfer learning techniques mainly focus on classification tasks for image and / or language datasets, and have limitations in solving regression problems or non-Euclidean space problems.
[0157] Especially when the training dataset is insufficient, the predictive performance for the above problems will inevitably decline. If multi-task processing that considers multiple task types is also required, the performance decline in related learning and prediction will be further aggravated.
[0158] Furthermore, most existing methods are optimized for processing data in Euclidean space, and therefore cannot work effectively in complex curved or nonlinear spaces.
[0159] Figure 7 An example of a concept diagram is shown to illustrate a multi-task processing model learning method provided in one embodiment of the present invention.
[0160] Therefore, as Figure 7 As shown, an embodiment of the present invention provides a computing system 1000 that offers a novel multi-task processing model learning method that can overcome the limitations of regression problems on small datasets and existing transfer learning techniques, as well as a multi-task processing execution method that utilizes a machine learning model learned based on this method.
[0161] In the following description provided in an embodiment of the present invention, for the sake of efficient explanation, the above-described substances are limited to "molecules" and their related domains are described based on "physical properties".
[0162] This is because molecular datasets typically have a small amount of data, include multiple task types, and primarily address regression problems.
[0163] In other words, molecular datasets have the following characteristics: they require handling multiple tasks related to a large number of physical properties, but the existing data is very limited, and the various physical properties are closely related or influence each other.
[0164] Taking these factors into account, molecular datasets, as data that are advantageous for application in multi-domain-based multitasking, can be a preferred example for illustrating the multitasking model learning method provided in an embodiment of the present invention, as well as the multitasking execution method utilizing the machine learning model learned by the method.
[0165] However, this is not a limitation; any embodiment applicable to multiple domains and multitasking can obviously be included in the embodiments of the present invention.
[0166] Hereinafter, with reference to the accompanying drawings, a multi-task processing model learning method and a multi-task processing execution method using a machine learning model learned by the method, according to an embodiment of the present invention, will be described in more detail.
[0167] Figure 8 A block diagram flowchart is shown to illustrate a multi-task processing model learning method provided in one embodiment of the present invention.
[0168] Reference Figure 8 An embodiment of the present invention provides a multi-task processing model learning method and a multi-task processing execution method using a machine learning model learned by the method, which may include: step S101, initializing the multi-task processing learning model (MtLM); step S103, acquiring experimental data; step S105, training the multi-task processing learning model (MtLM) based on the acquired experimental data; and step S107, providing the trained multi-task processing learning model (MtLM).
[0169] In detail, a computing system 1000 provided in one embodiment of the present invention can initialize a multi-task processing learning model (MtLM). (S101)
[0170] In other words, the multi-task processing learning model (MtLM) (Geometrically Aligned Transfer Encoder Model) provided by the embodiments of the present invention can be a machine learning model that aligns knowledge data (as an example, latent vectors, etc.) scattered in the latent space for each task in a comprehensive latent space (M) through geometric transfer to process multi-task outputs related to multiple domains.
[0171] That is, the Multi-Task Processing Learning Model (MtLM) provided in the embodiment not only learns knowledge data about various domains simultaneously, but also efficiently learns the relationships between the domains, thereby expanding the learning domain and performing effective multi-task processing learning that can learn the local patterns of each domain and the common principles between multiple domains at the same time.
[0172] In detail, in the embodiments, the computing system 1000 can perform initialization on the various components included in the multi-task processing learning model (MtLM) as described above.
[0173] As an example, the computing system 1000 can utilize random parameters ( Embedded networks within the Multi-Task Learning Model (MtLM) ), encoder network ( ), regressor (head) network ( ), migration network ( ) and / or reverse network ( Initialize using methods such as )
[0174] Furthermore, as an example, the computing system 1000 may be configured with a specified optimization algorithm applicable to the Multi-Task Processing Learning Model (MtLM).
[0175] For example, the computing system 1000 can set the AdamW (Decoupled Weight Decay Regularization) algorithm as an optimization algorithm, and can use the optimization algorithm to improve the weight decay process independently according to the embodiments.
[0176] Furthermore, the computing system 1000 provided in one embodiment of the present invention can acquire experimental data. (S103)
[0177] In other words, the experimental data provided by the embodiments of the present invention ( () refers to the learning data used to train the Multi-Task Processing Learning Model (MtLM), which can include information on the inherent properties of the material and specific information on the physical properties of the material.
[0178] In this case, the inherent property information of the substance provided in the embodiments may refer to information that specifically specifies the inherent properties possessed by a given substance. That is, in the embodiments, the inherent property information of the substance may be information that specifically specifies the inherent properties possessed by a given molecule.
[0179] For example, information about the inherent properties of a substance may include the specified substance name, molecular structural formula and / or chemical formula, etc.
[0180] Furthermore, the specific information on the physical properties of a substance provided in the embodiments may refer to information that specifically specifies the data values that a given substance possesses regarding a given physical property.
[0181] For example, specific information about the physical properties of a substance may include specified physical property (i.e., domain) values such as boiling point, melting point, refractive index, solubility, viscosity, surface tension, density, strength and / or thermal conductivity.
[0182] In detail, in the embodiments, the computing system 1000 can obtain the experimental data as described above based on specified user input and / or linkage with external servers.
[0183] Furthermore, in one embodiment of the present invention, the computing system 1000 can train a multi-task processing learning model (MtLM) based on the acquired experimental data. (S105)
[0184] Figure 9 A block diagram flowchart is shown to illustrate a training method for a multi-task processing learning model (MtLM) provided in one embodiment of the present invention. Figure 10 An example of a concept diagram is shown to illustrate a training method for a multi-task processing learning model (MtLM) provided in one embodiment of the present invention.
[0185] That is, refer to Figure 9 as well as Figure 10 In one embodiment, the computing system 1000 can perform pre-learning of the multi-task processing learning model (MtLM) based on the experimental data obtained as described above.
[0186] Specifically, in this embodiment, the computing system 1000 can set a training loop for the multi-task processing learning model (MtLM). (S201)
[0187] More specifically, in an embodiment, the computing system 1000 may set the number of epochs, the number of task iterations, and / or the number of batch iterations during training.
[0188] As an example, the computing system 1000 can set the training cycle in the following manner: during training, the number of rounds ´i´ is repeatedly executed ´1 to n (n≥1)´, the cycle is repeatedly executed for each task ´t´, and the cycle is repeatedly executed for each set configuration ´b´.
[0189] Furthermore, in this embodiment, the computing system 1000 can obtain a geometric alignment vector based on the experimental data obtained as described above. (S203)
[0190] In this invention, the geometric alignment vector provided in the embodiments can refer to various vectors obtained through a multi-task processing learning model (MtLM).
[0191] In an embodiment, the geometric alignment vector may include an embedding vector ( ), perturbation vector ( ), encoding vector, transfer vector, and inverse vector, etc.
[0192] In detail, in an embodiment, the computing system 1000 can input the acquired experimental data into a multi-task processing learning model (MtLM).
[0193] Furthermore, in the embodiment, the computing system 1000 can obtain the embedding vector based on the input multi-task processing learning model (MtLM), 1) with experimental data.
[0194] More specifically, the computing system 1000 can be linked with the embedding module (EBM) of the multi-task processing learning model (MtLM) to convert the input experimental data into embedding vectors through the embedding network.
[0195] Thus, the computing system 1000 is able to acquire the embedding vectors that project the experimental data into the specified embedding space and convert them into vector form.
[0196] Furthermore, in the embodiment, the computing system 1000, 2) can generate a perturbation vector based on the acquired embedding vector.
[0197] In detail, in an embodiment, the computing system 1000 may be linked with the perturbation module (PBM) of the multi-task processing learning model (MtLM) to generate multiple perturbation vectors (i.e., perturbation points) around a specified area based on the acquired embedding vector.
[0198] In this embodiment, the computing system 1000 can repeatedly perform the above-mentioned functional actions for each task and obtain the corresponding perturbation vector for each task.
[0199] As an example, the computing system 1000 can obtain the perturbation vector corresponding to task 't' and the perturbation vector corresponding to task 's'.
[0200] Furthermore, in the embodiment, the computing system 1000, 3) can obtain an encoded vector based on the generated perturbation vector and the embedding vector.
[0201] The encoding vector provided in the embodiments may include a latent vector generated according to a specified perturbation vector, i.e., a perturbation latent vector, and an original latent vector generated according to an embedding vector that is the original vector of the perturbation vector.
[0202] In detail, in an embodiment, the computing system 1000 can be linked with the encoder module of the multi-task processing learning model (MtLM) to project the generated perturbation vector into the latent space corresponding to the corresponding task through the encoder network and transform it into a latent vector.
[0203] Furthermore, in the embodiment, the computing system 1000 can be linked with the encoder module of the multi-task processing learning model (MtLM) to project the obtained embedding vectors into the latent space corresponding to the corresponding task through the encoder network and transform them into latent vectors.
[0204] In this way, in the embodiment, the computing system 1000 is able to obtain the perturbation potential vector and the original potential vector.
[0205] In this embodiment, the computing system 1000 can repeatedly perform the above-mentioned functional actions for each task and obtain the corresponding original potential vector and perturbation potential vector for each task.
[0206] As an example, the computing system 1000 can obtain the original latent vector corresponding to task 't' ( : Hereinafter referred to as the t-th original latent vector) and the perturbation latent vector corresponding to task ´t´ ( (Hereinafter referred to as the t-th perturbation potential vector).
[0207] Furthermore, the computing system 1000 can obtain the original latent vector corresponding to task 's' ( : Hereinafter referred to as the s-th original latent vector) and the perturbation latent vector corresponding to task 's' ( (Hereinafter referred to as the s-th perturbation potential vector).
[0208] Furthermore, in the embodiment, the computing system 1000 (4) can obtain a transfer vector based on the obtained encoding vector.
[0209] The migration vector provided in the embodiments may include a migration vector generated according to a specified perturbation potential vector, i.e., a perturbation migration vector, and a migration vector generated according to the original potential vector corresponding to the perturbation potential vector, i.e., an original migration vector.
[0210] In detail, in an embodiment, the computing system 1000 can work in conjunction with the transfer module (TFM) of the multi-task processing learning model (MtLM) to map the obtained perturbation latent vector and the original latent vector to the latent space of other tasks (in the embodiment, task 's' or task 't') through the transfer network, and convert them into transfer vectors.
[0211] In this way, the computing system 1000 can obtain the perturbation migration vector and the original migration vector.
[0212] In this embodiment, the computing system 1000 repeatedly performs the above-mentioned functional actions for each task and obtains the corresponding original migration vector and perturbation migration vector for each task.
[0213] As an example, the computing system 1000 can obtain the original migration vector corresponding to task 't'. : Hereinafter referred to as the t-th original transfer vector) and the perturbation transfer vector corresponding to task 't' ( (Hereinafter referred to as the t-th perturbation migration vector).
[0214] Furthermore, the computing system 1000 can obtain the original transfer vector corresponding to task 's'. : Hereinafter referred to as the s-th original transfer vector) and the perturbation transfer vector corresponding to task 's' ( (Hereinafter referred to as the s-th perturbation migration vector).
[0215] In this way, in the embodiment, the computing system 1000 is able to acquire geometric alignment vectors (i.e., embedding vectors, perturbation vectors, encoding vectors (including original latent vectors and perturbation latent vectors) and transfer vectors (including original transfer vectors and perturbation transfer vectors) based on experimental data.
[0216] Furthermore, in the embodiment, the computing system 1000 (5) can obtain the inverse vector based on the obtained migration vector.
[0217] The inverse vector provided in the embodiments may include an inverse vector generated according to a specified perturbation migration vector, i.e., a perturbation inverse vector, and an inverse vector generated according to the original migration vector corresponding to the perturbation migration vector, i.e., an original inverse vector.
[0218] In detail, in the embodiment, the computing system 1000 can be linked with the inverse transfer module (ITM) of the multi-task processing learning model (MtLM) to reconstruct the obtained perturbation transfer vector and the original transfer vector by remapping them back to the original latent space through the inverse network, and convert them into inverse vectors.
[0219] In this way, the computing system 1000 can obtain the inverse perturbation vector and the original inverse vector.
[0220] In this embodiment, the computing system 1000 can repeatedly perform the above-mentioned functional actions for each task and obtain the corresponding original inverse vector and perturbation inverse vector for each task.
[0221] As an example, the computing system 1000 can obtain the original inverse vector corresponding to task 't' ( : Hereinafter referred to as the t-th original inverse vector) and the perturbation inverse vector corresponding to task ´t´ ( (Hereinafter referred to as the t-th perturbation inverse vector).
[0222] Furthermore, the computing system 1000 can obtain the original inverse vector corresponding to task 's' ( : Hereinafter referred to as the s-th original inverse vector) and the perturbation inverse vector corresponding to task 's' ( (Hereinafter referred to as the s-th perturbation inverse vector).
[0223] In this way, in the embodiment, the computing system 1000 can acquire geometric alignment vectors (i.e., embedding vectors, perturbation vectors, encoding vectors (including original latent vectors and perturbation latent vectors), transfer vectors (including original transfer vectors and perturbation transfer vectors), and inverse vectors (including original inverse vectors and perturbation inverse vectors) based on experimental data.
[0224] Furthermore, in this embodiment, the computing system 1000 can calculate the geometric alignment loss based on the acquired geometric alignment vector. (S205)
[0225] In this invention, the geometric alignment loss provided in the embodiments can refer to various loss functions (Loss) calculated based on various vectors (i.e., geometric alignment vectors) obtained through the multi-task processing learning model (MtLM).
[0226] In an embodiment, the geometric alignment loss may include a regression loss ( Regression loss), autoencoder loss ( Autoencoder loss, consistency loss Consistency loss), mapping loss ( Mapping loss), distance loss Distance loss) and / or combined loss (e.g., integrated loss).
[0227] In the following description, for the sake of clarity, the case of calculating the geometric alignment loss based on task 't' will be explained.
[0228] Figure 11 as well as Figure 12 An example diagram is shown to illustrate a regression loss calculation method provided in one embodiment of the present invention.
[0229] In detail, refer to Figure 10 to Figure 12 In an embodiment, the computing system 1000 can calculate the regression loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0230] More specifically, in an embodiment, the computing system 1000 can calculate the predicted value based on the prediction made by the regressor module (RGM) according to the following [Equation 1]. ) and actual value ( The regression loss is calculated using the label value. The predicted value of [Equation 1] can also be expressed as ´ ´.
[0231] [Formula 1]
[0232]
[0233] That is, the computing system 1000 can calculate the regression loss by calculating the mean squared error (MSE) between the predicted value and the actual value.
[0234] In this embodiment, for each task, an independent regression loss is calculated based on the encoder module (ECM) and regressor module (RGM) matched to each task, and learning based on this is performed, thereby preventing mutual interference.
[0235] In this way, the computing system 1000 can easily evaluate the regression performance of the model by calculating the regression loss.
[0236] And, refer again Figure 10 In this embodiment, the computing system 1000 can calculate the autoencoder loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0237] In detail, in an embodiment, the computing system 1000 can calculate the autoencoder loss based on the original latent vector and the original inverse vector according to the following [Equation 2].
[0238] [Formula 2]
[0239]
[0240] That is, the computing system 1000 can calculate the autoencoder loss by calculating the mean square error (MSE) between the latent vector and the inverse vector.
[0241] In an embodiment, the computing system 1000 can improve the accuracy of the data migration process by calculating the autoencoder loss as described above.
[0242] Figure 13 An example diagram is shown for illustrating a comprehensive latent space (M) mapping method provided by an embodiment of the present invention.
[0243] On the other hand, refer to Figure 13In one embodiment, the computing system 1000 can learn a bidirectional transformation matrix (TM) that can be mapped to a common integrated latent space (M) for each task.
[0244] In detail, in an embodiment, the computing system 1000 can use knowledge data about the labels of both tasks to connect the potential space between tasks.
[0245] In this process, the computing system 1000 can calculate the consistency loss and mapping loss provided by the embodiment.
[0246] Figure 14 as well as Figure 15 An example diagram is shown to illustrate a consistency loss calculation method provided in one embodiment of the present invention.
[0247] In more detail, refer to Figure 10 , Figure 14 as well as Figure 15 In this embodiment, the computing system 1000 can calculate the consistency loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0248] Specifically, in an embodiment, the computing system 1000 can calculate the consistency loss based on the perturbation migration vector of task 't' and the perturbation migration vector of task 's' according to the following [Equation 3].
[0249] [Formula 3]
[0250]
[0251] That is, the computing system 1000 can calculate the consistency loss by calculating the mean square error (MSE) between the t-th perturbation migration vector and the s-th perturbation migration vector.
[0252] In this embodiment, the computing system 1000 can learn to derive a metric for calculating spatial distance from the transformation matrix (TM) and make the distances in the latent space of each task the same based on the derived metric.
[0253] As a result, the computing system 1000 is able to achieve geometric alignment between tasks more effectively.
[0254] Figure 16 as well as Figure 17 An example diagram is shown to illustrate a mapping loss calculation method provided in one embodiment of the present invention.
[0255] Furthermore, referring toFigure 10 , Figure 16 as well as Figure 17 In an embodiment, the computing system 1000 can calculate the mapping loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0256] In detail, in an embodiment, the computing system 1000 can calculate the mapping loss based on the actual value of task 't' and the value predicted based on the original inverse vector of task 's' according to the following [Equation 4].
[0257] [Formula 4]
[0258]
[0259] That is, the computing system 1000 can calculate the mapping loss by calculating the mean squared error (MSE) between the actual value of task 't' and the predicted value of task 's' based on the original inverse vector.
[0260] In an embodiment, the computing system 1000 can realize the transfer of latent vectors from the latent space of one task to the latent space of another task by calculating the mapping loss as described above, and perform learning of the other task based on the transferred vectors, thereby guiding the latent features to be similar to each other.
[0261] Thus, the computing system 1000 can evaluate the predictive performance of vectors transferred to the latent space of other tasks and guide learning in a direction that improves this performance.
[0262] And, refer again Figure 10 In an embodiment, the computing system 1000 can calculate the distance loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0263] In detail, in the embodiments, the computing system 1000 can, according to the following [Equation 5] and [Equation 6], and based on the distance between the original migration vector and the perturbed migration vector of each task ( (Hereinafter referred to as migration vector displacement) calculates the distance loss between tasks.
[0264] More specifically, in an embodiment, the computing system 1000 can calculate the distance between the t-th original migration vector and the t-th perturbed migration vector with respect to task ´t´ according to the following [Equation 5(a)] ( (Hereinafter referred to as the t-th migration vector displacement).
[0265] Furthermore, the computing system 1000 can calculate the distance between the s-th original migration vector and the s-th perturbation migration vector with respect to task 's' according to the following [Equation 5(b)] ( (Hereinafter referred to as the shift of the s-th migration vector).
[0266] [Formula 5]
[0267]
[0268] Furthermore, in an embodiment, the calculation system 1000 can calculate the distance loss by calculating the mean square error (MSE) between the t-th migration vector displacement and the s-th migration vector displacement according to the following [Equation 6].
[0269] [Formula 6]
[0270]
[0271] In [Equation 6], ´M´ represents the number of disturbance points.
[0272] In this embodiment, the computing system 1000 can define the t-th migration vector displacement and the s-th migration vector displacement as the displacement between the source task and the target task, respectively.
[0273] In this way, the computing system 1000 can analyze the displacements of the t-th and s-th migration vectors as lying in a flat Euclidean space, thus making it easier to calculate the distance between the original migration vector and the perturbed migration vector.
[0274] Therefore, the computing system 1000 is able to support a more complete maintenance of the consistency of the model in the latent space.
[0275] Figure 18 An example diagram is shown to illustrate a method for calculating comprehensive loss provided in one embodiment of the present invention.
[0276] Furthermore, referring to Figure 10 as well as Figure 18 In an embodiment, the computing system 1000 can calculate the integrated loss based on the multi-task processing learning model (MtLM) that has obtained the geometric alignment vector.
[0277] In detail, in an embodiment, the computing system 1000 can calculate the comprehensive loss obtained by weighted summation of the regression loss, autoencoder loss, consistency loss, mapping loss and distance loss mentioned above according to the following [Equation 7].
[0278] [Formula 7]
[0279]
[0280] In this embodiment, the computing system 1000 may apply weights to each loss function so that each loss function can be optimized for a specific aspect of the model.
[0281] Among them, [number 7]'s ´ ´ represents the weights of the autoencoder loss. ´ is the weight of the consistency loss, ´ ´ represents the weights of the mapping loss. ´ represents the weight of the distance loss.
[0282] In this embodiment, the computing system 1000 can use the weights described above to adjust the importance of the loss function corresponding to each weight during the model's learning process, thereby updating the parameters in the direction of minimizing the overall loss.
[0283] Return to Figure 9 In this embodiment, the computing system 1000 can also perform model optimization and parameter updates based on the geometric alignment loss calculated as described above. (S207)
[0284] In detail, in the embodiment, the computing system 1000 can perform optimization and parameter updates of the multi-task processing learning model (MtLM) based on the comprehensive loss described above.
[0285] As an example, the computing system 1000 can compute the gradient based on the comprehensive loss with respect to each parameter of the multi-task processing learning model (MtLM) via backpropagation.
[0286] Then, the computing system 1000 can use the calculated gradients and the pre-defined optimization algorithm (e.g., AdamW (Decoupled Weight Decay Regularization) algorithm) to perform parameter updates of the multi-task processing learning model (MtLM).
[0287] In this way, the computing system 1000 can optimize the multi-task processing learning model (MtLM) based on geometric alignment loss (especially integrated loss).
[0288] In this way, in the embodiment, the computing system 1000 is able to perform optimization and parameter update learning of the multi-task processing learning model (MtLM) by combining a variety of loss functions calculated from multiple perspectives.
[0289] At this point, each loss function can easily assist in enhancing model performance by correcting the accuracy, consistency, and / or distance of the knowledge data mapping.
[0290] Thus, the computing system 1000 can implement a multi-task processing model that provides improved performance to overcome the limitations of regression problems on small datasets and existing transfer learning techniques, and works more stably, thereby providing enhanced generalization performance.
[0291] Furthermore, in this embodiment, the computing system 1000 can terminate the training of the multi-task processing learning model (MtLM). (S209)
[0292] In detail, in the embodiment, when the set training termination condition is met, the computing system 1000 can terminate the above-mentioned multi-task processing learning model (MtLM) training process.
[0293] As an example, when the set training cycle is completed, the computing system 1000 can end the training of the multi-task processing learning model (MtLM).
[0294] Return to Figure 8 In one embodiment of the present invention, the computing system 1000 can also provide a trained multi-task processing learning model (MtLM). (S107)
[0295] That is, in the embodiment, the computing system 1000 can provide a multi-task processing learning model (MtLM) trained as described above in a prescribed manner.
[0296] As an example, the computing system 1000 can be linked with specified application services (e.g., raw material synthesis / evaluation service, material physical property prediction service, and / or optimal material recommendation service, etc.) to provide a multi-task processing learning model (MtLM) trained according to an embodiment of the present invention.
[0297] In this way, the computing system 1000 can effectively support the processing of various multitasking tasks performed using the performance-enhanced Multitasking Learning Model (MtLM).
[0298] In this way, in an embodiment, the computing system 1000, in order to process multi-tasks for making multi-domain related outputs, enables the knowledge data scattered in the latent space for each task to transfer and learn from each other through geometric alignment in a comprehensive latent space, thereby providing a multi-task processing learning model (MtLM) with improved performance and more stable operation that overcomes the regression problem of small datasets and the limitations of existing transfer learning techniques.
[0299] Therefore, the computing system 1000 can provide a transfer learning-based multi-task processing model that exhibits high generalization performance and works stably and robustly, even under conditions such as a small amount of data, multiple task types, or primarily solving regression problems.
[0300] In other words, the computing system 1000 is able to provide a multi-task processing learning model (MtLM) with enhanced predictive performance based on distillation knowledge, which is knowledge distilled through geometric alignment-based transfer learning performed in collaboration with other domains in multiple domains (physical properties in the embodiment), even in domains where experimental data (learning data) is insufficient.
[0301] For example, after the computing system 1000 pre-learns the multi-task processing learning model (MtLM) based on the first to tenth physical properties of each of the multiple molecular structures, if it receives an input of a first molecular structure containing only the first to fifth physical properties, it can more accurately predict the data values of the remaining sixth to tenth physical properties of the first molecular structure based on the knowledge data obtained through pre-learning transfer and distillation, and generate and provide output data based thereon.
[0302] In this way, the computing system 1000 provided by the embodiments of the present invention can provide a multi-task processing model that achieves effective transfer learning based on geometric alignment, ensures high generalization performance, improves prediction accuracy for regression problems, supports normalization by combining multiple loss functions, and ensures robust performance by executing stable learning procedures.
[0303] In summary, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: by transferring knowledge learned in the source task to the target task through transfer learning, the problem of insufficient data is eliminated, thereby enabling a multi-task processing model that maintains high performance even on small datasets.
[0304] Therefore, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: they can extend their application scope to fields where machine learning models are difficult to apply due to insufficient data or domain knowledge.
[0305] Furthermore, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: by providing a dedicated transfer learning technique that can be effectively applied to regression problems, it can achieve high prediction performance even on complex regression problems such as molecular datasets.
[0306] Furthermore, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: by optimizing the knowledge transfer between the source task and the target task in a Riemannian geometric manner, geometric consistency between tasks can be maintained, thereby improving the efficiency of transfer learning.
[0307] Furthermore, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: by combining multiple loss functions to normalize multiple aspects of the model, the generalization performance of the model can be further improved.
[0308] Therefore, the multi-task processing model learning method and the multi-task processing execution method using the machine learning model learned by the method provided in one embodiment of the present invention have the following effects: by providing a multi-task processing model that can be widely applied to a variety of materials (raw materials), the quality of the entire related industry can be improved.
[0309] On the other hand, the embodiments of the present invention described above can be implemented in the form of program instructions executable by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may contain program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded in the computer-readable recording medium are specifically designed and configured for the invention, but may also be program instructions known and available to those skilled in the art of computer software. Examples of computer-readable recording media include: magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical recording media, such as CD-ROMs (Compact Disc Read-Only Memory) and DVDs (Digital Video Discs); magneto-optical media, such as floppy disks; and hardware devices specifically configured to store and execute program instructions, such as ROMs (Read-Only Memory), RAMs (Random Access Memory), flash memory, etc. Examples of program instructions include not only machine language code generated by a compiler, but also high-level language code executable by a computer using an interpreter, etc. The hardware device can be changed to one or more software modules to perform the processing provided by this invention, and vice versa.
[0310] The specific implementation described in this invention is merely one embodiment and is not intended to limit the scope of the invention in any way. For the sake of simplicity, descriptions of conventional electronic structures, control systems, software, and other functional aspects of said systems may be omitted. Furthermore, the connections or connecting parts between the lines or components shown in the drawings are merely illustrative representations of functional and / or physical or electrical connections, and in actual devices, they may manifest as a wide variety of alternative or additional functional, physical, or electrical connections. Moreover, unless specifically mentioned as "essential" or "importantly," a component may not be essential to the application of this invention.
[0311] Furthermore, while the detailed description of the present invention has been provided with reference to preferred embodiments, those skilled in the art will understand that various modifications and alterations can be made to the invention without departing from the spirit and scope of the invention as set forth in the claims. Therefore, the scope of the present invention should not be limited to the contents described in the detailed specification, but rather should be defined by the claims.
[0312] [Industry Applicability]
[0313] This invention relates to a multi-task processing model learning method and a multi-task processing execution method using a machine learning model learned based on this method. Since it can be used in the artificial intelligence industry, it has industrial applicability.
Claims
1. A method of learning a multi-tasking model, the method causing a multi-tasking model to learn by a computing system including a memory and a processor, the method comprising the steps of: initializing a multi-tasking model for processing multi-tasks based on a plurality of domains; obtaining prescribed experimental data; and training the multi-tasking model based on the obtained experimental data, the step of training the multi-tasking model comprising the steps of: obtaining a geometric alignment vector that supports geometric alignment between data in a manifold based on the experimental data; calculating a geometric alignment loss based on the obtained geometric alignment vector; and updating parameters of the multi-tasking model based on the calculated geometric alignment loss.
2. The method of learning a multi-tasking model according to claim 1, wherein the step of obtaining the geometric alignment vector comprises the step of obtaining an embedding vector that is obtained by projecting the experimental data to a prescribed embedding space and converting to a vector form by an embedding module included in the multi-tasking model.
3. The method of learning a multi-tasking model according to claim 2, wherein the step of obtaining the geometric alignment vector further comprises the step of obtaining a plurality of perturbation vectors that are obtained by moving the embedding vector in a prescribed direction by a perturbation module included in the multi-tasking model.
4. The method of learning a multi-tasking model according to claim 3, wherein the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original latent vector that is obtained by projecting the embedding vector to a latent space of a first task by an encoder module included in the multi-tasking model; and obtaining a perturbation latent vector that is obtained by projecting the perturbation vector to the latent space of the first task by the encoder module included in the multi-tasking model.
5. The method of learning a multi-tasking model according to claim 4, wherein the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original transfer vector, which is a vector obtained by mapping the original latent vector into a latent space of a second task, Task2, according to a transfer module included in the multi-task processing model; and obtaining a perturbation transfer vector, which is a vector obtained by mapping the perturbation latent vector into the latent space of the second task according to the transfer module included in the multi-task processing model. 6.The multi-task processing model learning method of claim 5, wherein the step of obtaining the geometric alignment vector further comprises the steps of: obtaining an original inverse vector, which is a vector obtained by remapping the original transfer vector into the latent space of the first task according to an inverse transfer module included in the multi-task processing model; and obtaining a perturbation inverse vector, which is a vector obtained by remapping the perturbation transfer vector into the latent space of the first task according to the inverse transfer module included in the multi-task processing model. 7.The multi-task processing model learning method of claim 6, wherein the step of calculating the geometric alignment loss comprises the steps of: calculating a regression loss, an autoencoder loss, a consistency loss, a mapping loss, and a distance loss according to the geometric alignment vector; and calculating an integrated loss by weighted summing the calculated regression loss, autoencoder loss, consistency loss, mapping loss, and distance loss. 8.The multi-task processing model learning method of claim 7, wherein the step of updating the parameters of the multi-task processing model comprises a step of updating the parameters in a direction that minimizes the integrated loss. 9.The multi-task processing model learning method of claim 7, wherein the step of calculating the geometric alignment loss further comprises the steps of: obtaining a predicted value based on the latent vector according to a regressor module included in the multi-task processing model; calculating a mean squared error according to the obtained predicted value and a label value corresponding to the latent vector; and calculating the regression loss according to the calculated mean squared error. 10.The multi-task processing model learning method of claim 7, wherein The step of calculating the geometric alignment loss further includes the steps of: calculating a mean square error based on the original latent vector and the original inverse vector; and calculating the autoencoder loss based on the calculated mean square error.
11. The multi-task model learning method according to claim 7, wherein The step of calculating the geometric alignment loss further includes the steps of: calculating a mean square error based on the perturbed transfer vector obtained by mapping the latent space of the first task to the latent space of the second task and the perturbed transfer vector obtained by mapping the latent space of the second task to the latent space of the first task; and calculating the consistency loss based on the calculated mean square error.
12. The multi-task model learning method according to claim 7, wherein The step of calculating the geometric alignment loss further includes the steps of: calculating a mean square error based on the label value with respect to the first task and the predicted value based on the original inverse vector with respect to the second task; and calculating the mapping loss based on the calculated mean square error.
13. The multi-task model learning method according to claim 7, wherein The step of calculating the geometric alignment loss further includes the steps of: calculating a first transfer vector displacement, which is a distance between the original transfer vector with respect to the first task and the perturbed transfer vector; calculating a second transfer vector displacement, which is a distance between the original transfer vector with respect to the second task and the perturbed transfer vector; calculating a mean square error based on the calculated first transfer vector displacement and the second transfer vector displacement; and calculating the distance loss based on the calculated mean square error.
14. The multi-task model learning method according to claim 1, wherein The experimental data includes substance inherent property information and substance physical property specific information, the substance inherent property information is information that specifically specifies an inherent property possessed by a prescribed substance, and the substance physical property specific information is information that specifically specifies a data value possessed by a prescribed substance with respect to a prescribed physical property.
15. A multi-task model learning server comprising: at least one memory; and at least one processor that reads at least one application program stored in the memory, causing a multi-task model to learn, the instructions of the processor including instructions to perform the steps of: initializing a multi-tasking model that processes multi-tasks based on multiple domains, i.e., Domains; acquiring prescribed experimental data; acquiring a geometric alignment vector, which is a vector that supports geometric alignment between data in a manifold based on the acquired experimental data; According to the obtained geometric alignment vector, a geometric alignment loss is calculated. And According to the calculated geometric alignment loss, the parameters of the multi-task processing model are updated.