Method, computer program and apparatus for cross-domain structure mapping in machine learning processes

Cross-domain structure mapping using encoder-decoder models addresses the inefficiencies of traditional methods by facilitating efficient discovery and mapping of interdisciplinary connections, allowing researchers to leverage heterogeneous data for novel discoveries.

JP7740838B2Active Publication Date: 2025-09-17INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021206780
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-12-21
Publication Date
2025-09-17
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

Cross-domain innovation is hindered by the need for expert knowledge and high costs in identifying similar patterns across different domains, making it time-consuming and inefficient for researchers to discover new assets and applications.

Method used

A method and system for cross-domain structure mapping using encoder-decoder models to correlate heterogeneous data, calculating distributional distance metrics, and updating models to improve efficiency in discovering and mapping interdisciplinary connections.

Benefits of technology

Enhances the efficiency of interdisciplinary research by streamlining source discovery and mapping, enabling researchers to discover new assets and applications across domains without extensive domain-specific knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740838000024
    Figure 0007740838000024
  • Figure 0007740838000025
    Figure 0007740838000025
  • Figure 0007740838000026
    Figure 0007740838000026
Patent Text Reader

Abstract

To provide a method for a cross-domain structural mapping in machine learning processing.SOLUTION: A method of using a computing device executing to interrelate two or more corpuses of dissimilar data includes receiving input data from each of the two or more corpuses. The device computes a pass for each of the input data into two or more encoder-decoder models. The device further obtains a prediction of an identity mapping for each of different domains of knowledge from each of the two or more models. The device computes a distribution distance metric as an output from each low-dimensional embedding vector representation from each of the two or more models. The device still further computes a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metrics. The device updates the two or more models.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The field of embodiments of the present invention relates to machine learning models and systems for cross-domain structure mapping of file types of various domain entities. [Background technology]

[0002] Many problems in industry, science, and research can potentially be solved with inspiration from other orthogonal domains, with distinct domain data available through the Internet, private networks, and collections (e.g., documents, images, videos), enabling cross-domain innovation in designing solutions to problems in domains that may have similar concepts but in different contexts. However, the capacity scalability of cross-domain innovation requires experts in various domains and strong synergy in identifying these similar patterns across various domains, both of which are time-consuming and expensive. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure has been made in consideration of the above points, and aims to provide a method, a computer program, and an apparatus for cross-domain structure mapping in machine learning processing. [Means for solving the problem]

[0004] Embodiments relate to machine learning models and systems for cross-domain structure mapping. One embodiment provides a method using a computing device to correlate two or more corpora of heterogeneous data, the method including receiving input data from each of the two or more corpora of heterogeneous data. The computing device calculates a path for each of the input data to two or more encoder-decoder models. The computing device further obtains predictions of identity mappings for each of the various domains of knowledge from each of the two or more encoder-decoder models. The computing device additionally calculates a distributional distance metric as output from each of the low-dimensional embedded vector representations from each of the two or more encoder-decoder models. The computing device further calculates a function based on each of the predictions from each of the two or more encoder-decoder models and the distributional distance metric. The computing device additionally updates the two or more encoder-decoder models. Embodiments significantly improve the efficiency with which researchers can streamline interdisciplinary source discovery and matching without significant knowledge of other external domains. Some features contribute to the benefits of discovering new assets, extracting and relating various components of documents, and mapping these to related offerings and other products. Some other features contribute to the benefits of discovering new applications of researchers' research in new domains, encouraging greater utility of their research and assets through reuse in different fields.

[0005] One or more of the following features may be included. In some embodiments, the method may further include computing, by the computing device, a corresponding reconstruction loss for each of the two or more encoder-decoder models using the individual predictions and the input data from each of the two or more corpora of heterogeneous data. The computing device may further include extracting, from each of the two or more encoder-decoder models, a low-dimensional embedding vector of the input data representation.

[0006] In some embodiments, the method may further include the distribution distance metric being a pairwise mean relative survival time (MRLT) distribution distance metric and the function being a joint loss function.

[0007] In one or more embodiments, the method may further include calculating, by the computing device, a gradient of a loss from the joint loss function with respect to the model parameters for each of the two or more encoder-decoder models.

[0008] In some embodiments, the method may additionally include initializing, by a computing device, weights for each of the two or more encoder-decoder models. The computing device further performs pre-processing, transformation, and extraction of the input data into a fixed-dimensional feature vector. The computing device further performs feedforward processing for a feedforward pass for each in-domain sample of the input data to each respective one of the two or more encoder-decoder models. The computing device additionally generates corresponding output predictions for each in-domain sample of the input data using each respective one of the two or more encoder-decoder models. The computing device additionally calculates corresponding loss values ​​for a joint loss function of each of the two or more encoder-decoder models given the in-domain sample of the input data and the corresponding output predictions.

[0009] In one or more embodiments, the method may include calculating, by the computing device, a pairwise relative survival time (RLT) distribution distance metric based on a first RLT matrix and a second RLT matrix between each of the samples in the domain of the input data, and based on using a distribution distance between the first RLT matrix and the second RLT matrix.

[0010] In some embodiments, the method may further include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the first RLT matrix and the second RLT matrix between each of the samples in the domain of the input data, and based on using a squared loss function between the outputs of the first RLT matrix and the second RLT matrix.

[0011] In one or more embodiments, the method may include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the RLT matrix and the second RLT matrix between each of the samples in the domain of the input data, and using a Wasserstein distance determination of the distribution between the first RLT matrix and the second RLT matrix.

[0012] In some embodiments, the method may include two or more corpora of heterogeneous data including text, images, audio, and other data sources in various domains of knowledge.

[0013] These and other features, aspects, and advantages of the present embodiments will become apparent with reference to the following description, appended claims, and accompanying drawings. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 10 is a diagram of an example of leveraging topological representations for comparing representations of different domains with a defined structure for jointly learning appropriate mappings between these representations, according to one embodiment. [Figure 2] FIG. 1 is a diagram of a representative example of a simplex tricomplex for different distinct components. [Figure 3A] FIG. 1 is a diagram of a two-dimensional representation of an entity embedded collection. [Figure 3B] FIG. 10 is a representative diagram of defining epsilon (sphere radius), according to one embodiment. [Figure 3C] FIG. 4 is a representation of a simplex complex construction using the example from FIGS. 3A-3B, according to one embodiment. [Figure 4] FIG. 1 is a block diagram of a flow for a machine learning model architecture for cross-domain structure mapping, according to one embodiment. [Figure 5] FIG. 1 is a block diagram of a process for cross-domain structural mapping to interrelate two or more corpora of disparate data, according to one embodiment. [Figure 6] 1 is a diagram of a cloud computing environment, according to one embodiment. [Figure 7] FIG. 1 is a diagram of a set of abstraction model layers, according to one embodiment. [Figure 8] FIG. 1 is a diagram of a network architecture of a system for cross-domain structure mapping, according to one embodiment. [Figure 9] FIG. 7 is a diagram of a representative hardware environment that may be associated with the server and / or client of FIG. 6, according to one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating a distributed system for cross-domain structure mapping, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] While the description of various embodiments has been presented for illustrative purposes, it is not intended to be exhaustive or to be limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0016] Embodiments relate to machine learning models and systems for cross-domain structure mapping. One embodiment provides a method using a computing device to correlate two or more corpora of heterogeneous data, the method including receiving input data from each of the two or more corpora of heterogeneous data. The computing device calculates a path for each of the input data to two or more encoder-decoder models. The computing device further obtains predictions of identity mappings for each of the various domains of knowledge from each of the two or more encoder-decoder models. The computing device additionally calculates a distributional distance metric as output from each of the low-dimensional embedded vector representations from each of the two or more encoder-decoder models. The computing device further calculates a function based on each of the predictions from each of the two or more encoder-decoder models and the distributional distance metric. The computing device additionally updates the two or more encoder-decoder models. One or more of the following features may be included. In some embodiments, the method may further include computing, by the computing device, a corresponding reconstruction loss for each of the two or more encoder-decoder models using the individual predictions and the input data from each of the two or more corpora of heterogeneous data. The computing device may further include extracting a low-dimensional embedding vector of the input data representation from each of the two or more encoder-decoder models.

[0017] In some embodiments, the method may further include the distribution distance metric being a pairwise mean relative survival time (MRLT) distribution distance metric and the function being a joint loss function.

[0018] In one or more embodiments, the method may further include calculating, by the computing device, a gradient of a loss from the joint loss function with respect to the model parameters for each of the two or more encoder-decoder models.

[0019] In some embodiments, the method may additionally include initializing, by the computing device, weights for each of the two or more encoder-decoder models. The computing device further performs pre-processing, transformation, and extraction of the input data into a fixed-dimensional feature vector. The computing device further performs feedforward processing for a feedforward pass for each in-domain sample of the input data to each respective one of the two or more encoder-decoder models. The computing device additionally generates corresponding output predictions for each in-domain sample of the input data using each respective one of the two or more encoder-decoder models. The computing device additionally calculates corresponding loss values ​​for a joint loss function of each of the two or more encoder-decoder models given the in-domain samples of the input data and the corresponding output predictions. In one or more embodiments, the method may include calculating, by the computing device, a pairwise MRLT distribution distance metric based on a first relative survival time (RLT) matrix and a second RLT matrix between each of the in-domain samples of the input data and based on using a distribution distance between the first RLT matrix and the second RLT matrix. In some embodiments, the method may further include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the first and second RLT matrices between each of the samples within the domain of the input data and based on using a squared loss function between the outputs of the first and second RLT matrices. In one or more embodiments, the method may include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the RLT and second RLT matrices between each of the samples within the domain of the input data and based on using a Wasserstein distance determination of the distribution between the first and second RLT matrices.In some embodiments, the method may include two or more corpora of heterogeneous data including text, images, audio, and other data sources in different knowledge domains.

[0020] One or more embodiments include a model (e.g., model architecture 900 of FIG. 4) having an autoencoder (e.g., Autoencoder Model 1 920 and Autoencoder Model 2 925 (FIG. 4)) that employs one or more artificial intelligence (AI) models. The AI ​​models can include trained machine learning models (e.g., neural networks (NNs), convolutional NNs (CNNs), recurrent NNs (RNNs), long short-term memory (LSTM)-based NNs, gated recurrent unit (GRU)-based RNNs, tree-based CNNs, self-attention networks (e.g., NNs that utilize attention mechanisms as a fundamental building block; self-attention networks have been shown to be effective for sequence modeling tasks even without regression or convolution), models such as BiLSTMs (bidirectional LSTMs), etc.). Artificial NNs are groups of interconnected nodes or neurons.

[0021] The process of cross-domain mapping stems from the application of lateral thinking, a methodology for problem-solving using indirect and creative methods that employ non-obvious, analogy-guided thought processes. In some embodiments, an example of lateral thinking applied to innovation could be a restaurant that uses a conveyor belt to serve food selections, which could be based on airport baggage carousels through the core concept of transporting objects using a conveyor belt. Another example of cross-domain innovation is a gaming controller as a user interface, which could be based on automotive innovations that use display-directed controllers to control the car's functions through the concept of an intuitive controller. Both of these examples demonstrate similar conceptual principles in different contexts.

[0022] In some embodiments, the following example use case may be approached using cross-domain structural mapping. In one example, assume that a database of journals from two orthogonal domains, e.g., neuroscience and AI, is given and someone wants to find which pairs of article subsets are similar to each other. In one embodiment, by identifying pairs of article subsets, cross-domain structural mapping can help researchers identify and enhance their literature review process by discovering new idea combinations between the two domains. The approach of embodiments significantly improves the efficiency with which researchers streamline interdisciplinary source discovery and matching without significant knowledge of other external domains. As a result, embodiments significantly improve novel discoveries and connections between two or more different fields that may not have been easily identifiable manually and can be rapidly implemented by one or more embodiments.

[0023] In a bidding process, a client submits a request for proposal (RFP). Each process has multiple documents describing client requirements, bidding processes, logistics, etc. Competitors / bidders, who are service or product providers, must strive to understand and extract relevant client requirements and bidding information. These competitors / bidders then draft proposals for offerings that meet the client requirements based on the extracted information. Traditionally, this process is handled manually, requiring specialized knowledge and is error-prone and labor-intensive. Several state-of-the-art (SOTA) methods exist for cognitively extracting requirements from RFP documents. However, SOTA methods do not jointly extract requirements from multiple sources but instead populate a single response proposal. Some embodiments provide for jointly extracting and processing various components of an RFP document to associate them and map them to related offerings and other products / services (e.g., timelines, delivery methods, etc.) that may be included in the RFP document.

[0024] In another example, given a collection of business problems or business processes (e.g., a Component Business Model (CBM) portal) and a list of code bases and research assets (e.g., an entity's existing offerings or research assets), some embodiments assist by mapping requirements from clients to existing solutions provided by developers and researchers. The embodiment approach significantly improves discovery of new assets from the client side (including new use cases for a particular solution). Another benefit of the improvement is that researchers can discover new applications of their research in new domains, increasing the high-value utility of their research and assets through reuse across different fields.

[0025] FIG. 1 illustrates an example 600 that utilizes topological representations for comparing representations of different domains with a defined structure for jointly learning appropriate mappings between these representations, according to one embodiment. In some embodiments, key components to learning and comparing inter-domain structural similarities across various domains involve learning robust representations with the following properties: estimating the quality and diversity of representations for each domain; modeling the complex, nonlinear structure (relationships between entities) within each of the entities represented in the domain; and providing a method for comparing between two independent sets of entity domains. Example 600 illustrates a first domain (Domain 1 610) P D1 (X1) and the second domain (Domain 2 650) P D2 (X2). Domain 1 610 and Domain 2 650 each contain multiple documents.

[0026] A manifold is a topological space that resembles Euclidean space near each point. A topological space can be defined as a set of points as well as a set of point-wise neighborhoods, satisfying a set of axioms related to points and neighborhoods. Each point in an n-dimensional manifold has a neighborhood that is homeomorphic to Euclidean space of dimension n. The dimension of a mathematical space (or object) can be defined as the minimum number of coordinates required to identify any point in that dimension. Some embodiments define a low-dimensional manifold M data Data concentrated in p data Using the distribution of (x), the manifold has a complex nonlinear structure; data p data The novel features and patterns of (x) can be considered in terms of the properties of the manifold, e.g., M data Assume that domains can be represented by loops and high-dimensional holes in M. One embodiment leverages such topological representations to compare different representations of the domain using the defined structure, allowing for joint learning of appropriate mappings between these representations. In example 600, Domain 1 610 is represented as Manifold 1 620 (M1), and Domain 2 650 is represented as Manifold 2 640 (M2). In one embodiment, the process of cross-domain structure mapping (as described below) minimizes the topological distance 630 between Manifold 1 620 and Manifold 2 640.

[0027] FIG. 2 shows a representative example of a simplicial 3-complex 700 for different distinct components. Since there is no direct access to a topological representation of the data, one embodiment allows for learning an approximate representation of the manifold based on samples of data present for each domain. One exemplary embodiment uses a simpler representation space, such as a simplicial complex, which is a representation constructed by points, line segments, triangles, and higher-order tetrahedra. An n-dimensional simplicial is the convex hull of n+1 many affine points. In the exemplary simplicial 3-complex 700, it can be seen that the highest dimension of the simplicial complex is three (four sides) since a tetrahedron has four corners. Therefore, in the example, we consider it a "3"-complex of a simplicial.

[0028] Figure 3A shows a representative example of an entity embedding collection 810. Using this example representation, the process of cross-domain structure mapping considers the relevant ties between the pairwise distances between each of the samples.

[0029] FIG. 3B shows a two-dimensional representation 820 of defining epsilons (sphere radii) 825, according to one embodiment. In defining epsilons (sphere radii), one epsilon (sphere radii) 825 gives a single simplicial complex, and in one embodiment, different ranges of epsilons are considered. In one embodiment, a family of simplicial complexes is considered, and for each different epsilon, the cross-domain structure mapping process quantifies the properties of the simplicial complex by evaluating the occurrence of features or homologies, such as loops and multiple holes. For different epsilon values, there are different numbers of homologies, controlling which relationships are significant versus which relationships are noise. In one embodiment, these formed k-homologies (e.g., how many k-dimensional holes exist) are appropriately ranked by looking at the persistence barcode (a graphical representation of the simplicial components in a bar graph format) by observing the intersections of each of the persistence barcode's bars (where each bar in the persistence barcode represents a component of a simplicial complex; or grouping of data).

[0030] FIG. 3C shows a representative example of simplicial complex construction 830 using the example from FIGS. 3A-3B, according to one embodiment. In one embodiment, the parameter epsilon is utilized as a process to determine the optimal sphere radius, which allows for determining the best simplicial complex (representation of intra-domain relationships) to perform such mapping to inter-domain relationships. In some embodiments, persistence barcodes are determined in an efficient manner for large datasets, and comparisons between different persistence barcodes (which are often non-trivial) are required. In one embodiment, a small subset of landmark points are used to construct the simplicial complex, but considers nearby points, called witnesses. The witness complex is

number

[0031] One challenge with traditional methods is how to learn representations that are not clear. That is, the first step in learning pairwise distances to learn relationships between entities is not scalable and does not perform well, suggesting the inclusion of traditional "end-to-end" approaches. One traditional approach proposed comparing the quality and diversity of generative adversarial networks (GANs) for generated versus real data using topological properties of low-dimensional embedding vector representations of the data as a comparison between two probability distributions. In one embodiment, the process of cross-domain structure mapping attempts to learn a (good) topological representation (which can be expressed as a probability distribution) and perform cross-domain structure mapping on topological space (instead of geometric space); it uses a persistence barcode representation to represent the use of the RLT, where the RLT is used as a metric to indicate the number of holes, or equivalently, the number of connected components found in a dataset relative to the radius of a simplex sphere. One or more embodiments contribute to the benefits of using topological features as part of the loss function for cross-domain mapping. In one embodiment, geometric scores are used as part of a joint loss function, as opposed to an auxiliary scoring function as used in the prior art. Another advantage of one or more embodiments is that they require many fewer samples to process than the prior art, making the process of cross-domain structure mapping tractable for large datasets.

[0032] FIG. 4 illustrates a block diagram of a flow for a machine learning model architecture 900 for processing cross-domain structure mapping, according to one embodiment. In one embodiment, the machine learning model architecture 900 includes a domain corpus (D1) 905, a domain corpus (D2) 906, and input data 915 (

number

number

number

number

[0033] In one embodiment, the cross-domain structure mapping process of the machine learning model architecture extracts homology estimates using the RLT of each hole observed in a simplicial complex. The RLT is the ratio of the total time (considering epsilon as the time axis) over which a homology has existed to the maximum epsilon (the point at which the complex becomes a single entity). The RLT can also assess the confidence of how well the manifold representation is approximated. To ensure robustness, in one embodiment, the RLT is modeled probabilistically by considering the mean or mean RLT (MRLT), i.e., by selecting several landmarks or landmark points (based on a scalable complex suitable for large point sets, e.g., a witness complex). Through this, a probability distribution for each homology is derived, which then enables comparison against other manifold representations by means such as L2 error or Wasserstein distance. In one embodiment, the RLT is used as part of the (parameterized) loss function process 940, as opposed to an auxiliary metric as used in the prior art.

[0034] In one embodiment, the RLT process 950 can be expressed as follows:

number

number

number

[0035] In one embodiment, the MRLT is expressed as:

number

number

number

[0036] In one embodiment, the input data (input data 915(

number

number

number

number

[0037] In one embodiment, Autoencoder Model 1 920 and Autoencoder Model 2 925 can include four parts: an encoder (e.g., Encoder 921 and Encoder 926) that learns how the autoencoder model reduces input dimensionality and compresses the input data into an encoded representation; a bottleneck, which is a layer that contains a compressed representation of the input data (the lowest possible dimensionality of the input data); a decoder (e.g., Decoder 922, Decoder 927) that learns how the model reconstructs the data from the encoded representation as close as possible to the original input; and a reconstruction loss (Reconstruction Loss Process 935, Reconstruction Loss Process 936), which is a process that measures how well the decoder is performing and how close the output is to the original input. Training Autoencoder Model 1 920 and Autoencoder 2 925 involves the use of backpropagation to minimize the reconstruction loss.

[0038] In one embodiment, the autoencoder models (Autoencoder Model 1 920 and Autoencoder Model 2 925) may each include an RNN encoder-decoder, which includes two RNNs acting as an encoder-decoder pair. The encoders (e.g., Encoder 921, Encoder 926) map variable-length source sequences into fixed-length vectors, and the decoders (e.g., Decoder 922, Decoder 927) map that vector representation back to a variable-length target sequence. In some embodiments, the autoencoder models are each unsupervised artificial neural networks that learn how to efficiently compress and encode data and then how to reconstruct the data from the reduced encoded representation back to a representation that is as close as possible to the original input.

[0039] In one embodiment, as described below, processing for the machine learning model architecture 900 involves computing feedforward passes for the autoencoder models (Autoencoder Model 1 920 and Autoencoder Model 2 925), computing the RLT of the latent embedding (RLT processing 950), computing the RLT distribution distance loss metric, and computing the joint loss function (Loss Function processing 940), gradients, and update models.

[0040] In one embodiment, computing the feedforward paths for the autoencoder models (Autoencoder Model 1 920 and Autoencoder Model 2 925) includes:

number

number

[0041] In one embodiment, the feedforward path involves N distinct domain entities (D i), the process begins by feeding each sample, for each domain, to its respective autoencoder model (Autoencoder Model 1 920 and Autoencoder Model 2 925) to generate a prediction and calculate its loss value accordingly. In some embodiments, the autoencoder models (Autoencoder Model 1 920 and Autoencoder Model 2 925) may be shallow autoencoders or sequence-to-sequence encoder-decoder autoencoder models. Given a prediction, the machine learning model architecture 900 further calculates the corresponding loss value for each of the defined models.

[0042] In one embodiment, in the feedforward pass, for each domain autoencoder model (Autoencoder Model 1 920 and Autoencoder Model 2 925), weights are initialized for the corresponding autoencoder model. Then, the input data (Input Data 915 (

number

number

[0043] In one embodiment, the computation of RLTs for latent embeddings (RLT process 950) computes an RLT for each intra-domain set of embedding vectors, which is later used to compute the distance between two RLTs for sets of intra-domain embedding vectors. In one embodiment, the computation of RLTs for latent embeddings takes as input: a set X of embedding vector representations for each of the items in set D; L0, the number of landmarks to use; and an upper epsilon value (α max ) coefficient γ; parameter i for determining the upper persistence interval max the number of iterations n; the distance function dist(a,b) used to calculate the distance between sample a and sample b; the pairwise distance d, the maximum persistence value α, and a function witness(d,α,k) to calculate the family witness of complexes of maximal dimension for simplex k, and a function persistence(w,k) to calculate the persistence interval for families of dimension k. In one embodiment, the computation of the RLT of a latent embedding is performed using an n×i matrix that includes the RLT measurements as output. max Contains matrices of

[0044] In one embodiment, the computation of the RLT of the latent embedding is performed on a matrix of dimensions n×i to store the resulting computation of the RLT. max Then, the process randomly selects an embedding vector representation L0 from X and assigns it to L. Next, given L and X, the process calculates the given distance function using the defined distance metric dist(L,X) and assigns it as d. The maximum epsilon size is α max d and α max Then, the witness of the complex is given by the witness function witness(d,α max,2) and the output is assigned to W. Given W, the persistence value is calculated by calculating persistence(W,1) and assigned as I. Using the values ​​calculated above, the RLT (using RLT(i,X,L,K)) is calculated for each sample and [0,i max ], the RLT matrix is ​​populated.

[0045] In one embodiment, to calculate the RLT distribution distance loss metric, the input includes the distribution of RLT matrices from each intra-domain RLT matrix from different data sources, and the output is a measure that quantifies the distance between a first intra-domain RLT and a second intra-domain RLT. In this portion of the process for the machine learning model architecture 900, the distance between the two distributions of MRLTs from each intra-domain dataset is determined, and this distance is then used as part of the loss function process 940 that the machine learning model architecture 900 process minimizes.

[0046] In one embodiment, calculating the RLT distributed distance loss metric includes: (1) ) and the second RLT matrix RLT(X (2) ), we can compute the distance metric between two intra-domain embedded datasets. In one embodiment, the process is (1) ) and RLT(X (2) ) is used. In an alternative embodiment, RLT(X (1) ) and RLT(X (2) ) is implemented as a squared loss function between the outputs of

[0047] In one embodiment, to calculate the joint loss function (loss function process 940), gradients, and update the machine learning model architecture 900, the inputs include the reconstruction loss functions from each of the autoencoder models (Autoencoder Model 1 920 and Autoencoder Model 2 925) and the distance functions from each of the RLTs of the in-domain embedding vectors. The output includes the joint loss function from the machine learning model architecture 900 (from each of the different components) and the updated machine learning model architecture 900 parameters from the backpropagation process. In this part of the process for the machine learning model architecture 900, the joint loss function process 940 is determined from the calculations of all previously calculated components, followed by a stochastic gradient descent (SGD) process to update the machine learning model architecture 900 parameters. The process jointly optimizes both the reconstruction loss and the loss from the RLT metric distance. The SGD process includes an iterative process to optimize an objective function with favorable smoothness properties.

[0048] In one embodiment, given each of the loss functions from the reconstruction losses (reconstruction loss process 935, reconstruction loss process 936) from the autoencoder models (autoencoder model 1 920 and autoencoder model 2 925), and the distance function between the intra-domain embedding vectors RLT, each of the components are added together as a single loss value to compute the combined loss function (loss function process 940), gradients, and update the machine learning model architecture 900. The process then computes gradients and performs SGD processing to update the parameters of the machine learning model architecture 900 relative to the original encoder-decoder models (autoencoder model 1 920 and autoencoder model 2 925) of the machine learning model architecture 900.

[0049] One or more embodiments may apply the processing of the machine learning model architecture 900 to a wide range of industrial technologies as a usable technique to further accelerate and enhance research and innovation workflows. Embodiments may introduce techniques to enhance and bridge the gap between classical and symbolic methods of AI with deep learning methods to advance SOTA in a hybrid manner.

[0050] FIG. 5 illustrates a block diagram of a process 1000 for cross-domain structure mapping to correlate two or more corpora of heterogeneous data, according to one embodiment. In one embodiment, at block 1010, the process 1000 receives access to two or more corpora of heterogeneous data (e.g., domain corpus (D1) 905, domain corpus (D2) 906, FIG. 4 ) by a computing device (e.g., computing node 10 of FIG. 6 , hardware and software layer 60 of FIG. 7 , processing system 300 of FIG. 8 , system 400 of FIG. 9 , system 500 of FIG. 10 , machine learning model architecture 900 of FIG. 4 ). At block 1020, the process 1000 further provides for encoding, by the computing device, each of the two or more corpora of heterogeneous data (e.g., using encoder 921 of autoencoder model 1 920 and encoder 926 of autoencoder model 2 925, FIG. 4 ). At block 1030, process 1000 further provides for calculating, by a computing device, entity similarities in each of the two or more corpora of heterogeneous data (e.g., using RLT process 950 of FIG. 4). At block 1040, process 1000 additionally provides for generating, by a computing device, a corpus of interrelated entities based on the entity similarities calculated in each of the two or more corpora of heterogeneous data.

[0051] In one embodiment, the process 1000 may further include the feature that the two or more corpora of heterogeneous data are scientific journal articles from different domains of knowledge.

[0052] In one embodiment, the process 1000 may additionally include the feature of receiving input data from each of two or more corpora of heterogeneous data. The computing device further computes respective feedforward passes of the input data to two or more encoder-decoder models (e.g., autoencoder model 1 920 and autoencoder model 2 925 of FIG. 4). The computing device further obtains predictions of identity maps for each of the various domains of knowledge from each of the two or more encoder-decoder models. The computing device additionally generates individual predictions (e.g., output predictions 930 (

number

number

[0053] In one embodiment, process 1000 may further include the feature of extracting, by a computing device, a low-dimensional embedding vector of the input data representation (e.g., from domain 1 embedding vector processing 960 and domain 2 embedding vector processing 965 of FIG. 4 ) from each of the two or more encoder-decoder models. The computing device further calculates, as output, a pairwise MRLT distribution distance metric from each of the low-dimensional embedding vector representations from each of the two or more encoder-decoder models. The computing device further calculates a joint loss function (e.g., via loss function processing 940) based on each of the predictions from each of the two or more encoder-decoder models and the pairwise MRLT distribution distance metric. The computing device additionally calculates a gradient of the loss from the joint loss function with respect to the model parameters for each of the two or more encoder-decoder models. The computing device further updates the two or more encoder-decoder models.

[0054] In one embodiment, process 1000 may still additionally include initializing, by a computing device, weights for each of the two or more encoder-decoder models. The computing device further performs pre-processing, transformation, and extraction of the input data into a fixed-dimensional feature vector. The computing device further performs feedforward processing for a feedforward pass for each in-domain sample of the input data to each respective one of the two or more encoder-decoder models. The computing device additionally generates a corresponding output prediction (e.g., output prediction 930(

number

number

[0055] In one embodiment, the process 1000 may further include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the first RLT matrix and the second RLT matrix between each of the samples in the domain of the input data and based on using the distribution distance between the first RLT matrix and the second RLT matrix.

[0056] In one embodiment, the process 1000 may further include calculating, by the computing device, a pairwise MRLT distribution distance metric based on the first RLT matrix and the second RLT matrix between each of the samples in the domain of the input data and based on using a squared loss function between the outputs of the first RLT matrix and the second RLT matrix.

[0057] Although this disclosure includes a detailed description of cloud computing, it should be understood in advance that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.

[0058] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines (VMs), and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0059] Its features are as follows:

[0060] On-demand self-service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed without requiring human interaction with the service provider.

[0061] Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

[0062] Resource Pool: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. There is a sense of location independence in that consumers generally have no control or information about the exact location of the resources provided, although location may be specified at a higher level of abstraction (e.g., country, state, or data center).

[0063] Rapid Scalability: Capacity can be provisioned quickly and scalably, in some cases automatically, quickly scaled out, and quickly released and quickly scaled in. To the consumer, the capacity available for provisioning often appears unlimited, and any amount can be purchased at any time.

[0064] Service Metering: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at several levels of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active consumer accounts). Resource usage can be monitored, controlled, and reported, thereby providing transparency to both providers and consumers of the services being utilized.

[0065] The service model is as follows:

[0066] Software as a Service (SaaS): The capability offered to the consumer is the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the exception of limited consumer-specific application configuration settings.

[0067] Platform as a Service (PaaS): The capability offered to consumers is the ability to deploy consumer-created or acquired applications written using programming languages ​​and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application host environment configuration.

[0068] Infrastructure as a Service (IaaS): The ability offered to consumers is the ability to provision processing, storage, network, and other basic computing resources on which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating system, storage, deployed applications, and possibly limited control over select networking components (e.g., host firewalls).

[0069] The deployment model is as follows:

[0070] Private Cloud: Cloud infrastructure is operated exclusively for an organization, can be managed by that organization or a third party, and can exist on-premise or off-premise.

[0071] Community Cloud: Cloud infrastructure is shared by several organizations to support a specific community with shared objectives (e.g., mission, security requirements, policies, and compliance concerns). It may be managed by those organizations or a third party and can exist on-premises or off-premises.

[0072] Public Cloud: Cloud infrastructure is made available to the general public or large industry organizations and is owned by an organization that sells cloud services.

[0073] Hybrid Cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a unique entity, but are joined by standardized or proprietary technologies that allow data and application portability (e.g., cloud bursting for load balancing between clouds).

[0074] Cloud computing environments are service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0075] Referring now to FIG. 6, an exemplary cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 comprises one or more cloud computing nodes 10 that can communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N, or a combination thereof. The nodes 10 can communicate with each other. They can be grouped physically or virtually in one or more networks (not shown), such as private, community, public, or hybrid clouds, or a combination thereof, as described herein. This enables the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service, eliminating the need for cloud consumers to maintain resources on their local computing devices. It will be understood that the types of computing devices 54A-54N shown in FIG. 6 are intended to be exemplary only, and that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device over any type of network and / or network-addressable connection (e.g., using a web browser).

[0076] Referring to Figure 7, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 6) is shown. It should be understood at the outset that the components, layers, and functions shown in Figure 7 are intended to be merely exemplary, and the present embodiment is not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0077] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframes 61, reduced instruction set computer (RISC) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0078] The virtualization layer 70 provides an abstraction layer that can provide the following examples of virtual entities: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0079] In one example, the management layer 80 can provide the following functions: Resource provisioning 81 provides dynamic procurement of computing resources and other resources utilized to perform tasks within the cloud computing environment. Metering and billing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources can include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides allocation and management of cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides advance agreement on and procurement of cloud computing resources in anticipation of future demand according to SLAs.

[0080] The workload layer 90 provides examples of functionality for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom instruction delivery 93; data analysis processing 94; transaction processing 95; and cross-domain structure mapping processing 96 (see, e.g., system 500 of FIG. 10, machine learning model architecture for cross-domain structure mapping 900 of FIG. 4, and process 1000 of FIG. 5). As noted above, all of the foregoing examples with respect to FIG. 7 are merely illustrative, and embodiments are not limited to these examples.

[0081] Again, although this disclosure includes detailed descriptions of cloud computing, implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments may be implemented in any other type of clustered computing environment now known or later developed.

[0082] 8 illustrates a network architecture of a system 300 for cross-domain structure mapping, according to one embodiment. As shown in FIG. 8, multiple remote networks 302 are provided, including a first remote network 304 and a second remote network 306. A gateway 301 may be coupled between the remote network 302 and a proximal network 308. In the context of this network architecture 300, the networks 304, 306 may each take any form, including, but not limited to, a LAN, a WAN such as the Internet, a public switched telephone network (PSTN), an internal telephone network, etc.

[0083] In use, gateway 301 acts as an entrance point from remote network 302 to proximal network 308. As such, gateway 301 can function as a router that can direct a given packet of data arriving at gateway 301, as well as a switch that attaches the actual path to and from gateway 301 for a given packet.

[0084] Further included is at least one data server 314 coupled to the proximal network 308, which is accessible from the remote network 302 via the gateway 301. Note that the data server 314 can include any type of computing device / groupware. Coupled to each data server 314 are a plurality of user devices 316. Such user devices 316 can include desktop computers, laptop computers, handheld computers, printers, or any other type of logic-embedded device, or combinations thereof. Note that the user devices 316 can also be directly coupled to one of the networks in some embodiments.

[0085] A peripheral device 320 or a series of peripheral devices 320, such as a facsimile machine, a printer, a scanner, a hard disk drive, a storage unit or system that is networked or local, or both, may be coupled to one or more of the networks 304, 306, 308. It should be noted that databases and / or additional components may be utilized with or integrated into any type of network element coupled to the networks 304, 306, 308. In the context of this description, a network element may refer to any component of a network.

[0086] According to some approaches, the methods and systems described herein may be implemented using and / or on virtual systems and / or systems that emulate one or more other systems, such as a UNIX system emulating an IBM® z / OS environment, a UNIX system virtually hosting a MICROSOFT® WINDOWS® environment, a MICROSOFT® WINDOWS® system emulating an IBM® z / OS environment, etc. This virtualization and / or emulation may be implemented in some embodiments through the use of VMWARE® software.

[0087] 9 illustrates a representative hardware system 400 environment associated with the user device 316 and / or server 314 of FIG. 8 , according to one embodiment. In one example, the hardware configuration includes a workstation having a central processing unit 410, such as a microprocessor, and several other units interconnected via a system bus 412. The workstation illustrated in FIG. 9 may include an I / O adapter 418 for connecting peripheral devices, such as random access memory (RAM) 414, read-only memory (ROM) 416, and disk storage 420, to the bus 412; a user interface adapter 422 for connecting a keyboard 424, a mouse 426, speakers 428, a microphone 432, and / or other user interface devices, such as a touchscreen or a digital camera (not shown), to the bus 412; a communications adapter 434 for connecting the workstation to a communications network (e.g., a data processing network) 435; and a display adapter 436 for connecting the bus 412 to a display device 438.

[0088] In one example, a workstation may have an operating system resident thereon, such as the MICROSOFT® WINDOWS® Operating System (OS), MAC OS®, or UNIX® OS. In one embodiment, system 400 employs a POSIX®-based file system. It should be understood that other examples may be implemented on platforms and operating systems other than those mentioned. Such examples may include an operating system written in an object-oriented programming style using JAVA®, XML, C, C++, or other programming languages, or a combination thereof. Object-oriented programming (OOP), which is rapidly becoming used to develop complex applications, may also be used.

[0089] 10 is a block diagram illustrating a distributed system 500 for cross-domain structure mapping, according to one embodiment. In one embodiment, the system 500 includes a client device 510 (e.g., a mobile device, a smart device, a computing system, etc.), a cloud or resource sharing environment 520 (e.g., a public cloud computing environment, a private cloud computing environment, a data center, etc.), and a server 530. In one embodiment, the client device 510 is provided with cloud services from the server 530 through the cloud or resource sharing environment 520.

[0090] One or more embodiments may be a system, method, and / or computer program product at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to perform aspects of the embodiments.

[0091] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or grooved structures having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being ephemeral signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted over electrical wires.

[0092] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to an individual computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the individual computing / processing device.

[0093] The computer-readable program instructions for carrying out the operations of the present embodiments may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute computer readable program instructions to individualize the electronic circuitry by utilizing state information in the computer readable program instructions to implement aspects of the present embodiments.

[0094] Aspects of the present embodiments are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0095] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium, capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0096] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operable steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process that executes on a computer, other programmable apparatus, or other device to implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0097] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be performed as a single step, or may be executed concurrently, substantially concurrently, partially, or fully in a time-overlapping manner, or the blocks may sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0098] References to elements in the claims in the singular are not intended to mean "one and only one" unless expressly so stated, but rather "one or more." All structural and functional equivalents to the elements of the exemplary embodiments described above, now known or later known to those of ordinary skill in the art, are intended to be encompassed by the claims. No claim element herein shall be construed under the provisions of 35 U.S.C. § 112, paragraph 6, unless the element is expressly recited using the phrase "means for" or "step for."

[0099] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprise" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups or combinations thereof.

[0100] In the following claims, equivalent structures, materials, acts, and all equivalents of means- or step-plus-function elements are intended to include any structure, material, or acts for performing the function as specifically claimed in combination with other claimed elements. The description of the present embodiments has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments. The embodiments have been chosen and described in order to best explain the principles and practical applications of the present embodiments, and may enable others skilled in the art to appreciate the present embodiments for various embodiments with various modifications suited to the particular uses contemplated. [Explanation of symbols]

[0101] 600 examples 601 Domain 1 620 Manifold 1 630 phase distance 640 Manifold 2 650 Domain 2 700 Simplex 3 Complex 810 Entity Embedded Collections 820 Two-dimensional representative example 825 epsilon (sphere radius) 830 Simplicial Complex Construction 900 ML Model Architecture 905 Domain Corpus (D1) 906 Domain Corpus (D2) 915 Input Data 916 input data 920 Autoencoder Model 1 925 Autoencoder Model 2 926 Encoder 927 decoder 930 Output Forecast 931 Output Forecast 935 Reconstruction Loss Treatment 936 Reconstruction Loss Treatment 940 Loss Function Processing 945 Distribution Distance Metric Processing 950 RLT Processing 960 Domain 1 Embedded Vector Processing 965 Domain 2 Embedded Vector Processing Maximum alpha persistence a sample b Sample d pairwise distance i Persistence Interval k dimensions X dataset L Landmark μ average calculation p data (x) Data M data low-dimensional manifolds

Claims

1. 1. A method of using a computing device to correlate two or more corpora of heterogeneous data, comprising: receiving input data from each of two or more corpora of heterogeneous data; computing, by the computing device, respective passes of the input data through two or more encoder-decoder models; obtaining, by the computing device, predictions of identity mappings for each of the different domains of knowledge from each of the two or more encoder-decoder models; calculating, by the computing device, a distribution distance metric based on features occurring in a simplicial complex constructed over a range of parameters using each of the low-dimensional embedded vector representations from each of the two or more encoder-decoder models; calculating, by the computing device, a joint loss function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating, by the computing device, the two or more encoder-decoder models; A method comprising:

2. calculating, by the computing device, a corresponding reconstruction loss for each of the two or more encoder-decoder models using each of the predictions and the input data from each of the two or more corpora of heterogeneous data; extracting, by the computing device, a low-dimensional embedding vector of an input data representation from each of the two or more encoder-decoder models; The method of claim 1 further comprising:

3. A method described in claim 1 or 2, wherein the distribution distance metric measures the pairwise distance of the distribution of relative survival times, which indicates the length of time over which the observed feature exists within the range of the parameter.

4. A method using a computing device implemented to correlate two or more corpora of heterogeneous data, comprising: receiving input data from each of two or more corpora of heterogeneous data; computing, by the computing device, respective passes of the input data through two or more encoder-decoder models; obtaining, by the computing device, predictions of identity mappings for each of the different domains of knowledge from each of the two or more encoder-decoder models; computing, by the computing device, a distribution distance metric as output from each of the low-dimensional embedded vector representations from each of the two or more encoder-decoder models; calculating, by the computing device, a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating, by the computing device, the two or more encoder-decoder models; wherein the distribution distance metric is a pairwise mean relative survival time distribution distance metric and the function is a joint loss function.

5. calculating, by the computing device, a gradient of a loss from the joint loss function with respect to model parameters for each of the two or more encoder-decoder models.

5. The method of claim 3 or 4, further comprising:

6. initializing, by the computing device, weights for each of the two or more encoder-decoder models; performing, by said computing device, pre-processing, transformation and extraction of said input data into a fixed dimensional feature vector; performing, by the computing device, feedforward processing for a feedforward path for each in-domain sample of the input data to a respective one of the two or more encoder-decoder models; generating, by the computing device, a corresponding output prediction for each of the in-domain samples of the input data using a respective one of the two or more encoder-decoder models; calculating, by the computing device, a corresponding loss value for the joint loss function of each of the two or more encoder-decoder models given the in-domain samples of the input data and the corresponding output predictions; The method of any one of claims 3 to 5, further comprising:

7. calculating, by the computing device, the distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples in the domain of the input data and based on using a distribution distance between the two relative survival time metrics defined between the first relative survival time matrix and the second relative survival time matrix. The method of any one of claims 3 to 6, further comprising:

8. calculating, by the computing device, the distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples in the domain of the input data and based on using a squared loss function between outputs of the first relative survival time matrix and the second relative survival time matrix. The method of any one of claims 3 to 6, further comprising:

9. calculating, by the computing device, the distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples in the domain of the input data and based on using a Wasserstein distance determination of the distance between the first relative survival time matrix and the second relative survival time matrix. The method of any one of claims 3 to 6, further comprising:

10. 1. A computer program for interrelating two or more corpora of heterogeneous data, the computer program comprising: receiving input data from each of two or more corpora of heterogeneous data; Computing respective passes of the input data through two or more encoder-decoder models; obtaining a prediction of an identity mapping for each of the various domains of knowledge from each of the two or more encoder-decoder models; computing a distribution distance metric based on features occurring in a simplicial complex constructed over a range of parameters using each of the low-dimensional embedding vector representations from each of the two or more encoder-decoder models; calculating a joint loss function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating the two or more encoder-decoder models; A computer program for performing the following.

11. The computer program further causes the processor to: calculating a corresponding reconstruction loss for each of the two or more encoder-decoder models using each of the predictions and the input data from each of the two or more corpora of heterogeneous data; extracting a low-dimensional embedding vector of an input data representation from each of the two or more encoder-decoder models; This is to further 11. The computer program product of claim 10, wherein the two or more corpora of heterogeneous data include text, images, audio, and other data sources in different domains of knowledge.

12. A computer program for interrelating two or more corpora of heterogeneous data, comprising: receiving input data from each of two or more corpora of heterogeneous data; Computing respective passes of the input data through two or more encoder-decoder models; obtaining a prediction of an identity mapping for each of the various domains of knowledge from each of the two or more encoder-decoder models; computing, as an output, a distribution distance metric from each of the low-dimensional embedding vector representations from each of the two or more encoder-decoder models; calculating a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating the two or more encoder-decoder models; a computer program for causing the processor to perform the following, calculating a gradient of a loss from the joint loss function with respect to model parameters for each of the two or more encoder-decoder models; Further, The computer program product, wherein the distribution distance metric is a pairwise mean relative survival time distribution distance metric and the function is the joint loss function.

13. The computer program causes the processor to: initializing weights for each of the two or more encoder-decoder models; performing pre-processing, transformation and extraction of said input data into fixed dimensional feature vectors; performing feedforward processing for a feedforward pass for each in-domain sample of the input data to a respective one of the two or more encoder-decoder models; generating a corresponding output prediction for each of the in-domain samples of the input data using a respective one of the two or more encoder-decoder models; given the in-domain samples of the input data and the corresponding output predictions, calculating a corresponding loss value for the joint loss function of each of the two or more encoder-decoder models; 13. The computer program product of claim 12, further comprising:

14. The computer program causes the processor to: calculating the pairwise average relative survival time distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples within the domain of the input data and based on using a distribution of distances between the two relative survival time metrics defined between the first relative survival time matrix and the second relative survival time matrix; 14. The computer program product according to claim 12 or 13, further comprising:

15. The computer program causes the processor to: calculating, by the processor, the pairwise mean relative survival time distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples within the domain of the input data and based on using a squared loss function between the outputs of the first relative survival time matrix and the second relative survival time matrix.

14. The computer program product according to claim 12 or 13, further comprising:

16. The computer program causes the processor to: calculating the pairwise average relative survival time distribution distance metric based on a first relative survival time matrix and a second relative survival time matrix between each of the samples within the domain of the input data and based on using a Wasserstein distance determination of the distribution between the first relative survival time matrix and the second relative survival time matrix; 14. The computer program product according to claim 12 or 13, further comprising:

17. 1. An apparatus comprising: a memory configured to store instructions; Processor and wherein the processor: receiving input data from each of two or more corpora of heterogeneous data; Computing respective passes of the input data through two or more encoder-decoder models; obtaining a prediction of an identity mapping for each of the various domains of knowledge from each of the two or more encoder-decoder models; computing a distribution distance metric based on features occurring in a simplicial complex constructed over a range of parameters using each of the low-dimensional embedding vector representations from each of the two or more encoder-decoder models; calculating a joint loss function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating the two or more encoder-decoder models; an apparatus configured to execute the instructions to:

18. the processor: calculating a corresponding reconstruction loss for each of the two or more encoder-decoder models using each of the predictions and the input data from each of the two or more corpora of heterogeneous data; extracting a low-dimensional embedding vector of an input data representation from each of the two or more encoder-decoder models; and further configured to execute the instructions to:

20. The apparatus of claim 17, wherein the two or more corpora of heterogeneous data include text, images, audio, and other data sources in different domains of knowledge.

19. An apparatus comprising: a memory configured to store instructions; Processor and wherein the processor: receiving input data from each of two or more corpora of heterogeneous data; Computing respective passes of the input data through two or more encoder-decoder models; obtaining a prediction of an identity mapping for each of the various domains of knowledge from each of the two or more encoder-decoder models; computing, as an output, a distribution distance metric from each of the low-dimensional embedded vector representations from each of the two or more encoder-decoder models; calculating a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metric; updating the two or more encoder-decoder models; and wherein the processor is configured to execute the instructions to: calculating a gradient of a loss from the joint loss function with respect to model parameters for each of the two or more encoder-decoder models; and further configured to execute the instructions to: the distribution distance metric is a pairwise mean relative survival time distribution distance metric, and the function is the joint loss function.

20. the processor: initializing weights for each of the two or more encoder-decoder models; performing pre-processing, transformation and extraction of said input data into fixed dimensional feature vectors; performing feedforward processing for an in-domain sample-by-sample feedforward pass of the input data to each respective one of the two or more encoder-decoder models; generating a corresponding output prediction for each of the in-domain samples of the input data using a respective one of the two or more encoder-decoder models; Given the in-domain samples of the input data and the corresponding output predictions, calculating a corresponding loss value for the joint loss function of each of the two or more encoder-decoder models; based on a first relative survival time matrix and a second relative survival time matrix between each of the intra-domain samples of the input data; and using a distribution distance between the two relative survival time metrics defined between the first relative survival time matrix and the second relative survival time matrix; using a squared loss function between the outputs of the first relative survival time matrix and the second relative survival time matrix; or using a Wasserstein distance determination of the distribution between the first relative survival time matrix and the second relative survival time matrix. calculating the pairwise mean relative survival time distribution distance metric based on one or more of:

20. The apparatus of claim 19, further configured to execute the instructions to:

Citation Information

Patent Citations

  • Method and apparatus for multimodal prediction using trained statistical models

    JP2021526259A