Heterogeneous Graph Embedding for Financial Transaction Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing feature learning methods, such as Node2Vec, struggle to generate effective feature embeddings for high-dimensional, categorical transaction data in financial transactions due to their inability to handle complex interrelationships between multiple parties, leading to limited applicability across different use cases and models.

Innovation Solution

The proposed solution involves generating feature embeddings using a heterogeneous graph that connects users and merchants with different characteristics, employing a meta-path feature embedding technique and a jumping probability algorithm to derive semantic sequences of transaction pairs, enabling the representation of interparty relationships and optimizing downstream prediction tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing feature learning methods like Node2Vec are used to learn from low-dimensional representations of nodes on a homogeneous graph, then the method can generate vector representations of nodes, but it cannot handle complex interrelationships between multiple parties in heterogeneous networks and cannot digest information in bi-partite or tri-partite interrelationship networks

Engineering Contradiction:
Improveapplicability across different use cases and modelsVSAvoidcomplexity of handling heterogeneous graph structures
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the heterogeneous graph into multiple homogeneous subgraphs, each representing a specific relationship type or meta-path. By dividing the complex heterogeneous structure into manageable homogeneous components, the system can apply standard embedding techniques like Node2Vec to each subgraph while preserving the ability to capture diverse interrelationships through multiple subgraphs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces meta-paths as intermediary structures that connect different node types in the heterogeneous graph. These meta-paths serve as mediators that enable the system to capture complex multi-party relationships by transforming them into sequence-based representations that can be processed by embedding algorithms designed for homogeneous structures.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If feature embeddings are learned for high-dimensional categorical transaction data, then the embeddings can represent transaction features, but a vast amount of data is required and it is not feasible for prediction tasks with small sample sets

Engineering Contradiction:
Improveprecision of feature embeddingsVSAvoidamount of data required for training
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs preliminary construction of heterogeneous graph structures and meta-paths before applying embedding algorithms. By pre-organizing the data into graph structures with defined relationships and paths, the system reduces the data hunger of embedding algorithms, as the structural information provides additional constraints and guidance that reduce the sample size needed for effective learning.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If learned feature embeddings are generated for a specific business target, then the embeddings focus on that target, but they are not applicable or shareable with other use cases

Engineering Contradiction:
Improveaccuracy for specific prediction taskVSAvoidshareability across different models and use cases
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal heterogeneous graph embedding framework that can serve multiple downstream tasks. By learning embeddings that capture general structural and relational patterns in the heterogeneous graph through meta-paths, the resulting embeddings are transferable and applicable to various prediction tasks and business targets without requiring task-specific retraining, thus achieving multi-functionality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Device complexity

If homogeneous networks are used with one set of user nodes that share the same characteristics, then the network structure is simple, but it is difficult to learn from related multi-party connections and cannot understand complex transaction features

Engineering Contradiction:
Improvesimplicity of network structureVSAvoidloss of interparty relationship information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent transitions from a single-dimension homogeneous graph to a multi-dimensional heterogeneous graph structure by introducing multiple node types (users, merchants, items) and multiple relationship types (transactions, recommendations, purchases). This dimensional expansion allows the system to preserve and learn from complex multi-party connections while maintaining a structured framework through the use of meta-paths that organize these relationships.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11816718B2Heterogeneous graph embedding
Publication Date: 2023.11.14 INTUIT INC
  • US11816718B2 patent drawing
  • US11816718B2 patent drawing
  • US11816718B2 patent drawing

AI summary

A computer-implemented system and method for generating heterogeneous graph feature embeddings for feature learning and prediction. An application server may receive and process a plurality of feature datasets to generate a graph data structure comprising a plurality of interconnected transaction pairs. The application server processes the graph data structure to determine a first-order transaction pair corresponding to a maximum transaction frequency based on a user identifier; executes a jumping probability algorithm to process the graph data structure to determine a second-order transaction pair jumping from a first-order transaction pair; and generates a transaction sequence associated with the user identifier.