Identifier Embeddings for Accurate Digital Content Connections

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional relational models for digital content analysis are inaccurate, inefficient, and inflexible, leading to wasteful computing resource usage and user interaction, as they fail to adequately capture contextual information and are tied to specific data structures.

Innovation Solution

An identifier embedding system using a dual-branched embedding machine-learning model processes character and token embeddings to generate identifier embeddings, which are then used to determine digital connections between content items, leveraging ground truth data for training and flexibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional relational models are used to analyze digital content, then the system can generate digital content suggestions, but the accuracy of predictions and suggestions deteriorates due to insufficient contextual information extraction

Engineering Contradiction:
Improveprediction accuracyVSAvoidcontextual information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transforms identifiers from traditional fixed-dimensional representations into high-dimensional dense vector embeddings. This dimensional transformation allows the model to capture nuanced contextual information and semantic relationships that were previously inaccessible, directly resolving the contradiction between prediction accuracy and contextual information retention

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system changes the parameter representation of identifiers from discrete categorical values to continuous vector embeddings with learned semantic meanings. This parameter transformation enables the model to extract and preserve contextual information while improving prediction accuracy through sophisticated pattern recognition in the embedding space

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional relational systems generate and transmit suggestions to client devices, then digital content recommendations are provided, but computing resources and system bandwidth are wasted due to inaccurate predictions

Engineering Contradiction:
Improvesuggestion generation efficiencyVSAvoidcomputing resources
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system performs preliminary action by pre-computing and storing dense vector embeddings for all identifiers in the training dataset before making predictions. This preprocessing step creates a reusable knowledge base that accelerates subsequent prediction operations, improving productivity while reducing the computing resources needed for real-time suggestion generation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates compressed vector representations (embeddings) that copy and encode the essential contextual information of identifiers in a more efficient format. These embedding copies enable rapid comparison and matching operations, significantly reducing the computational energy required to generate and transmit suggestions

Inventive Principle:
Principle #26Copying

3Ease of operation

If conventional relation systems require user interactions to locate desired digital content, then users can identify content items, but significant time and computing resources are consumed due to inaccurate suggestions requiring dozens of interactions

Engineering Contradiction:
Improvecontent location easeVSAvoiduser interaction time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces the mechanical interaction-based content discovery system with an intelligent recommendation system based on vector embedding comparisons. Instead of requiring users to manually navigate and interact with content items, the system automatically computes semantic relationships and provides accurate suggestions, dramatically reducing both user interaction time and improving ease of operation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If conventional systems utilize models tied to specific fixed data structures, then historical user selections can be analyzed, but the system becomes rigid and inflexible, failing to analyze the wide variety of available information for extracting context

Engineering Contradiction:
Improvedata structure flexibilityVSAvoidavailable information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent implements universality by designing a data structure-agnostic embedding model that can process and represent identifiers from diverse sources and formats. The vector embedding approach serves multiple functions: it captures semantic meaning, preserves contextual relationships, and enables flexible querying across different data structures, thereby increasing adaptability while preventing information loss

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250355960A1Utilizing machine-learning models to generate identifier embeddings and determine digital connections between digital content items
Publication Date: 2025.11.20 DROPBOX INC
  • US20250355960A1 patent drawing
  • US20250355960A1 patent drawing
  • US20250355960A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that utilize machine learning models to generate identifier embeddings from digital content identifiers and then leverage these identifier embeddings to determine digital connections between digital content items. In particular, the disclosed systems can utilize an embedding machine-learning model that comprises a character-level embedding machine-learning model and a word-level embedding machine-learning model. For example, the disclosed systems can combine a character embedding from the character-level embedding machine-learning model and a token embedding from the word-level embedding machine-learning model. The disclosed systems can determine digital connections between the plurality of digital content items by processing these identifier embeddings for a plurality of digital content items utilizing a content management model. Based on the digital connections, the disclosed systems can surface one or more digital content suggestions to a user interface of a client device.