Phrase-Based Unstructured Content Parsing for CDN Edge Caching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current content delivery networks (CDNs) face challenges in optimizing content positioning at access points due to the increasing number of access points and varying user relevance over time, as existing techniques do not account for user-specific relevance or changing user interests, leading to inefficiencies in data caching and delivery.

Innovation Solution

A method that generates a forward materialization graph from unstructured content using an ontology, converts it into vector representations, computes inference paths, clusters based on relevance features, and places a structured representation at edge locations within the CDN to pre-position relevant data for anticipated user needs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of access points in CDN is increased to serve smaller geographic areas, then content delivery precision and user relevance are improved, but storage capacity at each access point is reduced and system complexity increases

Engineering Contradiction:
Improvecontent delivery precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments unstructured content into structured representations by extracting entities, relationships, and metadata. This segmentation allows precise content identification and routing to appropriate edge locations without requiring complex manual classification systems at each access point

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing system that converts unstructured content into structured representations with metadata tags. This intermediary layer handles the complexity of content analysis centrally, allowing simple access points to efficiently deliver content based on structured queries without needing complex local decision-making capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If content is cached at edge data centers closer to end users, then latency is reduced and content delivery speed is improved, but determining which content to cache requires more precision due to limited storage capacity

Engineering Contradiction:
Improvecontent delivery speedVSAvoidcontent selection precision
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by converting content to structured representations and pre-tagging with metadata before distribution to edge locations. This advance structuring enables rapid content selection at edge points without requiring complex real-time analysis, allowing fast content delivery while maintaining high selection precision through pre-computed structured data

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameters of content representation from unstructured format to structured format with extracted entities, relationships, and metadata. This parameter transformation enables precise content identification and matching at edge locations, improving both content selection precision and delivery speed by allowing efficient querying based on structured attributes

Inventive Principle:
Principle #35Parameter changes

3Productivity

If unstructured content is stored and delivered directly, then storage simplicity is maintained, but content relevance to user needs cannot be determined and delivery efficiency is reduced

Engineering Contradiction:
Improvedelivery efficiencyVSAvoidcontent processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments unstructured content into structured representations by extracting entities, relationships, and metadata components. This segmentation enables efficient content matching and delivery by allowing the system to query and route content based on specific extracted attributes without processing entire unstructured documents, thereby improving delivery efficiency while managing processing complexity through systematic decomposition

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12061627B2Phrase based unstructured content parsing
Publication Date: 2024.08.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12061627B2 patent drawing
  • US12061627B2 patent drawing
  • US12061627B2 patent drawing

AI summary

From an unstructured content using an ontology, a forward materialization graph is generated. The forward materialization graph is converted to a set of vector representations comprising multidimensional numbers representing elements of the forward materialization graph. A set of inference paths is computed for the set of vector representations. An inference path in the set of inference paths connecting a first vector representation with a second vector representation. Based on a set of features, the set of vector representations is formed into clusters, a feature in the set of features comprising a relevance probability, the relevance probability corresponding to a relevance of a portion of the unstructured content according to a relevance metric. A structured representation of the unstructured content is placed at an edge location of a content delivery network determined using the set of clusters.