Phrase-Based Unstructured Content Parsing for CDN Edge Caching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current content delivery networks (CDNs) face challenges in optimizing content positioning at access points due to the increasing number of access points and varying user relevance over time, as existing techniques do not account for user-specific relevance or changing user interests, leading to inefficiencies in data caching and delivery.
Innovation Solution
A method that generates a forward materialization graph from unstructured content using an ontology, converts it into vector representations, computes inference paths, clusters based on relevance features, and places a structured representation at edge locations within the CDN to pre-position relevant data for anticipated user needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of access points in CDN is increased to serve smaller geographic areas, then content delivery precision and user relevance are improved, but storage capacity at each access point is reduced and system complexity increases
Solution Approach 1:
The patent segments unstructured content into structured representations by extracting entities, relationships, and metadata. This segmentation allows precise content identification and routing to appropriate edge locations without requiring complex manual classification systems at each access point
Solution Approach 2:
The patent introduces an intermediary processing system that converts unstructured content into structured representations with metadata tags. This intermediary layer handles the complexity of content analysis centrally, allowing simple access points to efficiently deliver content based on structured queries without needing complex local decision-making capabilities
2Speed
If content is cached at edge data centers closer to end users, then latency is reduced and content delivery speed is improved, but determining which content to cache requires more precision due to limited storage capacity
Solution Approach 1:
The patent performs preliminary action by converting content to structured representations and pre-tagging with metadata before distribution to edge locations. This advance structuring enables rapid content selection at edge points without requiring complex real-time analysis, allowing fast content delivery while maintaining high selection precision through pre-computed structured data
Solution Approach 2:
The patent changes the parameters of content representation from unstructured format to structured format with extracted entities, relationships, and metadata. This parameter transformation enables precise content identification and matching at edge locations, improving both content selection precision and delivery speed by allowing efficient querying based on structured attributes
3Productivity
If unstructured content is stored and delivered directly, then storage simplicity is maintained, but content relevance to user needs cannot be determined and delivery efficiency is reduced
Solution Approach 1:
The patent segments unstructured content into structured representations by extracting entities, relationships, and metadata components. This segmentation enables efficient content matching and delivery by allowing the system to query and route content based on specific extracted attributes without processing entire unstructured documents, thereby improving delivery efficiency while managing processing complexity through systematic decomposition
Data Source
AI summary
From an unstructured content using an ontology, a forward materialization graph is generated. The forward materialization graph is converted to a set of vector representations comprising multidimensional numbers representing elements of the forward materialization graph. A set of inference paths is computed for the set of vector representations. An inference path in the set of inference paths connecting a first vector representation with a second vector representation. Based on a set of features, the set of vector representations is formed into clusters, a feature in the set of features comprising a relevance probability, the relevance probability corresponding to a relevance of a portion of the unstructured content according to a relevance metric. A structured representation of the unstructured content is placed at an edge location of a content delivery network determined using the set of clusters.


