Multi-Dimensional Feature-Space Node Assignment for Content Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing tools and methods for deriving insights from news articles about corporate entities are inefficient, requiring human intervention, compromising data integrity, and are computationally expensive, especially when dealing with complex and unstructured data sets.

Innovation Solution

A system utilizing label and feature-vector machine learning models to cluster and label content items from multiple sources, generating a multi-dimensional feature space representation to provide intuitive insights with minimal human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional spreadsheet approaches and human intervention are used to analyze news articles and derive insights, then data integrity and security are maintained, but the process is time-consuming and computationally expensive

Engineering Contradiction:
Improvedata integrityVSAvoidanalysis time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces machine learning models as intermediaries between the news articles and the analysis process. These models process the unstructured text data automatically, extracting entity relationships and generating insights without requiring direct human intervention, thus reducing time loss while maintaining reliability through controlled access to sensitive information.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical system of human analysts manually reviewing news articles with automated machine learning-based text processing. This substitution eliminates the time-consuming manual analysis while preserving data integrity through secure, automated access controls and processing pipelines.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If freely available automated tools are used to transform complex information, then analysis speed is improved, but data integrity and compliance are compromised due to access to secured information

Engineering Contradiction:
Improveanalysis speedVSAvoiddata integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces a secure, customized machine learning processing system as an intermediary between the news articles and analysis tools. This intermediary processes information locally with controlled access, enabling fast automated analysis while preventing unauthorized access to sensitive data, thus maintaining both productivity and data integrity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a multi-functional system that combines the speed of automated tools with the security of controlled access. The machine learning models perform multiple functions including text processing, entity extraction, relationship identification, and compliance checking within a single secure framework, achieving both high productivity and data integrity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Extent of automation

If machine learning algorithms are used to process unstructured data sets, then analysis automation is improved, but the ability to gather meaningful insights is reduced

Engineering Contradiction:
Improveautomation levelVSAvoidmeaningful insights
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent applies local quality by training machine learning models on domain-specific data and using specialized processing techniques tailored to different types of relationships (e.g., acquisition, partnership, competition). This localized approach ensures that automated processing extracts meaningful insights specific to each relationship type while maintaining high automation levels.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes parameters such as feature extraction methods, clustering algorithms, and relationship detection thresholds to optimize insight extraction from unstructured data. By adjusting these parameters based on the specific characteristics of the news articles and desired insights, the system maintains high automation while preserving meaningful information.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If complex data models are generated for unstructured data sets, then comprehensive insights are achieved, but computational expense and labor intensity increase

Engineering Contradiction:
Improvecomprehensive insightsVSAvoidcomputational expense
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent segments the complex analysis task into multiple smaller processing stages: initial text processing, entity identification, relationship extraction, and insight generation. Each stage uses appropriately sized computational models, avoiding the need for a single large complex model, thus reducing overall computational expense while maintaining comprehensive insights.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by generating insights at multiple levels of detail rather than creating a single comprehensive model. The system can provide summary-level insights quickly and drill down into detailed relationships only when needed, reducing computational expense by avoiding unnecessary processing of all data at maximum detail levels.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12423329B2Cluster based node assignment in multi-dimensional feature space
Publication Date: 2025.09.23 ORACLE INT CORP
  • US12423329B2 patent drawing
  • US12423329B2 patent drawing
  • US12423329B2 patent drawing

AI summary

A system and computer-implemented method includes accessing a set of clusters, where each cluster is defined to cover a portion of a multi-dimensional feature space, and each cluster is associated with a label of a plurality of labels. A plurality of content items is received from a plurality of data sources. The label is assigned to each of the plurality of content items using one or more label machine learning models. Using a feature-vector machine learning model, a set of feature vectors for the plurality of content items are identified. A feature vector is associated with a respective portion in the multi-dimensional feature space. Portions in the multi-dimensional feature space corresponding to the plurality of labels are identified. A cluster from the plurality of clusters is assigned to the feature vector based on a proximity of the feature vector to the plurality of clusters in the multi-dimensional feature space.