Entity Relationship Clustering via Semantic Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current entity recognition methods face challenges in efficiently extracting user behavioral habits and interests from unstructured online data due to varying text expression formats, leading to low extraction efficiency.

Innovation Solution

A method involving capturing social relationship data, performing entity recognition, reverse marking, and utilizing a pre-trained context semantic recognition model for entity relationship recognition, followed by character and semantic similarity calculations to cluster entity relationships, improving extraction efficiency through weighted calculations and nearest node algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If entity recognition is performed on unstructured online data using traditional entity recognition networks, then basic information such as personal attributes can be acquired easily, but extraction efficiency of behavioral habits and interests is low

Engineering Contradiction:
Improveease of acquiring basic informationVSAvoidextraction efficiency of behavioral habits
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent segments unstructured text data into structured formats by dividing text into preceding, middle, and following portions relative to entity positions. This segmentation enables systematic processing of behavioral habit expressions while maintaining ease of implementation through clear structural organization of the data flow.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If traditional entity recognition networks are used to process unstructured data with varying text expression formats, then the system is simple to implement, but extraction efficiency of entities related to behavioral habits is difficult to improve

Engineering Contradiction:
Improvesimplicity of entity recognition systemVSAvoidextraction efficiency of entities
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent introduces a new dimensional approach by calculating both character similarity and semantic similarity between entity pairs, then combining these dimensions through weighted summation. This multi-dimensional similarity assessment enables efficient extraction of behavioral habits while maintaining system simplicity through modular calculation components.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If entity relationships are extracted from unstructured data with different expression ways, then comprehensive coverage is achieved, but extraction efficiency decreases

Engineering Contradiction:
Improvecoverage of various text expressionsVSAvoidextraction efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent changes the parameter of similarity measurement by incorporating both character-level and semantic-level similarities with adjustable weighting coefficients. This parameter transformation allows the system to adapt to various text expression formats while maintaining high extraction efficiency through optimized similarity calculations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240311931A1Method, apparatus, device, and storage medium for clustering extraction of entity relationships
Publication Date: 2024.09.19 BEIJING HYDROPHIS NETWORK TECH CO LTD
  • US20240311931A1 patent drawing
  • US20240311931A1 patent drawing
  • US20240311931A1 patent drawing

AI summary

The present disclosure relates to an artificial intelligence technology, and discloses a method, an apparatus, a device, and a storage medium for clustering extraction of entity relationships. The method includes: capturing social relationship data of a user, performing entity recognition operation on the social relationship data to obtain an entity recognition result, and performing reverse marking operation on the social relationship data according to the entity recognition result to obtain a data marking sequence set; performing entity relationship recognition on various data marking sequences in the data marking sequence set to obtain a relationship recognition result, and constructing to obtain an entity-relationship group set; and calculating a character similarity between various entities and a semantic similarity between various entity relationships in the entity-relationship group set, and clustering various entity-relationship groups in the entity-relationship group set. The present disclosure may improve the extraction efficiency of entity relationships.