DeepWeb Entity Recognition via Uniqueness Constraint

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional DeepWeb entity recognition methods suffer from low accuracy due to incorrect attribute values and data conflicts, leading to errors in entity recognition, as they fail to effectively enforce the uniqueness constraint across disparate data sources.

Innovation Solution

A DeepWeb entity recognition method based on a uniqueness constraint, which involves acquiring an entity object set, performing structure conversion to obtain attribute sets, calculating matching degrees, filtering with a threshold, calculating object similarity, and merging clusters to recognize entities accurately using a uniqueness constraint.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional entity recognition methods are used, then the process is simple, but the accuracy is low due to incorrect attribute values and data conflicts

Engineering Contradiction:
Improveentity recognition accuracyVSAvoidrecognition process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the entity recognition process into distinct phases: data acquisition, attribute value correction, uniqueness constraint enforcement, and entity identification. This segmentation allows each phase to address specific issues independently, improving overall accuracy while managing complexity through structured processing steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by correcting attribute values and enforcing uniqueness constraints before performing entity recognition. This preliminary processing of data quality issues ensures that the subsequent recognition process operates on cleaned, reliable data, thereby improving accuracy without adding complexity to the core recognition algorithm.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If data from multiple sources are merged, then information completeness increases, but data conflicts and incorrect values increase

Engineering Contradiction:
Improveinformation completenessVSAvoiddata accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements feedback mechanisms that continuously validate attribute values against uniqueness constraints and source reliability criteria. When conflicts are detected, the system feedbacks to correct or reject problematic data entries, ensuring that information completeness does not compromise data accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameter of data validation by introducing uniqueness constraint checking and source reliability assessment. These parameter changes enable the system to distinguish between useful information and erroneous data, allowing merging of multiple sources while maintaining accuracy through enhanced validation parameters.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If attribute values are corrected, then entity recognition accuracy improves, but processing time increases

Engineering Contradiction:
Improveentity recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs attribute value correction as a preliminary action before entity recognition. By addressing data quality issues in advance, the system avoids time-consuming corrections during the recognition process itself, thereby improving accuracy without significantly increasing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies self-service principles by automatically detecting and correcting attribute value conflicts using predefined uniqueness constraints and source reliability rules. This automated self-correction reduces the need for manual intervention and minimizes processing time while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240386068A1Deepweb entity recognition method, apparatus, device, and medium based on uniqueness constraint
Publication Date: 2024.11.21 BEIJING HYDROPHIS NETWORK TECH CO LTD
  • US20240386068A1 patent drawing
  • US20240386068A1 patent drawing
  • US20240386068A1 patent drawing

AI summary

The present invention discloses a DeepWeb entity recognition method based on a uniqueness constraint, including: performing structure conversion on an entity object set to obtain an entity object attribute set of a DeepWeb; calculating a matching degree between entity objects in the entity object attribute set, and constructing a matching list of the entity object set according to the matching degree; filtering the matching list to obtain an entity class cluster; calculating an object similarity degree of each entity object in the entity class cluster, and merging the entity class cluster according to the object similarity degree to obtain the entity class set; searching a uniqueness constraint corresponding to each entity object in the entity object set according to the entity class set, and recognizing an entity object in the DeepWeb according to the uniqueness constraint. The present invention can improve the accuracy of entity recognition in DeepWebs.