Deep Level-wise XMLC Framework for Semantic Indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current semantic indexing methods for large-scale scientific literature retrieval are inefficient due to the high dimensionality of label spaces and the need for manual curation, which is time-consuming and costly, especially in domains like medicine where millions of labels are involved.

Innovation Solution

The Deep Level-wise Extreme Multi-label Learning and Classification (XMLC) framework decomposes labels into multiple levels using a category-based dynamic max-pooling methodology and a hierarchical pointer generation model to automatically select and merge labels, reducing the curse of dimensionality and eliminating the need for manual feature engineering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual curation is used to select keywords from domain ontology, then labeling accuracy is improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvelabeling accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic self-labeling of scientific literature through deep learning models that autonomously select keywords from domain ontologies without human intervention. The multi-label learning framework allows the system to automatically process and categorize large volumes of literature, eliminating the need for manual expert curation while maintaining high labeling accuracy through learned semantic representations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of expert labeling with an automated computational system. Deep neural networks and multi-label learning algorithms substitute human experts in the labeling process, using learned features and semantic understanding to automatically assign keywords from domain ontologies to scientific literature, thereby dramatically reducing time consumption while preserving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual indexing by domain experts is performed, then semantic indexing quality is improved, but productivity decreases due to the labor-intensive process

Engineering Contradiction:
Improvesemantic indexing qualityVSAvoidindexing throughput
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs automatic semantic indexing through multi-label learning frameworks that independently process scientific literature. The deep learning models automatically extract semantic information and assign keywords from domain ontologies without requiring human expert involvement, enabling high-throughput processing while maintaining indexing quality through learned semantic representations and hierarchical label structures.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the indexing process by changing from manual parameter-based labeling to automated learned feature-based labeling. Deep neural networks learn optimal labeling parameters and semantic representations from training data, automatically adjusting weighting and selection criteria to maintain high indexing quality while achieving scalable throughput across large literature collections.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If traditional multi-label learning is applied to extreme multi-label scenarios, then model simplicity is maintained, but performance degrades due to the curse of dimensionality

Engineering Contradiction:
Improvemodel simplicityVSAvoidmodel performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the extreme multi-label learning problem into hierarchical levels, organizing thousands of labels into structured hierarchies with parent-child relationships. This segmentation divides the vast label space into manageable subsets at different levels, allowing the model to process labels in a structured manner rather than as a single high-dimensional problem, thereby maintaining performance while controlling complexity through hierarchical decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces hierarchical structure as an additional dimension to organize the label space. Instead of treating all labels as a flat high-dimensional space, the system adds a hierarchical dimension that groups related labels together, enabling the model to exploit structural relationships and reduce the effective dimensionality of the problem while maintaining comprehensive label coverage.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Ease of operation

If feature engineering is performed manually for semantic indexing, then model interpretability is improved, but the process becomes time-consuming and requires expert knowledge

Engineering Contradiction:
Improvemodel interpretabilityVSAvoidfeature engineering time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual feature engineering with automated deep learning-based feature extraction. Neural networks automatically learn relevant semantic features and representations from raw text data, substituting the manual expert process with computational algorithms that automatically identify and extract meaningful features, thereby eliminating time-consuming manual feature engineering while maintaining model interpretability through learned feature hierarchies.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-featured engineering through deep learning models that autonomously extract and learn semantic features from scientific literature. The multi-label learning framework automatically identifies relevant features and representations without human intervention, enabling the system to adapt to different domains and ontologies while maintaining interpretability through the structured hierarchical label space and learned feature relationships.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11748613B2Systems and methods for large scale semantic indexing with deep level-wise extreme multi-label learning
Publication Date: 2023.09.05 BAIDU USA LLC
  • US11748613B2 patent drawing
  • US11748613B2 patent drawing
  • US11748613B2 patent drawing

AI summary

Described herein are embodiments for a deep level-wise extreme multi-label learning and classification (XMLC) framework to facilitate the semantic indexing of literatures. In one or more embodiments, the Deep Level-wise XMLC framework comprises two sequential modules, a deep level-wise multi-label learning module and a hierarchical pointer generation module. In one or more embodiments, the first module decomposes terms of domain ontology into multiple levels and builds a special convolutional neural network for each level with category-dependent dynamic max-pooling and macro F-measure based weights tuning. In one or more embodiments, the second module merges the level-wise outputs into a final summarized semantic indexing. The effectiveness of Deep Level-wise XMLC framework embodiments is demonstrated by comparing it with several state-of-the-art methods of automatic labeling on various datasets.