Machine Learning Information Extraction for Structured Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face inefficiencies and errors in manually extracting unstructured content from emails and documents into structured data formats for further processing, such as purchase applications and bill applications.

Innovation Solution

A method and apparatus for information extraction that determines target content and structured data objects based on user input, obtains structured information, and adds data items to corresponding fields, utilizing machine learning models to automate the extraction process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual extraction of unstructured content is used, then users can extract information from emails and documents, but the process is inefficient and error-prone

Engineering Contradiction:
Improveextraction efficiencyVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system enables self-service information extraction by automatically analyzing unstructured content and populating structured data objects without requiring manual user intervention for each extraction task, thereby improving both efficiency and accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual extraction process with an automated machine learning-based system that uses natural language processing and pattern recognition to extract information, eliminating human error and significantly improving productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated information extraction is implemented, then extraction efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveextraction efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system achieves universality by creating a multi-functional automated extraction platform that can handle various types of unstructured content (emails, documents, messages) and extract different kinds of information into multiple structured data formats, reducing the need for separate specialized tools

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies segmentation by breaking down the complex extraction task into distinct modular components: content parsing module, information identification module, structured data mapping module, and validation module, making the system more manageable and maintainable

Inventive Principle:
Principle #1Segmentation

3Stability of the object's composition

If structured data objects are used for information extraction, then data organization is improved, but the extraction process becomes more complex

Engineering Contradiction:
Improvedata structure organizationVSAvoidextraction process complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-defining structured data object templates with expected fields and data types before the extraction process begins, allowing the automated system to directly map extracted information to predefined structures without complex real-time decision-making

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250307555A1Information extraction
Publication Date: 2025.10.02 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250307555A1 patent drawing
  • US20250307555A1 patent drawing
  • US20250307555A1 patent drawing

AI summary

Embodiments of the disclosure provide a method, an apparatus, a device and a storage medium for information extraction. The method includes: determining, based on a user input indicating information extraction, a target content and a target structured data object; obtaining structured information of the target structured data object, the structured information indicating at least one field comprised in the target structured data object; determining, based on the target content and the structured information, at least one data item from the target content, the data item corresponding to one or more fields in the at least one field; and adding the at least one data item to corresponding one or more fields in the target structured data object, respectively. Thereby, it is possible to help a user in more efficiently organizing the information in the target content into various carriers.