Knowledge Graph Generation from Semi-Structured Webpage Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The establishment of knowledge graphs from semi-structured goods webpage information is time and labor-consuming, hindering the efficient use of this data in intelligent question-answering systems.
Innovation Solution
A knowledge graph generating apparatus and method that annotates entity-relationships in a template format, generates an extraction rule set, extracts product information from multiple webpages, classifies it using a product information classification model, and converts the data into semantic triples to automatically create and update knowledge graphs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual methods are used to establish knowledge graphs from goods webpage information, then the knowledge graph can be accurately constructed, but the process is time and labor-consuming
Solution Approach 1:
The system automatically extracts product information from webpages and generates knowledge graphs without human intervention. The extraction rule set autonomously identifies and extracts entities, attributes, and relationships from semi-structured webpage data, converting it into structured knowledge graph format. This self-service mechanism eliminates manual labor while maintaining construction accuracy through systematic processing rules.
Solution Approach 2:
The patent replaces manual mechanical operations with automated computational systems. Instead of human analysts manually parsing webpages and constructing knowledge graphs, the system uses computer-based extraction rules and algorithms to automatically process webpage information, transform semi-structured data into structured formats, and generate knowledge graphs programmatically.
2Productivity
If automated extraction methods are implemented, then processing efficiency is improved, but handling diverse webpage formats becomes more difficult
Solution Approach 1:
The extraction rule set is designed with universal applicability to handle multiple webpage formats and templates. The system can process different goods webpage structures (e.g., auction sites, e-commerce platforms) using the same extraction framework, adapting to various formats through configurable rules that identify common patterns in product information presentation across different website designs.
Solution Approach 2:
The system handles format diversity by dynamically adjusting extraction parameters and rules based on the detected webpage template type. When encountering different webpage formats, the system modifies extraction configurations, attribute mappings, and parsing strategies to suit the specific format while maintaining consistent output structure for the knowledge graph.
3Quantity of substance
If comprehensive product information is extracted from multiple webpages, then the knowledge graph becomes more complete, but the data processing complexity increases
Solution Approach 1:
The system segments the complex task of knowledge graph construction into distinct processing stages: webpage retrieval, information extraction, data classification, and knowledge graph generation. By dividing the process into manageable segments with specialized extraction rules for each stage, the system handles large volumes of data from multiple webpages systematically, reducing overall processing complexity while maintaining data completeness.
Solution Approach 2:
The patent introduces an intermediary classification layer that processes extracted product information before final knowledge graph construction. This intermediary step organizes raw extracted data into standardized formats, resolves inconsistencies, and prepares data for integration, thereby managing the complexity of processing comprehensive information from diverse sources while ensuring complete and consistent knowledge graph output.
Data Source
AI summary
A knowledge graph generating apparatus, method and non-transitory computer readable storage medium thereof are provided. The apparatus marks an entity-relationship of the template of goods information in the template of webpage according to the operating signal and generates an extraction rule set, wherein the template of webpage is one of multiple goods webpages and has a template format. The apparatus extracts a plurality of first product information of the first goods webpages according to the extraction rule set, wherein the first goods webpages have the template format and are selected from the goods webpages. The apparatus generates a classified goods information result through a product information classification model, wherein the product information classification model is generated based on the first product information and the entity-relationship of the template of goods information. The apparatus converts the classified goods information result into several semantic triples to generate a knowledge graph.


