Pre-trained Language Model Typography Structure Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pre-trained language models struggle to efficiently understand and process files with varying typography structures, leading to challenges in file semantic representation, classification, and knowledge element extraction across different industries.
Innovation Solution
A method for generating a pre-trained language model involves obtaining sample files, parsing them to extract typography structure information and text information, jointly training the model with task models, and fine-tuning it using the extracted information to enhance its understanding of file contents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a pre-trained language model is used for natural language processing, then language understanding capability is improved, but the ability to handle files with varying typography structures deteriorates
Solution Approach 1:
The patent segments the file processing task into multiple components: typography structure recognition, text extraction, and semantic understanding. The model processes different typography structures (tables, figures, references, etc.) as separate segments with specific parsing rules, then integrates them for comprehensive file understanding. This segmentation allows the model to handle diverse typography structures effectively while maintaining language understanding capability.
Solution Approach 2:
The patent changes the input parameters by incorporating typography structure information alongside text content. The model accepts structured inputs that include both the textual content and metadata about the typography structure (headers, paragraphs, lists, tables, etc.). This parameter expansion enables the model to adapt to various file formats and typography structures while preserving its language understanding strengths.
2Productivity
If general pre-trained language models are applied across different industries, then language processing capability is improved, but the cost and complexity of understanding industry-specific file contents increases
Solution Approach 1:
The patent applies preliminary action by pre-defining typography structure parsing rules and templates for different file types before processing actual content. The model is pre-configured with knowledge of common typography structures (headers, abstracts, conclusions, tables, figures, references) and their semantic meanings. This preliminary setup reduces the complexity of adapting to industry-specific files, as the model can leverage these pre-established patterns while maintaining high processing productivity.
Data Source
AI summary
A method for generating a pre-trained language model, includes: obtaining sample files; obtaining typography structure information and text information of the sample files by parsing the sample files; obtaining a plurality of task models of a pre-trained language model; obtaining a trained pre-trained language model by jointly training the pre-trained language model and the plurality of task models according to the typography structure information and the text information; and generating a target pre-trained language model by fine-tuning the trained pre-trained language model according to the typography structure information and the text information.


