Semantic Representation Model Knowledge Unit Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current natural language processing (NLP) methods, such as XLNet, fail to effectively model complete words and entities due to their unit-based approach, resulting in limited model performance.
Innovation Solution
A method and apparatus for generating a semantic representation model by performing recognition and segmentation on original texts to obtain knowledge units and non-knowledge units, followed by knowledge unit-level disorder processing and character attribute generation, which are used to create a training text set for training an initial semantic representation model using deep learning techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If word-level disorder processing is performed as in XLNet, then the modeling approach is simple, but complete words and entities cannot be effectively modeled
Solution Approach 1:
The patent segments text into knowledge units (complete words, entities, or phrases with semantic meaning) rather than processing at the character or word level. This segmentation enables the model to maintain complete semantic units while performing disorder processing, thereby improving modeling precision without excessive complexity
Solution Approach 2:
The patent introduces a new dimension of processing by operating at the knowledge unit level rather than traditional character or word levels. This dimensional change allows the model to capture complete semantic meanings while maintaining the benefits of disorder processing for position awareness
2Reliability
If knowledge unit-level disorder processing is performed, then complete words and entities can be modeled, but the processing complexity increases
Solution Approach 1:
The text is segmented into knowledge units that represent complete semantic entities. This segmentation approach improves model performance by ensuring complete words and entities are processed together, while the modular nature of segmentation keeps processing complexity manageable
Solution Approach 2:
Knowledge units are pre-identified and segmented before disorder processing. This preliminary action ensures that complete semantic units are established upfront, improving model reliability while avoiding the need for complex real-time processing during the main modeling phase
Data Source
AI summary
The disclosure discloses a method and an apparatus for generating a semantic representation model, and a storage medium. The detailed implementation includes: performing recognition and segmentation on the original text included in an original text set to obtain knowledge units and non-knowledge units in the original text; performing knowledge unit-level disorder processing on the knowledge units and the non-knowledge units in the original text to obtain a disorder text; generating a training text set based on the character attribute of each character in the disorder text; and training an initial semantic representation model by employing the training text set to generate the semantic representation model.


