Vertical domain natural language large model data processing method and device, equipment and medium
By performing security screening, anonymization, noise reduction, task clustering, and quality scoring on user question-and-answer data from a large-scale natural language model in a vertical domain, a high-quality training set is constructed, solving the problem of inconsistent data quality and improving model iteration efficiency and data processing efficiency.
Patent Information
- Authority / Receiving Office
- CN Β· China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing vertical domain large model data flywheel technology, the data quality is uneven, the marginal returns of low-quality data are significantly reduced, simply increasing the amount of data cannot continuously improve model performance, and the model iteration cycle is too long.
By performing security screening, anonymization, noise reduction, task clustering, and quality scoring on user question-and-answer data from a large-scale natural language processing model in a vertical domain within the target scenario, the classification of the data is determined. Based on the classification, data processing strategies are configured to construct high-quality training and testing sets, thereby achieving efficient interaction between data and the model.
It significantly improved model iteration efficiency, reduced manual intervention, increased data processing efficiency, enabled specialized model training, and created a snowball effect where data becomes better with use, models become more professional with training, and business value increases exponentially.
Smart Images

Figure CN122114153A_ABST