Vertical domain natural language large model data processing method and device, equipment and medium

By performing security screening, anonymization, noise reduction, task clustering, and quality scoring on user question-and-answer data from a large-scale natural language model in a vertical domain, a high-quality training set is constructed, solving the problem of inconsistent data quality and improving model iteration efficiency and data processing efficiency.

CN122114153APending Publication Date: 2026-05-29CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1

Patent Information

Authority / Receiving Office
CN Β· China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing vertical domain large model data flywheel technology, the data quality is uneven, the marginal returns of low-quality data are significantly reduced, simply increasing the amount of data cannot continuously improve model performance, and the model iteration cycle is too long.

Method used

By performing security screening, anonymization, noise reduction, task clustering, and quality scoring on user question-and-answer data from a large-scale natural language processing model in a vertical domain within the target scenario, the classification of the data is determined. Based on the classification, data processing strategies are configured to construct high-quality training and testing sets, thereby achieving efficient interaction between data and the model.

Benefits of technology

It significantly improved model iteration efficiency, reduced manual intervention, increased data processing efficiency, enabled specialized model training, and created a snowball effect where data becomes better with use, models become more professional with training, and business value increases exponentially.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114153A_ABST
    Figure CN122114153A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and provides a vertical domain natural language large model data processing method, device, equipment and medium, the method comprises: according to the task type under the target scene, the user question and answer data are clustered, and the clustering data corresponding to each task type is obtained;Determine the calling amount and scoring condition of the clustering data corresponding to each task type;According to the calling amount and historical scoring condition, the classification condition of the clustering data is determined;According to the classification condition, the data processing strategy is determined, and the clustering data is processed according to the data processing strategy.The present application realizes the efficient linkage of data and model, and makes the data specialized model training through data clustering, which significantly improves the model iteration efficiency.
Need to check novelty before this filing date? Find Prior Art