Text processing methods, devices and electronic devices

By using keywords and entity features for clustering and correction in text processing, the problem of unbalanced data distribution is solved, and the coverage of intelligent services and user experience are improved.

CN116955602BActive Publication Date: 2026-06-30CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2022-11-04
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing text processing technologies struggle to address the imbalance in data distribution within different scenario categories, causing intelligent services to ignore low-frequency queries, thus harming user experience and system usage frequency.

Method used

By obtaining keyword and entity features from the labeled text data, clustering is performed to determine preliminary sub-scenes. Vector similarity is used to retrieval space correction to finally determine the final sub-scenes and perform a balanced partitioning of the training and test sets.

Benefits of technology

It achieves a balanced distribution of data within scene categories, improves the coverage of intelligent services for low-frequency queries, and enhances user experience and system usage frequency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116955602B_ABST
    Figure CN116955602B_ABST
Patent Text Reader

Abstract

This invention provides a text processing method, apparatus, and electronic device. The method includes obtaining keyword features and entity features of text data labeled in a first scenario; performing clustering processing on the text data based on the keyword features and entity features to determine preliminary sub-scenarios in the first scenario; refining the preliminary sub-scenarios to determine final sub-scenarios; and partitioning the dataset based on the text data in the final sub-scenarios to obtain a training set and a test set. This invention enables a balanced data distribution within scenario categories, thereby allowing intelligent services trained based on this data distribution to cover less frequent queries, thus improving the overall user experience and the frequency of use of the intelligent system.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Fragment text processing method and device and electronic equipment

    CN111460096A

  • Text classification data processing method and device, storage medium and program product

    CN113722493A