A new word screening method, device and electronic equipment
By identifying new words through text processing and statistical parameters, the problem of new word identification in text processing and information mining is solved, and the efficiency of new word screening and the accuracy of speech recognition are improved.
CN117272986BActive Publication Date: 2026-07-10JUHAOKAN TECH CO LTD
Patent Information
- Application Number
- CN202211723261.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-07-10
- Estimated Expiration
- 2042-12-30
AI Technical Summary
Technical Problem
Existing technologies struggle to quickly and effectively identify new words in the fields of text processing and information mining.
Method used
By acquiring the text to be analyzed, text processing is performed to determine text segments. Based on statistical parameters of the text segments, such as frequency of occurrence, coagulation rate, left entropy, and right entropy, characters contained in text segments that meet preset conditions are identified as new words.
Benefits of technology
It enables rapid identification of new words in the fields of text processing and information mining, and improves the accuracy of speech recognition and the efficiency of updating the strong word list.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN117272986B_ABST
Abstract
The present disclosure provides a new word screening method and device and electronic equipment, relates to the technical field of data processing, and is used for solving the problem of how to quickly and effectively identify new words in the field of text processing and information mining. The method comprises the following steps: obtaining a text to be analyzed; performing text processing on the text to be analyzed to determine at least one text segment; determining a statistical parameter of each text segment according to the total number of text segments and the total number of each text segment; in the case that the statistical parameter meets a preset condition and the characters contained in the text segment meeting the preset condition are weak words, determining that the characters contained in the text segment meeting the preset condition are new words; wherein the weak words include words that have been used in the target field but are not nouns.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Word segmentation method and system for address standardized corpus
CN109858025A
New words determination method and device, computer equipment and medium
CN112329443A