Material chemistry field word segmentation method based on PSLL-BilSTM-CRF

By combining PSLL algorithm, BilSTM and CRF models, the problem of sparse high-quality labeling samples in Chinese word segmentation in the field of material chemistry is solved, significantly improving the accuracy of word segmentation and the effect of processing rare words.

CN119940356APending Publication Date: 2025-05-06YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452440.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the Chinese word segmentation task in the field of materials chemistry, high-quality labeling samples are sparse, resulting in poor performance of traditional word segmentation models in this field environment.

Method used

Combining the PSLL algorithm, BilSTM and CRF models, the PSLL-BilSTM-CRF algorithm is formed, and a large number of labeled samples are generated through pseudo-label learning to enhance the weight of domain vocabulary in the model.

Benefits of technology

The accuracy of word segmentation model in the field of material chemistry is improved, especially when dealing with a large number of rare words, which significantly improves word segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940356A_ABST
    Figure CN119940356A_ABST
Patent Text Reader

Abstract

The invention discloses a PSLL-BilSTM-CRF-based Chinese word segmentation method in a specific field, relates to the field of natural language processing, and combines a CRF model and a BiLSTM model to form a new PSLL-BilSTM-CRF algorithm to perform Chinese word segmentation in the field. According to the method, traditional jieba word segmentation is improved in thought, a word segmentation algorithm is constructed by using a method of mixing multiple modes so as to solve the task of field word segmentation, pseudo sample labeling random links are fused with BilSTM and CRF models, the problem of sparse high-quality labeled samples is solved, the weight of field vocabularies in the models is enhanced, and the field word segmentation efficiency is improved. And the performance of the current word segmentation model in the domain environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to Chinese word segmentation in a specific field. Background Art

[0002] Unlike most Western languages, there are no obvious spaces between words in written Chinese, and sentences appear in the form of strings. Therefore, the first step in processing Chinese is to automatically segment words, that is, to convert strings into word strings. In a certain limited field, such as finance, material chemistry, etc., it is very challenging to collect and organize professional vocabulary. The corpora or models existing in other fields often lack strong generalization, which further increases the cost of traditional Chinese word segmentation models.

[0003] In recent years, deep learning has achieved remarkable results in the field of NLP. More and more scholars in the field of NLP have begun to use a variety of mixed methods to construct word segmentation algorithms to solve the task of domain word segmentation, such as: using dictionaries or templates to extract domain word features, and then further introducing machine learning or deep learning algorithms to achieve domain word segmentation. In the era of big data, pseudo-labeling technology that seems to be suitable for small sample learning has gradually faded out of people's vision, but in fact, in the fields of samples and finance, medical images, materials chemistry, etc., pseudo-label learning is still a simple and effective means. The main idea of ​​(PseudoSample Labeled and Linked, referred to as PSLL) is to randomly obtain the current annotated corpus, judge the length of the current corpus and the size of the threshold, and finally extract vocabulary samples from the domain dictionary in a non-repeated manner when the sample length is less than the limit, and then arrange them in a random manner and link them with the current corpus. In view of the idea of ​​combining deep learning algorithms, the PSLL-BilSTM-CRF algorithm is proposed to solve the problem of sparse high-quality annotated samples, strengthen the weight of domain vocabulary in the model, and improve the performance of the current word segmentation model in this domain environment.

[0004] The present invention combines the PSLL algorithm with the BilSTM and CRF models to form a PSLL-BilSTM-CRF algorithm for word segmentation in the field of materials chemistry, which solves the problem of sparse high-quality labeled samples, strengthens the weight of domain-specific vocabulary in the model, and improves the performance of the current word segmentation model in this field environment. Summary of the invention

[0005] In order to solve the problem of sparse high-quality annotated samples in the field of materials chemistry, this paper proposes a Chinese word segmentation technology based on PSLL-BilSTM-CRF. This technology is based on the PSLL algorithm and the BilSTM and CRF models, and improves them to form a new PSLL-BilSTM-CRF algorithm to solve the problem of difficulty in modeling caused by the main characteristics of domain word segmentation.

[0006] The technical solution adopted by the present invention is:

[0007] Step 1: Construction of domain dictionary library;

[0008] Step 2: PSLL algorithm design;

[0009] Step 3: Word embedding design;

[0010] Step 4: Neural network layer design;

[0011] Step 5: CRF layer design;

[0012] Compared with the prior art, the present invention has the following advantages:

[0013] (1) Compared with traditional word segmentation algorithms, it has higher accuracy.

[0014] (2) For a large number of uncommon words in a domain, word segmentation is more effective. DETAILED DESCRIPTION

[0015] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation modes and the accompanying drawings.

[0016] In this specific implementation, the following processing steps are included for domain segmentation:

[0017] Step 1: Build a domain dictionary library

[0018] The data of the dictionary library in the field of materials chemistry is mainly obtained from dictionaries of the field written by experts. The more commonly used dictionaries include "Basic Terminology of Materials Science" and "Chemical Dictionary". Secondly, relevant financial literature is also an important source of vocabulary in the financial field, and the literature often contains the most cutting-edge professional vocabulary in the field. The original data are all natural sentences and cannot be used directly to build a dictionary library. Label filtering, null value removal, length detection, manual word extraction, cross-checking and other processes are required before they can be input into the dictionary library. See the process. Figure 1 .

[0019] Step 2: PSLL Algorithm Design

[0020] The main purpose of the PSLL algorithm is to generate a large amount of pseudo sample data based on the domain dictionary and some annotated training corpora. The main idea is to randomly obtain the current annotated corpus, determine the length L of the current corpus and the size of the threshold M, and randomly extract vocabulary samples from the domain dictionary S in a non-repeated manner when the sample length is less than M, and then arrange them in a random manner again and link them with the current corpus L. The algorithm principle can be expressed as the following formula. Among them, S represents the dictionary, M represents the length of the final output sentence, L represents the length of the currently extracted annotated corpus, R represents the current annotated corpus, O represents the output, r(x, y) represents randomly extracting y elements from the set x, f(x) represents randomly arranging the element group x, and c(x, y) represents randomly linking the x and y elements.

[0021] Step 3: Word Embedding Design

[0022] Word embedding mainly vectorizes the input training corpus, and the multi-dimensional representation of the vector retains the intrinsic meaning of the current word. This model uses BERT to vectorize the words in the training corpus. The vectorization process mainly maps the segmented words to high-dimensional vectors containing contextual information. BERT is based on a bidirectional language model and is an unsupervised training algorithm model. It uses a bidirectional Transformer encoder module as the main module. Transformer is completely based on the Attention mechanism for modeling, using stacked Attention fully connected layers, and most of the internal sequence processing uses the Encoder-Decoder structure. When using BERT for model fine-tuning, first fine-tune the model's training parameters, and then fine-tune the downstream parameters with variable control, so as to fully combine the Transformer structure and the existing high-quality model to handle the current task.

[0023] Step 4: Neural Network Layers

[0024] At the neural network layer, the commonly used LSTM algorithm mainly introduces a gating mechanism based on the RNN algorithm, which effectively solves the long-term dependency problem and gradient control problem of RNN in long-sequence text task training. The BiLSTM algorithm is essentially a variant of LSTM, mainly composed of a forward model and a backward model. It can fully consider the impact of past and future contextual information on the model learning at the current moment, and can more fully understand the meaning of the text. At the same time, BiLSTM requires fewer parameters than LSTM, which can save costs to a large extent as the number of samples gradually increases.

[0025] Step 5: CRF layer design

[0026] The full name of CRF is undirected probability graph model, and the model structure is as follows Figure 2 As shown in the structure diagram, y i The value of is related to the context and the entire observation sequence x. Therefore, the state sequence y i The value of not only considers the previous and next labels, but is also affected by the context information. If CRF is used to predict the state sequence, it not only uses the context information, but also solves the dependency problem between different labels to a certain extent. Therefore, the CRF algorithm is widely used in natural language processing tasks with serialized annotations. Using CRF to predict the dependency between adjacent labels, the mapping formula for the given input sequence x and the predicted state sequence result y is shown below.

[0027] BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Flowchart of domain dictionary library construction;

[0029] Figure 2 CRF structure diagram;

[0030] The above description is only a specific implementation mode of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all methods or steps in the process, except for mutually exclusive features and / or steps, can be combined in any way.

Claims

1. A method for word segmentation in the field of materials chemistry based on PSLL-BilSTM-CRF, characterized in that: The following steps are involved: Step 1: Construction of domain dictionary library; Step 2: PSLL algorithm design; Step 3: Word embedding design; Step 4: Neural network layer design; Step 5: CRF layer design.

2. The method according to claim 1, characterized in that: In step 3, BERT is used instead of Word2Vec to convert static word vectors into dynamic word vectors as parameters for learning. This solves the problem of polysemy or lexical ambiguity in the actual process and can effectively learn contextual information.

3. The method according to claim 1, characterized in that: Replacing LSTM with BiLSTM in step 4 can fully consider the impact of past and future contextual information on model learning at the current moment, and can more fully understand the meaning of the text. At the same time, BiLSTM requires fewer parameters than LSTM, which can greatly save costs as the number of samples gradually increases.