Deep learning-based Chinese word segmentation method

By combining global Transformer and capsule network, Chinese word segmentation features are dynamically extracted, which solves the problem of long dependence and noise in Chinese word segmentation model, improves word segmentation accuracy and efficiency, reduces dependence on labeled data, and promotes innovation in natural language processing technology.

CN120337919APending Publication Date: 2025-07-18SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510441060.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the Chinese word segmentation model has difficulties in dealing with unlogged words and ambiguity problems, and the convolutional extraction feature method is difficult to obtain long dependencies, and the global Transformer is susceptible to noise interference, which affects word segmentation efficiency and accuracy.

Method used

The method of combining global Transformer network with capsule network is adopted to extract multi-dimensional features through multi-head attention and dynamic routing mechanisms, dynamically filter n-gram mode, and optimize tag sequences with conditional random field layer to solve long dependence and noise problems.

Benefits of technology

It improves the accuracy and generalization ability of the Chinese word segmentation model, reduces the impact of noise, reduces the dependence on a large amount of labeled data, saves network resources, and improves word segmentation efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337919A_ABST
    Figure CN120337919A_ABST
Patent Text Reader

Abstract

The invention provides a Chinese word segmentation method based on deep learning, and relates to the technical field of natural languages, and the method comprises the steps: dividing a Chinese word segmentation public data set according to a proportion, and carrying out the pre-training, and obtaining the multi-dimensional features of the word granularity; and constructing a Chinese word segmentation model TC-CRF based on deep learning, and performing Chinese word segmentation processing on the multi-dimensional features by using the Chinese word segmentation model TC-CRF. According to the method, the problem that long dependence cannot be obtained by an existing convolution feature extraction method is solved, the problem that the global transformer obtains too large noise is solved, and Chinese word segmentation is effectively realized by using deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language technology, and in particular relates to a Chinese word segmentation method based on deep learning. Background Art

[0002] Chinese is an ideographic writing system with characters as the basic writing unit. There are no obvious delimiters such as spaces between words like in English. Therefore, Chinese word segmentation is a key step in preprocessing in Chinese natural language processing tasks.

[0003] In the early days, Chinese word segmentation mainly relied on dictionary-based mechanical matching methods, such as forward maximum matching (FMM), reverse maximum matching (RMM), bidirectional maximum matching, full segmentation, and other common matching algorithms. They matched strings through preset dictionaries, but were difficult to handle out-of-date (OOV) and ambiguity problems. With the rise of statistical learning, hidden Markov models (HMM) and conditional random fields (CRF) have been widely used in sequence labeling tasks to achieve word boundary prediction through probabilistic modeling. Although statistical methods can alleviate ambiguity and out-of-date word problems to a certain extent, the models constructed by such methods are relatively complex and require manual feature extraction.

[0004] Deep learning technology has promoted a paradigm shift in Chinese word segmentation. As a representative neural network in Chinese word segmentation, the BiLSTM-CRF model captures contextual features through a bidirectional long short-term memory network and combines the CRF layer to optimize the label sequence, becoming a representative architecture of the neural word segmentation model. However, the RNN-based model still faces the problems of insufficient long-distance dependency modeling and low training efficiency.

[0005] The self-attention-based Transformer model proposed in the prior art has completely changed the way of sequence modeling. Its multi-head self-attention mechanism can dynamically establish dependencies between arbitrary positions, providing a new idea for solving the long-distance dependency problem. In view of the fact that this model is easily disturbed by global attention noise in the Chinese word segmentation task, researchers have tried to combine local attention with global attention. The prior art also proposed that the convolutional layer can capture rich n-gram features, thereby obtaining contextual information within a limited perceptual field. However, this method is still insufficient in dynamically regulating feature selectivity and is difficult to adapt to the dynamic changes in the boundaries of Chinese words. Summary of the invention

[0006] In view of the above deficiencies in the prior art, the present invention provides a Chinese word segmentation method based on deep learning, which solves the problem that the existing convolutional feature extraction method cannot capture long dependencies and the problem of too much noise in the global transformer, and effectively realizes Chinese word segmentation using deep learning.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a Chinese word segmentation method based on deep learning, comprising the following steps:

[0008] For the publicly available Chinese word segmentation dataset, it is divided according to a certain proportion and pre-trained to obtain multi-dimensional features at the character level.

[0009] Construct a Chinese word segmentation model TC-CRF based on deep learning, and use the Chinese word segmentation model TC-CRF to perform Chinese word segmentation on the multi-dimensional features.

[0010] Further, the Chinese word segmentation model TC-CRF includes:

[0011] A global transformer network for processing 768-dimensional features to obtain 256-dimensional hidden layer features;

[0012] A capsule network for processing 768-dimensional features to extract features of different n-gram patterns, where each n-gram pattern is encapsulated as a capsule, and n-gram represents a sequence of n consecutive items in the text;

[0013] A fusion layer for concatenating the hidden layer features extracted by the global transformer network and the features extracted by the capsule network, and passing through a fully connected layer to obtain the score of the label;

[0014] A conditional random field layer for obtaining the structural information of the annotation sequence by modeling the dependencies between the annotation labels according to the score of the label, and completing the Chinese word segmentation process.

[0015] Still further, the processing of the 768-dimensional features to obtain 256-dimensional hidden layer features is specifically as follows:

[0016] According to the 768-dimensional features, key value K, query Q, and value V are obtained through linear transformation;

[0017] Using multi-head attention calculation, key value K, query Q, and value V are divided into h parts, and the calculation of attention for each head is as follows:

[0018]

[0019] Among them, Attention() represents attention calculation, d kRepresents the dimension of each head, represents the adjustment factor, h represents the number of heads, and T represents the transpose;

[0020] Concatenate the attention calculation results of each head:

[0021] M = concat(head0,..., head h )

[0022] where M represents the concatenation result, head h represents the attention calculation result of the h-th head, and concat() represents the concatenation operation;

[0023] Use a feed-forward network to perform a non-linear transformation on the concatenation result to obtain 256-dimensional hidden layer features.

[0024] Furthermore, the extraction of features of different n-gram patterns is specifically as follows:

[0025] Input the 768-dimensional features into a capsule network, and perform convolution operations with different convolution kernel sizes on the input features in the convolutional layer to obtain different n-gram patterns. Among them, the convolution operation is as follows:

[0026]

[0027] where represents the k-gram feature at the i" position, σ represents the activation function, x represents the sequence, i":i"+k represents the word vectors at the positions of the words in the range from i" to i"+k in the sequence, * represents the dot product, H k represents the convolution kernel with a width of k, and b represents the offset;

[0028] In the primary capsule layer, encapsulate each n-gram pattern into a capsule and send the capsule to the dynamic routing layer. Among them, each input capsule makes a prediction for the high-level capsule. Using the learnable matrix W, let the output of the i'-th capsule in the l-th layer be the vector u i' ∈R d , and the high-level capsule prediction generates a prediction vector u j|i' for the high-level capsule through an affine transformation:

[0029] u j|i' = W i'j u i'

[0030] where W i'j represents the learning parameter, maps the vector of the low-level capsule i' to the space of the high-level capsule j, and R d represents the vector space of size d;

[0031] Based on the prediction results of high-level capsules obtained from each low-level capsule, through affine transformation and voting on capsules composed of different n-gram patterns, the voting weights from low-level capsules to high-level capsules are continuously updated in an iterative manner. The prediction results are weighted and summed according to the weights to dynamically screen out the feature capsules useful for label score prediction, and the high-level capsules are obtained:

[0032]

[0033] Among them, u j|i represents the prediction vector of low-level capsule i' for high-level capsule j, c ij represents the weight, and s j represents the high-level capsule;

[0034] Use the flatten function to combine the number of output capsules and the output capsule dimension into the total number of capsule features to complete the unpacking;

[0035] Using unpacking, the feature capsules output by dynamic screening are restored to multi-dimensional features to complete the feature extraction of different n-gram patterns.

[0036] Furthermore, the expression of the loss function of the Chinese word segmentation model TC-CRF is as follows:

[0037]

[0038] Among them, Loss represents the loss function of the Chinese word segmentation model TC-CRF, N represents the total number of samples, p(y i |x i ) represents the probability that the sample belongs to the true label y i under the condition of the given input x i , y i represents the true label of the i-th sample, and x i represents the input of the i-th sample.

[0039] The beneficial effects of the present invention:

[0040] (1) The basic idea of the present invention is to extract the global dependencies of Chinese word segmentation based on the global Transformer, and at the same time, utilize the advantage of local feature extraction of the capsule network to propose a method of individual time-step independent routing to dynamically extract local dependencies of different lengths (that is, first extract n-gram patterns of different modes, encapsulate each n-gram pattern as a capsule, and use the dynamic routing algorithm to predict and vote on the high-level capsules to obtain more effective n-gram patterns), fuse the global and local features into more comprehensive features, and use the characteristics of the capsule structure to represent the label scores in vector form, retaining the direction of the features to better represent the labels, thereby solving the problem that directly using the global Transformer has excessive noise and affects the word segmentation efficiency.

[0041] (2) Aiming at the problem that the perception range of the global Transformer is too large, resulting in some irrelevant characters and features being mis-extracted, which further affects the sequence labeling decision of the Chinese word segmentation model, a method of dynamically extracting n-gram pattern features at a single time step using the capsule network is proposed, and then fused with the features extracted by the global Transformer to reduce the impact of noise on sequence labeling. After that, the fused features are fed into the conditional random field to solve the problem of adjacent character label dependencies.

[0042] (3) The Chinese word segmentation model TC-CRF proposed by the present invention extracts global context features through the Transformer, enabling the model to dynamically and selectively focus on important words and encode long-distance dependencies, improving the generalization ability of the model. The capsule network (CapsNet) can effectively screen out task-related features through the dynamic routing mechanism, suppress the weights of irrelevant information, making the model more robust in dealing with dynamic dependencies and reducing the impact of noise on model prediction. The CRF layer is used to optimize the prediction of word boundaries. This combination enables the present invention to more accurately capture complex features in Chinese word segmentation, thereby improving the accuracy of word segmentation.

[0043] (4) The present invention integrates the global perception ability of the Transformer and the dynamic feature screening mechanism of the capsule network through a parallel architecture, avoiding the inherent sequence processing bottleneck in the RNN model, improving the training and inference speed of the model, and thus saving network resources. Since the present invention can perform Chinese word segmentation more accurately, it reduces the dependence on a large amount of labeled data, thereby reducing the risk of data leakage and abuse and improving data security. As an advanced Chinese word segmentation technology, the present invention reduces the dependence on a large amount of labeled data, reduces the development cost, and can promote the innovation and application of natural language processing technology, especially in tasks such as machine translation, sentiment analysis, and text summarization, providing strong support for the development of the NLP field. Brief Description of the Drawings

[0044] Figure 1 This is the flowchart of the method of the present invention.

[0045] Figure 2 This is the structural diagram of the Chinese word segmentation model TC-CRF of the present invention.

[0046] Figure 3 This is the structural diagram of the global transformer encoder layer used in the present invention.

[0047] Figure 4 This is the structural diagram of the capsule network of the present invention. Detailed Embodiments

[0048] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

[0049] Embodiment

[0050] As Figure 1 and Figure 2 shown, the present invention provides a Chinese word segmentation method based on deep learning, and its implementation method is as follows:

[0051] S1. For the publicly available Chinese word segmentation dataset, divide it proportionally and perform pre-training to obtain multi-dimensional features at the character level;

[0052] S2. Construct a Chinese word segmentation model TC-CRF based on deep learning, and use the Chinese word segmentation model TC-CRF to perform Chinese word segmentation processing on the multi-dimensional features.

[0053] In this embodiment, the Chinese word segmentation model TC-CRF includes:

[0054] A global transformer network for processing 768-dimensional features to obtain 256-dimensional hidden layer features, and its implementation method is as follows:

[0055] According to the 768-dimensional features, obtain the key value K, query Q, and value V through linear transformation;

[0056] Use multi-head attention calculation to divide the key value K, query Q, and value V into h parts;

[0057] Concatenate the attention calculation results of each head;

[0058] The concatenation result is transformed nonlinearly using a feedforward network to obtain 256-dimensional hidden layer features;

[0059] The capsule network is used to process the 768-dimensional features and extract the features of different n-gram patterns, where each n-gram pattern is encapsulated as a capsule, where n-gram represents a sequence of n consecutive items in the text; for example, in the sentence "the cumulative annual output of spandex is 1,500 tons" in the ctb6 dataset, the word "计" is more dependent on the left-hand 2-gram pattern, while the word "一" is more dependent on the right-hand 4-gram feature, such as Figure 2 As shown in the figure, the capsule network includes a convolutional layer, a primary capsule layer and a digital capsule layer (i.e., a dynamic routing layer), and its implementation method is as follows:

[0060] The 768-dimensional features are input into the capsule network, and convolution operations with different convolution kernel sizes are performed on the input features in the convolution layer to obtain different n-gram patterns. The convolution operation is as follows:

[0061]

[0062] in, represents the k-gram feature at position i”, σ represents the activation function, x represents the sequence, i”:i”+k represents the word vector at the position of the word from i” to i”+k in the sequence, * represents the dot product, H k represents the convolution kernel with width k, and b represents the offset;

[0063] At the primary capsule layer, each n-gram pattern is encapsulated into a capsule and sent to the dynamic routing layer, where each input capsule predicts the high-level capsule. Using the learnable matrix W, let the output of the i'th capsule in the lth layer be the vector u i' ∈R d , the high-level capsule prediction generates the prediction vector u for the high-level capsule through affine transformation j|i' :

[0064] u j|i' =W i'j u i'

[0065] Among them, W i'j Represents the learning parameters, mapping the vector of the low-level capsule i' to the space of the high-level capsule j, R d represents a vector space of size d;

[0066] Based on the prediction results of high-level capsules obtained from each low-level capsule, through affine transformation and voting on capsules composed of different n-gram patterns, the voting weights from low-level capsules to high-level capsules are continuously updated in an iterative manner. The prediction results are weighted and summed according to the weights to dynamically screen out the feature capsules useful for predicting the label scores, obtaining high-level capsules (the high-level capsules are first unpacked into multi-dimensional features, then concatenated with the output of the transformer in the feature dimension, and finally transformed through a fully connected layer to obtain the final score of the label, which is sent to the conditional random field layer). Among them, for the obtained high-level capsules, each high-level capsule represents a segmented label (such as BMES label), the norm of each capsule is the score of the label, the features inside the capsule represent the features effective for this type of label, and each feature inside the capsule is a direction for this capsule:

[0067]

[0068] Among them, u j|i represents the prediction vector of the low-level capsule i' for the high-level capsule j, c ij represents the weight, s j represents the high-level capsule;

[0069] Use the flatten function to combine the number of output capsules and the output capsule dimension into the total number of features of the capsule to complete the unpacking;

[0070] Use unpacking to restore the feature capsules output by dynamic screening to multi-dimensional features and complete the feature extraction of different n-gram patterns;

[0071] The fusion layer is used to concatenate the hidden layer features extracted by the global transformer network and the features extracted by the capsule network, and through a fully connected layer, obtain the score situation of the label;

[0072] The conditional random field layer is used to, according to the score situation of the label, model and annotate the dependencies between labels to obtain the structural information of the annotation sequence and complete the word segmentation processing of Chinese.

[0073] In this embodiment, for the Chinese word segmentation public dataset ctb6, the training set, validation set, and test set are divided according to the ratio of 8:1:1 (the training set is used for the training process and participates in backpropagation, the validation set is used to adjust hyperparameters, and the test set is used to finally evaluate the performance of the Chinese word segmentation model TC-CRF. First, a vocabulary is established to obtain words and true labels, and then the sentences are converted into instances using the markings in the dataset. During training, the true labels are separated from the training process, and during validation, the true labels are compared with the predicted labels). For the Chinese word segmentation public data, first, pre-training is performed using bert-base to obtain a 768-dimensional feature space at the character level, that is, the feature of each character is 768

[0074] In this embodiment, the feature extraction part is innovatively composed of two parallel parts: a global transformer and a capsule network. The existing transformer structure is used, and at the same time, the capsule network is improved. Different n-gram patterns are combined as different capsules, and dynamic routing is performed at a single time step to make the capsule network more adaptable to the Chinese word segmentation pattern, and to reduce the time complexity of the model and the time expenditure.

[0075] In this embodiment, as Figure 2 and Figure 3 shown, first, the obtained 768-dimensional features are respectively fed into the global transformer and the capsule network, where: the global transformer consists of a self-attention network stacked by many identical encoding layers, and each layer is composed of a multi-head attention and a feed-forward network. Using multi-head attention can improve the defect that a single head will produce random errors. Assume that the input of the self-attention network layer is the word vector M, and through linear transformation, the key (K), query (Q), and value (V) are obtained:

[0076] Q = W Q M, K = W K M, V = W V M

[0077] where, W Q 、W K and W V all represent learnable weight matrices.

[0078] Multi-head calculation divides Q, K, and V into h parts, where h is the number of heads, and the calculation attention formula for each head is:

[0079]

[0080] where, Attention() represents the attention calculation, d k represents the dimension of each head, represents the adjustment factor, h represents the number of heads, and T represents the transpose.

[0081] After obtaining the calculation result, use the concatenation method, and its formula is:

[0082] M = concat(head0,..., head h )

[0083] Among them, M represents the concatenation result, head h represents the attention calculation result of the h-th head, and concat() represents the concatenation operation.

[0084] Use a feed-forward network to perform non-linear transformation on the result. Through two identical encoder layers, with 8 heads, the output is a hidden layer feature of 256 dimensions, and it is output to the fusion layer.

[0085] In this embodiment, as Figure 2 and Figure 4 shown, the pre-trained features are simultaneously fed into the capsule network. The capsule network extracts features of different n-gram patterns. The main implementation method is to use convolutional kernels of different sizes for convolution. The convolution operation formula is:

[0086]

[0087] Among them, represents the k-gram feature at the i” position, σ represents the activation function, x represents the sequence, i”:i”+k represents the word vectors at the positions of the words in the range from i” to i”+k on the sequence, * represents the dot product, H k represents the convolutional kernel with a width of k, and b represents the offset.

[0088] In this embodiment, each n-gram pattern is encapsulated into a capsule. For example, if there are 3 convolutional kernels of different sizes, then there will be 3 input capsules. At each word, there are 8 capsules, and each capsule has 8 dimensions. The extracted capsules are fed into the dynamic routing layer. Each input capsule predicts the high-level capsules, and gradually updates the voting weights from the low-level capsules to the high-level capsules through an iterative method, dynamically screening out the feature capsules that are more useful for label score prediction. The final output capsule for each is the specific situation of a label, and the norm length of each capsule is the score of that label. Using the method of de-encapsulation, the output capsules are restored to features and output to the fusion layer.

[0089] In this embodiment, the features extracted by the global transformer and the capsule network are concatenated, and after passing through a fully connected layer, the score situation of the final label is obtained.

[0090] In this embodiment, the score situation of the label is input into the conditional random field layer. The conditional random field layer CRF assumes that the optimal labeling sequence Y conforms to the Markov property, that is, the current position yi It is only related to the annotations in adjacent positions and not to those in more distant positions. By modeling the dependencies between annotations, the conditional random field layer CRF can effectively capture the structural information of the annotation sequence, thus solving the dependency problem in the annotation sequence.

[0091] In this embodiment, the negative log-likelihood loss function of the top-level conditional random field is used as the loss function of the model to maximize the conditional probability of the true label sequence. The performance of the model is evaluated by calculating the total probability of all possible paths and the score of the true path:

[0092]

[0093] Among them, Loss represents the loss function of the Chinese word segmentation model TC-CRF, N represents the total number of samples, p(y i |x i ) represents?y i represents the true label of the i-th sample, and x i represents the input of the i-th sample.

[0094] In summary, the present invention provides a method for dynamically extracting the features of Chinese word segmentation by a global transformer and a capsule network in parallel, which solves the problem that the existing method of feature extraction by convolution cannot obtain long dependencies, as well as the problem of too much noise obtained by the global transformer, and effectively realizes Chinese word segmentation using deep learning.

Claims

1. A Chinese word segmentation method based on deep learning, characterized in that, It includes the following steps: For the publicly available Chinese word segmentation dataset, divide it proportionally and perform pre-training to obtain multi-dimensional features at the character granularity; Construct a Chinese word segmentation model TC-CRF based on deep learning, and use the Chinese word segmentation model TC-CRF to perform Chinese word segmentation on the multi-dimensional features.

2. The Chinese word segmentation method based on deep learning according to claim 1, characterized in that The Chinese word segmentation model TC-CRF includes: A global transformer network for processing 768-dimensional features to obtain 256-dimensional hidden layer features; A capsule network for processing 768-dimensional features to extract features of different n-gram patterns, where each n-gram pattern is encapsulated into a capsule, and n-gram represents a sequence of n consecutive items in the text; A fusion layer for concatenating the hidden layer features extracted by the global transformer network and the features extracted by the capsule network, and passing through a fully connected layer to obtain the score of the label; A conditional random field layer for obtaining the structural information of the annotation sequence by modeling the dependency relationship between the annotation labels according to the score of the label, and completing the Chinese word segmentation process.

3. The Chinese word segmentation method based on deep learning according to claim 2, wherein The processing of the 768-dimensional features to obtain 256-dimensional hidden layer features is specifically as follows: According to the 768-dimensional features, linearly transform to obtain the key value K, query Q, and value V; Use multi-head attention calculation to divide the key value K, query Q, and value V into h parts, where the attention calculation for each head is as follows: Among them, Attention() represents attention calculation, and d k represents the dimension of each head, represents the adjustment factor, h represents the number of heads, and T represents transpose; Concatenate the attention calculation results of each head: M = concat(head0,..., head h ) Among them, M represents the splicing result, and head h represents the attention calculation result of the h-th head, and concat() represents the splicing operation; Use a feed-forward network to perform a non-linear transformation on the concatenated result to obtain 256-dimensional hidden layer features.

4. The Chinese word segmentation method based on deep learning according to claim 2, characterized in that, The extraction of features of different n-gram patterns is specifically as follows: Input the 768-dimensional features into the capsule network, and perform convolution operations with different convolution kernel sizes on the input features in the convolutional layer to obtain different n-gram patterns, where the convolution operation is as follows: Among them, represents the k-gram feature at the i" position, σ represents the activation function, x represents the sequence, i":i"+k represents the word vectors at the positions of the words in the range from i" to i"+k on the sequence, * represents dot product, H k represents the convolutional kernel with a width of k, and b represents the offset; In the primary capsule layer, each n-gram pattern is encapsulated into a capsule and the capsules are sent to the dynamic routing layer, where each input capsule makes a prediction for the higher-level capsules. Using the learnable matrix W, let the output of the i'-th capsule in the l-th layer be the vector u i' ∈R d , and the prediction of the higher-level capsules generates the prediction vector u for the higher-level capsules through an affine transformation j|i' : u j|i' = W i'j u i' Among them, W i'j represents the learning parameter that maps the vector of the lower-level capsule i' to the space of the higher-level capsule j, and R d represents a vector space of size d; Based on the prediction results of the high-level capsules obtained from each low-level capsule, through affine transformation and voting on the capsules composed of different n-gram patterns, continuously update the voting weights from the low-level capsules to the high-level capsules in an iterative manner, and perform weighted summation on the prediction results according to the weights to dynamically screen out the feature capsules useful for predicting the label scores to obtain high-level capsules. For the obtained high-level capsules, each high-level capsule represents a label of a word segment, the norm length of each capsule is the score of the label, the features inside the capsule represent the features effective for this type of label, and each feature inside the capsule is a direction for this capsule: Among them, u j|i represents the prediction vector of the lower-level capsule i' for the higher-level capsule j, and c ij represents the weight, and s j represents the higher-level capsule; Use the flatten function to combine the number of output capsules and the output capsule dimension into the total number of capsule features to complete the de-encapsulation; Use de-encapsulation to restore the feature capsules dynamically screened and output to multi-dimensional features to complete the extraction of features of different n-gram patterns.

5. The Chinese word segmentation method based on deep learning according to claim 2, wherein The expression of the loss function of the Chinese word segmentation model TC-CRF is as follows: Among them, Loss represents the loss function of the Chinese word segmentation model TC-CRF, N represents the total number of samples, p(y i |x i ) represents the probability that the sample belongs to the true label y i under the condition of the given input x i , y i represents the true label of the i-th sample, and x i represents the input of the i-th sample.