Text tokenization method, device, equipment and storage medium

By using the weight and confusion of the tag sequence in low resource scenarios for decoding, the problem that the word segmentation model cannot effectively predict unlogged words in low resource scenarios is solved, and the robustness and accuracy of word segmentation results are improved.

CN114218939BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111530194.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-06-10
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

In low-resource scenarios, word segmentation models usually fail to effectively predict unlogged words, resulting in inaccurate word segmentation results.

Method used

By obtaining multiple tag sequences corresponding to the text to be participled, encoding process is performed to obtain the weight of each tag sequence, and decoding the target tag sequence according to the weight and confusion degree to determine the word segmentation result.

Benefits of technology

In low resource scenarios, the robustness and accuracy of word segmentation effects are improved, the demand for a large number of computing resources is avoided, and it is suitable for industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218939B_ABST
    Figure CN114218939B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a text word segmentation method, apparatus, device, and storage medium. In the word segmentation decoding stage, the perplexity of each tag sequence is obtained according to the probability of each word segment in the word segmentation sequence corresponding to each tag sequence appearing under the condition that its previous word segment appears, so as to evaluate the rationality of the tag sequence. Combining with the weight corresponding to each tag sequence learned in the word segmentation encoding stage, the optimal tag sequence is selected, thereby ensuring the robustness and accuracy of the word segmentation effect in low-resource scenarios. Compared with the related art, the technical solution of the present application can complete word segmentation without a large amount of computing resources, is applicable to industrial scenarios, and does not need to promote the word segmentation performance of the word segmentation model by stacking models of multiple tasks, but realizes the above technical effects by improving the encoding algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of natural language processing, and in particular, to a text tokenization method, apparatus, device, and storage medium. Background Art

[0002] Text tokenization is a basic task in the field of natural language processing. As the basis for other natural language processing tasks, the tokenization model based on the neural network model needs to maintain good robustness in any scenario. However, for low-resource scenarios lacking training data, the tokenization model usually cannot well predict out-of-vocabulary words, resulting in inaccurate tokenization results.

[0003] For the above-mentioned low-resource scenarios, related technologies adopt models that stack multiple tasks to make up for the deficiencies of a single tokenization model. Essentially, the multi-task model learns the common knowledge between different standard data sets and the unique knowledge of a single data set, and then aggregates these two parts of knowledge to improve the robustness of the tokenization model and the accuracy of the tokenization results. However, this means that a powerful feature extraction layer is required to balance these two parts of information, the training requirements are high, and the obtained model is also difficult to achieve the expected goal. Summary of the Invention

[0004] The present disclosure provides a text tokenization method, apparatus, device, and storage medium to at least solve the problem in related technologies that for low-resource scenarios lacking training data, the tokenization model usually cannot well predict out-of-vocabulary words, resulting in inaccurate tokenization results. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a text tokenization method, including: obtaining a plurality of tag sequences corresponding to the text to be tokenized, where the tag sequences are used to segment the text to be tokenized into corresponding token sequences; performing an encoding process on the text to be tokenized to obtain weights corresponding to each of the tag sequences; decoding according to the weights and perplexity corresponding to each of the tag sequences to obtain a target tag sequence, so as to determine a tokenization result by using the target tag sequence; where the perplexity corresponding to each of the tag sequences is determined according to the statistical probability that each token appears under the condition that the previous token appears in the token sequence corresponding to the tag sequence.

[0006] Combined with the first aspect, in a possible implementation of the first aspect, decoding the target label sequence according to the weights and perplexities corresponding to each of the label sequences includes: for the word segmentation sequence corresponding to each of the label sequences, determining the statistical probability of the occurrence of the first word in the word segmentation sequence and the conditional statistical probability of each of the remaining words, where the conditional statistical probability of the k-th word refers to the statistical probability of the occurrence of the k-th word under the condition that the (k - 1)-th word appears, and k is a positive integer greater than 1; determining the perplexity of the label sequence according to the statistical probability of the occurrence of the first word and the conditional statistical probabilities of each of the remaining words; and determining the target label sequence according to the weights and perplexities corresponding to each of the label sequences.

[0007] Combined with the first aspect, in a possible implementation of the first aspect, determining the statistical probability of the occurrence of the first word in the word segmentation sequence and the conditional statistical probability of each of the remaining words includes: determining the probability of the occurrence of the first word in the preset corpus as the statistical probability of the occurrence of the first word; determining the probability of the consecutive occurrence of two adjacent words in the preset corpus among all the remaining words, and the probability of the occurrence of each word in the preset corpus; and determining the conditional statistical probability of the k-th word according to the probability of the consecutive occurrence of the k-th word and the (k - 1)-th word, and the probability of the occurrence of the (k - 1)-th word.

[0008] Combined with the first aspect, in a possible implementation of the first aspect, the preset corpus is constructed based on the standard corpus in the target domain and the associated corpus of the standard corpus. The associated corpus of the standard corpus refers to the corpus that is labeled according to the same character label annotation rule as the standard corpus, and the associated corpus belongs to a non-target domain. The character label annotation rule refers to the rule for annotating character labels for the text in the corpus, and the character labels are used to annotate the position of each character in the word segmentation of the text.

[0009] Combined with the first aspect, in a possible implementation of the first aspect, both the standard corpus and the associated corpus include a number of texts, the texts include a number of word segments, and the characters in the word segments are annotated with character labels; the number of the same word segments in the standard corpus and the associated corpus is greater than the preset number, and the character labels annotated for the characters in the same word segment in the standard corpus and in the associated corpus are the same.

[0010] Combined with the first aspect, in a possible implementation of the first aspect, the label sequence includes character labels corresponding to each character in the text to be segmented, and the character labels are used to annotate the position of the character in the word segmentation; the weight corresponding to the label sequence includes the emission weight and the state transition weight corresponding to each character label in the label sequence.

[0011] In combination with the first aspect, in a possible implementation manner of the first aspect, decoding the target label sequence according to the weights and perplexities corresponding to each of the label sequences includes: determining a first weight sum and a second weight sum corresponding to each label sequence, where the first weight sum is the sum of the emission weights corresponding to each character label in the label sequence, and the second weight sum is the sum of the state transition weights corresponding to each character label in the label sequence; determining a score of the label sequence according to the first weight sum, the second weight sum, and the perplexity corresponding to the label sequence; and determining the label sequence corresponding to the maximum score as the target label sequence.

[0012] In combination with the first aspect, in a possible implementation manner of the first aspect, encoding the text to be segmented to obtain weights corresponding to each of the label sequences includes: inputting the text to be segmented into an encoding model trained using the preset corpus, and outputting the emission weights and the state transition weights corresponding to each character label in each of the label sequences.

[0013] According to a second aspect of the embodiments of the present disclosure, there is provided a text segmentation device, including: a label sequence acquisition unit that acquires a plurality of label sequences corresponding to a text to be segmented, where the label sequences are used to segment the text to be segmented into corresponding segmented sequences; an encoding unit that encodes the text to be segmented to obtain weights corresponding to each of the label sequences; a decoding unit that decodes to obtain a target label sequence according to the weights and perplexities corresponding to each of the label sequences, so as to determine a segmentation result using the target label sequence; where the perplexity corresponding to each of the label sequences is determined according to the statistical probability of each segmented word appearing under the condition that the previous segmented word appears in the segmented sequence corresponding to the label sequence.

[0014] In combination with the second aspect, in a possible implementation manner of the second aspect, the decoding unit is specifically configured to: for the segmented sequence corresponding to each of the label sequences, determine the statistical probability of the first segmented word appearing in the segmented sequence and the conditional statistical probability of each of the remaining segmented words, where the conditional statistical probability of the kth segmented word refers to the statistical probability of the kth segmented word appearing under the condition that the (k - 1)th segmented word appears, and k is a positive integer greater than 1; determine the perplexity of the label sequence according to the statistical probability of the first segmented word appearing and the conditional statistical probabilities of each of the remaining segmented words; and determine the target label sequence according to the weights and perplexities corresponding to each of the label sequences.

[0015] In combination with the second aspect, in a possible implementation manner of the second aspect, the decoding unit is specifically configured to: determine the probability of occurrence of the first word segment in a preset corpus as the statistical probability of the occurrence of the first word segment; determine the probability of consecutive occurrence of two adjacent word segments among all the remaining word segments in the preset corpus, and the probability of occurrence of each word segment in the preset corpus; and determine the conditional statistical probability of the k-th word segment according to the probability of consecutive occurrence of the k-th word segment and the (k - 1)-th word segment, and the probability of occurrence of the (k - 1)-th word segment.

[0016] In combination with the second aspect, in a possible implementation manner of the second aspect, the preset corpus is constructed based on a standard corpus in the target domain and a related corpus of the standard corpus. The related corpus of the standard corpus refers to a corpus that is annotated according to the same character tag annotation rule as the standard corpus, and the related corpus belongs to a non-target domain. The character tag annotation rule refers to a rule for annotating character tags to the text in the corpus, and the character tags are used to annotate the position of each character in the word segmentation of the text.

[0017] In combination with the second aspect, in a possible implementation manner of the second aspect, both the standard corpus and the related corpus include a number of texts, and each text includes a number of word segments, and the characters in the word segments are annotated with character tags; the number of identical word segments in the standard corpus and the related corpus is greater than a preset number, and the character tags annotated to the characters in the same word segment in the standard corpus and in the related corpus are the same.

[0018] In combination with the second aspect, in a possible implementation manner of the second aspect, the tag sequence includes character tags corresponding to each character in the text to be segmented, and the character tags are used to annotate the position of the character in the word segmentation; the weight corresponding to the tag sequence includes the emission weight and the state transition weight corresponding to each character tag in the tag sequence.

[0019] In combination with the second aspect, in a possible implementation manner of the second aspect, the decoding unit is specifically configured to: determine the first weight sum and the second weight sum corresponding to each tag sequence, where the first weight sum is the sum of the emission weights corresponding to each character tag in the tag sequence, and the second weight sum is the sum of the state transition weights corresponding to each character tag in the tag sequence; determine the score of the tag sequence according to the first weight sum, the second weight sum, and the perplexity corresponding to the tag sequence; and determine the tag sequence corresponding to the maximum score as the target tag sequence.

[0020] In combination with the second aspect, in a possible implementation manner of the second aspect, the encoding unit is specifically configured to: input the text to be segmented into an encoding model trained by using the preset corpus, and output the emission weight and the state transition weight corresponding to each character label in each label sequence.

[0021] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the text segmentation method provided in the first aspect and any of its possible design manners.

[0022] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of a server, enabling the server to execute the text segmentation method provided in the first aspect and any of its possible design manners.

[0023] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes computer instructions, when the computer instructions run on a server, enabling the server to execute the text segmentation method provided in the first aspect and any of its possible design manners.

[0024] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In the segmentation decoding stage, the perplexity of a certain label sequence is obtained according to the probability of each segmented word in the segmented word sequence corresponding to the label sequence appearing under the condition that its previous segmented word appears, so as to be used to evaluate the rationality of the label sequence, and combined with the weight corresponding to each label sequence learned in the segmentation encoding stage, the optimal label sequence, that is, the target label sequence, is selected, thereby ensuring the robustness and accuracy of the segmentation effect in a low-resource scenario. Compared with the related art, the technical solution of the present application can complete segmentation without a large amount of computing resources, is applicable to industrial scenarios, and does not need to stack models of multiple tasks to promote the segmentation performance of the segmentation model, but realizes the above technical effects by improving the encoding algorithm.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0027] Figure 1 is a schematic diagram of a text segmentation system shown according to an exemplary embodiment;

[0028] Figure 2 It is a flowchart of a text word segmentation method shown according to an exemplary embodiment;

[0029] Figure 3 It is a schematic diagram of an encoding model structure shown according to an exemplary embodiment;

[0030] Figure 4 It is a block diagram of a text word segmentation device shown according to an exemplary embodiment;

[0031] Figure 5 It is a schematic diagram of a server structure shown according to an exemplary embodiment. Detailed implementation manners

[0032] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0034] In addition, in the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present disclosure, "a plurality of" means two or more.

[0035] A text is a continuous sequence of characters. Text word segmentation is the process of recombining a continuous sequence of characters into a sequence of words according to certain rules. In English writing, spaces are used as natural delimiters between words, while in Chinese, only characters, sentences, and paragraphs can be simply delimited by obvious delimiters, and there is no formal delimiter for words. Therefore, it is necessary to use certain rules to identify the boundaries between words in the text, and the process of identifying the boundaries between words in the text and finally obtaining a sequence of words is the text word segmentation process.

[0036] In the disclosed embodiment, the word sequence obtained by segmenting the text is called a segmentation sequence, and the segmentation sequence includes multiple segmentations, and the arrangement order of these segmentations is consistent with the position order in the text. For example, the segmentation of the text "This is my favorite song" can obtain the segmentation sequence [This / is / my / most / favorite / song].

[0037] The text segmentation method provided in the embodiment of the present disclosure can be applied to a text segmentation system, which is used to segment text to be processed. Figure 1 A diagram of a text segmentation system is shown in Figure 1. Figure 1 As shown, the text segmentation system includes a text segmentation device 11 for performing segmentation processing on a text to be processed and a server 12 which is communicatively connected to the text segmentation device 11 .

[0038] The text segmentation device 11 is used to execute the text segmentation method provided by the embodiment of the present disclosure, so as to segment the text to be processed into a sequence of segments through the text segmentation method. For example, multiple label sequences corresponding to the text to be segmented are obtained, and the label sequences are used to segment the text to be segmented into corresponding segmentation sequences; the text to be segmented is encoded to obtain weights corresponding to each of the label sequences; according to the weights and confusion corresponding to each of the label sequences, a target label sequence is decoded to determine the segmentation result using the target label sequence; wherein the confusion corresponding to each of the label sequences is determined according to the statistical probability of each segmentation appearing in the segmentation sequence corresponding to the label sequence under the condition that the previous segmentation appears.

[0039] It should be noted that the above characters refer to the individual characters that make up the words in the text. It should be understood that for words containing multiple individual characters, the position of each individual character in the word is different, that is, the word position of each individual character is different. The word position includes the first character, the middle character, and the last character, and for the characters that form a word alone, it is not necessary to distinguish their positions in the word by position. The above character labels refer to labels used to mark the positions of individual characters in the word. In an exemplary label system, the character labels can be B, M, E, and S, wherein B, M, and E represent the first character, the middle character, and the last character of the word, respectively, and S represents the single character word. Exemplarily, if the word "喜" in a text is marked with the label "B", it means that "喜" is the first character of a certain word. According to the labels marked for each word in the text, the text can be segmented into word sequences.

[0040] It should be noted that the above "B, M, E, S" are merely an exemplary tag system provided by the embodiments of the present disclosure. In other embodiments, words in the text can be labeled based on different tag systems. In one exemplary tag system, the character tags can be B and I, where B represents the first character of a word, and I represents the other characters (non-first characters) of a word. In another exemplary tag system, the character tags can be S, B, M1, M2, M, E, S, where B represents the first character of a word, M1 / M2 / M represents the middle characters of a word, E represents the ending character of a word, and S represents a single-character word. It should be understood that the tag system adopted can be a standard tag system in the field of natural language processing or a custom tag system to meet the requirements of the scenario.

[0041] The text tokenization device 11 can interact with the server 12 for data. For example, the text tokenization device 11 can obtain the text to be tokenized from the server 12. Also, for example, the text tokenization device 11 can send the tokenization result obtained by tokenizing the text to be tokenized to the server 12.

[0042] The server 12 can be a single server, or alternatively, can be a server cluster composed of multiple servers or a cloud computing service center, and the present disclosure does not make any limitation thereto. The server 12 is used to collect the text to be tokenized, such as receiving the text uploaded from the user side. In addition, the server 12 can also be used to receive the text tokenization result sent by the text tokenization device 11. The server 12 can also complete other natural language processing tasks based on the text tokenization result. For example, establish a thesaurus based on the text tokenization result, calculate the semantic similarity between texts, or search for other texts similar to a certain text.

[0043] It should be noted that the text tokenization device 11 and the server 12 can be independent devices, or can be integrated into the same device, and the present invention does not make specific limitations thereto.

[0044] In some embodiments, the text tokenization device 11 can be an electronic device or be included in an electronic device. Here, the electronic device includes but is not limited to various computer devices such as mobile phones, tablet computers, desktop computers, laptop computers, vehicle-mounted terminals, handheld terminals, augmented reality (AR) devices, virtual reality (VR) devices, etc. The embodiments of the present disclosure do not impose special restrictions on the specific form of the electronic device.

[0045] When the text tokenization device 11 and the server 12 are integrated into the same device, the communication method between the text tokenization device 11 and the server 12 is the communication between internal modules of the device. In this case, the communication process between the two is the same as "the communication process between the text tokenization device 11 and the server 12 when they are independent of each other".

[0046] In the following embodiments provided by the present invention, the text tokenization device 11 and the server 12 are mainly taken as an example where they are independently arranged for illustration.

[0047] The text tokenization method provided by the embodiments of the present disclosure can also be applied to various natural language processing methods related to the back-end implementation of various business scenarios, including but not limited to: lexical analysis methods such as new word discovery, part-of-speech tagging, and spelling correction, syntactic analysis methods such as constituent syntactic analysis, dependency syntactic tokenization, and sentence boundary detection, semantic analysis methods such as semantic disambiguation and semantic role labeling, and information extraction methods such as named entity recognition, entity disambiguation, sentiment analysis, and intention recognition. The text tokenization method provided by the embodiments of the present disclosure can be used as a basic link in the aforementioned natural language processing methods. For example, in the part-of-speech tagging method, the text tokenization method provided by the embodiments of the present disclosure is used to tokenize the text to be tagged to obtain a token sequence, and then the part-of-speech tagging rules are used to tag the words in the token sequence with part-of-speech tags. Another example is in the named entity recognition method, the text tokenization method provided by the embodiments of the present disclosure is used to tokenize the text to be processed to obtain a token sequence, and then the entity recognition rules are used to identify the named entities in the token sequence.

[0048] In some embodiments, a certain scale of training corpus is used to train an initial tokenization model based on a neural network model, and the model parameters are continuously optimized to enable the model to have the ability to learn the feature knowledge of the text and tokenize the text according to the learned features. After the tokenization model is trained, the text to be tokenized is input into the tokenization model to obtain a corresponding token sequence using the tokenization model. Among them, the text tokenization process can be divided into a tokenization encoding stage and a tokenization decoding stage. The tokenization encoding stage can be understood as the stage of extracting features from the text to be tokenized using the encoding layer of the model, and the tokenization decoding stage can be understood as the stage of processing the extracted features using a decoding algorithm to determine an optimal tokenization result.

[0049] As mentioned above, since text tokenization is a basic task in the field of natural language processing, and as the basis for other natural language processing tasks, it is required that the tokenization model needs to maintain good robustness in any scenario. However, for low-resource scenarios lacking training data, the tokenization model usually cannot well predict out-of-vocabulary words, resulting in inaccurate tokenization results.

[0050] For the above low-resource scenarios, related technologies train pre-trained models for word segmentation tasks in specific scenarios, thereby reducing the corpus data required for training. Alternatively, the deficiencies of a single word segmentation model are compensated by stacking models for multiple tasks, promoting the word segmentation performance of the word segmentation model. It can be seen that the foregoing solutions all improve the word segmentation encoding stage. However, since pre-trained models are often large and complex models, this means that when applying them to actual scenarios, a large amount of computing resources are required, so the applicable scenarios are limited, such as industrial scenarios with insufficient computing resources. The multi-task model essentially improves the robustness of the word segmentation model and the accuracy of the word segmentation results by learning the common knowledge between different standard data sets and the unique knowledge of a single data set, and then aggregating these two parts of knowledge, which means that a powerful feature extraction layer is required to balance these two parts of information, the training requirements are high, and the obtained model is also difficult to achieve the expected goal.

[0051] Embodiments of the present disclosure provide a text word segmentation method. This method introduces the perplexity information of the word segmentation sequence corresponding to each tag sequence in the word segmentation decoding stage. This perplexity information is an index used to evaluate the rationality of the word segmentation sequence, thereby enriching the basis for word segmentation decoding, achieving a good word segmentation effect without a large amount of computing resources, thus being applicable to industrial scenarios, and not promoting the word segmentation performance of the word segmentation model by stacking models for multiple tasks, thus not bringing additional training difficulties.

[0052] Figure 2 The flowchart of a text word segmentation method shown in an exemplary embodiment of the present disclosure is as Figure 2 shown, and this method may include:

[0053] S201, obtain multiple tag sequences corresponding to the text to be word segmented, where the tag sequences are used to segment the text to be word segmented into corresponding word segmentation sequences.

[0054] In the embodiments of the present disclosure, the text to be word segmented may be a sentence, a paragraph of text, etc. The characters in the text to be word segmented refer to single characters that can form words. The embodiments of the present disclosure are not limited to the number of characters in the text.

[0055] The tag sequence is composed of multiple character tags with a specific arrangement order. The number of character tags in the tag sequence is the same as the number of characters in the text to be word segmented and corresponds one by one to the characters included in the text to be word segmented. According to a given tag sequence, a word segmentation result can be obtained, that is, a word segmentation sequence can be obtained.

[0056] Actually, according to a preset tag system, the text to be word segmented corresponds to a tag matrix, and the tag matrix includes multiple tag sequences. For example, taking the text to be word segmented "I like listening to music" as an example, the corresponding tag matrix may be:

[0057]

[0058] Without constraint conditions, 4 6 tag sequences can be obtained based on the above tag matrix, such as [B B BB B B], [B M B M B B], etc., which are not listed one by one here. Different tag sequences mean different word segmentation results. For example, the word segmentation sequence corresponding to [S S S S S S] is [I / like / to / listen / to / music], while the word segmentation sequence corresponding to [S S S S B E] is [I / like / to / listening / to / music].

[0059] It is easy to understand that these 4 6 tag sequences include obviously unreasonable tag sequences, such as [B B B B B B]. Each tag sequence is evaluated based on a specific algorithm to obtain the evaluation score of each tag sequence. Finally, the optimal tag sequence can be obtained, and then the most reasonable word segmentation sequence can be obtained.

[0060] It should be noted that according to different adopted tag systems, the tag matrix corresponding to the same text to be segmented is different. The above example shows the tag matrix corresponding to the text to be segmented when using "B, M, E, S". When using "B, I", the tag matrix corresponding to the text to be segmented "I like to listen to music" is as follows:

[0061]

[0062] Without constraint conditions, 2 6 tag sequences can be obtained based on the above tag matrix, which are not listed one by one here. The word segmentation sequences obtained based on different tag sequences are also different.

[0063] In S201, the tag sequence corresponding to the obtained text to be segmented can be all the tag matrices included in the tag matrix corresponding to the text to be segmented, or the tag sequences whose rationality in all tag matrices meets certain conditions. For example, the tag matrix is processed based on a certain filtering rule to filter out unreasonable tag sequences and obtain tag sequences whose rationality meets certain conditions.

[0064] S202, perform encoding processing on the text to be segmented to obtain the weight corresponding to each tag sequence.

[0065] In some embodiments, by inputting the text to be segmented into an encoding model trained with a preset corpus, an emission weight matrix and a state transition weight matrix are output. The emission weight matrix and the state transition weight matrix have the same dimension as the label matrix corresponding to the text to be segmented. The emission weight matrix includes the emission weight of each character label for each character, that is, includes the emission weight corresponding to each character label in each label sequence. The state transition weight matrix includes the probability of transitioning from one character label to another character label, that is, includes the state transition weight corresponding to each character label in each label sequence, that is, the probability of this character label transitioning to the next character label.

[0066] It can be seen that in some implementation scenarios, by processing the text sequence to be segmented through the encoding model, the emission weight and the state transition weight corresponding to each character label in each label sequence can be output. Without changing the encoding model, in the word segmentation decoding stage, the present disclosure combines the perplexity corresponding to a certain label sequence with the emission weight and the state transition weight corresponding to each character label in each label sequence learned in the word segmentation encoding stage to determine the optimal label sequence, thereby ensuring the robustness and accuracy of the word segmentation effect in low-resource scenarios.

[0067] In some possible implementation manners, first, the character vectors of each character in the text to be segmented are obtained to obtain a character matrix corresponding to the text to be segmented. The character matrix includes the character vectors of each character in the text; then, using a preset encoding model, feature vectors are extracted from the character matrix, and an emission weight matrix is output according to the extracted feature vectors.

[0068] Specifically, in implementation, balanced corpus related to a specific scenario or field can be collected in advance, and the collected balanced corpus is preprocessed to filter out useless data, low-frequency words, and meaningless characters to obtain training data. Then, the preset model is trained using the training data to obtain a character vector model. Among them, the preset model can be a Skip-gram model; finally, a mapping dictionary of character vectors can be generated according to the character vector model. The mapping dictionary includes the mapping relationship between characters and character vectors. When it is necessary to obtain the character vectors of each character in the text to be segmented, the mapping dictionary of character vectors can be obtained first, and the character vectors of each character can be found from the mapping dictionary.

[0069] See Figure 3 , in some possible implementation manners, the preset encoding model may include: a Convolutional Neural Networks (CNN), a BiLSTM composed of two Long Short-Term Memory (LSTM) networks with opposite time series directions, and an output layer.

[0070] Among them, the feature vector of the character matrix can be obtained through a Convolutional Neural Network (CNN). CNN is a feedforward neural network, whose artificial neurons can respond to the surrounding units within a part of the coverage range, and can be applied to the field of natural language processing to achieve local connection, weight sharing, etc., and can effectively extract features. CNN includes a convolutional layer and a pooling layer. The convolutional layer is a feature extraction layer, and the input of each neuron is connected to the local receptive field of the previous layer and extracts the features of this local area. Once the features of this local area are extracted, the positional relationship between the extracted features and other features is also determined. The pooling layer is a feature mapping layer. Each computational layer of the network consists of multiple feature maps, and each feature map is a plane, and the weights of all neurons on the plane are equal. The feature mapping structure uses the Sigmoid function as the activation function of CNN, making the feature mapping have displacement invariance. In addition, since the neurons on a mapping surface share weights, the number of free parameters of the network is reduced.

[0071] When generating the emission weight matrix according to the feature vector, the feature vector can be input into two Long Short-Term Memory (LSTM) networks respectively to obtain the output vectors generated by each time node of the two LSTM networks within a preset time period, splice the output vectors formed at each time node to generate a spliced vector, transmit the spliced vector to the output layer, and synthesize the vector output by the output layer into the emission weight matrix. Among them, LSTM is an extension of the Recurrent Neural Network (RNN). The basic unit of the LSTM network can realize the memory function of information, and can control the memory, forgetting, and output of historical information through three structures: an input gate, a forget gate, and an output gate. It has a long-term memory function and can perfectly solve the problem of long-distance dependence.

[0072] The transition probability between character labels refers to the probability that a certain character label appears after another character label, that is, the probability of transitioning from another character label to this character label.

[0073] In some embodiments, the transition probability between character tags can be the number of times two character tags appear adjacent to each other before and after in the dataset divided by the total number of times the character tag that appears in the front appears in the dataset. For example, each character in each text in the dataset is labeled with a character tag, and then the total number of times each type of character tag appears in the dataset and the number of times any two character tags appear adjacent to each other before and after can be counted. Suppose the total number of times tag B appears in the dataset is 100, and the number of times tag B and tag I appear consecutively in the dataset with tag B in the front is 60, then the probability of tag B transitioning to tag I is 0.6. Furthermore, based on the dataset, a state transition weight matrix corresponding to the tag matrix can be obtained, and the state transition weight matrix includes the transition probabilities between character tags in each tag sequence.

[0074] In some implementation manners, when the above-mentioned preset encoding model outputs the emission weight matrix, it also outputs the state transition weight matrix.

[0075] S203. Decode to obtain the target tag sequence according to the weights and perplexities corresponding to each of the tag sequences, so as to determine the word segmentation result by using the target tag sequence; wherein, the perplexity corresponding to each tag sequence is determined according to the statistical probability of each word segment appearing under the condition that the previous word segment appears in the word segmentation sequence corresponding to the tag sequence.

[0076] Based on the Markov assumption, the possibility of the current word appearing depends on the previous word. Then, for a certain word segmentation sequence, except for the first word, the possibility of each of the remaining word segments appearing depends on its previous word segment. And if the probability of the kth word determined based on the (k - 1)th word is greater, it indicates that the kth word is more reasonable; conversely, if the probability of the kth word determined based on the (k - 1)th word is smaller, it indicates that the kth word is less reasonable. Further, based on the reasonableness of each word segment in the word segmentation sequence, the reasonableness of the word segmentation sequence can be obtained, and the reasonableness of the word segmentation sequence is equivalent to the reasonableness of the tag sequence.

[0077] Based on this, in the embodiments of the present disclosure, for the word segmentation sequence corresponding to each possible tag sequence, the probability of each word segment appearing under the condition that its previous word segment appears is used to obtain the perplexity of the tag sequence, so as to judge the reasonableness of the tag sequence based on the obtained perplexity, and in combination with the emission weight and state weight corresponding to each character tag in the tag sequence, the optimal tag sequence is determined, enriching the basis for word segmentation decoding and solving the problem of poor robustness and accuracy of the model word segmentation effect due to lack of training data in low-resource scenarios.

[0078] In some possible implementations, S203 may specifically include: for each token sequence corresponding to a tag sequence, determining the statistical probability of the first token appearing in the token sequence and the conditional statistical probability of each of the remaining tokens, where the conditional statistical probability of the k-th token refers to the statistical probability of the k-th token appearing under the condition that the (k - 1)-th token appears, the k-th token being any one of all the remaining tokens, and k being a positive integer greater than 1; determining the perplexity of the tag sequence according to the statistical probability of the first token appearing and the conditional statistical probabilities of each of the remaining tokens; and determining the target tag sequence according to the weight and perplexity corresponding to each tag sequence.

[0079] Exemplarily, the statistical probability of each token in the token sequence appearing under the condition that the previous token appears can be expressed as P(w k |w k-1 ), where w k-1 represents the (k - 1)-th token in the token sequence, w k represents the k-th token in the token sequence, k belongs to [2, n], and n represents the number of tokens in the token sequence.

[0080] In possible implementations, the product of the statistical probability of the first token appearing in the token sequence and the conditional statistical probabilities of each of the tokens from the 2nd to the n-th token can be used as the perplexity of the token sequence, that is, the perplexity of the tag sequence, as shown in the following formula 1:

[0081]

[0082] where PP(S) represents the perplexity of the tag sequence; P(w 1 ) represents the statistical probability of the first token appearing in the token sequence corresponding to the tag sequence; P(w k |w k-1 ) represents the conditional statistical probability of the k-th token in the token sequence.

[0083] In some implementations, for the token sequence corresponding to each tag sequence, the statistical probability of the first token appearing can be determined by determining the probability of the first token appearing in the preset corpus; by determining the probability of two adjacent tokens appearing consecutively in the preset corpus among all the remaining tokens, and the probability of each token appearing in the preset corpus; according to the probability of the k-th token and the (k - 1)-th token appearing consecutively, and the probability of the (k - 1)-th token appearing, the conditional statistical probability of the k-th token is determined. In the embodiments of the present disclosure, the probability of the first token in the token sequence appearing in the preset corpus and the conditional statistical probability of each of the remaining tokens in the preset corpus are introduced into the token decoding stage, and the perplexity information for evaluating the rationality of the token sequence is obtained, providing a new basis for the token decoding algorithm. Without changing the encoding model, the robustness and accuracy of the model's tokenization effect in low-resource scenarios can be improved.

[0084] Exemplarily, the conditional statistical probability of the k-th token can be determined by the following formula 2:

[0085]

[0086] where p(w k-1 , w k ) represents the probability that the k-th token appears after the (k - 1)-th token in the preset corpus, that is, the probability that the k-th token and the (k - 1)-th token appear consecutively;

[0087] p(w k-1 ) represents the probability that the (k - 1)-th token appears in the preset corpus.

[0088] In a possible implementation, a preset word library can be obtained according to the preset corpus. The preset word library includes a large number of words and the probability of each word appearing and the probability of each word appearing adjacent to other words in sequence based on the preset word library. Among them, the probability of each word appearing is determined according to the word frequency of each word in the word library and the total number of words in the word list, and the probability of each word appearing adjacent to other words in sequence is determined according to the number of times each word appears adjacent to other words in sequence and the total number of words in the word library.

[0089] In a possible implementation, the probability of each word appearing alone and the probability of it appearing adjacent to other words in sequence can be determined according to one or more preset word lists. For example, a first word list is constructed in advance according to the probability of each word appearing to record the correspondence between the word and the probability of its appearance, and a second word list is constructed according to the probability of each word appearing adjacent to other words in sequence to record the probability of the word appearing adjacent to other words in sequence. Thus, when the probability of a certain word appearing and the probability of it appearing adjacent to other words in sequence are required, the first word list and the second word list can be queried.

[0090] In the embodiments of the present disclosure, the preset corpus is pre-generated according to the standard corpus in the target field. Taking the target field as the search field as an example, the standard corpus in the target field can be the labeled search texts, and these labeled search texts constitute the standard corpus. These search texts can be from any search platform, which is not limited herein.

[0091] Considering the situation of insufficient standard corpus in some specific fields, in order to ensure that the data volume in the preset dataset is sufficient, in some other possible implementation manners, the preset corpus can be constructed according to the standard corpus in the target field and the associated corpus of the standard corpus. The associated corpus of the standard corpus refers to the corpus labeled based on the same character label annotation rule as the standard corpus, and the associated corpus belongs to a non-target field. For example, if the target field is the search field and the standard corpus in the search field is insufficient, the corpus of the machine question-answering field can be selected as the associated corpus of the search field to expand the preset corpus. By selecting the corpus labeled based on the same character label annotation rule as the standard corpus in other fields (non-target fields) as the associated corpus of the standard corpus, and generating the preset corpus based on the standard corpus and its associated corpus, that is, using the associated corpus from the non-target field to expand the preset corpus of the low-resource target field, so as to increase the data volume of the words in the preset corpus.

[0092] Exemplarily, if "Three Good Students" is labeled as "Three|S, Good|S, Student|B, Student|E" in a certain corpus and "Three|B, Good|E, Student|B, Student|E" in another corpus, it is considered that the label annotation rules based on these two corpora are different.

[0093] It is easy to understand that both the standard corpus in the target field and the labeled corpus in the non-target field include several words labeled with character tags. Based on this, it can be determined whether the standard corpus and the corpus in other fields are similar according to whether the character tags labeled for the same word in the standard corpus and the corpus in other fields are the same, and whether the number of such words is greater than the preset number. Exemplarily, if the number of the same words in the standard corpus and a certain corpus in the non-target field is greater than the preset number, and the character tags labeled for the same word in these two corpora are the same, it is determined that this corpus is the associated corpus of the standard corpus. Another exemplarily, traverse the words in the standard corpus to determine whether they are included in a certain corpus in the non-target field. If so, further determine whether the character tags labeled for the same word in these two corpora are the same. If so, determine this word as the target word. If the number of target words is greater than the preset number, determine the corpus in the non-target field as the associated corpus of the standard corpus.

[0094] It can be seen that considering the problem of insufficient scale of the standard corpus generated based on the corpus in the target domain in the low-resource scenario, a corpus annotated according to the same character label annotation rules as the standard corpus is selected in other domains (non-target domains) as the associated corpus of the standard corpus. Among them, when the number of the same words in a corpus in the non-target domain is greater than a preset number, and the labels annotated for these same words in the two corpora are the same, it is considered that the two corpora are annotated according to the same character label annotation rules. Furthermore, it is determined that this corpus in the non-target domain is the associated corpus of the standard corpus. Finally, a preset corpus is generated based on the standard corpus and its associated corpus, thereby increasing the data volume of the words in the preset corpus.

[0095] In some embodiments, the preset encoding model may be an encoding model trained based on the above-mentioned preset corpus. Furthermore, for the low-resource scenario, the feature extraction ability of the encoding model can be improved without changing the structure of the encoding model.

[0096] In some possible implementation manners of S203, first, the first weight sum and the second weight sum corresponding to each label sequence are determined. The first weight sum is the sum of the emission weights corresponding to each character label in the label sequence, and the second weight sum is the sum of the state transition weights corresponding to each character label in the label sequence; then, according to the first weight sum, the second weight sum, and the perplexity corresponding to the label sequence, the score of the label sequence is determined; finally, the label sequence corresponding to the maximum score is determined as the target label sequence.

[0097] Exemplarily, the score of a certain label sequence can be determined according to the following formula 3:

[0098]

[0099] where S i represents the score of the i-th label sequence;

[0100] e j represents the emission weight corresponding to the j-th character label in the i-th label sequence;

[0101] t j represents the state transition weight corresponding to the j-th character label in the i-th label sequence;

[0102] PP(S i ) represents the perplexity of the i-th label sequence.

[0103] In some possible implementation manners, the score of each label sequence can be determined according to the following formula 4:

[0104]

[0105] Among them, W e 、W t and W p are respectively preset weighting coefficients.

[0106] It can be intuitively seen from the above formulas 3 and 4 that the text word segmentation method provided by the embodiments of the present disclosure combines the emission weights, state transition weights corresponding to each label in the label sequence learned by the encoding model in the word segmentation encoding and decoding stage, and the perplexity corresponding to each label sequence, and selects the optimal label sequence, thereby ensuring the robustness and accuracy of the word segmentation effect in low-resource scenarios. Compared with the related art, the technical solution of the present application can complete word segmentation without a large amount of computing resources, is applicable to industrial scenarios, and does not need to stack models of multiple tasks to promote the word segmentation performance of the word segmentation model, but achieves the above technical effects by improving the encoding algorithm.

[0107] When specifically implemented, the Viterbi algorithm can be used to combine the emission weights, state transition weights corresponding to each character label in each label sequence, and P(w 1 ) and P(w k |w k-1 ) corresponding to the word segmentation sequence in each label sequence, and determine the target label sequence from all the label sequences. Based on the embodiments of the present disclosure, those skilled in the art clearly know how to apply the Viterbi algorithm to the embodiments of the present disclosure, so details are not described herein.

[0108] Simulate a low-resource scenario, and use multiple data sets to verify the text word segmentation method provided by the embodiments of the present disclosure. The results are shown in the following table. The models involved in the following table include pre-trained models trained by knowledge distillation (KD) combined with the Softmax activation function or conditional random fields (CRF). The data sets involved include the AS, PKU traditional Chinese data set, MSR, CITYU simplified Chinese data set, CTB data set, SXU data set, Weibo data set, and ZX data set.

[0109]

[0110] As can be seen from the above table, the text word segmentation method provided by the embodiments of the present disclosure, especially for low-resource scenarios, such as simulated low-resource scenarios with a data set utilization rate of 10% and 80%, shows good robustness and accuracy.

[0111] As can be seen from the above embodiments, the text tokenization method provided by the embodiments of the present disclosure calculates the perplexity of a certain tag sequence based on the probability of each token in the token sequence corresponding to the tag sequence appearing under the condition that its previous token appears during the token decoding stage, so as to evaluate the rationality of the tag sequence. By combining the weights of each tag sequence learned during the token encoding stage, the optimal tag sequence is selected, thereby ensuring the robustness and accuracy of the tokenization effect in low-resource scenarios. Compared with the related art, the technical solution of the present application can complete tokenization without a large amount of computing resources, is applicable to industrial scenarios, and does not need to stack models of multiple tasks to improve the tokenization performance of the tokenization model. Instead, the above technical effects are achieved by improving the encoding algorithm.

[0112] Figure 4 is a block diagram of a text tokenization device shown according to an exemplary embodiment, as Figure 4 shown, the text tokenization device provided by the embodiments of the present disclosure includes a tag sequence acquisition unit 401, an encoding unit 402, and a decoding unit 403.

[0113] The tag sequence acquisition unit 401 is configured to acquire multiple tag sequences corresponding to the text to be tokenized, and the tag sequences are used to split the text to be tokenized into corresponding token sequences. For example, as Figure 2 shown, the tag sequence acquisition unit 401 can be used to execute S201.

[0114] The encoding unit 402 is configured to perform encoding processing on the text to be tokenized to obtain weights corresponding to each tag sequence; for example, as Figure 2 shown, the encoding unit 402 can be used to execute S202.

[0115] The decoding unit 403 is configured to decode to obtain a target tag sequence according to the weights and perplexity corresponding to each tag sequence, so as to determine the tokenization result by using the target tag sequence; wherein, the perplexity corresponding to each tag sequence is determined according to the statistical probability of each token in the token sequence corresponding to the tag sequence appearing under the condition that its previous token appears. For example, as Figure 2 shown, the decoding unit 403 can be used to execute S203.

[0116] In some embodiments, the decoding unit 403 is specifically configured to: for each word segmentation sequence corresponding to the tag sequence, determine the statistical probability of the occurrence of the first word in the word segmentation sequence and the conditional statistical probability of each of the remaining words, where the conditional statistical probability of the k-th word refers to the statistical probability of the occurrence of the k-th word under the condition that the (k - 1)-th word appears, and k is a positive integer greater than 1; determine the perplexity of the tag sequence according to the statistical probability of the occurrence of the first word and the conditional statistical probabilities of each of the remaining words; and determine the target tag sequence according to the weight and perplexity corresponding to each tag sequence.

[0117] In some embodiments, the decoding unit 403 is specifically configured to determine the probability of the occurrence of the first word in the preset corpus as the statistical probability of the occurrence of the first word; determine the probability of the consecutive occurrence of two adjacent words among all the remaining words in the preset corpus and the probability of the occurrence of each word in the preset corpus; and determine the conditional statistical probability of the k-th word according to the probability of the consecutive occurrence of the k-th word and the (k - 1)-th word and the probability of the occurrence of the (k - 1)-th word.

[0118] In some embodiments, the preset corpus is constructed according to a standard corpus in the target domain and a related corpus of the standard corpus. The related corpus of the standard corpus refers to a corpus annotated according to the same character tag annotation rule as the standard corpus, and the related corpus belongs to a non-target domain. The character tag annotation rule refers to a rule for annotating character tags for the text in the corpus, and the character tags are used to annotate the position of each character in the text segmentation of the text.

[0119] In some embodiments, both the standard corpus and the related corpus include a plurality of texts, each text includes a plurality of word segments, and the characters in the word segments are annotated with character tags; the number of identical word segments in the standard corpus and the related corpus is greater than a preset number, and the character tags annotated for the characters in the same word segment in the standard corpus and in the related corpus are the same.

[0120] In some embodiments, the tag sequence includes character tags corresponding to each character in the text to be segmented, and the character tags are used to annotate the position of the character in the word segmentation; the weight corresponding to the tag sequence includes the emission weight and the state transition weight corresponding to each character tag in the tag sequence.

[0121] In some embodiments, the decoding unit 403 is specifically configured to: determine a first weight sum and a second weight sum corresponding to each tag sequence, where the first weight sum is the sum of the emission weights corresponding to each character tag in the tag sequence, and the second weight sum is the sum of the state transition weights corresponding to each character tag in the tag sequence; determine the score of the tag sequence according to the first weight sum, the second weight sum, and the perplexity corresponding to the tag sequence; and determine the tag sequence corresponding to the maximum score as the target tag sequence.

[0122] In some embodiments, the encoding unit 402 is specifically configured to: input the text to be segmented into an encoding model trained using the preset corpus, and output the emission weight and the state transition weight corresponding to each character tag in each tag sequence.

[0123] Based on the text segmentation device provided in the embodiments of the present disclosure, in the segmentation decoding stage, the perplexity of a tag sequence is obtained according to the probability that each segmented word in the segmented word sequence corresponding to the tag sequence appears under the condition that its previous segmented word appears, so as to evaluate the rationality of the tag sequence, and in combination with the weights of each tag sequence learned in the segmentation encoding stage, the optimal tag sequence is selected, thereby ensuring the robustness and accuracy of the segmentation effect in a low-resource scenario. Compared with the related art, the technical solution of the present application can complete segmentation without a large amount of computing resources, is applicable to industrial scenarios, and does not need to promote the segmentation performance of the segmentation model by stacking models of multiple tasks, but realizes the above technical effects by improving the encoding algorithm.

[0124] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0125] Figure 5 is a schematic structural diagram of a server provided by the present disclosure. As Figure 5 , the server 50 may include at least one processor 501 and a memory 503 for storing instructions executable by the processor. Wherein, the processor 501 is configured to execute the instructions in the memory 503 to implement the entity recognition method in the above embodiments.

[0126] In addition, the server 50 may further include a communication bus 502 and at least one communication interface 504.

[0127] The processor 501 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the solution of the present disclosure.

[0128] The communication bus 502 may include a path for transmitting information between the above components.

[0129] The communication interface 504 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0130] The memory 503 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processing unit through a bus. The memory can also be integrated with the processing unit.

[0131] Among them, the memory 503 is used to store instructions for executing the solution of the present disclosure and is controlled by the processor 501 to execute. The processor 501 is used to execute the instructions stored in the memory 503, thereby implementing the functions in the method of the present disclosure.

[0132] As an example, in combination with Figure 4 , the functions implemented by the tag sequence acquisition unit 401, the encoding unit 402, and the decoding unit 403 in the text tokenization device are the same as the functions of the processor 501 in Figure 5 .

[0133] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 5 the CPU0 and CPU1 in

[0134] In a specific implementation, as an embodiment, the server 50 may include multiple processors, such as Figure 5The processors 501 and 507 therein. Each of these processors can be a single-CPU processor or a multi-CPU processor. The processors herein can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0135] In a specific implementation, as an example, the server 50 may further include an output device 505 and an input device 506. The output device 505 communicates with the processor 501 and can display information in various ways. For example, the output device 505 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 506 communicates with the processor 501 and can accept user input in various ways. For example, the input device 506 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.

[0136] Those skilled in the art can understand that Figure 5 the structure shown in does not constitute a limitation on the server 50, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0137] In addition, the present disclosure also provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the server, the server can execute the text tokenization method as provided in the above embodiments.

[0138] In addition, the present disclosure also provides a computer program product, including computer instructions. When the computer instructions run on the server, the server executes the text tokenization method as provided in the above embodiments.

[0139] Those skilled in the art will readily think of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A text word segmentation method, characterized in that, it includes: Obtaining a plurality of tag sequences corresponding to the text to be segmented, where the tag sequences are used to segment the text to be segmented into corresponding word segmentation sequences; Performing encoding processing on the text to be segmented to obtain weights corresponding to each of the tag sequences; the tag sequence includes character tags corresponding to each character in the text to be segmented, and the character tags are used to mark the position of the character in the word segmentation; the weight corresponding to the tag sequence includes the emission weight and the state transition weight corresponding to each character tag in the tag sequence; Decoding according to the weights and perplexities corresponding to each tag sequence to obtain a target tag sequence, so as to determine the word segmentation result by using the target tag sequence; wherein, the perplexity corresponding to each tag sequence is determined according to the statistical probability of each word segmentation appearing under the condition that the previous word segmentation appears in the word segmentation sequence corresponding to the tag sequence; The decoding according to the weights and perplexities corresponding to each tag sequence to obtain a target tag sequence includes: determining the first weight sum and the second weight sum corresponding to each tag sequence, the first weight sum being the sum of the emission weights corresponding to each character tag in the tag sequence, and the second weight sum being the sum of the state transition weights corresponding to each character tag in the tag sequence; determining the score of the tag sequence according to the first weight sum, the second weight sum and the perplexity corresponding to the tag sequence; and determining the tag sequence corresponding to the maximum score as the target tag sequence.

2. The text word segmentation method according to claim 1, characterized in that, the decoding according to the weights and perplexities corresponding to each tag sequence to obtain a target tag sequence includes: For the word segmentation sequence corresponding to each tag sequence, determining the statistical probability of the first word segmentation appearing and the conditional statistical probability of each of the remaining word segmentations, where the conditional statistical probability of the kth word segmentation refers to the statistical probability of the kth word segmentation appearing under the condition that the (k - 1)th word segmentation appears, and k is a positive integer greater than 1; Determining the perplexity of the tag sequence according to the statistical probability of the first word segmentation appearing and the conditional statistical probabilities of each of the remaining word segmentations; Determining the target tag sequence according to the weights and perplexities corresponding to each tag sequence.

3. The text word segmentation method according to claim 2, characterized in that, the determining the statistical probability of the first word segmentation appearing and the conditional statistical probability of each of the remaining word segmentations includes: Determining the probability of the first word segmentation appearing in the preset corpus as the statistical probability of the first word segmentation appearing; Determining the probability of two adjacent word segmentations continuously appearing in the preset corpus among all the remaining word segmentations, and the probability of each word segmentation appearing in the preset corpus; Determining the conditional statistical probability of the kth word segmentation according to the probability of the kth word segmentation and the (k - 1)th word segmentation continuously appearing and the probability of the (k - 1)th word segmentation appearing.

4. The text word segmentation method according to claim 3, characterized in that, The preset corpus is constructed based on a standard corpus in the target domain and a related corpus of the standard corpus. The related corpus of the standard corpus refers to a corpus annotated based on the same character tag annotation rule as the standard corpus. The related corpus belongs to a non-target domain. The character tag annotation rule refers to a rule for annotating character tags to the text in the corpus. The character tag is used to annotate the position of each character in the text segmentation of the text.

5. The text segmentation method according to claim 4, wherein, both the standard corpus and the related corpus include a plurality of texts, each text includes a plurality of word segments, and the characters in the word segments are annotated with character tags; the number of the same word segments in the standard corpus and the related corpus is greater than a preset number, and the character tags annotated to the characters in the same word segment in the standard corpus and in the related corpus are the same.

6. The text segmentation method according to claim 5, wherein, the encoding process of the text to be segmented to obtain the weight corresponding to each tag sequence includes: inputting the text to be segmented into an encoding model trained by using the preset corpus, and outputting the emission weight and the state transition weight corresponding to each character tag in each tag sequence.

7. A text segmentation device, wherein, comprising: a tag sequence acquisition unit, configured to acquire a plurality of tag sequences corresponding to the text to be segmented, where the tag sequences are used to segment the text to be segmented into corresponding word segment sequences; an encoding unit, configured to perform an encoding process on the text to be segmented to obtain the weight corresponding to each tag sequence; the tag sequence includes character tags corresponding to each character in the text to be segmented, and the character tags are used to annotate the position of the character in the word segment; the weight corresponding to the tag sequence includes the emission weight and the state transition weight corresponding to each character tag in the tag sequence; a decoding unit, configured to decode to obtain a target tag sequence according to the weight and the perplexity corresponding to each tag sequence, so as to determine a word segmentation result by using the target tag sequence; wherein, the perplexity corresponding to each tag sequence is determined according to the statistical probability of each word segment appearing under the condition that the previous word segment appears in the word segment sequence corresponding to the tag sequence; The decoding unit is specifically configured to: determine a first weight sum and a second weight sum corresponding to each tag sequence, where the first weight sum is the sum of the emission weights corresponding to each character tag in the tag sequence, and the second weight sum is the sum of the state transition weights corresponding to each character tag in the tag sequence; determine the score of the tag sequence according to the first weight sum, the second weight sum and the perplexity corresponding to the tag sequence; and determine the tag sequence corresponding to the maximum score as the target tag sequence.

8. The text segmentation device according to claim 7, wherein, the decoding unit is specifically configured to: For each word segmentation sequence corresponding to the label sequence, determine the statistical probability of the first word in the word segmentation sequence and the conditional statistical probability of each of the remaining words, where the conditional statistical probability of the k-th word is the statistical probability of the k-th word occurring under the condition that the (k - 1)-th word has occurred, and k is a positive integer greater than 1; Determine the perplexity of the label sequence according to the statistical probability of the first word occurring and the conditional statistical probabilities of each of the remaining words; Determine the target label sequence according to the weight and perplexity corresponding to each label sequence; 9. The text word segmentation device according to claim 8, wherein, the decoding unit is specifically configured to: Determine the probability of the first word occurring in the preset corpus as the statistical probability of the first word occurring; Determine the probability of two adjacent words occurring continuously in the preset corpus among all the remaining words, and the probability of each word occurring in the preset corpus; Determine the conditional statistical probability of the k-th word according to the probability of the k-th word and the (k - 1)-th word occurring continuously and the probability of the (k - 1)-th word occurring; 10. The text word segmentation device according to claim 9, wherein, the preset corpus is constructed according to the standard corpus in the target field and the associated corpus of the standard corpus. The associated corpus of the standard corpus refers to a corpus labeled based on the same character label annotation rule as the standard corpus. The associated corpus belongs to a non-target field. The character label annotation rule refers to the rule for annotating character labels for the text in the corpus, and the character labels are used to annotate the position of each character in the text word segmentation of the text; 11. The text word segmentation device according to claim 10, wherein, both the standard corpus and the associated corpus include a number of texts, the texts include a number of word segmentations, and the characters in the word segmentations are annotated with character labels; the number of the same word segmentations in the standard corpus and the associated corpus is greater than a preset number, and the character labels annotated for the characters in the same word segmentation in the standard corpus and in the associated corpus are the same; 12. The text word segmentation device according to claim 11, wherein, the encoding unit is specifically configured to: Input the text to be word segmented into an encoding model trained using the preset corpus, and output the emission weight and the state transition weight corresponding to each character label in each label sequence; 13. An electronic device, wherein, it includes: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the text word segmentation method according to any one of claims 1 - 6; 14. A computer-readable storage medium, wherein, when the instructions in the computer-readable storage medium are executed by the processor of the server, the server is enabled to execute the text word segmentation method according to any one of claims 1 - 6; 15. A computer program product, including instructions, wherein, The computer program product includes computer instructions that, when run on a server, cause the server to execute the text word segmentation method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Label classification method and device of corpora, computer equipment and storage medium

    CN112084334A

  • Entity word extraction method and device and electronic equipment

    CN113743107A