Text tokenization method, device, computer equipment and computer-readable storage medium
By obtaining and determining the annotation information of the text to be divided, selecting appropriate word segmentation criteria, and performing word segmentation processing on the Chinese word segmentation model, solving the problems of single data fields, insufficient quantity and inconsistent standards in the existing model, and improving the accuracy of word segmentation.
Patent Information
- Application Number
- CN202210157392.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-02-21
AI Technical Summary
During the training process, the existing Chinese word segmentation model has problems such as single data field, insufficient data quantity and inconsistent word segmentation standards, resulting in inaccurate word segmentation.
By obtaining the annotation information of the text to be segmented, the target word segmentation standard type is determined, and the text is segmented according to the criteria. The method includes obtaining the text to be participle, determining the labeling information, extracting the target participle standard type, and performing word participle processing.
It improves the accuracy of word segmentation, and can select appropriate word segmentation standards based on the annotation information of different texts to ensure the accuracy of word segmentation results.
Smart Images

Figure CN114580395B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a text word segmentation method, device, computer device and computer-readable storage medium. Background Art
[0002] Chinese word segmentation is a basic technology in natural language processing and plays an important role in understanding sentences. Unlike English sentences with spaces in the middle, Chinese sentences have rich language expressions. For example, there are a large number of Chinese words with multiple meanings, and the meaning of a word needs to be understood in combination with the context, which poses challenges to Chinese word segmentation.
[0003] Currently, in the training process of Chinese word segmentation models, there are some problems with the training data sets used by Chinese word segmentation models. For example, the sources of the training data sets are relatively single. For example, all the training data sets come from the news field, which may lead to the problem that the Chinese word segmentation model cannot accurately segment the training data in non-news fields. For example, the number of training data in the training data set is small, which may lead to the problem of low accuracy of the Chinese word segmentation model. Another example is that the word segmentation standards among different training data sets are not the same. For example, for the word segmentation of Chinese names, some training data sets use the surname and the given name as separate word segmentation units for word segmentation, while some training data sets use the surname and the given name as a whole for word segmentation. That is, due to the lack of unified word segmentation standards, there are also problems with inaccurate word segmentation by the Chinese word segmentation model.
[0004] In summary, the existing Chinese word segmentation models have the problem of inaccurate word segmentation. Summary of the Invention
[0005] Embodiments of this application provide a text word segmentation method, device, computer device and computer-readable storage medium, which can improve the accuracy of word segmentation.
[0006] A text word segmentation method includes:
[0007] Obtain the text to be word segmented;
[0008] Determine the annotation information of the text to be word segmented according to the text to be word segmented;
[0009] Extract the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be word segmented according to the annotation information;
[0010] Segment the text to be word segmented according to the target word segmentation standard type.
[0011] Correspondingly, embodiments of this application provide a text word segmentation device, including:
[0012] An acquisition unit, which can be used to acquire the text to be segmented;
[0013] A determination unit, which can be used to determine the annotation information of the text to be segmented according to the text to be segmented;
[0014] An extraction unit, which can be used to extract the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information;
[0015] A word segmentation unit, which can be used to segment the text to be segmented according to the target word segmentation standard type.
[0016] In some embodiments, the word segmentation unit can specifically be used to perform feature extraction and word segmentation processing on the text to be segmented according to the target word segmentation standard type by using a text word segmentation model, wherein the word segmentation processing is performed according to the feature information obtained by the text word segmentation model for feature extraction of the text to be segmented.
[0017] In some embodiments, the text word segmentation device further includes a training unit. The training unit can specifically be used to acquire a candidate text data sample set, and the candidate text data sample set includes at least one candidate text data sample; determine a reference text data sample according to the candidate text data sample; mark the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type; acquire the label word segmentation information corresponding to the text data sample of each word segmentation standard type; generate a text data sample set according to the text data sample corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data sample.
[0018] In some embodiments, the training unit can specifically be used to acquire a candidate text data sample set, and the candidate text data sample set includes at least one candidate text data sample; determine a reference text data sample according to the candidate text data sample; mark the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type; acquire the label word segmentation information corresponding to the text data sample of each word segmentation standard type; generate a text data sample set according to the text data sample corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data sample.
[0019] In some embodiments, the training unit can specifically be used to perform word segmentation processing on the candidate text data sample to obtain the word segmentation result corresponding to the candidate text data sample; screen out the reference text data sample from at least one candidate text data sample according to the word segmentation result.
[0020] In some embodiments, the training unit can specifically be used to acquire at least two preset word segmentation strategies; perform word segmentation processing on the candidate text data sample according to the preset word segmentation strategies to obtain the word segmentation result corresponding to each preset word segmentation strategy.
[0021] In some embodiments, the training unit may be specifically configured to, for each candidate text data sample, obtain the number of different word segmentation results in the word segmentation result of the candidate text data sample; if the number of different word segmentation results is greater than or equal to a preset number threshold, then use the candidate text data sample as a reference text data sample.
[0022] In some embodiments, the training unit may be specifically configured to obtain reference word segmentation information corresponding to the word segmentation standard type; calculate word segmentation feature information according to the reference word segmentation information and the text data sample; and predict the predicted word segmentation information of the text data sample by using the text word segmentation model to be trained according to the word segmentation feature information.
[0023] In some embodiments, the training unit may be specifically configured to obtain a set of texts to be screened corresponding to the word segmentation standard type, where the set of texts to be screened includes at least one text to be screened; screen at least one text to be screened by using the trained text classification model to obtain a reference text; and perform feature extraction on the reference text to obtain reference word segmentation information corresponding to the word segmentation standard type.
[0024] In some embodiments, the training unit may be specifically configured to obtain a reference text data sample and a set of text data samples to be processed, where the set of text data samples to be processed includes at least one text data sample to be processed; determine positive text data samples and negative text data samples in at least one text data sample to be processed according to the reference text data sample; and converge the text classification model to be trained according to the positive text data samples and the negative text data samples to obtain a trained text classification model.
[0025] In some embodiments, the training unit may be specifically configured to perform feature extraction on the reference word segmentation information to obtain reference text feature information corresponding to the reference word segmentation information; perform feature extraction on the text data sample to obtain text feature information corresponding to the text data sample; and fuse the reference text feature information and the text feature information to obtain word segmentation feature information.
[0026] In some embodiments, the determination unit may be specifically configured to obtain at least two current word segmentation strategies; perform word segmentation processing on the text to be segmented according to the current word segmentation strategies to obtain current word segmentation results corresponding to each current word segmentation strategy; extract the number of identical current word segmentation results from the current word segmentation results, and determine the annotation information of the text to be segmented according to the number of identical current word segmentation results.
[0027] In some embodiments, a determination unit may be specifically configured to obtain a set of mapping relationships, where the set of mapping relationships includes the mapping relationships between preset annotation information and tokenization standard types for a text to be tokenized; and determine a target tokenization standard type corresponding to the annotation information from at least one tokenization standard type in the set of mapping relationships according to the set of mapping relationships and the annotation information.
[0028] In addition, an embodiment of the present application further provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute any text tokenization method provided by the embodiment of the present application.
[0029] In addition, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute any text tokenization method provided by the embodiment of the present application.
[0030] An embodiment of the present application can obtain a text to be tokenized; determine annotation information of the text to be tokenized according to the text to be tokenized; extract a target tokenization standard type corresponding to the annotation information from at least one tokenization standard type for the text to be tokenized according to the annotation information; tokenize the text to be tokenized according to the target tokenization standard type; since the embodiment of the present application can screen out the target tokenization standard type from the tokenization standard types of the text to be tokenized according to the annotation information of the text to be tokenized, the text to be tokenized can be accurately tokenized according to the target tokenization standard type, thereby improving the accuracy of tokenization. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0032] Figure 1 is a schematic diagram of the scenario of the text tokenization method provided by the embodiment of the present application;
[0033] Figure 2 is a schematic flowchart I of the text tokenization method provided by the embodiment of the present application;
[0034] Figure 3 is a schematic flowchart of training a text tokenization model to be trained provided by the embodiment of the present application;
[0035] Figure 4 is a schematic flowchart of determining a reference text data sample according to a candidate text data sample provided by the embodiment of the present application;
[0036] Figure 5 It is a schematic flowchart of the process of predicting the predicted word segmentation information of the text data sample by using the text word segmentation model to be trained for each word segmentation standard type provided by the embodiment of the present application;
[0037] Figure 6 It is the second schematic flowchart of the text word segmentation method provided by the embodiment of the present application;
[0038] Figure 7 It is the third schematic flowchart of the text word segmentation method provided by the embodiment of the present application;
[0039] Figure 8 It is the second schematic flowchart of the process of training the text word segmentation model to be trained provided by the embodiment of the present application;
[0040] Figure 9 It is the loss change diagram provided by the embodiment of the present application;
[0041] Figure 10 It is the schematic structural diagram of the word segmentation device provided by the embodiment of the present application;
[0042] Figure 11 It is the schematic structural diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0044] The embodiment of the present application provides a text word segmentation method, device, computer device and computer-readable storage medium. Among them, the word segmentation device can be integrated in the computer device, and the computer device can be a server or a terminal device, etc.
[0045] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions in this regard. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0046] Among them, the embodiments of this application may involve Artificial Intelligence (AI). Artificial intelligence is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0047] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0048] For example, referring to Figure 1 , taking the example that the text segmentation device is integrated in a computer device, the computer device can obtain the text to be segmented; determine the annotation information of the text to be segmented according to the text to be segmented; extract the target segmentation standard type corresponding to the annotation information from at least one segmentation standard type for the text to be segmented according to the annotation information; and segment the text to be segmented according to the target segmentation standard type.
[0049] Among them, the text to be segmented can be any text in any field, which can be text in a specific field. For example, the text to be segmented can be the text to be segmented in the news field, the text to be segmented in the film field, the text to be segmented in the medical field, etc.
[0050] Among them, the annotation information can be understood as the marking information of the text to be segmented for the word segmentation standard type.
[0051] Among them, the word segmentation standard type can refer to the type of word segmentation standard for text segmentation. The word segmentation standard type can include the national word segmentation standard type, the Peking University word segmentation standard type, and the user-defined word segmentation standard type. The word segmentation standard can include the national word segmentation standard, the Peking University word segmentation standard, and the user-defined word segmentation standard. The national word segmentation standard is the word segmentation standard corresponding to the national word segmentation standard type, and the national word segmentation standard can refer to the standard of Chinese information processing vocabulary. The Peking University word segmentation standard is the word segmentation standard corresponding to the Peking University word segmentation standard type, and the Peking University word segmentation standard can refer to the standard of the basic processing specification of the Modern Chinese Corpus of Peking University. The user-defined word segmentation standard is the word segmentation standard corresponding to the user-defined word segmentation standard type.
[0052] Of course, the word segmentation standard type in the embodiments of the present application is not limited to the national word segmentation standard type, the Peking University word segmentation standard type, and the user-defined word segmentation standard type.
[0053] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0054] This embodiment will be described from the perspective of the word segmentation device. The word segmentation device can be specifically integrated in a computer device, which can be a server or a terminal device, etc.; among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC, Personal Computer), a wearable device, a virtual reality device, or other intelligent devices that can obtain data, etc.
[0055] As Figure 2 shown, the specific process of the text word segmentation method is as follows:
[0056] S101. Obtain the text to be segmented.
[0057] Among them, the text to be segmented can be any text in any field, which can be text in a specific field. For example, the text to be segmented can be the text to be segmented in the news field, the text to be segmented in the film field, the text to be segmented in the medical field, etc.
[0058] Among them, the text to be segmented in the embodiments of the present application can be extracted from the database of the computer device or obtained online in real time.
[0059] S102: Determine the annotation information of the text to be segmented according to the text to be segmented.
[0060] The annotation information can be understood as the mark information of the text to be segmented for the segmentation standard type. The annotation information of the embodiment of the present application can be represented by an identifier, for example, the annotation information can be 1, 0, a, etc.
[0061] In the embodiment of the present application, there are multiple ways to determine the annotation information of the text to be segmented according to the text to be segmented, as detailed below:
[0062] For example, the computer device displays a tag information selection page, which includes at least one candidate tag information; the computer device selects tag information from the at least one candidate tag information in response to a selection operation on the candidate tag information; and uses the tag information as the tag information of the text to be segmented.
[0063] That is, the annotation information may be the annotation information selected by the user. Thus, the embodiment of the present application may select the corresponding annotation information according to the needs of the user, thereby accurately performing segmentation processing on the segmented text based on the target segmentation standard type corresponding to the annotation information.
[0064] In addition to the above, in order to improve the accuracy of word segmentation of the text to be segmented, the embodiment of the present application determines the way of marking information of the text to be segmented according to the text to be segmented, which can be specifically as follows:
[0065] For example, a computer device can obtain at least two current word segmentation strategies; perform word segmentation processing on the text to be segmented according to the current word segmentation strategies to obtain the current word segmentation results corresponding to each current word segmentation strategy; extract the number of identical current word segmentation results from the current word segmentation results, and determine the annotation information of the text to be segmented according to the number of identical current word segmentation results.
[0066] The current word segmentation strategy may refer to a strategy for word segmentation of a text to be segmented, and the strategy may be an algorithm or a word segmentation tool, etc. The current word segmentation strategy may at least include jieba (i.e., stammering) word segmentation strategy, StanfordCore NLP (i.e., Stanford Core NLP) word segmentation strategy, THULAC (i.e., THU Lexical Analyzer for Chinese) word segmentation strategy, PKUSEG word segmentation strategy, LTP-4.0 (i.e., Language Technology Platform-4.0) word segmentation strategy, TexSmart word segmentation strategy, and HanLP word segmentation strategy.
[0067] Among them, the computer device performs word segmentation processing on the text to be segmented according to each current word segmentation strategy, and obtains the current word segmentation result corresponding to each current word segmentation strategy.
[0068] Among them, the manner in which the embodiment of the present application determines the annotation information of the text to be segmented according to the number of the same current word segmentation results may be: the computer device determines the difficulty level of the text to be segmented according to the number of the same current word segmentation results; and extracts the annotation information corresponding to the difficulty level from several candidate annotation information according to the difficulty level.
[0069] Among them, the difficulty levels may correspond to the annotation information one by one. For example, the difficulty levels include a first difficulty level and a second difficulty level, the annotation information includes a and b, the first difficulty level corresponds to the annotation information a, and the second difficulty level corresponds to the annotation information b.
[0070] S103. According to the annotation information, extract the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented.
[0071] Among them, the word segmentation standard type may refer to the type of the word segmentation standard for performing word segmentation processing on the text, and the word segmentation standard type may include a national word segmentation standard type, a Peking University word segmentation standard type, and a user-defined word segmentation standard type.
[0072] The manner in which the embodiment of the present application extracts the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information may be as follows:
[0073] For example, the computer device obtains a set of mapping relationships, and the set of mapping relationships includes the mapping relationships between the preset annotation information and the word segmentation standard types for the text to be segmented; and determines the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type in the set of mapping relationships according to the set of mapping relationships and the annotation information.
[0074] For example, the word segmentation standard types include a first word segmentation standard type and a second word segmentation standard type, the annotation information includes "4" and "5", the first word segmentation standard type corresponds to the annotation information "4", and the second word segmentation standard information corresponds to the annotation information "5". Among them, the first word segmentation standard type may be a national word segmentation standard type, and the second word segmentation standard type may be a Peking University word segmentation standard type.
[0075] In this way, the embodiment of the present application can accurately perform word segmentation processing on the text to be segmented according to the target word segmentation standard type of the text to be segmented.
[0076] S104. Perform word segmentation on the text to be segmented according to the target word segmentation standard type.
[0077] In the embodiments of the present application, the text to be segmented can be segmented according to the target segmentation standard type by using the segmentation standard or segmentation tool corresponding to the target segmentation standard type.
[0078] In addition to the above, the embodiments of the present application can also segment the text to be segmented by using a text segmentation model, as follows:
[0079] For example, a computer device can extract features and segment the text to be segmented according to the target segmentation standard type by using a text segmentation model, where the segmentation process is based on the feature information obtained by extracting the features of the text to be segmented by the text segmentation model.
[0080] Among them, the text segmentation module can be a bert model, a lebert model, etc.
[0081] Among them, the text segmentation model of the embodiments of the present application can extract features and segment the text to be segmented based on the segmentation standard corresponding to the target segmentation standard type.
[0082] Among them, the text segmentation model of the embodiments of the present application can extract features from the text to be segmented according to the target segmentation standard type, so as to obtain the feature information corresponding to the target segmentation standard type; according to the feature information, use the text segmentation model to segment the text to be segmented to obtain the segmentation result corresponding to the target segmentation standard type.
[0083] For example, if the target segmentation standard type is the national segmentation standard type, the text segmentation model extracts features from the text to be segmented according to the national segmentation standard corresponding to the national segmentation standard type to obtain the feature information corresponding to the national segmentation standard; according to the feature information, use the text segmentation model to segment the text to be segmented to obtain the segmentation result corresponding to the national segmentation standard.
[0084] Before the embodiments of the present application extract features from the text to be segmented by using a text segmentation model according to the target segmentation standard type, the text segmentation model to be trained can also be trained. As Figure 3 shown, the process of training the text segmentation model to be trained by the embodiments of the present application can be as follows:
[0085] A1. Obtain a set of text data samples.
[0086] Among them, the set of text data samples includes text data samples corresponding to each segmentation standard type, and the labeled segmentation information corresponding to each text data sample.
[0087] Among them, the segmentation standard type can include multiple segmentation standard types, and the segmentation standard type can include the national segmentation standard type, the Peking University segmentation standard type, and the user-defined segmentation standard type.
[0088] Among them, the label word segmentation information can be understood as the label corresponding to the text data sample.
[0089] Among them, the text data samples can include text data samples in various fields. For example, the text data samples include the text to be segmented in the news field, the text to be segmented in the movie field, and the text to be segmented in the medical field.
[0090] The method for the embodiment of the present application to obtain the text data sample set can be as follows:
[0091] For example, a computer device can obtain a candidate text data sample set, and the candidate text data sample set includes at least one candidate text data sample; according to the candidate text data sample, determine a reference text data sample; mark the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type; obtain the label word segmentation information corresponding to the text data sample of each word segmentation standard type; generate a text data sample set according to the text data sample corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data sample.
[0092] Among them, since the candidate text data samples in the candidate text data sample set of the embodiment of the present application include candidate text data samples in various fields. For example, the candidate text data samples include candidate text data samples in the news field, candidate text data samples in the movie field, and candidate text data samples in the medical field, and the candidate text data samples in the candidate text data sample set include candidate text data samples that are difficult to segment and candidate text data samples that are easy to segment. Based on this, the embodiment of the present application needs to determine a reference text data sample from the candidate text data samples. Among them, the reference text data sample can be a candidate text data sample that is difficult to segment.
[0093] The candidate text data sample that is difficult to segment can refer to a candidate text data sample that is prone to word segmentation errors. The candidate text data sample that is easy to segment can refer to a candidate text data sample that is not prone to word segmentation errors.
[0094] Based on the above, as Figure 4 shown, the method for the embodiment of the present application to determine a reference text data sample according to the candidate text data sample can be as follows:
[0095] B1. Perform word segmentation processing on the candidate text data sample to obtain the word segmentation result corresponding to the candidate text data sample.
[0096] Among them, the embodiment of the present application can perform word segmentation processing on the candidate text data sample, so as to judge whether the candidate text data sample is a candidate text data sample that is difficult to segment or a candidate text data sample that is easy to segment according to the obtained word segmentation result.
[0097] Based on the above, the embodiment of the present application performs word segmentation processing on the candidate text data sample, and the method of obtaining the word segmentation result corresponding to the candidate text data sample can be as follows:
[0098] For example, the computer device may obtain at least two preset word segmentation strategies; perform word segmentation processing on the candidate text data sample according to the preset word segmentation strategies, and obtain the word segmentation results corresponding to each preset word segmentation strategy.
[0099] The preset word segmentation strategy may refer to a strategy for word segmentation processing of candidate text data samples, and the strategy may be an algorithm or a word segmentation tool, etc. The preset word segmentation strategy may at least include jieba (i.e., stammering) word segmentation strategy, Stanford Core NLP (i.e., Stanford Core NLP) word segmentation strategy, THULAC (i.e., THU Lexical Analyzer for Chinese, Chinese grammar analyzer) word segmentation strategy, PKUSEG word segmentation strategy, LTP-4.0 (i.e., Language Technology Platform-4.0) word segmentation strategy, TexSmart word segmentation strategy, and HanLP word segmentation strategy.
[0100] B2. Filter out a reference text data sample from at least one candidate text data sample according to the word segmentation result.
[0101] Among them, there are multiple word segmentation results of the candidate text data sample implemented in this application. Based on this, the embodiment of this application can screen out a reference text data sample from at least one candidate text data sample according to the word segmentation result as follows:
[0102] For example, the computer device may obtain, for each candidate text data sample, the number of different word segmentation results in the word segmentation results of the candidate text data sample; if the number of different word segmentation results is greater than or equal to a first preset number threshold, the candidate text data sample is used as a reference text data sample.
[0103] The first preset quantity threshold may be set to 2, but is not limited to 2.
[0104] For another example, the computer device may obtain the number of identical segmentation results for each candidate text data sample if there are identical segmentation results in the segmentation results of the candidate text data sample; if the number of identical segmentation results is less than a second preset number threshold, the candidate text data sample may be used as a reference text data sample.
[0105] The second preset quantity threshold may be set to 4, but is not limited to 4.
[0106] Among them, it can be understood that the reference text data sample is a candidate text data sample that is not easily segmented, and the candidate text data samples in the candidate text data sample set other than the reference text data sample are candidate text data samples that are easily segmented.
[0107] In addition to the above, the method for determining the reference text data sample according to the candidate text data sample in the embodiments of the present application can also be as follows:
[0108] For example, the computer device can obtain at least two preset word vectors, and there is a similarity between different preset word vectors; map the candidate text data sample to the vector space to obtain the target vector corresponding to the candidate text data sample; traverse the target vector according to the preset word vector; if there is a target part in the target word vector that is the same as each preset word vector in the at least two preset word vectors, then obtain the quantity of the target part; if the quantity of the target part is greater than the third preset quantity threshold, then use the candidate text data sample as the reference text data sample.
[0109] Among them, the similarity between the preset word vectors can be determined by comparing the target similarity between two preset word vectors; if the target similarity is greater than or equal to the preset target similarity threshold, it is determined that there is a similarity between the two preset word vectors.
[0110] Among them, the candidate text data sample can be a string with a length of 2-5. The third preset quantity threshold can be set to 4, but is not limited to 4.
[0111] For example, the computer device can perform word segmentation processing on the candidate text data sample to obtain the candidate word segmentation result; map the candidate word segmentation result to the vector space to obtain the word segmentation vector corresponding to the candidate word segmentation result; obtain the preset candidate word vector and compare the preset candidate word vector with the word segmentation vector; if the quantity of the word segmentation vectors that are the same as the preset candidate word vector is less than the fourth preset quantity threshold, then use the candidate text data sample as the reference text data sample.
[0112] Among them, the fourth preset quantity threshold can be set to 3, but is not limited to 3.
[0113] Another example is that the computer device can perform word segmentation processing on the candidate text data sample to obtain the target word segmentation result; perform named entity recognition on the candidate text data sample to obtain the named entity recognition result; traverse the target word segmentation result according to the named entity recognition result; if there is a difference between the named entity recognition result and the target word segmentation result, then use the candidate text data sample as the reference text data sample.
[0114] A2. For each word segmentation standard type, use the text data sample's predicted word segmentation information predicted by the text segmentation model to be trained.
[0115] Such as Figure 5As shown in the figure, for each word segmentation standard type in the embodiments of the present application, the method of using the text data sample prediction word segmentation information by the text segmentation model to be trained can be as follows:
[0116] C1. Obtain the reference word segmentation information corresponding to the word segmentation standard type.
[0117] The method for the embodiments of the present application to obtain the reference word segmentation information corresponding to the word segmentation standard type can be as follows:
[0118] For example, the computer device can obtain a set of texts to be screened corresponding to the word segmentation standard type, and the set of texts to be screened includes at least one text to be screened; use the trained text classification model to screen at least one text to be screened to obtain a reference text; perform feature extraction on the reference text to obtain the reference word segmentation information corresponding to the word segmentation standard type.
[0119] Among them, the texts to be screened in the embodiments of the present application may include phrases, and the phrases may include entries. Since there may be errors or the phrases may be incomplete in the phrases, based on this, the embodiments of the present application need to screen the texts to be screened in the set of texts to be screened.
[0120] Among them, the embodiments of the present application can use the embedding layer of the neural network to map the reference text to the vector space to obtain the reference word segmentation information.
[0121] Among them, before the embodiments of the present application use the trained text classification model to screen at least one text to be screened to obtain a reference text, the text classification model to be trained can be trained to obtain the trained text classification model. The method for the embodiments of the present application to train the text classification model to be trained can be as follows:
[0122] For example, the computer device can obtain a reference text data sample and a set of text data samples to be processed, and the set of text data samples to be processed includes at least one text data sample to be processed; determine the positive text data sample and the negative text data sample in at least one text data sample to be processed according to the reference text data sample; converge the text classification model to be trained according to the positive text data sample and the negative text data sample to obtain the trained text classification model.
[0123] Among them, the reference text data sample can be an entry with errors or incomplete, which is an entry that needs to be filtered. The reference text data sample can carry a mark, and based on this, the present application can identify the reference text data sample according to the mark of the reference text data sample.
[0124] Among them, the method for determining the positive text data samples and negative text data samples in at least one text data sample to be processed based on the reference text data sample of the embodiment of this application can be as follows: The computer device can map the reference text data sample to a vector space to obtain reference text features; map the text data sample to be processed to a vector space to obtain the text features to be processed; calculate the first similarity between the reference text features and the text features to be processed; if the first similarity is greater than or equal to the first preset similarity threshold, determine the text data corresponding to the text features to be processed as a positive text data sample; if the first similarity is less than the first preset similarity threshold, determine the text data corresponding to the text features to be processed as a negative text data sample.
[0125] Among them, the method for converging the text classification model to be trained based on the positive text data samples and negative text data samples in the embodiment of this application to obtain the trained text classification model can be as follows: Use the trained text classification model to classify and predict the positive text data samples to obtain the first predicted classification information corresponding to the positive text data samples; use the trained text classification model to classify and predict the negative text data samples to obtain the second predicted classification information corresponding to the negative text data samples.
[0126] The embodiment of this application can obtain the labels of the positive text data samples and the labels of the negative text data samples; calculate the first loss value between the first predicted classification information and the positive text data samples; calculate the second loss value between the second predicted classification information and the negative text data samples; converge the text classification model to be trained according to the first loss value and the second loss value to obtain the trained text classification model.
[0127] C2. Calculate the word segmentation feature information according to the reference word segmentation information and the text data sample.
[0128] The method for calculating the word segmentation feature information according to the reference word segmentation information and the text data sample in the embodiment of this application can be as follows:
[0129] For example, the computer device can extract features from the reference word segmentation information to obtain the reference text feature information corresponding to the reference word segmentation information; extract features from the text data sample to obtain the text feature information corresponding to the text data sample; fuse the reference text feature information and the text feature information to obtain the word segmentation feature information.
[0130] Among them, the embodiment of this application can use the embedding layer to extract features from the reference word segmentation information to obtain the reference text feature information corresponding to the reference word segmentation information; use the embedding layer to extract features from the text data sample to obtain the text feature information corresponding to the text data sample.
[0131] Among them, each text data sample in the embodiments of the present application may correspond to at least one reference word segmentation information, that is, each text feature information may correspond to at least one reference text feature information. Based on this, the embodiments of the present application can calculate the second similarity between the text feature information and the reference text feature information corresponding to the text feature information; determine the weight value of each reference text feature information corresponding to the text feature information according to the second similarity; perform weighted summation on the reference text feature information according to the weight value of each reference text feature information to obtain the weighted reference text feature information; and fuse the weighted reference text feature information and the corresponding text feature information to obtain the word segmentation feature information.
[0132] Among them, the way for the embodiments of the present application to fuse the weighted reference text feature information and the corresponding text feature information may be addition.
[0133] C3. According to the word segmentation feature information, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample.
[0134] Among them, the text word segmentation model to be trained may be a bert model, or a lebert model, etc.
[0135] A3. Converge the text word segmentation model to be trained according to the predicted word segmentation information and the labeled word segmentation information to obtain the text word segmentation model.
[0136] The embodiments of the present application can calculate the third loss value between the predicted word segmentation information and the labeled word segmentation information; converge the text word segmentation model to be trained according to the third loss value to obtain the text word segmentation model.
[0137] The embodiments of the present application can obtain the text to be segmented; determine the annotation information of the text to be segmented according to the text to be segmented; extract the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information; segment the text to be segmented according to the target word segmentation standard type; since the embodiments of the present application can screen out the target word segmentation standard type from the word segmentation standard types of the text to be segmented according to the annotation information of the text to be segmented, the text to be segmented can be accurately segmented according to the target word segmentation standard type, thereby improving the accuracy of word segmentation.
[0138] According to the method described in the above embodiments, the following will give further detailed examples.
[0139] In this embodiment, it is assumed that the word segmentation device is specifically integrated in a computer device, and the computer device may be a server or a terminal.
[0140] As Figure 6 shown, a text word segmentation method has the following specific process:
[0141] S201. The computer device obtains a set of text data samples.
[0142] Among them, the set of text data samples includes text data samples corresponding to each word segmentation standard type, and label word segmentation information corresponding to each text data sample.
[0143] Among them, the word segmentation standard type can include multiple word segmentation standard types, which can include national word segmentation standard types, Peking University word segmentation standard types, and user-defined word segmentation standard types.
[0144] In the national word segmentation standard corresponding to the national word segmentation standard type, the word segmentation methods for parts of speech such as nouns, verbs, adjectives, pronouns, numerals, quantifiers, adverbs, prepositions, conjunctions, auxiliary words, modal particles, interjections, and onomatopoeia are introduced. On the basis of the national word segmentation standard, the Peking University word segmentation standard corresponding to the Peking University word segmentation standard type supplements and adjusts the word segmentation methods of the national word segmentation standard.
[0145] Since the text to be segmented used in the embodiments of the present application is a text that is difficult to segment during the word segmentation process, based on this, the embodiments of the present application can customize two word segmentation standard types, and these two customized word segmentation standard types include a coarse-grained word segmentation standard and a fine-grained word segmentation standard. The granularity represents the degree of detail of word segmentation, and the degree of detail of word segmentation of the coarse-grained word segmentation standard is less than that of the fine-grained word segmentation standard.
[0146] For example, for the sentence "Nanshan District, Shenzhen City, Guangdong Province", when performing word segmentation using the coarse-grained word segmentation standard, the word segmentation result is "Guangdong Province / Shenzhen City / Nanshan District"; when performing word segmentation using the fine-grained word segmentation standard, the word segmentation result is "Guangdong / Province / Shenzhen / City / Nanshan / District".
[0147] Based on the above, the set of text data samples in the embodiments of the present application includes text data samples corresponding to the coarse-grained word segmentation standard, text data samples corresponding to the fine-grained word segmentation standard, and label word segmentation information corresponding to each text data sample. Based on this, using this set of text data samples to train the text segmentation model to be trained can enable the text segmentation model to be trained to learn the characteristics of the coarse-grained word segmentation standard and the fine-grained word segmentation standard.
[0148] Among them, the method for the embodiments of the present application to obtain the set of text data samples can be as follows:
[0149] For example, a computer device may obtain a set of candidate text data samples, where the set of candidate text data samples includes at least one candidate text data sample; determine a reference text data sample according to the candidate text data sample; label the reference text data sample to obtain text data samples corresponding to each word segmentation standard type; obtain label word segmentation information corresponding to the text data samples of each word segmentation standard type; and generate a set of text data samples according to the text data samples corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data samples.
[0150] Among them, the candidate text data sample in the embodiment of this application is Chinese corpus. As Figure 7 shown, obtaining the set of text data samples in the embodiment of this application can be understood as the preparation stage of Chinese corpus.
[0151] The embodiment of this application can prepare Chinese corpora in multiple fields. For example, the candidate text data samples in the set of candidate text data samples include candidate text data samples in multiple fields. For example, the candidate text data samples include candidate text data samples in the news field, candidate text data samples in the movie field, and candidate text data samples in the medical field. And the candidate text data samples in the set of candidate text data samples include candidate text data samples that are difficult to segment and candidate text data samples that are easy to segment. Based on this, the embodiment of this application needs to determine a reference text data sample from the candidate text data samples.
[0152] The candidate text data sample that is difficult to segment may refer to a candidate text data sample that is prone to word segmentation errors. The candidate text data sample that is easy to segment may refer to a candidate text data sample that is not prone to word segmentation errors.
[0153] Based on the above, the method for the embodiment of this application to determine a reference text data sample according to the candidate text data sample may be as follows:
[0154] For example, the computer device performs word segmentation processing on the candidate text data sample to obtain a word segmentation result corresponding to the candidate text data sample; and filters out the reference text data sample from at least one candidate text data sample according to the word segmentation result.
[0155] The embodiment of this application can perform word segmentation processing on the candidate text data sample, so that it can be determined according to the obtained word segmentation result whether the candidate text data sample is a candidate text data sample that is difficult to segment or a candidate text data sample that is easy to segment.
[0156] Based on the above, the method for the embodiment of this application to perform word segmentation processing on the candidate text data sample to obtain a word segmentation result corresponding to the candidate text data sample may be as follows:
[0157] For example, the computer device may obtain at least two preset word segmentation strategies; perform word segmentation processing on the candidate text data sample according to the preset word segmentation strategies, and obtain the word segmentation results corresponding to each preset word segmentation strategy.
[0158] The preset word segmentation strategy may refer to a strategy for word segmentation processing on the candidate text data sample, and the strategy may be an algorithm or a word segmentation tool, etc. The preset word segmentation strategy may at least include jieba (i.e., stammering) word segmentation strategy, StanfordCore NLP (i.e., Stanford Core NLP) word segmentation strategy, THULAC (i.e., THU Lexical Analyzer for Chinese, Chinese grammar analyzer) word segmentation strategy, PKUSEG word segmentation strategy, LTP-4.0 (i.e., Language Technology Platform-4.0) word segmentation strategy, TexSmart word segmentation strategy, and HanLP word segmentation strategy.
[0159] There are multiple word segmentation results for the candidate text data samples implemented in the present application. Based on this, the embodiment of the present application can select a reference text data sample from at least one candidate text data sample according to the word segmentation results as follows:
[0160] For example, the computer device can obtain the number of different word segmentation results in the word segmentation results of each candidate text data sample; if the number of different word segmentation results is greater than or equal to a first preset number threshold, the candidate text data sample is used as a reference text data sample.
[0161] The first preset quantity threshold may be set to 2, but is not limited to 2.
[0162] Furthermore, in the embodiment of the present application, if the number of different word segmentation results is greater than or equal to a first preset number threshold, the process of using the candidate text data sample as a reference text data sample can also be: if the number of different word segmentation results is greater than or equal to the first preset number threshold, the computer device calculates the third similarity between the different word segmentation results in the word segmentation results of the candidate text data sample; if the third similarity is greater than the second preset similarity threshold, the candidate text data sample is determined to be a reference text data sample.
[0163] That is, in the embodiment of the present application, different word segmentation results corresponding to the candidate text data samples have an intersection relationship, that is, there are common parts and different parts between the different word segmentation results corresponding to the candidate text data samples.
[0164] It can be understood that the reference text data sample is a candidate text data sample that is not easily segmented, and the candidate text data samples in the candidate text data sample set other than the reference text data sample are candidate text data samples that are easily segmented.
[0165] Among them, as Figure 7 shown, the classification stage of the embodiment of the present application can be understood as the embodiment of the present application marking the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type.
[0166] The process of the embodiment of the present application marking the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type can be: the computer device can obtain the word segmentation mark corresponding to each word segmentation standard type; according to the word segmentation mark, mark the reference text data sample to obtain the text data sample corresponding to each word segmentation standard type.
[0167] For example, the word segmentation standard types include a coarse-grained word segmentation standard and a fine-grained word segmentation standard. The word segmentation mark corresponding to the coarse-grained word segmentation standard can be 0, and the word segmentation mark corresponding to the fine-grained word segmentation standard can be 1. Based on this, the embodiment of the present application marks the reference text data as 0 to obtain the text data sample corresponding to the coarse-grained word segmentation standard; marks the reference text data as 1 to obtain the text data sample corresponding to the fine-grained word segmentation standard.
[0168] Among them, the labeled word segmentation information corresponding to the candidate text data sample can be the word segmentation information manually labeled according to each word segmentation standard type and stored in the database. The embodiment of the present application can extract the labeled word segmentation information corresponding to the text data sample of each word segmentation standard type according to the word segmentation mark corresponding to each word segmentation standard type.
[0169] S202. For each word segmentation standard type, the computer device uses the text segmentation model to be trained to predict the predicted word segmentation information of the text data sample.
[0170] As Figure 7 shown, the embodiment of the present application includes a training stage for the text segmentation model to be trained.
[0171] As Figure 5 shown, the manner in which the embodiment of the present application uses the text segmentation model to be trained to predict the predicted word segmentation information of the text data sample for each word segmentation standard type can be as follows:
[0172] C1. Obtain the reference word segmentation information corresponding to the word segmentation standard type;
[0173] The manner in which the embodiment of the present application obtains the reference word segmentation information corresponding to the word segmentation standard type can be as follows:
[0174] For example, a computer device can obtain a set of texts to be screened corresponding to a word segmentation standard type. The set of texts to be screened includes at least one text to be screened. The trained text classification model is used to screen at least one text to be screened to obtain a reference text. Feature extraction is performed on the reference text to obtain reference word segmentation information corresponding to the word segmentation standard type.
[0175] Among them, the text to be screened in the embodiments of the present application may include phrases, and the phrases may include entries. Since there may be errors or the phrases may be incomplete in the phrases, based on this, the embodiments of the present application need to screen the texts to be screened in the set of texts to be screened, and the screened reference text can be a complete and error-free text.
[0176] Among them, the embodiments of the present application can use the embedding layer of the neural network to map the reference text to the vector space to obtain the reference word segmentation information.
[0177] Among them, before the embodiments of the present application use the trained text classification model to screen at least one text to be screened to obtain a reference text, the text classification model to be trained can be trained to obtain the trained text classification model. The way for the embodiments of the present application to train the text classification model to be trained can be as follows:
[0178] For example, a computer device can obtain a reference text data sample and a set of text data samples to be processed. The set of text data samples to be processed includes at least one text data sample to be processed. According to the reference text data sample, positive text data samples and negative text data samples in at least one text data sample to be processed are determined. According to the positive text data samples and the negative text data samples, the text classification model to be trained is converged to obtain the trained text classification model.
[0179] Among them, the reference text data sample may be an entry with an error or incomplete, which is an entry that needs to be filtered. The reference text data sample may carry a mark. Based on this, the present application can identify the reference text data sample according to the mark of the reference text data sample.
[0180] Among them, the method for determining positive text data samples and negative text data samples in at least one text data sample to be processed based on the reference text data sample in the embodiments of the present application may be as follows: The computer device may map the reference text data sample to a vector space to obtain reference text features; map the text data sample to be processed to the vector space to obtain the text features to be processed; calculate the first similarity between the reference text features and the text features to be processed; if the first similarity is greater than or equal to the first preset similarity threshold, determine the text data corresponding to the text features to be processed as a positive text data sample; if the first similarity is less than the first preset similarity threshold, determine the text data corresponding to the text features to be processed as a negative text data sample.
[0181] Among them, the method for converging the text classification model to be trained based on the positive text data sample and the negative text data sample in the embodiments of the present application may be as follows: Use the trained text classification model to classify and predict the positive text data sample to obtain the first predicted classification information corresponding to the positive text data sample; use the trained text classification model to classify and predict the negative text data sample to obtain the second predicted classification information corresponding to the negative text data sample.
[0182] The embodiments of the present application may obtain the labels of the positive text data samples and the labels of the negative text data samples; calculate the first loss value between the first predicted classification information and the positive text data sample; calculate the second loss value between the second predicted classification information and the negative text data sample; converge the text classification model to be trained according to the first loss value and the second loss value to obtain the trained text classification model.
[0183] Among them, the text classification model to be trained may be a bert model.
[0184] The accuracy of the text classification model obtained by multiple iterations of the text classification model to be trained for the selected reference text reached 95%. The data volume of this reference text is larger and cleaner, because as many incorrect entries and incomplete entries as possible have been filtered out.
[0185] C2. Calculate the word segmentation feature information according to the reference word segmentation information and the text data sample.
[0186] The method for calculating the word segmentation feature information according to the reference word segmentation information and the text data sample in the embodiments of the present application may be as follows:
[0187] For example, the computer device may perform feature extraction on the reference word segmentation information to obtain the reference text feature information corresponding to the reference word segmentation information; perform feature extraction on the text data sample to obtain the text feature information corresponding to the text data sample; fuse the reference text feature information and the text feature information to obtain the word segmentation feature information.
[0188] Among them, in the embodiments of the present application, an embedding layer can be used to extract features from the reference word segmentation information to obtain reference text feature information corresponding to the reference word segmentation information; an embedding layer can be used to extract features from the text data sample to obtain text feature information corresponding to the text data sample.
[0189] Among them, each text data sample in the embodiments of the present application can correspond to at least one reference word segmentation information, that is, each text feature information can correspond to at least one reference text feature information. Based on this, the embodiments of the present application can calculate the second similarity between the text feature information and the reference text feature information corresponding to the text feature information; determine the weight value of each reference text feature information corresponding to the text feature information according to the second similarity; according to the weight value of each reference text feature information, perform weighted summation on the reference text feature information to obtain weighted reference text feature information; fuse the weighted reference text feature information and the corresponding text feature information to obtain word segmentation feature information.
[0190] Among them, the way of fusing the weighted reference text feature information and the corresponding text feature information in the embodiments of the present application can be addition.
[0191] C3. According to the word segmentation feature information, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample.
[0192] Such as Figure 8 As shown, during the training process, the text word segmentation model to be trained in the embodiments of the present application can be the lebert model. The embodiments of the present application can input the text data sample corresponding to each word segmentation standard type into the text word segmentation model to be trained in the form of "[CLS]+word segmentation mark+[SEP]+text data sample+[SEP]". [CLS] is a mark that enables the text word segmentation model to be trained to recognize the word segmentation mark, and [SEP] is a mark that enables the text word segmentation model to be trained to recognize the text data sample. The text data sample can be a Chinese sentence, and the embodiments of the present application can input the text data sample into the text word segmentation model to be trained in the form of a vector. The embodiments of the present application can also input the labeled word segmentation information corresponding to the text data sample into the text word segmentation model to be trained.
[0193] In the embodiments of the present application, the text word segmentation model to be trained recognizes the word segmentation mark. According to the word segmentation standard type corresponding to the word segmentation mark, based on the word segmentation standard corresponding to the word segmentation standard type of the word segmentation mark, a transformer layer can be used to extract features from the text data sample to obtain text feature information corresponding to the text data sample.
[0194] Based on the above, the computer device obtains reference word segmentation information, and extracts features from the reference word segmentation information to obtain reference text feature information corresponding to the reference word segmentation information.
[0195] In the embodiments of the present application, each text feature information may correspond to multiple reference text feature information. Based on this, the embodiments of the present application can fuse the text feature information and the reference text feature information through a dictionary adapter, and the dictionary adapter corresponds to the text feature information one by one.
[0196] The embodiments of the present application calculate a second similarity between the text feature information and the reference text feature information corresponding to the text feature information through a dictionary adapter; determine the weight value of each reference text feature information corresponding to the text feature information according to the second similarity; according to the weight value of each reference text feature information, perform weighted summation on the reference text feature information to obtain weighted reference text feature information; fuse the weighted reference text feature information and the corresponding text feature information to obtain word segmentation feature information.
[0197] Then, the embodiments of the present application transmit the word segmentation feature information to the transformer layer for feature extraction to obtain target word segmentation feature information, and then transmit the target word segmentation feature information to the CRF layer for prediction to obtain predicted word segmentation information.
[0198] S203. The computer device converges the text segmentation model to be trained according to the predicted word segmentation information and the labeled word segmentation information to obtain a text segmentation model.
[0199] Among them, the embodiments of the present application can calculate a third loss value between the predicted word segmentation information and the labeled word segmentation information; converge the text segmentation model to be trained according to the third loss value to obtain a text segmentation model.
[0200] In the embodiments of the present application, due to the adoption of the reference word segmentation information of the reference text, compared with the existing related technologies, the convergence speed of the text segmentation model to be trained in the embodiments of the present application is faster, and the expected result can be achieved in less time.
[0201] The embodiments of the present application use the reference word segmentation information obtained by feature extraction of the reference text to train the text segmentation model to be trained. The convergence speed of the text segmentation model to be trained in the embodiments of the present application is faster than that of the existing text segmentation model to be trained. As Figure 9 shown, by comparing the loss curve S1 of the text segmentation model to be trained in the embodiments of the present application and the loss curve S2 of the existing text segmentation model to be trained on the text data sample set, it can be clearly seen that the convergence speed of the text segmentation model to be trained in the embodiments of the present application is faster.
[0202] The present embodiment also uses a test set to test the text segmentation model of the present embodiment, and the F1 value of the text segmentation model of the present embodiment is 97.5%. The F1 value refers to the harmonic mean of the precision value and the recall rate.
[0203] S204: The computer device obtains the text to be segmented.
[0204] The text to be segmented in the embodiment of the present application may be extracted from a database of a computer device, or may be acquired online in real time.
[0205] S205: The computer device determines the tagging information of the text to be segmented according to the text to be segmented.
[0206] The annotation information can be understood as the marking information of the text to be segmented for the segmentation standard type. The annotation information of the embodiment of the present application can be represented by an identifier.
[0207] In the embodiment of the present application, there are multiple ways to determine the annotation information of the text to be segmented according to the text to be segmented, as detailed below:
[0208] For example, a computer device can obtain at least two current word segmentation strategies; perform word segmentation processing on the text to be segmented according to the current word segmentation strategies to obtain the current word segmentation results corresponding to each current word segmentation strategy; extract the number of identical current word segmentation results from the current word segmentation results, and determine the annotation information of the text to be segmented according to the number of identical current word segmentation results.
[0209] The current word segmentation strategy may refer to a strategy for word segmentation of a text to be segmented, and the strategy may be an algorithm or a word segmentation tool, etc. The current word segmentation strategy may at least include jieba (i.e., stammering) word segmentation strategy, StanfordCore NLP (i.e., Stanford Core NLP) word segmentation strategy, THULAC (i.e., THU Lexical Analyzer for Chinese) word segmentation strategy, PKUSEG word segmentation strategy, LTP-4.0 (i.e., Language Technology Platform-4.0) word segmentation strategy, TexSmart word segmentation strategy, and HanLP word segmentation strategy.
[0210] Among them, the method for determining the annotation information of the text to be segmented according to the number of the same current word segmentation results in the embodiment of the present application can be: the computer device determines the difficulty level of the text to be segmented according to the number of the same current word segmentation results; and according to the difficulty level, extracts the annotation information corresponding to the difficulty level from a number of candidate annotation information.
[0211] S206. The computer device extracts the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information.
[0212] In the embodiments of the present application, the method for extracting the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information may be as follows:
[0213] For example, the computer device obtains a set of mapping relationships, where the set of mapping relationships includes the mapping relationships between preset annotation information and word segmentation standard types for the text to be segmented; and determines the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type in the set of mapping relationships according to the set of mapping relationships and the annotation information.
[0214] In this way, the embodiments of the present application can accurately perform word segmentation processing on the text to be segmented according to the target word segmentation standard of the text to be segmented.
[0215] S207. Perform word segmentation on the text to be segmented by using a text word segmentation model according to the target word segmentation standard type.
[0216] In the embodiments of the present application, the method for performing word segmentation processing on the text to be segmented by using a text word segmentation model may be as follows:
[0217] For example, the computer device may extract features from the text to be segmented by using a text word segmentation model according to the target word segmentation standard type to obtain feature information corresponding to the target word segmentation standard type; and perform word segmentation processing on the text to be segmented by using the text word segmentation model according to the feature information.
[0218] Among them, the text word segmentation module may be a lebert model.
[0219] Among them, the text word segmentation model in the embodiments of the present application may extract features from the text to be segmented according to the target word segmentation standard type, so as to obtain feature information corresponding to the target word segmentation standard type; and perform word segmentation processing on the text to be segmented by using the text word segmentation model according to the feature information to obtain a word segmentation result corresponding to the target word segmentation standard type.
[0220] There may be multiple word segmentation standard types in the embodiments of the present application. The embodiments of the present application may take the word segmentation standard type including a coarse-grained word segmentation standard type and a fine-grained word segmentation standard type as an example for further elaboration.
[0221] Embodiments of the present application can obtain a text to be segmented; according to the text to be segmented, obtain annotation information corresponding to the coarse-grained word segmentation standard type of the text to be segmented; according to the annotation information corresponding to the coarse-grained word segmentation standard type, extract the coarse-grained word segmentation standard type from at least one word segmentation standard type for the text to be segmented; according to the coarse-grained word segmentation standard type, perform word segmentation on the text to be segmented to obtain a coarse-grained word segmentation result. The coarse-grained word segmentation result can be used as a new text to be segmented. Embodiments of the present application can annotate the coarse-grained word segmentation result to obtain annotation information of the coarse-grained word segmentation result. In the embodiments of the present application, the annotation information of the coarse-grained word segmentation result can be annotation information corresponding to the fine-grained word segmentation standard type.
[0222] Based on the above, embodiments of the present application can obtain annotation information corresponding to the fine-grained word segmentation standard type of the coarse-grained word segmentation result according to the coarse-grained word segmentation result; according to the annotation information corresponding to the fine-grained word segmentation standard type, extract the fine-grained word segmentation standard type from at least one word segmentation standard type for the coarse-grained word segmentation result; according to the fine-grained word segmentation standard, perform word segmentation on the coarse-grained word segmentation result to obtain a fine-grained word segmentation result.
[0223] That is, the present application can preset rules to perform word segmentation on the text to be segmented according to the coarse-grained word segmentation standard type to obtain a coarse-grained word segmentation result, and then perform word segmentation on the coarse-grained word segmentation result according to the fine-grained word segmentation standard type to obtain a fine-grained word segmentation result.
[0224] Based on the above, it can be understood that the text data sample set annotated by the embodiments of the present application has a greater difficulty, a wider coverage area, and a larger quantity, and can provide text data samples corresponding to each word segmentation standard type. Moreover, during the training process of the text word segmentation model to be trained, the embodiments of the present application can simultaneously use text data samples of different word segmentation standard types to train the text word segmentation model to be trained, which is more convenient and can reduce the number of training times.
[0225] Embodiments of the present application can obtain a text to be segmented; according to the text to be segmented, determine the annotation information of the text to be segmented; according to the annotation information, extract the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented; according to the target word segmentation standard type, perform word segmentation on the text to be segmented; since the embodiments of the present application can screen out the target word segmentation standard type from the word segmentation standard types of the text to be segmented according to the annotation information of the text to be segmented, the text to be segmented can be accurately segmented according to the target word segmentation standard type, thereby improving the accuracy of word segmentation.
[0226] To better implement the above method, an embodiment of the present application further provides a text tokenization device, which can be integrated in a computer device, such as a server or a terminal, etc. The terminal may include a tablet computer, a laptop computer, and / or a personal computer, etc.
[0227] For example, as Figure 10 shown, the tokenization device may include an acquisition unit 301, a determination unit 302, an extraction unit 303, a tokenization unit 304, and a training unit 305, as follows:
[0228] (1) Acquisition unit 301;
[0229] The acquisition unit 301 can be used to acquire the text to be tokenized.
[0230] (2) Determination unit 302;
[0231] The determination unit 302 can be used to determine the annotation information of the text to be tokenized according to the text to be tokenized.
[0232] In some embodiments, the determination unit 302 may specifically be used to acquire at least two current tokenization strategies; perform tokenization processing on the text to be tokenized according to the current tokenization strategies to obtain the current tokenization results corresponding to each current tokenization strategy; extract the number of the same current tokenization results from the current tokenization results, and determine the annotation information of the text to be tokenized according to the number of the same current tokenization results.
[0233] (3) Extraction unit 303;
[0234] The extraction unit 303 can be used to extract the target tokenization standard type corresponding to the annotation information from at least one tokenization standard type for the text to be tokenized according to the annotation information.
[0235] In some embodiments, the extraction unit 303 may specifically be used to acquire a mapping relationship set, where the mapping relationship set includes the mapping relationship between the preset annotation information and the tokenization standard type for the text to be tokenized; determine the target tokenization standard type corresponding to the annotation information from at least one tokenization standard type of the mapping relationship set according to the mapping relationship set and the annotation information.
[0236] (4) Tokenization unit 304;
[0237] The tokenization unit 304 can be used to tokenize the text to be tokenized according to the target tokenization standard type.
[0238] In some embodiments, the tokenization unit 304 may specifically be used to perform feature extraction and tokenization processing on the text to be tokenized according to the target tokenization standard type by using a text tokenization model, where the tokenization processing is performed according to the feature information obtained by the text tokenization model for feature extraction of the text to be tokenized.
[0239] (5) Training unit 305;
[0240] The training unit 305 can be used to obtain a set of text data samples, where the set of text data samples includes text data samples corresponding to each word segmentation standard type and label word segmentation information corresponding to each text data sample; for each word segmentation standard type, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample; converge the text word segmentation model to be trained according to the predicted word segmentation information and the label word segmentation information to obtain the text word segmentation model.
[0241] In some embodiments, the training unit 305 can specifically be used to obtain a set of candidate text data samples, where the set of candidate text data samples includes at least one candidate text data sample; determine a reference text data sample according to the candidate text data sample; mark the reference text data sample to obtain text data samples corresponding to each word segmentation standard type; obtain the label word segmentation information corresponding to the text data samples of each word segmentation standard type; generate a set of text data samples according to the text data samples corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data samples.
[0242] In some embodiments, the training unit 305 can specifically be used to perform word segmentation processing on the candidate text data sample to obtain the word segmentation result corresponding to the candidate text data sample; screen out the reference text data sample from at least one candidate text data sample according to the word segmentation result.
[0243] In some embodiments, the training unit 305 can specifically be used to obtain at least two preset word segmentation strategies; perform word segmentation processing on the candidate text data sample according to the preset word segmentation strategies to obtain the word segmentation result corresponding to each preset word segmentation strategy.
[0244] In some embodiments, the training unit 305 can specifically be used to, for each candidate text data sample, obtain the number of different word segmentation results in the word segmentation result of the candidate text data sample; if the number of different word segmentation results is greater than or equal to a preset number threshold, use the candidate text data sample as the reference text data sample.
[0245] In some embodiments, the training unit 305 can specifically be used to obtain the reference word segmentation information corresponding to the word segmentation standard type; calculate the word segmentation feature information according to the reference word segmentation information and the text data sample; predict the predicted word segmentation information of the text data sample using the text word segmentation model to be trained according to the word segmentation feature information.
[0246] In some embodiments, the training unit 305 may specifically be configured to obtain a set of texts to be screened corresponding to the word segmentation standard type, where the set of texts to be screened includes at least one text to be screened; screen at least one text to be screened by using the trained text classification model to obtain a reference text; and extract features from the reference text to obtain reference word segmentation information corresponding to the word segmentation standard type.
[0247] In some embodiments, the training unit 305 may specifically be configured to obtain a reference text data sample and a set of text data samples to be processed, where the set of text data samples to be processed includes at least one text data sample to be processed; determine positive text data samples and negative text data samples in at least one text data sample to be processed according to the reference text data sample; and converge the text classification model to be trained according to the positive text data samples and the negative text data samples to obtain a trained text classification model.
[0248] In some embodiments, the training unit 305 may specifically be configured to extract features from the reference word segmentation information to obtain reference text feature information corresponding to the reference word segmentation information; extract features from the text data sample to obtain text feature information corresponding to the text data sample; and fuse the reference text feature information and the text feature information to obtain word segmentation feature information.
[0249] As can be seen from the above, the obtaining unit 301 in the embodiments of the present application may be configured to obtain a text to be segmented; the determining unit 302 may be configured to determine annotation information of the text to be segmented according to the text to be segmented; the extracting unit 303 may be configured to extract a target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type for the text to be segmented according to the annotation information; the word segmentation unit 304 may be configured to segment the text to be segmented according to the target word segmentation standard type; since the embodiments of the present application can screen out the target word segmentation standard type from the word segmentation standard types of the text to be segmented according to the annotation information of the text to be segmented, the text to be segmented can be accurately segmented according to the target word segmentation standard type, thereby improving the accuracy of word segmentation.
[0250] The embodiments of the present application further provide a computer device, as Figure 11 shown, which shows a schematic structural diagram of the computer device involved in the embodiments of the present application. Specifically:
[0251] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 11 the structural diagram of the computer device shown in
[0252] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the computer device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, computer programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0253] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, computer programs required for at least one function (such as the sound playback function, image playback function, etc.), etc.; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0254] The computer device also includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0255] The computer device may further include an input unit 404, which can be used to receive input digital or character information for communication, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0256] Although not shown, the computer device may further include a display unit and the like, which will not be elaborated herein. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more computer programs into the memory 402 according to the following instructions, and the processor 401 will run the computer programs stored in the memory 402 to achieve various functions as follows:
[0257] Obtain the text to be segmented; determine the annotation information of the text to be segmented according to the text to be segmented; extract the target segmentation standard type corresponding to the annotation information from at least one segmentation standard type for the text to be segmented according to the annotation information; segment the text to be segmented according to the target segmentation standard type.
[0258] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.
[0259] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program or by controlling relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0260] Therefore, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute any one of the text segmentation methods provided by the embodiments of the present application.
[0261] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.
[0262] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0263] Since the instructions stored in the computer-readable storage medium can execute the steps in any one of the text segmentation methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the text segmentation methods provided by the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments, which will not be elaborated herein.
[0264] Wherein, according to an aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations provided in the foregoing embodiments.
[0265] The above has introduced in detail a text tokenization method, apparatus, computer device, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A text word segmentation method, characterized in that, Including: Obtain the text to be segmented; Determine the annotation information of the text to be segmented according to the text to be segmented; Obtain a set of mapping relationships, where the set of mapping relationships includes the mapping relationships between preset annotation information and the word segmentation standard types for the text to be segmented; Determine the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type in the set of mapping relationships according to the set of mapping relationships and the annotation information; Segment the text to be segmented according to the target word segmentation standard type; The determining the annotation information of the text to be segmented according to the text to be segmented includes: Obtain at least two current word segmentation strategies; Perform word segmentation processing on the text to be segmented according to the current word segmentation strategies to obtain the current word segmentation results corresponding to each current word segmentation strategy; Extract the number of the same current word segmentation results from the current word segmentation results, and determine the annotation information of the text to be segmented according to the number of the same current word segmentation results.
2. The text segmentation method according to claim 1, wherein The segmenting the text to be segmented according to the target word segmentation standard type includes: According to the target word segmentation standard type, use a text word segmentation model to perform feature extraction and word segmentation processing on the text to be segmented, where the word segmentation processing is performed according to the feature information obtained by the text word segmentation model for feature extraction of the text to be segmented.
3. The text tokenization method according to claim 2, characterized in that, Before using the text word segmentation model to perform feature extraction on the text to be segmented according to the target word segmentation standard type, the method further includes: Obtain a set of text data samples, where the set of text data samples includes text data samples corresponding to each word segmentation standard type, and the labeled word segmentation information corresponding to each text data sample; For each word segmentation standard type, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample; Converge the text word segmentation model to be trained according to the predicted word segmentation information and the labeled word segmentation information to obtain a text word segmentation model.
4. The text tokenization method according to claim 3, wherein, The obtaining the set of text data samples includes: Obtain a set of candidate text data samples, where the set of candidate text data samples includes at least one candidate text data sample; Determine a reference text data sample according to the candidate text data sample; Mark the reference text data sample to obtain text data samples corresponding to each word segmentation standard type; Obtain the labeled word segmentation information corresponding to the text data sample of each word segmentation standard type; Generate a set of text data samples according to the text data samples corresponding to each word segmentation standard type and the labeled word segmentation information corresponding to the text data sample.
5. The text segmentation method according to claim 4, wherein, The determining the reference text data sample according to the candidate text data sample includes: Perform word segmentation processing on the candidate text data sample to obtain the word segmentation result corresponding to the candidate text data sample; Filter out the reference text data sample from the at least one candidate text data sample according to the word segmentation result.
6. The text tokenization method according to claim 5, wherein The performing word segmentation processing on the candidate text data sample to obtain the word segmentation result corresponding to the candidate text data sample includes: Obtain at least two preset word segmentation strategies; According to the preset word segmentation strategy, perform word segmentation on the candidate text data sample to obtain the word segmentation result corresponding to each preset word segmentation strategy.
7. The text segmentation method according to claim 5, characterized in that, The screening of the reference text data sample from the at least one candidate text data sample according to the word segmentation result includes: For each candidate text data sample, obtain the number of different word segmentation results in the word segmentation result of the candidate text data sample; If the number of different word segmentation results is greater than or equal to the preset quantity threshold, use the candidate text data sample as the reference text data sample.
8. The text segmentation method according to claim 3, wherein For each word segmentation standard type, predicting the predicted word segmentation information of the text data sample by using the text word segmentation model to be trained includes: Obtain the reference word segmentation information corresponding to the word segmentation standard type; Calculate the word segmentation feature information according to the reference word segmentation information and the text data sample; Predict the predicted word segmentation information of the text data sample by using the text word segmentation model to be trained according to the word segmentation feature information.
9. The text segmentation method according to claim 8, characterized in that, The obtaining of the reference word segmentation information corresponding to the word segmentation standard type includes: Obtain the text set to be screened corresponding to the word segmentation standard type, where the text set to be screened includes at least one text to be screened; Use the trained text classification model to screen at least one text to be screened to obtain the reference text; Extract features from the reference text to obtain the reference word segmentation information corresponding to the word segmentation standard type.
10. The text tokenization method according to claim 8, wherein The calculating of the word segmentation feature information according to the reference word segmentation information and the text data sample includes: Extract features from the reference word segmentation information to obtain the reference text feature information corresponding to the reference word segmentation information; Extract features from the text data sample to obtain the text feature information corresponding to the text data sample; Fuse the reference text feature information and the text feature information to obtain the word segmentation feature information.
11. A text tokenization device, characterized in that, It includes: An obtaining unit, configured to obtain the text to be segmented; A determining unit, configured to determine the annotation information of the text to be segmented according to the text to be segmented; An extracting unit, configured to obtain a mapping relationship set, where the mapping relationship set includes the mapping relationship between the preset annotation information and the word segmentation standard type for the text to be segmented; determine the target word segmentation standard type corresponding to the annotation information from at least one word segmentation standard type in the mapping relationship set according to the mapping relationship set and the annotation information; A word segmentation unit, configured to perform word segmentation on the text to be segmented according to the target word segmentation standard type; The determining unit is specifically configured to: Obtain at least two current word segmentation strategies; Perform word segmentation on the text to be segmented according to the current word segmentation strategy to obtain the current word segmentation result corresponding to each current word segmentation strategy; Extract the number of the same current word segmentation results from the current word segmentation results, and determine the annotation information of the text to be segmented according to the number of the same current word segmentation results.
12. The text segmentation device according to claim 11, characterized in that, The word segmentation unit is specifically configured to: According to the target word segmentation standard type, a text word segmentation model is used to perform feature extraction and word segmentation processing on the text to be segmented, wherein the word segmentation processing is performed according to the feature information obtained by the text word segmentation model for the feature extraction of the text to be segmented.
13. The text tokenization device according to claim 12, wherein The text word segmentation device further includes a training unit, and the training unit is specifically used for: Obtain a text data sample set, where the text data sample set includes text data samples corresponding to each word segmentation standard type, and label word segmentation information corresponding to each text data sample; For each word segmentation standard type, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample; Converge the text word segmentation model to be trained according to the predicted word segmentation information and the label word segmentation information to obtain a text word segmentation model.
14. The text tokenization device according to claim 13, characterized in that, The training unit is specifically used for: Obtain a candidate text data sample set, where the candidate text data sample set includes at least one candidate text data sample; Determine a reference text data sample according to the candidate text data sample; Mark the reference text data sample to obtain text data samples corresponding to each word segmentation standard type; Obtain the label word segmentation information corresponding to the text data sample of each word segmentation standard type; Generate a text data sample set according to the text data sample corresponding to each word segmentation standard type and the label word segmentation information corresponding to the text data sample.
15. The text segmentation device according to claim 14, characterized in that, The training unit is specifically used for: Perform word segmentation processing on the candidate text data sample to obtain a word segmentation result corresponding to the candidate text data sample; According to the word segmentation result, screen out a reference text data sample from the at least one candidate text data sample.
16. The text tokenization device according to claim 15, characterized in that, The training unit is specifically used for: Obtain at least two preset word segmentation strategies; According to the preset word segmentation strategies, perform word segmentation processing on the candidate text data sample to obtain word segmentation results corresponding to each preset word segmentation strategy.
17. The text tokenization device according to claim 15, characterized in that, The training unit is specifically used for: For each candidate text data sample, obtain the number of different word segmentation results in the word segmentation result of the candidate text data sample; If the number of different word segmentation results is greater than or equal to a preset number threshold, use the candidate text data sample as the reference text data sample.
18. The text word segmentation device according to claim 13, characterized in that, The training unit is specifically used for: Obtain the reference word segmentation information corresponding to the word segmentation standard type; Calculate word segmentation feature information according to the reference word segmentation information and the text data sample; According to the word segmentation feature information, use the text word segmentation model to be trained to predict the predicted word segmentation information of the text data sample.
19. The text tokenization device according to claim 18, wherein The training unit is specifically used for: Obtain a text set to be screened corresponding to the word segmentation standard type, where the text set to be screened includes at least one text to be screened; Use the trained text classification model to screen at least one text to be screened to obtain a reference text; Perform feature extraction on the reference text to obtain the reference word segmentation information corresponding to the word segmentation standard type.
20. The text tokenization device according to claim 18, characterized in that, The training unit is specifically used for: Perform feature extraction on the reference word segmentation information to obtain reference text feature information corresponding to the reference word segmentation information; Extract features from the text data sample to obtain the text feature information corresponding to the text data sample; Fuse the reference text feature information and the text feature information to obtain the word segmentation feature information.
21. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the text word segmentation method according to any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the text word segmentation method according to any one of claims 1 to 10.
23. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps in the text word segmentation method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Chinese text word segmentation method and device and storage medium
CN112989819A