Word Segmentation Method and Device
The pre-trained language model uses word segmentation to process the corpus fragments, which solves the problems of inefficiency and high cost in the prior art, and realizes an efficient word segmentation scheme.
Patent Information
- Application Number
- CN202210195487.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-01
AI Technical Summary
The existing Chinese word segmentation method relies on manual construction of dictionaries or word segmentation texts, which leads to inefficient and costly, and urgently needs to improve word segmentation efficiency.
The pre-trained language model is used to process the word segmentation of the corpus fragments. By inserting the mask fragments and using the pre-trained language model to predict the corpus information, avoiding the use of dictionary or word segmentation text, and directly completing word segmentation.
It greatly reduces the construction cost of word segmentation scheme, improves word segmentation efficiency, and realizes efficient word segmentation processing.
Smart Images

Figure CN114676697B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a word segmentation method and apparatus. Background Art
[0002] With the development of artificial intelligence technology, Natural Language Processing (NLP) has become one of the important branches. In natural language processing, it is necessary to segment corpus data to provide a basis for subsequent semantic recognition.
[0003] Currently, there are mainly two methods for Chinese word segmentation: one is based on the dictionary segmentation algorithm, that is, matching the string to be matched with an artificially constructed dictionary. If a word corresponding to the string is found in the dictionary, it means the match is successful and the word can be recognized. For example, the forward maximum matching method, the backward maximum matching method, the bidirectional matching segmentation method, etc. The other method is the statistics-based word segmentation method, that is, based on a large-scale segmented text constructed manually, using a statistical machine learning model to label and train Chinese characters, so as to realize the segmentation of unknown text. For example, algorithms such as HMM, CRF, SVM, and deep learning. In the above methods, the dictionary or the segmented text is usually established manually. Due to the large scale of the dictionary and the segmented text, it requires a lot of manpower, low efficiency, and high establishment and maintenance costs.
[0004] In summary, how to improve the word segmentation efficiency has become an urgent technical problem to be solved. Summary of the Invention
[0005] The present disclosure provides a word segmentation method and apparatus to avoid the problems of reduced efficiency and excessive costs caused by manually constructing a dictionary or segmented text, reduce the construction cost of the word segmentation solution, and improve the word segmentation efficiency.
[0006] According to the first aspect of the embodiments of the present disclosure, the present disclosure provides a word segmentation method, including:
[0007] Dividing the corpus to be processed into multiple corpus segments according to a preset granularity;
[0008] Inserting mask segments between the multiple corpus segments, and inputting the corpus to be predicted including the multiple corpus segments and the mask segments into a pre-trained language model;
[0009] Predicting the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model;
[0010] Performing word segmentation processing on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result.
[0011] According to the second aspect of the embodiments of the present disclosure, the present disclosure provides a word segmentation apparatus, including:
[0012] A segmentation module, configured to segment a corpus to be processed into multiple corpus segments according to a preset granularity;
[0013] A prediction module, configured to insert mask segments between multiple corpus segments, and input a corpus to be predicted including multiple corpus segments and mask segments into a pre-trained language model; predict corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model;
[0014] A splitting module, configured to perform word segmentation on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result.
[0015] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which includes a processor and a memory. The memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the word segmentation method in the first aspect.
[0016] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by an electronic device, the electronic device can be enabled to execute at least the word segmentation method in the first aspect.
[0017] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the word segmentation method in the first aspect is implemented.
[0018] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0019] In the present disclosure, a corpus to be processed is segmented into multiple corpus segments according to a preset granularity; mask segments are inserted between the multiple corpus segments, and a corpus to be predicted including the multiple corpus segments and the mask segments is input into a pre-trained language model; corpus information in the mask segments adjacent to each of the multiple corpus segments is predicted through the pre-trained language model; word segmentation is performed on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result. In the present disclosure, the corpus information of the mask segments can be predicted through the pre-trained language model, so that the word segmentation of the corpus to be processed is completed through the predicted corpus information and the multiple corpus segments, and word segmentation can be completed without relying on a dictionary or a word segmentation text, avoiding the problems of reduced efficiency and excessive cost caused by manually constructing a dictionary or a word segmentation text in the related art, greatly reducing the construction cost of the word segmentation solution, and improving the word segmentation efficiency. Description of the Drawings
[0020] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0021] Figure 1 It is a schematic diagram of a word segmentation scenario shown according to an exemplary embodiment.
[0022] Figure 2 It is a schematic flowchart of a word segmentation method shown according to an exemplary embodiment.
[0023] Figure 3 It is a schematic structural diagram of a word segmentation device shown according to an exemplary embodiment.
[0024] Figure 4 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0025] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0027] As described above, natural language processing is one of the important branches in the field of artificial intelligence.
[0028] Natural language processing refers to that a computer accepts an input in the form of natural language from a user, and performs a series of operations such as processing and calculation through algorithms defined by humans inside, to simulate the understanding of natural language by humans, and returns the results expected by the user.
[0029] Word segmentation is one of the basic tasks of natural language processing. Briefly speaking, word segmentation is to decompose a sentence or a text paragraph into corpus data in units of characters or words, so as to perform processing and analysis on the words later.
[0030] At present, there are mainly two methods for Chinese word segmentation: one is the dictionary-based word segmentation algorithm, that is, matching the string to be matched with an artificially constructed dictionary. If a word corresponding to the string is found in the dictionary, it means the match is successful and the word can be recognized. For example, the forward maximum matching method, the backward maximum matching method, the bidirectional matching word segmentation method, etc. The other method is the statistics-based word segmentation method, that is, based on a large-scale segmented text constructed artificially, using a statistical machine learning model to label and train Chinese characters, so as to achieve the segmentation of unknown text. For example, algorithms such as HMM, CRF, SVM, and deep learning. In the above methods, the dictionary or segmented text is usually established manually. Due to the large scale of the dictionary and the segmented text, it requires a lot of manpower, has low efficiency, and high establishment and maintenance costs. In summary, how to improve the word segmentation efficiency has become a technical problem to be solved urgently.
[0031] To solve at least one technical problem existing in the related art, the present disclosure provides a word segmentation method and device.
[0032] The core idea of the above technical solution is: dividing the corpus to be processed into multiple corpus segments according to a preset granularity; inserting mask segments between the multiple corpus segments, and inputting the corpus to be predicted including the multiple corpus segments and the mask segments into a pre-trained language model; predicting the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model; performing word segmentation processing on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result. In this solution, the pre-trained language model can predict the corpus information of the mask segments, so as to complete the word segmentation processing of the corpus to be processed through the predicted corpus information and the multiple corpus segments, and the word segmentation can be completed without relying on a dictionary or segmented text, avoiding the problems of reduced efficiency and excessive cost caused by artificially constructing a dictionary or segmented text in the related art, greatly reducing the construction cost of the word segmentation solution, and improving the word segmentation efficiency.
[0033] In the present disclosure, the above solution can be implemented by an electronic device, and the electronic device can be a terminal device such as a robot, a mobile phone, a tablet computer, a wearable device (such as a smart bracelet, VR glasses, etc.), a PC, etc. Taking a robot as an example, it can be implemented by calling a dedicated application program installed in the robot, or by calling other application programs set in the robot, or by the robot calling a cloud server. Or the above solution can also be implemented by a server.
[0034] In the present disclosure, the above solution can also be implemented by multiple electronic devices in cooperation. For example, the server can send the execution result to the terminal device for the terminal device to display the execution result. The server can be a physical server including an independent host, or can be a virtual server hosted by a host cluster, or can be a cloud server, and the present disclosure does not limit this.
[0035] Taking Figure 1 the scenario shown as an example, the terminal device can transmit the recorded corpus to be processed to the server side, and the server executes the above solution to return the matching word segmentation results to the terminal device, so that the terminal device can display the feedback message. In Figure 1 , the terminal device can be one or more of a robot, a mobile phone, and a PC.
[0036] Based on the core idea introduced above, an embodiment of the present disclosure provides a word segmentation method. Figure 2 It is a schematic flowchart of the word segmentation method provided by an exemplary embodiment of the present disclosure. As Figure 2 shown, the method includes:
[0037] 201. Divide the corpus to be processed into multiple corpus segments according to a preset granularity;
[0038] 202. Insert mask segments between multiple corpus segments, and input the corpus to be predicted including multiple corpus segments and mask segments into a pre-trained language model;
[0039] 203. Predict the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model;
[0040] 204. Perform word segmentation processing on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain the target word segmentation result.
[0041] In the above method, the pre-trained language model can predict the corpus information of the mask segments, so as to complete the word segmentation processing of the corpus to be processed through the predicted corpus information and multiple corpus segments, and the word segmentation can be completed without relying on a dictionary or a word segmentation text, avoiding the problems of reduced efficiency and excessive cost caused by manually constructing a dictionary or a word segmentation text in the related art, greatly reducing the construction cost of the word segmentation solution, and improving the word segmentation efficiency.
[0042] The following introduces each step in the word segmentation method in combination with specific embodiments.
[0043] In 201, the corpus to be processed is divided into multiple corpus segments according to a preset granularity.
[0044] In the present disclosure, the preset granularity includes but is not limited to: characters (including single Chinese characters, letters, and symbols), words (including words, phrases, etc.), and sentences. For example, assuming that the corpus to be processed is "Nanjing Yangtze River Bridge" and the preset granularity is characters, then in 201, "Nanjing Yangtze River Bridge" can be divided into 6 corpus segments: "Nan", "Jing", "Chang", "Jiang", "Da", and "Qiao".
[0045] Furthermore, after dividing the to-be-processed corpus into multiple corpus segments, in 202, mask segments are inserted between the multiple corpus segments, and the to-be-predicted corpus including the multiple corpus segments and the mask segments is input into a pre-trained language model. In 203, the pre-trained language model is used to predict the corpus information in the mask segments adjacent to each of the multiple corpus segments.
[0046] In the present disclosure, the pre-trained language model refers to a pre-trained model (Pre-trained Model, PTM) based on a large corpus, which can learn general language representations, is beneficial to downstream NLP tasks, and can avoid training a model from scratch. With the development of computing power, the emergence of deep models (Transformer) and the enhancement of training techniques have continuously promoted the development of PTMs, making them deeper. It can be understood that the pre-trained model is an application of transfer learning, which uses a large amount of corpus to learn the context-related language representations of each member of the input information, and implicitly learns general grammar and semantic knowledge.
[0047] In practical applications, the pre-trained language model includes but is not limited to at least one of the first-generation pre-trained language model, the second-generation pre-trained language model, the knowledge-enhanced pre-trained model, the multi-language / cross-language pre-trained model, the pre-trained model for a specific language, the multi-modal pre-trained model, the pre-trained model for a specific domain, and the pre-trained model for a specific task. Optionally, the pre-trained language model can be BERT-Base (Chinese version) trained based on Chinese corpus.
[0048] In the present disclosure, the mask segment can be understood as a Prompt in the pre-training task. Among them, Prompt is a new natural language processing paradigm, which is mainly manifested in transforming the input-output form of the downstream task into the form in the pre-training task, that is, the form of the Masked Language Model (MLM).
[0049] Specifically, assume that the preset granularity is a character. In 202, first, mask segments are inserted between the multiple corpus segments, and then the to-be-predicted corpus including the multiple corpus segments and the mask segments is input into the pre-trained language model. Furthermore, the pre-trained language model outputs a set of candidate words corresponding to the mask segments adjacent to each of the multiple corpus segments. Among them, the set of candidate words includes multiple words. Optionally, the number of words included in the set of candidate words can be preset. Further, the preset number of words in the set of candidate words are sorted according to confidence. For example, if the preset number is K, then the set of candidate words corresponding to a certain mask segment includes the corpus information ranked top K in terms of confidence. In the present disclosure, the corpus information includes but is not limited to: characters (including single characters, letters, and symbols), words (including single words, phrases, etc.), and sentences.
[0050] Continuing with the above example, assume that the corpus to be processed is "Nanjing Yangtze River Bridge", and the preset granularity is characters. Therefore, the segmented corpus fragments are: "Nan", "Jing", "Chang", "Jiang", "Da", "Qiao". Based on the above assumption, in 202, mask fragments are inserted between the above 6 corpus fragments. Optionally, the corpus fragments are preprocessed to obtain "Nan ⅰ Jing ⅱ Shi ⅲ Chang ⅳ Jiang ⅴ Da ⅵ Qiao", and then, in the order of Latin numerals, the mask bits (i.e., mask fragments) are inserted between "Nan", "Jing", "Chang", "Jiang", "Da", "Qiao" to obtain the corpus to be predicted containing multiple corpus fragments and mask bits.
[0051] Furthermore, assume that the pre-trained language model is BERT-Base (Chinese version) trained based on Chinese corpus. Assume that the mask bit (i.e., mask fragment) is represented as
mask
mask
mask
mask
[0052] In 203, the BERT-Base model predicts the corpus information in the mask bits adjacent to each of the 6 corpus fragments. Optionally, the number of corpus information in the mask bits predicted by the BERT-Base model is preset. Taking the i-th position as an example, assume that the preset number is K. The BERT-Base model predicts the
mask
mask
mask
mask
[0053] In 204, the corpus to be processed is segmented based on multiple corpus fragments and corpus information to obtain the target segmentation result. For example, assume that the corpus to be processed is a piece of text. Then, the corpus to be processed is segmented based on the multiple corpus fragments corresponding to this text and the corpus information in the mask fragments, and the following target segmentation result can be obtained:
[0054] "A long time | a long time | ago | the whole | forest | was | lazy |
[0055] cicadas | only | when | hungry | chirp |
[0056] migratory birds | hide | in | caves | to | spend the winter |
[0057] bears | all year round | are | licking | honey |"
[0058] Specifically, an alternative embodiment of segmenting the to-be-processed corpus based on multiple corpus segments and corpus information to obtain the target segmentation result in 204 can be implemented as follows:
[0059] Compare the corpus information in each masked segment with the adjacent corpus segments; determine the first masked segment in which the corpus information does not match any of the adjacent corpus segments from each masked segment, and mark a segmentation identifier at the position of the first masked segment; segment the multiple corpus segments in the to-be-processed corpus based on the segmentation identifier to obtain the target segmentation result.
[0060] It can be understood that first, the corpus information in each masked segment is compared with the adjacent corpus segments. This step aims to determine whether to segment the position where the masked segment is located based on the matching result between the corpus information in the masked segment and the adjacent corpus segments. Taking "Nan
mask
mask
mask
mask
mask
mask
[0061] Furthermore, based on the above judgment process, determine the first masked segment in which the corpus information does not match any of the adjacent corpus segments from each masked segment, and mark a segmentation identifier at the position of the first masked segment. Taking "Nanjing City
mask
mask
mask
[0062] For example, assume that the preset granularity is a character. Optionally, the set of candidate characters includes a preset number of characters; the preset number of characters in the set of candidate characters are sorted according to confidence.
[0063] Based on the above introduction, in an alternative embodiment, in 203, the set of candidate characters corresponding to the masked segments adjacent to each of the multiple corpus segments is output through a pre-trained language model. The set of candidate characters includes multiple characters.
[0064] Based on this, in the above steps, optionally, the step of determining the first masked segment whose corpus information does not match that of the adjacent corpus segments from each masked segment can be implemented as follows: In the set of candidate characters corresponding to each masked segment, query whether there are characters that are the same as those in the adjacent corpus segments; if there are no characters in the set of candidate characters that are the same as those in the adjacent corpus segments, it is determined that the corpus information of the current masked segment does not match that of the adjacent corpus segments, and the current masked segment is used as the first masked segment.
[0065] Still taking "Nanjing
mask
mask
mask
mask
mask
mask
mask
[0066] Optionally, if there are characters in the set of candidate characters that are the same as those in any of the adjacent corpus segments, it is determined that the corpus information of the current masked segment matches that of the adjacent corpus segments, and the corpus segments adjacent to the current masked segment are merged into one word. Still taking "South
mask
mask
mask
mask
mask
mask
[0067] Optionally, the merged word can be input into the pre-trained language model as a newly added corpus segment, as one of the bases for predicting the corpus information in the masked segments adjacent to each of the multiple corpus segments. For example, after "South" and "Nanjing" are merged into "Nanjing", the corpus to be predicted becomes "Nanjing
mask
mask
mask
mask
[0068] In practical applications, after testing the word segmentation results with the Tsinghua word segmentation corpus test sets (pku_test.utf8, pku_test_gold.utf8) as the annotation sets, the F1 value obtained by the above steps is 65.6%, approaching 70%. F1 is the harmonic mean of the two metrics of precision and recall, and is generally used to evaluate the performance of models for classification problems with biased samples. Usually, the higher the F1 value, the better the model balances precision and recall.
[0069] In the Figure 2 word segmentation method shown, the pre-trained language model can predict the corpus information of the masked segments, so as to complete the word segmentation process of the corpus to be processed through the predicted corpus information and multiple corpus segments, and the word segmentation can be completed without relying on a dictionary or a word segmentation text, avoiding the problems of reduced efficiency and excessive cost caused by manually constructing a dictionary or a word segmentation text in the related technology, greatly reducing the construction cost of the word segmentation scheme, and improving the word segmentation efficiency.
[0070] Figure 3 A word segmentation device provided by an embodiment of the present disclosure. As Figure 3 shown, the word segmentation device includes:
[0071] A partitioning module 301, configured to partition the corpus to be processed into multiple corpus segments according to a preset granularity;
[0072] A prediction module 302, configured to insert masked segments between the multiple corpus segments, and input the corpus to be predicted including the multiple corpus segments and the masked segments into a pre-trained language model; predict the corpus information in the masked segments adjacent to each of the multiple corpus segments through the pre-trained language model;
[0073] A splitting module 303, configured to perform word segmentation on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result.
[0074] Optionally, in the process of the splitting module 303 performing word segmentation on the corpus to be processed based on the multiple corpus segments and the corpus information to obtain a target word segmentation result, it is specifically configured to:
[0075] Compare the corpus information in each masked segment with the adjacent corpus segments;
[0076] Determine a first masked segment in which the corpus information does not match the adjacent corpus segments from each masked segment, and mark a splitting identifier at the first masked segment;
[0077] Split the multiple corpus segments in the corpus to be processed based on the splitting identifier to obtain the target word segmentation result.
[0078] Optionally, the preset granularity is a character.
[0079] In the process of predicting the corpus information in the respective adjacent masked segments of multiple corpus segments by the prediction module 302 through the pre-trained language model, it is specifically configured as follows:
[0080] Output a set of candidate characters corresponding to the respective adjacent masked segments of multiple corpus segments through the pre-trained language model, where the set of candidate characters includes multiple characters.
[0081] Optionally, in the process of the segmentation module 303 determining the first masked segment whose corpus information does not match the adjacent corpus segments from each masked segment, it is specifically configured as follows:
[0082] Query whether there are characters that are the same as those in the adjacent corpus segments in the set of candidate characters corresponding to each masked segment;
[0083] If there are no characters that are the same as those in the adjacent corpus segments in the set of candidate characters, it is determined that the corpus information of the current masked segment does not match the adjacent corpus segments, and the current masked segment is used as the first masked segment.
[0084] Optionally, it further includes a merging module, which is configured to, if there are characters that are the same as those in any of the adjacent corpus segments in the set of candidate characters, determine that the corpus information of the current masked segment matches the adjacent corpus segments, and merge the corpus segments adjacent to the current masked segment into one word.
[0085] Optionally, the set of candidate characters includes a preset number of characters, and the preset number of characters in the set of candidate characters are sorted according to confidence.
[0086] Optionally, the pre-trained language model includes at least one of a first-generation pre-trained language model, a second-generation pre-trained language model, a knowledge-enhanced pre-trained model, a multilingual / cross-lingual pre-trained model, a pre-trained model for a specific language, a multimodal pre-trained model, a pre-trained model for a specific domain, and a pre-trained model for a specific task.
[0087] The above word segmentation device can execute the system or method provided in the foregoing embodiments. For parts not described in detail in this embodiment, reference may be made to the relevant descriptions of the foregoing embodiments, which will not be elaborated here.
[0088] In a possible design, the structure of the above word segmentation device can be implemented as an electronic device. Such as Figure 4As shown, the electronic device may include: a processor 21 and a memory 22. Among them, executable code is stored on the memory 22. When the executable code is executed by the processor 21, it enables the processor 21 to implement at least the word segmentation method provided in the foregoing embodiments.
[0089] Among them, the structure of the electronic device may further include a communication interface 23 for communicating with other devices or communication networks.
[0090] In addition, the present disclosure also provides a computer-readable storage medium including instructions. Executable code is stored on the medium. When the executable code is executed by the processor of the wireless router, it enables the processor to execute the feature data processing method based on the neural network provided in the foregoing embodiments. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0091] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the feature data processing method based on the neural network provided in the foregoing embodiments. The computer program / instructions are implemented by a program running on a terminal or a server.
[0092] Those skilled in the art will readily think of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0093] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A word segmentation method, characterized in that, including: dividing the corpus to be processed into multiple corpus segments according to a preset granularity; inserting mask segments between the multiple corpus segments, and inputting the corpus to be predicted including the multiple corpus segments and the mask segments into a pre-trained language model; predicting the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model; comparing the corpus information in each mask segment with the adjacent corpus segments; querying whether there is a character consistent with the adjacent corpus segments in the set of candidate characters corresponding to each mask segment; if there is no character consistent with the adjacent corpus segments in the set of candidate characters, determining that the corpus information of the current mask segment does not match the adjacent corpus segments, taking the current mask segment as the first mask segment, and marking a segmentation identifier at the first mask segment; if there is a character consistent with any of the adjacent corpus segments in the set of candidate characters, determining that the corpus information of the current mask segment matches the adjacent corpus segments, and merging the corpus segments adjacent to the current mask segment into one word; performing segmentation on the multiple corpus segments in the corpus to be processed based on the segmentation identifier to obtain a target word segmentation result.
2. The method according to claim 1, wherein The preset granularity is a character; The predicting the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model includes: outputting, by the pre-trained language model, a set of candidate characters corresponding to the mask segments adjacent to each of the multiple corpus segments, where the set of candidate characters includes multiple characters.
3. The method according to any one of claims 1 to 2, characterized in that, The set of candidate characters includes a preset number of characters; The preset number of characters in the set of candidate characters are sorted according to confidence.
4. The method according to claim 1, characterized in that, The pre-trained language model includes at least one of a first-generation pre-trained language model, a second-generation pre-trained language model, a knowledge-enhanced pre-trained model, a multi-language / cross-language pre-trained model, a pre-trained model for a specific language, a multi-modal pre-trained model, a pre-trained model for a specific domain, and a pre-trained model for a specific task.
5. A word segmentation device, characterized in that, including: a dividing module configured to divide the corpus to be processed into multiple corpus segments according to a preset granularity; a predicting module configured to insert mask segments between the multiple corpus segments, and input the corpus to be predicted including the multiple corpus segments and the mask segments into a pre-trained language model; predicting the corpus information in the mask segments adjacent to each of the multiple corpus segments through the pre-trained language model; The segmentation module is configured to compare the corpus information in each mask segment with adjacent corpus segments, and query whether there are words that are the same as those in the adjacent corpus segments in the set of candidate words corresponding to each mask segment. If there are no words in the set of candidate words that are the same as those in the adjacent corpus segments, it is determined that the corpus information of the current mask segment does not match the adjacent corpus segments, and the current mask segment is used as the first mask segment, and a segmentation identifier is marked at the first mask segment. And if there are words in the set of candidate words that are the same as those in any of the adjacent corpus segments, it is determined that the corpus information of the current mask segment matches the adjacent corpus segments, and the adjacent corpus segments of the current mask segment are merged into one word, and the plurality of corpus segments in the to-be-processed corpus are segmented based on the segmentation identifier to obtain a target word segmentation result.
6. An electronic device, characterized in that, Including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the word segmentation method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by an electronic device, the electronic device is enabled to execute the word segmentation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Chinese character segmentation method and device
CN109255117A
Chinese named entity recognition method and device and electronic device
CN111126068A