Dictionary construction method, device, electronic device and storage medium
By using the trained category prediction model, selecting and attributing target words based on semantic information in the text, the accuracy problem of the N-Gram method when word segmentation is inaccurate is solved, and a more efficient dictionary construction is achieved.
Patent Information
- Application Number
- CN202111475744.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In the prior art, the N-Gram-based dictionary construction method results in inaccurate word segmentation, which leads to inaccurate extraction of new words, thereby reducing the accuracy of dictionary construction.
The trained category prediction model is used to determine candidate words based on the semantic information of the word in the text to be processed, and the target words are selected through the semantic information, and the target words are assigned to the corresponding category according to the probability value of the target words belonging to multiple word categories.
The accuracy of dictionary construction is improved, and through the use of semantic information, the named entities and categories in the text are accurately identified.
Smart Images

Figure CN114398482B_ABST
Abstract
Description
Background Art
[0002] When performing information recognition on text, a dictionary is usually constructed first, and words of various categories are added to the dictionary. Then, the constructed dictionary is used to recognize the text, so that the categories to which the words contained in the text belong can be quickly and accurately identified.
[0003] In the related art, the N-Gram-based method is often used for dictionary construction. The N-Gram-based dictionary construction method first uses the N-Gram algorithm to segment the text to obtain N words, and then uses multiple existing dictionaries for comparison, and selects new words from the N words to add to the dictionary.
[0004] Since the new words extracted by this method are related to the N words obtained by the N-Gram algorithm, when the words obtained by segmentation using the N-Gram algorithm are inaccurate, the new words extracted from the N words are not accurate enough, so the accuracy of the constructed dictionary is not high. Summary of the invention
[0005] In order to solve the technical problems existing in the related art, the embodiments of the present application provide a dictionary construction method, device, electronic device and storage medium, which can improve the accuracy of dictionary construction.
[0006] The specific technical solutions provided by the embodiments of this application are as follows:
[0007] A dictionary construction method, comprising:
[0008] Obtaining a text to be processed and a basic dictionary; wherein the basic dictionary contains multiple word categories;
[0009] Based on the trained category prediction model, determining at least one candidate word contained in the text to be processed, and semantic information of each of the at least one candidate word;
[0010] By using the category prediction model, at least one target word that meets the set semantic condition is selected according to the semantic information of each of the at least one candidate word, and a probability value of each of the at least one target word belonging to the multiple word categories is determined;
[0011] The at least one target word is respectively assigned to a word category whose corresponding probability value meets the set probability condition.
[0012] A dictionary construction device, comprising:
[0013] An acquisition module, used to acquire the text to be processed and a basic dictionary; wherein the basic dictionary contains multiple word categories;
[0014] A word recognition module, used to determine at least one candidate word contained in the text to be processed and the semantic information of each of the at least one candidate word based on the trained category prediction model;
[0015] A category identification module, configured to select at least one target word that meets a set semantic condition based on the semantic information of each of the at least one candidate word through the category prediction model, and determine a probability value that each of the at least one target word belongs to each of the multiple word categories;
[0016] The dictionary construction module is used to classify the at least one target word into word categories whose corresponding probability values meet the set probability conditions.
[0017] Optionally, the category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the word recognition module is specifically used to:
[0018] Based on the text to be processed, obtaining a word vector corresponding to each of at least one word contained in the text to be processed through the pre-trained language sub-model; wherein each word vector represents semantic information of the corresponding word;
[0019] Based on the word vector corresponding to each of the at least one word, the at least one word is combined through the named entity recognition sub-model to obtain at least one candidate word and semantic information of each of the at least one candidate word; wherein each candidate word contains at least one word.
[0020] Optionally, the category identification module is specifically used to:
[0021] By using the named entity recognition sub-model, at least one target word whose semantic information belongs to a named entity is selected from the at least one candidate word; the named entity is an entity name with specific semantics;
[0022] For the at least one target word, the following operations are performed respectively: by using the named entity recognition sub-model, according to the semantic information of the one target word, the probability values that the one target word belongs to the multiple word categories are determined.
[0023] Optionally, a model training module is also included, and the model training module is used to:
[0024] Acquire a training data set; the training data set includes a plurality of text data samples, and the text data samples are marked with set categories;
[0025] Based on the training data set, the category prediction model is iteratively trained until a set convergence condition is met, wherein one iterative training process includes:
[0026] Based on the text data sample extracted from the training data set, determining at least one target word in the text data sample through the category prediction model, and determining the target word category corresponding to each of the at least one target word;
[0027] According to the target word category and the set category, a corresponding loss value is determined, and according to the loss value, parameters of the category prediction model are adjusted.
[0028] Optionally, the category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the training data set includes encyclopedia text data samples, domain text data samples and named entity recognition text data samples; and the model training module is further used for:
[0029] Based on the text data samples extracted from the encyclopedia text data samples and the domain text data samples, the corresponding embedding vector samples are determined by the pre-trained language sub-model; the encyclopedia text data samples are text data without a targeted domain, and each domain text data sample is text data including a plurality of named entities of set categories and is annotated with a corresponding domain category;
[0030] Based on the vector samples extracted from the embedding vector samples and the word vector samples, at least one corresponding target word and its respective corresponding target word category are determined through the named entity recognition sub-model; the word vector sample is obtained based on the named entity recognition text data sample; the named entity recognition text data sample is text data including at least one named entity, and each named entity word is annotated with a corresponding named entity category.
[0031] Optionally, the model training module is also used to:
[0032] Determining a first loss value according to the target word category and the field category, and adjusting parameters of the pre-trained language sub-model according to the first loss value;
[0033] A second loss value is determined according to the target word category, the field category, and the named entity category, and parameters of the named entity recognition sub-model are adjusted according to the second loss value.
[0034] Optionally, the model training module is also used to:
[0035] The encyclopedia text data samples, the domain text data samples and the named entity recognition text data samples in the training data set are respectively preprocessed; the preprocessing operation includes at least one of data screening and format conversion.
[0036] Optionally, the dictionary construction module is specifically used to:
[0037] Based on the probability values of a target word belonging to the plurality of word categories, determining the word category corresponding to the maximum probability value, and taking the word category as the target category;
[0038] If the probability value of the target word belonging to the target category is greater than a first set threshold, the target word is assigned to the target category.
[0039] Optionally, the dictionary construction module is further used to:
[0040] If the probability values of the target word belonging to the plurality of word categories are not greater than a second set threshold, at least one word category having a probability value greater than a third set threshold is selected as a candidate category; the third set threshold is less than the second set threshold;
[0041] Based on the similarity between the one target word and the words respectively included in the at least one candidate category, a candidate category that meets the set similarity condition is selected as the target category, and the one target word is assigned to the target category.
[0042] An electronic device provided by an embodiment of the present application includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of any one of the above-mentioned dictionary construction methods.
[0043] An embodiment of the present application provides a computer-readable storage medium, which includes a program code. When the program code is executed on an electronic device, the program code is used to enable the electronic device to execute the steps of any one of the above-mentioned dictionary construction methods.
[0044] A computer program product provided in an embodiment of the present application includes a computer program / instruction, which, when executed on a computer, enables the computer to execute the above-mentioned dictionary construction method.
[0045] The beneficial effects of this application are as follows:
[0046] The embodiments of the present application provide a dictionary construction method, device, electronic device and storage medium. After obtaining the text to be processed and a basic dictionary containing multiple word categories, at least one target word whose semantic information belongs to a named entity can be selected from the text to be processed based on the trained category prediction model, and the probability value of at least one target word belonging to multiple word categories can be determined, and at least one target word can be respectively assigned to a word category whose corresponding probability value meets the set probability condition. Compared with the N-Gram-based dictionary construction method in the related art, since the present solution can accurately determine the target word belonging to the named entity in the text to be processed based on the semantic information of the word based on the category prediction model, and determine the target named entity category to which the target word belongs based on the probability value of the target word belonging to the named entity category in the basic dictionary, the accuracy of the dictionary construction can be improved.
[0047] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1a A schematic diagram of an application scenario in an embodiment of the present application;
[0049] Figure 1b This is a schematic diagram of another application scenario in an embodiment of the present application;
[0050] Figure 2a A schematic diagram of a process for training a category prediction model in an embodiment of the present application;
[0051] Figure 2b A schematic diagram of labeling a named entity recognition text data sample in an embodiment of the present application;
[0052] Figure 2c This is a schematic diagram of the output of the embedded vector sample in the embodiment of the present application;
[0053] Figure 2d A schematic diagram of training a pre-trained language sub-model in an embodiment of the present application;
[0054] Figure 2e This is another schematic diagram of training a pre-trained language sub-model in an embodiment of the present application;
[0055] Figure 2f A schematic diagram of training a named entity recognition sub-model in an embodiment of the present application;
[0056] Figure 3A schematic diagram of a process for analyzing and processing a training data set in an embodiment of the present application;
[0057] Figure 4a Schematic diagram of the process of constructing a dictionary in an embodiment of the present application;
[0058] Figure 4b A schematic diagram of determining candidate words in an embodiment of the present application;
[0059] Figure 4c A schematic diagram of determining the output result of the category prediction model in an embodiment of the present application;
[0060] Figure 4d A schematic diagram of a process for determining a target category in an embodiment of the present application;
[0061] Figure 5 A schematic diagram of another process for determining a target category in an embodiment of the present application;
[0062] Figure 6 A schematic diagram of a dictionary construction method in an embodiment of the present application;
[0063] Figure 7 is a structural schematic diagram of a dictionary construction device in an embodiment of the present application;
[0064] Figure 8 is a schematic diagram of the structure of another dictionary construction device in an embodiment of the present application;
[0065] Fig. 9 A schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application;
[0066] Fig.10 Schematic diagram of the structure of a computing device in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.
[0068] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.
[0069] Some of the terms used in the embodiments of the present application are explained below to facilitate understanding by those skilled in the art.
[0070] Dictionary: Extract words of predefined categories from the text to construct the required dictionary. When performing information recognition on a text, the constructed dictionary can be used to recognize the text, which can quickly and accurately identify the category to which the words contained in the text belong.
[0071] Named entity: an entity name with specific semantics in the text, mainly including names of people, places, organizations, proper nouns, etc.
[0072] The word “exemplary” is used hereinafter to mean “serving as an example, example, or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0073] The terms "first" and "second" in this document are used for descriptive purposes only and should not be understood as explicitly or implicitly indicating relative importance or implicitly indicating the number of technical features indicated. Therefore, features defined as "first" and "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0074] The embodiments of the present application relate to artificial intelligence (AI), machine learning (ML) technology and natural language processing (NLP), and are designed based on machine learning technology and natural language processing technology in artificial intelligence.
[0075] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0076] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, mechatronics and other technologies. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, smart transportation and other major directions.
[0077] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0078] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0079] The embodiment of the present application adopts a category prediction model based on machine learning to determine at least one target word in the text, and determines the probability value of the at least one target word belonging to each of the multiple word categories in the basic dictionary.
[0080] The following is a brief introduction to the design concept of the embodiment of the present application:
[0081] Dictionaries can be used to identify information in texts to quickly and accurately determine the predefined categories to which words in the text belong. In the related art, the N-Gram-based method is often used to construct dictionaries. This method first segments the text based on the N-Gram algorithm to obtain N words, and then uses multiple existing dictionaries for comparison, and selects new words from the N words to add to the dictionary. However, the N-Gram algorithm used in this method cannot accurately segment the words in the text, resulting in a low accuracy rate for the dictionary finally constructed.
[0082] In view of this, the embodiments of the present application provide a dictionary construction method, device, electronic device and storage medium, which can first select target words whose semantic information belongs to named entities from the text to be processed based on a trained category prediction model and according to the semantic information of the words in the text, and then assign the target words to the target named entity categories in the dictionary according to the probability values of the target words belonging to each named entity category in the dictionary to complete the construction of the dictionary, thereby improving the accuracy of dictionary construction.
[0083] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.
[0084] See also Figure 1a As shown, it is a schematic diagram of an application scenario in an embodiment of the present application. The application scenario diagram includes a terminal device 100 and a server 200. The terminal device 100 and the server 200 can communicate through a communication network. Optionally, the communication network can be a wired network or a wireless network. The terminal device 100 and the server 200 can be directly or indirectly connected via wired or wireless communication, and the present application does not limit this.
[0085] In the embodiment of the present application, the terminal device 100 is an electronic device used by the user, which may be a personal computer, a mobile phone, a tablet computer, a notebook, an e-book reader, a smart home, a vehicle-mounted terminal, etc. The server 200 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0086] Exemplarily, when a user browses a text on a terminal device 100, the browsed text can be sent to a server 200 through the terminal device 100. The server 200 can determine at least one candidate word contained in the text and the semantic information of at least one candidate word based on the trained category prediction model, and select at least one target word that meets the set semantic conditions according to the semantic information of at least one candidate word through the category prediction model, and determine the probability value that at least one target word belongs to multiple word categories in the basic dictionary. After determining the probability value, the following operations can be performed for at least one target word: based on the probability value that a target word belongs to multiple word categories, a word category that meets the set probability condition is selected as the corresponding target category, and a target word is assigned to the target category, thereby completing the construction and expansion of the dictionary.
[0087] In addition, after determining the target word contained in the text and the word category to which the target word belongs, the server 200 can also send the result to the terminal device 100 so that the terminal 100 displays the category to which the word contained in the text belongs to the user, thereby achieving the purpose of information recognition of the text.
[0088] It should be noted that Figure 1a This is an example introduction to the application scenarios of the dictionary construction method of the present application. In fact, the application scenarios to which the method in the embodiments of the present application can be applied are not limited to this.
[0089] In some embodiments, the application scenarios in the embodiments of the present application can also be as follows Figure 1b As shown, it includes a terminal device 100 and a server 200. The server 200 includes a dictionary, and the dictionary is constructed by the dictionary construction method in the present application.
[0090] Specifically, the terminal device 100 can send the text to be recognized to the server 200. After receiving the text to be recognized, the dictionary in the server 200 can recognize the named entities in the text to be recognized and send the obtained recognition result to the terminal device 100.
[0091] For example, if the text to be recognized is "Xiaohua climbed Mount Tai to watch the sunrise at 5 a.m. on the Mid-Autumn Festival", after the terminal device 100 sends the text to be recognized to the server 200, the text to be recognized is recognized through the dictionary in the server 200, and it can be recognized that the name of the person in the text to be recognized is "Xiaohua", the festival is "Mid-Autumn Festival", the time is "5 a.m.", and the place name is "Mount Tai", that is, "name: Xiaohua, festival: Mid-Autumn Festival, time: 5 a.m., place name: Mount Tai". After obtaining the recognition result, the server 200 can send the recognition result to the terminal device 100.
[0092] Furthermore, after receiving the recognition result, the terminal device 100 can analyze and manage the text to be recognized according to the named entities in the recognition result. For example, assuming that the text to be recognized is a news message, the named entities in the news message can be recognized through the dictionary, thereby extracting important information from the news message, and analyzing the news message to determine whether it is a fake news message or a news message containing sensitive words based on the extracted important information.
[0093] First, the training process of the category prediction model in the embodiment of the present application is described in detail. The process can be executed by the server, for example Figure 1a Server 200 in . Figure 2a The following is a schematic diagram of the training process of the category prediction model in the embodiment of the present application. Figure 2a , the training process of the category prediction model in the embodiment of the present application is elaborated in detail.
[0094] Step S201, obtaining a training data set.
[0095] The acquired training data set may include multiple text data samples. Specifically, the training data set may include encyclopedia text data samples, domain text data samples, and named entity recognition text data samples.
[0096] Among them, encyclopedia text data samples are text data in non-targeted fields. For example, the text "Shipping company Evergreen Marine leased container ships longer granted round the Suez Canal stranded re-float rescue 6 days, the canal restored traffic." and the text "A veteran actor from a certain region, who won the Best Supporting Actor award at the regional film awards, passed away at a hospital at the age of 66." can all be encyclopedia text data. When encyclopedia text data is used as sample data to train a category prediction model, it is not necessary to label each encyclopedia text data sample.
[0097] The domain text data sample is a text data that includes multiple set category named entities, that is, it is highly related to a specific named entity, or contains more of a certain named entity. For example, the text "After the CountryD government transferred the CountryC's territory of western islands to CountryA, CountryA began to control some islands." can be a domain text data, and the text is a place domain text data. Among them, it is assumed that CountryD, CountryC and CountryA represent three different place names. When using domain text data as sample data to train the category prediction model, it is necessary to annotate the corresponding domain category for each domain text data sample. For example, for the above-mentioned domain text data, since the text data contains more place named entities, the text data can be labeled as "Location".
[0098] A named entity recognition text data sample is text data that includes at least one named entity, that is, needs to be annotated with secondary named entities. For example, the text "A cruise ship carrying tourists on Qiandao Lake in Zhejiang Province in southeastern China." can be a named entity recognition text data. When using named entity recognition text data as sample data to train a category prediction model, it is necessary to annotate the corresponding named entity category for each named entity word included in each named entity recognition text data. For each named entity recognition text data sample, the "BIOES annotation method" can be used to annotate the named entity category and relative position of each word contained in the named entity recognition text data sample. For example, Figure 2b As shown in the figure, for the named entity recognition text data "A ship carrying tourists on Qiandao Lake in Zhejiang Province", the named entity recognition text data is annotated using the "BIOES annotation method" and the result is "OOOO B-LOC E-LOC O S-LOC O". Among them, B represents the starting position, I represents the middle position, E represents the ending position, S represents a single named entity, and the label behind it represents the entity category. For example, "-LOC" means that the named entity belongs to a place named entity, and O represents a non-named entity.
[0099] Step S202: extracting text data samples from encyclopedia text data samples and domain text data samples.
[0100] The category prediction model to be trained includes a pre-trained language sub-model and a named entity recognition sub-model. When training the pre-trained language sub-model, text data samples can be extracted from the encyclopedia text data samples and domain text data samples in the training data set as training sample data.
[0101] Step S203: input the extracted text data sample into the pre-trained language sub-model to be trained to obtain a corresponding embedding vector sample.
[0102] The extracted text data sample is input into the pre-trained language sub-model to be trained, and features are extracted from the text data sample based on the pre-trained language sub-model to obtain the corresponding embedding vector sample, which may include the word vectors corresponding to each word in the corresponding text data sample. For example, for a text data sample "After the Country D government transferred the Country C'sterritory", after passing through the pre-trained language sub-model, the following can be obtained: Figure 2c The embedded vector samples shown are E1, E2, E3, …E8.
[0103] Specifically, Figure 2d As shown in FIG. 1 , before the extracted text data sample is input into the pre-trained language sub-model to be trained, the text data sample needs to be processed into a [CLS]+word segmentation+[SEP] structure, that is, the text data sample is processed into a "[CLS], I1, I2, I3, ..., [SEP]" structure. Among them, the special tag [CLS] can be used for the recognition task of the domain category.
[0104] After processing the text data sample into the structure of "[CLS], I1, I2, I3, ..., [SEP]", the text data sample can also be randomly masked to process the text data sample into the structure of "[CLS], I1, [MASK], I3, ..., [SEP]". The random mask mechanism is used to process the text data sample input into the pre-trained language sub-model so that the pre-trained language sub-model can complete the mask prediction task. Among them, the random mask mechanism randomly covers or replaces any word or phrase in a sentence, and then allows the pre-trained language sub-model to predict the covered or replaced part through understanding of the context.
[0105] After the text data sample is randomly masked, the processed text data sample can be used to train the pre-trained language sub-model. First, embed the text data sample, i.e., vectorize it, to obtain "E[CLS], E1, E[MASK], E3, ..., E[SEP]", and then let the vectorized text data sample pass through multiple hidden layers in the pre-trained language sub-model to obtain the output result "C, T1, T[MASK], T3, ..., T[SEP]". Since the text data samples input into the pre-trained language sub-model also include domain text data samples marked with domain categories, by training the pre-trained language sub-model, the pre-trained language sub-model can predict the domain category of the input text data sample. Using the random mask mechanism and domain category recognition tasks at the same time to train the pre-trained language sub-model can enable the pre-trained language sub-model to learn more semantic information.
[0106] For example, the extracted text data sample may be “Shipping company Green Marine leased container ships longer granted round the Suez Canal stranded re-float rescue 6 days, the canal restored traffic.” Figure 2e As shown, before the text data sample is input into the pre-trained language sub-model, the text data sample can be processed into “[CLS] Shipping company Green Marine leased container ships longer granted round in the Suez Canal stranded re-float rescue 6 days, the canal restored traffic. [SEP]”, and the text data sample is randomly masked to obtain “[CLS] Shipping [MASK] Green [MASK] leased container… [SEP]”.
[0107] After the text data sample after random mask processing is input into the pre-trained language sub-model, the pre-trained sub-model can vectorize the text data sample to obtain "E[CLS], E1, E2, E[MASK], E4, E[MASK], E6, E[…], E[SEP]", and then obtain the output result "C, T1, T2, T[MASK], T4, T[MASK], T6, T[…], T[SEP]".
[0108] After the extracted text data samples are randomly masked, they are input into the pre-trained language sub-model for training. The pre-trained language sub-model can predict the masked or replaced parts through understanding the context. That is, for the input "[CLS]Shipping[MASK]Green[MASK]leased container…[SEP]", the pre-trained language sub-model can predict the first "[MASK]" as "company" and the second "[MASK]" as "Marine" through understanding the context.
[0109] Optionally, the pre-trained language sub-model in the embodiment of the present application can be implemented using Bidirectional Encoder Representations from Transformers (BERT).
[0110] In the embodiment of the present application, encyclopedia text data and domain text data with domain category annotations are used as training data to train the pre-trained language sub-model, so that the pre-trained language sub-model can learn more semantic information, so that when performing named entity recognition on text data, the semantic information in the text data can be used to identify the target words belonging to the named entities in the text data faster and better.
[0111] Step S204: extracting vector samples from the embedded vector samples and the word vector samples obtained based on the named entity recognition text data samples.
[0112] In this embodiment, the named entity recognition text data samples of the training data set can be vectorized to obtain corresponding word vector samples. When the named entity recognition text data samples are vectorized, the WordEmbedding method can be used to obtain the word vector samples corresponding to the named entity recognition text data samples. The embodiment of the present application does not limit the method of obtaining the word vector samples from the named entity recognition text data samples.
[0113] After obtaining the word vector samples corresponding to the named entity recognition text data samples, vector samples can be extracted from the embedded vector samples and word vector samples obtained based on the pre-trained language sub-model.
[0114] Step S205: input the extracted vector sample into the named entity recognition sub-model to obtain at least one corresponding target word and its corresponding target word category.
[0115] The extracted vector sample is input into the named entity recognition sub-model. Based on the named entity recognition sub-model, at least one target word belonging to a named entity in the text data sample and the named entity category corresponding to each target word can be identified.
[0116] Specifically, Figure 2f As shown in the figure, assuming that the extracted text data sample is "[CLS] Shipping company Green Marine leased container ships longer... [SEP]", the text data sample has the annotation "[CLS] OO B-ORG E-ORG OOOO... [SEP]". Among them, O represents a non-named entity, B represents the starting position of a named entity, E represents the ending position of a named entity, S represents a single named entity, and the following label represents the entity category, "-LOC" represents the location entity category, and "-ORG" represents the organization entity category.
[0117] The annotated text data sample is input into the pre-trained language sub-model, and the corresponding embedding vector samples "E[CLS], E1, E2, E3, E4, E5, E6, E7, E8, E…, E[SEP]" can be obtained through the pre-trained language sub-model.
[0118] The obtained embedding vector samples are input into the named entity recognition sub-model to train the model. The named entity recognition sub-model includes two sub-models, one is the boundary recognition sub-model (Border Segment Model), which is used to identify the named entity boundaries in the text data, and the other is the category recognition sub-model (Named Entity Recognition Model, NER Model), which is used to identify the named entity category to which the named entity words in the text data belong. After training the two sub-models included in the named entity recognition sub-model, the prediction results of the two sub-models are combined to obtain the final prediction result.
[0119] After the embedded vector samples “E[CLS], E1, E2, E3, E4, E5, E6, E7, E8, E…, E[SEP]” are input into the named entity recognition sub-model, the named entity boundaries in the text data sample can be identified based on the boundary recognition sub-model in the named entity recognition sub-model, that is, the text “[CLS]Shipping company Green Marine leased container ships longer…[SEP]” can be identified through the boundary recognition sub-model, and “OOBEOO OO…” can be obtained. Based on the identified named entity boundaries and embedded vector samples, the named entity category to which each named entity word belongs can be obtained through the category recognition sub-model in the named entity recognition sub-model, that is, the text “[CLS]Shipping company Green Marine leased container ships longer…[SEP]” can be identified through the category recognition sub-model, and “OO ORG ORG OOOO…O” can be obtained. Then, by combining the recognition results of the boundary recognition sub-model with the recognition results of the category recognition sub-model, we can obtain the prediction results of the text “[CLS]Shipping company Green Marine leased container ships longer…[SEP]”, that is, predicting the named entity words in the text “[CLS]Shipping company Green Marine leased container ships longer…[SEP]” and the named entity categories to which the named entity words belong.
[0120] Based on the named entity recognition sub-model, the text “[CLS]Shipping company Green Marineleased container ships longer…[SEP]” is recognized, and the prediction result is “[CLS]O OB-ORGE-ORG OOOO…[SEP]”. According to the prediction results, we can know that “Shipping”, “company”, “leased”, “container”, “ships” and “longer” in the text are all non-named entities, “Green Marine” is a named entity, and “Green Marine” is an organizational named entity.
[0121] In an embodiment of the present application, the named entity recognition sub-model is trained by using the embedding vector output by the pre-trained language sub-model and the word vector obtained based on the named entity recognition text sample with named entity category annotations as training data. The semantic information of multiple words as a whole in the text data learned in the pre-trained language sub-model stage can be utilized to more accurately identify the target words belonging to the named entities in the text data.
[0122] Using encyclopedia text data and domain text data with domain category annotations as training data to train the pre-trained language sub-model can allow the pre-trained language sub-model to learn more semantic information, so that when performing named entity recognition on text data, the semantic information in the text data can be used to identify the target words belonging to the named entities in the text data faster and better.
[0123] Step S206, determining a first loss value according to the target word category and the domain category.
[0124] According to the target word category corresponding to the target word identified by the named entity recognition sub-model and the field category marked by the text data sample where the target word is located, the first loss value can be determined. Generally, the loss value is used to determine the degree of closeness between the actual output and the expected output. The smaller the loss value, the closer the actual output is to the expected output.
[0125] Step S207, determine whether the first loss value converges to a preset target value; if not, execute step S208; if yes, execute step S209.
[0126] Determine whether the first loss value converges to the preset target value. If the first loss value is less than or equal to the preset target value, or the change amplitude of the first loss value obtained after N consecutive trainings is less than or equal to the preset target value, it is considered that the first loss value has converged to the preset target value, indicating that the first loss value has converged; otherwise, it means that the first loss value has not yet converged.
[0127] Step S208: adjusting the parameters of the pre-trained language sub-model to be trained according to the determined first loss value.
[0128] If the first loss value has not converged, the model parameters of the pre-trained language sub-model are adjusted. After the model parameters are adjusted, the process returns to step S202 to continue the next round of training.
[0129] Step S209, determining a second loss value according to the target word category, the domain category, and the named entity category.
[0130] According to the target word category corresponding to the target word identified by the named entity recognition sub-model and the field category or named entity category marked by the text data sample where the target word is located, a second loss value can be determined. Generally, the loss value is used to determine the degree of closeness between the actual output and the expected output. The smaller the loss value, the closer the actual output is to the expected output.
[0131] In an embodiment of the present application, the first loss value determined by using the target word category corresponding to the target word output by the named entity recognition sub-model and the domain category marked in the extracted text data sample is used to adjust the model parameters of the pre-trained language sub-model, and the second loss value determined by using the target word category corresponding to the target word output by the named entity recognition sub-model and the domain category or named entity category marked in the extracted vector sample is used to adjust the model parameters of the named entity recognition sub-model. The training process of the pre-trained language sub-model can be completed first, and after the training of the pre-trained language sub-model is completed, the weight of the pre-trained sub-model will no longer participate in the training of the named entity recognition sub-model, so as to reduce the training time of the model, and enable the named entity recognition sub-model to better combine the two sub-tasks of boundary recognition and category prediction, more accurately identify the target words belonging to the named entities in the text, and accurately attribute the target words to the corresponding named entity categories.
[0132] Step S210, determine whether the second loss value converges to a preset target value; if not, execute step S211; if yes, execute step S212.
[0133] Determine whether the second loss value converges to the preset target value. If the second loss value is less than or equal to the preset target value, or the change amplitude of the second loss value obtained after N consecutive trainings is less than or equal to the preset target value, it is considered that the second loss value has converged to the preset target value, indicating that the second loss value has converged; otherwise, it means that the second loss value has not yet converged.
[0134] Step S211, adjusting the parameters of the named entity recognition sub-model to be trained according to the determined second loss value.
[0135] If the second loss value has not converged, the model parameters of the named entity recognition sub-model are adjusted. After the model parameters are adjusted, the process returns to step S204 to continue the next round of training.
[0136] Step S212, ending the training to obtain the trained category prediction model.
[0137] If both the first loss value and the second loss value converge, the currently obtained category prediction model is used as the trained category prediction model.
[0138] In the embodiment of the present application, a training data set including encyclopedia text data samples, domain text data samples with domain category annotations, and named entity recognition text data samples with named entity category annotations is used to train the pre-trained language sub-model and the named entity recognition sub-model in the category prediction model, so that the model can learn more semantic information in the text data, and more accurately identify the named entities in the text data based on the semantic information in the text data, and classify the named entities into corresponding named entity categories. In addition, after the training of the category prediction model is completed, a basic dictionary can be constructed based on the existing training data set to complete the construction of the dictionary.
[0139] After obtaining the trained category prediction model, a basic dictionary can be constructed based on a training data set including encyclopedia text data samples, domain text data samples and named entity recognition text data samples. The basic dictionary contains multiple named entity categories, and each named entity category can contain multiple named entity words.
[0140] Specifically, after obtaining the embedding vector sample corresponding to the text data sample through the pre-trained language sub-model, the embedding vector sample is input into the named entity recognition sub-model. During the training of the model, the named entity recognition task of the named entity recognition sub-model mainly includes two sub-tasks, one is the recognition of the named entity boundary, and the other is the recognition of the named entity category. The BIO / BIOES annotation method can simultaneously annotate the boundary and category of a word. Based on this annotation method, the previous NER model processes the NER problem into a multi-label classification task, directly predicting the named entity position and category of each word. This processing method combines the two sub-tasks of boundary recognition and category prediction well, but there are still two major problems. First, this type of model performs well when there are fewer named entity categories, but it is difficult to accurately predict the entity category of each word when there are more categories; second, the semantic information of a single word can only be used in the model recognition process. For named entities composed of multiple words, the overall semantic information of multiple words cannot be used.
[0141] The named entity recognition sub-model in the embodiment of the present application is also trained on two sub-tasks, but the combination method is different from the previous model. First, the boundary annotation information is used to train the model to recognize the boundaries of the named entity, and the text is divided into different semantic units according to the prediction results. The words in the same unit are regarded as belonging to the same named entity, and their word vectors are averaged and then assigned to each word, thereby utilizing the semantic information of the named entity as a whole.
[0142] The training of the pre-trained language sub-model and the named entity recognition sub-model can be achieved in two ways, one is the fine-tune method, that is, all the parameters of the pre-trained language sub-model and the named entity recognition sub-model are uniformly adjusted; the other is the feature-base method, that is, the output of the pre-trained language sub-model is directly used as the embedding vector as the input of the named entity recognition sub-model, and the parameters of the pre-trained language sub-model are fixed, and only the named entity recognition sub-model needs to be trained. The training method adopted in the embodiment of the present application is the feature-base method. After the training process of the pre-trained language sub-model is completed, the weight of the pre-trained language sub-model is no longer involved in the training, which reduces the model training time and enables the named entity recognition sub-model to better combine the two sub-tasks of boundary recognition and category prediction.
[0143] In one embodiment, after the training data set is acquired, multiple text data samples in the training data set may be analyzed and processed. Figure 3 The specific process of analyzing and processing the training data in the training data set in the embodiment of the present application is shown, such as Figure 3 As shown, the process may include the following steps:
[0144] Step S301, obtaining a training data set including encyclopedia text data, domain text data and named entity recognition text data.
[0145] Among them, encyclopedia text data refers to text data without a targeted field, domain text data refers to selected text data that is highly relevant to a specific named entity category, or contains a large number of named entities of a certain category, and named entity recognition text data refers to text data that requires word-level named entity annotation.
[0146] Step S302, pre-processing operations are performed on the encyclopedia text data, the domain text data and the named entity recognition text data respectively.
[0147] After obtaining encyclopedia text data, domain text data and named entity recognition text data, preprocessing operations can be performed on the encyclopedia text data, domain text data and named entity recognition text data, respectively. Among them, the preprocessing operation may include at least one of data screening and format conversion. Specifically, after obtaining encyclopedia text data samples, domain text data samples and named entity recognition text data samples, data cleaning, denoising and serialization operations can be performed on each text data sample, respectively. Among them, data cleaning can include missing value cleaning, format content cleaning, logical error cleaning, non-demand data cleaning and correlation verification; data denoising can include outlier filling, distance-based detection, etc.; data serialization is to convert data into a standard format.
[0148] Step S303, obtaining corresponding encyclopedia text sequences, domain text sequences and named entity recognition text sequences.
[0149] After preprocessing the encyclopedia text data, domain text data and named entity recognition text data respectively, encyclopedia text sequences, domain text sequences and named entity recognition text sequences can be obtained accordingly.
[0150] In an embodiment of the present application, by performing data cleaning operations on the text data in the training data set, data that meets the data quality requirements can be obtained, and then data denoising and serialization operations are performed on the text data, and the text data can be converted into a standard format. Therefore, the category prediction model can be trained based on the text data in the standard format obtained after processing, which can improve the training efficiency and training results of the model.
[0151] After obtaining the trained category prediction model, a dictionary can be constructed based on the category prediction model. Figure 4a A flowchart of a dictionary construction method provided in an embodiment of the present application is shown. The method can be executed by a server, for example, Figure 1a The server 200 in FIG. Figure 4a , the dictionary construction process in the embodiment of the present application is elaborated in detail.
[0152] Step S40, obtaining the text to be processed and the basic dictionary.
[0153] The basic dictionary is a dictionary including multiple named entity categories obtained in the training phase of training the category prediction model. In addition, each named entity category in the basic dictionary contains multiple corresponding named entity words. For example, the basic dictionary may include multiple named entity categories such as names of people, places, and names of organizations. A name of a person may include multiple names, a place name may include multiple places, and an organization name may include multiple organizations.
[0154] In one embodiment, after the text to be processed is obtained, data cleaning, denoising and serialization operations may be performed on the text to be processed to obtain a text sequence corresponding to the text to be processed.
[0155] Step S41, based on the trained category prediction model, determining at least one candidate word contained in the text to be processed and the semantic information of the at least one candidate word.
[0156] Among them, the category prediction model may include a pre-trained language sub-model and a named entity recognition sub-model.
[0157] The text to be processed is input into the category prediction model, and based on the pre-trained language sub-model in the category prediction model, the word vector corresponding to at least one word contained in the text to be processed can be obtained. Each word vector can represent the semantic information of the corresponding word.
[0158] After obtaining the word vector corresponding to at least one word, based on the named entity recognition sub-model in the category prediction model, the at least one word can be combined to obtain at least one candidate word and the semantic information of at least one candidate word. Each candidate word can contain at least one word.
[0159] For example, Figure 4b As shown, the text to be processed is "Xiao Ming goes to Hangzhou for fun". Before inputting the text to be processed into the category prediction model, it is necessary to process the text to be processed into the structure of "[CLS]Xiao Ming goes to Hangzhou for fun[SEP]", and then input "[CLS]Xiao Ming goes to Hangzhou for fun[SEP]" into the category prediction model. Based on the pre-trained language sub-model, "E[CLS]E1 E2 E3 E4 E5 E6 E7 E[SEP]" can be obtained. Based on the obtained "E[CLS]E1 E2 E3 E4 E5 E6 E7 E[SEP]", "E[CLS]E1+E2 E1+E2 E3 E4+E5 E4+E5 E6 E7E[SEP]" can be obtained through the named entity recognition sub-model.
[0160] In an embodiment of the present application, based on a pre-trained language sub-model and a named entity recognition sub-model, at least one candidate word contained in the text to be processed and the semantic information of at least one candidate word can be obtained, so that the text to be processed can be divided into multiple candidate words according to the semantic information of the text to be processed. Then, when performing named entity recognition, the overall semantic information of multiple words in the text can be used to accurately identify the target words belonging to the named entities in the text to be processed.
[0161] Step S42, using the category prediction model, at least one target word that meets the set semantic conditions is selected according to the semantic information of at least one candidate word, and the probability value of the at least one target word belonging to each of the multiple word categories is determined.
[0162] After determining at least one candidate word contained in the text to be processed and the semantic information of at least one candidate word, at least one target word whose semantic information belongs to a named entity can be selected from the at least one candidate word through the named entity recognition sub-model. The named entity is an entity name with specific semantics.
[0163] For at least one target word, the following operations may be performed respectively: by using a named entity recognition sub-model, according to semantic information of a target word, the probability values of the target word belonging to a plurality of word categories are determined.
[0164] For example, Figure 4c As shown in the figure, assuming that the basic dictionary includes three named entity categories, namely, names of people, names of organizations and names of places, after obtaining the corresponding "E[CLS]E1+E2 E1+E2 E3 E4+E5E4+E5 E6 E7 E[SEP]" according to the to-be-processed text "Xiao Ming went to Hangzhou for a trip", it can be determined from the semantic information that the words corresponding to "E1+E2" and "E4+E5" are all target words with semantic information belonging to named entities. The probabilities of "E1+E2" and "E4+E5" belonging to names of people, names of organizations and names of places are predicted respectively, and the probability values P1, P2 and P3 of "E1+E2" belonging to names of people, names of organizations and names of places can be obtained, and the probability values P4, P5 and P6 of "E4+E5" belonging to names of people, names of organizations and names of places can be obtained respectively.
[0165] In an embodiment of the present application, based on the named entity recognition sub-model, target words belonging to named entities in the text to be processed can be identified, and the probability values of the target words belonging to each named entity category in the basic dictionary can be obtained, so that the named entities in the text to be processed can be accurately identified and the probability of the named entities relative to each named entity category can be obtained.
[0166] Step S43: at least one target word is respectively assigned to a word category whose corresponding probability value meets a set probability condition.
[0167] After determining the probability values that at least one target word belongs to multiple word categories, for each target word, based on the probability values that the target word belongs to multiple word categories, the word category that meets the set probability conditions can be selected as the target category to which the target word belongs, and the target word is added to the target category of the basic dictionary. For example, after determining the probability values P1, P2, and P3 that the named entity "Xiao Ming" in the to-be-processed text "Xiao Ming goes to Hangzhou for fun" belongs to the name of a person, the name of an organization, and the name of a place, respectively, and the probability values P4, P5, and P6 that "Hangzhou" belongs to the name of a person, the name of an organization, and the name of a place, respectively, assuming that the probability values P1 and P6 meet the set probability conditions, the name of a person can be used as the named entity category of the named entity "Xiao Ming", and the name of a place can be used as the named entity category of the named entity "Hangzhou", and "Xiao Ming" can be added to the name of a person in the basic dictionary, and "Hangzhou" can be added to the name of a place in the basic dictionary.
[0168] In the above step S43, for one of the at least one target words, the target category to which the target word belongs is determined, and the process of attributing the target word to the target category of the basic dictionary can be as follows: Figure 4d As shown, the following steps are included:
[0169] Step S431, based on the probability values of the target word belonging to multiple word categories, determine the word category corresponding to the maximum probability value, and use the word category as the target category.
[0170] Among them, the target word is a word belonging to a named entity contained in the text to be processed output by the category prediction model, and the word category is the named entity category contained in the basic dictionary. After the text to be processed is input into the category prediction model, the category prediction model can output the prediction result of at least one target word belonging to a named entity contained in the text to be processed, and the prediction result is the probability value of each target word belonging to each named entity category, forming a (category, probability) result pair.
[0171] After determining the probability value of each target word belonging to each named entity category, one of the target words can be taken as an example to explain in detail how to assign each target word to the corresponding named entity category in the basic dictionary.
[0172] Arrange the probability values of the target words belonging to each named entity category in descending order, that is, sort the probability values of the target words belonging to each named entity category in order from large to small, determine the maximum probability value, and take the named entity category corresponding to the maximum probability value as the target category.
[0173] Step S432: if the probability value that the target word belongs to the target category is greater than a first set threshold, the target word is assigned to the target category.
[0174] If the maximum probability value is greater than the first set threshold, the target word may be assigned to the target category, that is, the target word may be assigned to the target category in the basic dictionary, thereby completing the dictionary expansion.
[0175] For example, the target word is "Hangzhou", and the named entity categories included in the basic dictionary include names of people, organization names, and place names. The probability value of the target word "Hangzhou" belonging to a person name is 0.08, the probability value of the target word "Hangzhou" belonging to an organization name is 0.12, and the probability value of the target word "Hangzhou" belonging to a place name is 0.98, and the first set threshold is 0.8. The probability values of the target word belonging to a person name, an organization name, and a place name are arranged in descending order, and the maximum probability value can be determined to be 0.98, and the named entity category corresponding to the maximum probability value is a place name. Since the maximum probability value of 0.98 is greater than the first set threshold of 0.8, the target word "Hangzhou" can be classified as a place name in the basic dictionary.
[0176] In the embodiment of the present application, after obtaining the probability values of the target words belonging to each named entity category, based on the maximum probability value exceeding the set threshold, the target words can be assigned to the named entity category corresponding to the maximum probability value, thereby completing the dictionary expansion. Thus, the target words determined to belong to the named entity in the text can be classified into the named entity category with the highest probability value, thereby improving the accuracy of the dictionary construction.
[0177] Step S433: If the probability values of the target word belonging to multiple word categories are not greater than the second set threshold, at least one word category with a probability value greater than the third set threshold is selected as a candidate category.
[0178] The third set threshold is smaller than the second set threshold, and the second set threshold may be equal to the first set threshold or may not be equal to the first set threshold.
[0179] And when the second set threshold is equal to the first set threshold, it is equivalent to selecting at least one word category with a probability value greater than the third set threshold as a candidate category if the maximum probability value among the probability values of the target word belonging to multiple word categories is not greater than the first set threshold.
[0180] When the probability values of the target word belonging to multiple word categories are not greater than the second set threshold, the probability values of the target word belonging to each named entity category can be arranged in descending order, and at least one named entity category whose probability value is greater than the third set threshold is selected as a candidate category.
[0181] Step S434, based on the similarity between the target word and the words contained in at least one candidate category, a candidate category that meets the set similarity condition is selected as the target category, and the target word is classified into the target category.
[0182] After selecting at least one candidate category, the similarity between the target word and at least the words contained in each candidate category can be determined, and from the similarity, the candidate category with the highest similarity is selected as the target category, and the target word is assigned to the target category of the basic dictionary.
[0183] In the embodiment of the present application, after determining the probability values of the target words belonging to each named entity category, multiple candidate categories can be selected from the named entity category according to the probability values, and the target words are compared with the words contained in each candidate category for similarity, and then the target category is determined according to the similarity, and then the target word is assigned to the target category to complete the dictionary expansion. Therefore, based on the similarity judgment, the target words determined in the text belonging to the named entity can be divided into the target category corresponding to the words with the highest similarity to the target words, thereby improving the accuracy of the dictionary construction.
[0184] In one embodiment, for each candidate category, a WordNet dictionary tool can be used to determine the similarity between the target word and the existing word set contained in the candidate category, and after determining the similarity between the target word and the existing word set in each candidate category, the candidate category with the highest similarity is selected from each similarity as the target category, and the target word is assigned to the target category of the basic dictionary to complete the dictionary expansion.
[0185] For example, the target word is "New York", and the named entity categories included in the basic dictionary are person names, organization names, place names, holidays, product names, and time. The probability value of the target word "New York" belonging to a person name is 0.45, the probability value of the target word "New York" belonging to an organization name is 0.67, the probability value of the target word "New York" belonging to a place name is 0.8, the probability value of the target word "New York" belonging to a holiday is 0.32, the probability value of the target word "New York" belonging to a product name is 0.76, and the probability value of the target word "New York" belonging to a person name is 0.06. The second set threshold is 0.9, and the third set threshold is 0.5. Since the probability values of the target word "New York" belonging to a person name, an organization name, a place name, a holiday, a product name, and time are not greater than the second set threshold of 0.9, multiple named entity categories with probability values greater than the third set threshold of 0.5 can be selected as candidate categories, that is, the selected candidate categories are organization names, place names, and product names. Determine the similarity between the target word "New York" and the existing word sets contained in the organization name, place name and product name respectively. Assuming that the similarity between the target word "New York" and the existing word set contained in the organization name is 0.06, the similarity with the existing word set contained in the place name is 0.96, and the similarity with the existing word set contained in the product name is 0.21, then the place name can be used as the target category corresponding to the target word "New York", and the target word "New York" can be attributed to the place name in the basic dictionary.
[0186] In summary, according to the probability values of a target word belonging to each named entity category, the detailed process of attributing a target word to the corresponding target category in the basic dictionary can also be as follows: Figure 5 As shown, the following steps are included:
[0187] Step S501: determine the probability value of the target word belonging to each named entity category.
[0188] Step S502, sort the probability values in descending order and determine the maximum probability value.
[0189] Step S503, determine whether the maximum probability value is greater than a first set threshold; if not, execute step S504; if yes, execute step S506.
[0190] Step S504: select at least one named entity category whose probability value is greater than a third set threshold as a candidate category.
[0191] Step S505 , based on the similarity between the target word and the words respectively included in at least one candidate category, select the candidate category with the highest similarity as the target category, and assign the target word to the target category.
[0192] Step S506: taking the named entity category corresponding to the maximum probability value as the target category, and assigning the target word to the target category.
[0193] The dictionary construction method provided in the embodiment of the present application can train the category prediction model based on a training data set including encyclopedia text data, domain text data with domain category annotations, and named entity recognition text samples with named entity category annotations to obtain a trained category prediction model, and construct a basic dictionary including multiple named entity categories based on the training data set. After obtaining the trained category prediction model, the target words belonging to the named entity in the text to be processed and the probability values of the target words belonging to each named entity category in the basic dictionary are identified based on the category prediction model. After obtaining the probability value, based on the maximum probability value being greater than the set threshold, the target word is attributed to the named entity category corresponding to the maximum probability value to complete the expansion of the dictionary, or if the maximum probability value is not greater than the set threshold, multiple named entity categories are selected from them according to the probability value as candidate categories, and the similarity between the target word and the words contained in each candidate category is determined. Based on the similarity, the target word is attributed to the named entity category corresponding to the highest similarity to complete the expansion of the dictionary. This effectively avoids the reliance on complex rules designed by domain experts or a large amount of manual annotation work, effectively utilizes massive natural language texts and deep models to learn text semantic representation, and ultimately achieves efficient and high-quality dictionary construction and expansion. At the same time, it also makes it possible to judge new corpus information more flexible and efficient, and the dictionary matching speed is fast, which greatly improves the accuracy of dictionary construction and alleviates performance issues to a certain extent.
[0194] See also Figure 6 As shown, a specific application scenario is used below to further explain the above embodiment in detail:
[0195] Assume that the obtained text to be processed is "Mingming went to visit Huangshan", and the basic dictionary contains 6 named entity categories, namely, person name, organization name, place name, festival, product name and time.
[0196] After the text to be processed "Mingming went to visit Huangshan" is processed into "[CLS] Mingming went to visit Huangshan [SEP]", it is input into the pre-trained language sub-model of the category prediction model to obtain the corresponding embedding vector "E[CLS]E1, E2, E3, E4, E5, E6, E7, E[SEP]".
[0197] The embedding vector "E[CLS]E1, E2, E3, E4, E5, E6, E7, E[SEP]" is input into the named entity recognition sub-model of the category prediction model, and "E1+E2" and "E6+E7" are obtained as target words belonging to the named entity. At the same time, it can be obtained that the probability value of "E1+E2" belonging to a person's name is 0.98, the probability value of belonging to an organization name is 0.31, the probability value of belonging to a place name is 0.12, the probability value of belonging to a festival is 0.06, the probability value of belonging to a product name is 0.45, and the probability value of belonging to time is 0.03; the probability value of "E6+E7" belonging to a person's name is 0.68, the probability value of belonging to an organization name is 0.78, the probability value of belonging to a place name is 0.89, the probability value of belonging to a festival is 0.07, the probability value of belonging to a product name is 0.72, and the probability value of belonging to time is 0.04.
[0198] Assuming that the first set threshold is 0.95, the second set threshold is 0.9, and the third set threshold is 0.6, it can be determined that the named entity category of the target word "Mingming" corresponding to "E1+E2" is a human name, and the target word "Mingming" is added to the human names in the basic dictionary.
[0199] Since the probability values of the target word "Huangshan" in "E6+E7" that belong to the name of a person, the name of an organization, the name of a place, the name of a festival, the name of a product, and the name of time are all less than the first set threshold value of 0.95 and the second set threshold value of 0.9, and the named entity categories greater than the third set threshold value of 0.6 are the name of a person, the name of an organization, the name of a place, and the name of a product, the similarity between the target word "Huangshan" and the words contained in the name of a person, the name of an organization, the name of a place, and the name of a product can be determined, and the similarity between the target word "Huangshan" and the words contained in the name of a person is 0.06, the similarity with the words contained in the name of an organization is 0.14, the similarity with the words contained in the place name is 0.76, and the similarity with the words contained in the product name is 0.04. It can be determined that the named entity category to which the target word "Huangshan" belongs is a place name, and the target word "Huangshan" is added to the place names in the basic dictionary.
[0200] In some embodiments, the dictionary construction method proposed in this application can be compared with the N-Gram-based dictionary construction method, the word2Vec-based sentiment dictionary construction method, the textrank-based dictionary construction method, and the domain dictionary construction method for product evaluation object mining in the related art. The comparison results can be shown in Table 1 below:
[0201] Table 1
[0202] Model Precision Recall F1-score Based on N-Gram 40.0 — — Based on word2Vec 60.3 37.5 46.2 Based on TextRank 59.8 58.6 59.2 Product evaluation object mining 64.5 69.2 66.8 Our Method 85.6 86.2 85.9
[0203] Among them, Precision represents accuracy, Recall represents recall, and F1-score is a comprehensive evaluation indicator. The higher the Precision, Recall, and F1-score, the better the performance of the method.
[0204] As can be seen from the above table, compared with the dictionary construction method based on N-Gram, the sentiment dictionary construction method based on word2Vec, the dictionary construction method based on textrank, and the domain dictionary construction method for product evaluation object mining, the Precision, Recall and F1-score of the dictionary construction method proposed in this application are the highest, indicating that the dictionary construction method proposed in this application has the best performance.
[0205] and Figure 4a The dictionary construction method shown is based on the same inventive concept, and a dictionary construction device is also provided in the embodiment of the present application, which can be deployed in a server or a terminal device. Since the device is a device corresponding to the dictionary construction method of the present application, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation of the above method, and the repeated parts will not be repeated.
[0206] Figure 7 A schematic diagram of the structure of a dictionary construction device provided in an embodiment of the present application is shown. Figure 7 As shown, the dictionary construction device includes an acquisition module 701, a word recognition module 702, a category recognition module 703 and a dictionary construction module 704.
[0207] The acquisition module 701 is used to acquire the text to be processed and a basic dictionary; wherein the basic dictionary contains multiple word categories;
[0208] The word recognition module 702 is used to determine at least one candidate word contained in the text to be processed and the semantic information of the at least one candidate word based on the trained category prediction model;
[0209] The category identification module 703 is used to select at least one target word that meets the set semantic conditions according to the semantic information of at least one candidate word through the category prediction model, and determine the probability value of the at least one target word belonging to multiple word categories respectively;
[0210] The dictionary construction module 704 is used to classify at least one target word into a word category whose corresponding probability value meets a set probability condition.
[0211] Optionally, the category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the word recognition module 702 is specifically used to:
[0212] Based on the text to be processed, a word vector corresponding to at least one word contained in the text to be processed is obtained by pre-training the language sub-model; wherein each word vector represents the semantic information of the corresponding word;
[0213] Based on the word vector corresponding to each of the at least one word, at least one word is combined through a named entity recognition sub-model to obtain at least one candidate word and semantic information of at least one candidate word; wherein each candidate word contains at least one word.
[0214] Optionally, the category identification module 703 is specifically used for:
[0215] By using a named entity recognition sub-model, at least one target word whose semantic information belongs to a named entity is selected from at least one candidate word; a named entity is an entity name with specific semantics;
[0216] For at least one target word, the following operations are performed respectively: by using a named entity recognition sub-model, according to the semantic information of a target word, the probability values of the target word belonging to multiple word categories are determined.
[0217] Optional, such as Figure 8 As shown, the above device may further include a model training module 801, which is used to:
[0218] Obtain a training data set; the training data set includes a plurality of text data samples, and the text data samples are marked with set categories;
[0219] Based on the training data set, the category prediction model is iteratively trained until the set convergence condition is met. One iterative training process includes:
[0220] Based on the text data samples extracted from the training data set, at least one target word in the text data sample is determined through a category prediction model, and a target word category corresponding to each of the at least one target word is determined;
[0221] According to the target word category and the set category, the corresponding loss value is determined, and the parameters of the category prediction model are adjusted according to the loss value.
[0222] Optionally, the category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the training data set includes encyclopedia text data samples, domain text data samples and named entity recognition text data samples; the model training module 801 is also used for:
[0223] Based on the text data samples extracted from the encyclopedia text data samples and the domain text data samples, the corresponding embedding vector samples are determined by pre-training the language sub-model; the encyclopedia text data samples are text data without a targeted domain, and each domain text data sample is text data including multiple named entities of set categories and is annotated with corresponding domain categories;
[0224] Based on the vector samples extracted from the embedding vector samples and the word vector samples, at least one corresponding target word and its respective corresponding target word category are determined through the named entity recognition sub-model; the word vector samples are obtained based on the named entity recognition text data samples; the named entity recognition text data samples are text data including at least one named entity, and each named entity word is annotated with a corresponding named entity category.
[0225] Optionally, the model training module 801 is also used for:
[0226] Determine a first loss value according to the target word category and the field category, and adjust parameters of the pre-trained language sub-model according to the first loss value;
[0227] A second loss value is determined according to the target word category, the domain category, and the named entity category, and parameters of the named entity recognition sub-model are adjusted according to the second loss value.
[0228] Optionally, the model training module 801 is also used for:
[0229] Preprocessing operations are performed on encyclopedia text data samples, domain text data samples and named entity recognition text data samples in the training data set respectively; the preprocessing operations include at least one of data screening and format conversion.
[0230] Optionally, the dictionary construction module 704 is specifically used for:
[0231] Based on the probability values of a target word belonging to multiple word categories, determine the word category corresponding to the maximum probability value, and use the word category as the target category;
[0232] If the probability value of a target word belonging to the target category is greater than a first set threshold, the target word is assigned to the target category.
[0233] Optionally, the dictionary construction module 704 is further used to:
[0234] If the probability value of a target word belonging to multiple word categories is not greater than the second set threshold, at least one word category with a probability value greater than a third set threshold is selected as a candidate category; the third set threshold is less than the second set threshold;
[0235] Based on the similarity between a target word and words contained in at least one candidate category, a candidate category that meets the set similarity condition is selected as the target category, and a target word is assigned to the target category.
[0236] After introducing the dictionary construction method and apparatus according to the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.
[0237] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to as "circuit", "module" or "system" herein.
[0238] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application, referring to Fig. 9 As shown, it is a schematic diagram of a hardware structure of an electronic device using an embodiment of the present application, and the electronic device 900 may include at least a processor 901 and a memory 902. Among them, the memory 902 stores program codes, and when the program codes are executed by the processor 901, the processor 901 executes the steps of any of the above-mentioned dictionary construction methods.
[0239] In some possible implementations, the computing device according to the present application may include at least one processor and at least one memory. The memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the dictionary construction method according to various exemplary implementations of the present application described above in this specification. For example, the processor may execute the following steps: Figure 4a Follow the steps shown in .
[0240] Refer to the following Fig.10 The computing device 1000 according to this embodiment of the present application is described below. Fig.10 As shown, the computing device 1000 is in the form of a general computing device. The components of the computing device 1000 may include but are not limited to: at least one processing unit 1001, at least one storage unit 1002, and a bus 1003 connecting different system components (including the storage unit 1002 and the processing unit 1001).
[0241] Bus 1003 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a processor, or a local bus using any of a variety of bus architectures.
[0242] The storage unit 1002 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 10021 and / or a cache memory unit 10022 , and may further include a read-only memory (ROM) 10023 .
[0243] The storage unit 1002 may also include a program / utility 10025 having a set (at least one) of program modules 10024, such program modules 10024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include the implementation of a network environment.
[0244] The computing device 1000 may also communicate with one or more external devices 1004 (e.g., keyboards, pointing devices, etc.), may also communicate with one or more devices that enable a user to interact with the computing device 1000, and / or communicate with any device that enables the computing device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1005. In addition, the computing device 1000 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1006. As shown, the network adapter 1006 communicates with other modules for the computing device 1000 via a bus 1003. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0245] Based on the same inventive concept as the above method embodiment, various aspects of the dictionary construction method provided in the present application can also be implemented in the form of a program product, which includes a program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the dictionary construction method according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device can execute the following steps: Figure 4a Follow the steps shown in .
[0246] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0247] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0248] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A dictionary construction method, characterized in that: include: Obtaining a text to be processed and a basic dictionary; wherein the basic dictionary contains multiple word categories; Based on the trained category prediction model, determining at least one candidate word contained in the text to be processed, and semantic information of each of the at least one candidate word; By using the category prediction model, at least one target word that meets the set semantic condition is selected according to the semantic information of each of the at least one candidate word, and a probability value of each of the at least one target word belonging to the multiple word categories is determined; Attributing the at least one target word to a word category whose corresponding probability value meets a set probability condition; The attributing the at least one target word to a word category whose corresponding probability value meets a set probability condition includes: For the at least one target word, the following operations are performed respectively: If the probability value of a target word belonging to each of the plurality of word categories is not greater than a second set threshold, at least one word category having a probability value greater than a third set threshold is selected as a candidate category; the third set threshold is less than the second set threshold; Based on the similarity between the one target word and the words respectively included in the at least one candidate category, a candidate category that meets the set similarity condition is selected as the target category, and the one target word is assigned to the target category.
2. The method according to claim 1, characterized in that The category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the method of determining at least one candidate word contained in the text to be processed and the semantic information of each of the at least one candidate word based on the trained category prediction model includes: Based on the text to be processed, obtaining a word vector corresponding to each of at least one word contained in the text to be processed through the pre-trained language sub-model; wherein each word vector represents semantic information of the corresponding word; Based on the word vector corresponding to each of the at least one word, the at least one word is combined through the named entity recognition sub-model to obtain at least one candidate word and semantic information of each of the at least one candidate word; wherein each candidate word contains at least one word.
3. The method according to claim 2, characterized in that The step of selecting at least one target word that meets a set semantic condition based on the semantic information of each of the at least one candidate word through the category prediction model, and determining a probability value that each of the at least one target word belongs to the multiple word categories, includes: By using the named entity recognition sub-model, at least one target word whose semantic information belongs to a named entity is selected from the at least one candidate word; the named entity is an entity name with specific semantics; For the at least one target word, the following operations are performed respectively: by using the named entity recognition sub-model, according to the semantic information of the one target word, the probability values that the one target word belongs to the multiple word categories are determined.
4. The method according to any one of claims 1 to 3, characterized in that: The training process of the category prediction model includes: Acquire a training data set; the training data set includes a plurality of text data samples, and the text data samples are marked with set categories; Based on the training data set, the category prediction model is iteratively trained until a set convergence condition is met, wherein one iterative training process includes: Based on the text data sample extracted from the training data set, determining at least one target word in the text data sample through the category prediction model, and determining the target word category corresponding to each of the at least one target word; According to the target word category and the set category, a corresponding loss value is determined, and according to the loss value, parameters of the category prediction model are adjusted.
5. The method according to claim 4, characterized in that The category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the training data set includes encyclopedia text data samples, domain text data samples and named entity recognition text data samples; The step of determining at least one target word in the text data sample based on the text data sample extracted from the training data set and determining the target word category corresponding to each of the at least one target word by using the category prediction model comprises: Based on the text data samples extracted from the encyclopedia text data samples and the domain text data samples, the corresponding embedding vector samples are determined by the pre-trained language sub-model; the encyclopedia text data samples are text data without a targeted domain, and each domain text data sample is text data including a plurality of named entities of set categories and is annotated with a corresponding domain category; Based on the vector samples extracted from the embedding vector samples and the word vector samples, at least one corresponding target word and its respective corresponding target word category are determined through the named entity recognition sub-model; the word vector sample is obtained based on the named entity recognition text data sample; the named entity recognition text data sample is text data including at least one named entity, and each named entity word is annotated with a corresponding named entity category.
6. The method according to claim 5, characterized in that The step of determining a corresponding loss value according to the target word category and the set category, and adjusting parameters of the category prediction model according to the loss value, includes: Determining a first loss value according to the target word category and the field category, and adjusting parameters of the pre-trained language sub-model according to the first loss value; A second loss value is determined according to the target word category, the field category, and the named entity category, and parameters of the named entity recognition sub-model are adjusted according to the second loss value.
7. The method according to claim 5, characterized in that After obtaining the training data set, the method further includes: The encyclopedia text data samples, the domain text data samples and the named entity recognition text data samples in the training data set are respectively preprocessed; the preprocessing operation includes at least one of data screening and format conversion.
8. The method according to any one of claims 1 to 3, characterized in that: The attributing the at least one target word to a word category whose corresponding probability value meets the set probability condition, further includes: For the at least one target word, the following operations are performed respectively: Based on the probability values of a target word belonging to the plurality of word categories, determining the word category corresponding to the maximum probability value, and taking the word category as the target category; If the probability value of the target word belonging to the target category is greater than a first set threshold, the target word is assigned to the target category.
9. A dictionary construction device, characterized in that: include: An acquisition module, used to acquire the text to be processed and a basic dictionary; wherein the basic dictionary contains multiple word categories; A word recognition module, used to determine at least one candidate word contained in the text to be processed and the semantic information of each of the at least one candidate word based on the trained category prediction model; A category identification module, configured to select at least one target word that meets a set semantic condition based on the semantic information of each of the at least one candidate word through the category prediction model, and determine a probability value that each of the at least one target word belongs to each of the multiple word categories; A dictionary construction module, used to classify the at least one target word into word categories whose corresponding probability values meet the set probability conditions; The dictionary construction module is specifically used to: for the at least one target word, perform the following operations respectively: if the probability value of a target word belonging to the multiple word categories is not greater than a second set threshold, then select at least one word category with a probability value greater than a third set threshold as a candidate category; the third set threshold is less than the second set threshold; based on the similarity between the one target word and the words contained in the at least one candidate category, select a candidate category that meets the set similarity condition as the target category, and assign the one target word to the target category.
10. The device according to claim 9, characterized in that The category prediction model includes a pre-trained language sub-model and a named entity recognition sub-model; the word recognition module is specifically used to: Based on the text to be processed, obtaining a word vector corresponding to each of at least one word contained in the text to be processed through the pre-trained language sub-model; wherein each word vector represents semantic information of the corresponding word; Based on the word vector corresponding to each of the at least one word, the at least one word is combined through the named entity recognition sub-model to obtain at least one candidate word and semantic information of each of the at least one candidate word; wherein each candidate word contains at least one word.
11. The device according to claim 10, characterized in that The category identification module is specifically used for: By using the named entity recognition sub-model, at least one target word whose semantic information belongs to a named entity is selected from the at least one candidate word; the named entity is an entity name with specific semantics; For the at least one target word, the following operations are performed respectively: by using the named entity recognition sub-model, according to the semantic information of the one target word, the probability values that the one target word belongs to the multiple word categories are determined.
12. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of any one of the methods in claims 1 to 8.
13. A computer-readable storage medium, characterized in that: The method comprises a program code, and when the program code is run on an electronic device, the program code is used to enable the electronic device to execute the steps of any method described in claims 1 to 8.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of any method described in claims 1 to 8 are implemented.
Citation Information
Patent Citations
Semantic meaning-based specific task text keyword extraction method
CN107193803A
Legal named entity recognition method based on cascade model and data enhancement
CN113609857A