Natural language processing method, language model training method and related equipment

The corpus is processed through multiple feature extraction models and clustering models, and the language model is trained using reinforcement learning, which solves the problems of low training efficiency and high resource consumption in the existing technology, and achieves faster training time and lower resource consumption.

CN114781611BActive Publication Date: 2025-05-06华润数字科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210423601.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-05-06
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

In the prior art, language models based on deep neural networks face huge challenges in training and deployment, including long training time and large resource consumption.

Method used

By obtaining the corpus, using multiple feature extraction models (such as implicit features, topic features and entity features) to extract the corpus to obtain semantic vectors, and then using clustering models to divide semantic clusters, and each semantic cluster is trained separately for language model to finally determine the final language model.

Benefits of technology

By using the semantic correlation strength and weakness of the corpus, parallel training of semantic clusters is divided and reinforcement learning ideas are adopted, so that the language model can learn more and deeper language rules as soon as possible, shortening the training time and reducing the training overhead of the language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114781611B_ABST
    Figure CN114781611B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and discloses a natural language processing method, a language model training method, and related equipment. The language model training method includes: obtaining a corpus; extracting features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus; obtaining a semantic vector corresponding to each document based on the plurality of feature vectors corresponding to each document; clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain a plurality of semantic clusters; training the language model using reinforcement learning according to each semantic cluster, and finally obtaining the parameters of the trained language model corresponding to each semantic cluster; determining the final language model according to the parameters of the trained language model corresponding to each semantic cluster. The present application achieves the improvement of the training efficiency of the language model and reduces the resource consumption during the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a natural language processing method, a language model training method and related equipment. Background Art

[0002] At present, language models based on deep neural networks have achieved outstanding results in many fields such as semantic understanding and text generation. However, most of the current language models rely on supervised learning models supported by massive data. As a result, the capacity of the model is getting larger and larger, which brings huge challenges to training and deploying the model. One aspect is that the amount of training sample data often reaches TB level; and in the existing technology, the existing training corpus is often used directly, resulting in a significant increase in training time and high resource consumption. Therefore, how to improve the efficiency of model training, shorten the training time, and reduce resource consumption has become an urgent problem to be solved. Summary of the invention

[0003] The present application provides a natural language processing method, a language model training method and related equipment to solve the problems of low efficiency and high consumption of model training in the prior art.

[0004] In a first aspect, the present application provides a language model training method, comprising:

[0005] Get the corpus;

[0006] Extracting features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus;

[0007] Based on the multiple feature vectors corresponding to the documents, obtaining a semantic vector corresponding to each document;

[0008] Clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters;

[0009] According to each semantic cluster, the language model is trained using reinforcement learning, and finally the parameters of the trained language model corresponding to each semantic cluster are obtained;

[0010] The final language model is determined according to the parameters of the trained language model corresponding to each semantic cluster.

[0011] Furthermore, the multiple feature extraction models include an implicit feature extraction model, a topic feature extraction model, and an entity feature extraction model. The multiple feature extraction models are used to extract features from the corpus to obtain multiple feature vectors corresponding to each document in the corpus, including:

[0012] Performing implicit feature extraction on each of the documents in the corpus using the implicit feature extraction model to obtain a first feature vector corresponding to each of the documents;

[0013] Using the topic feature extraction model to extract topic features from each document in the corpus, to obtain a second feature vector corresponding to each document;

[0014] The entity feature extraction model is used to extract entity features from each document in the corpus to obtain a third feature vector corresponding to each document.

[0015] Furthermore, the subject feature extraction model is used to extract subject features from each document in the corpus to obtain a second feature vector corresponding to each document, including:

[0016] Extracting subject words from each of the documents in the corpus using the subject feature extraction model to obtain a plurality of subject words and arrange them;

[0017] The arranged multiple subject words are vectorized through the Bert model under the subject feature extraction model to obtain the second feature vector corresponding to each of the documents.

[0018] Furthermore, the extracting entity features of each document in the corpus using the entity feature extraction model to obtain a third feature vector corresponding to each document includes:

[0019] Identify entities in each of the documents and relationships between entities using named entity recognition technology and relationship extraction technology in an entity feature extraction model;

[0020] Based on the entities and the relationships between the entities, a knowledge graph is constructed;

[0021] The knowledge graph is subjected to feature extraction through a graph convolutional neural network in an entity feature extraction model to obtain a third feature vector.

[0022] Furthermore, obtaining the semantic vector corresponding to each document based on the plurality of feature vectors corresponding to each document includes:

[0023] Obtaining weights of the first eigenvector, the second eigenvector, and the third eigenvector based on the analytic hierarchy process;

[0024] According to the weights of the first feature vector, the second feature vector, and the third feature vector, a weighted sum is performed on the first feature vector, the second feature vector, and the third feature vector to obtain a semantic vector corresponding to the document.

[0025] Furthermore, the training of the language model using reinforcement learning according to each semantic cluster includes:

[0026] In each training cycle, when the performance index of the language model corresponding to a semantic cluster reaches a preset threshold, the state information of the language model at this time is obtained, and the state information of the language model is broadcast to the language models corresponding to each semantic cluster;

[0027] After receiving the state information, the language model corresponding to each semantic cluster updates its own parameters and selects a processing path according to a selection probability; wherein the selection probability is obtained by processing a plurality of semantic vectors used in the training cycle through a deep learning neural network;

[0028] Different benefits are given according to the processing path selected by the language model corresponding to each of the semantic clusters;

[0029] According to the income of each language model, the total income of this training cycle is obtained;

[0030] The deep learning neural network adjusts parameters according to the total benefit and is trained through multiple training cycles until the total benefit converges.

[0031] Furthermore, determining the final language model according to the parameters of the trained language model corresponding to each semantic cluster includes:

[0032] When all training cycles are completed, the final gradient data of the language model corresponding to each semantic cluster is aggregated to the trainer corresponding to the same language model;

[0033] The trainer performs average processing on the final gradient data corresponding to all language models to obtain an average gradient;

[0034] The average gradient is sent to the language model corresponding to each of the semantic clusters to update its own parameters to obtain the final language model.

[0035] In a second aspect, the present application further provides a natural language processing method, the method comprising:

[0036] Get the text data to be processed;

[0037] According to the final language model described above, the text data to be processed is processed to obtain a processing result corresponding to the text data to be processed.

[0038] In a third aspect, the present application further provides a language model training device, the device comprising:

[0039] The acquisition module is used to obtain the corpus;

[0040] A feature extraction module, used to extract features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus;

[0041] A merging module, used for obtaining a semantic vector corresponding to each of the documents based on the multiple feature vectors corresponding to each of the documents;

[0042] A clustering module, used for clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters;

[0043] A training module is used to train the language model using reinforcement learning according to each semantic cluster, and finally obtain the parameters of the trained language model corresponding to each semantic cluster;

[0044] The determination module is used to determine the final language model according to the parameters of the trained language model corresponding to each semantic cluster.

[0045] In a fourth aspect, the present application further provides a computer device, comprising:

[0046] at least one processor; and,

[0047] a memory communicatively connected to the at least one processor; wherein,

[0048] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the language model training method as described above.

[0049] In a fifth aspect, the present application also provides a non-volatile computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the language model training method as described above is implemented.

[0050] A natural language processing method, a language model training method and related devices provided by the embodiments of the present application have at least the following beneficial effects compared with the prior art:

[0051] By acquiring a corpus, using multiple feature extraction models to extract features from the corpus, multiple feature vectors corresponding to each document in the corpus are obtained, and multiple feature vectors corresponding to each document in the corpus are obtained, so as to realize multi-dimensional extraction of text features in the corpus; based on the multiple feature vectors corresponding to each document, a semantic vector corresponding to each document is obtained, and by combining multiple feature vectors, a corresponding semantic vector is obtained to realize integration of text features, and the semantic vectors corresponding to each document in the corpus are clustered using a clustering model to obtain multiple semantic clusters, and the language model is trained using reinforcement learning according to each semantic cluster, and the parameters of the trained language model corresponding to each semantic cluster are obtained, and the final language model is determined according to the parameters of the trained language model corresponding to each semantic cluster. By using the strength of the semantic association in the corpus, different semantic clusters are divided for parallel training, and the reinforcement learning idea is adopted, so that the language model can learn more and deeper language rules as early as possible, shortening the training time, thereby accelerating the convergence of the model and reducing the training cost of the language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the scheme in the present application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 A flow chart of a language model training method provided in one embodiment of the present application;

[0054] Figure 2 A flowchart of a language model training method provided in another embodiment of the present application;

[0055] Figure 3 for Figure 2 A flowchart of a specific implementation of step S220 in FIG.

[0056] Figure 4 for Figure 2 A flowchart of a specific implementation of step S230 in FIG.

[0057] Figure 5 for Figure 1 A flowchart of another specific implementation of step S5 in FIG.

[0058] Figure 6 A schematic diagram of a module of a language model training device provided in one embodiment of the present application;

[0059] Figure 7 A schematic diagram of the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of the present application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0061] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly or implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0062] This application provides a language model training method. Figure 1 As shown, Figure 1 A flowchart of a language model training method provided in one embodiment of the present application.

[0063] In this embodiment, the language model training method includes:

[0064] S1. Obtain corpus;

[0065] Specifically, the present application can directly obtain the corpus from the database, or connect with other systems to directly obtain the corpus from other systems. The corpus is a collection of document corpora used for language model training when the documents are labeled.

[0066] S2, extracting features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus;

[0067] Specifically, by using an implicit feature extraction model, a topic feature extraction model and an entity feature extraction model to extract features from the corpus, a first feature vector, a second feature vector and a third feature vector corresponding to each document in the corpus are obtained respectively.

[0068] Further, such as Figure 2As shown, the multiple feature extraction models include an implicit feature extraction model, a topic feature extraction model and an entity feature extraction model, and the multiple feature extraction models are used to extract features from the corpus to obtain multiple feature vectors corresponding to each document in the corpus, including:

[0069] S210, performing implicit feature extraction on each of the documents in the corpus using the implicit feature extraction model to obtain a first feature vector corresponding to each of the documents;

[0070] S220, extracting topic features from each document in the corpus using the topic feature extraction model to obtain a second feature vector corresponding to each document;

[0071] S230: Utilize the entity feature extraction model to extract entity features from each document in the corpus to obtain a third feature vector corresponding to each document.

[0072] Specifically, the implicit feature extraction model directly extracts the implicit features of the document to obtain a first feature vector, where the first feature vector is a document-level feature vector. The implicit feature extraction model can be trained based on a TextCNN model.

[0073] The topic feature extraction model extracts keywords from the document, sorts the keywords, and then vectorizes them to obtain a second feature vector, which is an overall vector of the document's topic words.

[0074] The entity feature extraction model extracts entities from each document to construct a knowledge graph, and then extracts feature vectors from the knowledge graph using a graph convolutional neural network to obtain the third feature vector.

[0075] By using a variety of feature extraction models to process documents, feature vectors of multiple dimensions can be obtained, which can better reflect the text features.

[0076] Furthermore, if Figure 3 As shown, the topic feature extraction model is used to extract topic features from each document in the corpus to obtain a second feature vector corresponding to each document, including:

[0077] S221, extracting subject words from each of the documents in the corpus using the subject feature extraction model to obtain a plurality of subject words and arrange them;

[0078] S222, vectorizing the arranged multiple subject words through the Bert model under the subject feature extraction model to obtain a second feature vector corresponding to each of the documents.

[0079] Specifically, only topic words are extracted from each document in the corpus through a topic feature extraction model, and the number of topic words is set according to needs. A plurality of topic words are concatenated and sorted to form a sequence, which is input into a trained Bert model. After conversion by the Bert model, the second feature vector is output.

[0080] By using the topic feature extraction model to extract key words and, after sorting, inputting them into the Bert model for vectorization processing, we can obtain the topic features of the document, so that the model can learn more inherent language rules earlier in subsequent training.

[0081] Furthermore, if Figure 4 As shown, the entity feature extraction model is used to extract entity features from each document in the corpus to obtain a third feature vector corresponding to each document, including:

[0082] S231, identifying entities in each of the documents and relationships between entities through named entity recognition technology and relationship extraction technology in an entity feature extraction model;

[0083] S232, constructing a knowledge graph based on the entities and the relationships between the entities;

[0084] S233. Perform feature extraction on the knowledge graph through the graph convolutional neural network in the entity feature extraction model to obtain a third feature vector.

[0085] Specifically, through named entity recognition technology, Bert-Bi_LSTM-CRF is used in this application to identify entities in each document, and the relationship between entities is extracted by relationship extraction technology. After obtaining the relationship between entities, TransE and its subsequent improvements can be used to calculate the embedding vector of the relationship between entities.

[0086] The knowledge graph is constructed by using triples such as the relationship between entities.

[0087] The knowledge graph is feature extracted by the graph convolutional neural network in the entity feature extraction model to obtain a third feature vector, wherein the number of layers of the graph convolutional neural network can be set as needed; for example, when there are n vertices in the knowledge graph, that is, n entities, the embedding vector dimension of each vertex is m, and the matrix X∈R is defined n×m , define the vector Among them, A is the adjacency matrix of the nodes in the knowledge graph, M is the in-degree matrix, pass L 0 =X to calculate, j represents the number of layers of the graph convolutional network, D ggA represents the data in the g-th row and g-th column of the in-degree matrix, that is, the data on the diagonal; ge represents the data in the gth row and eth column of the adjacency matrix, W0 is the weight matrix, σ is the activation function (Relu, Sigmoid, etc. can be used), and the vector of the last layer is the third eigenvector.

[0088] By extracting the relationships between entities in the document, we obtain a knowledge graph, and use a graph convolutional neural network to extract features from the knowledge graph, so that the model can learn more intrinsic language rules earlier in subsequent training.

[0089] S3, obtaining a semantic vector corresponding to each of the documents based on the multiple feature vectors corresponding to each of the documents;

[0090] Specifically, a weighted sum is performed based on the first feature vector, the second feature vector and the third feature vector corresponding to each document to obtain the semantic vector corresponding to each document.

[0091] Furthermore, obtaining the semantic vector corresponding to each document based on the plurality of feature vectors corresponding to each document includes:

[0092] Obtaining weights of the first eigenvector, the second eigenvector, and the third eigenvector based on the analytic hierarchy process;

[0093] According to the weights of the first feature vector, the second feature vector, and the third feature vector, a weighted sum is performed on the first feature vector, the second feature vector, and the third feature vector to obtain a semantic vector corresponding to the document.

[0094] Specifically, the weights of the first feature vector, the second feature vector, and the third feature vector are obtained by using the Analytic Hierarchy Process (AHP), which refers to a decision-making method that decomposes elements that are always related to decision-making into levels such as goals, criteria, and plans, and performs qualitative and quantitative analysis on this basis. According to the weights of the first feature vector, the second feature vector, and the third feature vector, the first feature vector, the second feature vector, and the third feature vector are weighted and summed to obtain the semantic vector corresponding to the document.

[0095] The weights are obtained based on the hierarchical analysis method, and based on the weights, the feature vectors are weighted summed to obtain the semantic vector corresponding to the document, thereby achieving complete extraction of document features, so that the model can learn more intrinsic language rules earlier during subsequent training.

[0096] S4, clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters;

[0097] Specifically, the semantic vectors corresponding to each document are clustered using a clustering model. In this application, since there may be a limit on the number of trainers in the future, the clustering model used in this application is a K-means clustering model, where K is the number of trainers. When there is no limit on the number of trainers, a mean shift clustering model can be used to perform processing, and clustering is performed based on the actual situation of the semantic vectors corresponding to each document.

[0098] The K-means clustering model is an iterative clustering analysis algorithm, which is divided into K groups in advance, and then randomly selects K objects as the initial cluster centers, and then calculates the distance between each object and each seed cluster center, and assigns each object to the cluster center closest to it. The cluster centers and the objects assigned to them represent a cluster.

[0099] The mean shift clustering model is a sliding window-based algorithm to find dense areas of data points. This is a centroid-based algorithm that updates the candidate points of the center point to the mean of the points in the sliding window to locate the center point of each group / class. Then, similar windows are removed from these candidate windows to finally form a center point set and corresponding grouping.

[0100] S5, training the language model using reinforcement learning according to each semantic cluster, and finally obtaining the parameters of the trained language model corresponding to each semantic cluster;

[0101] Specifically, the semantic vectors in each semantic cluster are used to train the language model respectively. That is, when there are multiple semantic clusters, multiple language models are trained at the same time, and the training method is to adopt reinforcement learning, so that the total benefit converges and is maximized, thereby finally obtaining the parameters of the trained language model corresponding to each semantic cluster.

[0102] Further, such as Figure 5 As shown, the training of the language model using reinforcement learning according to each semantic cluster includes:

[0103] S501, in each training cycle, when the performance index of a language model corresponding to a semantic cluster reaches a preset threshold, obtaining the state information of the language model at this time, and broadcasting the state information of the language model to the language models corresponding to each semantic cluster;

[0104] S502, after receiving the state information, the language model corresponding to each of the semantic clusters updates its own parameters and selects a processing path according to a selection probability; wherein the selection probability is obtained by processing a plurality of semantic vectors used in the training cycle through a deep learning neural network;

[0105] S503, providing different benefits according to the processing path selected by the language model corresponding to each semantic cluster;

[0106] S504, obtaining the total revenue of this training cycle according to the revenue of each language model;

[0107] S505. The deep learning neural network adjusts parameters according to the total benefit, and trains for multiple training cycles until the total benefit converges.

[0108] Specifically, a language model is trained for each semantic cluster, that is, each language model is trained using semantic vectors in different semantic clusters. In each training cycle, the performance index of each language model is constantly detected during the training process. When the performance index reaches a preset threshold, the state information of the language model at this time is obtained, and the state information includes the number of samples Ns used when the training of this language cluster ends and the gradient information at this time; the state information is broadcast to the language model corresponding to each semantic cluster;

[0109] Other language models have three processing paths to choose from after receiving status information:

[0110] 1) Immediately end the training of this cycle and record the performance index value of the team at this moment (defined as F1 o ) and the number of samples used Nt; calculate the average of the gradient sent by the other party and the gradient of the current side, update the loss function of the side, and get the new performance index value (defined as F1 N ); if F1 o Less than F1 N , then the income is given Otherwise, give profit

[0111] 2) Record the performance index value of your team at this moment (defined as F2 o ), calculate the average of the gradient sent by the other party and the gradient of the current party, update the loss function of the party, continue training until the end of the training cycle, and get the new performance index value (F2 N ); Assume that the number of samples used in a training cycle is Nb, if F2 o Less than F2 N , then the income is given Otherwise, give profit

[0112] 3) Record your own performance index value at this moment (F3 o ), calculate the average of the gradient sent by the other party and the gradient of the current party, and update the loss function of the party; based on the number of samples Nt used in the current training, retrain and randomly select samples of ΔN to end the training cycle and obtain the new performance index value (F3 N ); if F3o Smaller than F3 N , then the income is given Otherwise, give profit

[0113] Selecting a processing path according to a selection probability; wherein the selection probability is obtained by processing a plurality of semantic vectors used in the training cycle through a deep learning neural network; specifically including:

[0114] The policy gradient method in the field of reinforcement learning is used for optimization to maximize the total training benefit. The specific process is as follows. Train a multi-layer neural network corresponding to the above three actions. Taking a two-layer neural network as an example, the input vector v (that is, the semantic vector used in this training cycle), the hidden layer weight matrix is ​​set to w1, the relu activation function is used, the bias is b1, and the output is o1=relu(w1*v+b1); the second hidden layer weight matrix is ​​set to w2, the bias is b2, the output is o2=relu(w2*o1+b2), and then the softmax layer is used to obtain o3, o3 is the probability of taking a processing path each time. In actual use, more hidden layers can be used to obtain better results. Specific benefits can be obtained by selecting the processing path with the highest probability for processing;

[0115] After a training cycle, the benefits corresponding to each language model are summarized to obtain the total training benefit. The total training benefit is obtained according to the following formula:

[0116]

[0117] Where, γ is the profit decay coefficient, n is the number of training cycles, i = 1 to (n-1), S t is the total training benefit obtained in the tth training cycle.

[0118] The deep learning neural network is adjusted according to the total revenue of each training cycle, and multiple training cycles are trained until the total revenue converges. The optimization of the upgraded network can be optimized by SGD, Adam and other methods.

[0119] By utilizing the strength of the semantic associations between training corpora, the training samples are divided into different semantic clusters for parallel training, and reinforcement learning is used to enable the model to learn more intrinsic language rules as early as possible, thereby accelerating model convergence and reducing the training overhead of the language model.

[0120] In other embodiments of the present application, a training agent may be set for the language model corresponding to each semantic cluster, and the agent is responsible for the execution of the training algorithm program, resource application and communication with other agents during the entire training process.

[0121] S6. Determine a final language model according to the parameters of the trained language model corresponding to each semantic cluster.

[0122] The parameters of the trained language models corresponding to each semantic cluster, mainly gradient data, are sent to the same language model for aggregation and averaging to obtain the average gradient, which is then sent to the language models corresponding to each semantic cluster for parameter update. The updated language models are aggregated to obtain the final language model.

[0123] The final language model can complete tasks such as machine translation, part-of-speech tagging, syntactic analysis, and classification.

[0124] Furthermore, determining the final language model according to the parameters of the trained language model corresponding to each semantic cluster includes:

[0125] When all training cycles are completed, the final gradient data of the language model corresponding to each semantic cluster is aggregated to the trainer corresponding to the same language model;

[0126] The trainer performs average processing on the final gradient data corresponding to all language models to obtain an average gradient;

[0127] The average gradient is sent to the language model corresponding to each of the semantic clusters to update its own parameters to obtain the final language model.

[0128] Specifically, when all training cycles are completed, the final gradient data of the language models corresponding to each semantic cluster are summarized to the trainer corresponding to the same language model; specifically, they can be summarized to the trainer corresponding to the language model with the lowest workload at this moment, and the average gradient is calculated for all gradients. The average gradient is sent to the language models corresponding to each of the semantic clusters to update their own parameters, and the updated language models are summarized to obtain the final language model.

[0129] By calculating the average gradient at the end, each language model also uses the average gradient to update its own parameters, so that the language model is further optimized.

[0130] This application obtains a corpus, uses a variety of feature extraction models to extract features from the corpus, obtains multiple feature vectors corresponding to each document in the corpus, and obtains multiple feature vectors corresponding to each document in the corpus, so as to realize multi-dimensional extraction of text features in the corpus; based on the multiple feature vectors corresponding to each document, obtains the semantic vector corresponding to each document, and obtains the corresponding semantic vector by combining multiple feature vectors to realize the integration of text features, clusters the semantic vectors corresponding to each document in the corpus using a clustering model, obtains multiple semantic clusters, and trains the language model using reinforcement learning according to each semantic cluster, obtains the parameters of the trained language model corresponding to each semantic cluster, and determines the final language model according to the parameters of the trained language model corresponding to each semantic cluster. By utilizing the strength of the semantic association in the corpus, different semantic clusters are divided for parallel training, and the reinforcement learning idea is adopted, so that the language model can learn more and deeper language rules as soon as possible, shortening the training time, thereby accelerating the convergence of the model and reducing the training cost of the language model.

[0131] The present application also provides a natural language processing method, the method comprising:

[0132] Get the text data to be processed;

[0133] According to the final language model described above, the text data to be processed is processed to obtain a processing result corresponding to the text data to be processed.

[0134] Specifically, obtain the text data to be processed, and process the text data to be processed according to the above-mentioned trained final language model. Specifically, all models under the final language model can be used to process the text data to be processed, or the text data to be processed can be first classified to determine to which semantic cluster the text data to be processed belongs. Based on the semantic cluster corresponding to the text data to be processed, the corresponding language model under the final language model is used to process the text data to obtain the corresponding processing result.

[0135] Furthermore, the language model training method can learn more intrinsic language rules more quickly by performing corresponding training according to the labels of each text in the corpus, thereby accelerating model convergence and reducing the training overhead of the language model. Depending on the labels, the final language model can perform tasks such as machine translation, part-of-speech tagging, syntactic analysis and classification, and obtain corresponding processing results.

[0136] By adopting the final language model, the output processing results are better and faster.

[0137] This embodiment also provides a language model training device, such as Figure 6The figure shows a functional module diagram of the language model training device of the present application.

[0138] The language model training device 100 described in the present application can be installed in an electronic device. According to the functions implemented, the language model training device 100 may include an acquisition module 101, a feature extraction module 102, a merging module 103, a clustering module 104, a training module 105 and a determination module 106. The module described in the present application may also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.

[0139] In this embodiment, the functions of each module / unit are as follows:

[0140] An acquisition module 101 is used to acquire a corpus;

[0141] A feature extraction module 102 is used to extract features from the corpus using multiple feature extraction models to obtain multiple feature vectors corresponding to each document in the corpus;

[0142] Further, the multiple feature extraction models include an implicit feature extraction model, a topic feature extraction model and an entity feature extraction model, and the feature extraction module 102 includes a first extraction submodule, a second extraction submodule and a third extraction submodule;

[0143] The first extraction submodule is used to extract implicit features from each of the documents in the corpus through the implicit feature extraction model to obtain a first feature vector corresponding to each of the documents;

[0144] The second extraction submodule is used to extract topic features from each document in the corpus using the topic feature extraction model to obtain a second feature vector corresponding to each document;

[0145] The third extraction submodule is used to extract entity features from each document in the corpus using the entity feature extraction model to obtain a third feature vector corresponding to each document.

[0146] By cooperating with the first extraction submodule, the second extraction submodule and the third extraction submodule and using a variety of feature extraction models to process the document, feature vectors of multiple dimensions can be obtained, which can better reflect the text features.

[0147] Furthermore, the second extraction submodule also includes a topic extraction unit and a vectorization unit;

[0148] The topic extraction unit is used to extract topic words from each of the documents in the corpus through the topic feature extraction model to obtain multiple topic words and arrange them;

[0149] The vectorization unit is used to vectorize the arranged multiple subject words through the Bert model under the subject feature extraction model to obtain the second feature vector corresponding to each of the documents.

[0150] Through the cooperation of the topic extraction unit and the vectorization unit, the topic feature extraction model is used to extract topic words. After sorting, they are input into the Bert model for vectorization processing to obtain the topic features of the document, which makes it easier for the model to learn more intrinsic language rules earlier in subsequent training.

[0151] Furthermore, the third extraction submodule also includes an entity extraction unit, a construction unit, and a graph convolution extraction unit;

[0152] The entity extraction unit is used to identify entities in each of the documents and relationships between entities through named entity recognition technology and relationship extraction technology in the entity feature extraction model;

[0153] The construction unit is used to construct a knowledge graph based on the entities and the relationships between the entities;

[0154] The graph convolution extraction unit is used to extract features from the knowledge graph through a graph convolutional neural network in an entity feature extraction model to obtain a third feature vector.

[0155] Through the cooperation of the entity extraction unit, the construction unit, and the graph convolution extraction unit, the entities and the relationships between entities in the document are extracted to obtain the knowledge graph, and the graph convolution neural network is used to extract features of the knowledge graph, so that the model can learn more intrinsic language rules earlier in subsequent training.

[0156] A merging module 103, configured to obtain a semantic vector corresponding to each of the documents based on the multiple feature vectors corresponding to each of the documents;

[0157] Furthermore, the merging module 103 includes a weight acquisition submodule and a weighted summation submodule;

[0158] The weight acquisition submodule is used to obtain the weights of the first eigenvector, the second eigenvector, and the third eigenvector based on the hierarchical analysis method;

[0159] The weighted summation submodule is used to perform weighted summation on the first feature vector, the second feature vector, and the third feature vector according to the weights of the first feature vector, the second feature vector, and the third feature vector to obtain the semantic vector corresponding to the document.

[0160] Through the cooperation of the weight acquisition submodule and the weighted summation submodule, the weights are obtained based on the hierarchical analysis method, and based on the weights, the feature vectors are weighted summed to obtain the semantic vector corresponding to the document, thereby achieving complete extraction of document features, so that the model can learn more intrinsic language rules earlier in subsequent training.

[0161] A clustering module 104 is used to cluster the semantic vectors corresponding to each document in the corpus using a clustering model to obtain a plurality of semantic clusters;

[0162] The training module 105 is used to train the language model using reinforcement learning according to each semantic cluster, and finally obtain the parameters of the trained language model corresponding to each semantic cluster;

[0163] Furthermore, the training module 105 includes a broadcast submodule, a path selection submodule, a corresponding processing submodule, a revenue calculation submodule and a parameter adjustment submodule;

[0164] The broadcast submodule is used to obtain the state information of the language model at this time when the performance index of the language model corresponding to a semantic cluster reaches a preset threshold in each training cycle, and broadcast the state information of the language model to the language models corresponding to each semantic cluster;

[0165] The path selection submodule is used for the language model corresponding to each semantic cluster to update its own parameters after receiving the state information, and select a processing path according to a selection probability; wherein the selection probability is obtained by processing a plurality of semantic vectors used in the training cycle through a deep learning neural network;

[0166] The corresponding processing submodule is used to provide different benefits according to the processing path selected by the language model corresponding to each of the semantic clusters;

[0167] The profit calculation submodule is used to obtain the total profit of this training cycle according to the profit of each language model;

[0168] The parameter adjustment submodule is used to adjust the parameters of the deep learning neural network according to the total benefit, and trains for multiple training cycles until the total benefit converges.

[0169] Through the cooperation of the broadcast submodule, path selection submodule, corresponding processing submodule, benefit calculation submodule and parameter adjustment submodule, the training samples are divided into different semantic clusters for parallel training by utilizing the strength of the semantic association between the training corpora. Reinforcement learning is adopted to enable the model to learn more intrinsic language rules as early as possible, thereby accelerating model convergence and reducing the training overhead of the language model.

[0170] The determination module 106 is used to determine the final language model according to the parameters of the trained language model corresponding to each semantic cluster.

[0171] Further, the determination module 106 includes a summarization submodule, an averaging submodule and a sending submodule;

[0172] The aggregation submodule is used to aggregate the final gradient data of the language model corresponding to each semantic cluster to the trainer corresponding to the same language model after all training cycles are completed;

[0173] The averaging submodule is used for the trainer to perform averaging processing on the final gradient data corresponding to all language models to obtain an average gradient;

[0174] The sending submodule is used to send the average gradient to the language model corresponding to each semantic cluster to update its own parameters to obtain the final language model.

[0175] By cooperating with the summary submodule, the average submodule and the sending submodule, the average gradient is finally calculated. Each language model also uses the average gradient to update its own parameters, so that the language model is further optimized.

[0176] By adopting the above-mentioned device, the language model training device 100 uses the acquisition module 101, the feature extraction module 102, the merging module 103, the clustering module 104, the training module 105 and the determination module 106 in coordination, obtains a corpus, uses multiple feature extraction models to extract features from the corpus, and obtains multiple feature vectors corresponding to each document in the corpus, thereby realizing multi-dimensional extraction of text features in the corpus; based on the multiple feature vectors corresponding to each document, a semantic vector corresponding to each document is obtained, and the corresponding semantic vector is obtained by combining the multiple feature vectors to realize integration of text features, and the semantic vectors corresponding to each document in the corpus are clustered using a clustering model to obtain multiple semantic clusters, and the language model is trained using reinforcement learning according to each semantic cluster to obtain the parameters of the trained language model corresponding to each semantic cluster, and the final language model is determined according to the parameters of the trained language model corresponding to each semantic cluster. By utilizing the strength of the semantic associations in the corpus, dividing the corpus into different semantic clusters for parallel training, and adopting reinforcement learning ideas, the language model can learn more and deeper language rules as early as possible, shortening the training time, thereby accelerating model convergence and reducing the training overhead of the language model.

[0177] This embodiment also provides a natural language processing device, and the natural language processing device described in this application can be installed in an electronic device. According to the functions implemented, the natural language processing device may include a data acquisition module and a processing module. The module described in this application may also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.

[0178] In this embodiment, the functions of each module / unit are as follows:

[0179] The data acquisition module is used to acquire the text data to be processed;

[0180] The processing module is used to process the text data to be processed according to the final language model as described above, and obtain a processing result corresponding to the text data to be processed.

[0181] Through the cooperation of the data acquisition module and the processing module, the final language model is used to make the output processing result better and faster.

[0182] The present application also provides a computer device. Figure 7 , Figure 7 This is a basic structural block diagram of the computer device in this embodiment.

[0183] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0184] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0185] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the language model training method, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0186] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer-readable instructions stored in the memory 41 or process data, such as computer-readable instructions for running the language model training method.

[0187] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0188] This embodiment implements the steps of the language model training method of the above embodiment when the processor executes the computer-readable instructions stored in the memory, obtains a corpus, extracts features from the corpus using multiple feature extraction models, obtains multiple feature vectors corresponding to each document in the corpus, and obtains multiple feature vectors corresponding to each document in the corpus, so as to realize multi-dimensional extraction of text features in the corpus; based on the multiple feature vectors corresponding to each document, obtains the semantic vector corresponding to each document, and obtains the corresponding semantic vector by combining multiple feature vectors to realize integration of text features, clusters the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters, and trains the language model using reinforcement learning according to each semantic cluster, obtains the parameters of the trained language model corresponding to each semantic cluster, and determines the final language model according to the parameters of the trained language model corresponding to each semantic cluster. By utilizing the strength of the semantic association in the corpus, different semantic clusters are divided for parallel training, and the reinforcement learning idea is adopted, so that the language model can learn more and deeper language rules as early as possible, shortening the training time, thereby accelerating the convergence of the model and reducing the training cost of the language model.

[0189] An embodiment of the present application also provides a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the language model training method as described above, by acquiring a corpus, using multiple feature extraction models to extract features from the corpus, and obtaining multiple feature vectors corresponding to each document in the corpus, thereby realizing multi-dimensional extraction of text features in the corpus; based on the multiple feature vectors corresponding to each of the documents, a semantic vector corresponding to each of the documents is obtained, and the corresponding semantic vector is obtained by combining the multiple feature vectors to realize integration of text features, and the semantic vectors corresponding to each document in the corpus are clustered using a clustering model to obtain multiple semantic clusters, and the language model is trained using reinforcement learning according to each semantic cluster to obtain the parameters of the trained language model corresponding to each semantic cluster, and the final language model is determined according to the parameters of the trained language model corresponding to each semantic cluster. By utilizing the strength of the semantic associations in the corpus, dividing the corpus into different semantic clusters for parallel training, and adopting reinforcement learning ideas, the language model can learn more and deeper language rules as early as possible, shortening the training time, thereby accelerating model convergence and reducing the training overhead of the language model.

[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0191] The language model training apparatus, computer device, and computer-readable storage medium of the above-mentioned embodiments of the present application have the same technical effects as the language model training method of the above-mentioned embodiments, and will not be elaborated here.

[0192] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A language model training method, characterized in that: The method comprises: Get the corpus; Extracting features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus; Based on the multiple feature vectors corresponding to the documents, obtaining a semantic vector corresponding to each document; Clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters; According to each semantic cluster, the language model is trained using reinforcement learning, and finally the parameters of the trained language model corresponding to each semantic cluster are obtained; Determine the final language model according to the parameters of the trained language model corresponding to each semantic cluster; The multiple feature extraction models include an implicit feature extraction model, a topic feature extraction model, and an entity feature extraction model. The multiple feature extraction models are used to extract features from the corpus to obtain multiple feature vectors corresponding to each document in the corpus, including: Performing implicit feature extraction on each of the documents in the corpus using the implicit feature extraction model to obtain a first feature vector corresponding to each of the documents; Using the topic feature extraction model to extract topic features from each document in the corpus, to obtain a second feature vector corresponding to each document; Using the entity feature extraction model to extract entity features from each document in the corpus, to obtain a third feature vector corresponding to each document; The obtaining, based on the plurality of feature vectors corresponding to the documents, a semantic vector corresponding to each document comprises: A weighted sum is performed on the first feature vector, the second feature vector and the third feature vector corresponding to each of the documents to obtain a semantic vector corresponding to each of the documents.

2. The language model training method according to claim 1, characterized in that: The subject feature extraction model is used to extract subject features from each document in the corpus to obtain a second feature vector corresponding to each document, which includes: Extracting subject words from each of the documents in the corpus using the subject feature extraction model to obtain a plurality of subject words and arrange them; The arranged multiple subject words are vectorized through the Bert model under the subject feature extraction model to obtain the second feature vector corresponding to each of the documents.

3. The language model training method according to claim 1, characterized in that: The extracting entity features of each document in the corpus using the entity feature extraction model to obtain a third feature vector corresponding to each document includes: Identify entities in each of the documents and relationships between entities using named entity recognition technology and relationship extraction technology in an entity feature extraction model; Based on the entities and the relationships between the entities, a knowledge graph is constructed; The knowledge graph is subjected to feature extraction through a graph convolutional neural network in an entity feature extraction model to obtain a third feature vector.

4. The language model training method according to claim 1, characterized in that: The obtaining, based on the plurality of feature vectors corresponding to the documents, a semantic vector corresponding to each document comprises: Obtaining weights of the first eigenvector, the second eigenvector, and the third eigenvector based on the analytic hierarchy process; According to the weights of the first feature vector, the second feature vector, and the third feature vector, a weighted sum is performed on the first feature vector, the second feature vector, and the third feature vector to obtain a semantic vector corresponding to the document.

5. The language model training method according to claim 1, characterized in that: The training of the language model by using reinforcement learning according to each semantic cluster includes: In each training cycle, when the performance index of the language model corresponding to a semantic cluster reaches a preset threshold, the state information of the language model at this time is obtained, and the state information of the language model is broadcast to the language models corresponding to each semantic cluster; After receiving the state information, the language model corresponding to each semantic cluster updates its own parameters and selects a processing path according to a selection probability; wherein the selection probability is obtained by processing a plurality of semantic vectors used in the training cycle through a deep learning neural network; Different benefits are given according to the processing path selected by the language model corresponding to each of the semantic clusters; According to the income of each language model, the total income of this training cycle is obtained; The deep learning neural network adjusts parameters according to the total benefit and is trained through multiple training cycles until the total benefit converges.

6. The language model training method according to claim 1, characterized in that: Determining the final language model according to the parameters of the trained language model corresponding to each semantic cluster includes: When all training cycles are completed, the final gradient data of the language model corresponding to each semantic cluster is aggregated to the trainer corresponding to the same language model; The trainer performs average processing on the final gradient data corresponding to all language models to obtain an average gradient; The average gradient is sent to the language model corresponding to each of the semantic clusters to update its own parameters to obtain the final language model.

7. A natural language processing method, characterized in that: The method comprises: Get the text data to be processed; According to the final language model as claimed in any one of claims 1 to 6, the text data to be processed is processed to obtain a processing result corresponding to the text data to be processed.

8. A language model training device, characterized in that: The device comprises: The acquisition module is used to obtain the corpus; A feature extraction module, used to extract features from the corpus using a variety of feature extraction models to obtain a plurality of feature vectors corresponding to each document in the corpus; A merging module, used for obtaining a semantic vector corresponding to each of the documents based on the multiple feature vectors corresponding to each of the documents; A clustering module, used for clustering the semantic vectors corresponding to each document in the corpus using a clustering model to obtain multiple semantic clusters; A training module is used to train the language model using reinforcement learning according to each semantic cluster, and finally obtain the parameters of the trained language model corresponding to each semantic cluster; A determination module, used to determine a final language model according to the parameters of the trained language model corresponding to each semantic cluster; Wherein, the multiple feature extraction models include an implicit feature extraction model, a topic feature extraction model and an entity feature extraction model, and the feature extraction module includes a first extraction submodule, a second extraction submodule and a third extraction submodule; The first extraction submodule is used to extract implicit features from each of the documents in the corpus through the implicit feature extraction model to obtain a first feature vector corresponding to each of the documents; The second extraction submodule is used to extract topic features from each document in the corpus using the topic feature extraction model to obtain a second feature vector corresponding to each document; The third extraction submodule is used to extract entity features from each document in the corpus using the entity feature extraction model to obtain a third feature vector corresponding to each document; The merging module is further used to perform weighted summation on the first feature vector, the second feature vector and the third feature vector corresponding to each of the documents to obtain a semantic vector corresponding to each of the documents.

9. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the language model training method as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the language model training method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Semantic recognition method and device, electronic equipment and computer readable storage medium

    CN111125331A

  • Text similarity calculation method and device and computer equipment

    CN113987117A

  • Method and system of creating and summarizing unstructured natural language sentence clusters for efficient tagging

    US20200394364A1