An Industry Text Matching Model Method and Device Based on Deep Learning
Through the industry text matching model based on deep learning, large-scale cross-industry data and multiple Chinese-style pre-trained models, the problem of low accuracy of the automatic question-and-answer system in semantic matching is solved, and the effect of semantic precise matching in different industries is achieved.
Patent Information
- Application Number
- CN202111369472.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-15
AI Technical Summary
The existing automatic question-and-answer system has problems such as low accuracy in semantic matching and inability to accurately identify question and rhetorical questions, resulting in user input requests not being accurately replied, affecting user's learning and work efficiency.
The industry text matching model method based on deep learning is adopted. By introducing large-scale cross-industry data as the training set and integrating multiple pre-trained models with Chinese characteristics (such as NEZHA, RoBERTa and ERNIE-Gram), the model encoding method, training method and output results are optimized to improve the accuracy of semantic matching.
It realizes semantic precise matching in different sub-industry, improves the semantic matching accuracy of the automatic question-and-answer system, and can accurately understand and reply to user requests without the need for professional training data in the industry.
Smart Images

Figure CN114282592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic matching, and particularly to a method and device for an industry text matching model based on deep learning. Background Art
[0002] With the rapid development of Internet technology and the wide use of intelligent interaction applications, convenience, rapidity, and instantaneity have become the main features of the current Internet society. In the current fast-paced study and work, various social tools and question-and-answer systems have become increasingly relied-upon "necessities" for people. These constantly innovating intelligent interaction applications and various social tools all have a common communication bridge with humans, including text, voice, images, etc. Among them, the most stable and mainstream is still text communication, and mainly short-text communication. How to quickly and accurately grasp the language characteristics of different groups is a very concerned issue for current major intelligent interaction applications, that is, how the system can quickly and accurately understand the meaning of the input text of different groups. Only on this basis can it make a correct response. When using a search engine to input a question, the similar questions automatically matched by the system will be automatically presented to the user.
[0003] For example, when people use the Internet for information search, trading, shopping and other learning and living activities, for the same standard question recognized by the system, different users will have different ways of expressing it. In order to adapt to the quickness and convenience in the Internet ecosystem, many online shopping platforms have launched automatic reply systems such as intelligent customer service. By allowing users to select options similar to their questions or performing similar matching based on the user's input and then replying, but such systems have obvious deficiencies such as limited reply range and inability to accurately identify interrogative sentences and rhetorical questions. For the requests input by users, the content automatically replied by the system according to the recognized similar sentences is not what the user really wants in many cases, because many systems cannot understand the user's questions from the semantic perspective, and thus cannot accurately judge the needs of the user for the question-and-answer system, thereby reducing the learning and work efficiency of some users and causing unnecessary trouble to the development of society and people's daily lives. Therefore, it is particularly important to further improve the semantic matching function of the automatic question-and-answer system and increase the accuracy of semantic matching.
[0004] In addition, not only for emerging Internet industries, but also more and more traditional industries, such as physical industries like healthcare, power, banking, transportation, etc., are committed to developing Q&A systems for their respective industries. According to social surveys, most of the problems encountered by different people in the above public places or platforms are similar. At the same time, these problems are characterized by a high repetition rate of being asked, diverse ways of expression but consistent answers. If a traditional manual service desk is adopted, it is easy to encounter service saturation and an inability to meet customer needs during peak user periods. Therefore, in the face of the current scenario of explosive data growth and customers' real-time data requirements, traditional customer service teams mainly relying on manual labor urgently need the support of an automatic Q&A system.
[0005] In the future, the intelligent Q&A systems in these industries will not only be active in people's actual offline lives, but their core semantic matching technology will also play its role in various aspects such as the search systems, knowledge base query systems, and intelligent online customer service in these industries.
[0006] To solve the above existing problems and improve the accuracy of semantic matching in the Q&A system, so that customers can quickly and accurately search for the information they need from a huge amount of information, the intelligent automatic Q&A system based on short text semantic similarity matching has developed rapidly, and this method is an important pillar for realizing the intelligent Q&A system. According to deep learning algorithms, it can fully understand the semantic information of customers. At the same time, the intelligent Q&A system provides a way for natural language communication between humans and machines. It can, on the premise of quickly and accurately analyzing and understanding customer needs, provide customers with correct answers, especially for conventional short text questions with a relatively high repetition rate, which has the characteristics of high efficiency, convenience, and speed.
[0007] For a system serving tens of millions or even hundreds of millions of users, especially with the emergence and wide application of deep learning, the language information processing process can be transformed from the traditional vector space of words to the vector space of the word Embedding layer, or even more complex neural network hidden layer spaces. This way well makes up for the shortcomings of short texts in the word vector space, such as sparsity and high noise, and can seamlessly combine the unsupervised learning and supervised learning processes, opening up a new direction for natural language processing based on Q&A systems.
[0008] The current question-and-answer systems are mainly divided into question analysis and answer matching. The research conducted in this paper focuses on the question analysis part. Based on the original keyword matching of questions, the extraction and matching functions of semantic features are enriched. At present, there is still a large room for development in the semantic similarity analysis of Chinese questions. This paper conducts research from two aspects. Firstly, it strengthens the extraction of semantic information features of questions; secondly, it improves the semantic similarity matching results of two questions. Conducting research from these two aspects has, to a certain extent, solved such problems existing in the question-and-answer system, further improved the system function, and enhanced the user experience. Generally speaking, it is a meaningful research.
[0009] Through research, it is found that for the current English-based automatic question-and-answer systems, the related technologies have been relatively mature. Because English grammar is relatively simple and word segmentation is also easier, while Chinese is a semantic language. It is difficult for computers to understand semantics and analyze syntax based on machine language, which has led to the slow progress of the research on Chinese question-and-answer systems. Traditional Chinese question-and-answer systems only consider the literal meaning of sentences and do not conduct deeper mining of the actual semantics, thus easily deviating from the correct answer.
[0010] The domestic research on semantic similarity matching algorithms suitable for Chinese texts mainly includes the following categories:
[0011] Text similarity matching based on knowledge base: mainly based on the reference semantic dictionary How Net, it can be divided into three types of similarity calculations: sense, sememe, and word. Among them, the sememe is defined as the smallest unit in the dictionary, and it mainly calculates the text similarity according to the distance of words in the How Net at the sememe level. The sense is determined according to the word, and each word can have one or more senses, and each sense is composed of one or more sememes. Therefore, the similarity calculation of senses can be equivalent to the similarity calculation between sememes. Sort according to the combination of all senses and take the maximum value as the similarity calculation result of the word.
[0012] Text similarity matching based on deep learning: Domestic researchers such as Chen improved the matching method of word vector pairs by integrating a single string vector into the Word2vec and proposed a Chinese character feature enhancement model based on Chinese. Yu et al. proposed a joint learning word embedding model based on the splitting method, splitting each individual Chinese character into multiple independent fonts composed of radicals, and then fusing the independent font vectors and word vectors.
[0013] Semantic Similarity Matching Based on BERT Pre-trained Model: Domestic researchers Wu Yan et al. proposed a Chinese semantic matching algorithm based on the BERT model (Bidirectional Encoder Representations from Transformers). This algorithm converts sentences into feature vector representations, combines the Attention mechanism, and calculates the semantic similarity of two sentences for matching. Through comparative experiments with traditional semantic matching models such as BiLSTM (Bi-directional Long Short-Term Memory), ESIM (Enhanced Sequential Inference Model), and BiMPM (Bilateral Multi-Perspective Matching), the experimental results of the Chinese semantic matching algorithm based on BERT are better than those of the above semantic matching models on the test set.
[0014] For the text similarity matching algorithm based on the knowledge base, this method is too dependent on the corpus, and both require context for vectorization description. If there are too many repeated sentences in the corpus, problems such as excessive computational complexity and overly sparse calculation results may occur. For the text similarity matching algorithm based on deep learning, this method also requires a large amount of professional data for network training and cannot achieve good industry transferability. In addition, due to other characteristics such as the depth of the network layer and the structure design, the traditional deep neural network still performs poorly in the ability to understand semantics. For the semantic similarity matching algorithm based on the BERT pre-trained model, although this algorithm can better represent context information by using the BERT model to replace the commonly used Word2vec model for sentence vector representation, due to the fact that the BERT model does not consider the characteristics of Chinese corpus enough in its design and the existing training data is also insufficient, there is still much room for improvement in its effect. Summary of the Invention
[0015] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0016] To this end, an object of the present invention is to propose a method for an industry text matching model based on deep learning. By introducing large-scale cross-industry data as the training set and integrating and applying the advantages of multiple pre-trained models with Chinese characteristics, the present invention can solve semantic matching problems in various application fields such as automotive production line technology reference in the manufacturing industry, patient consultation in the medical industry, and transaction search in the commercial field. More importantly, the finally applied model can achieve accurate semantic matching tasks for different industries without the need for professional training data within the industry.
[0017] Another object of the present invention is to provide an industry text matching model device based on deep learning.
[0018] To achieve the above object, on the one hand, the present invention provides an industry text matching model method based on deep learning, including the following steps: obtaining a preset number of cross-industry data as a training set to obtain sentences to be matched; inputting the sentences to be matched into an industry text matching model NERB based on deep learning, and respectively inputting them into optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram after data preprocessing; wherein, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and optimizer; based on the optimized pre-trained model, three text matching results are output after matching by the optimized pre-trained model; according to the comprehensive judgment of the three text matching results, when any two text matching results or all three text matching results are output as similar, the output result of the industry text matching model is judged as similar, otherwise it is judged as dissimilar.
[0019] The industry text matching model method based on deep learning according to the embodiments of the present invention obtains a preset number of cross-industry data as a training set to obtain sentences to be matched; inputs the sentences to be matched into an industry text matching model NERB based on deep learning, and respectively inputs them into optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram after data preprocessing; wherein, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and optimizer; based on the optimized pre-trained model, three text matching results are output after matching by the optimized pre-trained model; according to the comprehensive judgment of the three text matching results, when any two text matching results or all three text matching results are output as similar, the output result of the industry text matching model is judged as similar, otherwise it is judged as dissimilar. By introducing large-scale cross-industry data as a training set and integrating and applying the advantages of multiple pre-trained models with Chinese characteristics, the present invention can solve semantic matching problems in various application fields such as automotive production line technology reference in the manufacturing industry, patient consultation in the medical industry, and transaction search in the commercial field. More importantly, the finally applied model can achieve accurate semantic matching tasks for different industries without the need for industry-specific professional training data.
[0020] In addition, the industry text matching model method based on deep learning according to the above embodiments of the present invention may further have the following additional technical features:
[0021] Further, in an embodiment of the present invention, the optimization of the functional relative position encoding includes: the pre-trained model NEZHA outputs a sine function related to relative positions involved in the calculation of attention scores by adopting functional relative position encoding. The formula for functional relative position encoding is as follows:
[0022]
[0023] Further, in an embodiment of the present invention, the optimization of the full word coverage includes: the pre-trained model NEZHA adopts a full word coverage strategy. When a Chinese character is covered, other Chinese characters belonging to the same Chinese character are covered together.
[0024] Further, in an embodiment of the present invention, the optimization of the mixed precision training includes: the pre-trained model NEZHA adopts mixed precision training. In each training iteration, the main weights are rounded to the half-precision floating-point format, and the forward and backward passes are performed using the weights, activations, and gradients stored in the half-precision floating-point format; the gradients are converted to the single-precision floating-point format, and the main weights are updated using the single-precision floating-point format gradients.
[0025] Further, in an embodiment of the present invention, the optimization of the optimizer includes: the pre-trained model NEZHA adopts the LAMB optimizer, and the adaptive strategy adjusts the learning rate for each parameter in the LAMB optimizer.
[0026] Further, in an embodiment of the present invention, the optimized pre-trained model RoBERTa includes:
[0027] Multiple model parameter quantities and training data; pre-adjusting the optimizer hyperparameters; the pre-trained model RoBERTa selects a preset number of training sample numbers; removing the next sentence prediction task, and the data is continuously obtained from one document; using dynamic masking, obtaining multiple copies of data by copying a training sample, each copy of data using a different mask, and increasing the copying fraction, and generating a new mask pattern each time a sequence is input to the pre-trained model RoBERTa; using full word masking.
[0028] Further, in an embodiment of the present invention, the optimized pre-trained model RoBERTa further includes:
[0029] Text encoding, the pre-trained model RoBERTa is trained using a BPE vocabulary of a preset level of bytes during the text encoding process, and no additional preprocessing or tokenization is performed on the input.
[0030] Further, in an embodiment of the present invention, the optimized pre-trained model ERNIE-Gram includes:
[0031] The ERNIE-Gram model learns n-gram granularity language information through an explicit n-gram masked language model by explicitly introducing language granularity knowledge. Based on the explicit n-gram masked language model, the pre-trained model ERNIE-Gram performs multi-level n-gram language granularity masked learning.
[0032] Further, in an embodiment of the present invention, the method further includes: validating the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram, including:
[0033] For the industry text matching model NERB, when the output results of any two or more pre-trained models are "similar", the output result of the industry text matching model NERB is judged as "similar"; otherwise, it is judged as "dissimilar". Then the accuracy rate of the industry text matching model NERB is:
[0034] P = p1 * p2 * (1 - p3) + p1 * p3 * (1 - p2) + p2 * p3 * (1 - p1) + p1 * p2 * p3
[0035] = p1 * p2 + p1 * p3 + p2 * p3 – 2 * p1 * p2 * p3
[0036] Where p1, p2, and p3 are the accuracy rates of the three semantic matching models NEZHA, RoBERTa, and ERNIE-Gram of the pre-trained model during semantic matching, respectively.
[0037] If the three semantic matching models can correctly judge whether the second preset number of samples match in a dataset containing the first preset number of samples, and the remaining third preset number of samples that cannot be correctly judged are sorted and in a continuous subsequence.
[0038] To achieve the above object, on the other hand, the present invention proposes an industry text matching model device based on deep learning, including: an acquisition module for acquiring a preset number of cross-industry data as a training set to obtain sentences to be matched; a training module for inputting the sentences to be matched into an industry text matching model NERB based on deep learning, and respectively inputting them into optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram after data preprocessing; wherein, the optimized pre-trained model NEZHA includes optimizations on functional relative position encoding, full-word coverage, mixed-precision training, and optimizer; an output module for outputting three text matching results based on the optimized pre-trained models after being matched by the optimized pre-trained models; a judgment module for making a comprehensive judgment according to the three text matching results. When any two text matching results or all three text matching results are output as similar, the output result of the industry text matching model is judged as similar, otherwise it is judged as dissimilar.
[0039] The industry text matching model device based on deep learning according to the embodiment of the present invention can solve semantic matching problems in various application fields such as automotive production line technology reference in the manufacturing industry, patient consultation in the medical industry, and transaction search in the commercial field by introducing a large amount of cross-industry data as a training set and integrating and applying the advantages of multiple pre-trained models with Chinese characteristics. More importantly, the finally applied model can achieve accurate semantic matching tasks for different industries without the need for industry-specific professional training data.
[0040] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Description of the Drawings
[0041] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, wherein:
[0042] Figure 1 It is a flowchart of an industry text matching model method based on deep learning according to an embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of the number and length distribution of sentences in the training set according to an embodiment of the present invention;
[0044] Figure 3 It is a schematic diagram of the number and length distribution of sentences in the validation set according to an embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of a semantic modeling example of BERT and ERNIE models according to an embodiment of the present invention;
[0046] Figure 5 Schematic diagram of continuous n-gram masked language model vs explicit n-gram masked language model according to an embodiment of the present invention;
[0047] Figure 6 Schematic diagram of n-gram multi-level language granularity masked learning according to an embodiment of the present invention;
[0048] Figure 7 Schematic diagram of semantic matching comprehensive model according to an embodiment of the present invention;
[0049] Figure 8 Schematic diagram of an extreme performance of three models in the same sample set according to an embodiment of the present invention;
[0050] Figure 9 Schematic diagram of another extreme performance of three models in the same sample set according to an embodiment of the present invention;
[0051] Figure 10 Schematic diagram of the change of loss rate during training of different pre-trained models according to an embodiment of the present invention;
[0052] Figure 11 Schematic diagram of the change of accuracy rate during training of different pre-trained models according to an embodiment of the present invention;
[0053] Figure 12 Schematic diagram of the device structure of the industry text matching model based on deep learning according to an embodiment of the present invention. Detailed implementation manners
[0054] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0055] The method and device for the industry text matching model based on deep learning according to an embodiment of the present invention will be described below with reference to the drawings. First, the method for the industry text matching model based on deep learning according to an embodiment of the present invention will be described with reference to the drawings.
[0056] Figure 1 Flowchart of the method for the industry text matching model based on deep learning according to an embodiment of the present invention.
[0057] As Figure 1 shown, the method for the industry text matching model based on deep learning includes the following steps:
[0058] Step S1: Obtain a preset number of cross-industry data as a training set to obtain the sentences to be matched.
[0059] It can be understood that in the present invention, a preset number of cross-industry data is obtained as a training set, which can be set by those skilled in the art according to actual needs, and the present invention does not make any limitations.
[0060] Specifically, for various application fields such as the automotive production line technology reference in the manufacturing industry, patient consultation in the medical industry, and transaction search in the commercial field, by introducing a large amount of cross-industry data as a training set and integrating the advantages of multiple pre-trained models with Chinese characteristics, semantic precise matching for different industries can be finally achieved without the need for industry-specific training data.
[0061] It can be understood that LCQMC (A Large-scale Chinese Question Matching Corpus) is a Chinese question matching data set in the Baidu Zhidao field. This data set is extracted and constructed from user questions in different fields of Baidu Zhidao. Its appearance solves the problem of the lack of large-scale question matching data sets in the Chinese field.
[0062] The samples in the data set all appear in the form of sentence pairs. There are also tags "0" and "1" indicating whether these two sentences are similar after each sample in the training set. Here, "0" represents dissimilar, and "1" represents similar, that is, semantic matching. After statistics, it is found that the number of samples (number of sentence pairs) in the training set reaches 238,766 pairs, there are 8,802 samples in the validation set, and the number of samples in the test set also reaches 12,500 pairs. Such a large-scale semantic matching data set also makes a basic foundation for the good performance of the entire model in cross-industry sentence matching later.
[0063] In order to select more appropriate input parameters, such as "sentence length", when designing the model, the present application statistically analyzed the lengths of the samples in the training set and the validation set, and counted the number distribution of the sentences in the training set at different lengths as Figure 2 shown. Similarly, the number and length distribution of the sentences in the validation set can be obtained as Figure 3 shown. From this, we can see that the distribution of the sample lengths in the training set and the validation set is basically the same, and the lengths of most samples are between 5 and 15, which basically conforms to the habitual input lengths of people in various interactive applications.
[0064] Step S2: Input the sentence to be matched into the industry text matching model NERB based on deep learning. After data preprocessing, it is respectively input into the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram. Among them, the optimized pre-trained model NEZHA includes the optimization of functional relative position encoding, full-word coverage, mixed-precision training, and the optimizer.
[0065] Specifically, the introduction and related improvements of the pre-trained model NEZHA (Ne Zha).
[0066] NEZHA (Ne Zha) implies "omnipotent and can solve different tasks". On the model NEZHA, parallel training based on multiple GPUs and multiple machines is implemented, and the training process is optimized to improve training efficiency. Finally, the pre-trained model NEZHA for multiple Chinese NLP tasks is obtained. This invention mainly introduces its other four improvements, namely: functional relative position encoding, full-word coverage, mixed-precision training, and improved optimizer.
[0067] Functional position encoding; there are two types of position encoding: functional and parametric. Functional position encoding can be directly calculated by defining a function. In parametric position encoding, position encoding involves two concepts: distance and dimension. Word Embedding generally has several hundred dimensions, and each dimension has a value. The value of a position encoding is determined by two parameters: position and dimension.
[0068] Compared with the absolute position encoding of Transformer, the pre-trained model NEZHA adopts functional relative position encoding. Its output and the calculation of attention scores involve the sine function of relative positions, which solves a series of resource occupancy problems caused by the fact that words in Transformer do not know the distance between each other.
[0069] In the NEZHA model, both distance and dimension are derived from the sine function and are fixed during model training. That is to say, each dimension of the position encoding corresponds to a sine, and the sine functions of different dimensions have different wavelengths. Choosing a fixed sine function can make the model have stronger scalability; that is, when it encounters a sequence longer than the sequence length in training, it can still play a role. The formula for functional relative position encoding is as follows:
[0070]
[0071] Full Word Masking; Research on "Pre-Training with Whole Word Masking for Chinese BERT" shows that replacing randomly masked words with whole word masking can effectively improve the performance of pre-trained models. That is, if one Chinese character is masked, all other Chinese characters belonging to the same Chinese word are masked together. The NEZHA pre-trained model adopts the Whole Word Masking (WWM) strategy. When one Chinese character is masked, all other Chinese characters belonging to the same Chinese word are masked together. This strategy has been proven to be more effective than the random masking training in BERT (i.e., each symbol or Chinese character is randomly masked).
[0072] In the WWM implementation of NEZHA, the researchers used a tokenization tool, Jieba2, for Chinese word segmentation (i.e., finding the boundaries of Chinese words). In the WWM training data, each sample contains multiple masked Chinese characters, and the total number of masked Chinese characters accounts for about 12% of its length, while the randomly replaced ones account for 1.5%. Although this increases the computational difficulty of predicting the whole word, the final effect is better.
[0073] Mixed Precision Training; Traditional deep neural network training uses FP32 (i.e., single-precision floating-point format) to represent all variables involved in training (including model parameters and gradients); while mixed precision training uses multiple precisions in training. Specifically, it focuses on ensuring a single-precision copy of the weights in the model (referred to as the main weights). That is, in each training iteration, the main weights are rounded to FP16 (i.e., half-precision floating-point format), and the forward and backward passes are performed using the weights, activations, and gradients stored in FP16 format; finally, the gradients are converted to FP32 format, and the main weights are updated using the FP32 gradients.
[0074] The NEZHA model adopts the mixed precision training technique in pre-training. This technique can increase the training speed by 2 - 3 times and also reduce the memory consumption of the model, thus enabling the use of larger batch sizes.
[0075] Improved Optimizer (LAMB Optimizer); Usually, when the Batch Size in deep neural network training is very large (exceeding a certain threshold), it will have a negative impact on the generalization ability of the model. The LAMB optimizer adopts a general adaptive strategy to adjust the learning rate for each parameter, enabling the model to maintain its performance when the Batch Size is very large, allowing the model training to use a very large Batch Size, and thus greatly improving the training speed.
[0076] The experiment tests the performance of the pre-trained model by fine-tuning on various natural language understanding (NLU) tasks, and compares the NEZHA model with other top Chinese pre-trained language models including Google BERT (Chinese version), BERT-WWM, and ERNIE (the detailed parameters can be found in the paper). The final results are shown in Table 1 as follows:
[0077] Table 1 NEZHA Experiment Results
[0078]
[0079] It can be seen that NEZHA achieved relatively better performance in most cases; especially in the PD-NER task, NEZHA reached a score of 97.87 at its highest. Another model with relatively outstanding performance is ERNIE Baidu 2.0, which shows a trend of surpassing NEZHA. Regarding this situation, the author in the paper also explained that due to possible differences in experimental settings or fine-tuning methods, the comparison may not be completely fair. After the new versions of other models are released, they will evaluate them under the same settings and update this report.
[0080] Introduction and Related Improvements of the Pre-trained Model RoBERTa
[0081] As an improved version of the pre-trained model BERT - RoBERTa, compared with BERT, RoBERTa mainly has the following improvements in terms of model scale, computing power, data, and training methods:
[0082] Larger model parameters (trained on Cloud TPU v3-256 for 24 hours, which is equivalent to training for one month on TPU v3-8 (128G video memory)).
[0083] Larger quantity and more diverse training data. It is trained using 30G of Chinese data, including 300 million sentences and 10 billion words (i.e., tokens). It includes news, microblogs, community discussions, online books, and multiple encyclopedias, etc., covering hundreds of thousands of topics.
[0084] Adjusted hyperparameters such as the optimizer.
[0085] Larger batch size. RoBERTa uses a larger batch size during training. Batch sizes ranging from 256 to 8000 have been tried, and the RoBERTa in this application uses a batch size of 8000.
[0086] Removed the next sentence prediction (NSP) task, and the data is continuously obtained from one document.
[0087] Dynamic masking. The original BERT obtains a static mask by performing masking once during data preprocessing. RoBERTa, on the other hand, duplicates a training sample to obtain multiple copies of data, each with a different mask, and increases the duplication fraction, so that a new mask pattern is generated each time a sequence is input to the model. During the continuous input of a large amount of data, the model gradually adapts to different masking strategies and learns different language representations, that is, the dynamic masking effect is achieved.
[0088] Whole word masking is used. In whole word masking, if a partial WordPiece subword of a complete word is masked, the other parts belonging to the same word will also be masked, that is, whole word masking.
[0089] Text encoding. Byte-Pair Encoding (BPE) is a hybrid of character-level and word-level representations, which supports handling many common words in natural language corpora. The original BERT implementation uses a character-level BPE vocabulary of size 30K, which is learned after preprocessing the input using heuristic tokenization rules. RoBERTa is trained using a larger byte-level BPE vocabulary of 50K subword units and does not perform any additional preprocessing or tokenization on the input.
[0090] When conducting the baseline test for Chinese Simplified reading comprehension, the dataset used is the Chinese Machine Reading Comprehension Data - CMRC 2018 released by the Joint Laboratory of Harbin Institute of Technology and iFlytek. The task is that given a question, the system needs to extract a fragment from the passage as the answer, in the same form as SQuAD. The test results are shown in Table 2.
[0091] Table 2 Performance of different models on CMRC 2018
[0092] Model Development Set Test Set Challenge Set BERT 65.5(64.4) / 84.5(84.0) 70.0(68.7) / 87.0(86.3) 18.6(17.0) / 43.3(41.3) ERNIE 65.4(643) / 84.7(84.2) 69.4(68.2) / 86.6(86.1) 19.6(17.0) / 44.3(42.8) BERT-wwm 66.3(65.0) / 85.6(84.7) 70.5(69.1) / 87.4(86.7) 21.0(193) / 47.0(439) BERT-wwm-ext 67.1(65.6) / 85.7(85.0) 71.4(70.0) / 87.7(87.0) 24.0(20.0) / 47.3(44.6) RoBERTa-wwm-ext 67.4(66.5) / 87.2(86.5) 72.6(71.4) / 89.4(88.8) 26.2(24.6) / 51.0(49.1)
[0093] Introduction and related improvements of the pre-trained model ERNIE-Gram
[0094] In recent years, deep neural network pre-trained models for unsupervised text have significantly improved the performance of various NLP tasks. Compared with the early focus on context-independent word vector modeling work, later proposed models such as Cove, ELMo, and GPT have constructed sentence-level semantic representations. In particular, the BERT model proposed by Google has achieved better semantic representation effects by predicting masked words or Chinese characters and utilizing the multi-layer self-attention bidirectional modeling ability of Transformer.
[0095] Although the BERT model has stronger semantic representation capabilities, its modeling object is also mainly focused on the original language signal, and less on the use of semantic knowledge units for modeling. As we all know, many of the Chinese expression systems are based on semantic knowledge units such as words. Therefore, the application of these pre-trained models in the Chinese field exposes a very obvious problem. For example, when BERT processes Chinese tasks, it models by predicting a Chinese character. At this time, it is difficult for the model to learn word-level semantic units, which affects the cognitive ability of complete semantic representation. For example, for the words Heilongjiang, Badminton, and Qidouyan, the BERT model can easily infer the masked word information through the collocation of words, but does not explicitly model and recognize the semantic concept units (such as Heilongjiang, Badminton, and Zhengqidouyan) and their corresponding semantic relationships.
[0096] From this point of view, if we can use the potential knowledge contained in massive texts to let the model learn word-level semantic units, it will undoubtedly further improve the results of various NLP tasks. Therefore, Baidu proposed the ERNIE model based on knowledge enhancement. This model can learn semantic relationships in the Chinese context by modeling prior semantic knowledge such as entity concepts in massive data. Specifically in the design method of the model, the ERNIE model masks semantic units such as words and entities, allowing the model to learn the semantic representation of complete concepts. Therefore, it can be said that compared with the original language signal learned by BERT, ERNIE directly models prior semantic knowledge units, enhancing the semantic representation ability of the model.
[0097] Here is an example of studying the sentence "Diaoyu Islands and its affiliated islands are an inseparable part of China."
[0098] Learned by BERT: Diaoyu Islands are China's inherent territory.
[0099] Learned by ERNIE: The Diaoyu Islands are inherent to [mask][mask] [mask].
[0100] like Figure 4 As shown in the BERT model on the left, the word "fish" can be identified by the local co-occurrence of "fish" and "island", but the model has not learned the knowledge related to "Diaoyu Island". Figure 4 The ERNIE on the right side learns the expression of words and entities, enabling the model to model the relationship between "Diaoyu Islands" and "China", and then learn that "Diaoyu Islands" is the inherent territory of "China".
[0101] It can be seen that, compared with BERT, the ERNIE model can model the combined semantics of words and has stronger generality and scalability. For example, when modeling words representing colors such as red, yellow, and purple, ERNIE can learn the semantic relationships between different words through different semantic combinations of the same characters.
[0102] In addition, ERNIE has made an extended improvement in the training corpus and introduced multi-source data knowledge. In addition to modeling encyclopedia articles, it also models news and forum dialogue data. Since there are often cases where the Query semantics corresponding to the same reply are similar in the dialogue context, the learning and modeling of dialogue data have become an important way to improve the semantic representation ability. Based on this hypothesis, ERINE uses DLM (Dialogue Language Model) to model the Query-Response dialogue structure, takes the dialogue Pair as the input, introduces the DialogueEmbedding to identify the roles of the dialogue, and uses the Dialogue Response Loss to learn the implicit relationship of the dialogue. Through this method of modeling, the semantic representation ability of the model is further improved.
[0103] To sum up, ERNIE has made improvements in the learning of entity concept knowledge and the expansion of the training corpus, enhancing the semantic representation ability of the model in the Chinese context. To verify the knowledge learning ability of ERNIE, the researchers investigated the model through a variety of tasks, including semantic similarity tasks, sentiment analysis tasks, named entity recognition tasks, and retrieval-based question-answering matching tasks, etc. Here, we take the performance of the BERT and ERNIE models on the semantic similarity task LCQMC as an example. Among them, LCQMC is a question semantic matching dataset constructed by Harbin Institute of Technology at the international top conference COLING2018 in natural language processing, and its goal is to judge whether the semantics of two questions are the same. The experimental results are shown in Table 3. It can be seen that the performance of the ERNIE model has been significantly improved both in the development set and in the test set.
[0104] Table 3 Performance of BERT and ERNIE on LCQMC
[0105]
[0106] At the Deep Learning Developer Summit WAVE SUMMIT held on May 20, 2021, in response to the existing difficulties and pain points of current pre-trained models, Baidu Wenxin ERNIE open-sourced the latest pre-trained model: the multi-granularity language knowledge enhanced model ERNIE-Gram.
[0107] Since the birth of the ERNIE model, Baidu researchers have introduced knowledge into the pre-trained model and improved the knowledge expression ability of the semantic model through knowledge enhancement methods. The ERNIE-Gram model precisely enhances the model's performance by explicitly introducing language granularity knowledge. Specifically, ERNIE-Gram proposes an explicit n-gram masked language model to learn language information at the n-gram granularity, such as Figure 5 As shown, the relative continuous n-gram masked language model significantly reduces the semantic learning space (V^n→V_(n-gram), where V is the vocabulary size and n is the length of the modeled gram), and significantly improves the convergence speed of the pre-trained model.
[0108] In addition, based on the explicit n-gram semantic granularity modeling, ERNIE-Gram proposes multi-level n-gram language granularity masked learning. The specific structure is as Figure 6 shown. Using the two-stream self-attention mechanism, it realizes the simultaneous learning of fine-grained semantic knowledge within the n-gram language unit and coarse-grained semantic knowledge between n-gram language units, achieving multi-level language granularity knowledge learning.
[0109] Without increasing any computational complexity, ERNIE-Gram significantly outperforms the mainstream open-source pre-trained models in multiple typical Chinese tasks such as natural language inference tasks, short text similarity tasks, and reading comprehension tasks. In addition, the ERNIE-Gram English pre-trained model also outperforms the mainstream models in general language understanding tasks and reading comprehension tasks.
[0110] Step S3, based on the optimized pre-trained model, three text matching results are output after being matched by the optimized pre-trained model.
[0111] Step S4, make a comprehensive judgment according to the three text matching results. When any two or all three text matching results are output as similar, the output result of the industry text matching model is judged as similar; otherwise, it is judged as dissimilar.
[0112] It can be understood that the present invention designs a comprehensive matching model - NERB. The overall structure of the model is as Figure 7As shown, after the sentence pairs to be matched are input into the model, after preprocessing operations such as word segmentation, they enter different improved versions of the BERT pre-trained model - NEZHA, RoBERTa, and ERNIE-Gram respectively, giving full play to the different advantages of the three in semantic expression ability. After being matched by different pre-trained models, three results are output. Next, the three output results are integrated. When any two or all three results are output as "similar", the output result of the overall model is judged as similar; otherwise, it is judged as "dissimilar". In this way, the unique semantic understanding ability of different models is utilized, and the output effect of the entire model is further improved through result integration and judgment.
[0113] Although the three pre-trained models NEZHA, RoBERTa, and ERNIE-Gram are basically improved from the BERT pre-trained model, since their improvement ideas in many aspects such as training methods and model structures are very different, the semantic understanding abilities of the three pre-trained models also have their own characteristics and are mutually superior and inferior. If the performance effect of each pre-trained model is regarded as "basic performance effect + characteristic additional effect", then the idea of the comprehensive matching model is to, on the basis of the "basic performance effects" of the three models, give full play to the "characteristic additional effects" of different models after comprehensive integration, so as to further enhance the final performance of the entire model. To further prove the inventive concept of this application, it will be verified from two aspects below.
[0114] Suppose the accuracies of the three pre-trained models NEZHA, RoBERTa, and ERNIE-Gram in semantic matching are p1, p2, and p3 respectively, and the three models are independent of each other. For the comprehensive model, when the results of any two or more pre-trained models are output as "similar", the output result of the overall model is judged as "similar"; otherwise, it is judged as "dissimilar". From this, the accuracy of the overall comprehensive model can be obtained as follows:
[0115] P = p1 * p2 * (1 - p3) + p1 * p3 * (1 - p2) + p2 * p3 * (1 - p1) + p1 * p2 * p3
[0116] = p1 * p2 + p1 * p3 + p2 * p3 – 2 * p1 * p2 * p3
[0117] If p1 = p2 = p3 = 0.9 and substitute it into the formula, we get P = 0.972. Thus, it can be seen that the accuracy of the comprehensive model has obvious advantages and improvements compared with the sub-models.
[0118] Three semantic matching models are given, namely Model 1, Model 2, and Model 3. For the convenience of understanding and calculation, assume that all three models can correctly judge whether 90 out of 100 samples in a dataset match, and the remaining 10 samples that cannot be correctly judged are sorted and in a continuous subsequence. First, consider an extreme case, such as Figure 8 As shown, at this time, Model 1, Model 2, and Model 3 make mistakes in judging 10 samples in the same sequence. At this time, if any sample is randomly selected, according to the judgment rule of the comprehensive model (the results of any 2 or more pre-trained models are judged correctly), the judgment accuracy rate of the comprehensive model is 90%, the same as that of the three sub-models. However, note that at this time, three completely different pre-trained models have exactly the same sample set of judgment errors, and this probability is very, very low.
[0119] Such as Figure 9 As shown is another extreme case. At this time, the sample sets in which Model 1, Model 2, and Model 3 make mistakes are different. At this time, if any sample is randomly selected, according to the judgment rule of the comprehensive model (the results of any 2 or more pre-trained models are judged correctly), the judgment accuracy rate of the comprehensive model is 100%, far higher than that of the three sub-models.
[0120] Based on the above analysis, although it is very difficult for the comprehensive model to achieve a judgment accuracy rate of 100% in practice, considering the randomness of the sample distribution and the characteristics of the three sub-models, the judgment accuracy rate of the comprehensive model will be higher than that of any sub-model, and it has obvious advantages and improvements compared with the sub-models.
[0121] To test the application effect of the model in the industry, the present invention respectively selects data from the production line technology field in the manufacturing industry and the disease consultation field in the medical industry for testing. After data cleaning, part of the data with missing indicators such as no labels is filtered out, and 10 pieces of data are randomly selected. Tables 4 and 5 respectively show partial data displays of the production line technology field in the manufacturing industry and the disease consultation field in the medical industry. Each sample is a sentence pair, either similar or dissimilar, and the model outputs a judgment result after input.
[0122]
[0123] Table 4 Partial data of the production line technology field in the manufacturing industry
[0124]
[0125] Table 5 Partial data display of the disease consultation field in the medical industry
[0126] Experimental environment and experimental settings
[0127] The experimental environment of the present invention is as follows: The operating system is Windows 10, the CPU is Intel(R) Core(TM) i7-10510U CPU @ 1.80GHz 2.30GHz, the GPU is v100, the video memory size is 32GB, the Paddle deep learning framework is adopted, programming is implemented using the Python language, and the development tool used is Notebook.
[0128] The number of iterations for the experiment is set to 3, and the accuracy rate and loss rate are used as the evaluation criteria for the experiment. Let the total number of samples be N, and the number of correctly classified samples be n, then the accuracy rate (Accuracy) is:
[0129]
[0130] The experimental model of the present invention can be regarded as a three-layer network structure, including a text vector input layer, a semantic matching layer, and an output layer. There are many hyperparameters in the model that need to be set and adjusted. In the present invention, through multiple experiments, after each iteration is completed, the hyperparameters are set and adjusted according to the accuracy rate and loss rate of the experiment. After multiple iterative experiments, the hyperparameters set by the model are shown in Table 6.
[0131] Table 6 Hyperparameters of each neural network
[0132] Parameter Value Parameter Value Batch Size 100 Number of Epochs 3 Maximum Sequence Length 128 Learning Rate 5e-5 Weight Decay Coefficient 0.01 Learning Rate Warmup Ratio 0.1
[0133] 2) Experimental comparison and result analysis
[0134] The NEZHA, ERNIE-Gram, and RoBERTa models used in the present invention are compared with the basic version of the BERT, Tiny-BERT, and ALBERT pre-trained models.
[0135] In this experiment, the loss function adopted for each is the "cross-entropy" loss function. When selecting the optimizer, SGD, AdaGrad, RMSProp, and Adam optimizers are considered. Among them, since the Adam optimizer combines the advantages of the AdaGrad and RMSProp optimization algorithms, comprehensively considers the first moment estimation of the gradient (First Moment Estimation, that is, the mean of the gradient) and the second moment estimation (Second Moment Estimation, that is, the uncentered variance of the gradient), and continuously iteratively updates the network parameters, it is very suitable for applications in scenarios of large-scale data and parameters. Therefore, this solution selects Adam as the optimization function for each model. At the same time, considering that an important reason for the poor generalization performance of Adam is that its use of the L2 regularization term is not as effective as in SGD, the original definition of Weight Decay is combined to correct this problem.
[0136] In addition, considering the training costs and the improvement amplitudes of various pre-trained models comprehensively, it is found through multiple experiments that after the pre-trained models are trained for 3 epochs, neither the decrease in the loss rate nor the increase in the accuracy rate is obvious. Taking the three pre-trained models of Tiny-BERT, NEZHA, and ERNIE-Gram as examples, after three epochs of training, the change curves of the loss rates and accuracy rates of the three pre-trained models are respectively as Figure 10 and Figure 11 shown. It can be seen from this that after three epochs of training, neither the loss rate nor the accuracy rate of the model changes significantly. Therefore, in the experiment of this application, the number of training times "Epoch" of each model is set to 3 times, and the activation function uses the "Relu" function.
[0137] In the experiment, after the pre-trained models have experienced 3 epochs of training, in order to fully test each model, this application selects the semantic matching dataset of lcqmc in the experiment, which contains a total of 12,500 samples to be tested. After each model is tested on the test set, the corresponding accuracy rate and loss rate on the test set are obtained. The experimental test results of each model are finally compared as shown in Table 7.
[0138] Table 7 Comparison of the results of each model and the novel hybrid model proposed in this paper
[0139]
[0140]
[0141] It can be seen from Table 7 that the ERNIE-Gram, NEZHA, and RoBERTa models have improvements in both the accuracy rate and the loss rate compared with other pre-trained models (except for the loss rate of NEZHA). Whether it is the BERT, ALBERT, or Tiny-BERT model, the sub-models that make up the NERB semantic matching comprehensive model have better effects, which further ensures the accuracy rate and flexibility of the comprehensive semantic matching model when making judgments. Through learning and iteration on different pre-trained models with large-scale datasets, the entire NERB semantic matching comprehensive model can not only be applied for cross-industry migration, but also exert the unique semantic understanding capabilities of different sub-models, and further improve the output effect of the entire model through comprehensive judgment of the results.
[0142] According to the method of the industry text matching model based on deep learning in an embodiment of the present invention, a preset number of cross-industry data is obtained as a training set to obtain a statement to be matched; the statement to be matched is input into the industry text matching model NERB based on deep learning, and after data preprocessing, it is respectively input into the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram; among them, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and optimizer; based on the optimized pre-trained model, three text matching results are output after being matched by the optimized pre-trained model; according to the comprehensive judgment of the three text matching results, when any two text matching results or all three text matching results are output as similar, the output result of the industry text matching model is judged as similar, otherwise it is dissimilar. The present invention can solve the semantic matching problems in various application fields such as automotive production line technology reference in the manufacturing industry, patient consultation in the medical industry, and transaction search in the business field by introducing large-scale cross-industry data as a training set and integrating and applying the advantages of multiple pre-trained models with Chinese characteristics. More importantly, the finally applied model can achieve accurate semantic matching tasks for different industries without the need for industry-specific professional training data.
[0143] Next, a device for an industry text matching model based on deep learning according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0144] Figure 12 It is a schematic structural diagram of a device for an industry text matching model based on deep learning according to an embodiment of the present invention.
[0145] As Figure 12 shown, the device 10 for the industry text matching model based on deep learning includes: an acquisition module 100, a training module 200, an output module 300, and a judgment module 400.
[0146] The acquisition module 100 is used to obtain a preset number of cross-industry data as a training set to obtain a statement to be matched;
[0147] The training module 200 is used to input the statement to be matched into the industry text matching model NERB based on deep learning, and after data preprocessing, it is respectively input into the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram; among them, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and optimizer;
[0148] The output module 300 is used to output three text matching results based on the optimized pre-trained model after being matched by the optimized pre-trained model;
[0149] A judgment module 400 is configured to make a comprehensive judgment based on three text matching results. When any two or all three text matching results are output as similar, the output result of the industry text matching model is judged to be similar; otherwise, it is judged to be dissimilar.
[0150] According to the industry text matching model device based on deep learning in the embodiments of the present invention, by introducing a large-scale cross-industry data as the training set and integrating and applying the advantages of multiple pre-trained models with Chinese characteristics, semantic matching problems in various application fields such as technical references for automobile production lines in the manufacturing industry, patient consultations in the medical industry, and transaction searches in the commercial field can be solved. More importantly, the finally applied model can achieve accurate semantic matching tasks for different industries without the need for in-industry professional training data.
[0151] It should be noted that the foregoing explanations of the method embodiments of the industry text matching model based on deep learning are also applicable to this device, and will not be elaborated herein.
[0152] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0153] In the present invention, unless otherwise clearly defined and limited, the terms such as "installed", "connected", "connected to", "fixed" and the like should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0154] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0155] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0156] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for an industry text matching model based on deep learning, characterized in that, it includes the following steps: Obtain a preset number of cross-industry data as a training set to obtain sentences to be matched; Input the sentences to be matched into the industry text matching model NERB based on deep learning, and after data preprocessing, input them into the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram respectively; wherein, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and optimizer; Based on the optimized pre-trained model, after being matched by the optimized pre-trained model, output three text matching results; Make a comprehensive judgment according to the three text matching results. When any two or three text matching results are output as similar, the output result of the industry text matching model is judged as similar, otherwise it is judged as dissimilar; The optimization of the functional relative position encoding includes: the pre-trained model NEZHA adopts functional relative position encoding, and the output involves a sine function of relative position in the calculation of attention scores. The functional relative position encoding formula is as follows: The optimization of the full word coverage includes: the pre-trained model NEZHA adopts a full word coverage strategy. When a Chinese character is covered, other Chinese characters belonging to the same Chinese character are covered together; The optimization of the mixed precision training includes: the pre-trained model NEZHA adopts mixed precision training. In each training iteration, round the main weights to the half-precision floating-point format, and perform forward and backward passes using the weights, activations, and gradients stored in the half-precision floating-point format; convert the gradients to the single-precision floating-point format, and update the main weights using the single-precision floating-point format gradients; The optimization of the optimizer includes: the pre-trained model NEZHA adopts the LAMB optimizer, and the adaptive strategy adjusts the learning rate for each parameter in the LAMB optimizer; The method further includes: verifying the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram, including: For the industry text matching model NERB, when the results of any two or more pre-trained models are output as "similar", the output result of the industry text matching model NERB is judged as "similar", otherwise it is judged as "dissimilar". Then the accuracy rate of the industry text matching model NERB is: P = p1*p2*(1 - p3) + p1*p3*(1 - p2) + p2*p3*(1 - p1) + p1*p2*p3 = p1*p2 + p1*p3 + p2*p3 – 2*p1*p2*p3 wherein, p1, p2, and p3 are the accuracy rates of the three semantic matching models NEZHA, RoBERTa, and ERNIE-Gram of the pre-trained model during semantic matching respectively; If the three semantic matching models can correctly judge whether the second preset number of samples in a dataset containing the first preset number of samples match, and the remaining third preset number of samples that cannot be correctly judged are sorted and in a continuous subsequence.
2. The method for an industry text matching model based on deep learning according to claim 1, wherein, the optimized pre-trained model RoBERTa includes: Multiple model parameter quantities and training data; pre-adjusting optimizer hyperparameters; the pre-trained model RoBERTa selecting a preset number of training samples; removing the next sentence prediction task, and the data being continuously obtained from one document; using dynamic masking, obtaining multiple copies of data by copying a training sample, each copy of data using a different mask, and increasing the copying fraction, and generating a new mask pattern each time a sequence is input to the pre-trained model RoBERTa; using whole word masking.
3. The method for an industry text matching model based on deep learning according to claim 2, wherein, the optimized pre-trained model RoBERTa further includes: Text encoding, the pre-trained model RoBERTa being trained using a BPE vocabulary of a preset level of bytes during text encoding, and no additional preprocessing or tokenization being performed on the input.
4. The method for an industry text matching model based on deep learning according to claim 1, wherein, the optimized pre-trained model ERNIE-Gram includes: The ERNIE-Gram model explicitly introduces language granularity knowledge, an explicit n-gram masked language model, learns n-gram granularity language information, and based on the explicit n-gram masked language model, the pre-trained model ERNIE-Gram performs multi-level n-gram language granularity masking learning.
5. An industry text matching model device based on deep learning, wherein, it includes: An acquisition module, configured to acquire a preset number of cross-industry data as a training set to obtain sentences to be matched; A training module, configured to input the sentences to be matched into the industry text matching model NERB based on deep learning, and after data preprocessing, input them into the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram respectively; wherein, the optimized pre-trained model NEZHA includes: optimization of functional relative position encoding, full word coverage, mixed precision training, and the optimizer. An output module, configured to output three text matching results based on the optimized pre-trained model after being matched by the optimized pre-trained model; A judgment module, configured to perform a comprehensive judgment according to the three text matching results. When any two or all three text matching results are output as similar, the output result of the industry text matching model is judged to be similar, otherwise it is judged to be dissimilar; The optimization of the functional relative position encoding includes: The pre-trained model NEZHA outputs a sine function of the relative position involved in the calculation of the attention score by adopting the functional relative position encoding. The formula of the functional relative position encoding is as follows: The optimization of the whole-word coverage includes: The pre-trained model NEZHA adopts the whole-word coverage strategy. When a Chinese character is covered, other Chinese characters belonging to the same Chinese character are covered together; The optimization of the mixed-precision training includes: The pre-trained model NEZHA adopts the mixed-precision training. In each training iteration, the main weights are rounded to the half-precision floating-point format, and the forward and backward passes are performed using the weights, activations, and gradients stored in the half-precision floating-point format; The gradients are converted to the single-precision floating-point format, and the main weights are updated using the single-precision floating-point format gradients; The optimization of the optimizer includes: The pre-trained model NEZHA adopts the LAMB optimizer, and the adaptive strategy adjusts the learning rate for each parameter in the LAMB optimizer; The device further includes: verifying the optimized pre-trained models NEZHA, RoBERTa, and ERNIE-Gram, including: For the industry text matching model NERB, when the output results of any two or more pre-trained models are "similar", the output result of the industry text matching model NERB is judged as "similar", otherwise it is "dissimilar". Then the accuracy rate of the industry text matching model NERB is: P = p1*p2*(1-p3) + p1*p3*(1-p2) + p2*p3*(1-p1) + p1*p2*p3 = p1*p2 + p1*p3 + p2*p3 – 2*p1*p2*p3 Where p1, p2, and p3 are the accuracy rates of the three semantic matching models of the pre-trained models NEZHA, RoBERTa, and ERNIE-Gram respectively when performing semantic matching; If the three semantic matching models can correctly judge whether the second preset number of samples match in the dataset containing the first preset number of samples, and the remaining third preset number of samples that cannot be correctly judged are sorted and in a continuous subsequence.
Citation Information
Patent Citations
A method and device for generating a text matching model
CN109947919A
Intelligent semantic matching method and device based on deep hierarchical coding
CN111325028A
Cited By
Industrial model matching method based on natural language remote sensing information
CN122310135A