Text Matching Model Training Method, Apparatus, Electronic Device, and Storage Medium

By setting the dropout layer in the BERT model and performing unsupervised learning, the problem of traditional text matching models depend on labeled data is solved, efficient and accurate text matching model training is achieved, and user experience is improved.

CN113850383BActive Publication Date: 2025-07-25PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111134466.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-07-25
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

Traditional text matching model training relies on a large amount of labeled data, which leads to high time cost and is difficult to achieve efficient training in the absence of data.

Method used

Unsupervised learning method is adopted, and the dropout layer in the pre-trained BERT model is optimized by inputting the same data multiple times to reduce the dependence on the label data, and the randomness of the dropout layer and the activation ratio of the activation ratio is less than 1 to optimize the model parameters.

Benefits of technology

It reduces the cost of model training, improves training efficiency and matching accuracy, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850383B_ABST
    Figure CN113850383B_ABST
Patent Text Reader

Abstract

This application relates to artificial intelligence and provides a method for training a text matching model, including: obtaining a text matching model, where the text matching model includes a pre-trained BERT model, and the pre-trained BERT model includes a dropout layer; obtaining training data, where the data samples in the training data do not include labels; inputting the training data into the text matching model at least twice to respectively obtain the output results of the text matching model; obtaining a similarity representation between the output results; optimizing the parameters in the model based on the loss function corresponding to the similarity representation to obtain a trained text matching model. By setting a dropout layer with a determined activation ratio in the BERT model and inputting the same data twice, and performing reverse parameter optimization based on the differences between the two output results, since there is no need to use data with labels, the cost of model training is reduced and the efficiency of model training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to, but are not limited to, the field of artificial intelligence, and in particular, to a method, device, electronic device, and computer-readable storage medium for training a text matching model. Background Art

[0002] Text matching is an important basic problem in natural language processing and can be applied to a large number of natural language processing tasks, such as information retrieval, dialogue systems, etc. Taking an intelligent dialogue system as an example, by accurately matching the user's question with the preset questions in the dialogue library, the most suitable reply can be given to the user.

[0003] Most traditional text matching models rely on word and word frequency information, while deep learning models can mine deeper text information, thereby improving the accuracy of text matching. However, the training of deep learning models requires a large amount of labeled data, which not only requires a large time cost but also is difficult to implement in some practical cases where labeled data is lacking. Summary of the Invention

[0004] The embodiments of the present application provide a method, device, electronic device, and computer-readable storage medium for training a text matching model. By training the text matching model through unsupervised learning, the model training cost is reduced, and the model training efficiency and the accuracy of the final text matching are improved.

[0005] In a first aspect, an embodiment of the present application provides a method for training a text matching model, including: obtaining a text matching model, where the text matching model includes a pre-trained BERT model, and the pre-trained BERT model includes a dropout layer, and the activation ratio of the dropout layer is less than 1; obtaining training data, where the data samples in the training data do not include labels; inputting the training data into the text matching model multiple times to respectively obtain multiple output results of the text matching model; obtaining a similarity representation between the output results; and optimizing the parameters in the text matching model based on the loss function corresponding to the similarity representation to obtain a trained text matching model. The method for training a text matching model provided in this embodiment reduces the cost of model training and improves the efficiency of model training by setting a dropout layer with a determined activation ratio in the BERT model, inputting the same data twice, and performing reverse parameter optimization based on the differences between the two output results. Since there is no need to use data with labels, the cost of model training is reduced.

[0006] In a second aspect, an embodiment of the present application provides a text matching device, including: a text acquisition module for acquiring text to be matched; and a text matching module for performing text matching according to the method for training a text matching model described in the first aspect.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the processor. The programs include those for executing the text matching model training method as described in the first aspect.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including: program instructions that can be run by a processor, and the program instructions are used to execute the text matching model training method as described in the first aspect.

[0009] In the embodiment of the present application, by setting a dropout layer with a determined activation ratio in the BERT model and using the same batch of data for training and parameter optimization, since there is no need to use labeled data, the cost of model training is reduced, the efficiency of model training is improved. At the same time, the trained text matching model has a high matching accuracy during use, improving the user experience.

[0010] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings

[0011] Figure 1 It is a schematic flowchart of the text matching model method provided by the embodiment of the present application;

[0012] Figure 2 It is a schematic diagram of the dropout layer structure in the BERT model provided by the embodiment of the present application;

[0013] Figure 3 It is a schematic flowchart of the text matching method provided by the embodiment of the present application;

[0014] Figure 4 It is a schematic flowchart of the text matching method provided by another embodiment of the present application;

[0015] Figure 5 It is a mapping relationship diagram of matching text and target text provided by the embodiment of the present application. Detailed Embodiments

[0016] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application. Without conflict, the embodiments and features in the embodiments of the present application can be combined arbitrarily with each other.

[0017] Terms such as "first", "second", etc. in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that if it involves orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., it is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present application.

[0018] It should be noted that the meaning of "at least one" is one or more, the meaning of "multiple" is two or more, "greater than", "less than", "exceeding", etc. are understood as not including the present number, and "above", "below", "within", etc. are understood as including the present number. If "first" and "second" are described only for the purpose of distinguishing technical features, they cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0019] In the prior art, dialogue systems can generally be divided into three categories: chitchat-bot, IR-bot, and task-bot. For any of the above dialogue systems, natural language from users needs to be processed. In a dialogue system or a question-and-answer system, it is manifested as recognizing the natural language input by the user, analyzing and processing the natural language, and matching it with the text preset in the system.

[0020] Generally, a question-and-answer system will give a list of standard texts most relevant to the query text input by the user, where the standard text indicates the user's intention or the user's question. If the user clicks on a certain standard text in the list, the solution corresponding to the marked text will be shown to the user.

[0021] The BERT (Bidirectional Encoder Representations from Transformers) model is a deep semantic understanding model. The embodiments of the present application are based on the Bert model and additional structures and training methods are combined to achieve more accurate model output.

[0022] Due to the limitations of artificial intelligence, in order to ensure the basic user experience, each dialogue or question-and-answer system requires a large amount of labeled data for training. Considering the diversity of user input, obtaining a large amount of labeled data itself requires a relatively high cost and is sometimes even difficult to achieve.

[0023] Based on this, the present application proposes a method for training a text matching model.

[0024] In a first aspect, Figure 1 is a schematic flowchart of a method for training a text matching model provided in another embodiment of the present application. Refer to Figure 1 The method for training a text matching model provided in the embodiment at least includes:

[0025] Step S110: Obtain a text matching model, where the text matching model includes a pre-trained BERT model, and the pre-trained BERT model includes a dropout layer, and the activation ratio of the dropout layer is less than 1.

[0026] Dropout refers to randomly removing neural network training units from the network with a certain probability during the training process of a machine learning model. Since it is random removal, each batch of data is training a different network, and even when the same batch of data is input for training, it will still result in different output results. Therefore, the model also changes dynamically after each batch of data training.

[0027] In some embodiments, as Figure 2 shown, a dropout function is added after each network layer to improve the generalization ability of the model by randomly discarding some training parameters in each iteration.

[0028] In some embodiments, the activation ratio of dropout is 0.5, that is, the training parameters of half of the neural units are randomly removed.

[0029] In some embodiments, as long as the activation ratio of dropout is less than 1, it can meet the requirements of the embodiments of the present application, that is, as long as the randomness of removing neural units is ensured.

[0030] Step S120: Obtain training data, and the data samples in the training data do not include labels.

[0031] In some embodiments, using data without labels for model training belongs to unsupervised learning.

[0032] Step S130: Input the training data into the text matching model multiple times to obtain multiple output results of the text matching model respectively.

[0033] The essence of this step is to input the same data into the text matching model twice. Since the model has a dropout mechanism, the two output results of the same data will be different. The purpose of the embodiments of the present application is to make the two output results of the same data as close as possible or even the same through parameter fine-tuning and optimization.

[0034] In some embodiments, the same data is input into the model twice, and the output results are obtained respectively.

[0035] In some embodiments, the same data is input into the model three times or even more times, and the output results are obtained separately.

[0036] Step S140: Obtain the similarity representation between the output results.

[0037] The output result is a vector representation of the text, and there are various representation methods for the similarity representation of two output results.

[0038] In some embodiments, Cosine similarity is used for similarity representation.

[0039] In some embodiments, Euclidean distance is used for similarity representation.

[0040] In some embodiments, Jaccard similarity, Pearson correlation coefficient, Mahalanobis distance, etc. are used for similarity representation.

[0041] It should be noted that if the same data is input into the model more than twice in step S140, the similarity representation can be calculated between any two of the multiple output results.

[0042] Step S150: Optimize the parameters in the text matching model based on the loss function corresponding to the similarity representation.

[0043] Each similarity representation can correspond to its loss function, and the model parameters can be optimized by minimizing the loss function.

[0044] In some embodiments, the BP backpropagation algorithm is used for model iteration and optimization.

[0045] The text matching model training method provided in this embodiment reduces the cost of model training and improves the efficiency of model training by setting a dropout layer with a determined activation ratio in the BERT model, inputting the same data twice, and performing reverse parameter optimization based on the differences between the two output results, since there is no need to use labeled data.

[0046] In some embodiments, when the model training stage ends and the output of the model is relatively stable, the training data is input into the trained text matching model again, and the output result is obtained as the standard text data. In subsequent model usage, the output of the test data is compared with the standard text data, and the text corresponding to the standard text data that is closest is selected for output.

[0047] The following uses a specific embodiment to illustrate the text matching model training method:

[0048] Assume that the training data is {x i}, i = 1, 2,..., N, xi Represents an independent sentence. Using the pre-trained model BERT, the BERT model has a dropout structure. During the training process, a certain proportion of units in the dropout layer are randomly activated. In this embodiment, assuming that the dropout layer has 12 units and the activation ratio is 50%, then during the training process, each unit has only a 50% probability of participating in the operation, that is, only 6 units participate in the calculation. Pass all the training data through the BERT model twice. Due to the randomness of the dropout layer, two different sentence representations of the same sentence can be obtained, denoted as {s i,1 , s i,2}, i = 1, 2,..., N, where s i,1 , s i,2 are the final output results of the sentence x i input into BERT twice respectively.

[0049] In this embodiment, cosine similarity is used for representation. The purpose of model training is to make the vector representations generated by the same sentence twice close enough in distance. Therefore, for the sentence x i , its corresponding loss function is defined as That is, for the i-th sentence, the vector representations output by BERT twice are close enough, but the vector representations output by BERT of the i-th sentence and other sentences are far enough apart. Therefore, the loss function of the model is defined as Optimize the model parameters through the BP backpropagation algorithm to make the loss function approach the minimum value as much as possible.

[0050] Figure 3 The flowchart of the text matching model training method provided by another embodiment of this application. The text matching model training method provided by this embodiment at least includes:

[0051] Step S210: Obtain the text to be matched.

[0052] In some embodiments, the text to be matched is a sentence or even a keyword input by the user. The sentence input by the user can be in multiple languages, such as Chinese or English.

[0053] In some embodiments, it is also possible to identify the user's voice input, convert the voice into text, and then perform the matching.

[0054] Step S220: Obtain the text matching model.

[0055] In some embodiments, the text matching model is obtained by training with the text matching model training method provided in the embodiments of the present application. However, those skilled in the art should be aware that, based on the model training method provided in the embodiments of the present application, a text matching model obtained by adding other parameter fine-tuning methods is also within the protection scope of the embodiments of the present application.

[0056] Step S230: Input the text to be matched into the text matching model.

[0057] In some embodiments, the input method of the text to be matched is the same as that of the training data. For a specific application scenario, such as a question-and-answer system, a question query is input as the text to be matched.

[0058] Step S240: Use Cosine to calculate the similarity between the text to be matched and the standard text data.

[0059] In some embodiments, the trained BERT model can output the vector representation of the text to be matched, and use Cosine to represent the distance between the vector of the text to be matched and the preset standard text data.

[0060] Step S250: Sort the similarities and select the standard text data with the highest similarity as the matching text.

[0061] In some embodiments, since the standard text data is a set that contains the model output vectors corresponding to all preset texts, it is necessary to calculate the similarity between the text to be matched and each standard text data one by one, and select the standard text data with the highest similarity as the matching text.

[0062] In some embodiments, output the target text corresponding to the matching text. As Figure 5 shown, there is a mapping relationship between the target text and the matching text.

[0063] The text matching method provided in this embodiment is based on the text matching model obtained by training in the first aspect, calculates the similarity between the text to be matched and the standard text, and can accurately and efficiently select the correct matching text, and then display the target text.

[0064] Figure 4 The flowchart of the text matching model training method provided in another embodiment of the present application. The text matching model training method provided in this embodiment at least includes:

[0065] Step S310: Obtain the text to be matched.

[0066] In some embodiments, the text to be matched is a sentence or even a keyword input by the user, and the sentence input by the user can be in multiple languages, such as Chinese or English.

[0067] In some embodiments, it is also possible to identify the user's voice input, convert the voice into text, and then perform the matching.

[0068] Step S320: Obtain a text matching model.

[0069] In some embodiments, the text matching model is trained by the text matching model training method provided in this embodiment of the application. However, those skilled in the art should be aware that a text matching model obtained by adding other parameter fine-tuning methods based on the model training method provided in this embodiment of the application is also within the protection scope of this embodiment of the application.

[0070] Step S330: Input the text to be matched into the text matching model.

[0071] In some embodiments, the input method of the text to be matched is the same as that of the training data. For a specific application scenario, such as a retrieval system, a retrieval keyword or phrase is input as the text to be matched.

[0072] Step S340: Use the Euclidean distance cosine to calculate the similarity between the text to be matched and the standard text data.

[0073] In some embodiments, the trained BERT model can output the vector representation of the text to be matched, and the Euclidean distance is used to represent the distance between the vector of the text to be matched and the preset standard text data.

[0074] Step S250: Select the standard text data with a similarity greater than the preset threshold as the matching text.

[0075] In some embodiments, since the standard text data is a set that contains the model output vectors corresponding to all preset texts, it is necessary to calculate the similarity between the text to be matched and each standard text data one by one, and select the standard text data with a similarity greater than the preset threshold as the matching text.

[0076] It should be noted that the preset rules for similarity judgment are not limited to the selection of the maximum similarity or the selection of greater than the threshold, and can also be set according to the rules learned during the training process or the user's preferences.

[0077] The text matching method provided in this embodiment is based on the text matching model trained in the first aspect, calculates the similarity between the text to be matched and the standard text, and can accurately and efficiently select the correct matching text, and then display the target text.

[0078] In a second aspect, a text matching device provided by an embodiment of the present application at least includes: a text acquisition module for acquiring a text to be matched, and a text matching module for performing text matching according to the text matching model training method described in the first aspect.

[0079] The text acquisition module can not only acquire the text information typed by the user, but also acquire the voice information or even the image information of the user.

[0080] The text matching module is used to execute the trained text matching method, and the text matching model is stored in the text matching module.

[0081] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory: the memory is used to store a program; the processor is used to execute the program to execute the text matching model training method of any of the above embodiments.

[0082] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including:

[0083] The computer-readable storage medium stores a program, and the program is executed by a processor to complete the text matching model training method of any of the above embodiments.

[0084] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0087] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. The mobile terminal device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle terminal device, a wearable device, an ultra-mobile personal computer, a netbook, a personal digital assistant, a CPE, a UFI (wireless hotspot device), etc.; the embodiments of the present invention are not specifically limited. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0088] In the above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them: Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features: and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a text matching model, comprising: Obtaining a text matching model, the text matching model including a pre-trained BERT model, the pre-trained BERT model including a dropout layer, wherein the activation ratio of the dropout layer is less than 1; Obtaining training data, the data samples in the training data not including labels; Inputting the training data into the text matching model more than twice, respectively obtaining multiple output results of the text matching model; Obtain multiple output results {s of the above training data input into the text matching model more than twice i,1 , s i,2 ,... s i,k}, where i represents the label of each training data, i = 1, 2,..., N, and k represents the number of times this training data is input into the text matching model, and k is greater than 2; Use the Cosine formula to calculate the distance Cosin(s i,m , s i,n ) between any two output results, where m and n represent the labels of any two output results; the distance Cosin(s i,m , s i,n ) represents the similarity between the two output results; Based on the loss function corresponding to the similarity representation, optimizing the parameters in the text matching model through the backpropagation algorithm to obtain a trained text matching model; wherein, a dropout layer is added after each network layer of the text matching model, and part of the training parameters are randomly discarded through the dropout layer in each iteration; Inputting the training data into the trained text matching model, and using the output result as standard text data.

2. The text matching model training method according to claim 1, wherein The similarity representation corresponding to the Cosine cosine definition, and the loss function is The loss function of the text matching model is defined as 3. The text matching model training method according to claim 1, wherein Further comprising: Obtaining a text to be matched; Inputting the text to be matched into the text matching model; Calculating the similarity between the text to be matched and the standard text data; Selecting a matching text for the text to be matched according to a preset rule for similarity judgment.

4. The text matching model training method according to claim 3, wherein The selecting a matching text for the text to be matched according to the preset rule for similarity judgment includes: Sorting the similarities, and selecting the text corresponding to the standard text data with the highest similarity as the matching text; Or, selecting the text corresponding to the standard text data with a similarity greater than a preset threshold as the matching text.

5. The method for training a text matching model according to claim 3 or 4, characterized in that After obtaining the matching text, outputting a target text corresponding to the matching text; wherein the matching text includes a question text, and the target text includes a reply text corresponding to the question text.

6. A text matching device, characterized in that, Comprising: A text acquisition module, configured to obtain a text to be matched; A text matching module, configured to perform text matching according to the text matching model training method according to any one of claims 3-5.

7. An electronic device, the electronic device including a processor, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the processor, the programs including for executing the text matching model training method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, Storing program instructions executable by a processor, the program instructions for executing the text matching model training method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Text matching method and device, equipment and storage medium

    CN113297353A