Question and answer model training method and device, text question and answer method and device

By training the Q&A model based on the initial text sample set division under low supervision resource conditions, the problem of difficulty in effectively training the model in professional field knowledge Q&A in the existing technology is solved, and efficient and accurate Q&A model training is achieved.

CN117251540BActive Publication Date: 2025-06-27ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311062126.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-06-27
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively train Q&A models in low-supervised resource scenarios in professional field knowledge Q&A, especially when building high-quality question-answer sets, requiring a lot of professional knowledge to label people and time.

Method used

By obtaining the initial text sample set, identifying the target text sample, and dividing it into query text sample and main text sample, and negative text sample, a question-and-answer model is obtained based on these samples. This method does not require any labeled data and can train a Q&A model under low supervised resource conditions.

Benefits of technology

The training of Q&A model under low supervision resources is achieved, reducing labor costs and greatly enriching the scale of training samples, thereby improving the accuracy of the Q&A model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117251540B_ABST
    Figure CN117251540B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method and device for training a question-and-answer model, and a method and device for text question-and-answer. The method for training the question-and-answer model includes: obtaining an initial text sample set, and sequentially determining each initial text sample in the initial text sample set as a target text sample; in the case where the target text sample is determined to be a preset text sample, determining the first text in the target text sample as a query text sample, and the second text as the positive text sample corresponding to the query text sample; determining any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample; training a question-and-answer model according to the query text sample, the positive text sample, and the first negative text sample; enabling the construction of the entire training sample without any labeled data, not only reducing the labor cost, but also greatly enriching the scale of the training sample, and being able to obtain a more accurate question-and-answer model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for training a question-answering model. Background Art

[0002] In some professional field scenarios, users will ask questions about professional field knowledge; for example, users may have problems during the use of certain products and need to seek answers from professionals in the professional field. Answering these questions often requires rich professional knowledge and experience, and also requires a lot of human resources.

[0003] Currently, the mainstream question-answering models are all constructed based on deep learning neural network models. In the field of deep learning, question-answering models usually include two types: generative question-answering models and retrieval-based question-answering models. Among them, the generative question-answering model constructs a generative model and then predicts the answer word by word according to the question. The retrieval-based question-answering model needs to prepare a set of answer candidates for the question in advance, represent both the question and the answers in the set of answer candidates as semantic vectors, and then find the answer that best matches the question from the set of answer candidates through vector retrieval.

[0004] However, the generative question-answering model is widely used in free conversation scenarios. In professional question-answering scenarios, when a higher accuracy of answers is required, the generative question-answering model is difficult to be competent due to the uncontrollability of its prediction process. Therefore, the retrieval-based question-answering model is mainly used. And the retrieval-based question-answering model in deep learning requires a large number of training samples (also called supervised data) during the training process, that is, a set of question-answer pairs. In practical applications, it is very difficult to construct a high-quality set of question-answer pairs, and a large number of annotators with professional knowledge need to spend a lot of time and effort for annotation.

[0005] Therefore, there is an urgent need for a question-answering model for professional field knowledge that can adapt to low-supervised resource scenarios (data resources that only contain a very small scale of labeled data). Summary of the Invention

[0006] In view of this, the embodiments of this specification provide a method for training a question-answering model and a method for text question-answering. One or more embodiments of this specification also relate to a device for training a question-answering model, a device for text question-answering, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.

[0007] According to the first aspect of the embodiments of this specification, a method for training a question-answering model is provided, including:

[0008] Obtain an initial text sample set, and sequentially determine each initial text sample in the initial text sample set as a target text sample;

[0009] In the case where the target text sample is determined to be a preset text sample, the first text in the target text sample is determined as the query text sample, and the second text is determined as the positive text sample corresponding to the query text sample, where the first text is a preset number of texts obtained from the start position of the text in the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length;

[0010] Any other initial text sample in the initial text sample set except the target text sample is determined as the first negative text sample corresponding to the query text sample;

[0011] A question-answering model is trained based on the query text sample, the positive text sample, and the first negative text sample.

[0012] According to the second aspect of the embodiments of the present specification, a question-answering model training device is provided, including:

[0013] A target text sample determination module, configured to obtain an initial text sample set, and sequentially determine each initial text sample in the initial text sample set as a target text sample;

[0014] A text sample determination module, configured to, in the case where the target text sample is determined to be a preset text sample, determine the first text in the target text sample as the query text sample, and the second text as the positive text sample corresponding to the query text sample, where the first text is a preset number of texts obtained from the start position of the text in the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length;

[0015] A first negative text sample determination module, configured to determine any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample;

[0016] A model obtaining module, configured to train a question-answering model based on the query text sample, the positive text sample, and the first negative text sample.

[0017] According to the third aspect of the embodiments of the present specification, a text question-answering method is provided, including:

[0018] Determine a question query text and the answer text set corresponding to the question query text;

[0019] Input the problem query text and the set of answer texts into the question-answering model to obtain the target answer corresponding to the problem query text, where the question-answering model is trained by the above-mentioned question-answering model training method.

[0020] According to the fourth aspect of the embodiments of the present specification, a text question-answering device is provided, including:

[0021] A text determination module configured to determine a problem query text and a set of answer texts corresponding to the problem query text;

[0022] An answer obtaining module configured to input the problem query text and the set of answer texts into the question-answering model to obtain the target answer corresponding to the problem query text, where the question-answering model is trained by the above-mentioned question-answering model training method.

[0023] According to the fifth aspect of the embodiments of the present specification, a computing device is provided, including:

[0024] A memory and a processor;

[0025] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned question-answering model training method and / or the above-mentioned text question-answering method are implemented.

[0026] According to the sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned question-answering model training method and / or the above-mentioned text question-answering method are implemented.

[0027] According to the seventh aspect of the embodiments of the present specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned question-answering model training method and / or the above-mentioned text question-answering method.

[0028] The Q&A model training method provided by an embodiment of this specification includes obtaining an initial text sample set, and sequentially determining each initial text sample in the initial text sample set as a target text sample; in the case where the target text sample is determined to be a preset text sample, determining the first text in the target text sample as a query text sample, and the second text as the positive text sample corresponding to the query text sample, where the first text is a preset number of texts obtained from the start position of the text of the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length; determining any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample; and training a Q&A model according to the query text sample, the positive text sample, and the first negative text sample.

[0029] Specifically, this method traverses the initial text sample set, determines a query text sample, a positive text sample corresponding to the query text sample, and a negative text sample according to each initial text sample in the initial text sample set; subsequently, a triple training sample constructed by the query text sample, the positive text sample corresponding to the query text sample, and the negative text sample is used to train the Q&A model, so that the construction of the entire training sample does not require any labeled data, which not only reduces the labor cost, but also greatly enriches the scale of the training sample, and a more accurate Q&A model can be obtained. Description of the Drawings

[0030] Figure 1 is a schematic diagram of the application process of a Q&A model training method provided by an embodiment of this specification;

[0031] Figure 2 is a flowchart of a Q&A model training method provided by an embodiment of this specification;

[0032] Figure 3 is a schematic diagram of the model training process of a Q&A model training method provided by an embodiment of this specification;

[0033] Figure 4 is a flowchart of a text Q&A method provided by an embodiment of this specification;

[0034] Figure 5 is a process flowchart of a text Q&A method provided by an embodiment of this specification;

[0035] Figure 6 is a schematic diagram of the structure of a Q&A model training device provided by an embodiment of this specification;

[0036] Figure 7 It is a schematic structural diagram of a text Q&A device provided by an embodiment of this specification;

[0037] Figure 8 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific embodiments

[0038] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0039] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.

[0040] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".

[0041] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or refuse.

[0042] First, the noun terms involved in one or more embodiments of this specification are explained.

[0043] BERT: Bidirectional Encoder Representation from Transformers, a bidirectional encoder representation based on a large-scale corpus pre-training model, used to represent words or sentences in a sentence or article as vectors.

[0044] Triplet Loss: Triplet loss, commonly used as the training objective for contrastive learning models.

[0045] Dropout: Dropout, a training technique method that sets some parameters in the neural network to 0, usually used to prevent overfitting in neural network models.

[0046] There are mainly two existing neural network question-and-answer models: generative question-and-answer models and retrieval question-and-answer models. When applying generative question-and-answer models, there is no need to pre-specify the range of answers corresponding to questions, and they can answer some open-domain questions and can be used for daily conversation services. Their disadvantage is that the process of generating answers corresponding to questions is uncontrollable, so the predicted answers often do not match the questions or the answer content is inaccurate, and it is difficult to be used for professional field question-and-answer. Retrieval question-and-answer models mainly use contrastive learning to train neural networks. Since retrieval question-and-answer models only give answers to questions within a specified range, they are commonly used in professional question-and-answer applications. Usually, retrieval question-and-answer models require a large number of question-answer data pairs as training samples. However, in actual applications, it is very difficult to obtain such a high-quality large-scale question-answer data pair training set, and this requirement for training data also limits the large-scale application of existing retrieval question-and-answer models.

[0047] In this specification, a method for training a question-and-answer model and a method for text question-and-answer are provided. This specification also relates to a device for training a question-and-answer model, a device for text question-and-answer, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0048] See Figure 1 , Figure 1 shows a specific processing flowchart of a method for training a question-and-answer model according to an embodiment of this specification.

[0049] Figure 1 includes a server 106 and a terminal device 110. Among them, the server 106 can be understood as a physical server or a cloud server, and this specification does not limit this; the terminal device 110 can be understood as various electronic devices, which can be screen devices or screenless devices. Including but not limited to smartphones, laptops, tablets, smart speakers, smart TVs, PCs (Personal Computers), wearable devices, and so on.

[0050] For ease of understanding, in the embodiments of this specification, the server 106 is taken as a physical server and the terminal device 110 is taken as a laptop computer as examples to describe the question-and-answer model training method in detail.

[0051] Specifically, when implementing, the question-and-answer model 104 is trained in the server 106. Specifically, first, the training sample 102 is obtained.

[0052] The training sample 102 is obtained by traversing the pre-prepared answer set, determining an answer text from the answer set. When the answer text is a long text, the first few sentences in the answer text can be used as the anchor point, and the remaining sentences can be used as positive samples; when the answer text is not a long text, the answer text is input into the question-and-answer model, and the answer text is masked in the question-and-answer model to obtain two masked vectors. One of the masked vectors is used as the anchor point, and the other masked vector is used as the positive sample. Of course, when the answer text is a long text, this method can also be used to obtain the anchor point and positive samples.

[0053] Any other answer text in the answer set except the determined answer text is used as a common negative sample; alternatively, the entities in the determined answer text can be replaced with other entities of the same category, and the replaced text is used as a difficult negative sample.

[0054] The above determined anchor point, positive sample, common negative sample, and difficult negative sample are used as the training sample 102, and the question-and-answer model 104 is trained in the server 106 through the training sample 102.

[0055] In one common scenario, the user can use the terminal device 110 to interact with the question-and-answer model 104 set in the server 106. Various applications can be installed on the terminal device 110, such as voice interaction applications, web browser applications, communication applications, etc.

[0056] The terminal device 110 receives the user input question text 108. The terminal device 110 calls the question-and-answer model 104 in the server 106, inputs the question text 108 and the corresponding answer text set (i.e., the above pre-prepared answer set) into the question-and-answer model 104, obtains the target answer 112 corresponding to the question text 108 in the question-and-answer model 104, and can return the target answer 112 to the terminal device 110 so that the user can obtain the target answer 112.

[0057] In another common scenario, the Q&A model 104 trained in the server 106 can also be deployed to the terminal device 110. Then, the terminal device 110 receives the user input question text 108, and inputs the question text 108, the corresponding answer text set (i.e., the pre-prepared answer set mentioned above), and the Q&A model 104 deployed in the terminal device 110 into the Q&A model 104 to obtain the target answer 112 corresponding to the question text 108 in the Q&A model 104.

[0058] The user can input the question in the form of text or voice. If the voice form is adopted, the server 106 will also include corresponding voice processing parts, such as voice parsing, voice-to-text conversion, voice synthesis and other modules, which are not limited in this specification.

[0059] It should be understood that Figure 1 the number of terminal devices and Q&A models in

[0060] Through the application process of the Q&A model training method in the embodiments of this specification, after obtaining the training samples, a Q&A model can be pre-trained in advance, the target answer corresponding to the user's question text can be obtained based on the Q&A model, and the target answer can be returned to the user.

[0061] The following combines the attached Figure 2 , and further describes the Q&A model training method. Among them, Figure 2 FIG. shows a flowchart of a Q&A model training method provided by an embodiment of this specification, which specifically includes the following steps.

[0062] Step 202: Obtain an initial text sample set, and sequentially determine each initial text sample in the initial text sample set as a target text sample.

[0063] Among them, the initial text sample set can be understood as a set containing multiple initial text samples. Among them, the initial text sample can be understood as a document of any type, any length, and any professional field, such as a document for product usage instructions, a professional knowledge document in a certain field, etc.

[0064] The target text sample can be understood as traversing each initial text sample in the initial text sample set and determining the sample to be trained when the current initial text sample is traversed.

[0065] In specific implementations, the text collection containing multiple texts is determined to be different according to different application scenarios. For example, in the scenario where a user has a problem using a product and seeks answers or help, the text collection can be understood as a collection of help documents that explain the product; in the scenario where a user learns knowledge in a certain field and queries a certain knowledge point in the field, the text collection can be understood as a collection of professional knowledge documents in the field.

[0066] The following embodiments all take the scenario where the question-answering model training method is applied to a user who has questions when using a product and seeks answers or help as an example to explain the question-answering model training method in detail.

[0067] Then, first determine a help document set for explaining the product, the help document set is a pre-prepared document set, traverse each document in the help document set, and use the currently traversed document as a to-be-trained sample for the training model.

[0068] In one or more embodiments of the present specification, after determining the target text sample, by calculating the text length of the target text sample, if the text length is greater than or equal to the preset text length, the target text sample is determined to be a preset text sample, so as to process the target text sample according to the processing method of the preset text sample. The specific implementation method is as follows:

[0069] After obtaining the initial text sample set and determining each initial text sample in the initial text sample set as a target text sample in sequence, the method further includes:

[0070] The text length of the target text sample is determined, and when the text length is greater than or equal to a preset text length, the target text sample is determined to be a preset text sample.

[0071] Among them, the text length can be understood as the number of text characters in the text; the preset text length can be understood as the number of text characters in the preset text. Specifically, the preset text length can be customized or determined by calculation rules; the preset text sample can be understood as a text sample with a text length greater than or equal to the preset text length, such as a long text sample.

[0072] In practical applications, a preset text length is set in advance. The preset text length can be a custom text length or a text length threshold obtained by analyzing the text lengths of the documents in the help document set. For example, if the text length threshold is 10 characters, then when the text length of the target text sample is 20 characters, it is determined that the target text sample is a preset text sample. Specifically, the text length of the target text sample is determined. When the text length of the target text sample exceeds the custom text length or the text length threshold, it is determined that the target text sample is a long text sample.

[0073] A question-and-answer model training method provided by an embodiment of this specification calculates the text length of a target text sample, determines whether the text length of the target text sample exceeds the preset text length, and when the text length of the target text sample exceeds the preset text length, determines that the target text sample is a long text sample, and constructs a positive text sample in a manner of quickly constructing a positive text sample according to the long text sample.

[0074] Step 204: When it is determined that the target text sample is a preset text sample, the first text in the target text sample is determined as the query text sample, and the second text is determined as the positive text sample corresponding to the query text sample, where the first text is a preset number of texts obtained from the start position of the text of the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to the preset text length.

[0075] Among them, the query text sample can be understood as the anchor for training the question-and-answer model; the positive text sample can be understood as the positive instance for training the question-and-answer model, and the positive text sample is related to the query text sample; the preset number can be understood as the number of text characters set in advance or the number of text characters randomly selected. This specification does not limit this, but the preset number is less than the text length of the target text document. For example, when the text length of the target text document is 20 characters, the preset number can be 5 characters or 10 characters.

[0076] For example, it is determined that the preset text length is 10 characters; when the text length of the target text sample is 20 characters, that is, when the text length of the target text sample is greater than the preset text length, the target text sample can be determined as a long text sample. In the case of determining the target text sample as a long text sample, taking the preset quantity as 5 characters as an example, the text corresponding to the 5 text characters obtained from the start position of the text in the target text sample is determined as the query text sample, and the text corresponding to the remaining 15 text characters in the target text sample is determined as the positive text sample corresponding to the query text sample.

[0077] In one or more embodiments of this specification, after determining the target text sample, in the case of determining that the target text sample is the preset text sample, another method can be used to input the target text model into the question-and-answer model, perform two different masking processes in the question-and-answer model, obtain two different vectors, and use one of the vectors as the query text sample vector corresponding to the query text sample, and the other vector as the positive text sample vector corresponding to the positive text sample. The specific implementation method is as follows:

[0078] After obtaining the initial text sample set and sequentially determining each initial text sample in the initial text sample set as the target text sample, it further includes:

[0079] In the case of determining that the target text sample is the preset text sample, input the target text sample into the question-and-answer model;

[0080] Perform a masking process on the target text sample in the question-and-answer model to obtain a first text mask vector and a second text mask vector;

[0081] Determine the first text mask vector as the query text sample vector and the second text mask vector as the positive text sample vector.

[0082] Among them, the mask can be understood as a vector that randomly sets certain positions in the vector corresponding to the text to 0. Randomly using the mask twice can obtain different mask vectors, such as the dropout mask.

[0083] For example, in the case of determining that the target text sample is a long text sample, input the target text sample into the question-and-answer model, use two different dropout masks in the question-and-answer model, and by randomly setting different positions in the vector corresponding to the input target text sample to 0, obtain two different vectors, and determine one of the vectors as the query text sample vector and the other vector as the positive text sample vector.

[0084] A question-and-answer model training method provided by an embodiment of this specification. When it is determined that the target text sample is a preset text sample, two methods can be randomly selected to determine the anchor point and the positive text sample. By these two methods, the anchor point and the positive text sample are constructed, the training samples for training the question-and-answer model are increased, data augmentation is achieved, and larger-scale and more diverse training data that does not require data annotation is obtained.

[0085] In one or more embodiments of this specification, after the target text sample is determined, when the target text sample is not a preset text sample, the target text model is input into the question-and-answer model, and two different masking processes are performed in the question-and-answer model to obtain two different vectors. One of the vectors is used as the vector corresponding to the anchor point, and the other vector is used as the positive text sample vector. The specific implementation method is as follows:

[0086] After each initial text sample in the initial text sample set is sequentially determined as the target text sample, it further includes:

[0087] When it is determined that the target text sample is not the preset text sample, the target text sample is input into the question-and-answer model;

[0088] The target text sample is masked in the question-and-answer model to obtain a first text mask vector and a second text mask vector;

[0089] The first text mask vector is determined as the query text sample vector, and the second text mask vector is determined as the positive text sample vector.

[0090] Specifically, when the target text sample is not a preset text sample, that is, not a long text sample, the target text sample can be input into the question-and-answer model and processed through masking. The specific implementation method is similar to the above embodiment and will not be elaborated here.

[0091] A question-and-answer model training method provided by an embodiment of this specification can also construct the anchor point and the positive text sample when the target text sample is not a preset text sample. That is to say, for any target text sample, the anchor point and the positive text sample can be constructed, the training samples for training the question-and-answer model are increased, and data augmentation is achieved.

[0092] Step 206: Determine any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample.

[0093] Among them, the first negative text sample can be understood as an ordinary negative text sample (negative instance) that has no association relationship with the corresponding query text sample.

[0094] Specifically, after determining the target text sample, any other initial text sample in the initial text sample set except the target text sample is used as the ordinary negative text sample of the query text sample.

[0095] Step 208: Train a question-and-answer model based on the query text sample, the positive text sample, and the first negative text sample.

[0096] Specifically, the query text sample, the positive text sample, and the first negative text sample are used as training samples to train a question-and-answer model. The question-and-answer model is trained by narrowing the distance between the query text sample and the positive text sample and expanding the distance between the query text sample and the negative text sample to achieve a better training effect.

[0097] The obtained question-and-answer model can use any convex optimization algorithm (such as gradient descent method, Adam) to achieve end-to-end (starting from the original data input for training directly and then directly outputting the final result) training.

[0098] In one or more embodiments of this specification, after determining the first negative text sample, determine the first negative text sample vector corresponding to the first negative text sample, and train a question-and-answer model based on the first negative text sample vector, the query text sample vector obtained by the above masking processing method, and the positive text sample vector. The specific implementation is as follows:

[0099] After determining any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample, it further includes:

[0100] Determine the first negative text sample vector corresponding to the first negative text sample;

[0101] Train a question-and-answer model based on the query text sample vector, the positive text sample vector, and the first negative text sample vector.

[0102] In practical applications, the triplet loss can be calculated by the following formula:

[0103]

[0104] where E Q represents the query text sample vector, E Q - represents the first negative text sample vector, E Q + represents the positive text sample vector; cos(·) represents the cosine similarity. The greater the cosine similarity, the more similar the text samples are, [·] +denotes the rectified linear unit function, that is, only the values greater than 0 are taken. m represents the error margin, which is a preset constant. Setting m can prevent the problem of overfitting.

[0105] Specifically, the question - answering model is trained through triplet loss. Calculate the cosine similarity between the query text sample and the positive text sample, and calculate the cosine similarity between the query text sample and the first negative text sample. When the difference between the two cosine similarities is greater than m, the value of the overall Triplet Loss formula is negative, and the model training stops. At this time, it can be considered that the question - answering model has been trained.

[0106] A method for training a question - answering model provided by an embodiment of this specification. By determining the first negative text sample vector corresponding to the first negative text sample, and training the question - answering model according to the query text sample vector, the positive text sample vector, and the first negative text sample vector, it is not necessary to train the question - answering model through high - quality question - answer paired data, thus realizing a question - answering model in a low - supervision resource scenario.

[0107] In one or more embodiments of this specification, before model training, in addition to obtaining the first negative text sample corresponding to the query text sample, another way can be used to obtain the second negative text sample corresponding to the query text sample, so as to amplify the negative text samples corresponding to the query text sample and increase the data volume and diversity of the training samples. The specific implementation method is as follows:

[0108] Before training the question - answering model according to the query text sample, the positive text sample, and the first negative text sample, it further includes:

[0109] Determine the text entity of the target text sample, and replace the text entity with other text entities, where the other text entities are text entities of the same category as the text entity;

[0110] Determine the target text sample after entity replacement as the second negative text sample corresponding to the query text sample;

[0111] Correspondingly, after determining the target text sample after entity replacement as the second negative text sample corresponding to the query text sample, it further includes:

[0112] Train the question - answering model according to the query text sample, the positive text sample, and the second negative text sample; or

[0113] Train the question - answering model according to the query text sample, the positive text sample, the first negative text sample, and the second negative text sample.

[0114] Among them, a text entity can be understood as a product name or a term. For example, in the text "How does MySQL (a product name, a database) ensure the correct character encoding of the database", the text entities are "MySQL", "database (a term)", etc.; in the case where the text entity is "MySQL", other text entities can be "Redis (a product name, another database)".

[0115] The second negative text sample can be understood as a hard negative instance that has no association with the corresponding query text sample. A hard negative instance can be understood as a text sample that is similar to the query text sample but has no association. A hard negative instance may be considered by the text training model to be relevant to the query text sample, but in fact, this text sample has no association with the query text sample.

[0116] In practical applications, a named entity tool (such as Spacy, LTP) can be used to determine the text entities in the query text sample, replace the text entities in the query text sample with other text entities of the same category, and use the target text sample after replacing the text entities as the hard negative text sample corresponding to the query text sample.

[0117] In the case of obtaining the second negative text sample, one can choose to train a question-and-answer model based on the query text sample, the positive text sample, and the second negative text sample, or one can also choose to train a question-and-answer model based on the query text sample, the positive text sample, the first negative text sample, and the second negative text sample.

[0118] A method for training a question-and-answer model provided by an embodiment of this specification can, by determining the second negative text sample corresponding to the query text sample, achieve data augmentation for text contrast learning based on text entities. The question-and-answer model trained with the second negative text sample can more accurately judge the key information in the text sample and distinguish it.

[0119] In one or more embodiments of this specification, after obtaining the second negative text sample, by determining the query text sample vector corresponding to the query text sample, the positive text sample vector corresponding to the positive text sample, and the second negative text sample vector corresponding to the second negative text sample, a question-and-answer model is trained based on the above text sample vectors. The specific implementation method is as follows:

[0120] Training a question-and-answer model according to the query text sample, the positive text sample, and the second negative text sample includes:

[0121] Determining the query text sample vector corresponding to the query text sample, the positive text sample vector corresponding to the positive text sample, and the second negative text sample vector corresponding to the second negative text sample;

[0122] Train a question - answering model based on the query text sample vector, the positive text sample vector, and the second negative text sample vector;

[0123] Correspondingly, training the question - answering model according to the query text sample, the positive text sample, the first negative text sample, and the second negative text sample includes:

[0124] Determine the query text sample vector corresponding to the query text sample, the positive text sample vector corresponding to the positive text sample, the first negative text sample vector corresponding to the first negative text sample, and the second negative text sample vector corresponding to the second negative text sample;

[0125] Train a question - answering model based on the query text sample vector, the positive text sample vector, the first negative text sample vector, and the second negative text sample vector.

[0126] The specific implementation method is similar to that in the above - mentioned embodiment of training a question - answering model according to the query text sample vector, the positive text sample vector, and the first negative text sample vector, and will not be elaborated here.

[0127] A method for training a question - answering model provided in an embodiment of this specification is obtained by replacing text entities in a query text sample based on a second negative text sample, and training a question - answering model through a query text sample vector, a positive text sample vector, and a second negative text sample vector, or training a question - answering model through a query text sample vector, a positive text sample vector, a first negative text sample vector, and a second negative text sample vector, which can improve the accuracy of the question - answering model.

[0128] In one or more embodiments of this specification, by extracting query text entities in a query text sample and setting corresponding entity tags for each character in the query text sample, input the query text sample and the entity tags corresponding to each character in the query text sample into the question - answering model for vector processing to determine the query text sample vector corresponding to the query text sample. The specific implementation method is as follows:

[0129] The determination of the query text sample vector corresponding to the query text sample includes:

[0130] Extract the query text entities in the query text sample, and set corresponding entity tags for each character in the query text sample according to the query text entities;

[0131] Input the query text sample and the entity tags corresponding to each character in the query text sample into the question - answering model for vector processing to obtain the word vectors corresponding to each character and carrying entity tags;

[0132] Determine a query text sample vector corresponding to the query text sample according to the word vectors corresponding to the respective characters and carrying entity tags.

[0133] Among them, a query text entity can be understood as a product name or term included in the query text sample; an entity tag can be understood as an entity mark determined for each text entity, and the text entity can be classified through this mark; after determining the entity mark for each text entity, set the corresponding entity tag for each character in the text entity.

[0134] In practical applications, use a named entity tool to extract query text entities in the query text sample, and provide an entity tag for each character in the query text sample according to the entity extraction tool, which can be represented by [0-T], where 0 represents that a character does not belong to any entity, the character belonging to the first type of entity is marked as 1, the character belonging to the second type of entity is marked as 2, and so on. T represents the number of types of text entities involved.

[0135] Input the query text sample and the entity tags corresponding to the respective characters in the query text sample into the question-answering model, and the input sequence is:

[0136] [CLS], q1, q2,..., q D , [SEP]

[0137] Among them, q represents each character in the query text sample, D represents the text length of the query text sample, [CLS] and [SEP] are special symbols used by the question-answering model to represent the beginning and end of the input. Expand the number of types of the original text entities to T + 1 in the way of the embedding vector corresponding to the word (token) type in the question-answering model input, and use the entity tags corresponding to each character

[0138]

[0139] as the serial number of the type of the word to input, and the output is expressed as:

[0140] h represents a vector, that is, obtain word vectors corresponding to the respective characters and carrying entity tags; input this output into a layer of attention pooling layer to obtain a vector h Q , the pooling layer is to convert the above multiple word vectors into one vector; input the vector h Q into a layer of feed-forward network to obtain the vector representation of the final query text sample, that is, the query text sample vector, denoted as E Q .

[0141] Alternatively, the query text sample can be first input into the Q&A model, and then the named entity tool can be called to extract the query text entities in the query text sample. This specification does not make any limitations in this regard.

[0142] A Q&A model training method provided by an embodiment of this specification extracts query text entities in a query text sample, sets corresponding entity labels for each character in the query text sample, inputs the query text sample and the entity labels corresponding to each character in the query text sample into the Q&A model for vector processing, determines a query text sample vector corresponding to the query text sample. The query text sample vector carries text entity label information, and based on the text entities, the accuracy of the Q&A model can be improved.

[0143] In one or more embodiments of this specification, by extracting positive text entities in a positive text sample, setting corresponding entity labels for each character in the positive text sample, and inputting the positive text sample and the entity labels corresponding to each character in the positive text sample into the Q&A model for vector processing, a positive text sample vector corresponding to the positive text sample is determined. The specific implementation manner is as described below:

[0144] Determining the positive text sample vector corresponding to the positive text sample includes:

[0145] Extracting positive text entities in the positive text sample, and setting corresponding entity labels for each character in the positive text sample according to the positive text entities;

[0146] Inputting the positive text sample and the entity labels corresponding to each character in the positive text sample into the Q&A model for vector processing to obtain word vectors corresponding to each character and carrying entity labels;

[0147] Determining the positive text sample vector corresponding to the positive text sample according to the word vectors corresponding to each character and carrying entity labels.

[0148] The specific implementation is similar to the above embodiment and will not be elaborated here.

[0149] In one or more embodiments of this specification, by extracting first negative text entities in a first negative text sample, setting corresponding entity labels for each character in the first negative text sample, and inputting the first negative text sample and the entity labels corresponding to each character in the first negative text sample into the Q&A model for vector processing, a first negative text sample vector corresponding to the first negative text sample is determined. The specific implementation manner is as described below:

[0150] Determining the first negative text sample vector corresponding to the first negative text sample includes:

[0151] Extract the first negative text entity in the first negative text sample, and set corresponding entity labels for each character in the first negative text sample according to the first negative text entity;

[0152] Input the first negative text sample and the entity labels corresponding to each character in the first negative text sample into a question-answering model for vector processing to obtain word vectors carrying entity labels corresponding to each character;

[0153] Determine the first negative text sample vector corresponding to the first negative text sample according to the word vectors carrying entity labels corresponding to each character.

[0154] The specific implementation is similar to the above embodiment and will not be elaborated here.

[0155] In one or more embodiments of this specification, by extracting the second negative text entity in the second negative text sample, setting corresponding entity labels for each character in the second negative text sample, inputting the second negative text sample and the entity labels corresponding to each character in the second negative text sample into a question-answering model for vector processing, and determining the second negative text sample vector corresponding to the second negative text sample. The specific implementation manner is as follows:

[0156] The determination of the second negative text sample vector corresponding to the second negative text sample includes:

[0157] Extract the second negative text entity in the second negative text sample, and set corresponding entity labels for each character in the second negative text sample according to the second negative text entity;

[0158] Input the second negative text sample and the entity labels corresponding to each character in the second negative text sample into a question-answering model for vector processing to obtain word vectors carrying entity labels corresponding to each character;

[0159] Determine the second negative text sample vector corresponding to the second negative text sample according to the word vectors carrying entity labels corresponding to each character.

[0160] The specific implementation is similar to the above embodiment and will not be elaborated here.

[0161] A question-answering model training method provided in an embodiment of this specification can construct positive text samples based on query text samples in two ways, construct a first negative text sample and a second negative text sample based on the query text samples, realizing data augmentation, and can implement a data augmentation method based on entities. The trained question-answering model can better identify key information and distinguish it. Therefore, compared with other contrastive learning algorithms, the question-answering model training method provided in the embodiment of this specification can achieve a higher accuracy rate. For example, using a relatively good unsupervised contrastive learning algorithm currently can achieve an accuracy rate of 62%, while the question-answering model training method provided in the embodiment of this specification can achieve an accuracy rate of 84%.

[0162] See Figure 3 , Figure 3 which shows a schematic diagram of the model training process of a question-answering model training method provided in an embodiment of this specification.

[0163] It includes a PM (PretrainedModel, pre-trained model) 302. Among them, the pre-trained model can be understood as a model using BERT as the basic structure, or other pre-trained word vectors such as word2vec (a method of representing words as vectors), RoBERTa (a variant of the BERT model), etc. as the basic structure of the model.

[0164] Taking the help document set of a certain database product as an example of the initial text sample set, the question-answering model training method will be described in detail.

[0165] Specifically, during the model training process, traverse the help document set, and regard each document in the help document set as a document to be trained (i.e., the target text sample in the above embodiment). For example, it is determined that the document to be trained is "This article mainly introduces how RDS MySQL ensures the correct database character encoding. Details...".

[0166] In the case where the document to be trained is a long text, "how MySQL ensures the correct database character encoding" in the document to be trained can be determined as the anchor point 304, and the document to be trained is determined as the positive sample 320; regard any other piece of text such as "What is the difference between security auditing and SQL insight?" as the first negative sample 306, and any other piece of text such as "What is the impact of changing the configuration?" as the second negative sample 308; replace the entity "MySQL" in the anchor point 304 with the same category "Redis", and determine "how Redis ensures the correct database character encoding" after replacement as the hard negative sample 322.

[0167] The PM model needs to be trained using the anchor point 304, positive sample 320, first negative sample 306, second negative sample 308, and hard negative sample 322 as training samples.

[0168] Among them, the anchor point 304 can be understood as the query text sample in the above embodiments; inputting the anchor point 304 into the PM, the corresponding anchor point vector 310 of the anchor point 304 is obtained, and the anchor point vector 310 can be understood as the query text sample vector in the above embodiments;

[0169] The positive sample 320 can be understood as the positive text sample in the above embodiments; inputting the positive sample 320 into the PM, the corresponding positive sample vector 316 of the positive sample 320 is obtained, and the positive sample vector 316 can be understood as the positive text sample vector in the above embodiments;

[0170] The first negative sample 306 can be understood as the first negative text sample in the above embodiments; inputting the first negative sample 306 into the PM, the corresponding first negative sample vector 312 of the first negative sample 306 is obtained, and the first negative sample vector 312 can be understood as the first negative text sample vector in the above embodiments;

[0171] The second negative sample 308 can be understood as the first negative text sample in the above embodiments; inputting the second negative sample 308 into the PM, the corresponding second negative sample vector 314 of the second negative sample 308 is obtained, and the second negative sample vector 314 can be understood as the first negative text sample vector in the above embodiments;

[0172] The hard negative sample 322 can be understood as the second negative text sample in the above embodiments; inputting the hard negative sample 322 into the PM, the corresponding hard negative sample vector 318 of the hard negative sample 322 is obtained, and the hard negative sample vector 318 can be understood as the second negative text sample vector in the above embodiments.

[0173] Calculate the triplet loss of the obtained positive sample vector 316, first negative sample vector 312, second negative sample vector 314 and anchor point vector 310 to complete the training of the question-answering model.

[0174] The question-answering model training method provided in the embodiments of this specification realizes data augmentation by constructing positive text samples based on query text samples, constructing first negative text samples and second negative text samples based on query text samples. Without the need for high-quality data pairs, the model training can be completed, realizing a question-answering model and training target in a scenario with almost no supervised resources. Moreover, based on the entity-based data augmentation method, the trained question-answering model can better identify key information and distinguish it.

[0175] See Figure 4 , Figure 4The flowchart showing a text Q&A method provided by an embodiment of this specification specifically includes the following steps.

[0176] Step 402: Determine the question query text and the set of answer texts corresponding to the question query text.

[0177] Among them, the question query text can be understood as the text input by the user for a certain question; the set of answer texts can be understood as the set of answer texts corresponding to the text input by the user for a certain question.

[0178] In practical applications, when training the question text model above, since there was no question provided by the user input at that time, the answer texts in the set of answer texts were used for model training. Therefore, the set of answer texts can be understood as the initial text set in the above embodiment.

[0179] Step 404: Input the question query text and the set of answer texts into the Q&A model to obtain the target answer corresponding to the question query text, where the Q&A model is trained by the Q&A model training method described above.

[0180] Among them, the target answer can be understood as the answer text that matches the question query text.

[0181] For example, when a user encounters a problem while using a certain product and enters the problem through the search window, the question entered by the user and the set of help documents corresponding to the product prepared in advance are input into the Q&A model to obtain the answer corresponding to the question entered by the user.

[0182] In one or more embodiments of this specification, the set of answer texts may include answer texts. The question query text and each answer text in the set of answer texts are input into the Q&A model to obtain the question query text vector corresponding to the question query text and the answer text vectors corresponding to each answer text. The target answer corresponding to the question query text is obtained through the similarity between the two vectors. The specific implementation method is as follows:

[0183] The set of answer texts includes answer texts;

[0184] Correspondingly, the step of inputting the question query text and the set of answer texts into the Q&A model to obtain the target answer corresponding to the question query text includes:

[0185] Input the question query text and each answer text in the set of answer texts into the Q&A model to obtain the question query text vector corresponding to the question query text and the answer text vectors corresponding to each answer text;

[0186] Query the similarity between the text vector corresponding to the question and the answer text vectors corresponding to the respective answer texts, and obtain the target answer corresponding to the question query text.

[0187] Among them, the similarity can be understood as the degree of similarity between the question query text and each answer text.

[0188] Specifically, for the implementation manner of inputting the text into the question-answering model to obtain the corresponding vector, reference can be made to the above-mentioned embodiments, and details will not be elaborated here.

[0189] After obtaining the question query text vector and the answer text vectors corresponding to the respective answer texts, according to the matching relationship between the question query text vector and the answer text vectors corresponding to the respective answer texts, obtain the answer text that best matches the question query text, and output this answer text as the target answer.

[0190] A text question-answering method provided by an embodiment of this specification obtains a question-answering model through the above training, outputs the question query text and each answer text in the answer text set as vectors, and determines the answer text that best matches the question query text by calculating the matching degree of the two vectors.

[0191] In one or more embodiments of this specification, before determining the question query text, it is also possible to determine the answer text set, input each answer text in the answer text set into the question-answering model, and obtain the answer text vectors corresponding to the respective answer texts, so as to pre-obtain the answer text vectors corresponding to the respective answer texts before determining the question query text. The specific implementation manner is as follows:

[0192] Before determining the question query text and the answer text set corresponding to the question query text, it further includes:

[0193] Determine the answer text set, and input each answer text in the answer text set into the question-answering model to obtain the answer text vectors corresponding to the respective answer texts.

[0194] For the specific implementation, reference can be made to the above-mentioned embodiments, and details will not be elaborated here.

[0195] In one or more embodiments of this specification, in the case of pre-obtaining the answer text vectors corresponding to the respective answer texts, the answer text set includes the answer text vectors corresponding to the answer texts. Input the question query text and the answer text vectors corresponding to the respective answer texts in the answer text set into the question-answering model. At this time, the question-answering model only needs to calculate the question query text vector corresponding to the question query text, which improves the service rate. The specific implementation manner is as follows:

[0196] The answer text set includes the answer text vectors corresponding to the answer texts;

[0197] Correspondingly, inputting the question query text and the set of answer texts into the question-answering model to obtain the target answer corresponding to the question query text includes:

[0198] Inputting the question query text and the answer text vectors corresponding to each answer text in the set of answer texts into the question-answering model to obtain the question query text vector corresponding to the question query text;

[0199] Obtaining the target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to each answer text.

[0200] For the specific implementation, reference can be made to the above embodiments. However, at this time, through the question-answering model, only the question query text needs to be output as the corresponding question query text vector, which will not be elaborated here.

[0201] A text question-answering method provided in an embodiment of this specification, by pre-obtaining the answer text vectors corresponding to each answer text, when determining the question query text input by the user, only needs to output the question query text as the corresponding question query text vector through the question-answering model, and independently calculates the representation vectors of the question query text input by the user and the answer text respectively, improving the service efficiency.

[0202] In one or more embodiments of this specification, the target answer corresponding to the question query text can be determined by calculating the similarity between the question query text vector and the answer text vectors corresponding to each answer text. The specific implementation method is as follows:

[0203] The obtaining of the target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to each answer text includes:

[0204] Calculating the similarity between the query text vector and the answer text vectors corresponding to each answer text;

[0205] Determining the target answer corresponding to the question query text based on the similarity.

[0206] Among them, the calculation of the similarity can be performed using the cosine similarity method, or other methods, such as Pearson similarity, dynamic time warping, Hamming distance, Euclidean distance, etc. This specification does not make a limitation on this.

[0207] Specifically, taking the calculation of the cosine similarity between the query text vector and the answer text vectors corresponding to each answer text as an example, based on the calculated cosine similarity value, the greater the cosine similarity, the more similar the question query text is to the corresponding answer text, that is, the higher the matching degree. The answer text with the largest cosine similarity to the question query text among each answer text is determined as the target answer.

[0208] A text Q&A method provided by an embodiment of this specification can quickly determine a target answer corresponding to a question query text based on similarity by calculating the similarity between a query text vector and answer text vectors corresponding to each answer text.

[0209] See Figure 5 , Figure 5 which shows a process flow chart of a text Q&A method provided by an embodiment of this specification.

[0210] In the case of applying a Q&A model, the question input by the user and each answer in a pre-prepared answer set are encoded into vectors of a fixed dimension, the question vector corresponding to the question and the answer vector corresponding to the answer are output, and the answer with the highest similarity to the question is retrieved by calculating the similarity between the question vector and the answer vector. To improve the calculation efficiency, the calculation processes of the question and the answer are independent of each other, but both use the same set of model parameters and have exactly the same calculation process.

[0211] Step 1: Input question 502 and answer 504 into a pre-trained model 506.

[0212] Among them, question 502 can be understood as the question query text in the above embodiment, answer 504 can be understood as the answer text set in the above embodiment, and pre-trained model 506 is the PM in the above embodiment. Taking answer 504 as the answer text set in the above embodiment, where the answer text set includes answer text vectors, the process of applying the Q&A model will be described in detail.

[0213] In actual application, input question 502 and answer 504 into the Q&A model. Among them, the input sequence of question 502 is:

[0214] [CLS], q1, q2,..., q D , [SEP]

[0215] Among them, q represents each character of question 502, D represents the text length of question 502, and [CLS] and [SEP] are special symbols used by the Q&A model to represent the beginning and end of the input.

[0216] Input the answer text vectors in answer 504 into the pre-trained model 506.

[0217] Step 2: In the pre-trained model 506, output word vectors corresponding to each character of question 502.

[0218] Use a named entity tool to extract entities in the text of input question 502. For details, see the above embodiment and will not be elaborated here.

[0219] Question 502 passes through the pre-trained model 506, and the final output is expressed as:

[0220]

[0221] h represents a vector, that is, the word vectors corresponding to each character of question 502 are obtained, and the pre-trained model 506 does not process the answer text vector in answer 504.

[0222] Step three: Input the word vectors corresponding to each character of question 502 into the attention pooling layer 508.

[0223] Input the word vectors corresponding to each character of question 502 obtained through the pre-trained model 506 into a layer of attention pooling layer 508 to obtain an intermediate vector h corresponding to question 502 Q , h Q is a vector converted from the word vectors. The attention pooling layer 508 is used to convert the output of the pre-trained model 506 into a fixed-length vector.

[0224] The attention pooling layer 508 does not process the answer text vector in answer 504.

[0225] Step four: Input an intermediate vector h corresponding to question 502 Q , into the feed-forward network.

[0226] Input the h obtained by the above attention pooling layer 508 Q , and the answer text vector of the same fixed length as h Q into a layer of feed-forward network to obtain the vector representation of the final question 502, that is, the question vector 510, and the vector representation of the final answer 504, that is, the answer vector 512.

[0227] Among them, the question vector 510 can be understood as the question query text vector in the above embodiment, and the answer vector 512 can be understood as the answer text vector in the above embodiment.

[0228] Step five: Calculate the cosine similarity between the question vector 510 and the answer vector 512.

[0229] After obtaining the question vector 510 and the answer vector 512, compare the similarity between the question vector 510 and the answer vector 512 by calculating the cosine similarity between the question vector 510 and the answer vector 512, and use the answer corresponding to the answer vector 512 that is most similar to the question vector 510 as the target answer.

[0230] Through a text Q&A method provided by the embodiments of this specification, a question and an answer set are input into a Q&A model, corresponding question vectors and answer vectors are output in the Q&A model, the similarity between the question vectors and the answer vectors is calculated, and a target answer is obtained. The answer set may include answer vectors obtained in advance through the Q&A model, thereby achieving only calculating the question vectors corresponding to questions during the service, improving the service efficiency, and achieving accurate Q&A by controlling the answer range through preparing the answer set in advance.

[0231] Corresponding to the above method embodiments, this specification also provides embodiments of a Q&A model training device. Figure 6 The structure diagram of a Q&A model training device provided by an embodiment of this specification is shown. As Figure 6 shown, the device includes:

[0232] A target text sample determination module 602, configured to obtain an initial text sample set and sequentially determine each initial text sample in the initial text sample set as a target text sample;

[0233] A text sample determination module 604, configured to, when determining that the target text sample is a preset text sample, determine the first text in the target text sample as a query text sample and the second text as the positive text sample corresponding to the query text sample, where the first text is a preset number of texts obtained from the start position of the text of the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length;

[0234] A first negative text sample determination module 606, configured to determine any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample;

[0235] A model obtaining module 608, configured to train a Q&A model according to the query text sample, the positive text sample, and the first negative text sample.

[0236] The device further includes:

[0237] A first vector determination module, configured to, when determining that the target text sample is the preset text sample, input the target text sample into the Q&A model;

[0238] Perform masking processing on the target text sample in the Q&A model to obtain a first text mask vector and a second text mask vector;

[0239] Determine the first text masked vector as the query text sample vector, and determine the second text masked vector as the positive text sample vector.

[0240] The device further includes:

[0241] A second vector determination module, configured to input the target text sample into a question-and-answer model when it is determined that the target text sample is not the preset text sample;

[0242] Perform masking processing on the target text sample in the question-and-answer model to obtain a first text masked vector and a second text masked vector;

[0243] Determine the first text masked vector as the query text sample vector, and determine the second text masked vector as the positive text sample vector.

[0244] The device further includes:

[0245] A first model acquisition module, configured to determine a first negative text sample vector corresponding to the first negative text sample;

[0246] Train a question-and-answer model based on the query text sample vector, the positive text sample vector, and the first negative text sample vector.

[0247] The device further includes:

[0248] A second model acquisition module, configured to determine the text entity of the target text sample, and replace the text entity with another text entity, where the other text entity is a text entity of the same category as the text entity;

[0249] Determine the target text sample after entity replacement as the second negative text sample corresponding to the query text sample;

[0250] Optionally, the model acquisition module 608 is further configured to:

[0251] Train a question-and-answer model based on the query text sample, the positive text sample, and the second negative text sample; or

[0252] Train a question-and-answer model based on the query text sample, the positive text sample, the first negative text sample, and the second negative text sample.

[0253] Optionally, the model acquisition module 608 is further configured to:

[0254] Determine a query text sample vector corresponding to the query text sample, a positive text sample vector corresponding to the positive text sample, and a second negative text sample vector corresponding to the second negative text sample;

[0255] Train a question-answering model based on the query text sample vector, the positive text sample vector, and the second negative text sample vector.

[0256] Optionally, the model obtaining module 608 is further configured to:

[0257] Determine the query text sample vector corresponding to the query text sample, the positive text sample vector corresponding to the positive text sample, the first negative text sample vector corresponding to the first negative text sample, and the second negative text sample vector corresponding to the second negative text sample.

[0258] Train a question-answering model based on the query text sample vector, the positive text sample vector, the first negative text sample vector, and the second negative text sample vector.

[0259] Optionally, the model obtaining module 608 is further configured to:

[0260] Extract the query text entities from the query text sample, and set corresponding entity labels for each character in the query text sample according to the query text entities.

[0261] Input the query text sample and the entity labels corresponding to each character in the query text sample into the question-answering model for vector processing to obtain word vectors corresponding to each character and carrying entity labels.

[0262] Determine the query text sample vector corresponding to the query text sample according to the word vectors corresponding to each character and carrying entity labels.

[0263] Optionally, the model obtaining module 608 is further configured to:

[0264] Extract the positive text entities from the positive text sample, and set corresponding entity labels for each character in the positive text sample according to the positive text entities.

[0265] Input the positive text sample and the entity labels corresponding to each character in the positive text sample into the question-answering model for vector processing to obtain word vectors corresponding to each character and carrying entity labels.

[0266] Determine the positive text sample vector corresponding to the positive text sample according to the word vectors corresponding to each character and carrying entity labels.

[0267] Optionally, the model obtaining module 608 is further configured to:

[0268] Extract the first negative text entities in the first negative text sample, and set corresponding entity labels for each character in the first negative text sample according to the first negative text entities;

[0269] Input the first negative text sample and the entity labels corresponding to each character in the first negative text sample into the question - answering model for vector processing to obtain word vectors corresponding to each character and carrying entity labels;

[0270] Determine the first negative text sample vector corresponding to the first negative text sample according to the word vectors corresponding to each character and carrying entity labels.

[0271] Optionally, the model obtaining module 608 is further configured to:

[0272] Extract the second negative text entities in the second negative text sample, and set corresponding entity labels for each character in the second negative text sample according to the second negative text entities;

[0273] Input the second negative text sample and the entity labels corresponding to each character in the second negative text sample into the question - answering model for vector processing to obtain word vectors corresponding to each character and carrying entity labels;

[0274] Determine the second negative text sample vector corresponding to the second negative text sample according to the word vectors corresponding to each character and carrying entity labels.

[0275] The device further includes:

[0276] A preset text determining module, configured to determine the text length of the target text sample, and determine that the target text sample is a preset text sample when the text length is greater than or equal to a preset text length.

[0277] Corresponding to the above - mentioned method embodiment, this specification also provides an embodiment of a text question - answering device, Figure 7 showing a structural schematic diagram of a text question - answering device provided in an embodiment of this specification. As Figure 7 shown, the device includes:

[0278] A text determining module 702, configured to determine a question query text and a set of answer texts corresponding to the question query text;

[0279] An answer obtaining module 704, configured to input the question query text and the set of answer texts into the question - answering model to obtain a target answer corresponding to the question query text, where the question - answering model is trained by the above - mentioned question - answering model training method.

[0280] Optionally, the answer obtaining module 704 is further configured to:

[0281] Input the question query text and each answer text in the answer text set into a question-answering model to obtain a question query text vector corresponding to the question query text and answer text vectors corresponding to the respective answer texts;

[0282] Obtain a target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to the respective answer texts.

[0283] Optionally, the answer obtaining module 704 is further configured to:

[0284] Input the question query text and the answer text vectors corresponding to the respective answer texts in the answer text set into a question-answering model to obtain a question query text vector corresponding to the question query text;

[0285] Obtain a target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to the respective answer texts.

[0286] Optionally, the answer obtaining module 704 is further configured to:

[0287] Calculate the similarity between the query text vector and the answer text vectors corresponding to the respective answer texts;

[0288] Determine a target answer corresponding to the question query text based on the similarity.

[0289] The apparatus further includes:

[0290] An answer text vector obtaining module, configured to determine the answer text set and input each answer text in the answer text set into a question-answering model to obtain answer text vectors corresponding to the respective answer texts.

[0291] Optionally, the answer obtaining module 704 is further configured to:

[0292] Extract question query text entities in the question query text and set corresponding entity tags for each character in the question query text according to the question query text entities;

[0293] Input the question query text and the entity tags corresponding to each character in the question query text into a question-answering model for vector processing to obtain word vectors corresponding to the respective characters and carrying entity tags.

[0294] Determine a question query text vector corresponding to the question query text according to the word vectors corresponding to the respective characters and carrying entity tags; and

[0295] Extract answer text entities in each answer text in the answer text set, and set corresponding entity tags for each character in each answer text according to the answer text entities;

[0296] Input each answer text in the answer text set and the entity tags corresponding to each character in each answer text into a question and answer model for vector processing to obtain word vectors corresponding to the respective characters and carrying entity tags;

[0297] Determine an answer text vector corresponding to each answer text according to the word vectors corresponding to the respective characters and carrying entity tags.

[0298] The above is a schematic solution of a question and answer model training device according to this embodiment. It should be noted that the technical solution of this question and answer model training device and the technical solution of the above question and answer model training method belong to the same concept. For the details not described in detail in the technical solution of the question and answer model training device, reference can be made to the description of the technical solution of the above question and answer model training method.

[0299] Figure 8 FIG. shows a structural block diagram of a computing device 800 according to an embodiment of the present specification. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 through a bus 830, and a database 850 is used to store data.

[0300] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0301] In one embodiment of the present specification, the above components of the computing device 800, as well as Figure 8 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 8 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0302] The computing device 800 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 800 can also be a mobile or stationary server.

[0303] Among them, the processor 820 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above question-and-answer model training method are implemented.

[0304] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above question-and-answer model training method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above question-and-answer model training method.

[0305] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above question-and-answer model training method.

[0306] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above question-and-answer model training method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above question-and-answer model training method.

[0307] An embodiment of this specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above question-and-answer model training method.

[0308] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above question-and-answer model training method belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above question-and-answer model training method.

[0309] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0310] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0311] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0312] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0313] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A method for training a question-answering model, comprising: Obtaining an initial text sample set, and sequentially determining each initial text sample in the initial text sample set as a target text sample; When determining that the target text sample is a preset text sample, determining the first text in the target text sample as a query text sample, and the second text as the positive text sample corresponding to the query text sample, wherein the first text is a preset number of texts obtained from the start position of the text of the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length; Determining any other initial text sample in the initial text sample set except the target text sample as the first negative text sample corresponding to the query text sample; Training a question-answering model according to the query text sample, the positive text sample, and the first negative text sample.

2. The method for training a question-answering model according to claim 1, after obtaining the initial text sample set and sequentially determining each initial text sample in the initial text sample set as a target text sample, further comprising: When determining that the target text sample is the preset text sample, inputting the target text sample into the question-answering model; Performing a masking process on the target text sample in the question-answering model to obtain a first text mask vector and a second text mask vector; Determining the first text mask vector as a query text sample vector, and the second text mask vector as a positive text sample vector.

3. The method for training a question-answering model according to claim 1, after obtaining the initial text sample set and sequentially determining each initial text sample in the initial text sample set as a target text sample, further comprising: When determining that the target text sample is not the preset text sample, inputting the target text sample into the question-answering model; Performing a masking process on the target text sample in the question-answering model to obtain a first text mask vector and a second text mask vector; Determining the first text mask vector as a query text sample vector, and the second text mask vector as a positive text sample vector.

4. The method for training a question-answering model according to any one of claims 1-3, before training a question-answering model according to the query text sample, the positive text sample, and the first negative text sample, further comprising: Determining the text entity of the target text sample, and replacing the text entity with another text entity, wherein the other text entity is a text entity of the same category as the text entity; Determining the target text sample after entity replacement as the second negative text sample corresponding to the query text sample; Correspondingly, after determining the target text sample after entity replacement as the second negative text sample corresponding to the query text sample, further comprising: Training a question-answering model according to the query text sample, the positive text sample, and the second negative text sample; or A question-and-answer model is trained based on the query text sample, the positive text sample, the first negative text sample, and the second negative text sample.

5. The method for training a question-and-answer model according to claim 4, wherein the training of the question-and-answer model based on the query text sample, the positive text sample, and the second negative text sample includes: Determining a query text sample vector corresponding to the query text sample, a positive text sample vector corresponding to the positive text sample, and a second negative text sample vector corresponding to the second negative text sample; Training a question-and-answer model based on the query text sample vector, the positive text sample vector, and the second negative text sample vector; The training of the question-and-answer model based on the query text sample, the positive text sample, the first negative text sample, and the second negative text sample includes: Determining a query text sample vector corresponding to the query text sample, a positive text sample vector corresponding to the positive text sample, a first negative text sample vector corresponding to the first negative text sample, and a second negative text sample vector corresponding to the second negative text sample; Training a question-and-answer model based on the query text sample vector, the positive text sample vector, the first negative text sample vector, and the second negative text sample vector.

6. The method for training a question-and-answer model according to claim 5, wherein the determination of the query text sample vector corresponding to the query text sample includes: Extracting query text entities from the query text sample, and setting corresponding entity labels for each character in the query text sample according to the query text entities; Inputting the query text sample and the entity labels corresponding to each character in the query text sample into a question-and-answer model for vector processing to obtain word vectors corresponding to each character and carrying entity labels; Determining a query text sample vector corresponding to the query text sample according to the word vectors corresponding to each character and carrying entity labels.

7. A text question-and-answer method, including: Determining a question query text and a set of answer texts corresponding to the question query text; Inputting the question query text and the set of answer texts into a question-and-answer model to obtain a target answer corresponding to the question query text, wherein the question-and-answer model is trained by the method for training a question-and-answer model according to any one of claims 1-6.

8. The text question-and-answer method according to claim 7, wherein the set of answer texts includes answer texts; The inputting of the question query text and the set of answer texts into a question-and-answer model to obtain a target answer corresponding to the question query text includes: Inputting the question query text and each answer text in the set of answer texts into a question-and-answer model to obtain a question query text vector corresponding to the question query text and answer text vectors corresponding to each answer text; Obtaining a target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to each answer text.

9. The text Q&A method according to claim 7, wherein the answer text set includes answer text vectors corresponding to the answer texts; The step of inputting the question query text and the answer text set into a Q&A model to obtain the target answer corresponding to the question query text includes: Inputting the question query text and the answer text vectors corresponding to the respective answer texts in the answer text set into a Q&A model to obtain a question query text vector corresponding to the question query text; Obtaining the target answer corresponding to the question query text according to the similarity between the question query text vector and the answer text vectors corresponding to the respective answer texts.

10. The text Q&A method according to claim 8, wherein the step of inputting the question query text and the respective answer texts in the answer text set into a Q&A model to obtain a question query text vector corresponding to the question query text and answer text vectors corresponding to the respective answer texts includes: Extracting question query text entities in the question query text, and setting corresponding entity labels for each character in the question query text according to the question query text entities; Inputting the question query text and the entity labels corresponding to the respective characters in the question query text into a Q&A model for vector processing to obtain word vectors corresponding to the respective characters and carrying the entity labels; Determining a question query text vector corresponding to the question query text according to the word vectors corresponding to the respective characters and carrying the entity labels; And Extracting answer text entities in the respective answer texts in the answer text set, and setting corresponding entity labels for each character in the respective answer texts according to the answer text entities; Inputting the respective answer texts in the answer text set and the entity labels corresponding to the respective characters in the respective answer texts into a Q&A model for vector processing to obtain word vectors corresponding to the respective characters and carrying the entity labels; Determining answer text vectors corresponding to the respective answer texts according to the word vectors corresponding to the respective characters and carrying the entity labels.

11. A Q&A model training device, comprising: A target text sample determination module configured to obtain an initial text sample set and sequentially determine each initial text sample in the initial text sample set as a target text sample; A text sample determination module configured to, when determining that the target text sample is a preset text sample, determine the first text in the target text sample as a query text sample and the second text as a positive text sample corresponding to the query text sample, wherein the first text is a preset number of texts obtained from the start position of the text of the target text sample, the second text is the other text in the target text sample except the first text, and the preset text sample is a text sample with a text length greater than or equal to a preset text length; A first negative text sample determination module configured to determine any other initial text sample in the initial text sample set except the target text sample as a first negative text sample corresponding to the query text sample; A model acquisition module, configured to train and obtain a question-and-answer model according to the query text sample, the positive text sample, and the first negative text sample.

12. A text question-and-answer device, comprising: A text determination module, configured to determine a question query text and a set of answer texts corresponding to the question query text; An answer acquisition module, configured to input the question query text and the set of answer texts into a question-and-answer model to obtain a target answer corresponding to the question query text, wherein the question-and-answer model is trained by the above-mentioned question-and-answer model training method.

13. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the question-and-answer model training method according to any one of claims 1 to 6 are implemented, or the steps of the text question-and-answer method according to any one of claims 7 to 10 are implemented.

14. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the question-and-answer model training method according to any one of claims 1 to 6 are implemented, or the steps of the text question-and-answer method according to any one of claims 7 to 10 are implemented.

Citation Information

Patent Citations

  • Query processing model generation method and device and electronic equipment

    CN114117183A

  • Semantic representation model training method and device, storage medium and computer equipment

    CN116484220A