Text detection method, training method, device, equipment, medium and program product
By combining discriminant models and generative models, text detection is used to use deep features, the problem of inaccurate detection of answers and questions in the prior art is solved, and higher detection accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202410339038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-03-22
AI Technical Summary
The prior art is difficult to accurately detect the degree of matching answers with questions, resulting in poor user experience and prone to 'error' or 'answer'.
A model training method is adopted to detect deep features such as the logical correlation between the question text and the answer text through the combination of discriminant model and the generative model. The specific steps include: using the initial model to process the sample text pairs and generate the second sample detection result; adjust the model parameters based on the loss function to obtain the trained target model; using the target model to detect the text pairs to be detected, and generate the target detection result.
It improves the detection accuracy of the model, can more accurately represent the answer matching between the answer text and the question text, and improves the user experience.
Smart Images

Figure CN118261248B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technologies such as large language models, deep learning, text processing, etc. Specifically, it relates to a text detection method, a training method, a device, a device, a medium, and a program product. Background Art
[0002] With the popularization of Internet technologies, users can publish questions through the Internet in order to obtain answers provided by other users for that question.
[0003] Due to the wide range of fields involved in the questions, phenomena such as "wrong answer" or "off-topic answer" often occur. Therefore, there is an urgent need for a method that can accurately detect the matching degree between the answer and the question to improve the user experience. Summary of the Invention
[0004] The present disclosure provides a text detection method, a training method, a device, a device, a medium, and a program product.
[0005] According to one aspect of the present disclosure, there is provided a model training method, including: in response to the confidence of the first sample detection result being less than a predetermined threshold, using an initial model to process a sample text pair and a sample label to obtain a second sample detection result; wherein, the confidence of the first sample detection result is obtained by using a first model to detect the sample text pair; the sample text pair includes a corresponding sample question text and a sample answer text, the second sample detection result includes a discrimination result for characterizing the reply matching degree between the sample answer text and the sample question text and a discrimination reason corresponding to the discrimination result; the sample label includes a discrimination result label and a discrimination reason label; the initial model is of a different type from the first model; based on a first loss function, according to the second sample detection result and the sample label, obtaining a first loss value; and based on the first loss value, adjusting the model parameters of the initial model to obtain a trained target model.
[0006] According to another aspect of the present disclosure, there is provided a text detection method, including: using a first model to detect a plurality of text pairs to be detected to obtain a first detection result and a confidence corresponding to the first detection result, the first model being a discriminative model, the plurality of text pairs to be detected including a corresponding plurality of question texts and a plurality of answer texts; according to the confidence, determining a target text pair to be detected from the plurality of text pairs to be detected, wherein the confidence of the detection result corresponding to the target text pair to be detected is less than a predetermined threshold; using a second model to detect the target text pair to be detected to obtain a second detection result, wherein the second model is the trained target model obtained by the above training method; and according to the first detection result and the second detection result, generating a target detection result, the target detection result characterizing the reply matching degree between the answer text and the corresponding question text.
[0007] According to another aspect of the present disclosure, a model training device is provided, including: a first processing module, a first loss calculation module, and a first adjustment module.
[0008] The first processing module is configured to, in response to the confidence level of the first sample detection result being less than a predetermined threshold, process the sample text pair and the sample label by using an initial model to obtain a second sample detection result; wherein, the confidence level of the first sample detection result is obtained by detecting the sample text pair by using a first model; the sample text pair includes a corresponding sample question text and a sample answer text, the second sample detection result includes a discrimination result for characterizing the reply matching degree between the sample answer text and the sample question text and a discrimination reason corresponding to the discrimination result; the sample label includes a discrimination result label and a discrimination reason label; the type of the initial model is different from that of the first model.
[0009] The first loss calculation module is configured to obtain a first loss value based on a first loss function according to the second sample detection result and the sample label.
[0010] The first adjustment module adjusts the model parameters of the initial model based on the first loss value to obtain a trained target model.
[0011] According to another aspect of the present disclosure, a text detection device is provided, including: a second detection module, a second determination module, a third detection module, and a generation module.
[0012] The second detection module is configured to detect a plurality of text pairs to be detected by using a first model to obtain a first detection result and a confidence level corresponding to the first detection result, the first model being a discriminant model, and the plurality of text pairs to be detected including a plurality of corresponding question texts and a plurality of answer texts.
[0013] The second determination module is configured to determine a target text pair to be detected from the plurality of text pairs to be detected according to the confidence level, wherein the confidence level of the detection result corresponding to the target text pair to be detected is less than a predetermined threshold.
[0014] The third detection module is configured to detect the target text pair to be detected by using a second model to obtain a second detection result, wherein the second model is the trained target model obtained by the training method described above.
[0015] The generation module is configured to generate a target detection result according to the first detection result and the second detection result, and the target detection result characterizes the reply matching degree between the answer text and the corresponding question text.
[0016] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described above.
[0018] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described above when executed by a processor.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0021] Figure 1 Schematically shows an exemplary system architecture to which the model training method or text detection method and apparatus according to the embodiments of the present disclosure can be applied;
[0022] Figure 2 Schematically shows a flowchart of the model training method according to the embodiments of the present disclosure;
[0023] Figure 3A Schematically shows a schematic diagram of model training according to the embodiments of the present disclosure;
[0024] Figure 3B Schematically shows a schematic diagram of model training according to another embodiment of the present disclosure;
[0025] Figure 4 Schematically shows a flowchart of the text detection method according to the embodiments of the present disclosure;
[0026] Figure 5 Schematically shows a schematic diagram of text detection according to the embodiments of the present disclosure;
[0027] Figure 6 Schematically shows a block diagram of the model training apparatus according to the embodiments of the present disclosure;
[0028] Figure 7 Schematically shows a block diagram of the text detection apparatus according to the embodiments of the present disclosure; and
[0029] Figure 8 A block diagram of an electronic device suitable for implementing a model training method or a text detection method according to an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] In the scenario of judging the quality of text questions and answers, the question texts in the sample data cover a wide range of fields. During the multi-round training of the discriminative model based on a large amount of sample data, it is easy to cause the model to overfit.
[0032] When the input question-and-answer text pair to be detected deviates slightly from the sample data distribution, the inference performance of the model will fluctuate significantly, resulting in a significant decrease in the accuracy of the model output result.
[0033] In addition, the models in related examples for detecting the quality of text questions and answers are based on surface features such as the language fluency of the answer text and the semantic relevance between the question and the answer, lacking the processing of deep features such as the logical relevance between the question and the answer, which also leads to low model accuracy.
[0034] Generative large language models (LLMs) are natural language processing models with logical reasoning ability pre-trained using deep learning techniques.
[0035] Therefore, the embodiments of the present disclosure provide a model training method, using the discrimination result and the discrimination reason as labels to guide the initial model to perform detection using deep features such as the logical relevance between the question text and the answer text, thereby improving the model accuracy to achieve accurate detection of sample text pairs that are difficult to detect for the first model and obtaining a quality detection result of question-and-answer text pairs with a higher confidence level.
[0036] Figure 1 An exemplary system architecture to which the model training method or the text detection method and apparatus according to the embodiments of the present disclosure can be applied is schematically shown.
[0037] It should be noted that Figure 1The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the model training method and apparatus can be applied may include a terminal device, but the terminal device can implement the model training method and apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0038] As Figure 1 shown, the system architecture 100 according to this embodiment may include a terminal device 110, a first server 150_1, and a second server 150_2. The terminal device 110 may be various electronic devices with processing functions, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and smart wearable devices, etc.
[0039] The first server 150_1 may be used to execute the training method of the first model in the model training method provided by the embodiments of the present disclosure to obtain a discriminant model 140_1. The second server 150_2 may be used to execute the training method of the target model in the model training method provided by the embodiments of the present disclosure to obtain a generative model 140_2.
[0040] The terminal device may load the trained discriminant model 140_1 and the trained generative model 140_2 to process the question text answer text 120 according to the loaded discriminant model 140_1 to obtain a preliminary detection result. For the question text answer text with a confidence level lower than a predetermined threshold in the preliminary detection result, the generative model 140_2 may be used for processing to obtain a detection result 130.
[0041] It should be noted that the model training method or text detection method provided by the embodiments of the present disclosure can generally be executed by the terminal device 110. Correspondingly, the model training apparatus or text detection apparatus provided by the embodiments of the present disclosure can also be provided in the terminal device 110.
[0042] Alternatively, the model training method or text detection method provided by the embodiments of the present disclosure can generally also be executed by the first server 150_1 and the second server 150_2. Correspondingly, the model training device or text detection device provided by the embodiments of the present disclosure can generally be arranged in the execution of the first server 150_1 and the second server 150_2. The model training method or text detection method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the first server 150_1 and the second server 150_2 and capable of communicating with the terminal device 110 and / or the first server 150_1 and the second server 150_2. Correspondingly, the model training device or text detection device provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the first server 150_1 and the second server 150_2 and capable of communicating with the terminal device 110 and / or the first server 150_1 and the second server 150_2.
[0043] It should be understood that Figure 1 the numbers of the terminal devices and servers in
[0044] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the user's personal information all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs.
[0045] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.
[0046] Figure 2 Schematically shows a flowchart of a model training method according to an embodiment of the present disclosure.
[0047] As Figure 2 shown, the method 200 includes operations S210 to S230.
[0048] In operation S210, in response to the confidence level of the first sample detection result being less than a predetermined threshold, the initial model is used to process the sample text pair and the sample label to obtain a second sample detection result.
[0049] In operation S220, based on the first loss function, according to the second sample detection result and the sample label, a first loss value is obtained.
[0050] In operation S230, based on the first loss value, the model parameters of the initial model are adjusted to obtain a trained target model.
[0051] According to an embodiment of the present disclosure, the initial model and the first model are of different types. The initial model may be a generative model, such as: Wenxin model, Large Language Model Meta AI (LlaMA), GPT (Generative Pre-trained Transformer) model, etc. The first model may be a discriminative model, such as: convolutional neural network model, Enhanced Language Representation with Informative Entities (ERNIE) model.
[0052] According to an embodiment of the present disclosure, the first model and the initial model can be trained in two stages. In the first stage, the first model can be supervised-trained using the initial sample text pairs so that the loss value between the result output by the first model and the manually labeled result label reaches the convergence condition, and the trained first model is obtained. The convergence condition may be the maximum number of training times or a predetermined loss value, and the present disclosure does not make specific limitations thereto.
[0053] After completing the first-stage training, the trained first model processes the initial sample text pairs to obtain the first sample detection results. The first sample detection results may include the discrimination results corresponding to the initial sample text pairs and the confidence levels corresponding to the discrimination results. The confidence level corresponding to the discrimination result may characterize the discrimination difficulty of the first model for the initial sample text pair.
[0054] For example: The initial sample text pairs may include multiple text pairs composed of sample question texts and sample answer texts: Q1-A1, Q2-A2, …, Q n -A n . Processing the initial sample text pairs using the first model, the obtained first sample detection results may be: Q1-A1 (I, 95%), Q2-A2 (II, 55%), …, Q n -A n (III, 75%). Among them, I may indicate that the answer text A1 does not solve the question text Q1, II may indicate that the answer text A2 partially solves the question text Q2, and III may indicate that the answer text A n perfectly solves the question text Q n . 95%, 55% and 75% respectively represent the confidence levels of the first model for the detection results corresponding to their respective text pairs. The higher the confidence level, the lower the discrimination difficulty of the first model for the text pair, and the more accurate the detection result.
[0055] Therefore, by setting a predetermined threshold, based on the first sample detection result, sample text pairs with a confidence level less than the predetermined threshold can be screened out for training the generative model in the second stage.
[0056] For example: the predetermined threshold can be 80%, and the sample text pairs screened based on the first sample detection result can at least include Q2 - A2 (II, 55%), …, Q n -A n (III, 75%).
[0057] According to an embodiment of the present disclosure, before starting the training in the second stage, the discriminant result labels I~III in digital form with manual annotation corresponding to the screened sample text pairs can be converted into corresponding text - form labels: <Did not solve the problem>, <Partially solved the problem>, and <Perfectly solved the problem> to meet the input requirements of the generative model.
[0058] To guide the initial model to detect deep features such as logical relevance in the sample text pairs, the discriminant reason corresponding to the discriminant result label can be used as a label for training the generative model. Exemplarily, the discriminant reason label can be manually annotated according to the discriminant result label.
[0059] For example: for the sample text pair Q2 - A2, the discriminant result label can be “Partially solved the problem”, and the discriminant reason label can be “A2 only solved the first part of the problem in Q2 and did not solve the second part of the problem”.
[0060] According to an embodiment of the present disclosure, the initial model can be used to process the sample text pairs and sample labels to obtain the second sample detection result. The second sample detection result can include a discriminant result for characterizing the reply matching degree between the sample answer text and the sample question text and the discriminant reason corresponding to the discriminant result. For example:
Partially solved the problem; Reason: A2 only solved the first part of the problem in Q2
[0061] According to an embodiment of the present disclosure, based on the first loss function, according to the second sample detection result and the sample label, a first loss value can be obtained. The first loss function can be a cross - entropy loss function. For example:
[0062] (1)
[0063] Wherein, represents the t - th word, represents the first t - 1 words, and T represents the text sequence length of the sum of the text - form discriminant result label and the discriminant reason.
[0064] According to an embodiment of the present disclosure, the model parameters of the initial model can be adjusted based on the first loss value until the convergence condition is met, and a trained target model is obtained. The convergence condition can be the maximum number of training times, or the loss value reaches a predetermined threshold, or the loss value converges. The embodiments of the present disclosure do not specifically limit the convergence condition.
[0065] According to an embodiment of the present disclosure, using the discrimination result and the discrimination reason as labels to guide the initial model to detect using deep features such as the logical relevance between the question text and the answer text, thereby improving the model accuracy, so as to achieve accurate detection of sample text pairs that are difficult to detect for the first model, and obtain a quality detection result of a question-and-answer text pair with a relatively high confidence.
[0066] According to an embodiment of the present disclosure, for the first model, fine-tuning can be performed based on the sample text pairs in a specific application scenario to improve the discrimination accuracy of the first model.
[0067] For example: The first model can be the ERNIE (Enhanced Representation from knowledge Integration) model, which is constructed based on the encoder of Transformer. In the pre-training stage, the pre-training tasks performed can include the Masked Language Modeling (MLM) task and the Next Sentence Prediction (NSP) task.
[0068] In the embodiment of the present disclosure, a branch for outputting the confidence can be added to the output layer of the ERNIE model to determine the discrimination difficulty of the ERNIE model for the input sample text pair.
[0069] Then, the pre-trained model can be used to detect the sample text pair, generate a sample detection result and the confidence of the sample detection result; based on the second loss function, according to the sample detection result and the sample discrimination result label, obtain the second loss value; and based on the second loss value, adjust the model parameters of the pre-trained model to obtain the first model.
[0070] According to an embodiment of the present disclosure, the construction form of the sample text pair can be: <question text><separator><answer text>. The sample discrimination result label can be manually labeled.
[0071] According to an embodiment of the present disclosure, the second loss function can be the cross-entropy loss function. For example:
[0072] (2)
[0073] where C represents the category of the sample discrimination result; Indicates the sample discrimination result label; is the probability corresponding to the true label.
[0074] In the embodiments of the present disclosure, the categories of the sample discrimination results may include 3, and the sample discrimination result labels are respectively I - the problem is not solved; II - the problem is partially solved; III - the problem is perfectly solved.
[0075] It should be noted that the categories of the sample discrimination results can be set according to the needs of the actual application scenario, and the present disclosure does not make specific limitations on this.
[0076] According to the embodiments of the present disclosure, the pre-trained model may at least include a feature extraction layer and a feature processing layer. When adjusting the model parameters of the pre-trained model based on the second loss value, to improve the model training efficiency, the model parameters of the feature extraction layer can be fixed first, and the model parameters of the feature processing layer can be fine-tuned. When the convergence speed of the second loss value is slow, the model parameters of the feature extraction layer and the feature processing layer can be adjusted in parallel to improve the convergence speed of the second loss value and shorten the model training cycle.
[0077] According to the embodiments of the present disclosure, by adding a confidence output branch, the discrimination difficulty of the first model for the sample text pair can be determined based on the confidence, so as to facilitate screening out the sample text pairs with higher discrimination difficulty for training the generative model.
[0078] The following refers to Figure 3A and Figure 3B , and further illustrate the method shown in Figure 2 in combination with specific embodiments.
[0079] Figure 3A Schematically shows a schematic diagram of model training according to an embodiment of the present disclosure.
[0080] As Figure 3A shown, in this embodiment 300A, it may include an ERNIE model 310 and a Llama model 320.
[0081] First, input the Q&A (Query & Answer) text pairs [Q1 - A1,..., Q i -A i ,..., Q I -A I 311 into the ERNIE model 310, and output the detection result 312 and the confidence 313. The detection result 312 and the confidence 313 correspond to the Q&A text pairs. For example: the detection result for the Q1 - A1 text pair is I (the problem is not solved), and the confidence is 67%. The detection result for the Q1 - A1 text pair is II (the problem is partially solved), and the confidence is 97%.
[0082] Then, based on a predetermined threshold of confidence, e.g., 90%, the text pairs with confidence lower than the predetermined threshold are screened out as sample data for training the LlaMA model 320. Q&A text pairs [Q1 - A 1、…、 Q j -A j、…、 Q J -A J 314 can be obtained. Since the confidence of the detection results corresponding to the screened Q&A text pairs is low, it indicates that the ERNIE model 310 has difficulty in discriminating these text pairs. Correspondingly, the accuracy of the detection results of the ERNIE model 310 for these text pairs is also low.
[0083] Next, the manually labeled tag 321 and the Q&A text pairs [Q1 - A 1、…、 Q j -A j、…、 Q J -A J 314 are jointly input into the LlaMA model 320 for training to generate discrimination results + discrimination reasons 323.
[0084] According to an embodiment of the present disclosure, using an initial model to process sample text pairs and sample labels to obtain second sample detection results may include the following operations: concatenating the sample text pairs and sample labels to generate a sample target text; tokenizing the sample target text to generate a sample word sequence; encoding each word in the sample word sequence according to its arrangement position to generate a sample feature sequence; and processing the sample feature sequence based on an attention mechanism to generate a second sample detection result.
[0085] According to an embodiment of the present disclosure, the sample target text may be a text paragraph composed of sample text pairs, sample discrimination result labels, and sample discrimination reason labels. Tokenizing the sample target text can tokenize the sample target text in the form of words or phrases to form a sample word sequence. Then, based on the index value of the word or phrase in the corpus, the sample word sequence is converted into an integer index sequence. Next, the integer index sequence can be mapped to a real-valued vector and encoded based on the position of each word or phrase in the Token sequence to generate a sample feature sequence.
[0086] According to an embodiment of the present disclosure, the Llama model 320 is constructed based on a Transformer decoder and uses an autoregressive manner, that is, each Token (text unit) in the output sequence is generated one by one. The text unit can be a word, punctuation mark, number, or other language elements. During the decoding process, each time a Token is generated, the previously generated content is used as context to help predict the next Token. The generated Token sequence passes through an output layer, usually a linear transformation plus a Softmax function, which converts the probability distribution at each position into the probability of the corresponding Token. According to the probability, the Token with the highest probability is selected as the prediction result of the model. In a generation task, the above autoregressive generation process can be repeated to generate multiple Tokens until a termination marker (such as a full stop or end symbol) is encountered or a preset maximum output length is reached, and the second detection result including the discrimination result and the discrimination reason is output.
[0087] According to an embodiment of the present disclosure, the discrimination result label and the discrimination reason label can be concatenated to generate a label text; and based on the first loss function, the first loss value is obtained according to the label text and the second sample detection result.
[0088] For example: for Q1-A1, the discrimination result label can be "unsolved problem", and the discrimination reason can be "the answer has nothing to do with the question". "Unsolved problem" and "the answer has nothing to do with the question" can be concatenated to generate the label text "unsolved problem, because the answer has nothing to do with the question". The second detection result output by the Llama model 320 is a text including the discrimination result and the discrimination reason, so there is no need to perform a secondary concatenation operation. For example: the second detection result can be "unsolved problem, because the answer is irrelevant to the question". Then, based on the formula (1) described above, the loss value 324 is calculated according to the label text and the discrimination result + discrimination reason 323.
[0089] According to an embodiment of the present disclosure, since the manually annotated label 321 includes the discrimination result label 321_1 and the discrimination reason label 321_2, during the training process of the Llama model 320, the Llama model 320 can learn not only the discrimination result but also the discrimination reason corresponding to the discrimination result based on the context learning of the input data, enabling the Llama model 320 to have the ability to reason and discriminate on deep logical features.
[0090] In the scenario of Q&A text quality detection, the technical fields involved in Q&A are very extensive. However, the reasons for manual annotation are limited to the knowledge reserves and subjective cognitions of relevant personnel, resulting in uneven accuracy of the reasons for annotation. Therefore, there is a certain accuracy bottleneck in the Llama model 320 trained with the reasons for annotation tags of manual annotation.
[0091] The generative large language model is pre-trained based on rich corpora and has strong logical reasoning ability. Therefore, the embodiments of the present disclosure can use the reasons for discrimination output by the generative large language model as the reasons for discrimination tags to train the Llama model 320 to break through the accuracy bottleneck of the Llama model 320.
[0092] According to an embodiment of the present disclosure, a prompt text can be constructed according to a sample text pair, a first discrimination result tag corresponding to the sample text pair, and a reference text, where the reference text includes a text pair related to the target sample text, a second discrimination result tag corresponding to the text pair, and a discrimination reason tag corresponding to the text pair, and the first discrimination result tag and the second discrimination result tag are the same; use the large language model to process the prompt text to generate a second discrimination reason corresponding to the sample text pair; and determine the second discrimination reason as the discrimination reason tag.
[0093] Figure 3B Schematically shows a schematic diagram of model training according to another embodiment of the present disclosure.
[0094] As Figure 3B shown, in this embodiment 300B, it may include an ERNIE model 310, a Llama model 320, and a large language model 330. The process of using the ERNIE model 310 to process the Q&A text pair 311 to obtain a detection result 312 and screening out the Q&A text pair 314 for training the Llama model 320 from the Q&A text pair 311 based on the confidence 313 is the same as that in the embodiment 300A and will not be elaborated here.
[0095] Then, a prompt text 331 is constructed according to the screened Q&A text pair 314, the manually annotated discrimination result, and the reference text. The reference text may include other text pairs Q t -A t and the manually annotated discrimination result + discrimination reason corresponding to the text pair Q t -A t For example: the text pair Q
[0096] j -A j It can be questions and answers in the field of autonomous driving technology, 1-3 text pairs randomly selected from other question-answer text pairs in the field of autonomous driving technology, and the discrimination results and reasons for discrimination are marked for each of the 1-3 text pairs. For the text pair Q j -A j and the manually marked discrimination result corresponding to the text pair Q j -A j The corresponding discrimination results, the above 1-3 text pairs, and the discrimination results and reasons for discrimination corresponding to the above 1-3 text pairs together form a Prompt (prompt text).
[0097] It should be noted that the discrimination result label types of the 1-3 text pairs are the same as those of the text pair Q j -A j . For example: if the discrimination result label of the text pair Q j -A j is "I-Unresolved Problem", then the discrimination result label types of the 1-3 text pairs are also "I-Unresolved Problem", so that the generative large model can refer to the reasons for discrimination of the 1-3 text pairs and output a reason label that matches the discrimination result.
[0098] Next, input the prompt text 331 into the large language model 330 to output the reason for discrimination corresponding to the text pair Q j -A j , and determine this reason for discrimination as the discrimination reason label 321_3. Then, use the discrimination reason label 321_3 combined with the manually marked discrimination result label to train the Llama model 320. The training process is the same as that described in the previous embodiment 300A and will not be elaborated here.
[0099] According to the embodiments of the present disclosure, the discrimination results and reasons for discrimination of other text pairs related to the field of the text pair are jointly input into the large model as reference texts to output the discrimination reason labels for training the Llama model 320. By making full use of the rich corpus and logical reasoning ability of the generative large model, the logical reasoning ability of the generative large model is learned by the Llama model 320 in a way similar to result distillation. This can improve the model accuracy while reducing the model's demand for hardware resources, enabling the Llama model 320 to efficiently complete the text pair quality detection task under limited hardware resources.
[0100] In the scenario of quality detection for question-answer text pairs, not only the accuracy requirements of the detection results need to be met, but also the demand of the model operation for hardware resources needs to be considered. Therefore, the embodiments of the present disclosure provide a text detection method that combines a discriminative model and a generative model.
[0101] Figure 4A flowchart of a text detection method according to an embodiment of the present disclosure is schematically shown.
[0102] As Figure 4 shown, the text detection method 400 includes operations S410 to S440.
[0103] In operation S410, a first model is used to detect a plurality of text pairs to be detected, and a first detection result and a confidence level corresponding to the first detection result are obtained.
[0104] In operation S420, according to the confidence level, a target text pair to be detected is determined from the plurality of text pairs to be detected.
[0105] In operation S430, a second model is used to detect the target text pair to be detected, and a second detection result is obtained.
[0106] In operation S440, according to the first detection result and the second detection result, a target detection result is generated.
[0107] According to an embodiment of the present disclosure, the first model may be a discriminative model, for example: the ERNIE model. The second model may be a target model trained by using the model training method described above, for example: the Llama model.
[0108] According to an embodiment of the present disclosure, the plurality of text pairs to be detected include a plurality of question texts and a plurality of answer texts corresponding to each other.
[0109] According to an embodiment of the present disclosure, the first model may be used to detect all the text pairs to be detected, and a first detection result of each text pair and a confidence level corresponding to the first detection result are obtained. For example: the first detection result of the text pair Q1-A1 is II (partially solve the problem), and the confidence level is 95%. The first detection result of the text pair Q2-A2 is II (partially solve the problem), and the confidence level is 55%.
[0110] According to an embodiment of the present disclosure, the predetermined threshold of the confidence level may be set to 80%. Since the confidence level of the first detection result of the text pair Q1-A1 is greater than 80%, the first detection result of the text pair Q1-A1 may be used as the final detection result. However, since the confidence level of the first detection result of the text pair Q2-A2 is less than 80%, the text pair Q2-A2 may be determined as the target text pair and used as the input data of the second model. In order to reduce the data processing amount of the second model, the detection efficiency can be improved.
[0111] According to an embodiment of the present disclosure, by using the second model, the second detection result obtained by detecting the target text pair includes both the discriminant result category of the text pair and the discriminant reason corresponding to the discriminant result category.
[0112] For example, processing the text pair Q2-A2 using the second model, the second detection result obtained can be [I (unsolved problem), because the answer A2 is irrelevant to the question Q2].
[0113] Finally, combine the first detection results with a confidence level greater than a predetermined threshold output by the first model and the second detection results output by the second model as the target detection result. The target detection result characterizes the reply matching degree between the answer text and the corresponding question text.
[0114] According to an embodiment of the present disclosure, by comprehensively using a discriminative model and a generative model, compared with a single discriminative model in related examples, it can not only solve the problem of low detection accuracy caused by data overfitting of the discriminative model and improve the robustness of the model, but also, when hardware resources are limited, use the generative model to only detect text pairs with low confidence levels, reduce the data processing volume of the generative model, and improve the detection efficiency.
[0115] Figure 5 Schematically shows a schematic diagram of text detection according to an embodiment of the present disclosure.
[0116] As Figure 5 shown, in Embodiment 500, it may include an ERNIE model 510 and a Llama model 520.
[0117] First, input the Q&A text pair 501 "Q1-A1, Q2-A2, Q3-A3" into the ERNIE model 510, and output the detection result 502 "Q1-A1_ partially solves the problem_ 80%; Q2-A2_ perfectly solves the problem_ 60%; Q3-A3_ unsolved problem_ 50%". Based on the confidence level, screen out the Q&A text pair 503 "Q2-A2, Q3-A3" from the Q&A text pair 501.
[0118] Then, input the Q&A text pair 503 "Q2-A2, Q3-A3" into the Llama model 520, and output the detection result 504 "Q2-A2_ perfectly solves the problem_ discriminant reason XXX; Q3-A3_ partially solves the problem_ discriminant reason YYY".
[0119] For example, the text "Q2-A2" to be detected can be tokenized to generate a sequence of words. Then, based on the index values of the words or phrases in the corpus, the sample word sequence is converted into a sequence of integer indices. Next, the sequence of integer indices can be mapped to real-valued vectors, and encoded based on the position of each word or phrase in the Token sequence to generate a feature sequence. Next, in an autoregressive manner, based on the attention mechanism, each Token (text unit) in the output sequence is generated one by one. The text unit can be a word, punctuation mark, number, or other language element. During the decoding process, when generating each Token, the previously generated content is used as context to help predict the next Token. The generated Token sequence passes through an output layer, typically a linear transformation followed by a Softmax function, which converts the probability distribution at each position into the probability of the corresponding Token. Based on the probabilities, the Token with the highest probability is selected as the prediction result of the model. In a generation task, the above autoregressive generation process can be repeated to generate multiple Tokens until a termination marker (such as a period or end symbol) is encountered or a preset maximum output length is reached, and the second detection result can be "Q2-A2_ perfectly solves the problem_ discriminant reason XXX".
[0120] Finally, the results with a confidence level of the detection result 502 greater than a predetermined threshold are merged with the detection result 504 to obtain the detection result 505 "Q1-A1_ partially solves the problem_ 80%; Q2-A2_ perfectly solves the problem_ discriminant reason XXX; Q3-A3_ partially solves the problem_ discriminant reason YYY."
[0121] Figure 6 A block diagram of a model training apparatus according to an embodiment of the present disclosure is schematically shown.
[0122] As Figure 6 shown, the model training apparatus 600 may include a first processing module 610, a first loss calculation module 620, and a first adjustment module 630.
[0123] The first processing module 610 is configured to, in response to the confidence level of the first sample detection result being less than a predetermined threshold, process the sample text pair and the sample label using an initial model to obtain a second sample detection result. Wherein, the confidence level of the first sample detection result is obtained by detecting the sample text pair using a first model; the sample text pair includes a corresponding sample question text and a sample answer text, the second sample detection result includes a discrimination result for characterizing the reply matching degree between the sample answer text and the sample question text and a discrimination reason corresponding to the discrimination result; the sample label includes a discrimination result label and a discrimination reason label; the initial model is of a different type from the first model.
[0124] The first loss calculation module 620 is configured to obtain a first loss value based on a first loss function according to the second sample detection result and the sample label.
[0125] The first adjustment module 630 adjusts the model parameters of the initial model based on the first loss value to obtain a trained target model.
[0126] According to an embodiment of the present disclosure, the first loss calculation module may include: a first splicing sub-module and a calculation sub-module. The first splicing sub-module is configured to splice the discriminant result label and the discriminant reason label to generate a label text. The calculation sub-module is configured to obtain a first loss value based on the first loss function according to the label text and the second sample detection result.
[0127] According to an embodiment of the present disclosure, the first processing module may include: a first splicing sub-module, a first word segmentation sub-module, a first encoding sub-module, and a first processing sub-module.
[0128] The first splicing sub-module is configured to splice the sample text pair and the sample label to generate a sample target text. The first word segmentation sub-module is configured to perform word segmentation on the sample target text to generate a sample word sequence. The first encoding sub-module is configured to encode each word in the sample word sequence according to the arrangement position to generate a sample feature sequence. The first processing sub-module is configured to process the sample feature sequence based on an attention mechanism to generate a second sample detection result.
[0129] According to an embodiment of the present disclosure, the above training device further includes: a construction module, a second processing module, and a first determination module.
[0130] The construction module is configured to construct a prompt text according to the sample text pair, the first discriminant result label corresponding to the sample text pair, and the reference text, where the reference text includes a text pair related to the target sample text, a second discriminant result label corresponding to the text pair, and a discriminant reason label corresponding to the text pair, and the first discriminant result label and the second discriminant result label are the same.
[0131] The second processing module is configured to process the prompt text by using a large language model to generate a second discriminant reason corresponding to the sample text pair.
[0132] The first determination module is configured to determine the second discriminant reason as the discriminant reason label.
[0133] According to an embodiment of the present disclosure, the above training device further includes: a first detection module, a second loss calculation module, and a second adjustment module. The first detection module is configured to detect a sample text pair by using a pre-trained model, and generate a sample detection result and a confidence level of the sample detection result. The second loss calculation module is configured to obtain a second loss value based on a second loss function according to the sample detection result and the sample discrimination result label. The second adjustment module is configured to adjust the model parameters of the pre-trained model based on the second loss value to obtain a first model.
[0134] Figure 7 FIG. schematically shows a block diagram of a text detection device according to an embodiment of the present disclosure.
[0135] As Figure 7 shown, the text detection device 700 may include: a second detection module 710, a second determination module 720, a third detection module 730, and a generation module 740.
[0136] The second detection module 710 is configured to detect a plurality of text pairs to be detected by using the first model, and obtain a first detection result and a confidence level corresponding to the first detection result. The first model is a discriminative model, and the plurality of text pairs to be detected include a plurality of question texts and a plurality of answer texts corresponding to each other.
[0137] The second determination module 720 is configured to determine a target text pair to be detected from the plurality of text pairs to be detected according to the confidence level, wherein the confidence level of the detection result corresponding to the target text pair to be detected is less than a predetermined threshold.
[0138] The third detection module 730 is configured to detect the target text pair to be detected by using the second model, and obtain a second detection result, wherein the second model is a trained target model obtained by the training method described above
[0139] The generation module 740 is configured to generate a target detection result according to the first detection result and the second detection result, and the target detection result characterizes the reply matching degree between the answer text and the corresponding question text.
[0140] According to an embodiment of the present disclosure, the third detection module may include: a second word segmentation sub-module, a second encoding sub-module, and a second processing sub-module.
[0141] The second word segmentation sub-module is configured to segment the text of the target text pair to be detected to generate a word sequence. The second encoding sub-module is configured to encode each word in the word sequence according to the arrangement position to generate a feature sequence. The second processing sub-module is configured to process the feature sequence based on an attention mechanism to generate a second detection result.
[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0143] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0144] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described above.
[0145] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method described above when executed by a processor.
[0146] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0147] As Figure 8 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0148] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disc, etc.; and communication unit 809, such as a network card, modem, wireless communication transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0149] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 executes the various methods and processes described above, such as the model training method or the text detection method. For example, in some embodiments, the model training method or the text detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the model training method or the text detection method described above can be executed. Alternatively, in other embodiments, computing unit 801 can be configured to execute the model training method or the text detection method by any other suitable means (e.g., by means of firmware).
[0150] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0152] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0154] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0155] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.
[0156] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0157] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A model training method, comprising: In response to the confidence level of the first sample detection result being less than a predetermined threshold, the sample text pair and the sample label are processed using the initial model to obtain a second sample detection result; The confidence of the first sample detection result is obtained by detecting the sample text pair using the first model; the sample text pair includes a sample question text and a sample answer text corresponding to each other; the second sample detection result includes a discrimination result for characterizing the reply matching degree between the sample answer text and the sample question text and a discrimination reason corresponding to the discrimination result; the sample label includes a discrimination result label and a discrimination reason label; the initial model is of a different type from the first model; the discrimination reason label is manually annotated according to the discrimination result label, or output by a generative large language model; Based on the first loss function, obtaining a first loss value according to the second sample detection result and the sample label; and Based on the first loss value, the model parameters of the initial model are adjusted to obtain a trained target model.
2. The method according to claim 1, wherein: The obtaining, based on the first loss function and according to the second sample detection result and the sample label, a first loss value includes: splicing the discrimination result label and the discrimination reason label to generate a label text; and Based on the first loss function, the first loss value is obtained according to the label text and the second sample detection result.
3. The method according to claim 1, wherein: The method of using the initial model to process the sample text pair and the sample label to obtain a second sample detection result includes: Concatenating the sample text pair and the sample label to generate a sample target text; Segmenting the sample target text to generate a sample word sequence; Encoding each word in the sample word sequence according to its arrangement position to generate a sample feature sequence; and Based on the attention mechanism, the sample feature sequence is processed to generate the second sample detection result.
4. The method according to any one of claims 1 to 3, further comprising: Constructing a prompt text according to the sample text pair, the first discrimination result label corresponding to the sample text pair, and the reference text, wherein the reference text includes a text pair related to the target sample text, a second discrimination result label corresponding to the text pair, and a discrimination reason label corresponding to the text pair, and the first discrimination result label and the second discrimination result label are the same; Processing the prompt text using a large language model to generate a second discrimination reason corresponding to the sample text pair; and The second determination reason is determined as the determination reason label.
5. The method according to any one of claims 1 to 3, further comprising: Using the pre-trained model to detect the sample text pair, generating a sample detection result and a confidence level of the sample detection result; Based on the second loss function, obtaining a second loss value according to the sample detection result and the sample discrimination result label; as well as Based on the second loss value, adjust the model parameters of the pre-trained model to obtain the first model.
6. A text detection method, comprising: Using a first model to detect multiple pairs of texts to be detected, obtaining a first detection result and a confidence level corresponding to the first detection result, wherein the first model is a discriminant model, and the multiple pairs of texts to be detected include multiple question texts and multiple answer texts corresponding to each other; Determining a target text pair to be detected from the plurality of text pairs to be detected according to the confidence, wherein the confidence of the detection result corresponding to the target text pair to be detected is less than a predetermined threshold; Detecting the target text pair to be detected using a second model to obtain a second detection result, wherein the second model is a trained target model obtained by the training method according to any one of claims 1 to 4; and A target detection result is generated based on the first detection result and the second detection result, and the target detection result represents the response matching degree between the answer text and the corresponding question text.
7. The method according to claim 6, wherein: The detecting the target text pair to be detected by using the second model to obtain a second detection result includes: Segmenting the target text pair to be detected to generate a word sequence; Encoding each word in the word sequence according to its arrangement position to generate a feature sequence; and Based on the attention mechanism, the feature sequence is processed to generate the second detection result.
8. A model training device, comprising: A first processing module, configured to process the sample text pair and the sample label using the initial model to obtain a second sample detection result in response to the confidence level of the first sample detection result being less than a predetermined threshold; The confidence of the first sample detection result is obtained by detecting the sample text pair using the first model; the sample text pair includes a sample question text and a sample answer text corresponding to each other; the second sample detection result includes a discrimination result for characterizing the reply matching degree between the sample answer text and the sample question text and a discrimination reason corresponding to the discrimination result; the sample label includes a discrimination result label and a discrimination reason label; the initial model is of a different type from the first model; the discrimination reason label is manually annotated according to the discrimination result label, or output by a generative large language model; a first loss calculation module, configured to obtain a first loss value according to the second sample detection result and the sample label based on a first loss function; and A first adjustment module adjusts model parameters of the initial model based on the first loss value to obtain a trained target model.
9. The device according to claim 8, wherein: The first loss calculation module includes: A first splicing submodule is used to splice the discrimination result label and the discrimination reason label to generate a label text; and A calculation submodule is used to obtain the first loss value based on the first loss function, according to the label text and the second sample detection result.
10. The device according to claim 8, wherein: The first processing module comprises: A first concatenation submodule, used for concatenating the sample text pair and the sample label to generate a sample target text; A first word segmentation submodule is used to segment the sample target text to generate a sample word sequence; A first encoding submodule is used to encode each word in the sample word sequence according to its arrangement position to generate a sample feature sequence; and The first processing submodule is used to process the sample feature sequence based on the attention mechanism to generate the second sample detection result.
11. The device according to any one of claims 8 to 10, further comprising: A construction module, configured to construct a prompt text according to the sample text pair, a first discrimination result label corresponding to the sample text pair, and a reference text, wherein the reference text includes a text pair related to the target sample text, a second discrimination result label corresponding to the text pair, and a discrimination reason label corresponding to the text pair, and the first discrimination result label and the second discrimination result label are the same; A second processing module is used to process the prompt text using a large language model to generate a second discrimination reason corresponding to the sample text pair; and The first determining module is configured to determine the second determination reason as the determination reason label.
12. The device according to any one of claims 8 to 10, further comprising: A first detection module, used to detect the sample text pair using a pre-trained model, and generate a sample detection result and a confidence level of the sample detection result; A second loss calculation module, used to obtain a second loss value based on a second loss function and according to the sample detection result and the sample discrimination result label; as well as The second adjustment module is used to adjust the model parameters of the pre-trained model based on the second loss value to obtain the first model.
13. A text detection device, comprising: A second detection module is used to detect a plurality of text pairs to be detected using a first model to obtain a first detection result and a confidence level corresponding to the first detection result, wherein the first model is a discriminant model, and the plurality of text pairs to be detected include a plurality of question texts and a plurality of answer texts corresponding to each other; A second determination module is used to determine a target text pair to be detected from the multiple text pairs to be detected according to the confidence, wherein the confidence of the detection result corresponding to the target text pair to be detected is less than a predetermined threshold; A third detection module is used to detect the target text pair to be detected using a second model to obtain a second detection result, wherein the second model is a trained target model obtained by the training method according to any one of claims 1 to 4; and A generation module is used to generate a target detection result based on the first detection result and the second detection result, and the target detection result represents the response matching degree between the answer text and the corresponding question text.
14. The device according to claim 13, wherein: The third detection module comprises: A second word segmentation submodule is used to segment the target text pair to be detected to generate a word sequence; A second encoding submodule is used to encode each word in the word sequence according to its arrangement position to generate a feature sequence; and The second processing submodule is used to process the feature sequence based on the attention mechanism to generate the second detection result.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Visual question and answer method and device based on deep learning model, medium and equipment
CN113656570A
Method and device for training generative large language model based on knowledge base feedback
CN117009490A