Question and answer corpus generation method and device, computer equipment and readable storage medium

By extracting and labeling the multi-dimensional characteristics of the initial questions of the target discipline, screening the questions to be answered, and generating the question-and-answer corpus with the accuracy of the answers, the problem of insufficient accuracy of corpus generation in the prior art is solved, and the accuracy and reliability of the question-and-answer corpus is improved.

CN119938818APending Publication Date: 2025-05-06ZHEJIANG LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411763562.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, there is a problem of insufficient accuracy in the generation of corpus in the scientific research field using large language models, especially in vertical fields, which is difficult to meet the needs of high accuracy and professionalism.

Method used

Provide a method for generating a question-and-answer corpus, by obtaining the initial question set of the target subject, extracting multi-dimensional features and annotating them, detecting and filtering questions to be answered, obtaining the answer set and their accuracy, and finally generating the target question-and-answer corpus pair.

Benefits of technology

Through multi-dimensional feature annotation and preset condition screening, we ensure that the key information of the question is fully captured and the accuracy of the initial question is improved; based on the accuracy of the answer, the quality and reliability of the answers are ensured, and the accuracy and reliability of the generated question-and-answer corpus pairs are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938818A_ABST
    Figure CN119938818A_ABST
Patent Text Reader

Abstract

The invention relates to a question and answer corpus generation method and device, computer equipment and a readable storage medium. Obtaining an initial question set corresponding to the target subject; extracting multi-dimensional features of each initial question in the initial question set, and determining a labeling mode corresponding to each dimensional feature; for each initial question, labeling each dimension feature according to a respective corresponding labeling mode to obtain a labeling result corresponding to the initial question; detecting the initial question according to the labeling result, and determining the initial question as a question to be answered under the condition that the initial question meets a preset condition; obtaining an answer set corresponding to the to-be-answered question and the accuracy of each answer in the answer set; generating a target question and answer corpus pair according to the to-be-answered question, the answer set and the accuracy of each answer; the accuracy and reliability of the generated target question and answer corpus pair can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of corpus generation, and in particular to a question-answer corpus generation method, apparatus, computer device, and readable storage medium. Background Art

[0002] With the rapid development of deep learning and natural language processing technologies, large language models have shown great application potential in many fields. However, applying large language models to vertical fields, especially scientific research, requires higher accuracy and professionalism. Most of the current model training and tuning methods are based on open corpora, which are obviously insufficient in terms of corpus accuracy.

[0003] There is currently no effective solution to the problem of low corpus accuracy in the prior art. Summary of the invention

[0004] Based on this, it is necessary to provide a question and answer corpus generation method, device, computer equipment and readable storage medium to address the above technical problems.

[0005] In a first aspect, the present application provides a method for generating a question-answer corpus, the method comprising:

[0006] Obtain the initial set of questions corresponding to the target subject;

[0007] Extracting multi-dimensional features of each initial question in the initial question set, and determining a labeling method corresponding to each of the dimensional features;

[0008] For each of the initial questions, annotate each of the dimensional features according to the corresponding annotation method to obtain an annotation result corresponding to the initial question;

[0009] Detecting the initial question according to the marking result, and determining the initial question as a question to be answered if the initial question meets a preset condition;

[0010] Obtaining an answer set corresponding to the question to be answered, and the accuracy of each answer in the answer set;

[0011] A target question-answer corpus pair is generated according to the question to be answered, the answer set, and the accuracy of each answer.

[0012] In one embodiment, the multi-dimensional features include first dimensional features and second dimensional features; for each of the initial questions, each of the dimensional features is annotated according to the corresponding annotation method, and the annotation result corresponding to the initial question is obtained, including:

[0013] Pre-labeling the first dimensional features corresponding to the initial question using a large language model to obtain a first result;

[0014] If the first result satisfies a first preset screening condition, determining the initial question corresponding to the first result as a candidate question;

[0015] The first dimensional feature and the second dimensional feature corresponding to the candidate question are labeled according to the labeling methods corresponding to the first dimensional feature and the second dimensional feature, to obtain a second result corresponding to the first dimensional feature and a third result corresponding to the second dimensional feature;

[0016] A labeling result corresponding to the initial question is determined according to the second result and the third result.

[0017] In one embodiment, the detecting the initial question according to the labeling result and determining the initial question as a question to be answered if the initial question meets a preset condition includes:

[0018] If the second result satisfies the second preset screening condition, quantizing the third result to obtain a quantized annotation result;

[0019] Calculating the reliability corresponding to the initial question according to the quantitative labeling result;

[0020] If the reliability is greater than the preset reliability threshold, the initial question is determined as a question to be answered.

[0021] In one embodiment, obtaining the answer set corresponding to the question to be answered includes:

[0022] For each of the questions to be answered, determining a plurality of target processing objects that match the question to be answered;

[0023] Extracting a first answer input by each target processing object to the question to be answered;

[0024] Processing the question to be answered by using a large language model to obtain a second answer corresponding to the question to be answered;

[0025] According to the plurality of the first answers and the second answers, an answer set corresponding to each of the questions to be answered is obtained.

[0026] In one embodiment, the determining of a plurality of target processing objects matching the question to be answered includes:

[0027] When the question to be answered is to be assigned, obtaining the assignment waiting time corresponding to the question to be answered;

[0028] If the allocation waiting time is less than the preset time threshold, and there is a first processing object that accepts the question to be answered, the question to be answered is allocated to the first processing object, and the first processing object is used as the target processing object; the subject level corresponding to the subject to which the first processing object belongs is the same as the subject level corresponding to the target subject;

[0029] If the assigned waiting time is greater than or equal to the preset time threshold, and there is no first processing object that accepts the question to be answered, the question to be answered will be assigned to the second processing object, and the second processing object will be used as the target processing object; the subject level corresponding to the subject to which the second processing object belongs is higher than the subject level corresponding to the target subject.

[0030] In one embodiment, obtaining the accuracy of each answer in the answer set includes:

[0031] For each of the first answers among the plurality of the first answers, obtaining a first accuracy corresponding to each of the first answers;

[0032] Obtaining a second accuracy corresponding to the second answer;

[0033] The accuracy of each answer in the answer set is obtained according to the first accuracies and the second accuracies.

[0034] In one embodiment, generating a target question-answer corpus pair according to the question to be answered, the answer set, and the accuracy of each answer includes:

[0035] Determine the highest accuracy from a plurality of the first accuracies and the second accuracies, and use the answer corresponding to the highest accuracy as the target answer;

[0036] Generate a target question-answer corpus pair according to the target answer and the question to be answered corresponding to the target answer;

[0037] The target question-answer corpus pair is stored in a preset format.

[0038] In a second aspect, the present application further provides a question-answer corpus generation device, the device comprising:

[0039] A question acquisition module is used to obtain an initial set of questions corresponding to the target subject;

[0040] A feature extraction module, used to extract multi-dimensional features of each initial question in the initial question set, and determine a labeling method corresponding to each of the dimensional features;

[0041] A labeling module, used for labeling each of the dimensional features according to the corresponding labeling method for each of the initial questions, to obtain a labeling result corresponding to the initial question;

[0042] A determination module, configured to detect the initial question according to the marking result, and determine the initial question as a question to be answered if the initial question meets a preset condition;

[0043] An answer acquisition module, used to acquire an answer set corresponding to the question to be answered, and the accuracy of each answer in the answer set;

[0044] A generation module is used to generate a target question-answer corpus pair based on the question to be answered, the answer set and the accuracy of each answer.

[0045] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the embodiments of the first aspect are implemented.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of the embodiments of the first aspect above are implemented.

[0047] In a fifth aspect, the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect above.

[0048] The above-mentioned question-answer corpus generation method, device, computer equipment and readable storage medium obtain an initial question set corresponding to the target subject; extract multi-dimensional features of each initial question in the initial question set, and determine the labeling method corresponding to each dimensional feature; by labeling each initial question in the initial question set with multi-dimensional features, the labeling result corresponding to the initial question can be obtained, which can ensure that the key information of each question is fully captured to improve the accuracy of the initial question, and lay a foundation for improving the accuracy of the question-answer corpus; further, the initial question is detected according to the labeling result, and when the initial question meets the preset conditions, the initial question is determined as the question to be answered; by setting the preset conditions to screen the questions to be answered, the questions that do not meet the preset conditions can be effectively excluded, ensuring the quality of the questions finally selected for generating the question-answer corpus pair, which helps to improve the pertinence and reliability of the answers; further, the answer set corresponding to the question to be answered and the accuracy of each answer in the answer set are obtained; based on the accuracy of the answer, the quality and reliability of the answer can be further guaranteed; finally, according to the question to be answered, the answer set and the accuracy of each answer, the target question-answer corpus pair is generated, which can effectively improve the accuracy and reliability of the generated target question-answer corpus pair. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0050] Figure 1 is an application environment diagram of a question-answer corpus generation method in one embodiment;

[0051] Figure 2 Schematic diagram of a process for generating question-answer corpus in one embodiment;

[0052] Figure 3 A schematic diagram of a flow chart of a step of determining a marking result in an embodiment;

[0053] Figure 4 A flowchart of an answer set acquisition step in one embodiment;

[0054] Figure 5 A schematic diagram of a subject setting in an embodiment;

[0055] Figure 6 A schematic diagram of marking based on a marking method in an embodiment;

[0056] Figure 7is a structural block diagram of a question-answer corpus generating device in one embodiment;

[0057] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] The question-answer corpus generation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0060] In an exemplary embodiment, Figure 2 As shown, Figure 2 1 is a flow chart of a method for generating question-answer corpus in an embodiment; this embodiment uses the method applied to a terminal as an example for illustration, and it is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0061] Step S201, obtaining an initial question set corresponding to the target subject.

[0062] The target discipline is any one of the multi-level disciplines; the multi-level disciplines are set up according to the standard discipline classification; the multi-level disciplines may include but are not limited to first-level disciplines, second-level disciplines, and third-level disciplines.

[0063] Among them, the first-level disciplines may include but are not limited to sociology, statistics, political science, law, economics, management, environmental science and technology, chemical engineering, water conservancy engineering, energy science and technology, agronomy, forestry, mining engineering technology, biology, earth science, etc.; the second-level disciplines may include but are not limited to atmospheric science, solid earth physics, space physics, geochemistry, geology, hydrology, marine science, etc.; the tertiary disciplines may include but are not limited to regional geography, urban geography, world geography, tourism geography, disaster geography, remote sensing science, geographic information science, surveying and cartography, geographic big data and spatial intelligence, geographic observation and simulation, other disciplines of geography, soil geography, historical geography, population geography, etc.

[0064] The initial question set includes multiple initial questions.

[0065] In an exemplary embodiment, the initial question set may be stored in a server or database in the form of text, table, etc., but not limited to. For example, each initial question in the initial question set includes multiple fields, and the multiple fields may include but are not limited to question ID, question text, creation time, tags, and multi-dimensional features; the multiple fields corresponding to each initial question are stored in a text file; wherein each line of the text file represents an initial question, and each field is separated by a comma.

[0066] Exemplarily, the method for obtaining the initial question set corresponding to the target subject may be: determining a storage path of the initial question set in the server, and reading the initial question set corresponding to the target subject from the server based on the storage path.

[0067] Step S202: extracting multi-dimensional features of each initial question in the initial question set, and determining a labeling method corresponding to each dimensional feature.

[0068] Among them, the multi-dimensional features may include, but are not limited to, first-dimensional features and second-dimensional features; among them, the first-dimensional features may include, but are not limited to, question quality, subject classification, question type, and negative type. Among them, question quality refers to whether the text content is a question or a task, and whether it is suitable for starting a question-answering dialogue package, for example, whether it is a spam question; subject classification refers to whether the subject attribution is correct; question type refers to the correct question type, including Brainstorming, Classification, Extraction, Generation, Rewrite, Chat, Closed QA, Open QA, Summarization, Others, etc. Negative types include non-target language, inappropriateness, inclusion of personally identifiable information, factual errors, logical errors, calculation errors, repetitiveness, and non-domain relevance. Among them, non-target language refers to the use of a language different from the target language unless otherwise specified; inappropriateness refers to the lack of clear requests in questions or tasks; inclusion of personally identifiable information refers to the inclusion of non-public personally identifiable information in the information, which can be used to determine the identity of the user or a private third party; factual errors refer to factual errors in questions unless otherwise specified; logical errors refer to logical errors in questions unless otherwise specified; calculation errors refer to calculation errors in questions unless otherwise specified; repetitiveness refers to the expression of a point of view repeatedly; and non-domain relevance refers to the irrelevance of the questions raised to the target subject area.

[0069] Among them, the second dimension features may include, but are not limited to, question difficulty, question creativity, reasoning and calculation complexity, content completeness, layout neatness, seriousness, politeness, and harmfulness. Among them, question difficulty refers to the difficulty of answering this question or task; question creativity refers to whether the question method and question content are rare, and the possibility of the question being asked; reasoning and calculation complexity refers to whether the question or task requires complex reasoning and calculation; content completeness refers to whether the question has rich descriptive content and can clearly express the intention of the question; layout neatness refers to whether the layout is used correctly, such as whether paragraphs, tables, formulas, etc. are expressed in the correct form; seriousness refers to whether the information contains irony, word games or other humorous modifications; politeness refers to the attitude towards the recipient of the information, for example, whether words like "please" are used, or whether an unfriendly attitude is shown to the other party; harmfulness refers to whether the information contains a clear description of violent, illegal, and immoral behavior, and whether such behavior is glorified.

[0070] Among them, the labeling method refers to recording multi-dimensional features in a standardized form for subsequent processing and analysis. It should be noted that the choice of labeling method depends on the nature of the features of each dimension and the application scenario, and is not specifically limited here. For example, for whether the quality of the question is a junk question, the corresponding labeling method can be: yes or no; for the difficulty of the question, the corresponding labeling method can be: very easy, easy, uncertain, difficult, very difficult.

[0071] In an exemplary embodiment, assuming that the initial question set is stored in the form of text, the text is parsed to determine multiple initial questions; entity recognition is performed on the initial questions to determine the target entity corresponding to the initial question, and the multi-dimensional features of the initial question are determined based on the target entity. Further, based on the multi-dimensional features, the labeling method corresponding to each dimensional feature is determined respectively.

[0072] Step S203: for each initial question, annotate each dimensional feature according to its corresponding annotation method to obtain an annotation result corresponding to the initial question.

[0073] Exemplarily, taking question quality and question difficulty as an example, according to the labeling method corresponding to the question quality, the question quality is labeled whether it is a junk question, and the labeling result can be "yes"; according to the labeling method corresponding to the question difficulty, the question difficulty is labeled, and the labeling structure can be obtained as "difficult"; the labeling method of other multi-dimensional features is the same as the labeling method recorded in the above embodiment, and will not be repeated here.

[0074] Step S204, detecting the initial question according to the marking result, and determining the initial question as a question to be answered when the initial question meets a preset condition.

[0075] Among them, the preset conditions need to be set according to the actual problem screening requirements and are not specifically limited here.

[0076] Among them, the method for detecting the initial problem according to the annotation results can be: performing data analysis according to the annotation results to obtain corresponding analysis results; comparing the analysis results with preset conditions to determine whether the analysis results meet the preset conditions. If the analysis results meet the preset conditions, it is determined that the initial problem meets the preset conditions; if the analysis results do not meet the preset conditions, it is determined that the initial problem does not meet the preset conditions.

[0077] Step S205, obtaining an answer set corresponding to the question to be answered and the accuracy of each answer in the answer set.

[0078] The answer set includes multiple answers; each answer has a corresponding accuracy; and the accuracy is used to characterize the correctness and reliability of the answer.

[0079] The answer set and the accuracy of each answer in the answer set may be stored in a server or database in the form of text, table, etc., but is not limited to the form.

[0080] Step S206, generating a target question-answer corpus pair based on the question to be answered, the answer set, and the accuracy of each answer.

[0081] In an exemplary embodiment, the answer with the highest accuracy is determined based on the question to be answered, the answer set, and the accuracy of each answer; and a target question-answer corpus pair is generated based on the answer with the highest accuracy and the question to be answered corresponding to the answer.

[0082] In this embodiment, by performing multi-dimensional feature annotation on each initial question in the initial question set, the annotation results corresponding to the initial questions are obtained, which can ensure that the key information of each question is fully captured, so as to improve the accuracy of the initial questions and lay the foundation for improving the accuracy of the question and answer corpus; further, the initial questions are detected according to the annotation results, and when the initial questions meet the preset conditions, the initial questions are determined as questions to be answered; by setting preset conditions to screen the questions to be answered, questions that do not meet the preset conditions can be effectively excluded, ensuring the quality of the questions finally selected for generating the question and answer corpus pairs, which helps to improve the pertinence and reliability of the answers; further, the answer set corresponding to the questions to be answered and the accuracy of each answer in the answer set are obtained; based on the accuracy of the answers, the quality and reliability of the answers can be further guaranteed; finally, based on the questions to be answered, the answer set and the accuracy of each answer, the target question and answer corpus pair is generated, which can effectively improve the accuracy and reliability of the generated target question and answer corpus pair.

[0083] In one embodiment, the multi-dimensional features include first-dimensional features and second-dimensional features; Figure 3 As shown, Figure 3 The flowchart of the step of determining the labeling result in one embodiment is as follows; for each initial question, each dimension feature is labeled according to its corresponding labeling method to obtain the labeling result corresponding to the initial question, including the following steps:

[0084] Step S301: Pre-label the first dimensional features corresponding to the initial question through a large language model to obtain a first result.

[0085] The large language model is a pre-trained language model, which may be but is not limited to GPT (Generative Pre-trained Transformer).

[0086] Step S302: If the first result satisfies the first preset screening condition, the initial question corresponding to the first result is determined as a candidate question.

[0087] Among them, the first preset screening condition needs to be set according to the actual problem screening strategy and is not specifically limited here.

[0088] Step S303: annotate the first dimensional features and the second dimensional features corresponding to the candidate question according to the respective corresponding annotation methods of the first dimensional features and the second dimensional features, and obtain the second result corresponding to the first dimensional features and the third result corresponding to the second dimensional features.

[0089] Step S304: Determine the labeling result corresponding to the initial question according to the second result and the third result.

[0090] Exemplarily, firstly, the first dimension feature corresponding to the initial question is pre-labeled through the large language model to obtain the first result. Based on the first result, the initial question is preliminarily screened to preliminarily filter out low-quality and meaningless questions; specifically, it is determined whether the first result meets the first preset screening condition. If the first result meets the first preset screening condition, the initial question corresponding to the first result is determined as a candidate question; if the first result does not meet the first preset screening condition, the initial question corresponding to the first result is discarded; further, according to the respective corresponding labeling methods of the first dimension feature and the second dimension feature, the first dimension feature and the second dimension feature corresponding to the candidate question are labeled to obtain the second result corresponding to the first dimension feature and the third result corresponding to the second dimension feature. Based on the second result and the third result, further screening of the initial question can be achieved to ensure the quality and reliability of the question to be answered.

[0091] In this embodiment, based on the large language model, the initial question is preliminarily screened to obtain multiple corresponding candidate questions, which can effectively improve the accuracy of question screening. The first dimensional features and the second dimensional features corresponding to the candidate questions are further labeled according to the corresponding labeling methods of the first dimensional features and the second dimensional features, which lays the foundation for further screening of candidate questions to ensure high-quality questions to be answered.

[0092] In one embodiment, the initial question is detected according to the labeling result, and when the initial question meets the preset conditions, the initial question is determined as a question to be answered, including the following steps:

[0093] Step 1: If the second result meets the second preset screening condition, the third result is quantified to obtain a quantified labeling result.

[0094] Among them, the second preset screening condition needs to be set according to the actual problem screening strategy and is not specifically limited here.

[0095] The method of quantizing the third result to obtain the quantitative labeling result may be: mapping the third result to a corresponding numerical value according to a preset mapping relationship to obtain the quantitative labeling result. The preset mapping relationship is a mapping relationship between the third result and the numerical value; for example, assuming that the third result is any one of very easy, easy, uncertain, difficult, and very difficult, very easy can be mapped to 1, easy can be mapped to 2, uncertain can be mapped to 3, difficult can be mapped to 4, and very difficult can be mapped to 5.

[0096] Step 2: Calculate the reliability corresponding to the initial question based on the quantitative labeling results.

[0097] Step 3: If the reliability is greater than the preset reliability threshold, the initial question is determined as the question to be answered.

[0098] The reliability is used to characterize the correctness and reliability of the initial question. The preset reliability threshold needs to be set according to actual needs and is not specifically limited here.

[0099] In an exemplary embodiment, it is assumed that the first dimension features include question quality, subject classification, question type, and negative type, and the second dimension features include question difficulty, question creativity, reasoning and calculation complexity, content completeness, typesetting neatness, seriousness, politeness, and harmfulness; the question quality, subject classification, question type, and negative type are annotated according to the annotation method corresponding to the first dimension features to obtain the second result corresponding to the first dimension features; the question difficulty, question creativity, reasoning and calculation complexity, content completeness, typesetting neatness, seriousness, politeness, and harmfulness are annotated according to the annotation method corresponding to the second dimension features to determine the third result corresponding to the second dimension features; it is determined whether the second result meets the second preset filtering condition; if the second result meets the second preset filtering condition, the third result is quantified to obtain a quantitative annotation result; according to the quantitative annotation result, the reliability corresponding to the initial question is calculated, and the specific calculation is shown in formula (1).

[0100]

[0101] Among them, S represents the reliability corresponding to the initial question; N represents the number of annotations; S 难,i The quantitative annotation result indicating the difficulty of the question; S 创,i The quantitative annotation result representing the creativity of the problem; S 复,i The quantitative annotation result indicating the complexity of inference calculation; S 完,i The quantitative annotation result indicating the completeness of the content; S 严,i The quantitative labeling result indicating the seriousness; S 礼,i The quantitative annotation result of politeness; S 害,i A quantitative labeling result indicating the harmfulness.

[0102] Further, determine whether the reliability S corresponding to the initial question is greater than a preset reliability threshold; if the reliability S is greater than the preset reliability threshold, the initial question is determined as the question to be answered; if the reliability S is less than or equal to the preset reliability threshold, the initial question is discarded.

[0103] In this embodiment, when the second result satisfies the second preset screening condition, the third result is quantified to obtain a quantitative labeling result; then, based on the quantitative labeling result, the reliability corresponding to the initial question is calculated; the reliability is compared with the preset reliability threshold, and the initial question is further screened, thereby effectively improving the reliability and accuracy of the questions to be answered.

[0104] In one embodiment, Figure 4 As shown, Figure 4 The flowchart of the answer set acquisition step in one embodiment is as follows; acquiring the answer set corresponding to the question to be answered includes the following steps:

[0105] Step S401: for each question to be answered, determine a plurality of target processing objects that match the question to be answered.

[0106] The target processing object is an expert user who is good at the target subject.

[0107] Step S402: extracting the first answer input by each target processing object for the question to be answered.

[0108] Among them, the first answer refers to the professional answer given by the target processing object to the question to be answered.

[0109] It is understandable that the first answer can be, but is not limited to, stored in a memory in the form of text or in a data structure; in an exemplary embodiment, the first answer is stored in a text in Markdown format. In another exemplary embodiment, the first answer is stored in a data structure, and when generating a question-answering corpus, the first answer input by each target processing object for the question to be answered can be extracted based on the data structure to obtain multiple first answers corresponding to the question to be answered.

[0110] Step S403: Process the question to be answered by using the large language model to obtain a second answer corresponding to the question to be answered.

[0111] Exemplarily, by calling the large language model to process the question to be answered, a second answer corresponding to the question to be answered can be obtained. For example, by calling the large language model with specially designed key prompts, a highly accurate second answer can be obtained; wherein the key prompts may include, but are not limited to, subject prompts, question types, answer content depth, and the like.

[0112] Step S404, obtaining an answer set corresponding to each question to be answered based on the multiple first answers and the second answers.

[0113] In this embodiment, for each question to be answered, based on the first answers input by multiple target processing objects and the second answers obtained through the large language model, the answer set corresponding to each question to be answered can be determined, laying the foundation for generating accurate and reliable question-answer corpus pairs.

[0114] In one embodiment, determining a plurality of target processing objects that match the question to be answered includes the following steps:

[0115] Step 1: When the question to be answered is in the state of being to be assigned, obtain the assignment waiting time corresponding to the question to be answered.

[0116] The status of the question to be answered includes the state to be assigned and the state to be accepted. The assignment waiting time refers to the time interval for the question to be answered to switch from the state to be assigned to the state to be accepted.

[0117] Exemplarily, the method for obtaining the allocated waiting time corresponding to the question to be answered may be: timing the allocated waiting time corresponding to the question to be answered in real time by means of a counter.

[0118] Step 2: If the assigned waiting time is less than the preset time threshold, and there is a first processing object that accepts the question to be answered, the question to be answered is assigned to the first processing object, and the first processing object is used as the target processing object.

[0119] The subject level corresponding to the subject to which the first processing object belongs is the same as the subject level corresponding to the target subject. For example, if the subject level corresponding to the target subject is a third-level subject, the subject level corresponding to the subject to which the first processing object belongs is also a third-level subject.

[0120] Among them, the preset time threshold needs to be set according to the actual allocation requirements of the questions to be answered, and is not specifically limited here.

[0121] In an exemplary embodiment, when the question to be answered is received by the first processing object, a corresponding flag bit is generated. Based on the flag bit, it can be determined whether the status of the question to be answered has changed; for example, by identifying the flag bit corresponding to the question to be answered, it is determined whether the status of the question to be answered has switched from a to-be-assigned state to an accepted state.

[0122] Step 3: If the assigned waiting time is greater than or equal to the preset time threshold, and there is no first processing object accepting the question to be answered, the question to be answered is assigned to the second processing object, and the second processing object is used as the target processing object.

[0123] Among them, the subject level corresponding to the subject of the second processing object is higher than the subject level corresponding to the target subject. Exemplarily, the subject levels are first-level subject, second-level subject, and third-level subject from high to low; assuming that the subject level corresponding to the target subject is a third-level subject, then the subject level corresponding to the subject of the second processing object is higher than the third-level subject, that is, the subject of the second processing object can be a second-level subject or a first-level subject.

[0124] Exemplarily, assuming that the subject level corresponding to the target subject is a third-level subject, when the question to be answered is to be assigned, obtain the assigned waiting time corresponding to the question to be answered. Determine whether the assigned waiting time is less than the preset time threshold. If the assigned waiting time is less than the preset time threshold, and there is a first processing object of the third-level subject that accepts the question to be answered, the question to be answered is assigned to the first processing object, and the first processing object is used as the target processing object; if the assigned waiting time is greater than or equal to the preset time threshold, and there is no first processing object of the third-level subject that accepts the question to be answered, the question to be answered is assigned to the second processing object of the second-level subject, and the second processing object of the second-level subject is used as the target processing object. Further, if there is no second processing object of the second-level subject that accepts the question to be answered, the question to be answered is assigned to the second processing object of the first-level subject, and the second processing object of the first-level subject is used as the target processing object. Based on this, discipline diffusion is achieved to ensure that the question to be answered is effectively accepted.

[0125] It should be noted that, for the above-mentioned first processing object and second processing object, it is necessary to set in advance the subjects that the first processing object and the second processing object are respectively good at, so as to lay a foundation for the accurate allocation of questions to be answered.

[0126] In this embodiment, when the allocation waiting time is greater than or equal to the preset time threshold and there is no first processing object to accept the question to be answered, the question to be answered is assigned to the second processing object through the subject diffusion mechanism, and the second processing object is used as the target processing object. Based on this, when the question to be answered is not received by the first processing object, it can be reasonably spread from a small range to a larger range to match the corresponding second processing object, thereby increasing the exposure surface of the question to be answered and improving the probability of the question to be answered being accepted.

[0127] In one embodiment, obtaining the accuracy of each answer in the answer set includes the following steps:

[0128] Step 1: for each first answer among multiple first answers, obtain a first accuracy corresponding to each first answer.

[0129] The first accuracy is used to characterize the correctness and reliability of the first answer.

[0130] In an exemplary embodiment, the method for obtaining the first accuracy corresponding to each first answer may be: determining the accuracy calculation method corresponding to the first answer, performing accuracy analysis on the first answer according to the accuracy calculation method, and obtaining the first accuracy corresponding to the first answer.

[0131] Among them, the accuracy calculation method is related to the nature of the first answer and is not specifically limited here. For example, the accuracy of the first answer can be determined from multiple dimensions; illustratively, the multiple dimensions include: whether the answer is spam, that is, whether the text content is a reply irrelevant to the content of the question to be answered; determining whether the question contains negative labels; correctness, that is, whether the answer content is correct; creativity, that is, whether the answer is given with novel ideas and methods; completeness, that is, the degree of incomplete content in the answer content, for example, the knowledge points provided are not comprehensive; conciseness, that is, whether there is content that deviates from the topic of the answer; typesetting neatness, that is, whether the typesetting and punctuation are used correctly; for example, whether paragraphs, chapters, tables, formulas, etc. are expressed in the correct form; seriousness, that is, whether the answer is serious or humorous; politeness, that is, whether the answer is polite, whether it shows an unfriendly attitude towards the other party of the conversation; harmfulness, that is, whether the information contains a clear description of the harmful behavior, whether it beautifies, encourages or downplays the harmful behavior.

[0132] Step 2: Obtain a second accuracy corresponding to the second answer.

[0133] The second accuracy is used to characterize the correctness and reliability of the second answer.

[0134] It should be noted that the method for determining the second accuracy is the same as the method for determining the first accuracy in principle, and will not be described in detail here.

[0135] Step 2: Obtain the accuracy of each answer in the answer set based on the multiple first accuracies and the second accuracies.

[0136] For example, assuming that the answer set corresponding to the question to be answered includes three first answers and one second answer, the first accuracy corresponding to each of the three first answers and the second accuracy corresponding to the second answer are obtained, and the accuracy of each answer in the answer set is obtained according to the three first accuracies and the second accuracy. Specifically, according to formula (2), the first accuracy corresponding to the first answer can be determined.

[0137]

[0138] Among them, S j represents the first accuracy of the jth first answer; M represents the number of first answers; M represents the number of accuracy analyses; S 正,ij represents the accuracy of the jth first answer in the i-th accuracy analysis; S创,ij It represents the accuracy of the j-th first answer in the i-th accuracy analysis regarding creativity, and the remaining accuracies are similar and so on, which will not be repeated here.

[0139] Similarly, the calculation principle of the second accuracy corresponding to the second answer is the same as the calculation principle of the first accuracy corresponding to the first answer recorded in the above embodiment, and will not be repeated here.

[0140] In this embodiment, based on multiple dimensions, the first accuracy corresponding to each first answer and the second accuracy corresponding to the second answer can be accurately determined, which lays a foundation for improving the accuracy of the question and answer corpus.

[0141] In one embodiment, generating a target question-answer corpus pair according to the question to be answered, the answer set, and the accuracy of each answer includes the following steps:

[0142] Step 1: determine the highest accuracy from multiple first accuracies and second accuracies, and use the answer corresponding to the highest accuracy as the target answer.

[0143] Step 2: Generate a target question-answer corpus pair based on the target answer and the question to be answered corresponding to the target answer.

[0144] Step 3: Store the target question-answer corpus according to a preset format.

[0145] Among them, the preset format needs to be set according to actual needs and is not specifically limited here.

[0146] Exemplarily, based on multiple first accuracies and second accuracies, the highest accuracy is determined from the multiple first accuracies and the second accuracies, and the answer corresponding to the highest accuracy is used as the target answer. Further, based on the target answer and the question to be answered corresponding to the target answer, a target question-answer corpus pair is generated, and the target question-answer corpus pair is stored in a preset format to provide accurate question-answer corpus for model training.

[0147] In this embodiment, the answer corresponding to the highest accuracy is used as the target answer, and a target question-answer corpus pair is generated according to the target answer and the question to be answered corresponding to the target answer, thereby improving the accuracy and reliability of the target question-answer corpus pair.

[0148] In one of the embodiments, the question-answer corpus generation method is applied to a question-answer corpus generation system; the question-answer corpus generation system may include, but is not limited to, a user login module, a task processing module, a corpus generation module, and a corpus display module.

[0149] Among them, the user login module is used to handle user registration, login, and password management; the task processing module is used to obtain and display questions to be answered after the user logs in, and to submit the first answers corresponding to the questions to be answered; the corpus generation module is used to execute the question and answer corpus generation method in any of the above embodiments; the corpus display module is used to display the target question and answer corpus pair.

[0150] In a specific embodiment, see Figure 5 When each target processing object logs into the question and answer corpus generation system for the first time, it is necessary to set the corresponding subject in advance so that the corresponding questions to be answered can be accurately received during the subsequent question and answer corpus generation.

[0151] In a specific embodiment, each target processing object can provide corresponding initial questions for the corresponding subject to construct an initial question set, which lays a foundation for improving the accuracy of the question-answering corpus. It should be noted that the question-answering corpus generation system supports text in Markdown format.

[0152] It should be noted that in order to prevent certain popular disciplines from generating a large number of questions within a time window, which would cause the final corpus to be seriously skewed and affect the effect of large model training; the question-answering corpus generation system needs to control the proportion of questions in the same discipline when distributing questions to be answered, ensuring that the proportion of questions is not higher than the preset discipline balance threshold; the part exceeding the preset discipline balance threshold will be delayed until the next time window for distribution, based on which the discipline diversity of corpus production can be improved. Among them, the preset discipline balance threshold needs to be set according to the actual problem distribution situation, and is not specifically limited here. For example, the preset discipline balance threshold can be 30%.

[0153] In a specific embodiment, see Figure 6 ,For the initial question "What is the difference between a geyser and a hot spring? -2", multiple dimensional features such as question quality, subject classification, question type, negative type, question difficulty, question creativity, reasoning and calculation complexity, content completeness, layout neatness, seriousness, politeness, and harmfulness are annotated according to their corresponding annotation methods to obtain the final annotation result corresponding to the initial question.

[0154] In a specific embodiment, in order to support the training corpus for the multi-round dialogue capability of a large model, the system supports asking additional questions to the target question-answer corpus pair, i.e., question continuation. The question-answer corpus generation method recorded in any of the above embodiments is used to generate the target question-answer corpus pair corresponding to the additional question. The specific process will not be repeated here.

[0155] The above-mentioned question-answer corpus generation method obtains an initial question set corresponding to the target subject; extracts multi-dimensional features of each initial question in the initial question set, and determines the labeling method corresponding to each dimensional feature; by labeling each initial question in the initial question set with multi-dimensional features, the labeling result corresponding to the initial question can be obtained, which can ensure that the key information of each question is fully captured, so as to improve the accuracy of the initial question and lay a foundation for improving the accuracy of the question-answer corpus; further, the initial question is detected according to the labeling result, and when the initial question meets the preset conditions, the initial question is determined as the question to be answered; by setting the preset conditions to screen the questions to be answered, the questions that do not meet the preset conditions can be effectively excluded, and the quality of the questions finally selected for generating the question-answer corpus pairs can be ensured, which helps to improve the pertinence and reliability of the answers; further, the answer set corresponding to the question to be answered and the accuracy of each answer in the answer set are obtained; based on the accuracy of the answer, the quality and reliability of the answer can be further guaranteed; finally, according to the question to be answered, the answer set and the accuracy of each answer, the target question-answer corpus pair is generated, which can effectively improve the accuracy and reliability of the generated target question-answer corpus pair.

[0156] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0157] Based on the same inventive concept, the embodiment of the present application also provides a question-answer corpus generation device for implementing the question-answer corpus generation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more question-answer corpus generation device embodiments provided below can refer to the limitations of the question-answer corpus generation method above, and will not be repeated here.

[0158] In an exemplary embodiment, Figure 7 As shown, a question-answer corpus generation device is provided, including: a question acquisition module 701, a feature extraction module 702, a labeling module 703, a determination module 704, an answer acquisition module 705 and a generation module 706;

[0159] The question acquisition module 701 is used to acquire an initial question set corresponding to the target subject;

[0160] The feature extraction module 702 is used to extract the multi-dimensional features of each initial question in the initial question set, and determine the labeling method corresponding to each dimensional feature;

[0161] The labeling module 703 is used to label each dimension feature according to the corresponding labeling method for each initial question, so as to obtain the labeling result corresponding to the initial question;

[0162] A determination module 704 is used to detect the initial question according to the marking result, and determine the initial question as a question to be answered if the initial question meets a preset condition;

[0163] The answer acquisition module 705 is used to obtain the answer set corresponding to the question to be answered and the accuracy of each answer in the answer set;

[0164] The generation module 706 is used to generate a target question-answer corpus pair according to the question to be answered, the answer set and the accuracy of each answer.

[0165] The above-mentioned question-answer corpus generation device obtains an initial question set corresponding to the target subject; extracts multi-dimensional features of each initial question in the initial question set, and determines the labeling method corresponding to each dimensional feature; by labeling each initial question in the initial question set with multi-dimensional features, the labeling result corresponding to the initial question can be obtained, which can ensure that the key information of each question is fully captured, so as to improve the accuracy of the initial question and lay a foundation for improving the accuracy of the question-answer corpus; further, the initial question is detected according to the labeling result, and when the initial question meets the preset conditions, the initial question is determined as the question to be answered; by setting the preset conditions to screen the questions to be answered, the questions that do not meet the preset conditions can be effectively excluded, and the quality of the questions finally selected for generating the question-answer corpus pair can be ensured, which helps to improve the pertinence and reliability of the answers; further, the answer set corresponding to the question to be answered and the accuracy of each answer in the answer set are obtained; based on the accuracy of the answer, the quality and reliability of the answer can be further guaranteed; finally, according to the question to be answered, the answer set and the accuracy of each answer, the target question-answer corpus pair is generated, which can effectively improve the accuracy and reliability of the generated target question-answer corpus pair.

[0166] In one embodiment, the multi-dimensional features include first-dimensional features and second-dimensional features; the labeling module 703 is used to

[0167] Through the large language model, the first dimension features corresponding to the initial question are pre-labeled to obtain the first result;

[0168] If the first result satisfies the first preset screening condition, determining the initial question corresponding to the first result as a candidate question;

[0169] The first dimension features and the second dimension features corresponding to the candidate question are labeled according to the labeling methods corresponding to the first dimension features and the second dimension features, so as to obtain the second result corresponding to the first dimension features and the third result corresponding to the second dimension features;

[0170] According to the second result and the third result, a labeling result corresponding to the initial question is determined.

[0171] In one embodiment, the determination module 704 is further configured to

[0172] If the second result meets the second preset screening condition, the third result is quantified to obtain a quantified annotation result;

[0173] According to the quantitative annotation results, calculate the reliability corresponding to the initial question;

[0174] If the reliability is greater than a preset reliability threshold, the initial question is determined as a question to be answered.

[0175] In one embodiment, the answer acquisition module 705 is also used to

[0176] For each question to be answered, determine multiple target processing objects that match the question to be answered;

[0177] Extracting the first answer input by each target processing object to the question to be answered;

[0178] The question to be answered is processed through a large language model to obtain a second answer corresponding to the question to be answered;

[0179] According to the multiple first answers and the second answers, an answer set corresponding to each question to be answered is obtained.

[0180] In one embodiment, the answer acquisition module 705 is also used to

[0181] When the unanswered question is in the state of being to be assigned, obtaining the assignment waiting time corresponding to the unanswered question;

[0182] If the allocation waiting time is less than the preset time threshold, and there is a first processing object that accepts the question to be answered, the question to be answered is allocated to the first processing object, and the first processing object is used as the target processing object; the subject level corresponding to the subject to which the first processing object belongs is the same as the subject level corresponding to the target subject;

[0183] If the assigned waiting time is greater than or equal to the preset time threshold, and there is no first processing object accepting the question to be answered, the question to be answered will be assigned to the second processing object, and the second processing object will be used as the target processing object; the subject level corresponding to the subject to which the second processing object belongs is higher than the subject level corresponding to the target subject.

[0184] In one embodiment, the answer acquisition module 705 is also used to

[0185] For each first answer among the plurality of first answers, obtaining a first accuracy corresponding to each first answer;

[0186] Obtaining a second accuracy corresponding to the second answer;

[0187] According to the multiple first accuracies and the second accuracies, the accuracy of each answer in the answer set is obtained.

[0188] In one embodiment, the generating module 706 is used to determine the highest accuracy from the plurality of first accuracies and second accuracies, and use the answer corresponding to the highest accuracy as the target answer;

[0189] Generate a target question-answer corpus pair based on the target answer and the question to be answered corresponding to the target answer;

[0190] The target question-answer corpus is stored in a preset format.

[0191] Each module in the above-mentioned question-answer corpus generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0192] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to the generation of question and answer corpus. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for generating question and answer corpus is implemented.

[0193] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0194] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0195] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0196] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0197] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0198] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0199] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0200] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A question-answer corpus generation method, characterized in that: The method comprises: Obtain the initial set of questions corresponding to the target subject; Extracting multi-dimensional features of each initial question in the initial question set, and determining a labeling method corresponding to each of the dimensional features; For each of the initial questions, annotate each of the dimensional features according to the corresponding annotation method to obtain an annotation result corresponding to the initial question; Detecting the initial question according to the marking result, and determining the initial question as a question to be answered if the initial question meets a preset condition; Obtaining an answer set corresponding to the question to be answered, and the accuracy of each answer in the answer set; A target question-answer corpus pair is generated according to the question to be answered, the answer set, and the accuracy of each answer.

2. The method according to claim 1, characterized in that: The multi-dimensional features include first dimensional features and second dimensional features; for each of the initial questions, each of the dimensional features is annotated according to the corresponding annotation method, and the annotation result corresponding to the initial question is obtained, including: Pre-labeling the first dimensional features corresponding to the initial question using a large language model to obtain a first result; If the first result satisfies a first preset screening condition, determining the initial question corresponding to the first result as a candidate question; The first dimensional feature and the second dimensional feature corresponding to the candidate question are labeled according to the labeling methods corresponding to the first dimensional feature and the second dimensional feature, to obtain a second result corresponding to the first dimensional feature and a third result corresponding to the second dimensional feature; A labeling result corresponding to the initial question is determined according to the second result and the third result.

3. The method according to claim 2, characterized in that The detecting the initial question according to the marking result and determining the initial question as a question to be answered when the initial question meets a preset condition includes: If the second result satisfies the second preset screening condition, quantizing the third result to obtain a quantized annotation result; Calculating the reliability corresponding to the initial question according to the quantitative labeling result; If the reliability is greater than the preset reliability threshold, the initial question is determined as a question to be answered.

4. The method according to claim 1, characterized in that: The obtaining of the answer set corresponding to the question to be answered includes: For each of the questions to be answered, determining a plurality of target processing objects that match the question to be answered; Extracting a first answer input by each target processing object to the question to be answered; Processing the question to be answered by using a large language model to obtain a second answer corresponding to the question to be answered; According to the plurality of the first answers and the second answers, an answer set corresponding to each of the questions to be answered is obtained.

5. The method according to claim 4, characterized in that The determining of a plurality of target processing objects matching the question to be answered includes: When the question to be answered is to be assigned, obtaining the assignment waiting time corresponding to the question to be answered; If the allocation waiting time is less than the preset time threshold, and there is a first processing object that accepts the question to be answered, the question to be answered is allocated to the first processing object, and the first processing object is used as the target processing object; the subject level corresponding to the subject to which the first processing object belongs is the same as the subject level corresponding to the target subject; If the assigned waiting time is greater than or equal to the preset time threshold, and there is no first processing object that accepts the question to be answered, the question to be answered will be assigned to the second processing object, and the second processing object will be used as the target processing object; the subject level corresponding to the subject to which the second processing object belongs is higher than the subject level corresponding to the target subject.

6. The method according to claim 4, characterized in that Obtaining the accuracy of each answer in the answer set, including: For each of the first answers among the plurality of the first answers, obtaining a first accuracy corresponding to each of the first answers; Obtaining a second accuracy corresponding to the second answer; The accuracy of each answer in the answer set is obtained according to the first accuracies and the second accuracies.

7. The method according to claim 6, characterized in that The step of generating a target question-answer corpus pair according to the question to be answered, the answer set, and the accuracy of each answer includes: Determine the highest accuracy from a plurality of the first accuracies and the second accuracies, and use the answer corresponding to the highest accuracy as the target answer; Generate a target question-answer corpus pair according to the target answer and the question to be answered corresponding to the target answer; The target question-answer corpus pair is stored in a preset format.

8. A question-answer corpus generation device, characterized in that: The device comprises: A question acquisition module is used to obtain an initial set of questions corresponding to the target subject; A feature extraction module, used to extract multi-dimensional features of each initial question in the initial question set, and determine a labeling method corresponding to each of the dimensional features; A labeling module, used for labeling each of the dimensional features according to the corresponding labeling method for each of the initial questions, to obtain a labeling result corresponding to the initial question; A determination module, configured to detect the initial question according to the marking result, and determine the initial question as a question to be answered if the initial question meets a preset condition; An answer acquisition module, used to acquire an answer set corresponding to the question to be answered, and the accuracy of each answer in the answer set; A generation module is used to generate a target question-answer corpus pair based on the question to be answered, the answer set and the accuracy of each answer.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intelligent customer service system based on AI questions and answers and large language model

    CN120994785A