Method and device for finely determining double-closed-loop matrix type large model questioning verbal skill
By employing a dual-loop matrix approach, multiple questioning techniques are constructed and combined with evaluation functions and manual review, thus resolving the issue of fluctuating output results in large models and improving the output quality and stability of large models in specific tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, the method of arbitrarily determining the questioning wording by humans leads to significant fluctuations in the output results of large models, affecting the output quality and stability of large models in specific tasks.
A dual-closed-loop matrix approach is adopted. By constructing multiple questioning scripts, calling a large model to obtain the response result set, and combining evaluation functions and manual review, the target questioning scripts suitable for the target task are refined.
It enables a comprehensive and quantifiable evaluation of the effectiveness of questioning techniques, improves the output quality and stability of large models in specific tasks, and reduces the cost of manual intervention.
Smart Images

Figure CN121809686A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large model technology, and in particular to a method and apparatus for precisely determining the questioning language of a double-closed-loop matrix large model. Background Technology
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful capabilities in fields such as natural language processing, intelligent question answering, and text generation. Currently, various general-purpose large models are widely used in diverse industry scenarios. When using large models, questions are typically written manually and input into the model to obtain responses.
[0003] However, research has found that even slight differences in question wording can lead to significant fluctuations in model output. Therefore, the current method of arbitrarily determining question wording manually directly impacts the processing performance of large models. In practical engineering implementation, how to select question wording to guide large models to output accurate, stable, and business-logical results has become a key factor affecting system performance. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method and apparatus for finely determining the questioning scripts of a large model with a dual closed-loop matrix structure. This method can achieve dual-path evaluation of the effectiveness of questioning scripts for the large model and comprehensively evaluate the function value of each questioning script with respect to the evaluation function and the review conclusion, so as to finely determine the target questioning scripts suitable for the target task. Indirectly, the output quality and stability of the large model in a specific task are improved through the target questioning scripts.
[0005] This application provides a method for finely determining the questioning script of a large-scale dual-closed-loop matrix model, the method comprising: Construct multiple questioning scripts for a target task; wherein the target task is a batch task, and the batch task includes multiple task items; For each question, the large model is invoked based on the question to obtain the response result set output by the large model; wherein, the response result set includes the option results and analysis results corresponding to each task item; For the subset of option results included in the response result set corresponding to the question, the subset of option results is compared with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function; wherein, the standard answer set includes the correct option corresponding to each task item; the evaluation function is used to quantitatively evaluate the response effect corresponding to the question; Based on the subset of analysis results included in the response result set corresponding to the question, determine the review conclusion of the question; Based on the function value of each questioning phrase with respect to the evaluation function and the review conclusion of each questioning phrase, the target questioning phrase suitable for the target task is determined from the multiple questioning phrases.
[0006] This application embodiment also provides a device for finely determining the questioning script of a dual-closed-loop matrix-style large model, the device comprising: A construction module is used to construct multiple questioning scripts for a target task; wherein the target task is a batch task, and the batch task includes multiple task items; The acquisition module is used to call the large model for each question and obtain the response result set output by the large model; wherein, the response result set includes the option results and analysis results corresponding to each task item; The first evaluation module is used to compare the subset of option results in the response result set corresponding to the question with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function; wherein, the standard answer set includes the correct option corresponding to each task item; the evaluation function is used to quantitatively evaluate the response effect corresponding to the question; The second evaluation module is used to determine the review conclusion of the question based on the subset of analysis results included in the response result set corresponding to the question. The determination module is used to determine the target questioning script suitable for the target task from the plurality of questioning scripts based on the function value of each questioning script with respect to the evaluation function and the review conclusion of each questioning script.
[0007] This application embodiment also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the fine determination method for questioning techniques in the dual-closed-loop matrix large model described above are performed.
[0008] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the method for finely determining the questioning script of the double-closed-loop matrix large model described above.
[0009] The method and apparatus for finely determining questioning scripts in a dual-closed-loop matrix-style large model provided in this application embodiment achieve a comprehensive and quantifiable evaluation of script quality. On one hand, for each subset of option results corresponding to a questioning script, the function value of each questioning script with respect to the evaluation function is determined by comparing it with the standard answer set, thereby achieving a quantitative evaluation of the effectiveness of the questioning scripts for the large model. On the other hand, for each subset of analysis results corresponding to a questioning script, the review conclusion of each questioning script is determined. Based on the dual-path evaluation of the effectiveness of the questioning scripts, target questioning scripts suitable for the target task are comprehensively determined; indirectly, the output quality and stability of the large model in a specific task are improved through target questioning scripts.
[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart is shown below illustrating a method for finely determining questioning techniques in a dual-closed-loop matrix-style large model, as provided in an embodiment of this application. Figure 2 This illustration shows a structural diagram of a system for finely determining questioning techniques in a dual-closed-loop matrix-style large model, as provided in an embodiment of this application. Figure 3 This illustration shows a structural schematic diagram of a device for finely determining questioning techniques in a dual-closed-loop matrix-style large model provided in an embodiment of this application. Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.
[0014] Research has shown that with the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful capabilities in natural language processing, intelligent question answering, and text generation. Currently, various general-purpose large models are widely used in various industry scenarios. When using large models, questions are typically written manually and input into the model to obtain responses.
[0015] However, research has found that even slight differences in question wording can lead to significant fluctuations in model output. Therefore, the current method of arbitrarily determining question wording manually directly impacts the processing performance of large models. In practical engineering implementation, how to select question wording to guide large models to output accurate, stable, and business-logical results has become a key factor affecting system performance.
[0016] Based on this, this application provides a method for finely determining the questioning scripts of a large model with a dual-loop matrix structure, so as to realize the dual-path evaluation of the effect of questioning scripts on the large model, and comprehensively evaluate the function value of each questioning script with respect to the evaluation function and the review conclusion, so as to finely determine the target questioning scripts suitable for the target task; indirectly, the output quality and stability of the large model in a specific task are improved through the target questioning scripts.
[0017] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for finely determining questioning techniques in a dual-closed-loop matrix-style large model, as provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the method includes: S101. Construct multiple questioning scripts for the target task.
[0018] The target task is a batch task, which includes multiple task items, meaning that multiple task items need to be processed one by one in the same way.
[0019] In one example, the objective task is to record the test results according to the evaluation indicators of each item in the preset inspection report, and to fill in the inspection conclusion in the preset inspection report based on the test results.
[0020] The inspection reports require a significant amount of repetitive processing, with approximately 700 reports per year, each 500 pages long. These reports contain various issues requiring manual review, which is time-consuming and labor-intensive. One type of question involves multiple-choice questions, with each report averaging 3000 questions (3000 items). After completion, the answers must be manually verified, and errors corrected to ensure the accuracy meets the standards. Since manual processing is too time-consuming and labor-intensive, a large-scale model is used for processing. Table 1 below provides an example of an inspection report.
[0021] Table 1 Inspection Report
[0022] As shown in Table 1 above, the objective is to record whether the evaluation index S2 for each item in the pre-defined inspection report meets the requirements (S1), using the degree of compliance as the inspection result, and filling the result into the inspection conclusion S3 in the pre-defined inspection report. The options for the degree of compliance include: compliant, partially compliant, non-compliant, and inapplicable. This process requires a large-scale model to analyze the objective task and the inspection report; the large-scale model will output the analysis process and the final selected answer.
[0023] In this step, multiple question scripts tailored to the target task can be manually constructed; alternatively, a basic question script can be manually constructed, and then multiple question scripts can be derived from the basic question script using natural language processing techniques. In this embodiment, the question script refers to an instruction template for the specific structured task of the target task. Multiple question scripts may have similar or identical semantics, but differ in their wording.
[0024] For example, regarding the above target task, Script 1 is: "Help me determine whether the work done in statement 2 meets the requirements of statement 1? There are only four possible answers: meets, partially meets, does not meet, and does not apply. You only need to choose one answer for me. Statement 1: [S1]. Statement 2: [S2]." Script 2 is: "Help me determine the relationship between the work done in statement 2 and statement 1? There are only four possible answers: meets, partially meets, does not meet, and does not apply. You only need to choose one answer for me. Statement 1: [S1]. Statement 2: [S2]."
[0025] S102. For each question, call the large model according to the question and obtain the response result set output by the large model.
[0026] The response result set includes the option result and analysis result for each task item. The option result for each task item refers to the option selected by the large model from multiple preset options; the analysis result refers to the analytical reasoning process by which the large model obtains the option result.
[0027] In this step, corresponding to the example above, for each question, the question, the evaluation indicators and result records in the preset inspection report are combined into a question statement; the question statement is input into the large model, and the large model, for each item in the preset inspection report, understands the requirements of the evaluation indicators and the content of the result records according to the question, and generates the response result set by comparing the requirements of the evaluation indicators and the content of the result records.
[0028] Here, the evaluation indicators and results of each task item can be filled into the reserved positions of the question script in sequence to form a question statement; then, the question statement is submitted to the large model, that is, the large model is called through API to process the question statement and obtain the option results and analysis results.
[0029] In one example, the processing procedure for a large model is as follows: Answer format: <thinking>*** <thinking / > Option Results <thinking>Okay, let's analyze it step by step.
[0030] 1. Understand the requirements of statement 1. Statement 1 states: The site should be selected in a building with earthquake resistance, wind resistance, and rain protection capabilities.
[0031] This is a site selection or site condition requirement, emphasizing that the building itself possesses these three protective capabilities.
[0032] 2. Understand the content checked in statement 2. Statement 2 provides the following information: The computer room acceptance documents include descriptions of its earthquake resistance, wind resistance, and rain resistance capabilities (with written proof).
[0033] On-site inspection: The roof was not leaking (indicating that the rainproof capability was effective at the time of inspection).
[0034] There are no windows facing outwards (which reduces the possibility of rain and wind entering).
[0035] The accumulation of soil due to wind due to the absence of doors and windows (indicating effective wind protection).
[0036] The walls have no minor damage (which may indirectly indicate that the building structure is intact and is related to its wind and earthquake resistance).
[0037] 3. Comparison and judgment Rainproof: Documentation explains + no leaks on site → Meets requirements.
[0038] Windproof: Documentation provides instructions + No soil accumulation + No windows or windows are properly closed → Meets requirements.
[0039] Earthquake resistance: There are only documented descriptions, without details of on-site structural inspections or seismic assessments. However, if the document is an acceptance document, it can usually prove that the building was constructed according to earthquake-resistant design.
[0040] Statement 1 requires that "the site should be selected in a building with the capability to..." During the verification, based on the acceptance documents (proving the capability) and the absence of any abnormalities on site (indirectly proving the continued validity of the capability), it can be considered that the overall requirements are met.
[0041] Since there is no obvious missing or refutation of evidence, it does not fall under the categories of "partially conforming" or "not conforming".
[0042] 4. Conclusion Documented evidence + no opposing evidence on site → meets the requirements of "having earthquake resistance, wind resistance, and rain resistance".
[0043] <thinking / > Output option result: Matched In this manner, each candidate large model processes each task item one by one, summarizes the test results, and forms the response result set of each candidate large model.
[0044] S103. For the subset of option results included in the response result set corresponding to the question, compare the subset of option results with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function.
[0045] The standard answer set includes the correct options for each task item. Corresponding to the example above, the correct standard answer set can be obtained from historical inspection reports, or a portion of the current inspection report can be manually filled out to obtain the correct standard answer set. The evaluation function is used to quantitatively assess the response effectiveness corresponding to the questioning script.
[0046] In this step, comparing the subset of option results with the standard answer set of the target task allows for the quantification of the accuracy parameter of the option results output by the large model. Based on this accuracy parameter, the function value of the question wording with respect to the evaluation function is then determined. For example, when the average accuracy is in the first interval and the standard deviation of the accuracy is in the second interval, the function value of the evaluation function is determined to be a first preset value.
[0047] In one possible implementation, step S103 may include: S1031. Compare the subset of option results with the standard answer set of the target task to construct a result distribution matrix.
[0048] Wherein, the result distribution matrix Middle position element at This indicates that for the standard answer being the first... The row corresponds to the task item with the preset options, and the option result is the first one. The column corresponds to the number of task items in the preset options.
[0049] In one example, assuming there are four preset options: "Compliant," "Partially Compliant," "Not Compliant," and "Not Applicable," a matrix is used to count the correct and incorrect judgments made by the large model, resulting in a distribution matrix. It is expressed as follows:
[0050] but, arrive The result distribution matrix is formed. The first line , , Using examples to illustrate the meaning of a matrix: The count of cases where the true and correct answer to a question about the wording in the inspection report is "compliant" indicates that the question is being addressed. This indicates that the true and correct answer to the question about the wording in the inspection report is "compliant", and the judgment result output by the large model is also the count of cases where "compliant" is correct. In other words, it is the count of cases where the large model correctly judges the "compliant" option. This indicates that the correct answer to the question regarding the wording in the inspection report is "compliant," but the overall model incorrectly judged it as "partially compliant." In other words, this counts instances where the overall model incorrectly judged the "compliant" option as "partially compliant." Similarly, if the overall model incorrectly judges it as "non-compliant," this is also counted. If mistakenly judged as "not applicable", it will be counted. .
[0051] S1032. Based on the result distribution matrix, construct the accuracy-error rate matrix.
[0052] Wherein, the accuracy-error rate matrix Middle position element at This indicates that the correct option is the [number]. The row corresponds to the task item with the preset options, and the option result is the first one. The columns represent the percentage of tasks corresponding to the preset options. Corresponding to the example above, the accuracy-error rate matrix is represented as follows:
[0053] Similarly, for the accuracy-error rate matrix , with the first row , , Using examples to illustrate the meaning of a matrix: This indicates the percentage of cases where the true and correct answer in the report is "matched," and so on.
[0054] This indicates that the true and correct answer to the question about the wording in the inspection report is "compliant," representing the percentage of correct judgments output by the large model. This indicates the percentage of instances in the inspection report where the correct answer to a question about the wording is "compliant," but the larger model misclassifies it as "partially compliant," and so on.
[0055] Clearly, the accuracy-error rate matrix It contains detailed results of the overall statistics of the accuracy of the large model for each task item in the entire report, including the accuracy-error rate matrix. The sum of the diagonals is the overall accuracy result, also known as the total accuracy. .
[0056] S1033. Based on the accuracy-error rate matrix, determine the function value of the questioning phrase with respect to the evaluation function.
[0057] In this step, you can directly use the overall accuracy rate. Determine the function value of the question regarding the evaluation function.
[0058] In another possible implementation, the evaluation function includes a sovereign weight and secondary weights; the sovereign weight is used to characterize the importance of correct judgments for each category; the secondary weights are used to penalize specific types of misjudgments; based on the sovereign and secondary weights and the accuracy-error rate matrix... The element values corresponding to each weight in the algorithm can be used to determine the function value of the evaluation function through weighted summation. Therefore, S1033 may also include: When the preset options are 4, the evaluation function The formula is expressed as:
[0059] in, , representing the overall accuracy rate; Representing the accuracy-error rate matrix Middle position element at The corresponding weights; Substitute the corresponding element in the accuracy-error rate matrix into the evaluation function. The formula yields the function value of the question phrase with respect to the evaluation function.
[0060] Among them, sovereignty A value greater than 0 indicates the level of importance attached to the condition of "meeting" the criteria. Set to 1. Sovereignty , , Similarly, settings can be adjusted according to the actual situation. For each situation where sovereignty is of greater importance, the corresponding sovereignty weight can be increased accordingly.
[0061] Secondary weight >0 indicates the severity of the penalty for misjudging "complies" as "partially meets"; typically... , and It can be taken as 0.1* In addition, the corresponding secondary weight is increased based on which situation warrants a greater penalty.
[0062] Therefore, by combining a total of 16 primary and secondary weights, the emphasis or penalty level for corresponding accuracy can be flexibly adjusted, ultimately affecting the function value of the current script's evaluation function, thus obtaining a quantitative evaluation of the script's effectiveness. Clearly... This reflects the overall accuracy evaluation results obtained from the current script. A larger value indicates a better response effect from the current questioning technique. Defining an evaluation function by combining multiple elements and their corresponding weights in the accuracy-error rate matrix helps to further refine the evaluation of the questioning technique's effectiveness in large-scale models.
[0063] S104. Based on the subset of analysis results included in the response result set corresponding to the question, determine the review conclusion of the question.
[0064] Although selective answers are easy for programs to process, they are essentially discrete classification results and are difficult to reflect the quality of the thinking process of a large model. Therefore, this application also introduces an evaluation of analytical results.
[0065] In this step, the analysis results of the large model for each task under the question are submitted for manual review. Domain experts manually review the thought process returned by the large model, such as judging whether the reasoning is reasonable, the logic is rigorous, and the evidence is sufficient, and finally obtain the manual review conclusion for the question. The manual review conclusion can be expressed as a text description, ranking, or scoring.
[0066] S105. Based on the function value of each questioning phrase with respect to the evaluation function and the review conclusion of each questioning phrase, determine the target questioning phrase suitable for the target task from the multiple questioning phrases.
[0067] In this step, the function value of each question phrase with respect to the evaluation function and the review conclusion can be combined to select the question phrases that perform well in both aspects as the target question phrases suitable for the target task. By sorting the question phrases, the top one or more best phrases are usually selected as the target question phrases.
[0068] Furthermore, following S103, the method also includes: Based on the function value of each question phrase with respect to the evaluation function, a number of first question phrases are initially selected from the multiple question phrases; for the subset of analysis results included in the response result set corresponding to each first question phrase, the review conclusion of each first question phrase is determined; based on the review conclusion of each first question phrase, the target question phrase suitable for the target task is determined from the multiple first question phrases.
[0069] This embodiment of the application takes into account that, since computer processing of analytical answers is very poor, manually analyzing the correctness of the analysis results corresponding to the current question in the large model processing is more reliable; however, manual analysis is slow, time-consuming, and costly. Therefore, in this implementation, a small number of high-quality first question questions can be initially screened out by the function value of the evaluation function; then, the subset of analysis results of the first question questions is reviewed, and the target question questions are finally selected from the first question questions, generally the top 3 best questions. At this time, avoiding the review of the analysis results of all question questions can reduce the workload of review and improve the efficiency and effectiveness of determining the question questions in the large model.
[0070] In this way, the results of the large model are filtered through two types of closed loops, namely selective loop and analytical loop, so that the evaluation of the effectiveness of the rhetoric is both realistic and in line with human experience and intuition.
[0071] Furthermore, the method also includes: Based on the function value of each question phrase with respect to the evaluation function and / or the conclusion of the manual review, the question phrases are adjusted; from the adjusted question phrases, the target question phrase suitable for the target task is re-determined.
[0072] In this way, based on the function value of the evaluation function and / or the conclusion of manual review, the adjustment of the questioning script is guided by a closed-loop feedback mechanism. After the adjustment, the target questioning script suitable for the target task is re-determined in the aforementioned manner, thereby achieving iterative optimization of the questioning script and helping to obtain the best form of script expression.
[0073] Please see Figure 2 , Figure 2 This is a schematic diagram illustrating the structure of a system for finely determining questioning techniques in a dual-closed-loop matrix-based large-scale model, as provided in an embodiment of this application. Figure 2 As shown in the figure, the system provided in this application embodiment includes a forward processing path, a feedback processing loop A, and a feedback processing loop B.
[0074] In the forward processing path, the script editor provides a human-computer interaction interface where a person manually edits and sets a series of N different question scripts based on the characteristics of the inspection report being processed. Then, a computer program synthesizes the question scripts and the content of each task item in the inspection report into the actual question statement. Next, the question statement is sent to the large model for processing using the API format publicly available in the large model, i.e., calling the large model's API. The large model used in this embodiment is Qwen3-4B. Finally, the large model returns the response result set to the application via HTTP protocol. The application continues processing through two feedback processing loops A and B, feeding back the effect of each question script to the script editor and updating the effect ranking of each question script.
[0075] Feedback processing loop A is a processing pathway for selective answers (the option results corresponding to each task item). Since selective answers are highly suitable for computer-programmed processing, this loop can be automated without manual intervention, making it particularly suitable for batch tasks with a large number of scripts, reports, and items within those reports. Feedback processing loop A can first perform a coarse screening of the question scripts, obtaining a small number of high-quality scripts. Specifically, feedback processing loop A uses matrices to achieve fine-grained script evaluation and defines an evaluation function to quantitatively assess the response effect produced by the question scripts. Furthermore, feedback processing loop A can also perform closed-loop feedback adjustments to the question scripts.
[0076] Feedback processing loop B is a feedback processing pathway for analytical answers (the analysis results corresponding to each task item). It is particularly suitable for the final screening of the small number of high-quality scripts selected through loop A to obtain the final optimal script ranking. Furthermore, feedback processing loop B can also perform closed-loop feedback adjustments on question scripts.
[0077] This application provides a method for finely determining questioning scripts in a large-scale, dual-closed-loop matrix model. Through feedback processing of the response result set output by the large model using a dual-closed-loop structure, it achieves comprehensive, refined, quantifiable evaluation and continuous optimization of script quality. Specifically, it statistically analyzes the accuracy-error rate matrix for each category, overcoming the limitations of traditional single accuracy indicators. It defines a comprehensive evaluation function including sovereign and secondary weights, allowing for flexible adjustment of the importance of various correct judgments and the penalty for various misjudgments based on different business needs. This makes the quantitative evaluation of scripts more aligned with actual application scenarios and provides a more refined evaluation perspective. Furthermore, it supports continuous iterative updates of the script set, facilitating convergence to the optimal script format.
[0078] Furthermore, the evaluation of selective answers is fully automated, suitable for rapid initial screening of large batches of tasks and a large number of dialogues, significantly reducing the cost of manual intervention. The evaluation of analytical answers focuses on a small number of high-quality candidate dialogues, with manual review of the model's reasoning process to assess their logical rationality and sufficiency of evidence, ensuring that the final selected dialogues are not only highly accurate but also possess good interpretability and credibility. This achieves a tiered screening strategy of "coarse screening + fine screening," balancing efficiency and accuracy.
[0079] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a device for finely determining questioning techniques in a dual-closed-loop matrix-style large model, as provided in an embodiment of this application. Figure 3 As shown, the device 300 includes: The construction module 310 is used to construct multiple questioning scripts for a target task; wherein the target task is a batch task, and the batch task includes multiple task items; The acquisition module 320 is used to, for each question, invoke the large model according to the question and obtain the response result set output by the large model; wherein, the response result set includes the option results and analysis results corresponding to each task item; The first evaluation module 330 is used to compare the subset of option results included in the response result set corresponding to the question with the standard answer set of the target task, and determine the function value of the question with respect to the evaluation function; wherein, the standard answer set includes the correct option corresponding to each task item; the evaluation function is used to quantitatively evaluate the response effect corresponding to the question; The second evaluation module 340 is used to determine the review conclusion of the question based on the subset of analysis results included in the response result set corresponding to the question. The determination module 350 is used to determine the target questioning script suitable for the target task from the plurality of questioning scripts based on the function value of each questioning script with respect to the evaluation function and the review conclusion of each questioning script.
[0080] Furthermore, the option result corresponding to each task item refers to the option selected by the large model from multiple preset options; then, when the first evaluation module 330 compares the subset of option results with the standard answer set of the target task to determine the function value of the question phrase with respect to the evaluation function, the first evaluation module 330 is used to: By comparing the subset of option results with the standard answer set of the target task, a result distribution matrix is constructed; wherein, the result distribution matrix... Middle position element at This indicates that for the standard answer being the first... The row corresponds to the task item with the preset options, and the option result is the first one. The number of task items corresponding to the preset options; Based on the result distribution matrix, a precision-error rate matrix is constructed; wherein, the precision-error rate matrix... Middle position element at This indicates that the correct option is the [number]. The row corresponds to the task item with the preset options, and the option result is the first one. The percentage of tasks corresponding to the preset options; Based on the accuracy-error rate matrix, determine the function value of the questioning phrase with respect to the evaluation function.
[0081] Furthermore, when the first evaluation module 330 determines the function value of the questioning phrase with respect to the evaluation function based on the accuracy-error rate matrix, the first evaluation module 330 is used to: When the preset options are 4, the evaluation function The formula is expressed as:
[0082] in, , representing the overall accuracy rate; Representing the accuracy-error rate matrix Middle position element at The corresponding weights; Substitute the corresponding element in the accuracy-error rate matrix into the evaluation function. The formula yields the function value of the question phrase with respect to the evaluation function.
[0083] Furthermore, when determining the review conclusion of the question / question based on the subset of analysis results included in the response result set corresponding to the question / question, the second evaluation module 340 is used to: The analysis results of the large model under the question were submitted to human review to obtain the human review conclusion of the question.
[0084] Furthermore, the device 300 includes: a screening module; the screening module is used for: Based on the function value of each question word with respect to the evaluation function, a number of first question words are initially selected from the multiple question words; For each first question response set, the analysis result subset is used to determine the review conclusion for each first question response. Based on the review conclusion of each first questioning script, a target questioning script suitable for the target task is determined from the plurality of first questioning scripts.
[0085] Furthermore, the device 300 includes: an adjustment module; the adjustment module is used for: The question wording is adjusted based on the function value of each question wording with respect to the evaluation function and / or the conclusion of the manual review. From the adjusted questioning scripts, a new target questioning script suitable for the stated objective task is determined.
[0086] Furthermore, the target task is to record the inspection results according to the evaluation indicators of each item in the preset inspection report, and to fill in the inspection conclusion in the preset inspection report based on the inspection results; then, when the acquisition module 320 is used to call the large model for each question and obtain the response result set output by the large model, the acquisition module 320 is used for: For each question, the question, the evaluation indicators and result records in the preset inspection report are combined into a question statement; The question is input into the large model, which then interprets the requirements of the evaluation indicators and the recorded content of the result records for each item in the preset inspection report according to the question wording. By comparing the requirements of the evaluation indicators and the recorded content of the result records, the large model generates the response result set.
[0087] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.
[0088] The memory 420 stores machine-readable instructions that can be executed by the processor 410. When the electronic device 400 is running, the processor 410 and the memory 420 communicate via the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the fine determination method of the double closed-loop matrix large model questioning script as described in the above method embodiment can be executed. For specific implementation, please refer to the method embodiment, which will not be repeated here.
[0089] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the fine determination method for the double-closed-loop matrix large model questioning script as described in the above method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0090] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0091] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0092] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0094] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0095] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / thinking> < / thinking>
Claims
1. A method for precisely determining questioning techniques in a dual-closed-loop matrix-based large-scale model, characterized in that, The method includes: Construct multiple questioning scripts for a target task; wherein the target task is a batch task, and the batch task includes multiple task items; For each question, the large model is invoked based on the question to obtain the response result set output by the large model; wherein, the response result set includes the option results and analysis results corresponding to each task item; For the subset of option results included in the response result set corresponding to the question, the subset of option results is compared with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function; wherein, the standard answer set includes the correct option corresponding to each task item; the evaluation function is used to quantitatively evaluate the response effect corresponding to the question; Based on the subset of analysis results included in the response result set corresponding to the question, determine the review conclusion of the question; Based on the function value of each questioning phrase with respect to the evaluation function and the review conclusion of each questioning phrase, the target questioning phrase suitable for the target task is determined from the multiple questioning phrases.
2. The method according to claim 1, characterized in that, The option result corresponding to each task item refers to the option selected by the large model from multiple preset options; then, the subset of option results is compared with the standard answer set of the target task to determine the function value of the question wording with respect to the evaluation function, including: By comparing the subset of option results with the standard answer set of the target task, a result distribution matrix is constructed; wherein, the result distribution matrix... Middle position element at This indicates that for the standard answer being the first... The row corresponds to the task item with the preset options, and the option result is the first one. The number of task items corresponding to the preset options; Based on the result distribution matrix, a precision-error rate matrix is constructed; wherein, the precision-error rate matrix... Middle position element at This indicates that the correct option is the [number]. The row corresponds to the task item with the preset options, and the option result is the first one. The percentage of tasks corresponding to the preset options; Based on the accuracy-error rate matrix, determine the function value of the questioning phrase with respect to the evaluation function.
3. The method according to claim 2, characterized in that, Based on the accuracy-error rate matrix, determine the function value of the questioning phrase with respect to the evaluation function, including: When the preset options are 4, the evaluation function The formula is expressed as: in, , representing the overall accuracy rate; Representing the accuracy-error rate matrix Middle position element at The corresponding weights; Substitute the corresponding element in the accuracy-error rate matrix into the evaluation function. The formula yields the function value of the question phrase with respect to the evaluation function.
4. The method according to claim 2, characterized in that, Based on the subset of analysis results included in the response set corresponding to this question, the review conclusion for this question is determined, including: The analysis results of the large model under this question for each task item are submitted for manual review to obtain the manual review conclusion of this question.
5. The method according to claim 4, characterized in that, After comparing the subset of option results in the response result set corresponding to the question with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function, the method further includes: Based on the function value of each question word with respect to the evaluation function, a number of first question words are initially selected from the multiple question words; For each first question response set, the analysis result subset is used to determine the review conclusion for each first question response set. Based on the review conclusion of each first questioning script, a target questioning script suitable for the target task is determined from the plurality of first questioning scripts.
6. The method according to claim 4, characterized in that, The method further includes: The question wording is adjusted based on the function value of each question wording with respect to the evaluation function and / or the conclusion of the manual review. From the adjusted questioning scripts, a new target questioning script suitable for the stated objective task is determined.
7. The method according to claim 1, characterized in that, The objective task is to record the test results according to the evaluation indicators of each item in the preset inspection report, and to fill in the inspection conclusion in the preset inspection report based on the test results. For each question, the large model is invoked based on that question to obtain the response result set output by the large model, including: For each question, the question, the evaluation indicators and result records in the preset inspection report are combined into a question statement; The question is input into the large model, which then interprets the requirements of the evaluation indicators and the recorded content of the result records for each item in the preset inspection report according to the question wording. By comparing the requirements of the evaluation indicators and the recorded content of the result records, the large model generates the response result set.
8. A device for precisely determining the questioning script of a large-scale dual-closed-loop matrix model, characterized in that, The device includes: A construction module is used to construct multiple questioning scripts for a target task; wherein the target task is a batch task, and the batch task includes multiple task items; The acquisition module is used to call the large model for each question and obtain the response result set output by the large model; wherein, the response result set includes the option results and analysis results corresponding to each task item; The first evaluation module is used to compare the subset of option results in the response result set corresponding to the question with the standard answer set of the target task to determine the function value of the question with respect to the evaluation function; wherein, the standard answer set includes the correct option corresponding to each task item; the evaluation function is used to quantitatively evaluate the response effect corresponding to the question; The second evaluation module is used to determine the review conclusion of the question based on the subset of analysis results included in the response result set corresponding to the question. The determination module is used to determine the target questioning script suitable for the target task from the plurality of questioning scripts based on the function value of each questioning script with respect to the evaluation function and the review conclusion of each questioning script.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the fine determination method for the double-closed-loop matrix large model questioning script as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for finely determining the questioning script of a dual-closed-loop matrix-style large model as described in any one of claims 1 to 7.