Multi-agent based automatic subjective question scoring method and device, and storage medium

By employing a multi-agent collaborative scoring method, which utilizes the different perspectives of an overview agent and a detail review agent, and through logical verification and debate by the supervisory agent, the problem of low accuracy in subjective question scoring is solved, achieving higher scoring accuracy and logical consistency.

CN120471596BActive Publication Date: 2026-04-17人力资源和社会保障部人事考试中心
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
人力资源和社会保障部人事考试中心
Filing Date
2025-07-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for automatically scoring subjective questions lack logical consistency, resulting in low scoring accuracy.

Method used

A multi-agent collaborative scoring method is adopted, including an overview agent, a detail review agent, a logic verification agent, and a supervisory agent. The agents score from different perspectives and debate the results to generate the final score.

Benefits of technology

It improves the accuracy and logic of subjective question scoring, and ensures the rationality of the scoring through multiple safeguards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471596B_ABST
    Figure CN120471596B_ABST
Patent Text Reader

Abstract

The application discloses a multi-agent-based automatic subjective question scoring method and device and a storage medium. In the method, a to-be-scored answer is obtained, a general overview agent is used to score the to-be-scored answer according to a scoring standard and give a scoring reason, a first score and a corresponding scoring reason are obtained, a detail review agent is used to score the accuracy, word choice and clarity of the to-be-scored answer, a second score and a corresponding scoring reason are obtained, a logic verification agent is used to logically verify the score and the corresponding scoring reason, in the case that the score and the corresponding scoring reason pass the logical verification, the general overview agent and the detail review agent are used to debate, a supervision agent is used to supervise the debating process between the general overview agent and the detail review agent, and in the case that the agents reach an agreement, a final score of the to-be-scored answer is generated, thereby improving the accuracy of scoring subjective question answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of intelligent agents and automatic test scoring technology, and in particular to an automatic subjective question scoring method, device and storage medium based on multi-agent systems. Background Technology

[0002] With the continuous development of information technology, various industries have begun the process of intelligentization, but there are still certain challenges in the application of intelligentization in the scenario of automatic scoring of subjective questions.

[0003] Because manually scoring subjective questions from examinees is inefficient and costly, existing technologies include methods for automatically scoring subjective questions. For example, CN118861522A discloses a method for grading subjective questions based on a large model. This method collects and preprocesses subjective question answer data from an existing exam database to construct a high-quality dataset. It designs detailed scoring criteria, manually annotates dimensions such as accuracy, logic, and language expression of the answers, and generates weighted average score labels. Finally, it uses the constructed dataset to fine-tune and train a large-scale grading model based on the Transformer architecture, enabling the model to score subjective questions.

[0004] For example, CN119579123A discloses a leaderless group interview system and method based on a large-scale model multi-agent system. This method includes: generating interview questions by a company's human resources manager and corresponding agents; developing scoring criteria, generating role descriptions and role-reference answer schemes, and generating role prompts; after the interviewee receives the invitation and begins the leaderless group interview process, the process sequentially proceeds to a preparation stage, an independent presentation stage, a free discussion stage, and a concluding statement stage, generating the final interview record; the company's human resources manager sends the interviewee's name, interview questions, the final interview record, and scoring criteria to the scoring agent, and reviews and revises the interview scores and evaluations output by the scoring agent; finally, the backend management terminal automatically sends the interview scores and evaluations to the interviewee, completing the leaderless group interview based on a large-scale model multi-agent system.

[0005] It can be seen that in the existing technology, the method of automatically scoring subjective questions is usually to set the scoring criteria manually. By fine-tuning the large model or letting the large model learn autonomously, the large model scores subjective questions according to the set scoring criteria. That is, the existing technology simply scores subjective questions by the large model, which lacks a certain logic and thus the accuracy of subjective question scoring is low.

[0006] There is currently no effective solution to the technical problem of low accuracy in automatic scoring of subjective questions due to a lack of logical consistency in the existing technologies. Summary of the Invention

[0007] The embodiments of this disclosure provide a method, apparatus, and storage medium for automatic subjective question scoring based on multi-agent systems, in order to at least solve the technical problem in the prior art where automatic subjective question scoring lacks logical consistency, resulting in low accuracy of subjective question scoring.

[0008] According to one aspect of the present disclosure, an automatic subjective question scoring method based on multi-agent technology is provided, comprising: obtaining an answer to be scored, wherein the answer to be scored is an answer given by a target user for a preset subjective question; using an overview agent to score the answer to be scored based on a scoring standard set for the preset subjective question and provide a scoring reason, thereby obtaining a first score and a corresponding first scoring reason; using a detail review agent to score the accuracy, word choice, and clarity of the answer to be scored, thereby obtaining a second score and a corresponding second scoring reason; using a logic verification agent to logically verify the first scoring reason and the second scoring reason; if the first scoring reason and the second scoring reason pass the logic verification, using an overview agent and a detail review agent to conduct a debate and using a supervisory agent to supervise the debate process between the overview agent and the detail review agent; and generating a final score for the answer to be scored if the supervisory agent determines that the overview agent and the detail review agent have reached an agreement during the debate.

[0009] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0010] According to another aspect of the present disclosure, an automatic subjective question scoring device based on multi-agent technology is also provided, comprising: an acquisition module for acquiring an answer to be scored, wherein the answer to be scored is the answer given by a target user for a preset subjective question; an initial scoring module for scoring the answer to be scored and providing a scoring reason based on a scoring standard set for the preset subjective question through an overview agent, thereby obtaining a first score and a corresponding first scoring reason, and scoring the accuracy, word choice, and clarity of the answer to be scored through a detail review agent, thereby obtaining a second score and a corresponding second scoring reason; a verification module for logically verifying the first scoring reason and the second scoring reason through a logic verification agent; a supervision module for conducting a debate between the overview agent and the detail review agent and supervising the debate process between the overview agent and the detail review agent when the first scoring reason and the second scoring reason pass the logical verification; and a final scoring module for generating a final score for the answer to be scored when the supervision agent determines that the overview agent and the detail review agent have reached an agreement during the debate.

[0011] According to another aspect of the present disclosure, an automatic subjective question scoring device based on multi-agent technology is also provided, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions to perform the following processing steps: obtaining an answer to be scored, the answer to be scored being an answer given by a target user for a preset subjective question; scoring the answer to be scored and providing a scoring reason based on a scoring standard set for the preset subjective question by an overview agent, obtaining a first score and a corresponding first scoring reason; and scoring the accuracy, word choice, and clarity of the answer to be scored by a detail review agent, obtaining a second score and a corresponding second scoring reason; performing logical verification on the first scoring reason and the second scoring reason by a logic verification agent; if the first scoring reason and the second scoring reason pass the logical verification, engaging in debate by an overview agent and a detail review agent, and supervising the debate process between the overview agent and the detail review agent by a supervisory agent; and generating a final score for the answer to be scored if the supervisory agent determines that the overview agent and the detail review agent have reached an agreement during the debate.

[0012] In this embodiment, according to the technical solution, the overview agent and the detail review agent provide scores from different perspectives. The logic verification agent can verify the logical rationality of the scores provided by both agents (i.e., the aforementioned logic verification). If the verification passes, the overview agent and the detail review agent engage in debate. A supervisory agent oversees this debate, and once a consensus is reached, a final score is derived. Therefore, compared to the prior art described in the background section, this application provides multiple safeguards for the logical rationality of the final score through the overview agent and the detail review agent providing scores from different perspectives, engaging in debate, supervising the debate, and the logic verification agent's logic verification. This, to a certain extent, improves the accuracy of the scoring of the answers. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation thereof. In the drawings:

[0014] Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure;

[0015] Figure 2This is a flowchart illustrating the automatic subjective question scoring method based on multiple agents according to the first aspect of Embodiment 1 of this disclosure;

[0016] Figure 3 This is a schematic diagram of a multi-agent framework for collaboratively scoring subjective question answers provided in Embodiment 1 of this disclosure;

[0017] Figure 4 This is a schematic diagram of an automatic subjective question scoring device based on a multi-agent system according to the first aspect of Embodiment 2 of this disclosure; and

[0018] Figure 5 This is a schematic diagram of an automatic subjective question scoring device based on multiple agents, according to the first aspect of Embodiment 3 of this disclosure. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] Example 1

[0022] According to this embodiment, an embodiment of a method for automatic subjective question scoring based on multi-agent intelligence is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0023] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a multi-agent-based automatic subjective question scoring method is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0024] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0025] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multi-agent-based automatic subjective question scoring method in this embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the multi-agent-based automatic subjective question scoring method of the aforementioned application. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0026] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0027] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0028] It should be noted here that, in some optional embodiments, the above... Figure 1 The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0029] Under the aforementioned operating environment, according to the first aspect of this embodiment, an automatic subjective question scoring method based on multi-agent systems is provided. This method can be implemented by... Figure 1 The computing device implementation is shown. Figure 2 A flowchart illustrating the method is shown below. (Refer to...) Figure 2 As shown, the method includes:

[0030] S202: Obtain the answers to be scored, which are the answers given by the target users to the preset subjective questions;

[0031] S204: Through the overview agent, based on the scoring criteria set for the preset subjective questions, the answer to be scored is scored and the scoring reasons are given, resulting in a first score and the corresponding first scoring reasons; and through the detail review agent, the accuracy, word choice and clarity of the answer to be scored are scored, resulting in a second score and the corresponding second scoring reasons.

[0032] S206: The logic verification agent performs logical verification on the first and second scoring reasons;

[0033] S208: If the first and second reasons for scoring are logically verified, the debate is conducted by the overview agent and the detail review agent, and the debate process between the overview agent and the detail review agent is supervised by the supervisory agent.

[0034] S210: If the supervising agent determines that the overview agent and the detail review agent have reached a consensus during the debate, generate the final score for the answer to be scored.

[0035] The computing device can obtain the answer to be scored, which may refer to the answer given by the target user to the preset subjective question (S202). Here, the target user may refer to the user who needs to answer the preset subjective question, for example, the student who needs to answer the preset subjective question.

[0036] Then, the computing device can use the overview agent to score the answer to be scored based on the scoring criteria set for the above-mentioned preset subjective questions and give the scoring reasons, thus obtaining a first score and the corresponding first scoring reasons. Additionally, the computing device can use the detail review agent to score the accuracy, word choice, and clarity of the answer to be scored, thus obtaining a second score and the corresponding second scoring reasons (S204).

[0037] The aforementioned overview agent and detail review agent score the answers from different perspectives. The overview agent assesses the overall structure and logical coherence of the answer, ensuring that each scoring point in the scoring criteria is fully explained and clearly argued. The overview agent strictly adheres to the scoring criteria and may award points only if the requirements are fully met. Conversely, the detail review agent scores based on the accuracy, word choice, and clarity of the answer, ensuring rigor through meticulous scoring.

[0038] The scoring criteria (i.e., the scoring criteria set for the preset subjective questions mentioned above) can include several scoring points, each corresponding to a certain number of points. For example, assuming the preset subjective question has a total of 12 points, the scoring criteria includes 4 scoring points, each of which can correspond to 3 points. If the answer to be scored only fully satisfies one scoring point, the overview agent can give the answer 3 points. For a preset subjective question, the corresponding scoring criteria can be pre-labeled manually.

[0039] In this application, subjective questions are scored through the collaboration of multiple agents (in addition to the overview agent and detail review agent mentioned above, there are also logic verification agents and supervisory agents, which will be mentioned later). To enable different agents to perform their respective functions, prompts adapted to their functions can be pre-set for each agent. When an agent needs to perform a task set for it (for example, the overview agent needs to score the answer to be scored based on the scoring criteria, and the detail review agent needs to score the answer to be scored in terms of accuracy, word choice, and clarity), the computing device can input the pre-set prompts for that agent, thereby enabling the agent to perform the task it needs to perform.

[0040] For example, for an overview agent, the prompt can be set to "As a scoring assistant, you need to score the answers to be scored. The full score for this question is 12 points. Scoring must be strictly carried out in accordance with the following guidelines. Only when the requirements are fully met will a score be given" (the guidelines in the example can be given manually based on the scoring criteria).

[0041] For example, for a detail review agent, the prompt can be set to "As a scoring assistant, you need to score the answer to be scored. The full score for this question is 12 points. Scoring must be strictly carried out in accordance with the following guidelines. Only when the requirements are fully met will a score be given" (the guidelines in the example can be given manually based on the scoring direction of the detail review agent).

[0042] Then, the computing device can logically verify the first and second rating reasons through the logical verification agent (S206).

[0043] The logical verification mentioned here refers to verifying whether the first rating reason corresponding to the first rating and the second rating reason corresponding to the second rating are logically reasonable. The specific steps for performing logical verification will be explained in detail in the following steps.

[0044] If the first and second reasons for scoring are logically verified, the computing device can conduct debates between the overview agent and the detail review agent, and supervise the debate process between the overview agent and the detail review agent through a supervisory agent (S208).

[0045] In this debate, the detail review agent can challenge the initial rating and its corresponding reasons (including updated initial ratings and reasons) derived by the overview agent, thereby obtaining challenge information output by the detail review agent. During the debate, both the overview agent and the detail review agent can continuously update their ratings and corresponding reasons for the assigned answer.

[0046] The computing device can input prompts to both the overview agent and the detail review agent to initiate a debate. For example, the prompt to the overview agent could be: "You are a defender of the rating results, firmly upholding your rating results on a reasonable basis and not easily swayed" (the rating results mentioned here can include the rating and the reasons for the rating); the prompt to the overview agent could be: "You are a challenger of the rating results, questioning possible errors made by the overview agent and stating your own point of view."

[0047] Finally, if the supervising agent determines that the overview agent and the detail review agent have reached an agreement during the debate, the computing device can generate the final score of the answer to be scored by the supervising agent (S210).

[0048] The prompt designed for the supervising agent could be, for example: "You are responsible for controlling the debate process between the overview agent and the detail review agent. When both parties reach a consensus, stop the debate and generate the final formatted score." In this example, the formatted score could refer to the final score.

[0049] The formatted rating results can take the form shown below:

[0050] “[3] Reason for scoring: The answer mentions “...", so it gets 3 points.

[0051] [3] Reason for scoring: The answer mentions "..." and contains the keyword "...", therefore it receives 3 points.

[0052] [0] Reason for scoring: The answer did not mention "..." or express "...", therefore no points were awarded.

[0053] In summary, the score is 6 points.

[0054] Figure 3 This is a schematic diagram of a multi-agent framework for collaboratively scoring subjective question answers, provided in Embodiment 1 of this disclosure.

[0055] from Figure 3 As can be seen from the above, this application uses multiple intelligent agents to collaboratively score the answers to subjective questions. The overview agent and the detail review agent score the answers from different perspectives (the overview agent scores the answers as a whole according to a preset scoring standard, while the detail review agent scores the answers in detail). Then, the logic verification agent verifies whether the scores and reasons given by the two agents are reasonable. If the verification is successful, the overview agent and the detail review agent debate each other until the supervisory agent determines that the overview agent and the detail review agent have reached a consensus, and then the supervisory agent gives the final score.

[0056] As described in the background section, in the prior art, the method of automatically scoring subjective questions is usually to manually set the scoring criteria, and then to make the large model score the subjective questions according to the set scoring criteria by fine-tuning the large model or letting the large model learn autonomously. That is, the prior art simply scores subjective questions by the large model, which lacks a certain logic and thus the accuracy of the subjective question scoring is low.

[0057] In view of this, according to the technical solution of this embodiment, the overview agent and the detail review agent provide scores from different perspectives, and the logic verification agent can verify the logical rationality of the scores provided by the two agents respectively (i.e., the aforementioned logic verification). If the verification passes, the overview agent and the detail review agent engage in debate. A supervisory agent oversees the debate between the overview agent and the detail review agent, and once they reach a consensus, a final score is obtained. Therefore, compared to the prior art described in the background section, this application provides multiple safeguards for the logical rationality of the final score by having the overview agent and the detail review agent provide scores from different perspectives, the logic verification agent perform logic verification, the two agents used for scoring engage in debate, and the supervisory agent oversees the debate. This improves the accuracy of the scoring of the answers to be scored to a certain extent.

[0058] It should be noted that when the overview agent scores the answer to be scored, it can semantically match the answer with each scoring point in the scoring criteria. When the answer fully satisfies a scoring point, the score for that scoring point is awarded to the answer. Therefore, the first score and the corresponding scoring reason can be represented as a list of tuples: Result=[(Score i , R i ) | s i ∈ S], where S represents the scoring standard, Score i This indicates the scoring point s in the scoring criteria. i The corresponding score, R i Is with s i The relevant reasons for the rating.

[0059] A similar scoring method can be used for detail review agents. The detail review agent can semantically match the answer to be scored with each scoring point in the scoring criteria. When the detail review agent determines that the answer matches a scoring point in detail (i.e., it matches that scoring point well in terms of accuracy, word choice, and clarity), it assigns the score for that scoring point to the answer. Therefore, the second score and the corresponding scoring reason can also be represented in the above tuple list form: Result=[(Score i , R i ) | s i ∈ S], where S represents the scoring criteria set for the detail review agent, and Score i This indicates the scoring point s in the scoring criteria. i The corresponding score, R i Is with s iThe relevant reasons for the rating.

[0060] Optionally, the logical verification agent performs logical verification operations on the first score and its corresponding scoring reason, and the second score and its corresponding scoring reason. Specifically, this includes: verifying whether there is logical consistency between the scoring reason corresponding to the first score and the answer to be scored, and verifying whether there is internal consistency between the scoring reason of the first score; and verifying whether there is logical consistency between the scoring reason corresponding to the second score and the answer to be scored, and verifying whether there is internal consistency between the scoring reason of the second score.

[0061] In other words, verifying the logical consistency of a scoring reason means verifying whether the scoring reason and the answer to be scored are logically compatible. For example, if a scoring reason states, "The answer to be scored does not contain scoring point A, therefore the score corresponding to scoring point A will not be awarded to the answer to be scored," but the answer to be scored contains content corresponding to scoring point A, then the scoring reason and the answer to be scored are logically incompatible. Verifying the internal consistency of a scoring reason can refer to judging whether the logic of the scoring reason itself is reasonable, and whether the scoring reason matches the corresponding score.

[0062] For example, the prompts designed for a logic verification agent could be: "Your task is to check the logical and semantic consistency of the scoring reasons. The specific requirements are as follows: 1. Detect whether there are any contradictions or illogical content in the scoring reasons. 2. Detect whether the content described in the scoring reasons matches the candidate's answer and whether there are any discrepancies."

[0063] When the logic verification agent concludes that the first scoring reason fails to verify, the overview agent can re-evaluate the answer to be scored based on the output of the logic verification agent, so as to update the first score and the corresponding first scoring reason, and then provide it to the logic verification agent for logic verification.

[0064] The same applies to the second rating. When the logic verification agent concludes that the second rating reason fails to verify, the detail review agent can re-evaluate the answer to be rated based on the output of the logic verification agent, so as to update the second rating and the corresponding second rating reason, and then provide it to the logic verification agent.

[0065] Only when the logic verification agent concludes that the scoring reasons of both the overview agent and the detail review agent are verified can the subsequent debate between the overview agent and the detail review agent proceed.

[0066] Optionally, the method further includes:

[0067] The final score and the first score are input into the overview agent to obtain the reflection information output by the overview agent, which is used to indicate the reasons for the difference between the final score and the first score; the final score and the second score are input into the detail review agent to obtain the reflection information output by the detail review agent, which is used to indicate the reasons for the difference between the final score and the second score; and the reflection information output by the overview agent and the detail review agent is stored in the experience base.

[0068] In other words, after scoring an answer to be scored, the final score and the initial score given by the overview agent (or the detail review agent) can be input into the corresponding agent. This allows the agent to reflect on why its initial score and the final score are inconsistent, and obtain the reflection information output by the corresponding agent. The reflection information is recorded in the experience base. The reflection information stored in the experience base can be used to optimize the scoring of subjective question answers by the overview agent and the detail review agent in the future.

[0069] For an agent (including an overview agent and a detail review agent), the information input when it reflects can include not only the first score and the corresponding reason for the score (the first score for the overview agent and the second score for the detail review agent) and the final score and the corresponding reason for the score, but also intermediate information between the first score and the final score. This intermediate information can include: the output of the logic verification agent and the score and the corresponding reason for each update of the agent between the first score and the final score.

[0070] The aforementioned reflection information may include the corresponding agent's analysis and reasons for the difference between the first score and the final score, as well as the issues that need to be considered when scoring this type of pre-set subjective questions.

[0071] For example, prompts for the overview agent and the detail review agent to reflect could be: "Please analyze the differences between the initial and final ratings and the reasons for these differences. If there are differences, please provide reflections, clearly stating which reasons or conclusions (deviations from the original answer, pointing out aspects that need attention when rating such questions) are part of the rating assistant's experience base."

[0072] Optionally, the operation of debating through an overview agent and a detail review agent specifically includes: if the supervisory agent determines that the overview agent and the detail review agent have not reached a consensus, updating the first score based on the experience base by the overview agent to obtain an updated first score, and updating the second score based on the experience base by the detail review agent to obtain an updated second score; and wherein, if the supervisory agent determines that the overview agent and the detail review agent have reached a consensus during the debate, the operation of generating the final score of the answer to be scored specifically includes: if the supervisory agent determines that the overview agent and the detail review agent have reached a consensus based on the updated first score and the updated second score, generating the final score of the answer to be scored.

[0073] In other words, the experience base obtained in the above steps can be used to update the scores and corresponding reasons for the scores by the overview agent and the detail review agent during the debate process. When the supervisory agent determines that the overview agent and the detail review agent have not reached an agreement (which could mean that the first score and its corresponding reason are not agreed upon with the second score and its corresponding reason), the two agents (i.e., the overview agent and the detail review agent) can update the scores based on the experience base to obtain the updated scores and corresponding reasons for the scores (i.e., the updated first score and the updated second score mentioned above, where the updated first score can correspond to corresponding reasons for the scores, and the updated second score can correspond to corresponding reasons for the scores), and then continue the subsequent debate (the supervisory agent continues to determine whether the two agents have reached an agreement).

[0074] Optionally, the operation of debating through the overview agent and the detail review agent specifically includes: if the supervisory agent determines, based on the updated first score and the updated second score, that the overview agent further updates the updated first score, and the detail review agent further updates the updated second score.

[0075] In other words, in the above steps, the overview agent and the detail review agent update the first score and the second score respectively, resulting in updated first and second scores. The supervisory agent can then determine whether the overview agent and the detail review agent have reached a consensus. If the supervisory agent determines, based on the updated first and second scores, that the overview agent and the detail review agent have not yet reached a consensus, the overview agent further updates the updated first score, and the detail review agent further updates the updated second score (i.e., proceed to the next round of debate until the supervisory agent determines that the overview agent and the detail review agent have reached a consensus, ends the debate, and outputs the final score).

[0076] This can be understood as each update of the rating corresponding to a round of debate. Specifically, the supervisory agent oversees the debate process between the overview agent and the detail review agent. In a round of debate, the detail review agent can question the rating given by the overview agent, thus outputting questioning information, while the overview agent can insist on its own point of view to debate with the detail review agent (during which the overview agent can also output some opinion information). The supervisory agent can then determine whether the overview agent and the detail review agent have reached a consensus. If they have reached a consensus, the debate can end; if they have not reached a consensus, they can continue to the next round of debate. In the case of no consensus, in the next round of debate, the overview agent and the detail review agent can update the rating based on their experience bases. When the supervisory agent determines that the overview agent and the detail review agent have reached a consensus, it can output the final rating and the corresponding reasons for the final rating.

[0077] It should be noted that the supervisory agent can determine whether the overview agent and the detail review agent have reached a consensus in the following ways: for example, by judging whether the viewpoints of the overview agent and the detail review agent are in agreement based on the questioning information output by the detail review agent and the opinion information output by the overview agent; or by judging whether the score and the corresponding reasoning obtained by the overview agent and the detail review agent match.

[0078] It should also be noted that each intelligent agent can be instantiated from a pre-defined large language model. No specific limitations are imposed on the large language model here, and each intelligent agent can be deployed on a computing device. The computing device can enable these intelligent agents to perform corresponding functions by inputting prompts corresponding to the agents.

[0079] In addition, refer to Figure 1As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0080] Therefore, according to this embodiment, the logic and accuracy of scoring subjective question answers can be improved.

[0081] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0083] Example 2

[0084] Figure 4 An automated subjective question scoring device 400 based on multi-agent technology according to a first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 4As shown, the device 400 includes: an acquisition module 410, used to acquire answers to be scored, which are answers given by target users to preset subjective questions; an initial scoring module 420, used to score the answers to be scored and provide scoring reasons based on the scoring criteria set for the preset subjective questions through an overview agent, to obtain a first score and corresponding first scoring reasons, and to score the accuracy, word choice, and clarity of the answers to be scored through a detail review agent, to obtain a second score and corresponding second scoring reasons; a verification module 430, used to logically verify the first and second scoring reasons through a logic verification agent; a supervision module 440, used to conduct a debate between the overview agent and the detail review agent when the first and second scoring reasons pass the logical verification, and to supervise the debate process between the overview agent and the detail review agent through a supervision agent; and a final scoring module 450, used to generate a final score for the answers to be scored when the supervision agent determines that the overview agent and the detail review agent have reached an agreement during the debate.

[0085] Optionally, the verification module 430 is specifically used to verify whether there is logical consistency between the first scoring reason and the answer to be scored, and to verify whether there is internal consistency between the first scoring reason; and to verify whether there is logical consistency between the second scoring reason and the answer to be scored, and to verify whether there is internal consistency between the second scoring reason.

[0086] Optionally, the device 400 further includes: a reflection module 460, configured to input the final score and the first score into the overview agent, obtain reflection information output by the overview agent, the reflection information output by the overview agent being used to indicate the reasons for the difference between the final score and the first score; input the final score and the second score into the detail review agent, obtain reflection information output by the detail review agent, the reflection information output by the detail review agent being used to indicate the reasons for the difference between the final score and the second score; and store the reflection information output by the overview agent and the reflection information output by the detail review agent into an experience base.

[0087] Optionally, the supervision module 440 is specifically used to update the first score based on the experience base by the overview agent when the supervision agent determines that the overview agent and the detail review agent have not reached an agreement, thereby obtaining an updated first score, and to update the second score based on the experience base by the detail review agent, thereby obtaining an updated second score; and wherein, the final scoring module 450 is specifically used to generate a final score for the answer to be scored when the supervision agent determines that the overview agent and the detail review agent have reached an agreement based on the updated first score and the updated second score.

[0088] Optionally, the supervision module 440 is specifically used to further update the updated first score through the overview agent and the updated second score through the detail review agent when the supervision agent determines that there is no consensus between the overview agent and the detail review agent based on the updated first score and the updated second score.

[0089] Therefore, according to this embodiment, the logic and accuracy of scoring subjective question answers can be improved.

[0090] Example 3

[0091] Figure 5 An automatic subjective question scoring device 500 based on multi-agent technology according to a first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 5 As shown, the device 500 includes: a processor 510; and a memory 520, connected to the processor 510, for providing the processor with instructions to perform the following processing steps:

[0092] The process involves: acquiring answers to be scored, which are the target user's responses to preset subjective questions; using an overview agent to score the answers based on the scoring criteria set for the preset subjective questions and providing a reason for the score, resulting in a first reason for scoring; and using a detail review agent to score the accuracy, word choice, and clarity of the answers, resulting in a second reason for scoring; using a logic verification agent to logically verify the first and second reasons for scoring; if the first and second reasons for scoring pass the logic verification, a debate is conducted between the overview agent and the detail review agent, and the debate process is supervised by a monitoring agent; and if the monitoring agent determines that the overview agent and the detail review agent have reached a consensus during the debate, a final score for the answer to be scored is generated.

[0093] Optionally, the logical verification agent performs logical verification on the first and second scoring reasons, specifically including: verifying whether the first scoring reason is logically consistent with the answer to be scored, and verifying whether the first scoring reason has internal consistency; and verifying whether the second scoring reason is logically consistent with the answer to be scored, and verifying whether the second scoring reason has internal consistency.

[0094] Optionally, the memory 520 is also configured to provide the processor 510 with instructions to process the following steps: inputting the final score and the first score into the overview agent, obtaining reflection information output by the overview agent, the reflection information output by the overview agent being used to indicate the reason for the difference between the final score and the first score; inputting the final score and the second score into the detail review agent, obtaining reflection information output by the detail review agent, the reflection information output by the detail review agent being used to indicate the reason for the difference between the final score and the second score; and storing the reflection information output by the overview agent and the reflection information output by the detail review agent into an experience base.

[0095] Optionally, the operation of debating through an overview agent and a detail review agent specifically includes: if the supervisory agent determines that the overview agent and the detail review agent have not reached a consensus, updating the first score based on the experience base by the overview agent to obtain an updated first score, and updating the second score based on the experience base by the detail review agent to obtain an updated second score; and wherein, if the supervisory agent determines that the overview agent and the detail review agent have reached a consensus during the debate, the operation of generating the final score of the answer to be scored specifically includes: if the supervisory agent determines that the overview agent and the detail review agent have reached a consensus based on the updated first score and the updated second score, generating the final score of the answer to be scored.

[0096] Optionally, the operation of debating through the overview agent and the detail review agent specifically includes: if the supervisory agent determines, based on the updated first score and the updated second score, that the overview agent further updates the updated first score, and the detail review agent further updates the updated second score.

[0097] Therefore, according to this embodiment, the logic and accuracy of scoring subjective question answers can be improved.

[0098] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0099] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0100] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0101] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0103] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0104] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-agent based automatic subjective question scoring method, characterized in that, include: Obtain the answers to be scored, which are the answers given by the target users to preset subjective questions; The overview agent scores the answer to be scored based on the scoring criteria set for the preset subjective questions and provides a scoring reason to obtain a first score and a corresponding first scoring reason. The detail review agent scores the accuracy, word choice and clarity of the answer to be scored to obtain a second score and a corresponding second scoring reason. The overview agent is used to evaluate the overall structure and logical coherence of the answer to be scored and to score the answer to be scored based on whether the answer to be scored fully meets the scoring criteria. A logic verification agent is used to perform logical verification on the first rating reason and the second rating reason. Specifically, the operation of logical verification on the first rating reason and the second rating reason by the logic verification agent includes: verifying whether there is logical consistency between the first rating reason and the answer to be rated, and verifying whether there is internal consistency between the first rating reason; and verifying whether there is logical consistency between the second rating reason and the answer to be rated, and verifying whether there is internal consistency between the second rating reason, wherein the internal consistency is used to determine whether the rating reason matches the corresponding rating. If the first and second scoring reasons pass the logical verification, a debate is conducted between the overview agent and the detail review agent, and the debate process between the overview agent and the detail review agent is supervised by a supervisory agent; and If the supervising agent determines that the overview agent and the detail review agent have reached a consensus during the debate, a final score is generated for the answer to be scored. The method further includes: The final score and the first score are input into the overview agent to obtain the reflection information output by the overview agent. The reflection information output by the overview agent is used to indicate the reason for the difference between the final score and the first score. The final score and the second score are input into the detail review agent, and the reflection information output by the detail review agent is obtained. The reflection information output by the detail review agent is used to indicate the reasons for the difference between the final score and the second score; and The reflection information output by the overview agent and the reflection information output by the detail review agent are stored in an experience base, and the operation of debating through the overview agent and the detail review agent specifically includes: In a round of debate, the detail review agent challenges the score given by the overview agent, thereby outputting a challenge message. The overview agent, by adhering to its own viewpoint, outputs corresponding viewpoint information. The supervisory agent is used to determine whether the overview agent and the detail review agent have reached a consensus. If, as determined by the supervisory agent, there is no consensus between the overview agent and the detail review agent, the overview agent updates the first score based on the experience base to obtain an updated first score, and the detail review agent updates the second score based on the experience base to obtain an updated second score; and wherein, The operation of generating a final score for the answer to be scored, whereby the supervising agent determines that the overview agent and the detail review agent have reached a consensus during the debate, specifically includes: When the supervising agent determines that the overview agent and the detail review agent have reached an agreement based on the updated first score and the updated second score, the final score of the answer to be scored is generated.

2. The method of claim 1, wherein, The operation of debating through the overview agent and the detail review agent specifically includes: If the supervising agent determines, based on the updated first score and the updated second score, that there is no consensus between the overview agent and the detail review agent, the overview agent further updates the updated first score, and the detail review agent further updates the updated second score.

3. A storage medium, characterized by The storage medium includes a stored program, wherein the method described in any one of claims 1 to 2 is generated and executed by a processor when the program is run.

4. An input information processing apparatus characterized by comprising: include: The acquisition module is used to acquire answers to be scored, which are the answers given by the target user to preset subjective questions; An initial scoring module is used to score the answer to be scored and provide a scoring reason based on the scoring criteria set for the preset subjective questions through an overview agent, thereby obtaining a first score and a corresponding first scoring reason; and to score the accuracy, word choice and clarity of the answer to be scored through a detail review agent, thereby obtaining a second score and a corresponding second scoring reason. The overview agent is used to score the answer to be scored by evaluating the overall structure and logical coherence of the answer to be scored, and to score the answer to be scored based on whether the answer to be scored fully meets the scoring criteria. The verification module is used to perform logical verification on the first rating reason and the second rating reason through a logical verification agent. Specifically, the operation of performing logical verification on the first rating reason and the second rating reason through the logical verification agent includes: verifying whether there is logical consistency between the first rating reason and the answer to be rated, and verifying whether there is internal consistency between the first rating reason; and verifying whether there is logical consistency between the second rating reason and the answer to be rated, and verifying whether there is internal consistency between the second rating reason, wherein the internal consistency is used to determine whether the rating reason matches the corresponding rating. The supervision module is configured to, when the first and second scoring reasons pass the logical verification, conduct a debate between the overview agent and the detail review agent, and supervise the debate process between the overview agent and the detail review agent through a supervision agent; and The final scoring module is used to generate a final score for the answer to be scored when the supervising agent determines that the overview agent and the detail review agent have reached a consensus during the debate. The device further includes: a reflection module, used to input the final score and a first score to the overview agent, obtain reflection information output by the overview agent, the reflection information output by the overview agent indicating the reasons for the difference between the final score and the first score; input the final score and a second score to the detail review agent, obtain reflection information output by the detail review agent, the reflection information output by the detail review agent indicating the reasons for the difference between the final score and the second score; and store the reflection information output by the overview agent and the reflection information output by the detail review agent in an experience base. The supervising module is used to, in a round of debate... In this process, a detail review agent challenges the rating given by the overview agent, outputting challenge information. The overview agent, in turn, upholds its own viewpoint and outputs corresponding viewpoint information. A supervisory agent determines whether the overview agent and the detail review agent reach a consensus. If the supervisory agent determines that the overview agent and the detail review agent do not reach a consensus, the overview agent updates the first rating based on its experience base, obtaining an updated first rating, and the detail review agent updates the second rating based on its experience base, obtaining an updated second rating. The final rating module, when the supervisory agent determines that the overview agent and the detail review agent reach a consensus based on the updated first and updated second ratings, generates the final rating for the answer to be rated.

5. A multi-agent based automatic subjective question scoring apparatus, characterized by, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Obtain the answers to be scored, which are the answers given by the target user to preset subjective questions; The overview agent scores the answer to be scored based on the scoring criteria set for the preset subjective questions and provides a scoring reason to obtain a first score and a corresponding first scoring reason. The detail review agent scores the accuracy, word choice and clarity of the answer to be scored to obtain a second score and a corresponding second scoring reason. The overview agent is used to evaluate the overall structure and logical coherence of the answer to be scored and to score the answer to be scored based on whether the answer to be scored fully meets the scoring criteria. A logic verification agent is used to perform logical verification on the first rating reason and the second rating reason. Specifically, the operation of logical verification on the first rating reason and the second rating reason by the logic verification agent includes: verifying whether there is logical consistency between the first rating reason and the answer to be rated, and verifying whether there is internal consistency between the first rating reason; and verifying whether there is logical consistency between the second rating reason and the answer to be rated, and verifying whether there is internal consistency between the second rating reason, wherein the internal consistency is used to determine whether the rating reason matches the corresponding rating. If the first and second scoring reasons pass the logical verification, a debate is conducted between the overview agent and the detail review agent, and the debate process between the overview agent and the detail review agent is supervised by a supervisory agent; and If the supervisory agent determines that the overview agent and the detail review agent have reached a consensus during the debate, a final score is generated for the answer to be scored. The memory is also used to provide the processor with instructions to process the following steps: The final score and the first score are input into the overview agent to obtain the reflection information output by the overview agent. The reflection information output by the overview agent is used to indicate the reasons for the difference between the final score and the first score. The final score and the second score are input into the detail review agent, and the reflection information output by the detail review agent is obtained. The reflection information output by the detail review agent is used to indicate the reasons for the difference between the final score and the second score; and The reflection information output by the overview agent and the reflection information output by the detail review agent are stored in the experience base, and the operation of debating through the overview agent and the detail review agent specifically includes: In a round of debate, the detail review agent questions the score given by the overview agent and outputs questioning information, while the overview agent insists on its own point of view and outputs corresponding opinion information. The supervisory agent is used to determine whether the overview agent and the detail review agent have reached a consensus. If, as determined by the supervisory agent, there is no consensus between the overview agent and the detail review agent, the overview agent updates the first score based on the experience base to obtain an updated first score, and the detail review agent updates the second score based on the experience base to obtain an updated second score; and wherein, The operation of generating a final score for the answer to be scored when the supervising agent determines that the overview agent and the detail review agent have reached an agreement during the debate specifically includes: generating a final score for the answer to be scored when the supervising agent determines that the overview agent and the detail review agent have reached an agreement based on the updated first score and the updated second score.

Citation Information

Patent Citations

  • Subjective question scoring method and device based on examination big data and text semantics

    CN116629270A

  • Automatic scoring method and device for answers to subjective questions, electronic equipment and storage medium

    CN118627498A

  • Information reasoning method and device

    CN119990305A