Generative AI evaluation system

The generative AI evaluation system addresses the risk of misinformation by evaluating responses from multiple perspectives, ensuring reliable educational content through integrated scoring and teacher feedback.

JP7756468B1Active Publication Date: 2025-10-20小澤 暢吾

Patent Information

Application Number
JP2025108325
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-20
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Generative AI responses in educational settings may contain misinformation, posing a risk of providing incorrect knowledge to students.

Method used

A generative AI evaluation system that includes a subject consistency evaluation unit, semantic matching evaluation unit, factuality verification unit, and integrated evaluation unit to assess response reliability, with teacher and administrator controls for final confirmation and management.

Benefits of technology

Ensures the provision of highly reliable information by objectively evaluating responses from multiple perspectives, preventing misinformation and supporting educational instruction through teacher feedback and visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756468000001_ABST
    Figure 0007756468000001_ABST
Patent Text Reader

Abstract

The goal is to objectively and multifacetedly evaluate the reliability of the response sentences output by the generative AI. [Solution] The generative AI evaluation system 10 is a generative AI evaluation system that evaluates the reliability of response sentences output by a generative AI, and includes: a subject consistency evaluation unit 12 that calculates a Z-score based on the relationship between the response sentence and subject consistency; a semantic matching evaluation unit 14 that evaluates the semantic proximity between the response sentence and a knowledge base and calculates an RAG score; a factuality verification unit 16 that divides the response sentence into sentences and verifies its factuality by matching it with external information; an integrated evaluation unit 18 that calculates a reliability integrated score by weighted averaging based on the Z-score, RAG score, and the results of the factuality verification; and an output control unit 20 that controls whether to output the response sentence and the display format depending on the reliability integrated score.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a generative AI evaluation system. [Background technology]

[0002] In recent years, generative AI has been increasingly introduced into the field of education, where it is expected to be used for automatic responses to student questions and for creating teaching materials. However, there is a risk that the output of generative AI may contain misinformation or inaccurate descriptions, making it difficult to ensure quality for educational purposes.

[0003] As a technology related to the present invention, for example, Patent Document 1 discloses a computing device operated by at least one processor, which includes a target artificial intelligence model that learns at least one task, performs the task on an input medical image, and outputs a target result, and a confidence prediction model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Special Publication No. 2024-505213 Summary of the Invention [Problem to be solved by the invention]

[0005] The responses output by AI generators may contain false information, which poses a risk of giving students incorrect knowledge, especially in educational applications. Therefore, a system is needed that evaluates the reliability of responses from multiple perspectives and allows teachers to check or regenerate them as necessary.

[0006] The purpose of this invention is to objectively and multifacetedly evaluate the reliability of response sentences output by a generation AI. [Means for solving the problem]

[0007] The generative AI evaluation system of the present invention is a generative AI evaluation system that evaluates the reliability of response sentences output by a generative AI, and is characterized by comprising: a subject consistency evaluation unit that calculates a Z-score based on the relationship between the response sentence and subject consistency; a semantic matching evaluation unit that evaluates the semantic proximity between the response sentence and a knowledge base and calculates a RAG score; a factuality verification unit that divides the response sentence into sentences and verifies factuality by matching with external information; an integrated evaluation unit that calculates a reliability integrated score by weighted averaging based on the Z-score, the RAG score, and the results of the factuality verification; and an output control unit that controls whether to output the response sentence and the display format depending on the reliability integrated score.

[0008] In addition, it is preferable that the generative AI evaluation system of the present invention further includes a teacher evaluation unit that presents at least one of the Z score, the RAG score, the factuality verification result, and the reliability integrated score to a teacher and accepts evaluation operations from the teacher.

[0009] In addition, it is preferable that the generative AI evaluation system of the present invention further comprises a history management unit that can save the Z-score, the RAG score, the factuality verification result, the reliability integrated score, and the response sentence as a history and display it in graph or table format.

[0010] In addition, in the generative AI evaluation system of the present invention, it is preferable that the integrated evaluation unit is equipped with a score weight adjustment means that enables the weighting coefficients of the Z score, the RAG score, and the verification result of factuality to be arbitrarily set or changed.

[0011] In addition, in the generative AI evaluation system of the present invention, when used by multiple users, it is preferable to have an administrator control unit that allows an educational institution or administrator to centrally manage and apply score thresholds, weight settings, and operation policies.

[0012] In addition, it is preferable that the generation AI evaluation system of the present invention further includes a template generation unit that automatically adds a template prompt including a role instruction or tone specification to an input sentence to the generation AI based on user attribute information. [Effects of the Invention]

[0013] This invention makes it possible to objectively and multifacetedly evaluate the reliability of response sentences output by generation AI, thereby preventing the spread of misinformation in educational settings. Furthermore, through final confirmation by teachers and visualization of the history, it also contributes to feedback on learning instruction and individual responses. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a block diagram showing the overall configuration of a generation AI evaluation system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram showing a processing configuration for a teacher to check and evaluate reliability scores, etc., in a generative AI evaluation system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the following, similar elements in all drawings will be designated by the same reference numerals, and duplicate explanations will be omitted. Furthermore, in the description below, previously described reference numerals will be used as necessary.

[0016] Fig. 1 is a block diagram showing the overall configuration of a generative AI evaluation system 10 according to an embodiment of the present invention. Fig. 2 is a diagram showing a processing configuration for a teacher 6 to check and evaluate reliability scores and the like in the generative AI evaluation system 10 according to an embodiment of the present invention.

[0017] The generative AI evaluation system 10 includes a subject consistency evaluation unit 12, a semantic matching evaluation unit 14, a factuality verification unit 16, an integrated evaluation unit 18, an output control unit 20, a teacher evaluation unit 22, a history management unit 24, an administrator control unit 26, and a memory unit 28. The generative AI evaluation system 10 is connected to students 4 and teachers 6 via a network 2.

[0018] The generative AI evaluation system 10 may also include a template guidance unit that automatically adds template prompts to input to the generative AI based on attribute information such as the grade, subject, and level of understanding of the user, the student 4. The template prompts include tone, role, explanation style, and the like, and include instructions such as "Please explain it in an easy-to-understand way for elementary school students" or "Please answer as a science teacher for eighth-grade students." A specific configuration of the template guidance unit is to reflect the user's attributes (grade, subject, learning style, etc.) in selecting a template, which is then added to the beginning of the sentence input to the generative AI. Multiple template prompts are provided corresponding to user attributes, and a selection unit automatically selects the most appropriate style and tone and adds it to the beginning of the sentence input to the generative AI.

[0019] The subject consistency evaluation unit 12 compares the response sentence output by the generation AI with predefined subject-specific knowledge vectors, evaluates whether it is statistically valid, and calculates a Z-score. This Z-score is an index showing consistency with the target subject, and quantitatively evaluates whether the response content is suited to the task.

[0020] The subject consistency evaluation unit 12 may have a function to dynamically select a knowledge vector according to the grade and subject of the student 4. This allows for appropriate consistency evaluation according to the learning stage, improving the educational validity of the output results.

[0021] The semantic matching evaluation unit 14 analyzes the vocabulary and sentence structure contained in the response sentence and evaluates its semantic proximity to the knowledge base. Using the RAG (Retrieval-Augmented Generation) method or similar, it quantifies the similarity with a reliable information source and outputs it as an RAG score. This allows us to evaluate the degree to which the response sentence matches known knowledge.

[0022] The factuality verification unit 16 divides the response text into sentence units, compares each sentence with external knowledge, and determines the truth of the content. It uses a fact-checking method to verify whether the response is based on facts and improve reliability. The fact-checking utilizes natural language matching with external knowledge bases (Wikipedia, educational databases, paper search systems, etc.), and also includes configurations that use BLEU scores, EM scores, etc. as auxiliary indicators.

[0023] The integrated evaluation unit 18 integrates the Z-scores from the subject alignment evaluation unit 12, the RAG scores from the semantic matching evaluation unit 14, and the verification results from the factuality verification unit 16, and calculates a reliability integrated score by performing a weighted average or other calculation. Here, the weighting coefficients for each score can be arbitrarily set or changed by the score weight adjustment means, allowing for flexible adjustments according to the user's objectives and the educational institution's operational policies. The weighting can be adjusted according to the purpose and target.

[0024] The integrated evaluation unit 18 may also be configured to monitor the degree of discrepancy between the Z score and the RAG score, and if the difference exceeds a predetermined threshold, to determine that an evaluation abnormality has occurred and set a warning flag, thereby making it possible to detect cases where the consistency with knowledge is high but the fit with the subject is insufficient.

[0025] The output control unit 20 performs output control such as permitting the output of a response sentence, regenerating it, or displaying it with a warning, depending on the reliability integrated score calculated by the integrated evaluation unit 18. This prevents unreliable outputs from being presented as is.

[0026] The output control unit 20 may be configured to dynamically adjust the vocabulary difficulty, writing style, and availability of visual support of the response sentence according to the reliability integrated score. For example, if the score is 0.8 or higher, the original text is output, if the score is in the middle range (0.5 to 0.8), the vocabulary and expressions are simplified, and if the score is less than 0.5, output is withheld and alternative candidates are presented or confirmation by the teacher is prompted.

[0027] The output control unit 20 may include a vocabulary control unit and a style conversion unit that adjust the vocabulary level, writing style, and presentation format of the response sentence according to the integrated reliability score. The vocabulary control unit converts difficult technical terms and abstract words into simple expressions to promote understanding when the reliability score is medium (e.g., between 0.5 and 0.8). The style conversion unit temporarily suspends response output when the reliability score is low (e.g., below 0.5), and performs processes such as displaying alternative candidates or a simple summary or encouraging the teacher to review the response. Visual support may include configurations that add furigana, illustrations, and a reading function. These processes are realized using templated writing style conversion rules, a difficulty level dictionary, etc.

[0028] The teacher evaluation unit 22 presents the integrated evaluation result as well as each score and verification result to the teacher 6, and accepts operations such as approval, rejection, and comment input by the teacher. This makes it possible to manually supplement the educational validity of the response sentence.

[0029] Furthermore, the value of the teacher evaluation unit 22 as a learning support tool can be further enhanced by configuring it to have a function for displaying explanations with evidence for the response sentences and to add hints and explanation links that students can refer to.

[0030] The teacher evaluation unit 22 is provided to complement the limitations of automatic reliability evaluation by the generation AI, and particularly in educational settings, it is necessary for the teacher 6 to judge the educational appropriateness of the response sentence with his / her expert judgment. This enables the teacher 6 to make a final pass / fail decision based on the appropriateness of each score, rather than simply following the numerical values.

[0031] Furthermore, the teacher evaluation unit 22 may be configured to have an approval guidance control function that allows the selection of explaining the reason for the score and presenting feedback to the student 4 in addition to the decision options of approval, rejection, and regeneration.

[0032] Furthermore, the teacher evaluation unit 22 may be provided with a function that displays the basis for the evaluation score and allows feedback comments to be registered for students, in addition to the decision options of approval, rejection, regeneration, etc. This will further increase the reliability and transparency when using generative AI in educational settings, and will promote collaboration with educational instruction.

[0033] The history management unit 24 has a function to save the Z-score, RAG score, factual verification result, reliability integrated score, and response sentence as history and visualize it in graph or table format as needed. This makes it possible to use it as material for understanding the level of understanding and trends of the student 4.

[0034] The history management unit 24 may have a function to cumulatively record each user's learning history, response history, evaluation tendency, etc., and statistically analyze the correlation with the output result based on the reliability score. Furthermore, in cooperation with the output control unit 20, it is possible to control the presentation of responses that are individually optimized according to the user's attribute information (grade, comprehension evaluation, etc.). In this way, the meta-information acquired by the history management unit 24 can also be used as teaching aids for the teacher 6 and as feedback for system improvement.

[0035] The history management unit 24 may be configured to record and manage the user's learning history and attribute information (e.g., grade, proficiency level), and the output control unit 20 may be configured to adjust the content of the response sentences presented based on this information. This makes it possible to present appropriate information depending on the target user, even if the reliability score is the same. Furthermore, it is desirable to have a configuration in which output adjustment history, such as the type of template prompt, vocabulary conversion, style adjustment, and whether or not visual support was used, is also recorded and visualized as a response history for each user.

[0036] The history management unit 24 also has a function to statistically analyze the usage history of each student 4 and visualize the trends in Z scores, RAG scores, and fact check results using line graphs, heat maps, radar charts, etc. This enables a visual understanding of students' strengths and weaknesses and supports individualized instruction by teachers.

[0037] The history management unit 24 may be configured to have a graph generation function for comparing the progress of reliability scores for each student, a correlation display function with the teacher's feedback history, and a statistical processing function for each lesson. This allows teachers to visually grasp the trends in students' understanding of the entire class or a specific unit, which can be useful for improving learning instruction.

[0038] The history management unit 24 may be provided with a self-tuning configuration that analyzes the transition of the reliability integrated score and enables dynamic adjustment of the score threshold and the weight of the evaluation element used by the output control unit.

[0039] The administrator control unit 26 may be configured to collectively control the operation mode of the entire system (for classes, assignments, self-study, etc.) and to be able to switch threshold settings, display content control, log collection policies, etc. according to the mode.

[0040] The administrator control unit 26 has the function of creating and distributing evaluation templates (e.g., for elementary school students, junior high school students, and high school students), and can standardize and apply output policies, weight settings, thresholds, etc. for each educational institution. This enables unified operation of reliability evaluation according to the educational level of the user.

[0041] The administrator control unit 26 is designed for use with multiple users and has a function that allows educational institutions and administrators to centrally set and reflect score thresholds, weighting of evaluation items, and output policies, enabling flexible operation on an organizational basis.

[0042] The administrator control unit 26 may be configured to set a reliability score threshold in accordance with the operational policy of each educational institution, create templates for weighting each evaluation element, and distribute settings collectively through import / export. This enables the unified application of educational policies nationwide or across organizations, and is expected to lead to fair and efficient use of generative AI. Furthermore, the administrator control unit 26 may be equipped with a template management function that enables the collective setting and distribution of template prompts, controlled vocabulary dictionaries, visual support settings, etc., for each subject or grade.

[0043] The storage unit 28 is a storage device that stores data and control programs used or generated by each of the components, and is located locally or in a cloud environment.

[0044] In the generative AI evaluation system 10, each function other than the memory unit 28 can be realized by either a hardware configuration or a software configuration. For example, when realized by software, these functions can be realized by running application software stored in a recording medium such as RAM, ROM, or a hard disk on a server device that is actually configured with a CPU or MPU, RAM, ROM, etc.

[0045] Next, the effects of the generative AI evaluation system 10 configured as described above will be explained. When a student 4 uses the generative AI to obtain a response to a question or assignment, a prompt is input from the student's terminal to the generative AI, and the generative AI generates a response. The generated response is sent sequentially to the subject alignment evaluation unit 12, the semantic matching evaluation unit 14, and the factuality verification unit 16, which calculate a Z score, RAG score, and fact check result, respectively. These scores are aggregated in the integrated evaluation unit 18, which calculates an integrated reliability score.

[0046] The subject consistency evaluation unit 12 analyzes the vocabulary and sentence structure of the input response sentence and calculates a Z-score by comparing it with the standard knowledge vector defined for each grade and subject. The Z-score is an index that indicates how statistically appropriate the output response sentence is with respect to the learning content.

[0047] The semantic matching evaluation unit 14 compares the content of the response sentence with the knowledge base and evaluates semantic consistency. The RAG score quantifies the degree to which the generated response matches existing reliable information based on the semantic proximity to similar documents and reliable information sources.

[0048] The factuality verification unit 16 verifies the factuality of the response by breaking down the response sentence into sentence units and comparing each sentence with external knowledge. For example, by comparing the sentences with reliable information such as encyclopedias, textbooks, and research paper databases, it determines whether or not the response contains falsehoods or misinformation.

[0049] The integrated evaluation unit 18 may be configured to dynamically adjust the vocabulary level, writing style, and presentation format of the response sentence based on the calculated reliability integrated score. For example, if the score is high, it is output as is, if the score is medium, the vocabulary difficulty is adjusted, and if the score is low, it is regenerated or displayed with a warning. A difficulty dictionary is used for vocabulary adjustment, a template for writing style conversion, and furigana, illustrations, read-out, etc. are used for presentation format.

[0050] Furthermore, the integrated evaluation unit 18 may be configured to monitor the degree of deviation between the Z score and the RAG score, and if the deviation exceeds a predetermined standard, add a warning message to the response sentence, making it easier to identify cases where the knowledge is inconsistent even if it is subjectively appropriate, or where the knowledge is accurate but not in line with the subject context.

[0051] The output control unit 20 may have a mechanism for gradually adjusting the vocabulary level, writing style, and presentation format of the response sentence based on the numerical value of the reliability integrated score. For example, if the score is high, the response is output as is; if the score is medium, the response is converted to simple language and visual support such as diagrams is added; if the score is low, the response is withheld and regenerated or the teacher's approval is requested. This adjustment process includes dictionary-based vocabulary difficulty assessment, template conversion, and speech support.

[0052] The teacher evaluation unit 22 is designed to complement the limitations of automatic evaluation by generative AI, and the subjective and educational judgment of teachers is particularly important in educational settings. It is designed to visually present the reliability integrated score and its components, allowing teachers to ultimately judge and adjust the appropriateness of the response sentence.

[0053] The history management unit 24 accumulates and analyzes the generated AI usage history of each student 4, and this is used by teachers for instruction. For example, changes in a student's learning comprehension level can be visually grasped from the transition of past Z scores and RAG scores, which can be applied to individually optimized education.

[0054] This invention employs a three-tiered evaluation structure: subject-specific consistency evaluation using Z-scores, knowledge consistency evaluation using RAG scores, and fact-checking to verify factuality. Each score evaluates the reliability of the response from a different perspective and complements each other. For example, cases where a response is subject-specific but lacks factual consistency, or vice versa, are appropriately evaluated in the integrated evaluation section.

[0055] Furthermore, this invention allows teachers to input and record individual feedback on responses, enabling the accumulation of information that contributes to student learning support. Also, by utilizing trends in statistically compiled evaluation scores, it is possible to analyze the level of understanding of the entire class or grade level. Additionally, the system's academic mode, which is suitable for research purposes, allows educational institutions to make advanced use of the AI-generated evaluation system.

[0056] The above evaluation results are aggregated in the integrated evaluation unit 18, and a reliability integrated score is calculated by a weighted average. The weight of each score is set according to the purpose and the operational policy of the educational institution, and serves as an overall indicator of reliability.

[0057] In the output control unit 20, if the reliability integrated score is above a certain level, the response sentence is displayed as is, but if it is below a certain level, the response sentence is prompted to be regenerated or displayed with a warning attached, thereby reducing the risk that students receive incorrect information.

[0058] Furthermore, this system can be configured to link with a learning management system (LMS) or educational portal, so that evaluation results are automatically reflected in grade sheets and student guidance records, thereby improving the efficiency of administrative tasks in educational settings.

[0059] In addition, the teacher evaluation unit 22 provides the teacher 6 with an interface that displays a list of Z scores, RAG scores, fact check results, reliability integrated scores, etc., as shown in Figure 2. The teacher 6 can refer to these evaluation results to decide whether to approve or block the response sentence, and can enter comments as necessary.

[0060] Figure 2 shows the interface that teachers can use to view evaluation results, and includes an area where the numerical values ​​and trends of the Z-score, RAG score, fact-check results, and integrated reliability score are all displayed together. Additionally, the interface can also include an explanation of the cause of score fluctuations, instructions for regeneration, a comment input field, feedback history display, and a list of teacher approval history.

[0061] The history management unit 24 records the evaluation scores and output statements for each response in chronological order and visualizes them in the form of charts and tables. This makes it easier for the teacher 6 to grasp the student 4's growth and understanding trends, contributing to educational support.

[0062] Furthermore, the administrator control unit 26 allows for unified management of operational rules by upper-level administrators such as schools and boards of education, and can apply evaluation thresholds, weighting settings, whether or not to regenerate, etc. to all users at once. This allows the generation AI to be used under a consistent operational policy for each institution.

[0063] In an embodiment of the present invention, the administrator control unit 26 may be configured to have an API for linking with external educational systems, and to bidirectionally link the reliability evaluation results with a learning management system (LMS) or a grade evaluation system. This allows for the creation of an environment in which the generated AI evaluation results are used in an integrated manner with other educational activities.

[0064] As described above, according to the present invention, response sentences output by a generative AI are evaluated from multiple perspectives (subject alignment, knowledge verification, and factuality) in an integrated manner, and output is controlled based on the evaluation results, thereby enabling the provision of highly reliable information. Furthermore, a generative AI operation system suited to educational settings is realized in that it allows supplementary evaluation and management by teachers 6 and administrators. The above has described in detail the configuration and effects of the generative AI evaluation system 10. Below, examples are presented to further concretely understand the present invention.

[0065] Example 1 In this embodiment, the output control unit 20 dynamically controls the presentation method of the response sentence based on the reliability integrated score calculated by the integrated evaluation unit 18. If the reliability integrated score is in a predetermined high-reliability range (e.g., 0.8 or higher), the response sentence is presented as is. On the other hand, if the reliability integrated score is in an intermediate range (e.g., 0.5 or higher but less than 0.8), a warning message or supplemental explanation is added to the response sentence. If the score is in a low-reliability range (e.g., less than 0.5), the response sentence is temporarily suspended and presented with a message prompting confirmation by the teacher 6, or a regeneration process is performed. This gradual output control according to the score prevents the provision of inappropriate information to students 4 and ensures reliability in educational settings. This configuration may also include an output assistance mechanism that can dynamically adjust vocabulary level, writing style, and visual support format (furigana, illustrations, read-aloud, etc.) based on user attribute information. The system is web-based and may be configured to operate in a cloud or local environment. The output sentences and their evaluation results are saved as a history, which teachers can refer to visually in graph or table format.

[0066] Example 2 This embodiment employs a response optimization configuration that dynamically adjusts the vocabulary level, writing style, and presentation format of the response sentence based on the reliability integrated score obtained by the integrated evaluation unit 18. It is implemented using a web-based user interface, providing a platform-independent operating environment. Specifically, the following operations are performed: (1) High reliability (score 0.8 or higher): The response sentence is output as is, and visual support information such as diagrams, highlights, and furigana is added as necessary. (2) Medium reliability (score less than 0.5-0.8): Automatically adjusts the vocabulary difficulty of the response text and simplifies the writing style to improve readability. (3) Low reliability (score less than 0.5): The response sentence is temporarily withheld, and multiple alternatives are presented, or confirmation and approval by the teacher 6 is requested. Vocabulary control is performed based on a difficulty dictionary set for each grade, and style conversion is processed using pre-registered template patterns. Furthermore, historical information regarding response adjustment (reliability score, adjustment details, teacher judgment, etc.) is accumulated by the history management unit 24 and used for future evaluation improvements and educational feedback. The reliability integrated score by the integrated evaluation unit 18 is calculated based on subject consistency (Z-score), knowledge consistency (RAG score), and factual verification (fact-check results). Furthermore, visual support may be configured to include phonetic guides, diagrams, and a reading function.

[0067] Example 3 In this embodiment, when the reliability integrated score falls below a certain threshold, visual and audio supplementary information is automatically added to the response sentence to support the student's comprehension. Specifically, difficult words or technical terms not appropriate for the student's grade level are extracted from the response sentence based on a difficulty level dictionary, and furigana is added to the relevant phrases. Furthermore, if the response sentence contains a causal relationship or a procedural structure, that structure is visualized as an illustration and presented along with the response sentence. Furthermore, the entire response sentence can be converted into synthetic speech, and a UI (play button, stop / rewind button, etc.) that provides a reading function can be added. The application and combination of these support functions are controlled based on user profile information, including the student's grade, learning history, and support needs (whether or not special support is required), or settings set by the administrator control unit 26. Additionally, the operation history related to response adjustment and the addition of support functions (reliability score, adjustment type, whether or not the teacher 6 confirmed, etc.) is recorded and accumulated by the history management unit 24 and utilized as basic data for educational support records and learning feedback. This configuration makes it possible to provide AI responses that are safe and educationally considerate while supporting content comprehension, even for learners with reading and writing difficulties and students in the lower grades.

[0068] Example 4 In this embodiment, a template prompt to be added to the generation AI is dynamically selected and automatically inserted based on the grade, subject, and question content of the user, student 4. Template prompts are control statements designed to standardize the output quality of the generation AI and ensure educational validity. For example, they may be formatted as follows: "You are a science teacher for second-year junior high school students. Please answer the following questions clearly and factually." The components of a template prompt are defined by multiple parameters, including the educational level (elementary school, junior high school, etc.), subject (Japanese, mathematics, science, social studies, etc.), tone (politeness, standard), writing style (colloquial, written), and output policy (concise, detailed). The prompt selection unit automatically extracts a corresponding template from the knowledge base based on student 4's profile and input content, and prepends it to the input sentence to the generation AI. In this configuration, the type of template prompt applied, the timing of use, and the corresponding output score (Z-score, RAG score, fact-check result) are recorded in the history management unit 24 for later analysis and effectiveness verification. This allows us to quantitatively understand the educational effectiveness of template prompts and provide feedback to design optimal prompts.

[0069] Example 5 This embodiment employs a configuration that visualizes Student 4's learning history to support understanding of learning trends and educational intervention. The history management unit 24 chronologically records Student 4's questions, the AI's responses, various scores (Z-scores, RAG scores, fact-check results), and output control history. This history information is visually displayed in the following formats on the web-based teacher interface. Specifically, it includes a line graph showing the time series of each score, a bar graph showing the distribution of questions by subject, and a table display listing the teacher's approval and comment history. This allows Teacher 6 to easily grasp Student 4's progress in understanding, strengths and weaknesses by subject, and the need for intervention. Furthermore, this visualization function can be applied to Student 4's own self-reflection, explanations to parents, and progress reports, contributing to improved learning support.

[0070] Example 6 This embodiment employs a configuration that provides individually optimized support to students 4 who have learning difficulties. Specifically, the output control unit 20 selectively adds visual, audio, and structural support based on the student 4's profile information. Support functions include (1) automatic generation of speech-to-speech responses and display of a playback UI, (2) automatic addition of furigana to difficult words, and (3) generation of structural diagrams that diagram procedural explanations and classification structures. These functions are switched on and off depending on the student 4's support needs and controlled by profile information or administrator settings. Furthermore, history information, such as which support functions were applied and their effects, is recorded in the history management unit 24, enabling future support improvements and reflection in teaching plans. This builds a flexible AI support platform that can also be used in special needs education.

[0071] Example 7 In this embodiment, the system is configured to individually adjust and adapt trust assessment and output control parameters based on the student's 4 usage history. The history management unit 24 accumulates each student's past trust scores (Z-scores, RAG scores, fact-check results) and teacher evaluation results, and dynamically controls the following based on the results of statistical analysis: (1) adjusting the score threshold (relaxing for students with a confirmed level of trust and tightening for students prone to misinformation); (2) changing the content of template prompts (e.g., simplifying or inserting supporting sentences); and (3) adjusting the teacher notification level (optimizing notification frequency). This enables individually optimized output management based on the student's 4 trust tendency, preventing excessive intervention and ensuring appropriate instruction. The adjustment results are recorded as a history, contributing to future educational support design and system operation improvements.

[0072] Example 8 This embodiment employs a configuration that allows administrators, such as schools or boards of education, to apply various settings for the entire system in one go. Through the administrator control unit 26, settings such as grade-specific and subject-specific score thresholds, template prompts, dictionary data (vocabulary difficulty dictionaries, NG word lists, etc.), and operation modes (normal, academic, personal) can be centrally managed. This allows for unified assessment policies across multiple classes, grades, and schools, while significantly reducing the burden of individual setting work on teachers in the field. Furthermore, the system is configured to support monitoring of operational performance and ensure traceability through the history audit and operation log export functions on the management screen.

[0073] Example 9 This example demonstrates the application of generative AI not limited to educational applications, but also to public and specialized fields such as business support, government response, and research support. The system is equipped with switching functions, such as adjusting the honorific language of template prompts and adding technical terminology, depending on user attributes, making it adaptable to situations requiring a high level of tone and expertise in the response content. Furthermore, the system automatically checks the consistency of responses with laws, regulations, and business manuals, and presents source information as needed, thereby contributing to ensuring auditability and accountability. All setting change histories and operation logs are saved in chronological order and operated in cooperation with the history management unit 24 and administrator control unit 26 so that administrators can review and audit them at a later date. This enables the creation of a highly reliable and explainable AI-assisted environment in the public and business fields.

[0074] Example 10 In this embodiment, the integrated evaluation unit 18 is configured to customize the weighting of the components, Z-score, RAG score, and fact-check results, when calculating the reliability integrated score. The weighting of each score is optimized according to the purpose of use and the characteristics of the subject, enabling more appropriate evaluation. For example, in the Japanese language field, emphasis is placed on conformity with the subject's expression, so Z-score: 0.5, RAG score: 0.4, and fact-check: 0.1 are set. In contrast, in a subject such as social studies, where factuality is emphasized, a configuration such as Z-score: 0.2, RAG score: 0.3, and fact-check: 0.5 is selected. These weighting settings can be changed from the teacher's UI screen, and the change history is recorded in the history management unit 24, allowing for tracking of setting policies and revision of evaluation policies.

[0075] Example 11 In this embodiment, when a multi-turn dialogue between Student 4 and the generation AI takes place, the system evaluates whether each response is consistent with the context of the entire dialogue history and controls output based on the results. Specifically, the integrated evaluation unit 18 has a function for calculating a new "contextual consistency score" that quantifies the contextual consistency with the response content of the previous turn. If this score falls below a predetermined standard, the prompt reconstruction unit automatically inserts guidance and supplementary statements that are in line with the dialogue context to encourage the generation of a more consistent response. In addition, the history management unit 24 chronologically stores the questions, responses, and evaluation scores for all turns, providing a visualization of the dialogue flow and evaluation on the teacher interface. This enables reliable dialogue support while maintaining the progress of Student 4's understanding and consistency of instruction.

[0076] Example 12 In this example, the AI ​​generates responses by dividing them into sentences or clauses, assigning each section an individual reliability score (Z-score, RAG score, fact-check result), and visually displaying it. The display interface presents the score for each sentence as a color-coded bar or numerical value, and includes review support buttons that allow teachers 6 or administrators to perform localized operations such as "confirm," "correct," and "hold." The system also includes a function to input feedback comments for each sentence, allowing for educational supplementation and guidance. These local review results and operation history are recorded in the history management unit 24 and used to improve output quality in the future and provide individual feedback to students 4. This allows for accurate identification of problems with specific sections in addition to evaluating the overall output text, building a more refined review system. [Explanation of symbols]

[0077] 2 Network, 4 Students, 6 Teachers, 10 Generative AI Evaluation System, 12 Subject Consistency Evaluation Unit, 14 Semantic Matching Evaluation Unit, 16 Factuality Verification Unit, 18 Integrated Evaluation Unit, 20 Output Control Unit, 22 Teacher Evaluation Unit, 24 History Management Unit, 26 Administrator Control Unit, 28 Memory Unit.

Claims

1. A generation AI evaluation system that performs reliability evaluation on a response sentence output by a generation AI, a subject consistency evaluation unit that calculates a Z-score based on the relationship between the response sentence and subject consistency; a semantic matching evaluation unit that evaluates the semantic proximity between the response sentence and a knowledge base and calculates a RAG score; a factuality verification unit that divides the response sentence into sentences and verifies the factuality by comparing the response sentence with external information; an integrated evaluation unit that calculates a reliability integrated score by weighted averaging based on the Z score, the RAG score, and the verification result of the factuality; an output control unit that controls whether to output a response sentence and a display format of the response sentence according to the reliability integrated score; A generative AI evaluation system comprising:

2. 2. The generative AI evaluation system according to claim 1, A generative AI evaluation system further comprising a teacher evaluation unit that presents at least one of the Z score, the RAG score, the factuality verification result, and the reliability integrated score to a teacher and accepts evaluation operations by the teacher.

3. 2. The generative AI evaluation system according to claim 1, The generative AI evaluation system further comprises a history management unit that stores the Z score, the RAG score, the factuality verification result, the reliability integrated score, and the response sentence as a history and can display it in graph or table format.

4. 2. The generative AI evaluation system according to claim 1, The integrated evaluation unit is characterized by comprising a score weighting adjustment means that enables the weighting coefficients of the Z score, the RAG score, and the factuality verification result to be arbitrarily set or changed. A generative AI evaluation system.

5. 2. The generative AI evaluation system according to claim 1, A generative AI evaluation system characterized by having an administrator control unit that allows an educational institution or administrator to centrally manage and apply score thresholds, weight settings, and operational policies when used by multiple users.

6. 2. The generative AI evaluation system according to claim 1, A generative AI evaluation system further comprising a template generation unit that automatically adds a template prompt including a role instruction or tone specification to an input sentence to the generative AI based on user attribute information.

Citation Information

Patent Citations

  • System

    JP2025045681A

  • System

    JP2025048524A

  • Dynamic Document Reliability Formulation

    US20200301908A1

  • Dynamic Source Reliability Formulation

    US20200302336A1

  • Method and apparatus for providing confidence information for the results of artificial intelligence models

    JP2024505213A

Cited By

  • Content credit collection probability evaluation and structure optimization system oriented to generative search engine

    CN121764954A

  • Generative AI Thinking Evaluation System

    JP7883813B1

  • Information processing system based on anomaly detection and autonomous countermeasure inference by an AI agent

    JP7898792B1