Standard evaluation calculation method, apparatus, and program based on evaluation reconstruction by generative AI
Patent Information
- Application Number
- JP2026096693
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2026-02-17
- Filing Date
- 2026-06-10
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2046-06-10
AI Technical Summary
【0030】 本発明は、評価更新の反復過程における収束解を基準として利用する点に特徴を有する。 これにより、請求項1に記載の発明により、従来の統計的集計処理では実現困難であった評価の安定構造を情報処理的に構成することが可能となる。 また、本方法によれば、生成AIを評価更新作用として反復適用し、その不動固定点を標準評価として算出することにより、評価基準として利用可能な安定状態を構成することが可能となる。
Smart Images

Figure 0007927366000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing technology that reconstructs evaluations for a plurality of objects and calculates stable standard evaluations. In particular, the present invention relates to a method, an apparatus, and a program for calculating a standard evaluation based on evaluation reconstruction processing using a model (e.g., generative AI). In the present specification, generative AI refers to, but is not limited to, a machine learning model including a generative model. [Background Art]
[0002] Conventionally, statistical aggregation methods such as average value, median value, and weighted average have been widely used as methods for integrating evaluations of a plurality of objects. However, these methods have the following problems. ·Cannot consider contextual relationships between evaluations ·Cannot reflect semantic consistency ·Does not include knowledge correction ·Cannot handle the reinterpretation process of evaluation ·Cannot define the stable state of the evaluation system In particular, with the popularization of generative AI, evaluation is not merely aggregation but involves processes of reinterpretation and reconstruction. However, conventional technologies do not use this reconstruction process for standard generation. Therefore, there is a need for a method of calculating a standard evaluation that reflects the reconstruction process. [Prior Art Documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent No. 3668491, "Self-evaluation Ability Measurement System" (FI: G06Q 10 / 06) [Patent Document 2] Japanese Unexamined Patent Application Publication No. 2022-120876, "Learning Support Apparatus" (FI: G09B 19 / 00) [Patent Document 3] Japanese Unexamined Patent Application Publication No. 2025-078897, "Personnel Evaluation System" (FI: G06Q 10 / 06) [Patent Document 4] Japanese Patent Publication No. 2026-084864, "Integration Processing Device and Evaluation Integration Method for Multiple Evaluation Information" (FI:G06Q 10 / 1053) [Patent Document 5] Japanese Patent Publication No. 2018-036718, "Evaluation Support Device" (FI:G06Q 50 / 20) [Non-patent literature]
[0004] [Non-Patent Document 1] Sraffa, P., Production of Commodities by Means of Commodities, Cambridge University Press, 1960. [Non-Patent Document 2] Godel, K., Uber formal unentscheidbare Satze..., 1931.
[0005] [Description of prior art] Patent Document 1 (Patent No. 3668491) relates to an evaluation processing system for measuring self-evaluation ability, and discloses a configuration that measures evaluation ability using evaluation values obtained from multiple evaluators. However, the document focuses on measuring the evaluation results, and does not disclose a process for generating the criteria for the evaluation system itself, particularly a configuration for determining the convergence state of the evaluation reconstruction process as the evaluation criterion. This invention differs from the aforementioned document in its technical concept in that, prior to measuring evaluation ability, it generates the stable state of the evaluation system itself as a standard evaluation.
[0006] Patent Document 2 (Japanese Patent Publication No. 2022-120876) discloses a technology that performs automatic evaluation by comparing exemplary standard logs pre-recorded in the system with operation logs when the learner actually performs a task, and evaluates the learner's proficiency in the task by comparing the results of the learner's self-evaluation of the task they entered with the results of the automatic evaluation. However, the exemplary criteria log in that document is pre-recorded, and the structure for generating evaluation criteria by iteratively applying a restructured evaluation system is not disclosed. In particular, a configuration in which an evaluation transformation process is repeatedly applied and its convergence state is adopted as a reference value is not shown.
[0007] Patent Document 3 (Japanese Patent Publication No. 2025-078897) discloses a technology for performing evaluations that eliminate subjectivity by using evaluation items based on objective facts and pre-set evaluation points. However, this document calculates evaluation values that eliminate subjectivity using predetermined objective facts and evaluation points, and does not generate evaluation criteria by repeatedly reconstructing the evaluation system.
[0008] Patent document 4 (Japanese Patent Publication No. 2026-084864) discloses a technology that, for example, uses a generating AI to compare the self-assessment declared by a job candidate with actual performance data and evaluate its reliability. However, the document uses the candidates' own performance data to evaluate their self-assessments, and does not disclose a structure that uses a stable state achieved through the repeated application of an evaluation reconstruction process as a standard.
[0009] Patent Document 5 (Japanese Patent Publication No. 2018-036718) discloses a technology that supports accurate guidance by plotting objective evaluation scores, such as those scored by administrators, and self-scores entered by the person being evaluated about their own ability level on a graph, and visualizing the discrepancy between a pre-set evaluation index (the ideal correlation between objective evaluation and self-evaluation) and the actual evaluation results. However, the document evaluates actual evaluation results using pre-set evaluations, and does not disclose a configuration in which an evaluation conversion process is repeatedly applied and the convergence state of the process is determined as the standard evaluation. Furthermore, no technological concept has been presented that utilizes an evaluation reconstruction process, including semantic consistency correction, contextual integration, and knowledge correction using generative models, as a criterion generation mechanism.
[0010] Non-Patent Document 1 (Sraffa) theoretically demonstrates a structure in which a value scale cannot be uniquely determined only from within a production system. Non-Patent Document 2 (Gödel) shows that complete truth determination is impossible from within a formal system. These documents theoretically suggest a structure in which evaluation criteria or truth criteria cannot be determined only within the system, but do not disclose a technology for generating a stable state of an evaluation system by an information processing method.
[0011] [Difference Between the Present Invention and Prior Art] As described above, each prior art individually addresses any of the following. • Statistical integration of evaluation values • Evaluation estimation by machine learning • Numerical convergence by iterative calculation • Evaluation generation by generative models However, for example, no configuration is disclosed or suggested that uses the evaluation reconstruction process itself as the evaluation criterion generation process and adopts the convergence state of the process as the standard evaluation.
[0012] [Technical Features of the Present Invention] As a technical feature of the invention recited in the claims, the present invention adopts a configuration in which evaluations are expressed as vectors, evaluation reconstruction processing by generative AI is repeatedly applied, and the convergence state thereof is output as a standard evaluation. Further, the evaluation criteria are obtained by the following configuration. That is, Evaluation transformation function F : R n → R n Iterative update v_{t+1} = F(v_t) Fixed point condition F(v*) = v* The stable state satisfying the above is adopted as an evaluation criterion. Note that this configuration is an example for obtaining an evaluation criterion, and does not limit the configuration of the claims.
[0013] The key point here is that the convergence value is not merely a calculation result, but is standardized as a stable state that satisfies the semantic and contextual consistency of the evaluation system. In other words, the present invention is Evaluation generation process = reference generation process It is structured as follows. Furthermore, the evaluation criteria for the invention described in the claims are not limited to a stable state that satisfies the semantic and contextual consistency of the evaluation system, but may be any criteria that satisfy the predetermined conditions.
[0014] [Argument for inventiveness] In the prior art, Evaluation Estimation Evaluation Integration Parameter convergence All of these methods are used as means to obtain the value being calculated. However, the present invention, The evaluation and reconstruction process itself is used as a criterion generation mechanism. This technological concept does not exist in prior art.
[0015] Furthermore, the configuration in the dependent invention described above, which uses generative AI to perform semantic consistency correction, contextual integration, and knowledge correction to generate a stable state for the entire evaluation system, cannot be derived from a simple combination of statistical processing, machine learning estimation, numerical convergence estimation, or evaluation integration.
[0016] Therefore, the present invention is not something that can be easily conceived by simply combining prior art or changing design elements, and thus possesses inventive step.
[0017] [Relationship with No. 3668491] While Patent Document 1 describes a technology for measuring evaluation ability, the present invention is a technology for generating evaluation criteria. In other words, Patent Document 1 Evaluation target → Measurement This invention Evaluation criteria → Generation The two differ in their technical subject matter and technical function.
[0018] Therefore, the present invention is not an improvement on Patent Document 1, but belongs to an independent technical field that newly constitutes a criteria generation mechanism for an evaluation system. [Overview of the Initiative] [Problems that the invention aims to solve]
[0019] The present invention aims to provide a technique for calculating a standard evaluation based on the stable state of the evaluation system by utilizing the evaluation reconstruction process. [Means for solving the problem]
[0020] To solve the above problems, the present invention provides a computer that Obtain an evaluation vector containing evaluation values for multiple targets, The evaluation vector is input to the evaluation update model that updates the evaluation value, and the process of updating the evaluation value of each of the multiple targets included in the vector is repeated. The system is characterized by outputting a standard evaluation based on the evaluation value of the converged evaluation vector.
[0021] Preferably, the computer uses the generated AI as the evaluation update model to perform iterative processing of the evaluation vector and output the standard evaluation.
[0022] Preferably, the generating AI is a machine learning model that includes a large-scale language model. This structure allows for obtaining absolute evaluations that do not rely on subjective assessments.
[0023] Preferably, the computer constructs the evaluation update model based on external evaluation criteria for evaluating the evaluation value.
[0024] Preferably, the computer inputs at least one of contextual integration, semantic consistency correction, and knowledge correction into the generating AI to generate the external evaluation criteria.
[0025] Preferably, the computer performs the iterative process to calculate the standard evaluation.
[0026] Preferably, the computer determines convergence based on the norm difference of the evaluation vector before and after the update by the iterative process. This configuration allows for stable and absolute evaluation.
[0027] Preferably, the computer measures the learner's self-assessment ability based on the difference between the standard evaluation and the evaluation results entered by the learner, the instructor, and other learners from their respective terminals on the computer network regarding the learner's learning results. This configuration allows for the evaluation of learners' self-assessments and helps identify instances of underestimation or overestimation.
[0028] Preferably, the computer measures the learner's self-assessment ability based on the objective uniqueness of the assessment results.
[0029] Preferably, the present invention represents the evaluation as a vector, iteratively applies evaluation reconstruction by a generative AI, and adopts the convergence state as the standard evaluation. Specifically, it includes the following steps: 1. Process to obtain the evaluation vector of the subject to evaluation. 2. Process of performing evaluation and conversion using generation AI 3. The process of iterating on evaluation updates. 4. Process for determining convergence 5. The process of outputting the convergence value as the standard evaluation. [Mathematical representation] Evaluation vector v ∈ R n Evaluation and reconstruction effect F : R n → R n Iterative updates v_{t+1} = F(v_t) Standard evaluation (fixed point) F(v*) = v* [Effects of the Invention]
[0030] The present invention is characterized by utilizing the converged solution in the iterative process of evaluation updates as a reference. As a result, the invention described in claim 1 makes it possible to construct an information-process-based stable evaluation structure, which was difficult to achieve with conventional statistical aggregation processing. Furthermore, this method makes it possible to construct a stable state that can be used as an evaluation criterion by repeatedly applying the generated AI as an evaluation update function and calculating its fixed point as a standard evaluation.
[0031] Furthermore, the present invention as described in the other claims provides the following effects. • This enables evaluation processing that reflects the semantic relationships between the items being evaluated. • This allows for the reconstruction of context-dependent evaluations. • It becomes possible to incorporate the process of reinterpreting evaluations into the process of forming evaluation criteria. • It becomes possible to use the convergence point of the evaluation system as a reference point. • It becomes possible to generate AI evaluation criteria that integrate multiple evaluations. In this specification, "objective uniqueness" refers to the property that the evaluation results are reproducible and stably uniquely determined, and is a technical expression of the absolute nature of the evaluation. [Brief explanation of the drawing]
[0032] [Figure 1] Layout diagram of a network system showing the first embodiment of the present invention [Figure 2] Layout diagram of a network system showing a second embodiment of the present invention [Figure 3] Block diagram showing the hardware configuration of the server machine in Figure 1. [Figure 4] Block diagram showing the functional configuration of the system in the present invention [Figure 5]Evaluation Reconstruction Iteration Process Flowchart [Figure 6] Conceptual diagram of the evaluation convergence process [Modes for carrying out the invention]
[0033] Further details will be provided below with reference to the attached drawings. The drawings show preferred embodiments. However, many different forms are possible and the embodiments are not limited to those described herein.
[0034] For example, in this embodiment, the configuration and operation of the standard evaluation calculation system will be described, but similar configurations, devices, computer programs, etc., can achieve similar effects. Furthermore, the program may be stored on a recording medium. Using this recording medium, the program can be installed on a computer, for example, thereby configuring the standard evaluation calculation system. Here, the recording medium on which the program is stored may be a non-transient recording medium such as a CD-ROM.
[0035] <1. Overview of the Invention> (1) Subject to evaluation Assign evaluation values to multiple targets. Examples: academic papers, works of art, people, products, etc. (2) Generation of evaluation vector The evaluation values for each target are constructed as a numerical sequence (evaluation vector). Example: v = (v1, v2, …, vn) (3) Generative AI Evaluation and Reconstruction A generative AI (e.g., a large-scale language model) does the following: Contextual integration ·Semantic alignment ·Knowledge correction • Adjusting relative relationships Let this be the evaluation transformation F. (4) Iteration Set the initial evaluation v0, v_{t+1} = F(v_t) Repeat this process. (5) Convergence determination ||v_{t+1} - v_t|| < ε It converges when the following conditions are met. (6) Standard evaluation output The convergence value v* will be used as the standard evaluation. [Examples] For example, if the initial evaluation is (80, 90, 70), Output after AI reconstruction iterations: (82.1, 88.4, 75.2) This will be used as the standard evaluation.
[0036] In this invention, the multiple objects (evaluation targets) are multiple different evaluation entities (papers, works, etc.), but they may also be multiple different evaluation criteria for a single evaluation entity (for example, creativity, expressiveness, logical ability, etc.). In this case, the evaluation vector is a vector whose components are the evaluation values for each evaluation criterion for a single evaluation entity.
[0037] Furthermore, in the evaluation reconstruction according to the present invention, a classical or deterministic algorithm that satisfies predetermined conditions may be used as the evaluation transformation F.
[0038] <1.1. System Configuration> Figure 1 is a block diagram showing the configuration of a system according to one embodiment. As shown in Figure 1, the standard evaluation calculation system 1A comprises a standard evaluation calculation device 10A and a client machine 4, which are configured to communicate with each other via a line 2A.
[0039] The standard evaluation calculation device 10A functions as a server for the standard evaluation calculation system 1A and can utilize general-purpose server computers or personal computers. The standard evaluation calculation device 10A may also be composed of multiple computers capable of sending and receiving information via a network NW or another network.
[0040] Client device 4 is a terminal device used by the user to calculate the standard evaluation, and can utilize personal computers, smartphones, tablet devices, wearable devices, etc.
[0041] In this embodiment, line 2A is a LAN (Local Area Network), but as shown in Figure 2, it may also be network line 2B (for example, an IP (Internet Protocol) network). Furthermore, there are no restrictions on the type of communication protocol, nor on the type or size of the network.
[0042] The standard evaluation calculation system 1A may consist only of the standard evaluation calculation device 10A (server), or it may be configured with a network line 2B as shown in Figure 2 (standard evaluation calculation system 1B).
[0043] <1.2. Hardware Configuration> Figure 3 is a hardware configuration diagram of the system according to the present invention. As shown in Figure 3, the standard evaluation calculation device 10A comprises a control means 11, an input control means 15, an output control means 16, and storage means 17, 18, 18, 19.
[0044] The control means 11 has a processor such as a CPU capable of executing instruction sets, and controls the entire operation process of the standard evaluation calculation device 10A by executing the standard evaluation calculation program, OS, and other applications according to the present invention. The input control means 15 and output control means 16 have a communication interface device for connecting to a network, and perform communication control with the network NW to input and output information. The storage means 17, 18, 18, 19 include volatile memory such as RAM capable of storing instruction sets, and non-volatile recording media such as HDDs and SSDs capable of recording the OS, standard evaluation calculation programs, and various other information.
[0045] Furthermore, the operation / input means 21 and display means 22 shown in Figure 3 are provided in the client machine 4. The operation / input means 21 has an input device capable of input processing, such as a keyboard or touch panel, and the display means 22 has a display device capable of display processing, such as a display. In addition, the client machine 4 is equipped with control means, storage means, input control means, and output control means.
[0046] The control means for client machine 4 has a processor such as a CPU capable of executing instruction sets, and controls the entire operation of client machine 4 by running the OS and other applications. The memory means of client machine 4 includes volatile memory such as RAM capable of storing instruction sets, and non-volatile recording media such as HDDs or SSDs capable of storing the OS, etc. The input control means and output control means of client machine 4 have a communication interface device for connecting to a network, and perform input and output of information via line 2A or network line 2B.
[0047] <1.3. Functional Configuration of Standard Evaluation Calculation Device 10A> Figure 4 is a functional block diagram of the standard evaluation calculation device 10A. As shown in Figure 4, the standard evaluation calculation device 10A comprises an acquisition unit 101, an external standard generation unit 102, a model construction unit 103, an update processing unit 104, a convergence determination unit 105, a standard evaluation output unit 106, and a self-evaluation capability measurement unit 107. This represents the concrete realization of information processing by software (stored in the storage means 17) by hardware (control means 11).
[0048] <1.3.1. Acquisition part 101> The acquisition unit 101 acquires an evaluation vector that includes the evaluation values of multiple targets. The acquisition unit 101 receives input of evaluation values for each of the multiple targets (multiple evaluation targets) from the user and acquires an evaluation vector that includes the evaluation values of the multiple targets.
[0049] <1.3.2.External reference generation unit 102> The external criteria generation unit 102 generates external evaluation criteria for evaluating the evaluation values of multiple targets. The external criteria generation unit 102 generates external evaluation criteria based on at least one of a plurality of external knowledge sources. In this embodiment, the external criteria generation unit 102 inputs the external knowledge source into the generation AI to generate external evaluation criteria and obtains the generated external evaluation criteria. Here, external knowledge sources include contextual integration (indicators based on past high-scoring answers and model answers), semantic consistency correction (indicators based on institutional quality standards such as curriculum guidelines and teacher rubrics), and knowledge correction (indicators based on expert knowledge).
[0050] Here, contextual integration is an index used to evaluate whether the evaluation target follows the structure and logical flow of the correct answer data, for example. Semantic consistency correction is an index used to correct for variations in spelling between words included in each evaluation target, based on semantic features and using the words used in the learning system as a reference. By using external evaluation criteria based on such external knowledge sources, it is possible to reduce inconsistencies in word spelling and complexity of sentence structure within the evaluation target, thereby enabling a more accurate evaluation of the target.
[0051] <1.3.3. Model Construction Section 103> The model building unit 103 constructs an evaluation update model. The model building unit 103 constructs an evaluation update model based on the generated external evaluation criteria. In this embodiment, the model building unit 103 constructs an evaluation update model having evaluation update rules, evaluation reconstruction directions, and / or standard evaluations (fixed points) according to the external evaluation criteria. Specifically, the model building unit 103 adjusts the parameters of the evaluation transformation F according to the external evaluation criteria to construct an evaluation update model (generating AI or algorithm (function, matrix, etc.)) in which the slope and / or direction of the evaluation update and the standard evaluation differ for each external evaluation criterion.
[0052] <1.3.4. Update Processing Unit 104> The update processing unit 104 executes the process of updating the evaluation vector. The update processing unit 104 inputs the evaluation vector to the evaluation update model that updates the evaluation values and iteratively processes the process of updating the evaluation values of each component of the evaluation vector.
[0053] <1.3.5. Convergence Determination Unit 105> The convergence determination unit 105 determines the convergence of the update process. The convergence determination unit 105 determines convergence based on the norm difference of the evaluation vector before and after the update by the iterative process.
[0054] <1.3.6. Standard Evaluation Output Unit 106> The standard evaluation output unit 106 outputs the standard evaluation. The standard evaluation output unit 106 outputs the evaluation value of the converged evaluation vector as the standard evaluation.
[0055] <1.3.7. Self-assessment ability measurement section 107> The self-assessment ability measurement unit 107 measures the learner's self-assessment ability based on the objective uniqueness of the assessment results. The self-assessment ability measurement unit 107 measures the learner's self-assessment ability based on the difference between the assessment results entered from the learner's own, the instructor's, and / or other learners' client devices 4, 5 and the standard assessment for a given learner's learning results. Specifically, the self-assessment ability measurement unit 107 measures the learner's self-assessment ability according to the magnitude of the difference between the norm of each component of the assessment results entered from the learner's, the instructor's, and / or other learners' client devices 4, 5 and the norm of the standard assessment.
[0056] <2. Processing Flowchart> The standard evaluation calculation method of the present invention will be described below with reference to Figure 5. Figure 5 is a flowchart showing the process from obtaining the evaluation vector to outputting the standard evaluation using external evaluation criteria.
[0057] <2.1. Obtaining the Evaluation Vector> First, in step S101 (hereinafter, "step SX" will be simply referred to as "SX"), the acquisition unit 101 acquires an evaluation vector. In this embodiment, the acquisition unit 101 receives input of evaluation values from multiple users for each evaluation target of one user, and calculates an average value for each evaluation target based on the multiple evaluation values for each evaluation target for one user. Then, the acquisition unit 101 acquires an evaluation vector whose components are the average values for each evaluation target.
[0058] <2.2. Generation of External Evaluation Criteria> In S102, the external criterion generation unit 102 converts the external knowledge source to generate an external evaluation criterion. In this embodiment, the external criterion generation unit 102 inputs a plurality of external knowledge sources, the type of evaluation target (type of evaluation body and evaluation perspective), and instructions to the generation AI to integrate the plurality of external knowledge sources to generate a criterion, and generates an external evaluation criterion for each evaluation perspective. Specifically, for example, as an external evaluation criterion, a vector is used in which the component corresponding to the dimension (type of evaluation perspective) of each component of the evaluation vector acquired in S101 has parameters (criterion parameters) corresponding to the external knowledge source.
[0059] More specifically, the external criteria generation unit 102 generates external evaluation criteria for each type of evaluation body and evaluation perspective. Specifically, the external criteria generation unit 102 generates external evaluation criteria (vectors) for each type of evaluation body (e.g., for essay evaluation, mathematics answer evaluation, interview evaluation, patent examination evaluation, etc.), each having parameters (criteria parameters) corresponding to the external knowledge source in the components of the evaluation vector that correspond to the dimensions of each component.
[0060] <2.3. Evaluation Update Process> In S103, the update processing unit 104 executes an evaluation update process. In this embodiment, the model construction unit 103 generates an evaluation update model according to the external evaluation criteria generated in S102. The update processing unit 104 then inputs the evaluation vector obtained in S101 into the evaluation update model and obtains an updated evaluation vector in which each component has been updated by the evaluation update model.
[0061] Specifically, evaluation update models include a weight update matrix that has weights based on the criterion parameters of the external evaluation criteria and updates the weights for each component of the evaluation vector. Other evaluation update models include a model that updates evaluation values by referencing an external knowledge source database, and a trained knowledge model (generative AI) in which parameters for updating the evaluation vector based on the external evaluation criteria have been learned. Furthermore, evaluation update models that use evaluation values as a probability distribution and update that probability distribution are also used. Finally, evaluation update models that update evaluation values using an evaluation network graph representing the relationships between evaluation targets based on the criterion parameters of the external evaluation criteria are also employed.
[0062] <2.4. Convergence Criteria> In S104, the convergence determination unit 105 determines whether the update process has converged. In this embodiment, the convergence determination unit 105 compares each component of the evaluation vector that was updated immediately before with each component of the evaluation vector that was updated immediately before the updated evaluation vector, and determines whether the norm of the difference between each component of these evaluation vectors is less than or equal to a predetermined threshold.
[0063] More preferably, convergence may be determined only when the external evaluation criteria, which are external reference conditions, are met. Specifically, the convergence determination unit 105 may determine convergence using a convergence determination for evaluation updates according to the external evaluation criteria. More specifically, the convergence determination unit 105 determines convergence using a convergence determination with different thresholds for the convergence conditions according to the external evaluation criteria (see Figure 6).
[0064] If it is determined that convergence will occur (YES in S104), the process proceeds to S105. On the other hand, if it is determined that convergence will not occur (NO in S104), the process returns to S104, and the update processing unit 104 inputs the acquired updated evaluation vector into the evaluation update model.
[0065] Preferably, the update processing unit 104 calculates the updated evaluation vector by multiplying each component of the evaluation vector that was updated immediately before, and each component of the evaluation vector that was updated immediately before the updated evaluation vector, by a different coefficient. Specifically, the update processing unit 104 multiplies each component of the evaluation vector that was updated immediately before by 1-α, multiplies each component of the evaluation vector that was updated immediately before the updated evaluation vector by α, sums these values, and calculates the updated evaluation vector in which each component contains the sum.
[0066] Preferably, the update processing unit 104 uses a generation AI as the evaluation update model, inputs the updated evaluation vector converted to JSON format into the generation AI, and executes the update process.
[0067] <2.5. Output of Standard Evaluation> In S105, the standard evaluation output unit 106 outputs the standard evaluation. In this embodiment, the standard evaluation output unit 106 outputs the average value of the evaluation values of each component of the evaluation vector that converged in S104 as the standard evaluation. Alternatively, the standard evaluation output unit 106 may output the evaluation values of each component as the standard evaluation.
[0068] By executing steps S101 to S105 as described above, it is possible to generate stable evaluation criteria (standard evaluation) that are not influenced by the evaluator's subjectivity.
[0069] In a preferred embodiment of the present invention, notification is given based on the convergence rate of the standard evaluation. Specifically, the update processing unit 104 performs the standard evaluation calculation process multiple times for the same evaluation vector, and the notification unit of the standard evaluation calculation system 1A (not shown) determines that the evaluation vector is abnormal and gives a notification if the convergence rate, which is the ratio of the number of convergences to the number of calculation processes, is below a predetermined threshold.
[0070] Furthermore, notifications are made based on the convergence speed of the update process. Specifically, the convergence determination unit 105 determines whether the update process has been performed more than a predetermined number of times, and the notification unit determines that the evaluation vector is abnormal if the update process has been performed more than a predetermined number of times, and issues a notification.
[0071] In a preferred embodiment of this system, if multiple standard evaluations do not match, the system outputs a standard evaluation according to the degree of variation among those standard evaluations. Specifically, the update processing unit 104 inputs the same single evaluation vector obtained multiple times into the evaluation update model (e.g., the generation AI), and the standard evaluation output unit 106 determines whether the respective provisional standard evaluations obtained from the evaluation update model match. If the provisional standard evaluations match, the standard evaluation output unit 106 outputs that provisional standard evaluation as the standard evaluation. On the other hand, if the respective provisional standard evaluations obtained as a result of the calculation do not match, the standard evaluation output unit 106 determines whether the difference norm between those provisional standard evaluations is below a predetermined threshold. If the difference norm between the evaluation vectors (fixed points) is below a predetermined threshold, the standard evaluation output unit 106 outputs the average value of each component of those provisional standard evaluations as the standard evaluation. [Industrial applicability]
[0072] This invention is applicable to the following fields. • Educational evaluation Research evaluation • Art judging • Recruitment evaluation • Content Recommendations AI Ranking Product Reviews [Explanation of Symbols]
[0073] 1A,1B Network system, 2A Line, 2B Internet line, 4,5 Client machines, 10A,10B Server machines, 11 Control means, 15 Input control means, 16 Output control means, 17,18,19,20 Storage means, 21 Operation / input means, 22 Display means, 101 Acquisition unit, 102 External reference generation unit, 103 Model construction unit, 104 Update processing unit, 105 Convergence determination unit, 106 Standard evaluation output unit, 107 Self-evaluation capability measurement unit
Claims
1. Computers Obtain an evaluation vector containing evaluation values for multiple targets, An external knowledge source and instructions for generating criteria from the external knowledge source are input to a large-scale language model to generate an external evaluation criterion, which is a vector having criterion parameters corresponding to the external knowledge source for each component of the evaluation vector. The evaluation vector is input to an evaluation update model that updates the evaluation value based on the evaluation vector and the external evaluation criteria, and the process of updating the evaluation value of each of the multiple targets included in the evaluation vector is repeated. Based on the evaluation value of the evaluation vector whose evaluation value has converged through the above iterative process, a standard evaluation is output. A standard evaluation calculation method characterized by including a deterministic algorithm in which parameters relating to the amount and / or direction of updating the evaluation value are determined according to the reference parameters.
2. The standard evaluation calculation method according to Claim 1, wherein the external knowledge source is an index for evaluating the evaluation vector based on at least one of past highly-rated answers, model answers, learning systems, and expert knowledge.
3. The evaluation update model further includes a large-scale language model, The aforementioned computer, A method for calculating a standard evaluation according to claim 1, comprising: inputting the same evaluation vector 1 multiple times into the large-scale language model; determining whether the respective provisional standard evaluations obtained from the large-scale language model are the same; and outputting the provisional standard evaluation as the standard evaluation if the provisional standard evaluations are the same.
4. The standard evaluation calculation method according to Claim 1, characterized in that the evaluation update model is a weight update matrix having weights based on the reference parameters and updating the evaluation values for each component of the evaluation vector.
5. The standard evaluation calculation method according to claim 1, wherein the computer determines convergence based on the norm difference of the evaluation vector before and after the update by the iterative process.
6. A method for measuring a learner's self-assessment ability, wherein the computer measures the learner's self-assessment ability based on the difference between the standard evaluation described in claim 1 and the evaluation results entered by the learner, the instructor, and other learners from their respective terminals on a computer network regarding the learner's learning results.
7. An apparatus for performing the standard evaluation calculation method described in any one of claims 1 to 5, or the self-evaluation ability measurement method described in claim 6.
8. A program that causes a computer to execute the standard evaluation calculation method described in any one of claims 1 to 5, or the self-evaluation ability measurement method described in claim 6.
Citation Information
Patent Citations
FI2018-036718
FI2022-120876
FI2025-078897
FI3668491
FIQ10/1053