Universal artificial intelligence test method and device for intelligent agent and electronic equipment

Through random matching of test scenarios and generating test questions, comprehensive testing of agents is carried out, which solves the problem of difficulty in evaluating the general artificial intelligence characteristics of agents in the existing technology, and achieves efficient and objective testing and evaluation.

CN120045442AInactive Publication Date: 2025-05-27BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311522900.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the existing technology to comprehensively test whether an agent has the essential characteristics of general artificial intelligence. Traditional testing methods have problems such as strong subjectivity and irregular evaluation standards.

Method used

By obtaining test tasks, randomly match multiple sets of test scenarios for the task, multiple random test questions are generated, and the agent is tested based on these questions to determine whether it has universal artificial intelligence characteristics, such as completing unlimited tasks, autonomously generating tasks and value-driven.

Benefits of technology

A comprehensive test of the general artificial intelligence characteristics of the agent is realized, which can approach the infinite task test effect, provides a standard evaluation system, and improves the objectivity and accuracy of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045442A_ABST
    Figure CN120045442A_ABST
Patent Text Reader

Abstract

The invention provides a general artificial intelligence test method and device for an intelligent agent and electronic equipment. The method comprises the following steps: acquiring a test task; randomly matching multiple groups of test scenes for the test task to obtain multiple groups of random test questions; testing the intelligent agent based on the multiple groups of random test questions to obtain a test result; and according to the test result, determining whether the intelligent agent has universal artificial intelligence characteristics or not. The universal artificial intelligence features of the intelligent agent can be tested.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of general artificial intelligence technology, and particularly to a general artificial intelligence testing method, device, and electronic device for an intelligent agent. Background Art

[0002] With the emergence of large models such as ChatGPT and various embodied intelligent agent models, general artificial intelligence has developed in a very rapid manner and achieved important results in many sub - fields.

[0003] Currently, in the process of testing the artificial intelligence performance of an intelligent agent, often only one or a few tasks are tested. These tasks can only reflect the development level of the intelligent agent in specific ability dimensions and levels, and cannot reflect whether the intelligent agent has the essential characteristics of general artificial intelligence.

[0004] Currently, finding a feasible method for testing general artificial intelligence of an intelligent agent has become a research hotspot. Summary of the Invention

[0005] The present invention provides a general artificial intelligence testing method, device, and electronic device for an intelligent agent, which can realize the testing of the general artificial intelligence characteristics of the intelligent agent.

[0006] The present invention provides a general artificial intelligence testing method for an intelligent agent, the method comprising: obtaining a test task; randomly matching multiple groups of test scenarios to the test task to obtain multiple groups of random test questions; testing the intelligent agent based on the multiple groups of random test questions to obtain a test result; and determining whether the intelligent agent has general artificial intelligence characteristics according to the test result.

[0007] According to the general artificial intelligence testing method for an intelligent agent provided by the present invention, the determining whether the intelligent agent has general artificial intelligence characteristics according to the test result specifically comprises: when the test result is that the intelligent agent passes the test of the random test questions, determining that the intelligent agent has general artificial intelligence characteristics; and when the test result is that the intelligent agent fails the test of the random test questions, determining that the intelligent agent does not have general artificial intelligence characteristics.

[0008] A general artificial intelligence testing method for an agent provided by the present invention, wherein the general artificial intelligence feature includes the feature of being able to complete an infinite number of tasks; when the test task is to test whether the agent can complete an infinite number of tasks, testing the agent based on multiple groups of the random test questions to obtain a test result, specifically including: testing the agent based on multiple groups of first random test questions to obtain a first successful number of the agent successfully completing the first random test questions, wherein the first random test questions are random test questions matching the feature of being able to complete an infinite number of tasks; obtaining a first success rate of the agent based on the first successful number and a first number of questions of the first random test questions, wherein the first success rate matches the feature of being able to complete an infinite number of tasks; when the first success rate exceeds a success rate threshold, obtaining a first confidence level corresponding to the first success rate based on the first number of questions, and determining that the agent passes the first random test questions test with the first confidence level, and taking the agent passing the first random test questions test with the first confidence level as the test result; when the first success rate does not exceed the success rate threshold, determining that the agent fails the first random test questions test, and taking the agent failing the first random test questions test as the test result.

[0009] A general artificial intelligence testing method for an agent provided by the present invention, the method further includes: obtaining multiple groups of other first random test questions, wherein the other first random test questions are random test questions that belong to different dimensions from the first random test questions and match the feature of being able to complete an infinite number of tasks; sequentially performing the steps of testing the agent based on multiple groups of the other first random test questions until it is determined that the agent passes the other first random test questions test with another first confidence level, or it is determined that the agent fails the other first random test questions test, wherein the another first confidence level is determined based on the number of questions of the other first random test questions; respectively determining a first question difficulty weight corresponding to the first random test questions, and another question difficulty weight corresponding to the other first random test questions; determining a final test result based on the test result that the agent passes the first random test questions test with the first confidence level, the first question difficulty weight, the test result that the agent fails the other first random test questions test, and the another question difficulty weight.

[0010] According to a general artificial intelligence testing method for an agent provided by the present invention, the method further includes: determining a final test result based on the test result that the agent fails the first random test question, the first question difficulty weight, the test results that the agent passes other first random test questions with other first confidence levels, and the other question difficulty weights.

[0011] According to a general artificial intelligence testing method for an agent provided by the present invention, the general artificial intelligence feature includes the feature of being able to autonomously generate derivative tasks, wherein the derivative tasks are generated based on the test tasks; when the test task is to test whether the agent can autonomously generate derivative tasks, testing the agent based on multiple groups of the random test questions to obtain a test result, specifically including: testing the agent based on multiple groups of second random test questions in the absence of task instruction input to obtain the second successful number of the agent successfully completing the second random test questions, wherein the second random test questions are random test questions matching the feature of being able to autonomously generate derivative tasks; obtaining a second success rate of the agent based on the second successful number and the second question number of the second random test questions, wherein the second success rate matches the feature of being able to autonomously generate derivative tasks; when the second success rate exceeds a success rate threshold, obtaining a second confidence level corresponding to the second success rate based on the second question number and determining that the agent passes the second random test questions with the second confidence level, and taking the agent passing the second random test questions with the second confidence level as the test result; when the second success rate does not exceed the success rate threshold, determining that the agent fails the second random test questions and taking the agent failing the second random test questions as the test result.

[0012] According to a general artificial intelligence testing method for an agent provided by the present invention, determining that the agent successfully completes the second random test questions is characterized by the following method: determining a set of generated derivative tasks obtained by the agent based on the second random test questions; obtaining a pre-set set of expected derivative tasks matching the second random test questions; obtaining a task set intersection and union ratio based on the set of generated derivative tasks and the set of expected derivative tasks; when the task set intersection and union ratio exceeds an intersection and union ratio threshold, determining that the agent successfully completes the second random test questions; determining that the agent passes the second random test questions with the second confidence level is characterized by the following method: determining that the agent passes the second random test questions with the second confidence level according to the task set intersection and union ratio.

[0013] A general artificial intelligence testing method for an intelligent agent provided by the present invention, wherein the general artificial intelligence features include features that can be value-driven. When the test task is to test whether the intelligent agent can be value-driven, testing the intelligent agent based on multiple groups of the random test questions to obtain a test result, specifically including: testing the intelligent agent based on multiple groups of third random test questions to obtain multiple groups of value preferences of the intelligent agent in a preset value dimension during multiple groups of tests; determining multiple groups of value preference relationships of the intelligent agent among the preset value dimensions based on the multiple groups of value preferences, wherein the value preference relationship represents the importance degree of the value preference of the intelligent agent in the preset value dimension; obtaining an average value preference relationship based on the multiple groups of value preference relationships; when the average value preference relationship exceeds a value preference relationship threshold, determining that the intelligent agent has a stable value pattern, and the average value preference relationship is the stable value pattern; when the similarity between the average value preference relationship and a preset human value preference relationship is greater than a similarity threshold, obtaining a third confidence level based on the number of third questions of the third random test questions, and determining that the intelligent agent passes the third random test questions test with the third confidence level according to the preset human value preference relationship as the value-driven category, and taking the intelligent agent passing the third random test questions test with the third confidence level according to the preset human value preference relationship as the value-driven category as the test result; when the average value preference relationship does not exceed the value preference relationship threshold, determining that the intelligent agent fails the third random test questions test, and taking the intelligent agent failing the third random test questions test as the test result.

[0014] The present invention also provides a general artificial intelligence testing device for an intelligent agent, the device including: an acquisition module, configured to acquire a test task; a processing module, configured to randomly match multiple groups of test scenarios for the test task to obtain multiple groups of random test questions; a testing module, configured to test the intelligent agent based on the multiple groups of random test questions to obtain a test result; a generation module, configured to determine whether the intelligent agent has general artificial intelligence features according to the test result.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the general artificial intelligence testing method for the intelligent agent as described in any one of the above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the general artificial intelligence testing method for the intelligent agent as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the general artificial intelligence testing method of the agent as described in any one of the above.

[0018] For the general artificial intelligence testing method, device and electronic device provided by the present invention, a test task is obtained, and multiple groups of test scenarios are randomly matched to the test task to obtain multiple groups of random test questions. Since randomly matching multiple groups of test scenarios to the test task can approximate the test effect of infinite tasks, the agent is tested based on the multiple groups of random test questions to obtain a test result; and whether the agent has general artificial intelligence characteristics is determined according to the test result, so as to realize the testing of the general artificial intelligence characteristics of the agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of the general artificial intelligence testing method of the agent provided by the present invention;

[0021] Figure 2 It is one of the schematic flowcharts of testing the agent based on multiple groups of random test questions to obtain a test result provided by the present invention;

[0022] Figure 3 It is another schematic flowchart of testing the agent based on multiple groups of random test questions to obtain a test result provided by the present invention;

[0023] Figure 4 It is still another schematic flowchart of testing the agent based on multiple groups of random test questions to obtain a test result provided by the present invention;

[0024] Figure 5 It is a schematic structural diagram of the general artificial intelligence testing device of the agent provided by the present invention;

[0025] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will, in conjunction with the accompanying drawings of the present invention, clearly and completely describe the technical solutions in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0027] General artificial intelligence should possess three essential characteristics: whether it can complete an infinite number of tasks, whether it can autonomously generate tasks, and whether it is value-driven. These essential characteristics include whether it can complete an infinite number of tasks, whether it can autonomously generate tasks, and whether it is value-driven, etc. The traditional evaluation method for general artificial intelligence - the Turing test - has problems such as strong subjectivity and non-standard evaluation criteria. Therefore, based on the existing test platforms and datasets, a set of standard test methods and evaluation systems need to be constructed so as to analyze whether the intelligent agent under test possesses the basic characteristics of general artificial intelligence. The general artificial intelligence test method for the intelligent agent provided by the present invention can be a feasible test method that can achieve the foregoing characteristics.

[0028] Figure 1 It is a schematic flowchart of the general artificial intelligence test method for the intelligent agent provided by the present invention.

[0029] The following will be combined with Figure 1 to illustrate the process of the general artificial intelligence test method for the intelligent agent provided by the present invention.

[0030] In an exemplary embodiment of the present invention, in combination with Figure 1 it can be seen that the general artificial intelligence test method for the intelligent agent may include Step 110 to Step 140, and each step will be introduced separately below.

[0031] In Step 110, obtain test tasks.

[0032] In Step 120, randomly match multiple groups of test scenarios for the test tasks to obtain multiple groups of random test questions.

[0033] In Step 130, test the intelligent agent based on the multiple groups of random test questions to obtain test results.

[0034] In Step 140, determine whether the intelligent agent has the characteristics of general artificial intelligence according to the test results.

[0035] In one embodiment, the corresponding test tasks can be selected from a given test task pool. Among them, the test task pool can be a pre-given task pool covering 6 dimensions, where the 6 dimensions can include 5 ability dimensions, namely vision, language, motion, cognition, and learning, and 1 value dimension.

[0036] In another embodiment, different test tasks can be determined according to different general artificial intelligence features to be tested, where the test tasks can be to test whether the agent can complete infinite tasks, whether the agent can autonomously generate derivative tasks, and whether the agent can perform value-driven tasks.

[0037] In another embodiment, for each given test task, multiple groups of test scenarios can be randomly matched to the test task to obtain multiple groups of random test questions. Among them, a test question is composed of a test task and a test scenario.

[0038] Furthermore, the agent can be tested multiple times based on multiple groups of random test questions to obtain test results, and according to the test results, it can be determined whether the agent has general artificial intelligence features. Since randomly matching multiple groups of test scenarios to the test task can approximate the test effect of infinite tasks, testing the agent based on multiple groups of random test questions to obtain test results can implement the testing of the general artificial intelligence features of the agent.

[0039] The general artificial intelligence testing method for the agent provided by the present invention obtains test tasks, and randomly matches multiple groups of test scenarios to the test tasks to obtain multiple groups of random test questions. Since randomly matching multiple groups of test scenarios to the test task can approximate the test effect of infinite tasks, testing the agent based on multiple groups of random test questions to obtain test results; and determining whether the agent has general artificial intelligence features according to the test results can implement the testing of the general artificial intelligence features of the agent.

[0040] In another exemplary embodiment of the present invention, taking the embodiment described above as an example, determining whether the agent has general artificial intelligence features according to the test results (corresponding to step 140) can be implemented in the following manner:

[0041] In the case where the test result is that the agent passes the random test questions, it is determined that the agent has general artificial intelligence features;

[0042] In the case where the test result is that the agent fails the random test questions, it is determined that the agent does not have general artificial intelligence features.

[0043] In one embodiment, when it is determined that the agent passes the random test questions, it can be determined that the agent has general artificial intelligence features; when it is determined that the agent fails the random test questions, it can be determined that the agent does not have general artificial intelligence features.

[0044] It should be noted that when testing different general artificial intelligence features of the agent, different random test questions can be used.

[0045] Figure 2 It is one of the schematic flowcharts for testing an agent based on multiple groups of random test questions to obtain test results provided by the present invention.

[0046] The following will be combined with Figure 2 to illustrate the process of testing an agent based on multiple groups of random test questions to obtain test results in the scenario corresponding to the feature that the general artificial intelligence feature fails to complete infinite tasks.

[0047] In an exemplary embodiment of the present invention, when the test task is to test whether the agent can complete infinite tasks, combined with Figure 2 it can be known that testing the agent based on multiple groups of random test questions to obtain test results may include steps 210 to 240. Each step will be introduced separately below.

[0048] In step 210, the agent is tested based on multiple groups of first random test questions to obtain the first success quantity of the agent successfully completing the first random test questions.

[0049] In one embodiment, M test tasks (each task has a preset difficulty weight) can be selected in a test question bank (covering 6 dimensions and 5 levels). To cover all dimensions and levels, at least 6 * 5 = 30 tasks need to be selected. Then, multiple groups of test scenarios are randomly matched to the test tasks to obtain multiple groups of random test questions. Among them, the random test questions used to test whether the agent can complete infinite tasks can be called the first random test questions. The first random test questions are random test questions matching the feature of being able to complete infinite tasks.

[0050] In another example, each test task is repeatedly tested N times. The scenarios of the first random test questions are different each time, but the same task knowledge points are examined.

[0051] Furthermore, the agent is tested against the first random test questions. According to the completion standard of the specified questions, it is judged whether the current first random test question is successful, and the first success quantity of the agent successfully completing the first random test questions is obtained.

[0052] In step 220, based on the first success quantity and the first question quantity of the first random test questions, the first success rate of the agent is obtained.

[0053] In one embodiment, continuing with the embodiment described above, the number of successful times n in N repeated experiments can be counted. In other words, the first success quantity of the agent successfully completing the first random test questions is determined.

[0054] Further, based on the first number of successes and the first number of first random test questions, the first success rate of the agent can be obtained, and the first success rate can be expressed as n / N. Among them, the first success rate matches the characteristic of being able to complete infinite tasks.

[0055] In step 230, when the first success rate exceeds the success rate threshold, based on the first number of questions, the first confidence level corresponding to the first success rate is obtained, and it is determined that the agent passes the test of the first random test question with the first confidence level, and taking the agent passing the test of the first random test question with the first confidence level as the test result.

[0056] In step 240, when the first success rate does not exceed the success rate threshold, it is determined that the agent fails the test of the first random test question, and taking the agent failing the test of the first random test question as the test result.

[0057] In one embodiment, it can be stipulated that the success rate threshold for passing the test of the test task is T, that is, the success rate threshold is determined.

[0058] Further, the first confidence level corresponding to the first success rate can be obtained based on the first number of questions. In an example, the first confidence level can be expressed as sigmoid(G*N)*2 - 1, where sigmoid(.) represents a function; G is the relevant experience value corresponding to the task category; among them, 0 < G ≤ 1 means that a larger amount of tasks reaches a higher confidence level; G > 1 means that a smaller amount of tasks reaches a higher confidence level; N represents the first number of questions.

[0059] In one embodiment, when the first success rate exceeds the success rate threshold, it is determined that the agent passes the test of the first random test question with the first confidence level, and taking the agent passing the test of the first random test question with the first confidence level as the test result.

[0060] In other words, for the selected test task, the agent can pass the "complete infinite tasks" test with the first confidence level of sigmoid(G*N)*2 - 1, where the determination criterion for passing is: the first success rate of the agent successfully completing the task is greater than the success rate threshold.

[0061] In yet another embodiment, when the first success rate does not exceed the success rate threshold, it can be determined that the agent fails the test of the first random test question, and taking the agent failing the test of the first random test question as the test result.

[0062] In yet another exemplary embodiment of the present invention, continuing with the embodiment described above, the general artificial intelligence test method for the agent may further include the following steps:

[0063] Obtain multiple sets of other first random test questions, where the other first random test questions are random test questions that belong to different dimensions from the first random test questions and match the characteristics of being able to complete infinite tasks;

[0064] Execute the test on the intelligent agent based on multiple sets of other first random test questions step by step as described above until it is determined that the intelligent agent passes the test of the other first random test questions with another first confidence level, or it is determined that the intelligent agent fails the test of the other first random test questions, where the other first confidence level is determined based on the number of questions in the other first random test questions;

[0065] Respectively determine the first question difficulty weight corresponding to the first random test questions and the other question difficulty weights corresponding to the other first random test questions;

[0066] Determine the final test result based on the test result that the intelligent agent passes the test of the first random test questions with the first confidence level, the first question difficulty weight, the test result that the intelligent agent fails the test of the other first random test questions, and the other question difficulty weights.

[0067] In one embodiment, multiple sets of other first random test questions can also be obtained. It can be understood that the other first random test questions and the first random test questions are both random test questions that match the characteristics of being able to complete infinite tasks. In the application process, it can be determined that the intelligent agent passes the test of the other first random test questions with another first confidence level, or it is determined that the intelligent agent fails the test of the other first random test questions in the manner of steps 210 to 240 above.

[0068] Furthermore, respectively determine the first question difficulty weight corresponding to the first random test questions and the other question difficulty weights corresponding to the other first random test questions. And determine the final test result based on the test result that the intelligent agent passes the test of the first random test questions with the first confidence level, the first question difficulty weight, the test result that the intelligent agent fails the test of the other first random test questions, and the other question difficulty weights.

[0069] In other words, if the first question difficulty weight is much smaller than the other question difficulty weights, and if the test result is that the intelligent agent fails the test of the other first random test questions, then the proportion of this test result is larger, and then the test result that the intelligent agent fails the test of the other first random test questions will be mainly referred to to determine the final test result.

[0070] In another exemplary embodiment of the present invention, continuing with the embodiment described above as an example, the general artificial intelligence test method for an intelligent agent may further include the following steps:

[0071] Determine the final test result based on the test result that the agent fails the first random test question, the first question difficulty weight, the test results that the agent passes other first random test questions with other first confidence levels, and the other question difficulty weights.

[0072] In one example, if the first question difficulty weight is much smaller than the other question difficulty weights, and if the agent passes other first random test questions with other first confidence levels, then the proportion of this test result is larger, and then the test result that the agent passes other first random test questions with other first confidence levels will be mainly referred to to determine the final test result.

[0073] It should be noted that if the agent passes the first random test question test and the agent passes other first random test questions test, then the final test result is that the agent passes the test of the feature of being able to complete infinite tasks. If the agent fails the first random test question test and the agent fails other first random test questions test, then the final test result is that the agent fails the test of the feature of being able to complete infinite tasks.

[0074] Figure 3 It is the second flow diagram of the test result obtained by testing the agent based on multiple groups of random test questions provided by the present invention.

[0075] The following will be combined with Figure 3 Describe the process of testing the agent based on multiple groups of random test questions to obtain the test result in the scenario corresponding to the general artificial intelligence feature of being able to autonomously generate derivative tasks. Among them, the derivative task is generated based on the test task.

[0076] In an exemplary embodiment of the present invention, when the test task is to test whether the agent can autonomously generate derivative tasks, combined with Figure 3 It can be known that testing the agent based on multiple groups of random test questions to obtain the test result may include steps 310 to 340, and each step will be introduced separately below.

[0077] In step 310, in the absence of task instruction input, test the agent based on multiple groups of second random test questions to obtain the second success number of the agent successfully completing the second random test questions.

[0078] In one embodiment, 1 test task can be selected from the test question bank. Then randomly match multiple groups of test scenarios for the test task to obtain multiple groups of random test questions. Among them, the random test questions used to test whether the agent can autonomously generate derivative tasks can be called second random test questions. Among them, the second random test questions are random test questions that match the feature of being able to autonomously generate derivative tasks.

[0079] In yet another embodiment, a test is conducted. In the case where there is no explicit task instruction input, the derivative tasks generated by the agent are collected and compared with the expected derivative task set S. In one example, the IOU (Intersection over Union) value of the generated derivative task set K and the expected derivative task set S is used to quantify the matching effect. When the IOU is greater than the empirical threshold F, it is recorded that the agent has successfully completed the second random test question.

[0080] In step 320, based on the second success count and the second number of questions of the second random test questions, the second success rate of the agent is obtained.

[0081] In yet another embodiment, the test can be repeated N times. Each time, the task scenario is different and no explicit instruction is given. The number of successful times n in the N repeated experiments is counted, that is, the second success count is obtained.

[0082] In yet another example, the second success rate of the agent can be obtained based on the second success count and the second number of questions of the second random test questions. The second success rate can be expressed as n / N. Among them, the second success rate matches the characteristic of being able to autonomously generate derivative tasks.

[0083] In step 330, when the second success rate exceeds the success rate threshold, based on the second number of questions, the second confidence level corresponding to the second success rate is obtained, and it is determined that the agent passes the second random test question test with the second confidence level, and taking the agent passing the second random test question test with the second confidence level as the test result.

[0084] In step 340, when the second success rate does not exceed the success rate threshold, it is determined that the agent fails the second random test question test, and taking the agent failing the second random test question test as the test result.

[0085] In one embodiment, the success rate threshold for passing the test task can be specified as T, that is, the success rate threshold is determined.

[0086] Furthermore, the second confidence level corresponding to the second success rate can be obtained based on the second number of questions. In one example, the second confidence level can be expressed as sigmoid(G*N)*2 - 1, where sigmoid(.) represents a function; G is the relevant empirical value corresponding to the task category; among them, 0 < G ≤ 1 indicates that a larger amount of tasks reaches a higher confidence level; G > 1 indicates that a smaller amount of tasks reaches a higher confidence level; N represents the second number of questions.

[0087] In one embodiment, when the second success rate exceeds the success rate threshold, it is determined that the agent passes the second random test question test with a second confidence level, and taking the agent passing the second random test question test with the second confidence level as the test result.

[0088] In other words, for the selected test task, the agent can pass the "autonomous task generation" test with a second confidence level of sigmoid(G*N)*2 - 1, where the judgment criterion for passing is that the second success rate of the agent is greater than the success rate threshold T, and the IOU of the generated derivative task set K of the agent and the expected derivative task set S is greater than the empirical threshold F.

[0089] In another embodiment, when the second success rate does not exceed the success rate threshold, it can be determined that the agent fails the second random test question test, and taking the agent failing the second random test question test as the test result.

[0090] In another exemplary embodiment of the present invention, continue with the embodiment described above Figure 3 as an example for illustration. Among them, determining that the agent successfully completes the second random test question can be characterized by the following method:

[0091] Determine the generated derivative task set obtained by the agent based on the second random test question;

[0092] Obtain the expected derivative task set pre-set and matching the second random test question;

[0093] Based on the generated derivative task set and the expected derivative task set, obtain the intersection over union of the task sets;

[0094] When the intersection over union of the task sets exceeds the intersection over union threshold, determine that the agent successfully completes the second random test question.

[0095] In one embodiment, the generated derivative task set K generated by the agent can be collected and compared with the expected derivative task set S. In one example, the IOU (Intersection over Union) value of the generated derivative task set K and the expected derivative task set S is used to quantify the matching effect. Among them, when the IOU is greater than the empirical threshold F (corresponding to the intersection over union threshold), it is recorded that the agent successfully completes the second random test question.

[0096] Furthermore, determining that the agent passes the second random test question test with a second confidence level can be characterized by the following method:

[0097] Determine that the agent passes the second random test question test with a second confidence level according to the intersection over union of the task sets.

[0098] Figure 4This is the third schematic diagram of the process for testing an agent based on multiple sets of random test questions and obtaining test results provided by the present invention.

[0099] The following will be combined with Figure 4 to illustrate the process of testing an agent based on multiple sets of random test questions and obtaining test results in a scenario where the general artificial intelligence feature is a value-driven feature.

[0100] In an exemplary embodiment of the present invention, when the test task is to test whether an agent can be value-driven, combined with Figure 4 it can be known that testing an agent based on multiple sets of random test questions and obtaining test results may include steps 410 to 460. Each step will be introduced separately below.

[0101] In step 410, the agent is tested based on multiple sets of third random test questions, and multiple sets of value preferences of the agent in a preset value dimension during the multiple sets of tests are obtained.

[0102] In step 420, based on the multiple sets of value preferences, multiple sets of value preference relationships of the agent among the preset value dimensions are determined.

[0103] In one embodiment, a test task can be selected from the test question bank. Then, multiple sets of test scenarios are randomly matched for the test task to obtain multiple sets of random test questions. Among them, the random test questions used to test whether the agent can be value-driven can be called the third random test questions.

[0104] Furthermore, the agent can be tested based on multiple sets of third random test questions. According to multiple predefined value dimensions (corresponding to the preset value dimensions), the value preferences reflected in the execution of the task are collected and encoded as multiple sets of value preference relationships among the respective components (corresponding to the preset value dimensions) of the value vector. Among them, one component represents one dimension. The value preference relationship can represent the importance degree of the value preference of the agent under the preset value dimension.

[0105] In step 430, an average value preference relationship is obtained based on the multiple sets of value preference relationships.

[0106] In step 440, when the average value preference relationship exceeds the value preference relationship threshold, it is determined that the agent has a stable value pattern.

[0107] In one embodiment, the test can be repeated N times. Each time, the task scenario is different, but the same task knowledge points are examined. Calculate the multiple sets of value preference relationships in the N repeated experiments and statistically obtain the average value preference relationship.

[0108] In another embodiment, a value preference relationship threshold may be determined. In the case where the average value preference relationship exceeds the value preference relationship threshold, it may be determined that the agent has a stable value pattern, and the average value preference relationship is the stable value pattern.

[0109] In step 450, in the case where the similarity between the average value preference relationship and the preset human value preference relationship is greater than the similarity threshold, a third confidence level is obtained based on the third number of questions of the third random test questions, and it is determined that the agent passes the third random test questions test with the third confidence level in the value-driven category according to the preset human value preference relationship, and taking the agent passing the third random test questions test with the third confidence level in the value-driven category according to the preset human value preference relationship as the test result;

[0110] In step 460, in the case where the average value preference relationship does not exceed the value preference relationship threshold, it is determined that the agent fails the third random test questions test, and taking the agent failing the third random test questions test as the test result.

[0111] In one embodiment, the similarity between the average value preference relationship and the preset human value preference relationship may be calculated. In the case where the similarity between the average value preference relationship and the preset human value preference relationship is greater than the similarity threshold T, it is determined that the agent passes the third random test questions test in the value-driven category according to the preset human value preference relationship.

[0112] In another embodiment, a third confidence level may be obtained based on the third number of questions of the third random test questions. In one example, the third confidence level may be expressed as sigmoid(G*N)*2 - 1, where sigmoid(.) represents a function; G is the relevant experience value corresponding to the task category; where 0 < G ≤ 1 indicates that a larger amount of tasks reaches a higher confidence level; G > 1 indicates that a smaller amount of tasks reaches a higher confidence level; N represents the third number of questions.

[0113] In one embodiment, in the case where the similarity between the average value preference relationship and the preset human value preference relationship is greater than the similarity threshold, a third confidence level is obtained based on the third number of questions of the third random test questions, and it is determined that the agent passes the third random test questions test with the third confidence level in the value-driven category according to the preset human value preference relationship, and taking the agent passing the third random test questions test with the third confidence level in the value-driven category according to the preset human value preference relationship as the test result.

[0114] In other words, for the selected test task, the agent can pass the "value-driven" test with the third confidence level of sigmoid(G*N)*2 - 1. The passing criterion is that the agent uses the preset human value preference relationship as the value-driven category and the average value preference relationship as the value stability.

[0115] In yet another embodiment, when the average value preference relationship does not exceed the value preference relationship threshold, it is determined that the agent fails the test of the third random test question, and the fact that the agent fails the test of the third random test question is used as the test result.

[0116] In the general artificial intelligence test method for the agent provided by the present invention, based on a given multi-dimensional and multi-level general artificial intelligence task rating table (6 dimensions including vision, language, motion, cognition, learning, value, etc., and each dimension has 5 levels), it is proposed to use a classification selection task plus a randomized scenario setting to approximate the infinite task test effect and cover the general artificial intelligence task rating table;

[0117] Under the premise of a limited number of test tasks, statistically test and determine whether the agent "completes infinite tasks";

[0118] Based on the expectations of human society for the agent's tasks, test and determine whether the agent "autonomously generates tasks"; and

[0119] Based on the value pattern of human society, test and determine whether the agent is "value-driven". Thus, the general artificial intelligence characteristics of the agent can be comprehensively tested.

[0120] According to the above description, in the general artificial intelligence test method for the agent provided by the present invention, test tasks are obtained, and multiple groups of test scenarios are randomly matched to the test tasks to obtain multiple groups of random test questions. Since randomly matching multiple groups of test scenarios to the test tasks can approximate the infinite task test effect, the agent is tested based on the multiple groups of random test questions to obtain a test result; and whether the agent has general artificial intelligence characteristics is determined according to the test result, so as to realize the test of the general artificial intelligence characteristics of the agent.

[0121] Based on the same concept, the present invention also provides a general artificial intelligence test device for an agent.

[0122] The general artificial intelligence test device for the agent provided by the present invention is described below. The general artificial intelligence test device for the agent described below can be mutually referred to the general artificial intelligence test method for the agent described above.

[0123] Figure 5It is a schematic structural diagram of a general artificial intelligence testing device for an agent provided by the present invention.

[0124] In an exemplary embodiment of the present invention, in combination with Figure 5 it can be seen that the general artificial intelligence testing device for an agent may include an acquisition module 510, a processing module 520, a testing module 530, and a generation module 540. Each module will be introduced separately below.

[0125] The acquisition module 510 can be configured to acquire a test task;

[0126] The processing module 520 can be configured to randomly match multiple groups of test scenarios for the test task to obtain multiple groups of random test questions;

[0127] The testing module 530 can be configured to test the agent based on multiple groups of random test questions to obtain a test result;

[0128] The generation module 540 can be configured to determine whether the agent has general artificial intelligence characteristics according to the test result.

[0129] In an exemplary embodiment of the present invention, the generation module 540 can adopt the following method to determine whether the agent has general artificial intelligence characteristics according to the test result:

[0130] In the case where the test result is that the agent passes the random test questions, it is determined that the agent has general artificial intelligence characteristics;

[0131] In the case where the test result is that the agent fails to pass the random test questions, it is determined that the agent does not have general artificial intelligence characteristics.

[0132] In an exemplary embodiment of the present invention, the general artificial intelligence characteristics may include the characteristic of being able to complete an infinite number of tasks; in the case where the test task is to test whether the agent can complete an infinite number of tasks, the testing module 530 can adopt the following method to test the agent based on multiple groups of random test questions to obtain a test result:

[0133] Test the agent based on multiple groups of first random test questions to obtain the first successful number of the agent successfully completing the first random test questions, where the first random test questions are random test questions matching the characteristic of being able to complete an infinite number of tasks;

[0134] Based on the first successful number and the first number of questions of the first random test questions, obtain the first success rate of the agent, where the first success rate matches the characteristic of being able to complete an infinite number of tasks;

[0135] When the first success rate exceeds the success rate threshold, a first confidence level corresponding to the first success rate is obtained based on the number of the first test questions, and it is determined that the agent passes the test of the first random test question with the first confidence level, and taking the fact that the agent passes the test of the first random test question with the first confidence level as the test result;

[0136] When the first success rate does not exceed the success rate threshold, it is determined that the agent fails the test of the first random test question, and taking the fact that the agent fails the test of the first random test question as the test result.

[0137] In an exemplary embodiment of the present invention, the test module 530 can also be used for:

[0138] Obtain multiple groups of other first random test questions, where the other first random test questions are random test questions that belong to different dimensions from the first random test questions and match the feature of being able to complete infinite tasks;

[0139] Successively perform the test on the agent based on multiple groups of other first random test questions according to the foregoing steps until it is determined that the agent passes the test of the other first random test question with the other first confidence level, or it is determined that the agent fails the test of the other first random test question, where the other first confidence level is determined based on the number of the other first random test questions;

[0140] Respectively determine the first question difficulty weight corresponding to the first random test question and the other question difficulty weight corresponding to the other first random test questions;

[0141] Based on the test result that the agent passes the test of the first random test question with the first confidence level, the first question difficulty weight, the test result that the agent fails the test of the other first random test question, and the other question difficulty weight, determine the final test result.

[0142] In an exemplary embodiment of the present invention, the test module 530 can also be used for:

[0143] Based on the test result that the agent fails the test of the first random test question, the first question difficulty weight, the test result that the agent passes the test of the other first random test question with the other first confidence level, and the other question difficulty weight, determine the final test result.

[0144] In an exemplary embodiment of the present invention, the general artificial intelligence feature can include the feature of being able to autonomously generate derivative tasks, where the derivative tasks are generated based on the test tasks; when the test task is to test whether the agent can autonomously generate derivative tasks, the test module 530 can adopt the following method to implement the test on the agent based on multiple groups of random test questions to obtain the test result:

[0145] In the absence of task instruction input, the agent is tested based on multiple sets of second random test questions, and the second successful number of the agent successfully completing the second random test questions is obtained, where the second random test questions are random test questions matching the characteristics of being able to autonomously generate derivative tasks;

[0146] Based on the second successful number and the second number of questions of the second random test questions, the second success rate of the agent is obtained, where the second success rate matches the characteristics of being able to autonomously generate derivative tasks;

[0147] When the second success rate exceeds the success rate threshold, based on the second number of questions, the second confidence level corresponding to the second success rate is obtained, and it is determined that the agent passes the second random test questions test with the second confidence level, and taking the agent passing the second random test questions test with the second confidence level as the test result;

[0148] When the second success rate does not exceed the success rate threshold, it is determined that the agent fails the second random test questions test, and taking the agent failing the second random test questions test as the test result.

[0149] In an exemplary embodiment of the present invention, the test module 530 may characterize and determine that the agent successfully completes the second random test questions in the following manner:

[0150] Determine the set of generated derivative tasks obtained by the agent based on the second random test questions;

[0151] Obtain the pre-set set of expected derivative tasks matching the second random test questions;

[0152] Based on the set of generated derivative tasks and the set of expected derivative tasks, obtain the task set intersection-union ratio;

[0153] When the task set intersection-union ratio exceeds the intersection-union ratio threshold, it is determined that the agent successfully completes the second random test questions;

[0154] Determining that the agent passes the second random test questions test with the second confidence level can be characterized in the following manner:

[0155] Determine that the agent passes the second random test questions test with the second confidence level according to the task set intersection-union ratio.

[0156] In an exemplary embodiment of the present invention, the general artificial intelligence characteristics may include characteristics that can be value-driven. When the test task is to test whether the agent can be value-driven, the test module 530 may implement the test of the agent based on multiple sets of random test questions to obtain the test result in the following manner:

[0157] Testing the agent based on multiple sets of third random test questions, and obtaining multiple sets of value preferences of the agent under a preset value dimension reflected during the multiple sets of tests;

[0158] Based on the multiple sets of value preferences, determining multiple sets of value preference relationships between the preset value dimensions of the agent; wherein, the value preference relationship represents the importance degree of the value preference of the agent under the preset value dimension;

[0159] Based on the multiple sets of value preference relationships, obtaining an average value preference relationship;

[0160] In the case where the average value preference relationship exceeds the value preference relationship threshold, determining that the agent has a stable value pattern;

[0161] In the case where the similarity between the average value preference relationship and the preset human value preference relationship is greater than the similarity threshold, obtaining a third confidence level based on the number of the third test questions of the third random test questions, and determining that the agent passes the third random test questions test with the third confidence level according to the preset human value preference relationship as the value driving category, and taking that the agent passes the third random test questions test with the third confidence level according to the preset human value preference relationship as the value driving category as the test result;

[0162] In the case where the average value preference relationship does not exceed the value preference relationship threshold, determining that the agent fails the third random test questions test, and taking that the agent fails the third random test questions test as the test result.

[0163] Figure 6 Illustrated is a schematic diagram of the physical structure of an electronic device, as Figure 6 shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the general artificial intelligence test method of the agent, and the method includes: obtaining a test task; randomly matching multiple sets of test scenarios for the test task to obtain multiple sets of random test questions; testing the agent based on the multiple sets of random test questions to obtain a test result; and determining whether the agent has general artificial intelligence characteristics according to the test result.

[0164] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0165] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the general artificial intelligence test method of the agent provided by the above-mentioned various methods. The method includes: obtaining a test task; randomly matching multiple groups of test scenarios for the test task to obtain multiple groups of random test questions; testing the agent based on the multiple groups of random test questions to obtain a test result; and determining whether the agent has general artificial intelligence characteristics according to the test result.

[0166] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the general artificial intelligence test method of the agent provided by the above-mentioned various methods. The method includes: obtaining a test task; randomly matching multiple groups of test scenarios for the test task to obtain multiple groups of random test questions; testing the agent based on the multiple groups of random test questions to obtain a test result; and determining whether the agent has general artificial intelligence characteristics according to the test result.

[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0169] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present invention, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be advantageous.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A general artificial intelligence testing method for an agent, characterized in that, the method includes: Obtaining a test task; Randomly matching multiple groups of test scenarios for the test task to obtain multiple groups of random test questions; Testing the agent based on multiple groups of the random test questions to obtain a test result; According to the test result, determining whether the agent has general artificial intelligence characteristics.

2. The general artificial intelligence testing method for an agent according to claim 1, characterized in that, the determining whether the agent has general artificial intelligence characteristics according to the test result specifically includes: In the case where the test result is that the agent passes the random test question test, determining that the agent has general artificial intelligence characteristics; In the case where the test result is that the agent fails the random test question test, determining that the agent does not have general artificial intelligence characteristics.

3. The general artificial intelligence testing method for an agent according to claim 1 or 2, characterized in that, the general artificial intelligence characteristics include the characteristic of being able to complete an infinite number of tasks; in the case where the test task is to test whether the agent can complete an infinite number of tasks, the testing the agent based on multiple groups of the random test questions to obtain a test result specifically includes: Testing the agent based on multiple groups of first random test questions to obtain the first successful number of the agent successfully completing the first random test questions, where the first random test questions are random test questions matching the characteristic of being able to complete an infinite number of tasks; Based on the first successful number and the first number of questions of the first random test questions, obtaining the first success rate of the agent, where the first success rate matches the characteristic of being able to complete an infinite number of tasks; In the case where the first success rate exceeds the success rate threshold, obtaining a first confidence level corresponding to the first success rate based on the first number of questions, and determining that the agent passes the first random test question test with the first confidence level, and taking the agent passing the first random test question test with the first confidence level as the test result; In the case where the first success rate does not exceed the success rate threshold, determining that the agent fails the first random test question test, and taking the agent failing the first random test question test as the test result.

4. The general artificial intelligence testing method for an agent according to claim 3, characterized in that, the method further includes: Obtaining multiple groups of other first random test questions, where the other first random test questions are random test questions that belong to different dimensions from the first random test questions and match the characteristic of being able to complete an infinite number of tasks; Execute the test on the agent based on multiple groups of the other first random test questions in sequence according to the foregoing steps until it is determined that the agent passes the test of the other first random test questions with another first confidence level, or it is determined that the agent fails the test of the other first random test questions, where the other first confidence level is determined based on the number of questions of the other first random test questions; Determine the first question difficulty weight corresponding to the first random test question and the other question difficulty weight corresponding to the other first random test question respectively; Determine the final test result based on the test result that the agent passes the test of the first random test question with the first confidence level, the first question difficulty weight, the test result that the agent fails the test of the other first random test questions, and the other question difficulty weight.

5. The general artificial intelligence test method for an agent according to claim 4, wherein, the method further includes: Determine the final test result based on the test result that the agent fails the test of the first random test question, the first question difficulty weight, the test result that the agent passes the test of the other first random test questions with another first confidence level, and the other question difficulty weight.

6. The general artificial intelligence test method for an agent according to claim 1 or 2, wherein, the general artificial intelligence feature includes the feature of being able to autonomously generate derivative tasks, where the derivative tasks are generated based on the test tasks; in the case where the test task is to test whether the agent can autonomously generate derivative tasks, the test of the agent based on multiple groups of the random test questions to obtain a test result specifically includes: In the absence of task instruction input, test the agent based on multiple groups of second random test questions to obtain the second successful number of the agent successfully completing the second random test questions, where the second random test questions are random test questions matching the feature of being able to autonomously generate derivative tasks; Obtain the second success rate of the agent based on the second successful number and the second number of questions of the second random test questions, where the second success rate matches the feature of being able to autonomously generate derivative tasks; In the case where the second success rate exceeds the success rate threshold, obtain the second confidence level corresponding to the second success rate based on the second number of questions, and determine that the agent passes the test of the second random test questions with the second confidence level, and use the fact that the agent passes the test of the second random test questions with the second confidence level as the test result; In the case where the second success rate does not exceed the success rate threshold, determine that the agent fails the test of the second random test questions, and use the fact that the agent fails the test of the second random test questions as the test result.

7. The general artificial intelligence test method for an agent according to claim 6, wherein, Determining that the agent successfully completes the second random test question is characterized by the following: Determine the set of generated derivative tasks obtained by the agent based on the second random test question; Obtain a preset set of expected derivative tasks that match the second random test question; Based on the set of generated derivative tasks and the set of expected derivative tasks, obtain the intersection - union ratio of the task sets; In the case where the intersection - union ratio of the task sets exceeds the intersection - union ratio threshold, determine that the agent successfully completes the second random test question; Determining that the agent passes the test of the second random test question with the second confidence level is characterized by the following: Determine that the agent passes the test of the second random test question with the second confidence level according to the intersection - union ratio of the task sets.

8. The general artificial intelligence test method for an agent according to claim 1 or 2, wherein, The general artificial intelligence feature includes a feature that can be value - driven. When the test task is to test whether the agent can be value - driven, testing the agent based on multiple sets of the random test questions to obtain a test result specifically includes: Testing the agent based on multiple sets of third random test questions to obtain multiple sets of value preferences of the agent in a preset value dimension during the process of multiple sets of tests; Based on multiple sets of the value preferences, determine multiple sets of value preference relationships of the agent among the preset value dimensions; wherein, the value preference relationship characterizes the importance degree of the agent's value preference in the preset value dimension; Based on multiple sets of value preference relationships, obtain an average value preference relationship; In the case where the average value preference relationship exceeds the value preference relationship threshold, determine that the agent has a stable value pattern; in the case where the similarity between the average value preference relationship and the preset human value preference relationship is greater than the similarity threshold, obtain a third confidence level based on the number of the third questions in the third random test questions, and determine that the agent passes the test of the third random test questions with the third confidence level according to the preset human value preference relationship as the value - driven category, and take the fact that the agent passes the test of the third random test questions with the third confidence level according to the preset human value preference relationship as the value - driven category as the test result; In the case where the average value preference relationship does not exceed the value preference relationship threshold, determine that the agent fails the test of the third random test questions, and take the fact that the agent fails the test of the third random test questions as the test result.

9. A general artificial intelligence test device for an agent, wherein, The device includes: An acquisition module for acquiring a test task; A processing module for randomly matching multiple sets of test scenarios for the test task to obtain multiple sets of random test questions; A test module for testing the agent based on multiple sets of the random test questions to obtain a test result; A generation module for determining whether the agent has general artificial intelligence features according to the test result.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the general artificial intelligence test method of the agent according to any one of claims 1 to 8.