Intelligent level grade evaluation method and system based on restricted boltzmann machine
By using a method for assessing the intelligence level based on a restricted Boltzmann machine and calculating the ELO score using retrospective data, the objectivity and universality issues of intelligence level assessment in existing technologies are resolved, and an objective and unified assessment of the intelligence level of intelligent agents is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
- Filing Date
- 2023-06-29
- Publication Date
- 2026-04-21
AI Technical Summary
Existing intelligent level assessment technologies lack objectivity and universality, resulting in inconsistent assessment standards among different experts and making it difficult to adapt to the needs of different test scenarios.
An intelligence level assessment method based on Restricted Boltzmann Machines is adopted. By acquiring debriefing data from multiple tested agents, the ELO score is calculated using a trained Restricted Boltzmann Machine network, and the intelligence level of the agent is determined based on the ELO score.
It achieves an objective and universal assessment of the level of intelligence, freeing it from reliance on subjective assessment standards and providing a unified assessment framework.
Smart Images

Figure CN116738180B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence assessment technology, and in particular to a method and system for assessing the level of intelligence based on a restricted Boltzmann machine. Background Technology
[0002] In recent years, my country has vigorously developed the research and construction of intelligent systems / platforms employing virtual reality, simulation, and deductive reasoning technologies to achieve adversarial simulation, reasoning, and decision analysis. Among these efforts, assessing the intelligence level of intelligent systems / platforms is a crucial aspect of artificial intelligence research, playing a pivotal role in promoting the development of AI.
[0003] Based on the output format of the assessment results, existing intelligent level assessment technologies can be divided into qualitative assessment and quantitative assessment. Qualitative assessment involves subjectively describing the level of intelligence in tiers through expert or social consensus. Quantitative assessment involves determining the weights of factors that may affect the level of intelligence and then performing a weighted aggregation calculation to obtain the overall level value. Based on the method of determining weights during the assessment process, existing intelligent level assessment technologies can be divided into subjective assessment and objective assessment. Qualitative assessment is generally a subjective assessment method. Quantitative assessment, after obtaining the overall level value, also involves experts determining the level classification standards to ultimately determine the level of intelligence; therefore, quantitative assessment is also a subjective assessment method.
[0004] Clearly, existing intelligence level assessment technologies cannot escape subjective level classification standards. Different experts use different standards, and different test scenarios require the development of new standards, resulting in a lack of universality in the assessment technology. With the deepening application of intelligent systems / platforms, intelligence assessment needs to incorporate human-machine collaboration, intelligent and autonomous simulation training, integrated training, and adversarial training. Therefore, it is necessary to develop objective and universal intelligence level assessment methods to address the problems of subjective and non-universally applicable intelligence level assessments. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for assessing the intelligence level based on a restricted Boltzmann machine, which can achieve an objective and universal assessment of the intelligence level.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] A method for assessing the intelligence level based on a restricted Boltzmann machine, the method comprising:
[0008] The test involves acquiring debriefing data from pairwise battles between multiple tested agents in an experimental task. The amount of debriefing data is the same as the number of battles. The debriefing data includes adjudication data and the battle data of each of the two tested agents. The adjudication data is used to characterize the battle results, and the battle data is used to characterize the battle situation.
[0009] For each tested agent, all the replay data corresponding to the tested agent are determined; for each replay data corresponding to the tested agent, multiple evaluation indicators are calculated based on the tested agent's battle data in the replay data, and the ELO score of the tested agent is predicted using a trained restricted Boltzmann machine network with the multiple evaluation indicators as input; the average of all the ELO scores of the tested agent is calculated to obtain the final ELO score of the tested agent.
[0010] For each tested agent, the intelligence level of the tested agent is determined based on its final ELO score.
[0011] A system for assessing the intelligence level based on a restricted Boltzmann machine, the system comprising:
[0012] The data acquisition module is used to acquire replay data obtained from pairwise battles between multiple tested agents in an experimental task; the amount of replay data is the same as the number of battles; the replay data includes adjudication data and the battle data of each of the two tested agents; the adjudication data is used to characterize the battle result; the battle data is used to characterize the battle situation.
[0013] The integral prediction module is used to determine all the replay data corresponding to each tested agent; for each replay data corresponding to the tested agent, calculate multiple evaluation indicators based on the tested agent's combat data in the replay data, and use the multiple evaluation indicators as input to predict the ELO score of the tested agent using a trained restricted Boltzmann machine network; calculate the average of all the ELO scores of the tested agent to obtain the final ELO score of the tested agent.
[0014] The level determination module is used to determine the intelligence level of each tested intelligent agent based on its final ELO score.
[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0016] This invention provides a method and system for evaluating the intelligence level of multiple tested agents through pairwise battles in an experimental task. First, it acquires debriefing data from these battles, obtained by multiple tested agents. Then, for each tested agent, multiple evaluation indicators are calculated based on its battle data from each debriefing. Using these indicators as input, a trained Restricted Boltzmann Machine network predicts the ELO score of the tested agent, further calculating its final ELO score. Finally, for each tested agent, its intelligence level is determined based on its final ELO score. This eliminates the need for subjective level classification standards, achieving an objective and universal intelligence level evaluation. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the evaluation method provided in Embodiment 1 of the present invention;
[0019] Figure 2 This is a schematic diagram of the evaluation method provided in Embodiment 1 of the present invention;
[0020] Figure 3 This is a schematic diagram of the three-level indicator system provided in Embodiment 1 of the present invention;
[0021] Figure 4 This is a schematic diagram of the initial restricted Boltzmann machine network provided in Embodiment 1 of the present invention;
[0022] Figure 5 This is a system block diagram of the evaluation system provided in Embodiment 2 of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The purpose of this invention is to provide a method and system for assessing the intelligence level based on a restricted Boltzmann machine, which can achieve an objective and universal assessment of the intelligence level.
[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Example 1:
[0027] This embodiment provides a method for evaluating the intelligence level based on a restricted Boltzmann machine, such as... Figure 1 and Figure 2 As shown, the evaluation method includes:
[0028] S1: Obtain debriefing data from pairwise battles between multiple tested agents in an experimental task; the number of debriefing data points is the same as the number of battles; the debriefing data includes adjudication data and the battle data of each of the two tested agents; the adjudication data is used to characterize the battle results; the battle data is used to characterize the battle situation.
[0029] This embodiment can construct test tasks targeting all the capabilities of the tested intelligent agents in typical test scenarios. Typical test scenarios include scenarios for implementing the test tasks, and both the test scenarios and test tasks can be customized by the user. Here, this embodiment combines a specific airspace-land combat scenario to construct four test tasks targeting the perception, cognition, decision-making, and action capabilities of the tested intelligent agents. The perception task is as follows: Test agent B commands 5 attack drones to attack the protected target of test agent A and then retreats; test agent A commands 5 reconnaissance drones to track and search for the attack drones. The cognition task is as follows: Test agent A commands 1 infantry fighting vehicle to transport 2 infantry units to a designated position behind test agent B's defensive line; test agent A commands 2 reconnaissance drones... The unmanned vehicle support system anticipates the movements of five reconnaissance unmanned vehicles on the defense line of Agent B, preventing the infantry fighting vehicles from being detected by the enemy. The decision-making task is as follows: Agent A and Agent B each command 10 attack drones. The drone swarm commanded by Agent B will launch a surprise attack, while the drone swarm commanded by Agent A will retaliate after being attacked. The action task is as follows: Agent A will command 10 attack drones to attack the 10 air defense unmanned vehicles commanded by Agent B, while the unmanned vehicle swarm commanded by Agent B will conduct a rapid defense.
[0030] After defining the experimental tasks, all tested agents engaged in pairwise battles within those tasks, yielding post-battle data. It's important to note that these pairwise battles were repeated multiple times; the same two tested agents could repeatedly engage in battles. The total number of battles conducted by all tested agents equals the number of post-battle data sets. In other words, the number of post-battle data sets depends on the total number of battles conducted by all tested agents in the experimental tasks. If all tested agents engaged in n pairwise battles in the experimental tasks, then n sets of post-battle data sets can be obtained. Each set of post-battle data represents one battle, including adjudication data and the individual battle data of the two tested agents. The adjudication data characterizes the battle result, specifically including the final outcome of the battle between tested agent A and tested agent B, including wins, draws, and losses for tested agent A, and corresponding losses, draws, and wins for tested agent B. The combat data includes operator attribute data and adversarial data. Operator attribute data includes the operator's name, number, and opponent's number. The operator includes the model commanded by the tested intelligent agent, such as drones, unmanned vehicles, and infantry fighting vehicles. The adversarial data includes the operator's position coordinates, health, action type, and action target number at each moment. The action type includes maneuvering, stationary maneuvering, shooting, laser suppression, and infrared detection.
[0031] S2: For each tested agent, determine all the replay data corresponding to the tested agent; for each replay data corresponding to the tested agent, calculate multiple evaluation indicators based on the tested agent's battle data in the replay data, and use the multiple evaluation indicators as input to predict the tested agent's ELO score using a trained Restricted Boltzmann Machine Network; calculate the average of all the tested agent's ELO scores to obtain the tested agent's final ELO score;
[0032] The replay data corresponding to the tested agent is the replay data generated when the tested agent participates in the battle as one side. For example, if tested agent A and tested agent B play against each other and generate replay data C, then the replay data corresponding to tested agent A and tested agent B will both include this replay data C.
[0033] For each replay data point corresponding to the tested agent, multiple evaluation metrics are calculated based on the agent's combat data within the replay data. Specifically, multiple evaluation metrics are calculated based on the agent's operator attribute data and adversarial data within the replay data. Assuming this embodiment contains a total of g evaluation metrics, the user can customize the type of evaluation metric and the mathematical operation function I for each metric. i (x), where the mathematical operation function I iThe number of (x) is equal to the number of evaluation indicators, i = 1, 2, ..., g, where x is the collective term for the retrospective data. Then, based on the j-th group of retrospective data x... j The calculated evaluation indices I1, I2, ..., I g The value can be expressed as: I1(x) j ),I2(x j ),...,I g (x j ), j = 1, 2, ..., n.
[0034] The number of matches a tested agent participates in corresponds to the number of sets of replay data, which in turn allows for the calculation of a number of evaluation metrics. For each evaluation metric, multiple metrics are used as input, and a trained Restricted Boltzmann Machine (RBM) network is used to predict the ELO score of the tested agent. The number of ELO scores corresponds to the number of evaluation metrics; that is, when a tested agent participates in multiple matches, it will have multiple ELO scores. In this embodiment, the average of these multiple ELO scores is taken as the final ELO score of the tested agent. Through the above process, the final ELO score of each tested agent can be calculated.
[0035] S3: For each tested intelligent agent, the intelligence level of the tested intelligent agent is determined based on the final ELO score of the tested intelligent agent.
[0036] Before using multiple evaluation metrics as input to predict the ELO score of the tested agent using a trained Restricted Boltzmann Machine (RBM) network, the evaluation method in this embodiment further includes: training the initial RBM network to obtain a trained RBM network. This step may include:
[0037] (1) Obtain the dataset, which includes multiple samples. The samples include historical values of multiple evaluation metrics and historical values of the ELO scores of the tested agent.
[0038] Specifically, obtaining a dataset can include:
[0039] 1) Obtain historical replay data from multiple tested agents in pair battles during the experimental task. The number of historical replay data is the same as the number of battles. The historical replay data includes historical adjudication data and the historical battle data of each of the two tested agents in the battle. The historical adjudication data is used to characterize the battle results, and the historical battle data is used to characterize the battle situation.
[0040] 2) For each tested agent, determine all historical replay data corresponding to the tested agent; for each historical replay data corresponding to the tested agent, calculate the historical values of multiple evaluation indicators based on the historical battle data of the tested agent in the historical replay data; calculate the historical value of the tested agent's ELO score using the integral iteration formula based on the historical adjudication data in the historical replay data; the historical values of multiple evaluation indicators and the historical value of the tested agent's ELO score form a sample, and all samples form a dataset.
[0041] In calculating the ELO score, this embodiment can statistically analyze the expected win rates of all tested agents and convert them into ELO scores using an iterative integral formula. Specifically, the iterative integral formula includes:
[0042]
[0043] in, Let ELO score be the score of agent A after the nth battle. Let K be the ELO score of agent A after the (n-1)th battle, which is the ELO score after the previous battle; K is the game score change coefficient, which is a constant, and in this embodiment, K can be set to 20. This represents the actual score of agent A in the battle, with values of 1, 0.5, and 0 for victory, draw, and defeat, respectively. Let be the expected win rate of agent A in the nth battle; The ELO score of agent B after the nth battle; The ELO score of agent B after the (n-1)th battle is the same as the ELO score after the previous battle. This represents the actual score of agent B in the battle, with values of 1, 0.5, and 0 for victory, draw, and defeat, respectively. Let be the expected win rate of agent B in the nth battle.
[0044] All tested agents have the same initial ELO integral, i.e.
[0045] The formula for calculating the expected win rate is as follows:
[0046] The constant s = 400.
[0047] By iterating using the above integral iteration formula based on the historical decision data in the historical review data, the ELO scores of all tested agents can be updated, and the historical values of the ELO scores of the tested agents can be obtained.
[0048] Using the three sets of replay data shown in Table 1 as examples, the calculation process of the above ELO integral will be further explained:
[0049] Table 1
[0050]
[0051] All tested agents have the same initial ELO score, assuming the initial ELO score is 600, i.e.
[0052] In the first battle
[0053]
[0054]
[0055] In the second battle,
[0056]
[0057]
[0058] In the third battle,
[0059]
[0060]
[0061] Through the above process, the ELO scores of all tested agents can be updated.
[0062] (2) Construct an initial restricted Boltzmann machine network. The initial restricted Boltzmann machine network includes multiple index layers set in sequence and several restricted Boltzmann machine networks located between adjacent index layers. The index layers include a three-level index layer, a two-level index layer, a first-level index layer and an ELO integral layer from bottom to top. The first-level index layer includes first-level indicators used to reflect the agent's perception ability, cognition ability, decision-making ability and action ability. The second-level index layer includes second-level indicators used to reflect the agent's autonomy, learning and cooperation. The third-level index layer includes third-level indicators used to reflect the technical characteristics of the second-level indicators. The third-level indicators are the evaluation indicators. The ELO integral layer includes the ELO integral of the tested agent.
[0063] In this embodiment, the third-level indicators are the aforementioned evaluation indicators. Therefore, calculating the historical values of multiple evaluation indicators based on historical review data is equivalent to calculating the historical values of multiple third-level indicators based on historical review data. Assuming the third-level indicator layer contains g third-level indicators, the user can define a custom mathematical operation function I for the third-level indicators. i (x), mathematical operation function I i The number of (x) is equal to the number of third-level indicators, i = 1, 2, ..., g, where x is the collective term for the back-review data, and x is the j-th group of back-review data. j In the three-level indicators I1, I2, ..., Ig The numerical value in is represented as I1(x) j ),I2(x j ),...,I g (x j ), j = 1, 2, ..., n, denoted as For example, the three-level indicator layer contains 22 tertiary indicators, and users can define custom mathematical operation functions (I) for these tertiary indicators. i (x), i = 1, 2, ..., 22, the j-th set of review data x j In the three-level indicators I1, I2, ..., I 22 The numerical value in is represented as I1(x) j ),I2(x j ),...,I 22 (x j ), j = 1, 2, ..., n, denoted as
[0064] Assuming this embodiment has 9 tested agents and 22 tertiary indicators, some data is shown in Table 2 below:
[0065] Table 2
[0066]
[0067] This embodiment can pre-establish a three-level indicator system, such as Figure 3 As shown, from top to bottom, there are three levels of indicators: a primary indicator layer, a secondary indicator layer, and a tertiary indicator layer. One primary indicator in the primary indicator layer can correspond to several secondary indicators in the secondary indicator layer, and one secondary indicator in the secondary indicator layer can correspond to several tertiary indicators in the tertiary indicator layer. In this embodiment, the primary indicators can correspond to experimental tasks, that is, the number of experimental tasks is equal to the number of primary indicators included in the primary indicator layer.
[0068] Specifically, this embodiment can have four primary indicators: perception ability, cognitive ability, decision-making ability, and action ability. It can also have twelve secondary indicators. For the primary indicator of perception ability, the corresponding secondary indicators can be perception autonomy, perception learning, and perception coordination. For the primary indicator of cognitive ability, the corresponding secondary indicators can be cognitive autonomy, cognitive learning, and cognitive coordination. For the primary indicator of decision-making ability, the corresponding secondary indicators can be decision autonomy, decision learning, and decision coordination. For the primary indicator of action ability, the corresponding secondary indicators can be action autonomy, action learning, and action coordination. The number and type of tertiary indicators can be customized by the user. Here, this embodiment provides an example of tertiary indicators as follows:
[0069] For the secondary indicator of perception autonomy, the corresponding tertiary indicators are: 1) Target detection capability, which reflects the ability of friendly forces to detect enemy targets on the battlefield in a combat scenario. It is calculated as: (Total replay frames - Number of frames when the target was first detected) / Total replay frames, where one frame refers to a moment in time, and the total replay frames refer to the runtime. 2) Target tracking capability, which reflects the ability to continuously track a target after its initial detection in a combat scenario. It is calculated as: Number of frames successfully tracked / (Total replay frames - Number of frames when the target was first detected), where successful target tracking means continuously tracking the target for more than ten frames.
[0070] For the secondary indicator of perceptual learning, the corresponding tertiary indicator can be: 1) the ability to discover the evolution of the target, which is calculated as: (autonomy of the agent after reinforcement learning - autonomy of the agent before reinforcement learning) / autonomy of the agent before reinforcement learning * 0.5 + 0.5.
[0071] For the secondary indicator of perception collaboration, the corresponding tertiary indicator can be: 1) Collaborative target discovery capability, which reflects the ability of our side to discover enemy targets based on operators in combat scenarios, while hiding our own targets. The calculation method is: (Number of enemy targets discovered by our side / Number of enemy targets - Number of enemy targets discovered by our side / Number of our targets) * 0.5 + 0.5.
[0072] For the secondary indicator of cognitive learning, the corresponding tertiary indicator can be: 1) Global situation judgment evolution ability, which is calculated as: (value function of the agent before reinforcement learning - value function of the agent after reinforcement learning) / value function of the agent after reinforcement learning * 0.5 + 0.5.
[0073] For the secondary indicator of cognitive synergy, the corresponding tertiary indicator can be: 1) Threat ranking capability. In multi-aircraft air combat tactical decision-making, various indicators are combined to assess the degree of threat posed by enemy aircraft to our side, so as to allocate our targets and firepower. The calculation method is: the number of frames in which our side issues strike actions / the total number of frames in which our side issues all actions.
[0074] For the secondary indicator of decision-making autonomy, the corresponding tertiary indicator can be: 1) Decision-making benefit capability, which is calculated as: Decision-making benefit - Task autonomous execution capability, where decision-making benefit = (Task return after action - Task return before action) / Average of returns before action, and task autonomous execution capability is an action autonomy indicator, as defined below.
[0075] For the secondary indicator of decision learning, the corresponding tertiary indicator can be: 1) Decision imitation learning ability, which is calculated as: (cross-entropy after the agent performs imitation learning - cross-entropy before the agent performs imitation learning) / cross-entropy before the agent performs imitation learning * 0.5 + 0.5.
[0076] For the secondary indicator of decision-making coordination, the corresponding tertiary indicator can be: 1) Decision-making coordination capability, which is calculated as: elimination ratio - wasted resources, where elimination ratio = number of eliminated opponents / number of opponents discovered, and wasted resources = number of shots * expected benefit of shooting - target health.
[0077] For the secondary indicator of autonomy of action, the corresponding tertiary indicator can be: 1) Task autonomous execution capability, which is calculated as: (task completion reward - task start reward) / task completion reward.
[0078] For the secondary indicator of action learning, the corresponding tertiary indicator can be: 1) the ability to discover the evolution of the target, which is calculated as: (autonomy of the agent after reinforcement learning - autonomy of the agent before reinforcement learning) / autonomy of the agent before reinforcement learning * 0.5 + 0.5.
[0079] For the secondary indicator of action coordination, the corresponding tertiary indicator can be: 1) Task coordination execution capability, which is calculated as the average number of responses of a certain operator to requests from other operators.
[0080] After establishing the three-level indicator system, this embodiment further initializes a Restricted Boltzmann Machine (RBM) network between the four indicator layers: the third-level indicator layer, the second-level indicator layer, the first-level indicator layer, and the ELO integration layer. This initial RBM network comprises multiple sequentially set indicator layers and several RBM networks located between adjacent indicator layers. The number of RBM networks located between adjacent indicator layers is determined by the number of nodes in the previous indicator layer of the adjacent layer. Assuming the k-1 level indicator layer has q nodes, then the k-level and k-1 level indicator layers contain a total of q "Gaussian-Bernoulli-Gaussian" RBM networks. That is, the number of RBM networks located between adjacent indicator layers is the same as the number of nodes in the previous indicator layer of the adjacent layer. The number of nodes in the second-level indicator layer is the number of second-level indicators included in the second-level indicator layer. The number of nodes in the first-level indicator layer is the number of first-level indicators included in the first-level indicator layer. The ELO integration layer has 1 node. Figure 3As shown, this embodiment can have a total of 12 secondary indicators, 4 primary indicators, and 1 ELO integral. Therefore, it contains a total of 17 Gaussian-Bernoulli-Gaussian restricted Boltzmann machine networks between the tertiary indicator layer, secondary indicator layer, primary indicator layer, and ELO integral layer. In this embodiment, the number of input and output layer nodes can be flexibly set according to the number of nodes in each indicator layer, and the number of hidden layer nodes is set to 16 (other values can also be used as needed) to complete the initialization process of each Gaussian-Bernoulli-Gaussian restricted Boltzmann machine network. After all Gaussian-Bernoulli-Gaussian restricted Boltzmann machine networks are initialized, the initial restricted Boltzmann machine network can be obtained, as shown below. Figure 4 As shown, Figure 4 Each dashed box in the diagram represents a Gaussian-Bernoulli-Gaussian restricted Boltzmann machine network.
[0081] Specifically, assuming that one node in the k-1 level index layer corresponds to p nodes in the k level index layer, taking the initialization of the restricted Boltzmann machine network as an example: with the k level index layer as the input layer, containing p nodes... i = 1, 2, ..., p, bias i = 1, 2, ..., p, standard deviation σ k With hidden layer H k→k-1 This is the output layer, containing 16 nodes. j = 1, 2, ..., 16, bias j = 1, 2, ..., 16, input hidden layer weight matrix i = 1, 2, ..., p, j = 1, 2, ..., 16, forming a Gauss-Bernoulli restricted Boltzmann machine; then, with hidden layer H... k→k-1 This is the input layer, containing 16 nodes. j = 1, 2, ..., 16, bias j = 1, 2, ..., 16, with a k-1 level index layer as the output layer, containing 1 node v k-1 Bias c k-1 Standard deviation σ k-1 Hidden output inter-layer weight matrix j = 1, 2, ..., 16, forming a Bernoulli-Gaussian Restricted Boltzmann Machine (RBMM). Through the above process, the initialization process of a Gaussian-Bernoulli-Gaussian RBMM network can be completed. Repeating the above process yields the initial RBMM network. It should be noted that the superscript k→k-1 indicates the relationship between the k-level index layer and the k-1-level index layer; the superscript k→h indicates the relationship between the k-level index layer and the hidden layers between the k and k-1 level index layers; and the superscript h→k-1 indicates the relationship between the hidden layers between the k and k-1 level index layers and the k-1 level index layer.
[0082] (3) Use the dataset to train the initial restricted Boltzmann machine network to obtain the trained restricted Boltzmann machine network.
[0083] This embodiment uses a three-level index as input and the agent's ELO integral as output to train an initial Restricted Boltzmann Machine (RBM) network. Training the initial RBM network using a dataset to obtain a trained RBM network includes: using the dataset as input, solving the objective function using maximum likelihood estimation to obtain the optimal network parameters of the initial RBM network; updating the initial RBM network based on the optimal network parameters to obtain the trained RBM network. Specifically, this embodiment can sequentially train 12 RBM networks between the three-level and two-level index layers, 4 RBM networks between the two-level and one-level index layers, and 1 RBM network between the one-level index layer and the ELO integral layer. After training 17 RBM networks, the trained RBM network is obtained.
[0084] Assume a node v in the k-1 level index layer k-1 , denoted as V k-1 This corresponds to p nodes in the k-level indicator layer. Let V be the denoted V k Taking the training of a restricted Boltzmann machine network as an example:
[0085] First, train the k-level index layer and hidden layer H from bottom to top. k→k-1 Network parameters between The training objective is to maximize the objective function logP(V) = i = 1, 2, ..., p, j = 1, 2, ..., 16. k ,θ k→h ).
[0086] objective function
[0087] Among them, E(V) k H k→k-1 ) is the energy function; v k Represents a node in the k-level indicator layer; h k→k-1 This represents the node of the hidden layer between the k-level and k-1-level index layers.
[0088]
[0089] The optimal network parameters θ are solved using the maximum likelihood estimation method. k→h :
[0090]
[0091] Among them, P(Hk→k-1 |V k ) indicates that given v k Under the conditions, h k→k-1 The probability of occurrence; P(V) k H k→k-1 ) represents v k and h k→k-1 The joint probability; data This represents the expected distribution of the input layer data; <·> recon This represents the desired distribution of the reconstructed input layer data. In this embodiment, reconstructed data is obtained through Gibbs sampling, starting from the input layer data V. k Start, use Predict the hidden layer data, then use Predicting the reconstructed input layer data yields the reconstructed data.
[0092]
[0093]
[0094] Where sigmoid(x) = 1 / (1+exp(-x)); It satisfies the mean μ and variance σ 2 The Gaussian distribution.
[0095] renew
[0096]
[0097]
[0098]
[0099] Where α is the learning rate, (·) * The data is reconstructed using Gibbs sampling. Indicates that given reconstructed data Under the conditions, h k→k-1 Reconstructed data The probability of occurrence.
[0100] Then, keep θ k→h Keeping the hidden layer H unchanged, train it. k→k-1 Network parameters between the k-1 level indicator layer j = 1, 2, ..., 16, the training objective is to maximize the objective function logP(H k →k-1 ,θ h→k-1 ).
[0101] objective function
[0102] Among them, E(H k→k-1 V k-1 ) is the energy function;
[0103]
[0104] The optimal network parameters θ are solved using the maximum likelihood estimation method. h→k-1 :
[0105]
[0106] Among them, <·> data This represents the expected distribution of the hidden layer data; <·> recon This represents the desired distribution of the reconstructed hidden layer data. In this embodiment, reconstructed data is obtained through Gibbs sampling, starting from the hidden layer data H. k→k-1 To begin, use P(v) k-1 |H k →k-1 Sampling of the output layer, using Sampling is performed on the hidden layer to obtain the reconstructed data (H). k→k-1 ) * .
[0107]
[0108]
[0109] Where sigmoid(x) = 1 / (1+exp(-x)); It satisfies the mean μ and variance σ 2 The Gaussian distribution.
[0110] renew
[0111]
[0112]
[0113]
[0114] Where α is the learning rate, (·) * This is the reconstructed data obtained through Gibbs sampling.
[0115] The training process for each Restricted Boltzmann Machine (RBM) network is performed sequentially as described above. Once all RBMs are trained, a well-trained RBM network is obtained. Following these steps, this embodiment first initializes the RBM network, calculates the Level 3 indicators based on historical replay data as input, and uses the agent's ELO score as output, then trains the initial RBM network to obtain a well-trained RBM network. In application scenarios, the Level 3 indicators are calculated based on replay data as input, and the trained RBM network is used for prediction to obtain the agent's default Level 3 indicators, agent ELO score, and intelligence level.
[0116] Preferably, after acquiring the dataset, the evaluation method in this embodiment further includes: selecting the maximum value among the historical ELO scores of all tested agents as the upper limit, and selecting the minimum value among the historical ELO scores of all tested agents as the lower limit; uniformly dividing the range composed of the upper and lower limits to obtain multiple intervals, each interval corresponding to an intelligence level, the larger the maximum value of the interval, the higher the intelligence level corresponding to the interval. For example, in this embodiment, the highest and lowest historical ELO scores of all tested agents are used as the upper and lower limits, respectively, to uniformly divide five intervals, corresponding to five intelligence levels, which are recorded from low to high as level one, level two, level three, level four, and level five. In Table 2, the highest and lowest scores are 941 and 552, respectively, which can be uniformly divided into [552-629], [630-707], [708-785], [786-863], and [864-941], corresponding to the five intelligence levels of level one, level two, level three, level four, and level five, respectively.
[0117] In the application scenario, the three-level indicators are calculated based on the review data as input, and a trained restricted Boltzmann machine network is used for prediction to obtain the final ELO score of the tested agent and its corresponding intelligence level. Specifically, determining the intelligence level of the tested agent based on its final ELO score may include: judging the interval in which the final ELO score of the tested agent falls; and determining the intelligence level of the tested agent based on the intelligence level corresponding to the interval.
[0118] After obtaining multiple evaluation metrics, if several evaluation metrics have no values (i.e., there can be one or more evaluation metrics with no values), the evaluation method in this embodiment further includes: labeling the evaluation metrics with no values as unknown metrics, labeling the evaluation metrics with values as known metrics, and randomly assigning an initial value to the unknown metrics. Using the values of the known metrics and the initial values of the unknown metrics as input, the trained Restricted Boltzmann Machine (RBM) network is used to predict the values of the upper-level metrics corresponding to the unknown metrics. Using the values of the upper-level metrics corresponding to the unknown metrics as input, the trained RBM network is used to predict the values of the unknown metrics. It is determined whether the difference between the value of the unknown metrics and the initial value of the unknown metrics is less than a preset threshold. If yes, the iteration ends, and the value of the unknown metrics is used as the actual value of the unknown metrics. If no, the iteration continues, using the value of the unknown metrics as the initial value of the unknown metrics in the next iteration, and returning to the step of "using the values of the known metrics and the initial values of the unknown metrics as input, and using the trained RBM network to predict the values of the upper-level metrics corresponding to the unknown metrics".
[0119] In this embodiment, there are 22 tertiary indicators, 12 secondary indicators, 4 primary indicators, and 1 ELO score. Based on the technical characteristics of the secondary indicators reflected by the tertiary indicators, custom tertiary indicators I1 and I2 correspond to the secondary indicators. Custom Level 3 Indicator I3 corresponds to Level 2 Indicator Custom tertiary indicator I4 and custom tertiary indicator I5 correspond to secondary indicators. ...Custom Level 3 Indicator I 21 and custom three-level indicators I 22 Corresponding to secondary indicators The numerical representation of the third-level indicators is denoted as follows: Collectively denoted as V 3 The numerical representation of the secondary indicator is denoted as... Collectively denoted as V 2 The numerical representation of the primary indicator is denoted as... Collectively denoted as V 1 The ELO integral is denoted as v. 0 , denoted as V 0 .
[0120] Specifically, this embodiment can use a trained Restricted Boltzmann Machine network for prediction to obtain the agent's default three-level index. It is assumed that there is an index v in the second-level index layer. 2 , denoted as V 2 This corresponds to p indicators in the three-level indicator layer. Collectively denoted as V 3 During this period, the three-level indicators Default (i.e., no value), using the default third-level index. For example, the prediction:
[0121] For the default three-level indicators The user provides a random initial value, and the conditional probability formula is used. and P(v) 2 |H 3→2 ), thus inferring the upper-level secondary indicator v corresponding to the tertiary indicator. 2 Then use the conditional probability formula and The default third-level indicator is inferred. When the predicted default level 3 metric differs significantly from the user-provided initial value, the predicted default level 3 metric is substituted into the conditional probability formula. and P(v) 2 |H 3→2 The process continues until the difference between the predicted default level 3 indicator and the previous prediction is less than 0.0001, at which point the default level 3 indicator can be obtained.
[0122] For example, some of the agent's third-level metrics are shown in Table 3 below, with the custom third-level metric I2 being the default:
[0123] Table 3
[0124]
[0125] For agent O, the default three-level index The user provides a random initial value, and the conditional probability formula is used. and The upper-level secondary indicators corresponding to the tertiary indicators are inferred. Then, according to the conditional probability formula and The default third-level indicator is inferred. When the predicted default level 3 metric differs significantly from the user-provided initial value, the predicted default level 3 metric is substituted into the conditional probability formula. and The estimation is repeated until the difference between the estimated default third-level indicator and the previous estimation is less than 0.0001, at which point the default third-level indicator is obtained.
[0126] Furthermore, this embodiment can use an RBM (Restricted Boltzmann Machine) network for prediction to obtain the ELO score of the tested agent and its corresponding intelligence level. Specifically, based on known three-level indicators and inferred default three-level indicators, the conditional probability formula is used. and P(v) 2 |H 3→2 All secondary indicators were inferred using the conditional probability formula. and P(v) 1 |H 2→1 All primary indicators were inferred using the conditional probability formula. and P(v) 0 |H 1→0 By inferring the ELO score of the agent and mapping it to the five intervals, the level of intelligence of the agent can be determined.
[0127] In Table 3, the ELO integral of agent O is estimated to be 843, which corresponds to the fourth level of intelligence in the interval [786-863].
[0128] It should be noted that the above conditional probability formulas can all be determined based on a trained Restricted Boltzmann Machine (RBM) network. The process of making inferences using the conditional probability formulas belongs to the internal working principle of the trained RBM network.
[0129] The intelligence level evaluation method based on Restricted Boltzmann Machine (RBM) provided in this embodiment aims to objectively evaluate the intelligence level of an agent, and this evaluation method is highly versatile. To achieve this objective, the evaluation method of this embodiment first establishes a three-level index system. In typical experimental scenarios, experimental tasks targeting the agent's capabilities are constructed, with agents competing in pairs to obtain post-test data. The agent's ELO score is calculated using an integral iterative formula, and the highest and lowest scores are used as upper and lower limits to evenly divide the network into five intervals, corresponding to five intelligence level levels. Next, the RBM network is initialized, and the three-level indexes are calculated based on the post-test data as input, with the agent's ELO score as output, for training. In application scenarios, the three-level indexes are calculated based on the post-test data as input, and the trained RBM network is used for prediction to obtain the agent's default three-level indexes, agent ELO score, and intelligence level. Through this method, this embodiment objectively evaluates the intelligence level of an agent, and it is highly versatile.
[0130] Compared to existing evaluation methods, the evaluation method in this embodiment has the following advantages:
[0131] (1) The intelligent level assessment method based on restricted Boltzmann machine proposed in this embodiment is an objective method. It only needs to give the structure of the three-level index system, without specifying the aggregate weight of each index.
[0132] (2) The intelligent level evaluation method based on restricted Boltzmann machine proposed in this embodiment can infer the default three-level indicators in the application scenario when the three-level indicators are defaulted, and further infer the agent's ELO score.
[0133] (3) The intelligent level assessment method based on restricted Boltzmann machine proposed in this embodiment is a general method that can be used to assess the general intelligent level of different intelligent agents’ perception, cognition, decision-making and action capabilities.
[0134] Example 2:
[0135] This embodiment provides an intelligent level assessment system based on a restricted Boltzmann machine, such as... Figure 5 As shown, the evaluation system includes:
[0136] The data acquisition module M1 is used to acquire replay data obtained by multiple tested agents in pairwise battles during an experimental task; the amount of replay data is the same as the number of battles; the replay data includes adjudication data and the battle data of each of the two tested agents in the battle; the adjudication data is used to characterize the battle result; the battle data is used to characterize the battle situation.
[0137] The integral prediction module M2 is used to determine all the replay data corresponding to each tested agent; for each replay data corresponding to the tested agent, multiple evaluation indicators are calculated based on the tested agent's combat data in the replay data, and the ELO score of the tested agent is predicted using a trained restricted Boltzmann machine network with the multiple evaluation indicators as input; the average of all the ELO scores of the tested agent is calculated to obtain the final ELO score of the tested agent.
[0138] The level determination module M3 is used to determine the intelligence level of each tested intelligent agent based on its final ELO score.
[0139] Each embodiment in this specification focuses on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be found in the method section.
[0140] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for evaluating the intelligence level based on a restricted Boltzmann machine, characterized in that, The evaluation method includes: The test involves acquiring debriefing data from pairwise battles between multiple tested agents in an experimental task. The amount of debriefing data is the same as the number of battles. The debriefing data includes adjudication data and the battle data of each of the two tested agents. The adjudication data is used to characterize the battle results, and the battle data is used to characterize the battle situation. For each tested agent, all the replay data corresponding to the tested agent are determined; for each replay data corresponding to the tested agent, multiple evaluation indicators are calculated based on the tested agent's battle data in the replay data, and the ELO score of the tested agent is predicted using a trained restricted Boltzmann machine network with the multiple evaluation indicators as input; the average of all the ELO scores of the tested agent is calculated to obtain the final ELO score of the tested agent. For each tested agent, the intelligence level of the tested agent is determined based on the final ELO score of the tested agent; Before predicting the ELO score of the tested agent using a trained Restricted Boltzmann Machine (RBM) network with multiple evaluation metrics as input, the evaluation method further includes: training an initial RBM network to obtain a trained RBM network, specifically including: Obtain a dataset; the dataset includes multiple samples, the samples include historical values of multiple evaluation metrics and historical values of the ELO scores of the tested agent; An initial restricted Boltzmann machine (RBM) network is constructed. This initial RBM network comprises multiple sequentially arranged index layers and several RBM networks located between adjacent index layers. The index layers include, from bottom to top, a three-level index layer, a two-level index layer, a first-level index layer, and an ELO integral layer. The first-level index layer includes first-level indicators reflecting the agent's perception, cognition, decision-making, and action capabilities. The second-level index layer includes second-level indicators reflecting the agent's autonomy, learning, and collaboration. The third-level index layer includes third-level indicators reflecting the technical characteristics of the second-level indicators; these third-level indicators are the evaluation indicators. The ELO integral layer includes the ELO score of the tested agent. The initial Restricted Boltzmann Machine (RBM) network is trained using the dataset to obtain a trained RBM network.
2. The evaluation method according to claim 1, characterized in that, The acquisition of the dataset specifically includes: Historical replay data is obtained from pairwise battles between multiple tested agents in an experimental task; the number of historical replay data is the same as the number of battles; the historical replay data includes historical adjudication data and the historical battle data of each of the two tested agents in the battle; the historical adjudication data is used to characterize the battle results; the historical battle data is used to characterize the battle situation. For each tested agent, all historical replay data corresponding to the tested agent are determined; for each historical replay data corresponding to the tested agent, historical values of multiple evaluation indicators are calculated based on the historical battle data of the tested agent in the historical replay data; based on the historical adjudication data in the historical replay data, the historical value of the tested agent's ELO score is calculated using an integral iteration formula; the historical values of multiple evaluation indicators and the historical value of the tested agent's ELO score form a sample, and all samples form a dataset.
3. The evaluation method according to claim 2, characterized in that, The integral iteration formula includes: , ; in, For agent A in the first... ELO points after the first match; For agent A in the th ELO points after the first match; The coefficient of change of the game integral; This represents the actual score of agent A in the battle; For agent A in the first... Expected win rate in the next match; For agent B in the 1st ELO points after the first match; For agent B in the 1st ELO points after the first match; This represents the actual score of agent B in the battle. For agent B in the 1st Expected win rate in the next match; The formula for calculating the expected win rate is as follows: , ,constant .
4. The evaluation method according to claim 1, characterized in that, The number of restricted Boltzmann machine networks located between adjacent index layers is the same as the number of nodes in the index layer above the adjacent index layer; the number of nodes in the second-level index layer is the number of second-level indices included in the second-level index layer; the number of nodes in the first-level index layer is the number of first-level indices included in the first-level index layer; the number of nodes in the ELO integration layer is 1.
5. The evaluation method according to claim 1, characterized in that, The step of training the initial Restricted Boltzmann Machine (RBM) network using the dataset to obtain a trained RBM network specifically includes: Using the dataset as input, the objective function is solved using the maximum likelihood estimation method to obtain the optimal network parameters of the initial restricted Boltzmann machine network; The initial restricted Boltzmann machine network is updated according to the optimal network parameters to obtain a trained restricted Boltzmann machine network.
6. The evaluation method according to claim 1, characterized in that, After obtaining the dataset, the evaluation method further includes: The maximum value among the historical ELO scores of all the tested agents is selected as the upper limit, and the minimum value among the historical ELO scores of all the tested agents is selected as the lower limit. The range consisting of the upper limit and the lower limit is evenly divided to obtain multiple intervals; each interval corresponds to an intelligence level; the larger the maximum value of the interval, the higher the intelligence level of the interval.
7. The evaluation method according to claim 6, characterized in that, The process of determining the intelligence level of the tested agent based on its final ELO score specifically includes: Determine the interval in which the final ELO score of the tested agent lies; The intelligence level of the tested intelligent agent is determined based on the intelligence level corresponding to the interval.
8. The evaluation method according to claim 1, characterized in that, After obtaining multiple evaluation indicators, if several of the evaluation indicators have no values, the evaluation method further includes: The evaluation index with no value is marked as an unknown index, and the evaluation index with a value is marked as a known index; an initial value is randomly assigned to the unknown index. Using the known values of the indicators and the initial values of the unknown indicators as input, the trained Restricted Boltzmann Machine Network is used to predict the values of the upper-level indicators corresponding to the unknown indicators. Using the value of the upper-level index corresponding to the unknown index as input, the value of the unknown index is predicted by the trained restricted Boltzmann machine network. Determine whether the difference between the value of the unknown indicator and the initial value of the unknown indicator is less than a preset threshold; If yes, the iteration ends, and the value of the unknown indicator is taken as the actual value of the unknown indicator; if no, the iteration continues, and the value of the unknown indicator is taken as the initial value of the unknown indicator in the next iteration, returning to the step of "using the value of the known indicator and the initial value of the unknown indicator as input, and using the trained restricted Boltzmann machine network to predict the value of the upper-level indicator corresponding to the unknown indicator".
9. A system for assessing the level of intelligence based on a restricted Boltzmann machine, characterized in that, The evaluation system includes: The data acquisition module is used to acquire replay data obtained from pairwise battles between multiple tested agents in an experimental task; the amount of replay data is the same as the number of battles; the replay data includes adjudication data and the battle data of each of the two tested agents; the adjudication data is used to characterize the battle result; the battle data is used to characterize the battle situation. The integral prediction module is used to determine all the replay data corresponding to each tested agent; for each replay data corresponding to the tested agent, calculate multiple evaluation indicators based on the tested agent's combat data in the replay data, and use the multiple evaluation indicators as input to predict the ELO score of the tested agent using a trained restricted Boltzmann machine network; calculate the average of all the ELO scores of the tested agent to obtain the final ELO score of the tested agent. The level determination module is used to determine the intelligence level of each tested intelligent agent based on its final ELO score. Before predicting the ELO score of the tested agent using a trained Restricted Boltzmann Machine (RBM) network with multiple evaluation metrics as input, the evaluation system further includes: training an initial RBM network to obtain a trained RBM network, specifically including: Obtain a dataset; the dataset includes multiple samples, the samples include historical values of multiple evaluation metrics and historical values of the ELO scores of the tested agent; An initial restricted Boltzmann machine (RBM) network is constructed. This initial RBM network comprises multiple sequentially arranged index layers and several RBM networks located between adjacent index layers. The index layers include, from bottom to top, a three-level index layer, a two-level index layer, a first-level index layer, and an ELO integral layer. The first-level index layer includes first-level indicators reflecting the agent's perception, cognition, decision-making, and action capabilities. The second-level index layer includes second-level indicators reflecting the agent's autonomy, learning, and collaboration. The third-level index layer includes third-level indicators reflecting the technical characteristics of the second-level indicators; these third-level indicators are the evaluation indicators. The ELO integral layer includes the ELO score of the tested agent. The initial Restricted Boltzmann Machine (RBM) network is trained using the dataset to obtain a trained RBM network.
Citation Information
Patent Citations
Systems and methods for modeling probability distributions
CN111758108A
Automatic course training method for intelligent model playing chess with rules
CN111882072A