Coverage rate testing method and device for deep reinforcement learning, equipment and storage medium
Through Gaussian noise algorithm and hierarchical data structure to calculate the coverage rate of deep reinforcement learning, the problem of inaccurate coverage in high-dimensional continuous space is solved, and more accurate coverage quantification and effective exploration of agent state space are achieved.
Patent Information
- Application Number
- CN202510204053.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for the prior art to accurately calculate the coverage rate of deep reinforcement learning, especially in high-dimensional continuous spaces, where multi-grained coverage standards and approximate coverage evaluation algorithms have problems with inaccurate calculations.
The Gaussian noise algorithm is used to generate candidate states, and the Euclidean distance of the candidate state is calculated using hierarchical data structures such as HNSW and kernel density estimation. The candidate coverage is determined based on the target Euclidean distance, and the optimal test case is determined through kernel function and coverage comparison.
It realizes accurate calculation of coverage in high-dimensional continuous space, avoids discretization errors, improves the accuracy of coverage evaluation and the effectiveness of test cases, and improves the agent's ability to explore state space.
Smart Images

Figure CN120336993A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a coverage testing method, apparatus, device, and storage medium for deep reinforcement learning. Background Art
[0002] For deep reinforcement learning, coverage testing focuses on the exploration degree of the agent in the state space and action space. However, when performing coverage testing on deep reinforcement learning through the multi-granularity coverage standard method, it is necessary to rely on the statistics of the number of occurrences of elements such as states and actions to calculate the coverage rate. Since the number of states and actions in the continuous space is theoretically infinite and difficult to accurately count, it will lead to inaccurate calculation of the coverage rate.
[0003] In view of this, how to accurately perform coverage testing on deep reinforcement learning has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, an object of the present disclosure is to propose a coverage testing method, apparatus, device, and storage medium for deep reinforcement learning to solve or partially solve the above technical problems.
[0005] Based on the above object, a first aspect of the present disclosure proposes a coverage testing method for deep reinforcement learning, the method comprising:
[0006] Obtaining an initial state and a set of historical states of an agent in a deep reinforcement learning model;
[0007] Obtaining candidate states according to the initial state through a Gaussian noise algorithm, and constructing a hierarchical data structure by using the set of historical states;
[0008] Determining a target Euclidean distance from the candidate states to the set of historical states according to the hierarchical data structure, and determining a candidate coverage rate of the candidate states based on the target Euclidean distance;
[0009] Determining an optimal test case from the candidate states and the current state based on the candidate coverage rate and the current coverage rate of the current state.
[0010] Based on the same inventive concept, a second aspect of the present disclosure proposes a coverage testing apparatus for deep reinforcement learning, comprising:
[0011] An obtaining module, configured to obtain an initial state and a set of historical states of an agent in a deep reinforcement learning model;
[0012] A candidate state determining module, configured to obtain candidate states according to the initial state through a Gaussian noise algorithm, and construct a hierarchical data structure by using the set of historical states;
[0013] A candidate coverage rate determination module, configured to determine a target Euclidean distance from the candidate state to the set of historical states according to the hierarchical data structure, and determine a candidate coverage rate of the candidate state based on the target Euclidean distance;
[0014] An optimal test case determination module, configured to determine an optimal test case from the candidate state and the current state based on the candidate coverage rate and the current coverage rate of the current state.
[0015] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the method described above is implemented.
[0016] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described above.
[0017] As can be seen from the above, the present disclosure provides a coverage rate testing method, apparatus, device, and storage medium for deep reinforcement learning. The initial state of the agent in the deep reinforcement learning model and the set of historical states are obtained. A candidate state is obtained according to the initial state by the Gaussian noise algorithm, and a hierarchical data structure is constructed using the set of historical states. The target Euclidean distance from the candidate state to the set of historical states is determined according to the hierarchical data structure, and the candidate coverage rate of the candidate state is determined based on the target Euclidean distance, which can accurately calculate the candidate coverage rate of the candidate state in a high-dimensional continuous space and does not require discretization of the set of historical states, avoiding coverage errors caused by discretization and achieving more accurate quantification of the continuous state space coverage rate. Based on the candidate coverage rate and the current coverage rate of the current state, an optimal test case is determined from the candidate state and the current state, and the optimal test case can better measure the exploration of the agent in the state space. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of the coverage rate testing method for deep reinforcement learning according to an embodiment of the present disclosure;
[0020] Figure 2The first schematic diagram constructed for the two-dimensional kernel density estimation of the embodiments of the present disclosure;
[0021] Figure 3 The second schematic diagram constructed for the two-dimensional kernel density estimation of the embodiments of the present disclosure;
[0022] Figure 4 The structural schematic diagram of the coverage test device for deep reinforcement learning of the embodiments of the present disclosure;
[0023] Figure 5 The structural schematic diagram of the electronic device of the embodiments of the present disclosure. Detailed implementation manners
[0024] To make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0025] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be the general meanings understood by those of ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0026] Based on the description of the background technology, coverage testing was initially used in software testing to evaluate the sufficiency of testing by measuring the coverage of test cases for code or functions. With the development of artificial intelligence, the concept of coverage testing has been extended to the fields of deep neural networks (DNN) and deep reinforcement learning (DRL). In DNN testing, coverage testing measures the coverage of test inputs for the internal structure (such as neuron activation) or behavior (such as output distribution) of the model. In reinforcement learning, coverage testing focuses on the exploration degree of the agent for the state space and action space, ensuring that it fully covers the key areas in the environment and avoiding policy biases. Through state, action, trajectory and reward coverage analysis, coverage testing helps to optimize the training strategy and improve the robustness and generalization ability of the agent, and is an important evaluation tool in reinforcement learning.
[0027] Coverage testing usually consists of two core parts: the definition of coverage metrics and the design of test case generation methods. These two parts work together to ensure that test cases can effectively cover the key parts of the system, thereby discovering potential errors or vulnerabilities. Coverage metrics can reflect the exploration breadth of the agent in the environmental state and action space, measuring whether it has conducted comprehensive exploration to avoid falling into local optima. The test case generation method, based on environmental characteristics and algorithm expectations, designs diverse input scenarios to comprehensively test the performance of the algorithm. These two complement each other and provide important support for the optimization and reliable operation of the reinforcement learning algorithm. By evaluating the sufficiency of exploration through coverage and using this to guide the generation of test cases, more targeted and comprehensive test scenarios can be constructed more efficiently, not only helping the algorithm make robust decisions in complex environments but also significantly enhancing its adaptability and reliability in complex environments.
[0028] The coverage testing methods in related technologies include: multi-granularity coverage criteria, approximate coverage APC evaluation algorithm, and coverage maximization algorithm. The coverage testing methods are specifically described as follows:
[0029] The multi-granularity coverage criteria start from three levels of state, action, and policy, and design multi-granularity coverage metrics such as state boundary coverage (SBCov) and action boundary coverage (ABCov), considering multi-level and multi-granularity coverage requirements. In addition, these criteria are optimized by applying the principle of combinatorial coverage. The state coverage rate is used as a feedback mechanism to guide the selection of test cases for the DRL system.
[0030] Although the multi-granularity coverage criteria are innovative in exploring the test sufficiency of the DRL system, they have limitations. The multi-granularity coverage criteria method essentially relies on counting the number of occurrences of elements such as states and actions to calculate the coverage rate. When calculating metrics such as state boundary coverage (SBCov) and action boundary coverage (ASCov), it is based on counting the number of state dimension values and the number of partition regions covered. In continuous space scenarios, this approach faces difficulties. Since the number of states and actions in continuous space is theoretically infinite and difficult to accurately count, it leads to inaccurate calculations. For example, when dealing with high-dimensional continuous state spaces, determining whether a state is covered and calculating the coverage degree become extremely complex, and there will be a large number of states that are difficult to effectively count, affecting the accurate evaluation of the test sufficiency of the DRL system.
[0031] The approximate coverage APC evaluation algorithm uses t-SNE to project the high-dimensional state space into two dimensions and approximately calculates the coverage rate through an occupancy grid, solving the coverage rate evaluation problem caused by the high-dimensional state space and the synchronization of data generation and training in deep reinforcement learning. The coverage rate maximization algorithm is based on the ε-greedy strategy and combines the Rapidly-exploring Random Tree (RRT) to search for unknown initial states and perturb high-score initial states, realizing the feedback on the simulated environment, accelerating the improvement of the coverage rate, effectively balancing exploration and exploitation, and improving the coverage rate and rewards in standard tasks.
[0032] The approximate coverage APC evaluation algorithm has problems in many aspects: In the coverage rate evaluation link, although using t-SNE for dimensionality reduction preserves the local neighborhood structure, it may lose important information and affect the accuracy of coverage rate evaluation; In the two-dimensional grid approximation calculation method, the choice of the grid cell size is crucial, and it is difficult to balance between granularity and computational complexity, and different choices will lead to large differences in evaluation results. In the coverage rate maximization stage, the initialization of the ε value and the setting of the attenuation rate in the ε-greedy strategy lack theoretical basis and may not achieve the best balance between exploration and exploitation; The search efficiency of the RRT structure is greatly affected by the dimensionality and shape of the state space, and it is prone to problems such as slow search speed and getting stuck in local optima in high-dimensional complex spaces; The perturbation of high-score initial states is based on Gaussian noise, and the method is single, which may not fully explore the state space and limit the improvement of the algorithm performance.
[0033] As mentioned above, how to accurately conduct coverage rate testing for deep reinforcement learning has become an important research issue.
[0034] Based on the above description, as Figure 1 shown, the coverage rate testing method for deep reinforcement learning proposed in this embodiment includes:
[0035] Step 101, obtain the initial state and the set of historical states of the agent in the deep reinforcement learning model.
[0036] Specifically, the initial state s of the agent in the deep reinforcement learning model is obtained through random sampling init . The state space of the agent in the deep reinforcement learning model is obtained, and the set of historical states H = {s1′, s2′,..., s N ′} is determined from the state space.
[0037] In reinforcement learning, the state space is a core concept. The state space refers to the set of all situations or environmental states that the agent may be in, usually denoted by the symbol S.
[0038] Step 102: Obtain candidate states according to the initial state through the Gaussian noise algorithm, and construct a hierarchical data structure using the historical state set.
[0039] Specifically, when implemented, through the Gaussian noise algorithm, noise with a probability density function following a Gaussian distribution is added to the initial state s init to obtain a noise state s noise . The noise state s noise is added to the candidate set Candidates, and a candidate state s candidate is selected from the candidate set Candidates. The candidate coverage Coverage(s candidate ) of each candidate state s candidate is calculated respectively.
[0040] Use the historical state set H = {s1′, s2′, …, s N ′} to construct a hierarchical data structure T for quickly indexing states. Among them, the hierarchical data structure is the data structure of Hierarchical Navigable Small Worldgraphs (abbreviated as HNSW). The core idea is to quickly narrow the search range and accurately locate the target through hierarchical navigation. The HNSW data structure constructs a multi-layer graph from sparse to dense. The top layer is used for quick navigation, and the bottom layer is used for accurate search. Nodes are connected through a network to ensure that the shortest path between any two points is short. When querying, start from the top layer and gradually move down to the lower layer, and use greedy search to approach the target point, and finally find the nearest neighbor in the bottom layer. The HNSW data structure has the characteristics of high efficiency, high precision, and dynamic update ability, and is especially suitable for real-time application scenarios of high-dimensional data and large-scale data sets.
[0041] Step 103: Determine the target Euclidean distance from the candidate state to the historical state set according to the hierarchical data structure, and determine the candidate coverage of the candidate state based on the target Euclidean distance.
[0042] Specifically, when implemented, the candidate state s candidate is a d-dimensional continuous vector s = (s1, s2, …, s N ). According to the hierarchical data structure, the target Euclidean distance from the candidate state of each dimension to the historical state set is determined respectively. Based on the target Euclidean distance, the kernel function value of the candidate state of each dimension is determined, and the sum of the kernel function values of the candidate states of all dimensions is obtained through summation processing. The coverage of the candidate state is calculated according to the sum of the kernel functions.
[0043] Step 104: Determine the optimal test case from the candidate state and the current state based on the candidate coverage and the current coverage of the current state.
[0044] In specific implementation, the current state is the currently optimal state stored. By comparing and processing the candidate coverage rate and the current coverage rate, it is determined whether the currently optimal state needs to be updated, and the state with the larger coverage rate in the candidate state and the current state is used as the optimal test case.
[0045] Through the above embodiments, the initial state and the set of historical states of the agent in the deep reinforcement learning model are obtained. The candidate state is obtained from the initial state by the Gaussian noise algorithm, and the hierarchical data structure is constructed using the set of historical states. The target Euclidean distance from the candidate state to the set of historical states is determined according to the hierarchical data structure, and the candidate coverage rate of the candidate state is determined based on the target Euclidean distance, which can accurately calculate the candidate coverage rate of the candidate state in the high-dimensional continuous space, and there is no need to discretize the set of historical states, avoiding the coverage rate error caused by discretization, and realizing more accurate quantification of the continuous state space coverage rate. Based on the candidate coverage rate and the current coverage rate of the current state, the optimal test case is determined from the candidate state and the current state, and based on the optimal test case, the exploration of the state space by the agent can be better measured.
[0046] In some embodiments, step 102 includes:
[0047] Step 1021, adding noise with a probability density function following a Gaussian distribution to the initial state through the Gaussian noise algorithm to obtain a noise state.
[0048] Step 1022, adding the noise state to the candidate set and selecting a candidate state from the candidate set.
[0049] In specific implementation, through the Gaussian noise algorithm, noise with a probability density function following a Gaussian distribution is added to the initial state s init to obtain a noise state s noise .
[0050] s noise = s init + N(0, σ 2 )
[0051] where s noise is the noise state, s init is the initial state, and N(0, σ 2 ) is the noise with a probability density function following a Gaussian distribution.
[0052] The noise state s noise is added to the candidate set Candidates, and a candidate state s candidate is selected from the candidate set Candidates, and the candidate coverage rate Coverage(s candidate ) of each candidate state scandidate ), and use the candidate state with the largest coverage rate among multiple candidate states as the optimal test case.
[0053] Through the above solution, by using the Gaussian noise algorithm, noise with a probability density function obeying a Gaussian distribution can be added to the initial state, so as to obtain a noise state based on the initial state. By adding the noise state to the candidate set, it is convenient to sequentially select candidate states from the candidate set to calculate the coverage rate, and use the candidate state with the largest coverage rate among multiple candidate states as the optimal test case.
[0054] In some embodiments, step 103 includes:
[0055] Step 1031, determine the target Euclidean distance from the candidate state of each dimension to the historical state set according to the hierarchical data structure.
[0056] Step 1032, determine the kernel function value of the candidate state of each dimension based on the target Euclidean distance, and sum the kernel function values of the candidate states of all dimensions to obtain a kernel function sum.
[0057] Step 1033, obtain the coverage rate of the candidate state according to the kernel function sum.
[0058] Specifically, the HNSW data structure constructs a multi-layer graph from sparse to dense, with the top layer for fast navigation and the bottom layer for accurate search. Nodes are connected through a network to ensure that the shortest path between any two points is short. According to the hierarchical data structure, determine the target Euclidean distance from the candidate state of each dimension to the historical state set. Specifically, when querying, start from the top layer and gradually move down to the lower layer, using greedy search to approximate the target point, and finally find the nearest neighbor in the bottom layer. The HNSW data structure has the advantages of high efficiency, high precision, and dynamic update ability, and is especially suitable for real-time application scenarios of high-dimensional data and large-scale data sets.
[0059] Candidate state s candidate is a d-dimensional continuous vector s=(s1, s2, …, s N ). For example, for the candidate state s i of the i-th dimension, by querying the HNSW data structure, determine the candidate state s i of the i-th dimension to the historical state set H={s1′, s2′, …, s N ′} of the target Euclidean distance d i . Through the kernel density estimation algorithm, based on the target Euclidean distance d i determine the kernel function value K(d i) After calculating the kernel function values for all dimensions of the candidate state, sum up the kernel function values of the candidate states for all dimensions to obtain the kernel function sum S. Calculate the coverage Coverage(s) of the candidate state based on the kernel function sum S.
[0060] Kernel Density Estimation (KDE) is a non-parametric statistical method used to estimate the probability density function of a random variable. Kernel density estimation generates a smooth density curve by superimposing kernel functions around each data point. The basic idea of KDE is to consider each sample point as a center and use a kernel function (usually a symmetric probability density function) to measure the contribution of that sample point to other points. By superimposing the contributions of all sample points, the probability density function of the entire sample space can be obtained.
[0061] As a non-parametric statistical method, Kernel Density Estimation (KDE) has the characteristics of high flexibility and no need to assume the form of data distribution, and can effectively capture the probability density structure of complex data. Especially in high-dimensional continuous state spaces, KDE generates a smooth density curve through the superposition of kernel functions, which can accurately reflect the distribution characteristics of the state space. In addition, the kernel bandwidth parameter of KDE provides flexible control over the neighborhood range, enabling it to adapt to the exploration needs of state spaces at different scales. These characteristics make KDE an ideal tool for quantifying the exploration degree of an agent in the environment and the density of state access.
[0062] Based on the above advantages, the embodiments of the present disclosure adopt KDE as the core method to design the coverage metric (KS Coverage). By considering each historical state as a kernel function center and calculating the sum of the kernel function values between the target state and all historical states, the access density of the neighborhood where the target state is located can be quantified. Combining with the exponential mapping function, the output of KDE is normalized to a unified interval, thus intuitively reflecting the level of coverage. This method can not only accurately measure the exploration breadth of the agent in the state space, but also provide a reliable guiding basis for test case generation, ensuring that the generated test scenarios can effectively cover unexplored or sparsely accessed state regions, thereby enhancing the robustness and adaptability of the reinforcement learning algorithm.
[0063] Figure 2 The first schematic diagram constructed for the two-dimensional kernel density estimation of the embodiments of the present disclosure. As Figure 2 shown, the data points and their respective kernel functions are represented by gray dashed lines, where Figure 2 the black dots in are data points, Figure 2 the gray dashed lines around the black dots in are the kernel functions corresponding to the data points, Figure 2 the abscissa and ordinate in are the corresponding feature information. Figure 3A second schematic diagram of a two-dimensional kernel density estimation constructed according to an embodiment of the present disclosure. Figure 3 As shown, the sum of the kernel functions is the kernel density estimate, which reflects the coverage of all sample points in the entire sample space, where Figure 3 The black points in are sample points. Figure 3 The color range in the graph indicates the coverage of all sample points in the entire sample space. Figure 3 The horizontal and vertical coordinates are the corresponding feature information.
[0064] Through the above scheme, according to the hierarchical data structure, the target Euclidean distance from the candidate state of each dimension to the historical state set is determined, which can reduce the distance calculation complexity of high-dimensional states from the linear level to the logarithmic level, significantly improving the query efficiency of large-scale state sets. Based on the target Euclidean distance, the kernel function value of the candidate state of each dimension is determined, and the kernel function value of the candidate state of all dimensions is summed to obtain the kernel function sum, which can convert the discrete state distance into a continuous probability density and overcome the information loss caused by the grid discretization. The coverage rate of the candidate state is obtained according to the kernel function sum. The coverage rate growth rate can be adaptively adjusted.
[0065] In some embodiments, step 1031 includes:
[0066] Step 1031A, determining the current dimension from the multiple dimensions of the candidate states.
[0067] Step 1031B: determine a plurality of historical states from the historical state set.
[0068] Step 1031C: Determine multiple current Euclidean distances between the candidate state of the current dimension and each historical state according to the hierarchical data structure.
[0069] Step 1031D, performing mean processing on the multiple current Euclidean distances to obtain the target Euclidean distance from the candidate state of the current dimension to the historical state set.
[0070] Step 1031E, determine the next dimension from the multiple dimensions of the candidate state, repeat the above steps to obtain the target Euclidean distance from the candidate state of the next dimension to the historical state set, until the target Euclidean distance from the candidate state of each dimension to the historical state set is determined.
[0071] In specific implementation, the candidate state s candidate is a d-dimensional continuous vector s=(s1,s2,…,s N ). For example, from the candidate state s candidate The current dimension determined from the multiple dimensions of is the i-th dimension. From the historical state set H = {s1′, s2′, …, s NThe multiple historical states determined in {s1′, s2′, …, s N ′. Calculate the candidate state s for the i-th dimension respectively i The multiple current Euclidean distances between and each historical state. Perform a mean process on the multiple current Euclidean distances to obtain the target Euclidean distance d from the candidate state of the i-th dimension to the historical state set i . Determine the next dimension as the (i + 1)-th dimension from the multiple dimensions of the candidate state s candidate , repeat steps 1031C to 1031D, and calculate the target Euclidean distance d from the candidate state of the (i + 1)-th dimension to the historical state set i+1 , until the target Euclidean distances from the candidate states of each dimension to the historical state set are calculated.
[0072] Through the above solution, according to the hierarchical data structure, the target Euclidean distances from the candidate states of each dimension to the historical state set are determined, which can reduce the distance calculation complexity of high-dimensional states from linear level to logarithmic level, and significantly improve the query efficiency of large-scale state sets. In addition, by calculating the target Euclidean distances from the candidate states of each dimension to the historical state set, the kernel function values of the candidate states of each dimension can be calculated respectively, so as to calculate the kernel function sum based on the kernel function values of the candidate states of all dimensions.
[0073] In some embodiments, step 1032 includes:
[0074] Step 1032A, obtain the kernel function value of the candidate state of each dimension according to the target Euclidean distance and the kernel bandwidth parameter,
[0075]
[0076] where K(d i ) is the kernel function value of the candidate state of the i-th dimension, d i is the target Euclidean distance from the candidate state of the i-th dimension to the historical state set, δ is the kernel bandwidth parameter, and exp(·) is the natural exponential function.
[0077] Step 1032B, perform a summation process on the kernel function values of the candidate states of all dimensions to obtain the kernel function sum,
[0078]
[0079] where S is the kernel function sum and n is the total number of dimensions of the candidate state.
[0080] In specific implementation, through the kernel density estimation algorithm, based on the target Euclidean distance d i determine the kernel function value K(d i)。After calculating the kernel function values of all dimensions of the candidate state, sum up the kernel function values of the candidate states of all dimensions to obtain the kernel function sum S.
[0081] Through the above solution, based on the target Euclidean distance, determine the kernel function values of the candidate states of each dimension, and sum up the kernel function values of the candidate states of all dimensions to obtain the kernel function sum, which can convert the discrete state distance into a continuous probability density and overcome the information loss caused by grid discretization.
[0082] In some embodiments, step 1033 includes:
[0083] Step 1033A, obtain the coverage rate of the candidate state according to the kernel function sum,
[0084] Coverage(s) = 1 - exp(-αS)
[0085] where Coverage(s) is the coverage rate of the candidate state, α is the attenuation coefficient, and S is the kernel function sum.
[0086] In specific implementation, calculate the coverage rate Coverage(s) of the candidate state according to the kernel function sum S. Specifically, according to the kernel function sum S and the attenuation coefficient α, map the kernel function sum S to the interval [0, 1) through an exponential function to calculate the coverage rate Coverage(s) of the candidate state.
[0087] Through the above solution, according to the kernel function sum and the attenuation coefficient, map the kernel function sum to the interval [0, 1) through an exponential function to calculate the coverage rate of the candidate state. In this way, the coverage rate growth rate can be adaptively adjusted according to the attenuation coefficient.
[0088] In some embodiments, step 104 includes:
[0089] Step 1041, compare the candidate coverage rate with the current coverage rate of the current state.
[0090] Step 1042, in response to determining that the current coverage rate is greater than the candidate coverage rate, use the current state as the optimal test case.
[0091] Step 1043, in response to determining that the candidate coverage rate is greater than the current coverage rate, use the candidate state as the optimal test case.
[0092] In specific implementation, when the current coverage rate Coverage(0) is greater than the candidate coverage rate Coverage(candidate), it means that the coverage rate of the current state is large, and keep the current state s0 as the optimal test case s bestThe candidate coverage Coverage(candidate) is greater than the current coverage Coverage(0), indicating that the coverage of the candidate state is large. Keep the candidate state s candidate as the optimal test case s best .
[0093] Through the above solution, by comparing the current coverage and the candidate coverage, it is possible to accurately determine whether the current optimal state needs to be updated, and thus use the state with the larger coverage in the candidate state and the current state as the optimal test case. Based on the optimal test case, it is possible to better measure the exploration of the state space by the agent.
[0094] Through the above embodiments, the initial state and the set of historical states of the agent in the deep reinforcement learning model are obtained. The candidate state is obtained from the initial state according to the Gaussian noise algorithm, and the hierarchical data structure is constructed using the set of historical states. The target Euclidean distance from the candidate state to the set of historical states is determined according to the hierarchical data structure, and the candidate coverage of the candidate state is determined based on the target Euclidean distance. It is possible to accurately calculate the candidate coverage of the candidate state in the high-dimensional continuous space, and there is no need to discretize the set of historical states, avoiding the coverage error caused by discretization, and realizing more accurate quantification of the coverage of the continuous state space. Based on the candidate coverage and the current coverage of the current state, the optimal test case is determined from the candidate state and the current state. Based on the optimal test case, it is possible to better measure the exploration of the state space by the agent.
[0095] It should be noted that the embodiments of the present disclosure can also be further described in the following manner:
[0096] (1) KSCoverage metric calculation process
[0097] Algorithm 1: KSCoverage metric calculation algorithm
[0098] Input: The target state is a d-dimensional continuous vector s = (s1, s2,..., s N ). The set of historical states H = {s1′, s2′,..., s N ′}, which records all the states visited by the agent in one episode, where each s i ′ is a d-dimensional continuous vector. σ is the kernel bandwidth parameter that controls the neighborhood range of the Gaussian kernel function. α is the decay coefficient that controls the growth rate of the coverage.
[0099] Output: The coverage Coverage(s) of the target state s, and the value range is [0, 1).
[0100] Step 1, construct the HNSW data structure
[0101] Construct the HNSW data structure T using the historical state set H for fast indexing of states. Here, n is the number of historical states.
[0102] Step 2: Calculate the distance
[0103] Use the constructed HNSW data structure T to find the Euclidean distance d from the target state s to several neighboring states in the historical state set H i = ‖s i - s i ′‖, where i = 1, 2, 3, …, n.
[0104] Step 3: Calculate the kernel function value
[0105] For each calculated Euclidean distance d i , use the Gaussian kernel function value
[0106] Step 4: Calculate the sum of kernel functions
[0107] Sum all the calculated kernel function values K(d i ) to obtain the sum of kernel function values
[0108] Step 5: Calculate the coverage rate
[0109] According to the sum of kernel function values S and the decay coefficient α, use the exponential function to map it to the interval [0, 1), and calculate the coverage rate Coverage(s) = 1 - exp(-αS) of the target state s.
[0110] Specifically, through the efficient search in the multi-dimensional space and kernel density estimation, this algorithm realizes the accurate quantification of the coverage rate of the continuous state space. The core process is divided into five steps: First, use the HNSW data structure to perform spatial indexing on the historical state set, reducing the distance calculation complexity of high-dimensional states from the linear level O(N) to the logarithmic level O(logN), significantly improving the query efficiency of large-scale state sets; Second, through the Gaussian kernel function transform the discrete state distance into a continuous probability density, overcoming the information loss caused by traditional grid discretization; Finally, through the exponential mapping Coverage(s) = 1 - exp(-αS), dynamically map the sum of kernel function values to the interval [0, 1), where the α parameter can adaptively adjust the coverage rate growth rate.
[0111] After determining the coverage rate index, the embodiments of the present disclosure can use test case generation methods, such as test case generation methods based on adversarial attacks (such as FGSSM, CW, etc.), heuristic-based test case generation methods, fuzz testing based on Gaussian noise, etc., and use the proposed KSCoverage as the guiding target to generate test cases.
[0112] (2) DRL Coverage Testing Framework Based on KSCoverage
[0113] Algorithm 2: DRL Coverage Testing Framework Based on KSCoverage
[0114] Input: Set of initial states, maximum number of iterations T, perturbation step size, kernel bandwidth parameter, decay coefficient. List of adversarial attack methods (such as FGSM, CW, etc.).
[0115] Output: Generated test case set Result
[0116] Step 1, Initialization
[0117] Result set
[0118] Historical state set H = H0
[0119] Current best coverage best_coverage = 0
[0120] Step 2, Iterative optimization
[0121] For t = 1 to T:
[0122] Construct the HNSW data structure: Use the current H to construct the HNSW data structure T
[0123] Generate candidate test cases using the Gaussian noise method:
[0124] ① Randomly sample an initial state s init
[0125] ② Generate noise: s noise = s init + N(0, σ 2 )
[0126] ③ Add s noise to the candidate set Candidates
[0127] Evaluate candidate test cases:
[0128] For each s candidate ∈ Candidates, call Algorithm 1 to calculate its coverage Coverage(s candidate )
[0129] Select the optimal test case:
[0130] Select the optimal state s best = argmax sCoverage(s), and select the candidate state with the maximum coverage as the optimal state. If the coverage of the optimal state is greater than the current best coverage Coverage(s best ) > best_coverage, then update Coverage(s best ) as the optimal test case, and add the optimal state s best to the result set Result.
[0131] Specifically, taking Gaussian fuzz testing as an example, the coverage optimization framework based on directional exploration uses KSCoverage as the goal to guide the generation of test cases. Its iterative process includes the following stages:
[0132] In each iteration, first generate Gaussian noise to perturb the candidate state s noise = s init + N(0, σ 2 ), and control the exploration range through the kernel bandwidth parameter σ; then use the greedy strategy to select the candidate state with the highest coverage and add it to the result set, and dynamically update the HNSW data structure to accelerate subsequent calculations. The algorithm introduces a decay coefficient α to adaptively adjust the coverage growth rate with the exploration process, providing an extensible coverage optimization framework for the reinforcement learning system in high-dimensional continuous space.
[0133] Through the above embodiments, for the coverage test index in deep reinforcement learning, the design of the coverage index of deep reinforcement learning based on kernel density estimation (KSCoverage) is to accurately process the state calculation in the continuous state space, without the need for state discretization operations, reducing the errors caused by discretization and achieving more accurate quantification of the continuous state space coverage. In terms of state space indexing and calculation method, an efficient approximate nearest neighbor search algorithm HNSW based on a multi-layer graph structure is used. The core idea is to quickly narrow the search range and accurately locate the target through hierarchical navigation, constructing a multi-layer graph from sparse to dense. The top layer is used for fast navigation, and the bottom layer is used for accurate search. Nodes are connected through a network to ensure that the shortest path between any two points is short. When querying, start from the top layer and gradually move down to the lower layer, using greedy search to approximate the target point, and finally find the nearest neighbor in the bottom layer. The HNSW data structure has the advantages of high efficiency, high precision and dynamic update ability, and is particularly suitable for real-time application scenarios of high-dimensional data and large-scale data sets, and the index can reach the time complexity of O(logN). The embodiments of the present disclosure use a coverage test framework guided by KSCoverage, which improves the success rate of test case generation, and the generated test cases can better measure the exploration of the state space by the agent.
[0134] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0135] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0136] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a coverage testing device for deep reinforcement learning.
[0137] Reference Figure 4 , the coverage testing device for deep reinforcement learning includes:
[0138] An acquisition module 301, configured to acquire the initial state and the set of historical states of the agent in the deep reinforcement learning model;
[0139] A candidate state determination module 302, configured to obtain candidate states according to the initial state by a Gaussian noise algorithm and construct a hierarchical data structure by using the set of historical states;
[0140] A candidate coverage determination module 303, configured to determine the target Euclidean distance from the candidate states to the set of historical states according to the hierarchical data structure and determine the candidate coverage of the candidate states based on the target Euclidean distance;
[0141] An optimal test case determination module 304, configured to determine an optimal test case from the candidate states and the current state based on the candidate coverage and the current coverage of the current state.
[0142] In some embodiments, the candidate state determination module 302 includes:
[0143] A noise state determination unit, configured to add noise with a probability density function obeying a Gaussian distribution to the initial state by a Gaussian noise algorithm to obtain a noise state;
[0144] A candidate status determination unit, configured to add the noise status to a candidate set and select a candidate status from the candidate set.
[0145] In some embodiments, the candidate coverage determination module 303 includes:
[0146] A target Euclidean distance determination unit, configured to determine, according to the hierarchical data structure, the target Euclidean distance from the candidate status of each dimension to the historical status set;
[0147] A kernel function sum determination unit, configured to determine the kernel function value of the candidate status of each dimension based on the target Euclidean distance, and perform a summation process on the kernel function values of the candidate statuses of all dimensions to obtain a kernel function sum;
[0148] A coverage determination unit, configured to obtain the coverage of the candidate status according to the kernel function sum.
[0149] In some embodiments, the target Euclidean distance determination unit includes:
[0150] A current dimension determination subunit, configured to determine a current dimension from multiple dimensions of the candidate status;
[0151] A historical status determination subunit, configured to determine multiple historical statuses from the historical status set;
[0152] A current Euclidean distance determination subunit, configured to determine, according to the hierarchical data structure, multiple current Euclidean distances between the candidate status of the current dimension and each historical status;
[0153] A first target Euclidean distance determination subunit, configured to perform a mean process on the multiple current Euclidean distances to obtain the target Euclidean distance from the candidate status of the current dimension to the historical status set;
[0154] A second target Euclidean distance determination subunit, configured to determine a next dimension from multiple dimensions of the candidate status, and repeat the above steps to obtain the target Euclidean distance from the candidate status of the next dimension to the historical status set until the target Euclidean distance from the candidate status of each dimension to the historical status set is determined.
[0155] In some embodiments, the kernel function sum determination unit includes:
[0156] A kernel function value determination subunit, configured to obtain the kernel function value of the candidate status of each dimension according to the target Euclidean distance and a kernel bandwidth parameter,
[0157]
[0158] where K(di ) is the kernel function value of the candidate state for the i-th dimension, d i is the target Euclidean distance from the candidate state of the i-th dimension to the set of historical states, δ is the kernel bandwidth parameter, and exp(·) is the natural exponential function;
[0159] The kernel function and determination subunit is configured to sum up the kernel function values of the candidate states for all dimensions to obtain the kernel function sum,
[0160]
[0161] where S is the kernel function sum and n is the total number of dimensions of the candidate states.
[0162] In some embodiments, the coverage rate determination unit includes:
[0163] The coverage rate determination subunit is configured to obtain the coverage rate of the candidate state according to the kernel function sum,
[0164] Coverage(s) = 1 - exp(-αS)
[0165] where Coverage(s) is the coverage rate of the candidate state, α is the attenuation coefficient, and S is the kernel function sum.
[0166] In some embodiments, the optimal test case determination module 304 includes:
[0167] The comparison processing unit is configured to compare the candidate coverage rate with the current coverage rate of the current state;
[0168] The first optimal test case determination unit is configured to, in response to determining that the current coverage rate is greater than the candidate coverage rate, use the current state as the optimal test case;
[0169] The second optimal test case determination unit is configured to, in response to determining that the candidate coverage rate is greater than the current coverage rate, use the candidate state as the optimal test case.
[0170] For the convenience of description, when describing the above device, it is described by function as various modules respectively. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0171] The device in the above embodiments is used to implement the corresponding coverage rate test method of deep reinforcement learning in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0172] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the coverage rate testing method of deep reinforcement learning described in any one of the above embodiments.
[0173] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0174] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0175] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0176] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0177] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to enable communication and interaction between this device and other devices. The communication module can achieve communication through a wired method (such as USB (Universal Serial Bus), network cable, etc.) or through a wireless method (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0178] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0179] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0180] The electronic device in the above embodiment is used to implement the corresponding coverage rate test method of deep reinforcement learning in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0181] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the coverage rate test method of deep reinforcement learning as described in any of the foregoing embodiments.
[0182] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0183] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the coverage rate testing method of deep reinforcement learning as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0184] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a computer program product, including computer program instructions, which when running on a computer, cause the computer to execute the coverage rate testing method of deep reinforcement learning as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0185] It can be understood that before using the technical solutions of the various embodiments in the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0186] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0187] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0188] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0189] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the idea of the present disclosure, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0190] Additionally, for simplicity of explanation and discussion, and so as not to render the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.
[0191] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0192] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the present disclosure. Thus, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A coverage testing method for deep reinforcement learning, characterized in that, The method includes: Obtaining the initial state and the set of historical states of the agent in the deep reinforcement learning model; Obtaining candidate states according to the initial state through the Gaussian noise algorithm, and constructing a hierarchical data structure by using the set of historical states; Determining the target Euclidean distance from the candidate states to the set of historical states according to the hierarchical data structure, and determining the candidate coverage rate of the candidate states based on the target Euclidean distance; Determining the optimal test case from the candidate states and the current state based on the candidate coverage rate and the current coverage rate of the current state.
2. The method according to claim 1, wherein The obtaining candidate states according to the initial state through the Gaussian noise algorithm includes: Adding noise with a probability density function following a Gaussian distribution to the initial state through the Gaussian noise algorithm to obtain a noise state; Adding the noise state to the candidate set, and selecting candidate states from the candidate set.
3. The method according to claim 1, wherein The determining the target Euclidean distance from the candidate states to the set of historical states according to the hierarchical data structure, and determining the candidate coverage rate of the candidate states based on the target Euclidean distance includes: Determining the target Euclidean distance from the candidate states in each dimension to the set of historical states according to the hierarchical data structure; Determining the kernel function value of the candidate states in each dimension based on the target Euclidean distance, and performing a summation process on the kernel function values of the candidate states in all dimensions to obtain a kernel function sum; Obtaining the coverage rate of the candidate states according to the kernel function sum.
4. The method according to claim 3, wherein The determining the target Euclidean distance from the candidate states in each dimension to the set of historical states according to the hierarchical data structure includes: Determining the current dimension from multiple dimensions of the candidate states; Determining multiple historical states from the set of historical states; Determining multiple current Euclidean distances between the candidate states in the current dimension and each historical state according to the hierarchical data structure; Performing a mean process on the multiple current Euclidean distances to obtain the target Euclidean distance from the candidate states in the current dimension to the set of historical states; Determining the next dimension from multiple dimensions of the candidate states, and repeating the above steps to obtain the target Euclidean distance from the candidate states in the next dimension to the set of historical states until the target Euclidean distance from the candidate states in each dimension to the set of historical states is determined.
5. The method according to claim 1, wherein The determining the kernel function value of the candidate states in each dimension based on the target Euclidean distance, and performing a summation process on the kernel function values of the candidate states in all dimensions to obtain a kernel function sum includes: Obtaining the kernel function value of the candidate states in each dimension according to the target Euclidean distance and the kernel bandwidth parameter, where \(K(d i )\) is the kernel function value of the candidate state in the \(i\)-th dimension, \(d i \) is the target Euclidean distance from the candidate state in the \(i\)-th dimension to the set of historical states, \(\delta\) is the kernel bandwidth parameter, and \(\exp(\cdot)\) is the natural exponential function; Performing a summation process on the kernel function values of the candidate states in all dimensions to obtain a kernel function sum, where S is the kernel function sum and n is the total number of dimensions of the candidate states.
6. The method according to claim 1, characterized in that The obtaining the coverage rate of the candidate states according to the kernel function sum includes: Obtaining the coverage rate of the candidate states according to the kernel function sum, Coverage(s) = 1 - exp(-αS) where Coverage(s) is the coverage rate of the candidate states, α is the attenuation coefficient, and S is the kernel function sum.
7. The method according to claim 1, wherein Determining an optimal test case from the candidate state and the current state based on the candidate coverage rate and the current coverage rate of the current state includes: Performing a comparison process on the candidate coverage rate and the current coverage rate of the current state; In response to determining that the current coverage rate is greater than the candidate coverage rate, using the current state as the optimal test case; In response to determining that the candidate coverage rate is greater than the current coverage rate, using the candidate state as the optimal test case.
8. A coverage testing device for deep reinforcement learning, characterized in that, It includes: An acquisition module configured to acquire the initial state and the set of historical states of an agent in a deep reinforcement learning model; A candidate state determination module configured to obtain a candidate state according to the initial state by a Gaussian noise algorithm and construct a hierarchical data structure by using the set of historical states; A candidate coverage rate determination module configured to determine the target Euclidean distance from the candidate state to the set of historical states according to the hierarchical data structure and determine the candidate coverage rate of the candidate state based on the target Euclidean distance; An optimal test case determination module configured to determine an optimal test case from the candidate state and the current state based on the candidate coverage rate and the current coverage rate of the current state.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method described in any one of claims 1 to 7.