Software testing method and device, server, storage medium and program product

By dynamically adjusting the selection of coding rules using reinforcement learning models in software testing, the problems of limitations in coding rules selection and high false alarm rates in the prior art are solved, and more accurate problem positioning and more efficient static analysis of software units are achieved.

CN120066953APending Publication Date: 2025-05-30GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510032058.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There are limitations in the selection of coding rules in the prior art, and the false positive rate of static analysis results is high, resulting in inaccurate problem positioning.

Method used

By identifying the language characteristics and security levels of the source code of the target software, multiple candidate coding rules are selected, and input them and source code into the reinforcement learning model, output appropriate target coding rules, and perform static analysis.

Benefits of technology

Dynamically adjust the selection of coding rules, improve the pertinence and effectiveness of detection, reduce false positives in static analysis, improve the accuracy of problem positioning, and significantly improve the quality and efficiency of static analysis of software units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066953A_ABST
    Figure CN120066953A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of testing, in particular to a software testing method and device, a server, a storage medium and a program product.The method comprises the steps that language features and security levels of source codes of target software are recognized; selecting a plurality of candidate coding rules according to the language features and the security level, inputting the plurality of candidate coding rules and the source code into a reinforcement learning model, and outputting a target coding rule of the source code by the reinforcement learning model; and testing and analyzing the source code based on the target coding rule. Therefore, the problems of specific selection limitation of coding rules, relatively high false alarm rate of analysis results and the like in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of testing technologies, and particularly to a software testing method, apparatus, server, storage medium, and program product. Background Art

[0002] With the rapid development of information technology, the iteration cycle of software products has been significantly shortened, which poses higher requirements for software quality. Software testing is one of the key means to ensure software quality, and software unit static analysis runs through various testing stages of the entire software life cycle. Static analysis can detect and fix defects in the early stage of software development, thereby shortening the development cycle, reducing costs, and improving efficiency.

[0003] In related technologies, generally a set of software coding rules is selected, and the source code is analyzed line by line with the help of a static analysis tool to export the static analysis results. Although this static analysis method can effectively detect problems in the source code, it also has certain limitations. First, when selecting coding rules, since each set of coding standards has different focuses, various standards may have intersections: if a set of coding rules is selected, it may not be possible to comprehensively analyze the source code from multiple perspectives; if multiple sets of rules are selected, the analysis time will be extended, and there may be overlaps in the types of problems in the analysis results. Second, since traditional static analysis is mainly defined by experience and has a low degree of fit with specific code, the false positive rate of static analysis results is relatively high, and problem positioning is not accurate enough. Summary of the Invention

[0004] This application provides a software testing method, apparatus, server, storage medium, and program product to solve the specific limitations in the selection of coding rules in related technologies and the problems such as a relatively high false positive rate of analysis results.

[0005] The first aspect of the embodiments of this application provides a software testing method, including the following steps: identifying the language features and security levels of the source code of the target software; selecting multiple candidate coding rules according to the language features and security levels, inputting the multiple candidate coding rules and the source code into a reinforcement learning model, and the reinforcement learning model outputs the target coding rule of the source code; performing test analysis on the source code based on the target coding rule.

[0006] Optionally, the execution process of the reinforcement learning model includes: selecting a candidate coding rule from multiple candidate coding rules based on a target policy; performing static analysis on the source code based on the selected candidate coding rule to obtain an analysis result; determining the reward value of the selected candidate coding rule according to the analysis result; determining the confidence category of the candidate coding rule according to the reward value of the candidate coding rule, and selecting the candidate coding rule of the target confidence category as the target coding rule.

[0007] Optionally, determining the reward value of the selected candidate coding rule according to the analysis result includes: identifying the scores of all single coding rules in the analysis result; calculating the accuracy rate according to the scores of all single coding rules and the number of single coding rules; determining the reward value of the candidate coding rule according to the accuracy rate.

[0008] Optionally, the confidence categories include the first to the third categories, and the target confidence category is the first category, where the confidence of the first category is greater than the confidence of the second category, and the confidence of the second category is greater than the confidence of the third category.

[0009] Optionally, the target policy includes: selecting a random action according to ε, or selecting an action with the maximum reward value with a probability of 1 - ε, where the action is to perform static analysis on the source code based on the coding rule.

[0010] Optionally, the relationship between the action and ε is:

[0011]

[0012] where a t is the action taken at time t, s t is the state at time t, ε is the exploration probability, A is the set of all possible actions, and Q(s t , a t ) is the expected return of taking action a t in state s t .

[0013] An embodiment of the second aspect of this application provides a software testing device, including: an identification module, configured to identify the language features and security levels of the source code of the target software; an input / output module, configured to select multiple candidate coding rules according to the language features and security levels, input the multiple candidate coding rules and the source code into a reinforcement learning model, and the reinforcement learning model outputs the target coding rule of the source code; a testing module, configured to perform test analysis on the source code based on the target coding rule.

[0014] Optionally, the execution process of the reinforcement learning model includes: selecting a candidate coding rule from multiple candidate coding rules based on the target policy; performing static analysis on the source code based on the selected candidate coding rule to obtain an analysis result; determining the reward value of the selected candidate coding rule according to the analysis result; determining the confidence category of the candidate coding rule according to the reward value of the candidate coding rule, and selecting the candidate coding rule of the target confidence category as the target coding rule.

[0015] Optionally, the input / output module is further configured to identify the scores of all single coding rules in the analysis result; calculate the accuracy rate according to the scores of all single coding rules and the number of single coding rules; determine the reward value of the candidate coding rule according to the accuracy rate.

[0016] Optionally, the confidence categories include first to third categories, and the target confidence category is the first category, where the confidence of the first category is greater than that of the second category, and the confidence of the second category is greater than that of the third category.

[0017] Optionally, the target policy includes: selecting a random action according to ε, or selecting an action with the maximum reward value with a probability of 1 - ε, where the action is to perform static analysis on the source code based on the encoding rule.

[0018] Optionally, the relationship between the action and ε is:

[0019]

[0020] where a t is the action taken at time t, s t is the state at time t, ε is the exploration probability, A is the set of all possible actions, and Q(s t , a t ) is the expected return of taking action a t in state s t .

[0021] An embodiment of the third aspect of the present application provides a server, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the software testing method as in the above embodiment.

[0022] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed, it is used to implement the software testing method as in the above embodiment.

[0023] An embodiment of the fifth aspect of the present application provides a computer program product, including: a computer program or instruction, and when the computer program or instruction is executed, it is used to implement the software testing method as in the above embodiment.

[0024] Therefore, the present application has at least the following beneficial effects:

[0025] Embodiments of the present application can dynamically adjust the selection of encoding rules based on the characteristics and requirements of the source code, thereby improving the pertinence and effectiveness of detection, and adaptively outputting a target encoding rule suitable for the source code through a reinforcement learning model, which can gradually reduce false positives in static analysis and improve the accuracy of problem location, significantly improving the quality and efficiency of static analysis of software units. Thus, the limitations in the selection of encoding rules in the related art are solved, and problems such as a high false positive rate of analysis results are also solved.

[0026] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings

[0027] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, where:

[0028] Figure 1 It is a flowchart of a software testing method provided according to an embodiment of the present application;

[0029] Figure 2 It is an exemplary diagram of the execution process of reinforcement learning provided according to an embodiment of the present application;

[0030] Figure 3 It is an exemplary diagram of the update process of a status file provided according to an embodiment of the present application;

[0031] Figure 4 It is a block diagram of a software testing apparatus according to an embodiment of the present application;

[0032] Figure 5 It is a schematic structural diagram of a server according to an embodiment of the present application. Detailed Embodiments

[0033] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0034] The software testing method, apparatus, server, storage medium, and program product of embodiments of the present application will be described below with reference to the drawings. In view of the problems mentioned in the above background art, the present application provides a software testing method. In this method, based on the characteristics and requirements of the source code, the selection of coding rules is dynamically adjusted, thereby improving the pertinence and effectiveness of detection. Moreover, through the reinforcement learning model, the target coding rules suitable for the source code are adaptively output, which can gradually reduce false positives in static analysis and improve the accuracy of problem location, significantly improving the quality and efficiency of software unit static analysis. Thereby, the limitations in the selection of coding rules in the related art and the problem of a relatively high false positive rate of analysis results are solved.

[0035] Specifically, Figure 1 It is a schematic flowchart of a software testing method provided according to an embodiment of the present application.

[0036] As Figure 1 shown, the software testing method includes the following steps:

[0037] In step S101, the language features and security level of the source code of the target software are identified.

[0038] In the embodiments of the present application, the programming language used in the source code is analyzed, such as C, C++, Python, or Simulink models, etc. The embodiments of the present application can understand its syntax structure, common programming patterns, and potential security risks according to the characteristics of the programming language. Secondly, the embodiments of the present application can determine the application scenario and security requirements of the software, such as functional safety levels A, B, C, or D. According to the different security levels, suitable coding rules and security standards are selected.

[0039] In step S102, multiple candidate coding rules are selected according to the language features and security level, and the multiple candidate coding rules and the source code are input into the reinforcement learning model, and the reinforcement learning model outputs the target coding rule of the source code.

[0040] It can be understood that the embodiments of the present application can select appropriate coding rules according to the language characteristics and security level requirements of the source code. The selected coding rules are numbered and recorded in the IntR.c file, where IntR.c is the abbreviation of Intital rule.c and the file name can be defined by oneself. In the actual execution process, the selected coding rules include MISRA (focusing on functionality, stability, and maintainability), CERT (focusing on code style standardization), CWE (focusing on design and architecture defects and low-level coding and design errors), etc.

[0041] In the embodiments of the present application, the reinforcement learning model can be constructed by using the Markov decision process. First, the reinforcement learning algorithm is combined with the background of static analysis of software units to define state, action, policy, and reward elements, thereby constructing a reinforcement learning agent. Among them, the reward set of each coding rule is defined as state S, and all states are recorded in the Sta.c file. The recorded content is shown in Table 1.

[0042] Table 1

[0043]

[0044] The process of analyzing the source code using software coding rules is defined as action A, and whether a problem is found in the static analysis result output after executing action A is used as the reward value R, where -1 represents a problem that does not belong to a violation of the coding rule, that is, a false alarm, 0 represents that the coding rule does not find a problem, and 1 represents that the problem found in the code belongs to a violation of the coding rule problem.

[0045] In one embodiment of the present application, the execution process of the reinforcement learning model includes: selecting a candidate coding rule from multiple candidate coding rules based on a target policy; performing static analysis on the source code based on the selected candidate coding rule to obtain an analysis result; determining the reward value of the selected candidate coding rule according to the analysis result; determining the confidence category of the candidate coding rule according to the reward value of the candidate coding rule, and selecting the candidate coding rule of the target confidence category as the target coding rule.

[0046] Specifically, the embodiment of the present application can select a coding rule from the IntR.c file of the coding rule library by using the ε-greedy strategy to perform static analysis on the source code. Output the static analysis result to the tester for judging the correctness of the static analysis result and giving a reward value R, and record all the reward values in the Sta.c status file. In the actual execution process, the embodiment of the present application can identify the scores of all single coding rules in the analysis result; calculate the correct rate according to the scores of all single coding rules and the number of single coding rules; determine the reward value of the candidate coding rule according to the correct rate, and the calculation method of the correct rate is: the sum of all scores of a single coding rule is 1 and / the number of static analysis results ((1 + 1 + … + 1) / n).

[0047] In the embodiment of the present application, the confidence categories include the first to the third categories, and the target confidence category is the first category, where the confidence of the first category is greater than the confidence of the second category, and the confidence of the second category is greater than the confidence of the third category.

[0048] Among them, the first category can be a trustworthy category, and the result correct rate of the coding rules classified as the trustworthy category can be 90 - 100%; the second category can be a to-be-determined category, and the result correct rate of the coding rules classified as the to-be-determined category can be 20 - 90%. The third category can be an eliminated category, and the result correct rate of the coding rules classified as the to-be-determined category is less than 20%.

[0049] The embodiment of the present application can label and divide the confidence categories of the selected coding rules according to the correct rate calculation method and the annotation principle, update the reinforcement learning agent according to the Bellman equation, and repeat this step to make the reinforcement learning agent tend to generate a model that is easy to find coding rules with relatively high confidence. The learning process of the reinforcement learning is shown in the appendix Figure 2 After the learning agent has learned all the coding rules, output the coding rules with the confidence categories in the coding rule library to the coding rule library and label them as trustworthy rules, which are the finally selected coding rules. Thus, by using the reinforcement learning agent to perform an adaptive selection strategy to select a set of static analysis coding rules with higher code fitting degree and wider detection range, the software static analysis coverage and problem location accuracy can be effectively improved, and problems such as problem false alarm rate can be reduced.

[0050] It should be noted that after the coding rules are applied and feedback is generated, the coding rule library and the status file need to be updated. Among them, updating the coding rule library: adjust the confidence category according to the accuracy rate of the rules. Newly discovered valid rules are added to the rule library, while those marked as eliminated are removed from the rule library. Updating the status file: record information such as the execution history and reward value of each rule. This helps the subsequent decision-making process. The update process of the status file is as Figure 3 shown.

[0051] When updating the reinforcement learning agent, the embodiment of the present application can adopt the Bellman equation Q(s,a)←R+γmaxa'Q(s',a'), where the discount factor γ(0≤γ≤1) is used to control the degree of emphasis on future rewards.

[0052] In the embodiment of the present application, the target policy can be the ε-greedy policy. The embodiment of the present application can use the ε-greedy policy and the mechanism of maximizing rewards in reinforcement learning to optimize the static analysis optimization process and reduce the false positive rate of static analysis results, including: selecting a random action according to ε, or selecting an action with the maximum reward value with a probability of 1-ε, so as to balance the relationship between environment exploration and policy utilization in the environment learning process, where the action is to perform static analysis on the source code based on the coding rules. The relationship between the action and ε is:

[0053]

[0054] where a t is the action taken at time t, s t is the state at time t, ε is the exploration probability, A is the set of all possible actions, and Q(s t ,a t ) is the expected return of taking action a t in state s t .

[0055] In the early stage of training, the embodiment of the present application can set a random probability value, and select the action guided by the agent with a lower probability. As the training progresses, the knowledge of exploring the environment gradually accumulates, and the ε value continuously decays, and the action with the maximum behavior value is selected with a larger probability to utilize the learned knowledge.

[0056] In step S103, test analysis is performed on the source code based on the target coding rules.

[0057] It can be understood that the embodiment of the present application can use the reinforcement learning model to adaptively select coding rules to perform static analysis on the source code, which can more comprehensively cover different coding standards, thereby increasing the depth and breadth of static analysis.

[0058] The software testing method proposed according to the embodiments of the present application dynamically adjusts the selection of coding rules based on the characteristics and requirements of the source code, thereby improving the pertinence and effectiveness of detection. Moreover, by adaptively outputting the target coding rules suitable for the source code through the reinforcement learning model, it can gradually reduce the false positives in static analysis and improve the accuracy of problem location, significantly improving the quality and efficiency of static analysis of software units. Thus, the limitations in the selection of coding rules in the related art are solved, and problems such as a high false positive rate of the analysis results are addressed.

[0059] Next, a software testing device proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0060] Figure 4 It is a block diagram of the software testing device according to the embodiments of the present application.

[0061] As Figure 4 shown, the software testing device 10 includes: an identification module 100, an input / output module 200, and a testing module 300.

[0062] Among them, the identification module 100 is used to identify the language features and security levels of the source code of the target software; the input / output module 200 is used to select multiple candidate coding rules according to the language features and security levels, input the multiple candidate coding rules and the source code into the reinforcement learning model, and the reinforcement learning model outputs the target coding rules of the source code; the testing module 300 is used to perform test analysis on the source code based on the target coding rules.

[0063] In an embodiment of the present application, the execution process of the reinforcement learning model includes: selecting a candidate coding rule from multiple candidate coding rules based on the target policy; performing static analysis on the source code based on the selected candidate coding rule to obtain an analysis result; determining the reward value of the selected candidate coding rule according to the analysis result; determining the confidence category of the selected candidate coding rule according to the reward value of the candidate coding rule, and selecting the candidate coding rule of the target confidence category as the target coding rule.

[0064] In an embodiment of the present application, the input / output module 200 is further used to identify the scores of all single coding rules in the analysis result; calculate the correct rate according to the scores of all single coding rules and the number of single coding rules; determine the reward value of the candidate coding rule according to the correct rate.

[0065] In an embodiment of the present application, the confidence categories include the first to the third categories, and the target confidence category is the first category, where the confidence of the first category is greater than the confidence of the second category, and the confidence of the second category is greater than the confidence of the third category.

[0066] In one embodiment of the present application, the target policy includes: selecting a random action according to ε, or selecting an action with the maximum reward value with a probability of 1 - ε, where the action is to perform static analysis on the source code based on the encoding rule.

[0067] In one embodiment of the present application, the relationship between the action and ε is as follows:

[0068]

[0069] where a t is the action taken at time t, s t is the state at time t, ε is the exploration probability, A is the set of all possible actions, and Q(s t , a t ) is the expected return of taking action a t in state s t .

[0070] It should be noted that the foregoing explanation of the software testing method embodiment also applies to the software testing device of this embodiment, and will not be elaborated here.

[0071] According to the software testing device provided by the embodiments of the present application, based on the characteristics and requirements of the source code, the selection of the encoding rule is dynamically adjusted, thereby improving the pertinence and effectiveness of detection. Moreover, by adaptively outputting the target encoding rule suitable for the source code through the reinforcement learning model, the false positives in static analysis can be gradually reduced, and the accuracy of problem location can be improved, significantly enhancing the quality and efficiency of software unit static analysis. Thus, the limitations in the selection of encoding rules in the related art and the problems such as a relatively high false positive rate of the analysis results are solved.

[0072] Figure 5 The following is a schematic structural diagram of the server provided by the embodiments of the present application. The server may include:

[0073] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0074] When the processor 502 executes the program, it implements the software testing method provided in the foregoing embodiment.

[0075] Further, the server further includes:

[0076] A communication interface 503 for communication between the memory 501 and the processor 502.

[0077] The memory 501 is used to store a computer program executable on the processor 502.

[0078] The memory 501 may include a high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk memory.

[0079] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0080] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0081] The processor 502 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0082] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above software testing method is implemented.

[0083] The embodiments of the present application also provide a computer program product, including: a computer program or instruction, and when the computer program or instruction is executed, the software testing method as described in the above embodiments is implemented.

[0084] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0085] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0086] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0087] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0088] Those of ordinary skill in the technical field of this application can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0089] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A software testing method, characterized in that: The following steps are involved: Identify the language characteristics and security level of the target software's source code; Selecting a plurality of candidate coding rules according to the language features and the security level, inputting the plurality of candidate coding rules and the source code into a reinforcement learning model, and the reinforcement learning model outputting a target coding rule for the source code; The source code is tested and analyzed based on the target coding rules.

2. The software testing method according to claim 1, characterized in that: The execution process of the reinforcement learning model includes: Selecting a candidate encoding rule from the plurality of candidate encoding rules based on a target strategy; Performing static analysis on the source code based on the selected candidate coding rule to obtain an analysis result; Determining a reward value of the selected candidate coding rule according to the analysis result; The confidence category of the candidate coding rule is determined according to the reward value of the candidate coding rule, and the candidate coding rule of the target confidence category is selected as the target coding rule.

3. The software testing method according to claim 2, characterized in that: Determining the reward value of the selected candidate coding rule according to the analysis result includes: Identifying scores for all individual coding rules in the analysis results; The accuracy rate was calculated based on the scores of all single coding rules and the number of single coding rules; The reward value of the candidate encoding rule is determined according to the accuracy rate.

4. The software testing method according to claim 2, characterized in that: The confidence categories include first to third categories, and the target confidence category is the first category, wherein the confidence of the first category is greater than the confidence of the second category, and the confidence of the second category is greater than the confidence of the third category.

5. The software testing method according to claim 2, characterized in that: The target strategy includes: selecting a random action according to ε, or selecting an action with the maximum reward value according to probability 1-ε, wherein the action is to perform static analysis on the source code based on the coding rules.

6. The software testing method according to claim 5, characterized in that: The relationship between the action and ε is: Among them, a t is the action taken at time t, s t is the state at time t, ε is the exploration probability, A is the set of all possible actions, Q(s t ,a t ) is in state s t Take action a t expected return.

7. A software testing device, characterized in that: include: An identification module, used to identify the language characteristics and security level of the source code of the target software; An input-output module, configured to select a plurality of candidate coding rules according to the language features and the security level, input the plurality of candidate coding rules and the source code into a reinforcement learning model, and the reinforcement learning model outputs a target coding rule for the source code; A testing module is used to test and analyze the source code based on the target coding rule.

8. A server, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the software testing method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed, the software testing method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed, it implements the software testing method according to any one of claims 1 to 6.