Network testing method and device, computer equipment and storage medium
By combining basic operators, a random network is generated for quantitative compilation and testing, and using the central processor to calculate the truth value results to determine the abnormal network, the problem that traditional coverage testing methods cannot effectively test multiple operator optimization and abnormal networks, and the ability to evaluate the accuracy of the model and discover problems is realized.
Patent Information
- Application Number
- CN202510339477.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional coverage testing methods cannot effectively test whether the optimization of quantization and compilation of multiple operators is correct, and cannot verify the operational accuracy of the abnormal network structure, which affects the test progress.
By obtaining multiple basic operators supported by quantitative compilation, combining them into composite operators, and generating a random network based on the composite operator and basic operators, performing quantitative compilation and testing, and using the central processor to calculate the truth value results to determine the abnormal network.
It realizes the accuracy evaluation of the model during quantitative compilation and hardware operation, and can promptly discover and solve problems that arise during network construction, quantitative compilation or hardware operation, and improves the reliability and stability of the model.
Smart Images

Figure CN120186048A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of testing technologies, and particularly to a network testing method, apparatus, computer device, and storage medium. Background Art
[0002] With the continuous development of deep learning, more and more scenarios need to run on hardware, thus requiring quantization compilation. During the quantization compilation process, coverage testing is needed.
[0003] Traditional coverage testing methods can only verify whether a single operator can perform operations correctly, and cannot test whether the optimization (folding, fusion, etc.) of multiple operators by quantization and compilation is correct. Moreover, due to the possible existence of abnormal network structures, it is impossible to verify the operation correctness of some abnormal network structures, and it is not clear whether the test fails or is caused by the abnormal network structure, which will affect the test progress. Summary of the Invention
[0004] Based on this, it is necessary to provide a network testing method, apparatus, computer device, and storage medium for the above technical problems.
[0005] In a first aspect, the present disclosure provides a network testing method. The method includes:
[0006] Obtain a plurality of basic operators supported by quantization compilation, combine the plurality of basic operators to obtain at least one composite operator;
[0007] Based on the at least one composite operator and the basic operators, and according to a preset input tensor and a preset random network depth, generate at least one random network;
[0008] Perform quantization compilation on the at least one random network, and test the at least one quantized and compiled random network to obtain a test result;
[0009] Determine the true value result of the at least one random network, and determine an abnormal network based on the true value result and the test result, where the true value result is obtained by processing the at least one random network using a central processing unit.
[0010] In one embodiment, the combining the plurality of basic operators to obtain at least one composite operator includes:
[0011] Obtain the function information and parameter information of each basic operator;
[0012] Adjust the parameter information of each basic operator, and combine the parameter information and function information of at least two adjusted basic operators to obtain at least one composite operator
[0013] In one embodiment, generating at least one random network based on the at least one composite operator and the base operator, and according to a preset input tensor and a preset random network depth includes:
[0014] Select at least one initial operator from the at least one composite operator and the base operator, and update the sampling space of the at least one initial operator according to the preset input tensor;
[0015] Sample in the updated sampling space of the at least one initial operator to obtain at least one first operator;
[0016] Construct a random network based on the at least one first operator and the preset random network depth.
[0017] In one embodiment, constructing a random network based on the at least one first operator and the preset random network depth includes:
[0018] Construct an initial network based on the at least one first operator;
[0019] In response to the network depth of the initial network meeting the preset random network depth, determine the initial network as the random network;
[0020] In response to the network depth of the initial network not meeting the preset random network depth, select at least one successor operator from the at least one composite operator and the base operator;
[0021] Determine the predecessor operators of the at least one successor operator from the initial network, and update the sampling space of the at least one successor operator based on the output tensor of the predecessor operators;
[0022] Sample in the updated sampling space of the at least one successor operator, and add the sampled subsequent operators after the predecessor operators in the initial network until the initial network obtained after the addition meets the preset random network depth, to obtain a random network.
[0023] In one embodiment, when a successor operator requires multiple predecessor operators, determining the predecessor operators of the at least one successor operator from the initial network includes:
[0024] Determine the operators in the initial network with the same number of output channels in the output tensor as the predecessor operators of the at least one successor operator.
[0025] In one embodiment, after determining the abnormal network based on the true value result and the test result, the method further includes:
[0026] Obtain a target test network, perform quantization compilation on the target test network and then conduct on-board testing, and in response to the test result of the on-board testing being abnormal;
[0027] Determine the network structure of the target test network, and in response to the network structure of the target test network being the same as the network structure of the abnormal network, determine that there is an error in the network structure of the target test network.
[0028] In a second aspect, the present disclosure also provides a network testing device. The device includes:
[0029] An operator combination module, configured to obtain a plurality of basic operators supported by quantization compilation, combine the plurality of basic operators to obtain at least one composite operator;
[0030] A random network generation module, configured to generate at least one random network based on the at least one composite operator and the basic operators, and according to a preset input tensor and a preset random network depth;
[0031] A testing module, configured to perform quantization compilation on the at least one random network, and conduct on-board testing on the at least one quantized and compiled random network to obtain a test result;
[0032] An abnormal network determination module, configured to determine the true value result of the at least one random network, and determine an abnormal network based on the true value result and the test result, where the true value result is obtained by processing the at least one random network using a central processing unit.
[0033] In a third aspect, the present disclosure also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0034] In a fourth aspect, the present disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0035] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0036] In the above embodiments, by combining multiple basic operators into composite operators, the data transmission between operators and the storage of intermediate results are reduced, and the overhead during the calculation process is lowered. Generating a random network based on the composite operators and basic operators can explore different combinations of network structures. Calculating the true value results of the random network using a central processing unit (CPU) and comparing them with the test results can accurately evaluate the accuracy of the model during quantization compilation and hardware operation. Determining abnormal networks based on the true value results and test results helps to promptly discover and solve problems that occur during network construction, quantization compilation, or hardware operation. Conducting in-depth analysis of abnormal networks can identify the root causes of problems, such as unreasonable quantization parameter settings or inappropriate compilation optimization strategies, and make corresponding adjustments and improvements to enhance the reliability and stability of the model. Quantizing and compiling multiple random networks and conducting tests can comprehensively verify the performance and compatibility of the hardware and quantization compilation tools. Discovering and resolving problems at an early stage of development can avoid issues that occur during large-scale testing in the later stage, reduce repeated modifications and debugging during the development process, and improve development efficiency. The diversity of random networks enables testing to cover more network structures and calculation scenarios, thereby more comprehensively verifying the performance of the hardware and the model. This helps to discover potential problems and vulnerabilities and improve the quality and stability of the product. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 Schematic diagram of the application environment of the network testing method in an embodiment;
[0039] Figure 2 Schematic diagram of the flowchart of the network testing method in an embodiment;
[0040] Figure 3 Schematic diagram of the flowchart of step S202 in an embodiment;
[0041] Figure 4 Schematic diagram of the fusion of composite operators in an embodiment;
[0042] Figure 5 Schematic diagram of the flowchart of step S204 in an embodiment;
[0043] Figure 6 Schematic diagram of the flowchart of step S406 in an embodiment;
[0044] Figure 7 It is a schematic flow diagram after step S208 in an embodiment;
[0045] Figure 8 It is a schematic block diagram of the structure of a network testing device in an embodiment;
[0046] Figure 9 It is a schematic internal structure diagram of a computer device in an embodiment;
[0047] Figure 10 It is a schematic internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0048] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0049] It should be noted that the terms "first", "second", etc. in the description and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0050] In this article, the term "and / or" is only a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0051] As described in the background art, the NPU (Neural Processing Unit) is a computing chip specifically designed to accelerate neural network and deep learning algorithms. In the fields of image recognition, speech recognition, natural language processing, etc., the NPU can provide fast data processing and analysis capabilities. Currently, in end-side applications (such as autonomous driving, embodied intelligence, etc.), in order to reduce the inference latency and power consumption of neural networks, customized computational graph optimization technologies such as quantization and compilation can be deployed at the front end of the chip. These software-level technologies plus the chip itself are called NPU IP.
[0052] Among them, quantization is to convert the parameters in the neural network model from 32-bit floating-point numbers (FP32) to a lower-precision data type (such as INT8), which can reduce the storage requirements and computational volume of the model, thereby accelerating the inference of the model, while also ensuring a relatively low loss of model accuracy. The quantized output model is further optimized by the compiler: instruction generation techniques such as operator fusion, constant folding, and loop unrolling, and is converted into an executable file, so as to achieve operation on the hardware.
[0053] To ensure that the model after quantization and compilation can output correct calculation results, it is necessary to verify the entire quantization, compilation, and subsequent NPU board test process. Take the running result of the model on the CPU as the ground truth, and then execute the quantization, compilation, and actual NPU board running process. Compare the obtained calculation results (including the intermediate calculation results of each layer of the neural network) with the ground truth. If the comparison passes, the verification is successful; if the comparison fails, the verification fails, and it is necessary to analyze the problems existing in the quantization and compilation processes. Therefore, a large number of test cases need to be generated for verification, and this process is called coverage testing.
[0054] The existing coverage testing method generates test cases at the single-operator level. For example, to generate a test case for a Conv2d (2D convolution operator): nn.Conv2d(in_channels = 3, out_channels = 16,...), randomly sample each parameter of the operator (such as in_channels) within a certain numerical range, and a large number of test cases can be generated. Each test case needs to go through the complete steps of quantization, compilation, and board testing, and is considered verified successfully only after the comparison passes. Otherwise, any error in any intermediate link is considered a verification failure, and the engineer needs to locate and solve the corresponding problem and then verify again. This process is repeated until each test case can pass the verification, indicating that the NPU IP passes the test for this batch of test cases. However, the current coverage testing method can only verify whether a single operator can operate correctly, and cannot test whether the optimization (folding, fusion, etc.) of quantization and compilation for multiple operators is correct; it cannot verify some abnormal network structures.
[0055] Therefore, to solve the above problems, the embodiments of the present disclosure provide a network testing method, which can be applied to, for example, Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 can obtain multiple basic operators supported by quantization compilation in the server 104, combine the multiple basic operators to obtain at least one composite operator. The terminal 102 generates at least one random network based on the at least one composite operator and the basic operator, and according to a preset input tensor and a preset random network depth. The terminal 102 performs quantization compilation on the at least one random network, and performs on-board testing on the at least one quantized-compiled random network to obtain a test result. The terminal 102 determines the true value result of the at least one random network, and determines an abnormal network based on the true value result and the test result, where the true value result is obtained by processing the at least one random network using a central processing unit. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that this solution can also be applied to the terminal or the server alone to implement.
[0056] In one embodiment, as Figure 2 shown, a network testing method is provided, taking the method applied to Figure 1 the terminal 102 in
[0057] S202, obtain multiple basic operators supported by quantization compilation, and combine the multiple basic operators to obtain at least one composite operator.
[0058] Among them, quantization is to convert the parameters in the neural network model from a high-precision data type (such as 32-bit floating-point number) to a low-precision data type (such as 8-bit integer) to reduce the storage requirements and computational amount of the model and improve the inference speed. Compilation is the process of converting the model into code that can run efficiently on specific hardware (such as NPU). Quantization compilation combines these two steps, aiming to optimize the performance of the model on the hardware. Basic operator: is the most basic computational unit in the neural network, such as the convolution operator (nn.Conv2d), the batch normalization operator (nn.BatchNorm2d), the activation function operator (nn.ReLU), etc. These operators each complete specific mathematical operations and are the basis for constructing the neural network. A composite operator is a new operator composed of multiple basic operators. By fusing multiple related basic operators into one composite operator, the data transmission and intermediate result storage overhead in the calculation process can be reduced, and the calculation efficiency can be improved.
[0059] Specifically, the set of basic operators supported by the compilation tool can be clearly quantified. For example, for common deep learning frameworks (such as PyTorch) and corresponding quantization compilation tools, the supported basic operators may include various types of operators such as convolution, pooling, normalization, activation functions, etc. These operators can be used as basic operators. To increase the diversity and comprehensiveness of the tests, a random combination method can be adopted. Randomly select multiple basic operators from the list of supported basic operators for combination.
[0060] S204, based on the at least one composite operator and the basic operator, and according to a preset input tensor and a preset random network depth, generate at least one random network.
[0061] Among them, the preset input tensor: the input data preset before constructing the random network, usually a multi-dimensional tensor. For example, in image-related tasks, the input tensor may be four-dimensional, in the form of (batch size, channels, height, weight), providing the initial data for the network, and subsequent operators will perform calculations based on this input. The preset random network depth: the number of operators included in the preset random network, which determines the complexity and computational depth of the network. The greater the depth, the more complex the features the network may learn.
[0062] Specifically, all composite operators and basic operators can be aggregated into a set for convenient subsequent extraction. Set the preset input tensor and random network depth according to requirements. Create an empty network container to store the operators added later. According to the preset random network depth, perform a loop operation, and randomly select an operator from the operator set each time. According to the type of the selected operator and the number of channels of the current input tensor, randomly set the parameters of the operator, and then add it to the network. Finally, generate a random network. Repeat the above process to generate multiple random networks.
[0063] S206, perform quantization compilation on the at least one random network, and conduct on-board testing on the at least one quantized-compiled random network to obtain test results.
[0064] Among them, the test can be to deploy the quantized-compiled random network to an actual hardware development board (such as a development board containing an NPU) for running tests.
[0065] Specifically, according to the target hardware (such as NPU), determine the quantization method, such as 8-bit integer quantization, mixed-precision quantization, etc. Insert quantization nodes at appropriate positions in the computational graph of the stochastic network. These nodes will perform quantization and dequantization operations on the data during model runtime to ensure that the model can still work properly at low precision. The insertion positions of the quantization nodes are usually determined according to the network structure and the characteristics of the operators, for example, before and after the convolutional layer, fully connected layer, etc. (not limited in some embodiments of the present disclosure). Perform forward propagation on the stochastic network using a representative calibration dataset. During this process, collect statistical information of the data in each layer of the network, such as the maximum value, minimum value, etc., to determine the scaling factors and zeros required for quantization. According to the parameters obtained from calibration, convert the floating-point parameters and operations in the stochastic network into a quantized representation. At this time, data such as weights and activation values in the network are stored and calculated in a low-precision form. Optimize the computational graph of the quantized stochastic network. According to the instruction set and architecture characteristics of the target hardware, convert the optimized computational graph into machine code or intermediate representation that the target hardware can execute. For example, for an NPU chip, the compilation tool will generate code suitable for the NPU instruction set to ensure that the network can run efficiently on this hardware. Run the quantized and compiled stochastic network model on the development board, input the test data into the model for inference calculation, and record the output results of the model. During the running process, the performance metrics of the hardware, such as running time, power consumption, etc., can also be monitored.
[0066] S208, determine the true value result of the at least one stochastic network, and determine the abnormal network based on the true value result and the test result, wherein the true value result is obtained by processing the at least one stochastic network using a central processing unit.
[0067] Specifically, the random network is deployed on the central processing unit for processing. The CPU usually has high-precision computing power and a stable operating environment, making it suitable as a reference platform for obtaining true value results. The test data is sequentially input into the random network and calculated through the forward propagation process of the network. During the calculation process, the output results of each layer of the network are recorded, and finally the output results of the entire network for the test data are obtained, which are the true value results. Due to the high-precision characteristics of CPU computing, these results are considered relatively accurate reference values. Ensure that the data formats, dimensions, and orders of the true value results and the test results are consistent for accurate comparison. For example, if the true value results and the test results are in tensor form, it is necessary to ensure that their shapes and data types are the same. For each random network, compare its true value results with the test results and calculate the corresponding comparison metric values. Compare the calculated metric values with the set threshold: if the metric value exceeds the threshold, it indicates that there is a large difference between the test results and the true value results of the random network, and the random network is determined to be an abnormal network. If the metric value does not exceed the threshold, it is considered that the test results of the random network match the true value results and the network is normal. Among them, the metric values can include: accuracy, mean square error, cosine similarity, etc.
[0068] In the above network testing method, by combining multiple basic operators into composite operators, the data transmission between operators and the storage of intermediate results are reduced, and the overhead during the calculation process is lowered. Generating random networks based on composite operators and basic operators can explore different network structure combinations. Using the central processing unit (CPU) to calculate the true value results of the random network and comparing them with the test results can accurately evaluate the accuracy of the model during quantization compilation and hardware operation. Determining abnormal networks based on the true value results and the test results helps to promptly discover and solve problems that occur during network construction, quantization compilation, or hardware operation. Conducting in-depth analysis of abnormal networks can identify the root causes of problems, such as unreasonable quantization parameter settings and inappropriate compilation optimization strategies, and make corresponding adjustments and improvements to improve the reliability and stability of the model. Quantizing and testing multiple random networks can comprehensively verify the performance and compatibility of the hardware and the quantization compilation tool. Discovering and solving problems in the early stage of development can avoid problems that occur during large-scale testing in the later stage, reduce repeated modifications and debugging during the development process, and improve development efficiency. The diversity of random networks enables testing to cover more network structures and computing scenarios, thereby more comprehensively verifying the performance of the hardware and the model. This helps to discover potential problems and vulnerabilities and improve the quality and stability of the product.
[0069] In one embodiment, as Figure 3 shown, the combining of the multiple basic operators to obtain at least one composite operator includes:
[0070] S302, obtain the function information and parameter information of each of the basic operators.
[0071] S304, adjust the parameter information of each basic operator, and combine the parameter information and function information of at least two adjusted basic operators to obtain at least one composite operator.
[0072] Among them, function information: refers to the description of the specific calculation function and role of the basic operator. For example, the function information of the convolution operator is to perform a convolution operation on the input data to extract local features; the function information of the batch normalization operator is to normalize the mean and variance of the input data to accelerate the training and convergence of the network; the function information of the activation function operator is to perform a non-linear transformation on the input data so that the network can learn more complex patterns. Parameter information: the parameters that can be adjusted in the basic operator, and these parameters will affect the calculation result and behavior of the operator. For the convolution operator nn.Conv2d, the parameter information includes the number of input channels (in_channels), the number of output channels (out_channels), the convolution kernel size (kernel_size), the stride (stride), the padding (padding), etc.; for the batch normalization operator nn.BatchNorm2d, the parameter information is mainly the number of features (num_features); for the activation function operator nn.ReLU, possible parameter information such as whether to perform an in-place operation (inplace), etc.
[0073] Specifically, for each of the common deep learning frameworks (such as PyTorch, TensorFlow, etc.), the function information and parameter information of each basic operator can be obtained by referring to the official documentation. The documentation will describe in detail the purpose of the operator, the input and output requirements, as well as the meaning and value range of each parameter. Alternatively, in the code, the function information and parameter information can be understood by viewing the definition and related comments of the operator. For example, in PyTorch, view the definition code of nn.Conv2d to understand its parameter setting method and function. The parameter information of the basic operator can be adjusted randomly. For example, for the convolutional operator nn.Conv2d, randomly generate parameters such as the number of input channels, the number of output channels, and the kernel size that meet the value range. By randomly setting parameters, a large number of different operator combinations can be simulated, covering a wider range of usage scenarios, so as to comprehensively detect the correctness and stability of the NPU chip when processing different parameter values and discover potential problems. There are both quantization and compilation techniques to support the optimization of multiple operator combinations. Random parameter setting can test the effect of these techniques on operator fusion under different parameter conditions, verify their compatibility and optimization ability for diverse parameter configurations, and ensure that the techniques can effectively play their roles in various situations. Avoiding the limitations brought by fixed parameter settings, randomization can generate rich and diverse test cases, increase the randomness and unpredictability of the test, more realistically reflect the performance of the NPU chip in a complex and changing actual environment, and improve the effectiveness and reliability of the test results.
[0074] In an image recognition task, if more refined features need to be extracted, the stride of the convolutional kernel can be appropriately reduced, and the number of convolutional kernels can be increased, etc. According to the computational logic and requirements of the neural network, determine the combination order of at least two basic operators. In a convolutional neural network, the combination order can be convolutional layer -> batch normalization layer -> activation function layer. Integrate the adjusted parameter information and function information of the basic operator to construct a new composite operator. The composite operator can be implemented by defining a new class. In the class, call each basic operator in sequence according to the combination order and pass the corresponding parameters.
[0075] In some exemplary embodiments, such as Figure 4 shown Figure 4 illustrates the process of fusing three independent neural network operators (nn.Conv2d, nn.BatchNorm2d, nn.ReLU) into a composite operator ConvBNRelu.
[0076] On the left are three independent operators:
[0077] nn.Conv2d is a two-dimensional convolutional layer. in_channels = 3 indicates that the number of input channels is 3, out_channels = 8 indicates that the number of output channels is 8, kernel_size = 3 is the convolutional kernel size, stride = 1 is the stride, padding = 0 means no padding, dilation = 1 is the dilation rate, and groups = 1 indicates that the number of groups is 1.
[0078] nn.BatchNorm2d is a two-dimensional batch normalization layer. num_features = 8 represents the number of features, which is consistent with the number of output channels of the previous nn.Conv2d.
[0079] nn.ReLU is a rectified linear unit layer. inplace = True means performing the operation in-place to save memory.
[0080] The composite operator ConvBNRelu on the right integrates the functions and parameters of the previous three operators, combining convolution, batch normalization, and activation operations into one operation.
[0081] In this embodiment, combining multiple basic operators into a composite operator can reduce data transmission between operators and storage of intermediate results, reducing computational overhead. The composite operator ConvBNReLU completes convolution, batch normalization, and activation operations in one forward pass. Compared with performing these three operations separately, it reduces the number of times data is transferred between different operators, improving computational efficiency. By adjusting the parameter information of the basic operators and combining them, various composite operators with different functions and characteristics can be generated. These composite operators can better adapt to different application scenarios and network structure requirements, providing more choices and flexibility for the design of neural networks.
[0082] In one embodiment, as Figure 5 shown, generating at least one random network based on the at least one composite operator and the basic operator, and according to a preset input tensor and a preset random network depth includes:
[0083] S402, selecting at least one initial operator from the at least one composite operator and the basic operator, and updating the sampling space of the at least one initial operator according to the preset input tensor.
[0084] Among them, the initial operator is usually selected from composite operators and basic operators as the operator for the initial stage of constructing a random network, and the construction of the subsequent network will be gradually carried out based on these initial operators. For each operator, the sampling space is the range of values for its parameters (such as convolution kernel size, number of channels, stride, etc.), and this range of values constitutes the sampling space of the operator. When constructing a random network, it is necessary to update the sampling space according to information such as the dimension of the input tensor to ensure that the parameter settings of the operator are reasonable. The first operator is the specific operator obtained by sampling in the updated sampling space, which determines the specific parameter values and is an actual component of constructing the random network.
[0085] Specifically, at least one operator is randomly selected or selected according to a certain strategy (such as according to the type ratio of the operators, etc.) from the set of composite operators and basic operators as the initial operator. When constructing a random network for image recognition, several basic operators such as convolution, batch normalization, activation functions, and existing composite operators can be randomly selected as the initial operators. According to the preset dimension information of the input tensor (such as the number of channels, height, width, etc.), the parameter sampling space of each initial operator is updated. If the number of channels of the input tensor is 3, then for the convolution operator nn.Conv2d, the sampling space of its in_channels parameter is limited to 3, and reasonable value ranges for other parameters (such as out_channels, kernel_size, etc.) are also set according to the actual situation and experience.
[0086] S404, sample in the updated sampling space of the at least one initial operator to obtain at least one first operator.
[0087] Specifically, within the updated sampling space, each initial operator is randomly sampled to determine its specific parameter values, thereby obtaining the first operator. For the convolution operator nn.Conv2d, when the limited in_channels is 3, randomly sampling to obtain parameters such as out_channels being 16 and kernel_size being 3 to form the specific first operator of nn.Conv2d.
[0088] In some exemplary embodiments, a randomly input four-dimensional tensor is set, and its dimensions are: (batchsize, channels, height, weight). For example, a random tensor with dimensions (1, 4, 28, 28);
[0089] Randomly select an operator from all single operators and composite operators supported by the NPU, such as nn.Conv2d, and assume its parameter sampling space is:
[0090] Table 1. Example of the parameter sampling space of nn.Conv2d supported by NPU
[0091]
[0092]
[0093] Table 2 Updated parameter sampling space
[0094]
[0095] At this time, it is necessary to update the sampling space according to the dimensions of the input tensor (as shown in Table 2), because adding a new operator is not arbitrary. Some parameters of the new operator must be consistent with some dimensions of its input tensor. For example, for a convolutional operator nn.Conv2d, its in_channels must be consistent with the channels of its input tensor. If the dimensions of the input tensor are (1, 3, 16, 16) and its channels = 3, then the in_channels of nn.Conv2d must be equal to 3. Therefore, the sampling space of its in_channels should be updated from [1, 4096] to [3]. For example, the in_channels parameter must be equal to the number of input channels channels of the tensor (here it is 3), the groups cannot exceed channels, and must be a factor of channels.
[0096] Then sample in the updated parameter space to obtain an nn.Conv2d operator, such as nn.Conv2d(in_channels = 4, out_channels = 16, kernel_size = 5, stride = 2, padding = 3, dilation = 1, groups = 2).
[0097] S406. Construct a random network based on the at least one first operator and the preset random network depth.
[0098] Specifically, based on the obtained first operator, according to certain rules (such as topological order), and according to the preset random network depth, gradually add new operators (which can continue to be selected from the set of composite operators and basic operators). When adding a new operator, it is also necessary to update the sampling space of the new operator according to the output tensor dimensions of the previous operator and sample to determine the parameters until the number of operators in the network reaches the preset random network depth, and the construction of the random network is completed.
[0099] In this embodiment, by operations such as randomly selecting initial operators, updating the sampling space, and randomly sampling to determine operator parameters, a large number of random networks with different structures and parameter configurations can be generated. This diversity helps to explore different network architectures, discover better network structures, and improve the performance of the network in various tasks. The sampling space is updated according to the preset input tensor, so that the generated random network can better adapt to the characteristics of the input data. Different input data dimensions and distributions may require operators with different parameter configurations, and the constructed random network can process the input data more effectively, improving the generalization ability of the network. The construction process of the random network simulates various network structures and parameter settings that may be encountered in actual applications. By testing and analyzing these random networks, the processing capabilities of hardware (such as NPU) for different network structures can be better evaluated, potential problems can be discovered in advance and optimized, and the performance and stability of the hardware in actual applications can be improved.
[0100] In one embodiment, as Figure 6 shown, constructing a random network based on the at least one first operator and the preset random network depth includes:
[0101] S502, constructing an initial network based on the at least one first operator.
[0102] S504, in response to the network depth of the initial network satisfying the preset random network depth, determining the initial network as the random network.
[0103] S506, in response to the network depth of the initial network not satisfying the preset random network depth, selecting at least one successor operator from the at least one composite operator and the basic operator.
[0104] S508, determining the predecessor operator of the at least one successor operator from the initial network, and updating the sampling space of the at least one successor operator based on the output tensor of the predecessor operator.
[0105] S510, sampling in the updated sampling space of the at least one successor operator, and adding the sampled subsequent operator after the predecessor operator in the initial network until the initial network obtained after the addition satisfies the preset random network depth, to obtain a random network.
[0106] Wherein, when the network depth of the initial network does not satisfy the preset random network depth, the successor operator is a new operator selected from the composite operator and the basic operator for adding to the initial network to increase the network depth. In the initial network, the predecessor operator is the operator that provides input for the successor operator. The parameter settings and connections of the successor operator need to be determined based on the output tensor of the predecessor operator.
[0107] Specifically, based on at least one first operator, connect these operators in a certain order (such as topological order) to construct an initial network. Arrange multiple first operators in certain combinations to form a preliminary network structure. Calculate the number of operators included in the initial network, that is, the network depth. Compare the calculated network depth with a preset random network depth. When the network depth of the initial network does not meet the preset random network depth, randomly or according to a certain strategy select at least one successor operator from the set of composite operators and basic operators. Determine the predecessor operators of these successor operators in the initial network, and update the parameter sampling space of the successor operators according to the dimension (such as the number of channels, height, width, etc.) information of the output tensor of the predecessor operators. For the convolutional operator nn.Conv2d, if the number of output channels of the predecessor operator is 16, then the sampling space of the in_channels parameter of this convolutional operator is limited to 16. Within the updated sampling space, randomly sample each successor operator to determine its specific parameter values, and then add the sampled successor operator behind the corresponding predecessor operator in the initial network. Repeat the above process of selecting successor operators, determining predecessor operators, updating the sampling space, sampling, and adding operators until the network depth of the initial network obtained after adding the successor operators meets the preset random network depth.
[0108] In this embodiment, by presetting the random network depth and gradually adding operators according to the situation of the network depth, the structure scale of the random network can be accurately controlled. The network depth can be flexibly set according to actual requirements and hardware resources to construct a network with appropriate complexity, avoiding the network structure being too simple or too complex. Randomness is introduced in the process of selecting successor operators, updating the sampling space, and sampling to determine operator parameters, which makes the generated random network have rich diversity.
[0109] In one embodiment, when one of the successor operators requires multiple predecessor operators, the determining the predecessor operators of the at least one successor operator from the initial network includes:
[0110] Determine the operators in the initial network with the same number of output channels in the output tensor as the predecessor operators of the at least one successor operator.
[0111] Specifically, when some successor operators require multiple predecessor operators. This is the case for certain operators (such as torch.cat, which is used to concatenate tensors along a specified dimension) that require multiple inputs when performing operations. After determining that a successor operator requires multiple predecessor operators, traverse the output tensors of each operator in the initial network. Find those operators whose output tensors have the same number of output channels and identify them as the predecessor operators of the successor operator. In an initial network containing multiple convolutional layers and other operators, look for operators with 16 output channels and use these operators as the predecessor operators of a successor operator that requires multiple predecessor operators (such as torch.cat and requires the same number of input channels). After determining the predecessor operators, update the parameter sampling space of the successor operator according to other dimension information (such as height, width, etc.) of the output tensors of the predecessor operators. Sample the successor operator within the updated sampling space to determine its specific parameter values, and then add the successor operator after the corresponding predecessor operators in the initial network and continue to build the network until the network depth meets the preset random network depth.
[0112] In this embodiment, for a successor operator that requires multiple predecessor operator inputs, requiring the output tensors of the predecessor operators to have the same number of output channels ensures the consistency of data dimensions when the successor operator performs calculations. If the input tensors have different numbers of channels, the torch.cat operator cannot correctly concatenate tensors. In this way, calculation errors caused by mismatched data dimensions can be avoided, ensuring the correct execution of the network calculation logic. Selecting the predecessor operators according to the rule of the same number of output channels makes the construction of the network structure more reasonable.
[0113] In one embodiment, as Figure 7 shown, after determining the abnormal network based on the true value result and the test result, the method further includes:
[0114] S602, obtain a target test network, perform quantization compilation on the target test network and then conduct on-board testing, in response to the test result of the on-board testing being abnormal;
[0115] S604, determine the network structure of the target test network, and in response to the network structure of the target test network being the same as the network structure of the abnormal network, determine that there is an error in the network structure of the target test network.
[0116] Among them, the target test network can be the neural network model to be tested this time. It can be a randomly generated network through the previous steps or a specific network designed in advance, aiming to verify the performance and correctness on the target hardware (such as the NPU development board). The abnormal network can be a network that has been determined to have abnormal test results during the previous test process. There are significant differences between the test results of these networks and the true value results. The network structure refers to the specific architecture of the neural network, including the types of operators contained in the network (such as convolution, pooling, activation functions, etc.), the connection methods of the operators (such as which operators are sequential, which operators have multiple inputs, etc.), and the parameter settings of each layer (such as the convolution kernel size, number of channels, etc.).
[0117] Specifically, the target test network can be obtained. The target test network can be a randomly generated network previously or a network manually designed and generated according to specific requirements. According to the quantization and compilation steps mentioned above, perform quantization and compilation operations on the target test network to convert it into a form suitable for running on the target hardware. Deploy the quantized and compiled target test network to the hardware development board (usually the NPU), use the test data set for inference calculation, record the test results, and compare them with the true value results obtained by running on the CPU. If the difference between the two exceeds the preset threshold, the test result of the on-board test is determined to be abnormal. By analyzing the code or model definition file of the target test network, determine the types of operators contained in the network, the connection order of the operators, and the parameter settings of each layer and other information, so as to clarify the specific structure of the network. Compare the network structure of the target test network in detail with the network structure of the previously determined abnormal network. It can be judged whether the two are the same by comparing aspects such as the types, quantities, connection methods of the operators in the network, and the parameter settings of each layer. If the network structure of the target test network is the same as that of the abnormal network, it can be determined that there is an error in the test result of the target test network. This means that there may be inherent problems in this network structure, resulting in abnormalities during the quantization and compilation and on-board test processes. The abnormality that is not caused by the quantization and compilation and on-board test processes is caused by the network structure error and requires adjusting the network structure.
[0118] In this embodiment, by comparing network structures, it is possible to quickly determine whether a target test network has the same problems as known abnormal networks. When it is found that the network structures are the same, it can be directly determined that the network has errors, avoiding the consumption of time and energy for further in-depth analysis of the network and improving the efficiency of problem location. For the network structures determined to be incorrect, in-depth analysis and research can be carried out to find out the reasons for the anomalies, such as unreasonable combinations of certain operators, improper parameter settings, etc. Then, based on the analysis results, the network structure can be optimized and improved to avoid similar problems from occurring again in subsequent network design and testing, and to improve the performance and stability of the network. During the large-scale network testing process, a large number of networks need to be tested. By this method based on network structure comparison, networks that may have problems can be quickly screened out, and key attention and processing can be given to these networks, while networks with different structures can continue with the normal testing process, thus improving the efficiency of the entire testing process.
[0119] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.
[0120] Based on the same inventive concept, the embodiments of the present disclosure also provide a network testing device for implementing the above-mentioned network testing method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the network testing device provided below can refer to the limitations on the network testing method in the above text, and will not be repeated here.
[0121] In one embodiment, as Figure 8 shown, a network testing device 700 is provided, including: an operator combination module 702, a random network generation module 704, a testing module 706, and an abnormal network determination module 708, where:
[0122] The operator combination module 702 is configured to obtain a plurality of basic operators supported by quantization compilation, and combine the plurality of basic operators to obtain at least one composite operator;
[0123] A random network generation module 704, configured to generate at least one random network based on the at least one composite operator and the basic operator, and according to a preset input tensor and a preset random network depth;
[0124] A test module 706, configured to perform quantization compilation on the at least one random network, and perform on-board testing on the at least one quantized and compiled random network to obtain a test result;
[0125] An abnormal network determination module 708, configured to determine a true value result of the at least one random network, and determine an abnormal network based on the true value result and the test result, where the true value result is obtained by processing the at least one random network using a central processing unit.
[0126] In an embodiment of the device, the operator combination module 702 includes:
[0127] An information acquisition module, configured to acquire function information and parameter information of each of the basic operators.
[0128] A combination module, configured to adjust the parameter information of each basic operator, and combine the parameter information and function information of at least two adjusted basic operators to obtain at least one composite operator.
[0129] In an embodiment of the device, the random network generation module 704 includes:
[0130] A sampling space update module, configured to select at least one initial operator from the at least one composite operator and the basic operator, and update the sampling space of the at least one initial operator according to the preset input tensor;
[0131] A sampling module, configured to sample in the updated sampling space of the at least one initial operator to obtain at least one first operator;
[0132] A network construction module, configured to construct a random network based on the at least one first operator and the preset random network depth.
[0133] In an embodiment of the device, the network construction module includes:
[0134] An initial network construction module, configured to construct an initial network based on the at least one first operator.
[0135] A network determination module, configured to determine the initial network as a random network in response to the network depth of the initial network satisfying the preset random network depth.
[0136] A successor operator selection module, configured to select at least one successor operator from the at least one composite operator and the basic operator in response to the network depth of the initial network not meeting the preset random network depth.
[0137] An update module, configured to determine the predecessor operators of the at least one successor operator from the initial network, and update the sampling space of the at least one successor operator based on the output tensors of the predecessor operators.
[0138] An operator addition module, configured to sample in the updated sampling space of the at least one successor operator, and add the sampled subsequent operator after the predecessor operator in the initial network until the initial network obtained after the addition meets the preset random network depth, thereby obtaining a random network.
[0139] In an embodiment of the device, when one successor operator requires multiple predecessor operators, the update module is further configured to determine, from the initial network, operators with the same number of output channels in the output tensors as the predecessor operators of the at least one successor operator.
[0140] In an embodiment of the device, the device further includes: a network structure determination module, configured to obtain a target test network, perform on-board testing after quantization compilation of the target test network, and in response to the test result of the on-board testing being abnormal; determine the network structure of the target test network, and in response to the network structure of the target test network being the same as the network structure of the abnormal network, determine that the network structure of the target test network is incorrect.
[0141] Each module in the above network testing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to be called by the processor to execute the operations corresponding to the above respective modules.
[0142] In an embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 9As shown in the figure. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store basic operators and composite selections. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a network testing method.
[0143] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 10 the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a network testing method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0144] Those skilled in the art can understand that Figure 9 the structures shown in Figure 9 or 10 are merely block diagrams of some structures related to the solution of the present disclosure, and do not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0145] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0146] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0147] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.
[0148] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAMs), magnetoresistive random access memories (MRAMs), ferroelectric random access memories (FRAMs), phase change memories (PCMs), graphene memories, etc. Volatile memories can include random access memories (RAMs) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present disclosure can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present disclosure can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0150] The above-described embodiments merely represent several implementation manners of the present disclosure. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can still be made, and these all fall within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the appended claims.
Claims
1. A network testing method, characterized in that: The method comprises: Acquire multiple basic operators supported by quantization compilation, and combine the multiple basic operators to obtain at least one composite operator; Based on the at least one composite operator and the basic operator, at least one random network is generated according to a preset input tensor and a preset random network depth; Quantizing and compiling the at least one random network, and testing the at least one random network after quantizing and compiling to obtain a test result; Determine a true value result of the at least one random network, and determine an abnormal network based on the true value result and the test result, wherein the true value result is obtained after processing the at least one random network using a central processing unit.
2. The method according to claim 1, characterized in that The combining of the multiple basic operators to obtain at least one composite operator comprises: Obtaining function information and parameter information of each of the basic operators; The parameter information of each basic operator is adjusted, and the adjusted parameter information and function information of at least two basic operators are combined to obtain at least one composite operator.
3. The method according to claim 1, characterized in that: The generating at least one random network based on the at least one composite operator and the basic operator and according to a preset input tensor and a preset random network depth includes: Selecting at least one initial operator from the at least one composite operator and the basic operator, and updating a sampling space of the at least one initial operator according to the preset input tensor; Sampling is performed in the updated sampling space of the at least one initial operator to obtain at least one first operator; A random network is constructed based on the at least one first operator and the preset random network depth.
4. The method according to claim 3, characterized in that The step of constructing a random network based on the at least one first operator and the preset random network depth includes: constructing an initial network based on the at least one first operator; In response to the network depth of the initial network satisfying the preset random network depth, determining that the initial network is a random network; In response to the network depth of the initial network not satisfying the preset random network depth, selecting at least one successor operator from the at least one composite operator and the basic operator; Determine a predecessor operator of the at least one successor operator from the initial network, and update a sampling space of the at least one successor operator based on an output tensor of the predecessor operator; Sampling is performed in the sampling space of the at least one successor operator after the update, and the subsequent operator obtained after the sampling is added after the predecessor operator in the initial network until the initial network obtained after the addition meets the preset random network depth, thereby obtaining a random network.
5. The method according to claim 4, characterized in that When a successor operator requires multiple predecessor operators, determining the predecessor operators of at least one successor operator from the initial network includes: An operator having the same number of output channels as that contained in the output tensor is determined from the initial network as a predecessor operator of at least one successor operator.
6. The method according to claim 1, characterized in that After determining the abnormal network based on the true value result and the test result, the method further includes: Acquire a target test network, perform a board test on the target test network after quantization compilation, and respond in response to a test result of the board test being abnormal; The network structure of the target test network is determined, and in response to the network structure of the target test network being the same as the network structure of the abnormal network, it is determined that an error occurs in the network structure of the target test network.
7. A network testing device, characterized in that: The device comprises: An operator combination module, used to obtain multiple basic operators supported by quantization compilation, and combine the multiple basic operators to obtain at least one composite operator; A random network generation module, configured to generate at least one random network based on the at least one composite operator and the basic operator and according to a preset input tensor and a preset random network depth; A testing module, used to quantize and compile the at least one random network, and perform a board test on the at least one random network after quantization and compilation to obtain a test result; The abnormal network determination module is used to determine the true value result of the at least one random network, and determine the abnormal network based on the true value result and the test result, wherein the true value result is obtained by processing the at least one random network using a central processing unit.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Chip testing method and device, equipment, storage medium and program product
CN120994485A
Software testing method, related device and storage medium
CN121614407A