An Interpretability Method Based on the Dynamic Behavior of a Model and Related Devices
By acquiring and analyzing the asymmetric dynamic boundaries of image data in the convolutional neural network model, the problem of lack of transparency and interpretability of the model is solved, and the precise interpretation and transparency of the model decision-making process are achieved.
Patent Information
- Application Number
- CN202010683161.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-07-15
AI Technical Summary
The convolutional deep neural network model lacks transparency and interpretability. The existing interpretation methods cannot accurately explain the model decision process and reflect the dynamic behavior of the model regarding input characteristics, and cannot meet the needs of risk-sensitive tasks.
By obtaining the asymmetric dynamic boundary of the target pixel point in the image data, including the left dynamic boundary and the right dynamic boundary, the difference operation is performed, and the difference between the left and right dynamic boundaries is obtained, and the difference is analyzed to explain the decision mechanism of the neural network model.
In the case of faithfulness to the original model, accurately interpret the decision-making process of the neural network model, increase the transparency of the model, reduce deployment risks, and establish user trust relationships.
Smart Images

Figure CN113947178B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to an interpretability method based on the dynamic behavior of a model and related devices. Background Art
[0002] With the development of artificial intelligence and deep learning, convolutional deep neural networks have shone brightly in many fields such as computer vision, speech recognition, image recognition, and natural language processing due to their excellent performance. However, convolutional deep neural network models lack transparency and interpretability, which severely limits the development and application of convolutional deep neural network models in real-world tasks, especially risk-sensitive tasks (such as autonomous driving, healthcare, finance, etc.). In order to increase the transparency of convolutional neural networks, reduce the potential risks of the models in deployment applications, and at the same time establish a trust relationship with users, research has now begun on how to explain the decision-making process and classification results of the models, and a series of model interpretation methods have been proposed, such as feature reverse interpretation methods, local approximation interpretation methods, model simulation interpretation methods, and so on. However, many of the above methods using surrogate models cannot guarantee the accuracy and reliability of the interpretation methods, and based on the static decision-making behavior of the models, they cannot truly reflect the impact of data feature changes on the final decision of the models, far from meeting the requirements in risk-sensitive tasks.
[0003] Therefore, how to accurately interpret the neural network model while being faithful to the original model is an urgent problem to be solved. Summary of the Invention
[0004] An embodiment of this application provides an interpretability method based on the dynamic behavior of a model and related devices, which can accurately interpret the neural network model while being faithful to the original model.
[0005] In a first aspect, an embodiment of this application provides an interpretability method based on the dynamic behavior of a model, which may include:
[0006] Obtain the asymmetric dynamic boundaries of the target pixel points in the image data, where the image data is the input data of the target neural network model, the target pixel points are any of the multiple pixel points in the image data, the asymmetric dynamic boundaries include a left dynamic boundary and a right dynamic boundary, the left dynamic boundary is the lower boundary of the pixel value of the target pixel point, the right dynamic boundary is the upper boundary of the pixel value of the target pixel point, and when the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged; perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point; according to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively, obtain the target pixel point set of the image data, and the target pixel point set includes one or more pixel points among the multiple pixel points whose differences between the left and right dynamic boundaries exceed a first threshold.
[0007] Regarding the problems that the convolutional neural network model lacks transparency and interpretability, and the existing interpretability methods cannot accurately interpret the convolutional neural network model and reflect the dynamic behavior of the model regarding input features, etc., in the embodiments of the present application, for a given neural network model and image data (i.e., the input data of the neural network model), the asymmetric dynamic boundaries of the image data under the neural network model can be obtained. The asymmetric dynamic boundaries include a left dynamic boundary and a right dynamic boundary. Among them, the left dynamic boundary is the range boundary where reducing the pixel value will not change the classification result, and the right dynamic boundary is the range boundary where increasing the pixel value will not change the classification result. Therefore, perform a difference operation on the left dynamic boundary and the right dynamic boundary to obtain the difference between the left and right dynamic boundaries. Finally, analyze according to the difference between the left and right dynamic boundaries to obtain the target pixel point set of the image data. Among them, the input data is image data. This uses the dynamic behavior of the neural network model regarding the target pixel points in the input data (that is: increasing or decreasing the pixel value of the target pixel point and not changing its classification result) to explain the decision-making mechanism of the neural network model for each input data of the neural network model, thereby increasing the transparency of the neural network and establishing a trust relationship with users.
[0008] In a possible implementation manner, the target neural network model includes a convolutional layer and / or a fully connected layer. Implementing the embodiments of the present application can interpret the convolutional neural network model and reflect the dynamic behavior of the convolutional neural network model regarding input features. Among them, the target neural network model including a convolutional layer and / or a fully connected layer is a convolutional neural network.
[0009] In a possible implementation manner, the method further includes: normalizing the difference between the left and right dynamic boundaries of the target pixel point to obtain the differential dynamic boundary value of the target pixel point; the obtaining the target pixel point set of the image data according to the differences between the left and right dynamic boundaries corresponding to multiple pixel points in the image data includes: determining, from the differential dynamic boundary values corresponding to the multiple pixel points, the pixel points whose differential dynamic boundary values exceed a second threshold as the target pixel point set of the image data. Implementing the embodiments of the present application, by normalizing, the difference between the left and right dynamic boundaries of the target pixel point is processed from a relatively large difference value to a relatively small value, which is convenient for arithmetic processing and also makes the difference in the influence of each pixel point on the classification result more obvious, for example: processed within the range of 0-1. Among them, the larger the differential dynamic boundary value of the target pixel point, the more the classification result of the image data in the target neural network model depends on the features of the target pixel point. For example: normalizing the differential dynamic boundary to within [0,1], a larger value means that the predicted class tends to have this feature in the input data (that is, if this feature value becomes smaller, the class will change), otherwise, the predicted class tends to not have this feature. For example: the horizontal line above "7", that is, if this feature value becomes smaller, the class will change to 1; the horizontal line below "7", that is, if this feature value becomes larger, the class will also change to "2", indicating that the classification result of the image 7 data depends on the features of the pixel points where the horizontal line above "7" and the horizontal line below "7" are located.
[0010] In a possible implementation manner, the method further includes: obtaining a target feature map corresponding to the image data according to the differential dynamic boundary values corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a first pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a second pixel point. Among them, the differential dynamic boundary value of the first pixel point is greater than the differential dynamic boundary value of the second pixel point. Implementing the embodiments of the present application, after normalizing the differential dynamic boundary value, the decision-making mechanism of the target neural network model can be reflected. Secondly, an interpretation map, that is, a target feature map, is generated based on the numerical size of the differential dynamic boundary, and the difference in the influence of each pixel point on the classification result can be more intuitively discovered.
[0011] In a possible implementation manner, the method further includes: obtaining a target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a fourth pixel point, and the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point. Implementing the embodiments of the present application, the difference between the left and right dynamic boundaries can directly reflect the decision-making mechanism of the target neural network model. Secondly, the extracted features are relatively simple and easy to understand. Generating a target feature map based on the value of the difference between the left and right dynamic boundaries can more accurately discover the differential influence of each pixel point on the classification result.
[0012] In a possible implementation manner, the method further includes: obtaining the dynamic boundary range of the target pixel point set according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixel points; obtaining the robustness boundary range of the target neural network model according to the dynamic boundary range, where the robustness boundary range is within the dynamic boundary range and is used to determine the classification result after inputting into the target neural network model. Implementing the embodiments of the present application, the left dynamic boundary is the range limit for reducing the pixel point value without changing the classification result, and the right dynamic boundary is the range limit for increasing the pixel point value without changing the classification result. The robustness boundary range is within the left and right dynamic boundary ranges. From the perspective of the robustness of the target neural network model, the robustness boundary range is equivalent to the dynamic boundary range corresponding to the image data, and both are used to determine the classification result after inputting into the target neural network model.
[0013] In a second aspect, the embodiments of the present application provide an interpretability device based on the dynamic behavior of a model, including:
[0014] A first acquisition unit, configured to acquire the asymmetric dynamic boundary of a target pixel point in the image data, where the image data is the input data of the target neural network model, the target pixel point is any one of the multiple pixel points in the image data, the asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary, the left dynamic boundary is the lower boundary of the pixel point value of the target pixel point, the right dynamic boundary is the upper boundary of the pixel point value of the target pixel point, and when the pixel point value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged;
[0015] A difference unit, configured to perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point;
[0016] A pixel unit is configured to obtain a set of target pixels of the image data according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixels, where the set of target pixels includes one or more pixels among the multiple pixels whose difference between the left and right dynamic boundaries exceeds a first threshold.
[0017] In a possible implementation manner, the target neural network model includes a convolutional layer and / or a fully connected layer.
[0018] In a possible implementation manner, the device further includes: a normalization unit configured to normalize the difference between the left and right dynamic boundaries of the target pixels to obtain a differential dynamic boundary value of the target pixels; the pixel unit is specifically configured to: determine, from the differential dynamic boundary values respectively corresponding to the multiple pixels, the pixels whose differential dynamic boundary values exceed a second threshold as the set of target pixels of the image data.
[0019] In a possible implementation manner, the device further includes: a first feature map unit configured to obtain a target feature map corresponding to the image data according to the differential dynamic boundary values respectively corresponding to the multiple pixels and the image data, where the visual perception brightness corresponding to a first pixel among the multiple pixels in the target feature map is greater than the visual perception brightness corresponding to a second pixel, and wherein the differential dynamic boundary value of the first pixel is greater than the differential dynamic boundary value of the second pixel.
[0020] In a possible implementation manner, the device further includes: a second feature map unit configured to obtain a target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixels and the image data, where the visual perception brightness corresponding to a third pixel among the multiple pixels in the target feature map is greater than the visual perception brightness corresponding to a fourth pixel, and wherein the difference between the left and right dynamic boundaries of the third pixel is greater than the difference between the left and right dynamic boundaries of the fourth pixel.
[0021] In a possible implementation manner, the device further includes: a second acquisition unit configured to obtain a dynamic boundary range of the set of target pixels according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixels; a robustness unit configured to obtain a robustness boundary range of the target neural network model according to the dynamic boundary range, where the robustness boundary range is within the dynamic boundary range and is used to determine a classification result after inputting into the target neural network model.
[0022] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor configured to support the terminal device in implementing corresponding functions in the interpretability method based on model dynamic behavior provided in the first aspect. The terminal device may further include a memory for coupling with the processor to store necessary program instructions and data of the terminal device. The terminal device may further include a communication interface for the network device to communicate with other devices or communication networks.
[0023] In a fourth aspect, an embodiment of the present application provides a computer storage medium for storing computer software instructions used for the interpretability device based on model dynamic behavior provided in the second aspect above, which includes programs designed to execute the above aspects.
[0024] In a fifth aspect, an embodiment of the present application provides a computer program, which includes instructions that, when executed by a computer, enable the computer to execute the processes performed by the interpretability device based on model dynamic behavior in the second aspect above.
[0025] In a sixth aspect, the present application provides a chip system, which includes a processor for supporting a terminal device in implementing the functions involved in the first aspect above. For example, generating or processing information involved in the interpretability method based on model dynamic behavior above. In a possible design, the chip system further includes a memory for storing necessary program instructions and data of the data sending device. The chip system may be composed of chips or may include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required for use in the embodiments of the present application or the background art.
[0027] Figure 1 is a schematic diagram of the architecture of an interpretability system based on model dynamic behavior provided by an embodiment of the present application.
[0028] Figure 2 is a schematic diagram of the structure of an interpretability device based on model dynamic behavior provided by an embodiment of the present application.
[0029] Figure 3 is a schematic diagram of the operation process of a dynamic boundary quantization module provided by an embodiment of the present application.
[0030] Figure 4 is a schematic diagram of the operation process of a model interpretation module provided by an embodiment of the present application.
[0031] Figure 5It is a schematic flowchart of an interpretability method based on the dynamic behavior of a model provided by an embodiment of the present application.
[0032] Figure 6 It is a schematic diagram of dynamic key features provided by an embodiment of the present application.
[0033] Figure 7 It is a schematic structural diagram of another interpretability device based on the dynamic behavior of a model provided by an embodiment of the present application.
[0034] Figure 8 It is a schematic structural diagram of yet another interpretability device based on the dynamic behavior of a model provided by an embodiment of the present application. Detailed implementation manners
[0035] Next, the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.
[0036] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0037] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0038] As used in the specification of this application, terms such as "component", "module", "system", etc. are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media on which various data structures are stored. A component can communicate, for example, through local and / or remote processes according to a signal having one or more data packets (e.g., data from two components interacting with each other between a local system, a distributed system, and / or a network, e.g., the Internet interacting with other systems through a signal).
[0039] First, some terms in this application are explained to facilitate the understanding of those skilled in the art.
[0040] (1) Robustness refers to the property of a control system to maintain certain other performance characteristics under certain (structural, magnitude) parameter perturbations. According to different definitions of performance, it can be divided into stability robustness and performance robustness.
[0041] (2) Convolutional Neural Network (CNN) is a class of feedforward neural networks with convolutional calculations and a deep structure, and is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input information according to their hierarchical structure, so they are also called "Shift-Invariant Artificial Neural Networks (SIANN)".
[0042] (3) Convolutional layer: In a convolutional neural network, each convolutional layer consists of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. More layers of the network can iteratively extract more complex features from the low-level features.
[0043] (4) Fully connected layer. Each node in the fully connected layer is connected to all nodes in the previous layer and is used to synthesize the features extracted previously. Due to its fully connected property, the parameters of the fully connected layer are generally the most.
[0044] Secondly, to facilitate the understanding of the embodiments of the present application, the technical problems to be solved by the embodiments of the present application and the corresponding application scenarios are specifically analyzed below. In order to increase the transparency of the convolutional neural network, reduce the potential risks in the deployment and application of the model, and at the same time establish a trust relationship with users, it is currently necessary to explain the decision-making process and classification results of the model. Therefore, a series of model interpretation methods have been proposed, such as feature reverse interpretation method, local approximation interpretation method, model simulation interpretation method, etc. Exemplarily, in the interpretation framework reflecting the dynamic behavior of the model, most of them adopt the following several ways:
[0045] Prior art one: Based on the core features extracted by the feature reverse method, we can reasonably interpret the prediction results of the convolutional neural network. For example: by calculating the gradient of the model output result with respect to the input image, generating an interpretation saliency map, and then inferring the feature importance using the saliency map.
[0046] However, the features extracted by such methods are relatively rough and difficult to understand, and cannot meet the requirements of accurate interpretation. In addition, the zero-gradient problem makes it difficult for the gradient-based feature reverse method to track and locate the decision-making features required, resulting in the failure of the interpretation method.
[0047] Prior art two: Explain the decision-making process and decision-making basis of the deep learning model for each input instance through local analysis methods. For example: using influence functions to track the prediction results of the model and identify the training instances that have the greatest impact on the prediction results.
[0048] However, such methods usually adopt approximation means and cannot guarantee the accuracy of the interpretation results. In addition, such methods analyze the internal working mechanism of the convolutional neural network through local analysis, so they can only obtain the local features of the model and cannot accurately interpret the overall behavior of the convolutional neural network model, nor can they make consistent interpretations for similar samples.
[0049] Prior art three: Make an interpretation of the model prediction results by judging the features that need to appear and the features that should not appear for a given input. For example: using an autoencoder to search for the necessary features and unnecessary features for classification near the input sample manifold.
[0050] However, such methods usually introduce the generative model VAE (Variational Auto-Encoder) or the generative model GAN (Generative Adversarial Nets). Therefore, the accuracy of the interpretation results cannot be guaranteed, and it cannot be theoretically guaranteed to be close to the original samples on the data manifold.
[0051] Therefore, to address the above technical problems, through in-depth research on the internal working mechanism of convolutional neural networks and combined with the robustness verification of the model, the present application proposes an interpretability method and system based on the dynamic behavior of the model. For each input instance, the decision-making mechanism of the model is explained using the dynamic behavior of the model with respect to features, thereby increasing the transparency of the neural network and establishing a trust relationship with users. This solves the problem in the prior art that accurate interpretation cannot be satisfied, and realizes the decision-making mechanism of the dynamic behavior interpretation model for the feature data of each actual input.
[0052] Based on the above-mentioned proposed technical problems and for the convenience of understanding the embodiments of the present application, first, an interpretability system architecture based on the dynamic behavior of the model, on which the embodiments of the present application are based, will be described below. Please refer to Figure 1 , Figure 1 is a schematic diagram of an interpretability system architecture based on the dynamic behavior of the model provided by the embodiments of the present application. The interpretability system architecture based on the dynamic behavior of the model in the embodiments of the present application may include Figure 1 the service device 001 and the terminal 002 in Figure 1 . As shown in
[0053] The service device 001 includes, but is not limited to, cloud servers, background servers, data processing servers, etc. The service device 001 can be a service device that brings various conveniences to third-party use by obtaining, processing, analyzing, and extracting data based on interactive data. When the service device 001 is a server, the server can communicate with multiple movable terminals through the Internet, and a corresponding server-side program also needs to run on the server to provide corresponding detection and other services. For example, the server can obtain the asymmetric dynamic boundary of the target pixel point in the image data. The image data is the input data of the target neural network model. The target pixel point is any one of the multiple pixel points in the image data. The asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary. The left dynamic boundary is the lower boundary of the pixel value of the target pixel point, and the right dynamic boundary is the upper boundary of the pixel value of the target pixel point. When the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged; perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point; according to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively, obtain the target pixel point set of the image data. The target pixel point set includes one or more pixel points among the multiple pixel points whose differences between the left and right dynamic boundaries exceed a first threshold.
[0054] The terminal 002 can be a device at the outermost periphery of a computer network, such as a communication terminal installed with a display device, a portable terminal, a mobile device, a user terminal, a mobile terminal, a wireless communication device, a user agent, a user device, or a user equipment (UE). When the movable terminal 002 is a laptop computer with a display, it can receive and display the feature picture sent by the server, and intuitively display the key features in the input data that affect the classification result of the neural network model. For example, according to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively and the image data, obtain the target feature map corresponding to the image data. In the target feature map, the visual perception brightness corresponding to the pixel points with a large difference between the left and right dynamic boundaries among the multiple pixel points is greater than the visual perception brightness corresponding to the pixel points with a small difference between the left and right dynamic boundaries. In the embodiments of the present application, the movable terminal may also include, but is not limited to, any movable electronic product based on an intelligent operating system, which can perform human-computer interaction with the user through input devices such as a keyboard, a virtual keyboard, a touchpad, a touch screen, and a voice control device, such as a projector, a smart phone, a display, etc. The present application does not make specific limitations on the specific functions performed by the terminal 002.
[0055] It can be understood that Figure 1The interpretable system architecture based on the dynamic behavior of the model is only a partial exemplary implementation in the embodiments of this application. The interpretable system architecture based on the dynamic behavior of the model in the embodiments of this application includes but is not limited to the above interpretable system architecture based on the dynamic behavior of the model.
[0056] Based on the above-provided interpretable system architecture based on the dynamic behavior of the model, the embodiments of this application provide an interpretable device based on the dynamic behavior of the model applied to the above interpretable system architecture based on the dynamic behavior of the model. This device is applicable to the above Figure 1 server 001. Please refer to Figure 2 , Figure 2 is a schematic structural diagram of an interpretable device based on the dynamic behavior of the model provided by the embodiments of this application. This interpretable device based on the dynamic behavior of the model can be equivalent to the above Figure 1 server 001 shown. Among them, in the interpretable device based on the dynamic behavior of the model, a dynamic boundary quantization module 01 and a model interpretation module 02 can be included. Among them, the dynamic boundary quantization module 01 includes: a quantization modeling module 101 and an optimization solving module 102; the model interpretation module 02 includes: a boundary preprocessing module 103 and a visualization module 104. For a given input instance and a neural network, the interpretable device based on the dynamic behavior of the model can first perform asymmetric dynamic quantization on the boundary of the quantized input instance, and then use this boundary to interpret the model prediction mechanism, and the features required or not required for the classification result can be obtained.
[0057] Please refer to the appendix Figure 3 , the appendix Figure 3 is a schematic diagram of the operation process of a dynamic boundary quantization module provided by the embodiments of this application. As Figure 3 shown, the dynamic boundary quantization module 01 first obtains the input data x and the target neural network model; secondly, according to the above input data x and the target neural network model, performs dynamic quantization modeling; uses an optimization method to solve the dynamic quantization modeling, and obtains the asymmetric dynamic boundary of the input data x; then, determines whether the obtained left dynamic boundary and right dynamic boundary meet the pre-set optimization conditions. If they do not meet, continue to solve. If they meet, obtain the asymmetric dynamic boundary quantization result of the input data x under the given target neural network model.
[0058] Among them, the quantization modeling module 101 is used to quantize the asymmetric dynamic quantization boundary of the input instance, that is, according to the target neural network and the input data to be input into the target neural network model, performs dynamic quantization modeling to obtain the asymmetric dynamic boundary of the input data. For example: define a solution equation including the largest space of x, Among them, That is, z (N+M) [t] - z (N+M) [j] ≥ δ, j = 1, …, K and j ≠ t. Wherein, in this equation, x is the input data, ∈1 is the left dynamic boundary, ∈2 is the right dynamic boundary, i is the current i-th layer in the target neural network model, n is the number of all pixel points in the input data (that is, the number of pixel points in the image data, and the value of n is as large as the actual number), N is the number of convolutional layers in the target neural network model, M is the number of fully connected layers in the target neural network model, t is the correct class in the classification result of the target neural network model, and j represents a class different from t.
[0059] The optimization solving module 102 can use an optimization method to solve the modeling in the above quantization modeling module 101 to obtain the corresponding asymmetric dynamic boundary. It can also judge whether the obtained left and right dynamic boundaries meet the pre-set optimization conditions. If not, continue to solve; if so, output the asymmetric dynamic boundary quantization result of the above input data under the given target neural network model.
[0060] Please refer to Appendix Figure 4 , Appendix Figure 4 is a schematic diagram of the operation process of a model interpretation module provided in an embodiment of the present application. As Figure 4 shown, the model interpretation module 02 can first use the asymmetric dynamic boundaries (left dynamic boundary, right dynamic boundary) obtained by quantifying the dynamic boundaries as inputs, and then perform differential preprocessing on the left and right dynamic boundaries to obtain differential dynamic boundaries; secondly, the model interpretation module 02 judges whether the asymmetric dynamic boundaries of each pixel point of the input data x have been processed. If so, output the differential dynamic boundaries of all pixel points of the input data x, otherwise continue to process; finally, the model decision mechanism of the target neural network model can be reflected by the differential dynamic boundaries, that is, generate a feature interpretation map based on the numerical size of the differential dynamic boundaries.
[0061] Among them, the boundary preprocessing module 103 is used to perform differential preprocessing on the left and right dynamic boundaries corresponding to each pixel point in the input data (that is: image data) to obtain differential dynamic boundaries. Among them, the differential preprocessing is to perform a differential operation on the left and right dynamic boundaries (such as: subtracting the right dynamic boundary value from the left dynamic boundary value), and then perform a normalization process on the difference to obtain a differential dynamic boundary value within a certain range. It should be noted that only the differential operation can be performed on the left and right dynamic boundaries without normalization. The normalization process is to make the features affecting the classification result of the input data in the target neural network model more distinct in the generated interpretation map in the next step, and at the same time, the small differences in the numerical sizes between them are also helpful for generating pictures.
[0062] The visualization module 104 obtains a target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries corresponding to the multiple pixel points and the image data. The visual perception brightness corresponding to a third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a fourth pixel point, where the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point. For example: as Figure 2 shown, normalizing the differential dynamic boundary to [0, 1], a larger value means that the predicted class (such as Figure 2 "7") tends to have this feature (such as the horizontal line above "7", that is, if this feature value becomes smaller, the class will change), otherwise, the predicted class tends to not have this feature (such as the horizontal line below "7", making it become "2"). Then visualize, the feature corresponding to a larger differential dynamic boundary value has a brighter visual perception, otherwise it is darker. Use this difference to reflect the decision-making mechanism of model classification.
[0063] It can be understood that Figure 2 the interpretability device based on the dynamic behavior of the model in
[0064] is only an exemplary implementation manner in the embodiments of the present application. The interpretability device based on the dynamic behavior of the model in the embodiments of the present application includes but is not limited to the above-mentioned interpretability device based on the dynamic behavior of the model.
[0064] Based on Figure 1 the provided interpretability system architecture based on the dynamic behavior of the model, and Figure 2 the provided structure of the interpretability device based on the dynamic behavior of the model, combined with the interpretability method based on the dynamic behavior of the model provided in the present application, specifically analyze and solve the technical problems proposed in the present application.
[0065] See Figure 5 Figure 5 is a schematic flowchart of an interpretability method based on the dynamic behavior of the model provided by the embodiments of the present application. This method can be applied to the interpretability system architecture based on the dynamic behavior of the model described in the above Figure 1 Among them, the interpretability device based on the dynamic behavior of the model can be used to support and execute Figure 5 the method flow steps S501 - step S504 shown in
[0066] Step S501: Obtain the asymmetric dynamic boundary of the target pixel point in the image data.
[0067] Specifically, the interpretability device based on the dynamic behavior of the model obtains the asymmetric dynamic boundary of the target pixel in the image data. The image data is the input data of the target neural network model. The target pixel is any one of multiple pixels in the image data. The asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary. The left dynamic boundary is the lower boundary of the pixel value of the target pixel, and the right dynamic boundary is the upper boundary of the pixel value of the target pixel. When the pixel value of the target pixel changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged. It should be noted that the left dynamic boundary can be the range limit where reducing the pixel value will not change the classification result, and the right dynamic boundary can be the range limit where increasing the pixel value will not change the classification result. Therefore, when the pixel value of the target pixel changes between the range of the asymmetric dynamic boundary, the classification result of the image data in the target neural network model remains unchanged. It should also be noted that the input data is explained by using the influence of the dynamic increase and decrease of the pixel value on the classification result, based on the dynamic behavior of the target neural network model. Therefore, it can better reflect the decision-making mechanism of the target neural network model. At the same time, the classification result of the target neural network model is used to ensure the fidelity of the interpretation method. Therefore, the effect of improving the transparency and interpretability of the neural network is achieved.
[0068] Optionally, the target neural network model of the interpretability device based on the dynamic behavior of the model includes a convolutional layer and / or a fully connected layer. Implementing the embodiments of the present application can explain the convolutional neural network model and reflect the dynamic behavior of the convolutional neural network model regarding input features. Among them, the target neural network model includes a convolutional layer and / or a fully connected layer, which is a convolutional neural network.
[0069] Step S502: Perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel to obtain the difference between the left and right dynamic boundaries of the target pixel.
[0070] Specifically, the interpretability device based on the dynamic behavior of the model performs a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel to obtain the difference between the left and right dynamic boundaries of the target pixel. It should be noted that the difference operation can subtract the pixel value of the target pixel corresponding to the right dynamic boundary from the pixel value of the target pixel corresponding to the left dynamic boundary.
[0071] Optionally, normalize the difference between the left and right dynamic boundaries of the target pixel to obtain the differential dynamic boundary value of the target pixel.
[0072] Step S503: Obtain the set of target pixels of the image data according to the differences between the left and right dynamic boundaries corresponding to multiple pixels in the image data.
[0073] Specifically, the interpretability device based on the dynamic behavior of the model obtains a set of target pixel points of the image data according to the difference between the left and right dynamic boundaries corresponding to the multiple pixel points respectively. The set of target pixel points includes one or more pixel points among the multiple pixel points whose difference between the left and right dynamic boundaries exceeds a first threshold. It should be noted that the set of target pixel points can be regarded as the dynamic key features of the input image data, which are used to determine the classification result of the image data under the target neural network model. For example, a vehicle is determined to be a vehicle based on tire features. Without tire features, it may not be recognized as a vehicle; a fan is determined to be a fan due to blade features. Without blade features, it may not be recognized as a fan; a phone is determined to be a phone due to button features and microphone features. Without any of the button features and microphone features, it may not be recognized as a fan. Therefore, the difference between the left and right dynamic boundaries of the pixel point sets corresponding to the above tire features, blade features, button features, and microphone features exceeds the first threshold, which can be regarded as the set of target pixel points, that is, the dynamic key features. Please refer to the appendix Figure 6 , Figure 6 is a schematic diagram of dynamic key features provided by an embodiment of the present application. As Figure 6 shown: There are two wheels under the "vehicle" in the image data. When the pixel values of the pixel points corresponding to the wheels gradually decrease, the wheels of the vehicle gradually disappear. At this time, the category of the vehicle may change and may be recognized as a house by the target neural network model, indicating that the classification result of the image data "vehicle" has changed. Therefore, the image data "vehicle" depends on the features of the pixel points where the two wheels are located under the "vehicle". And the set of pixel points corresponding to the wheels can be regarded as the set of target pixel points of the "vehicle", that is, the dynamic key features, which are used to determine the classification result of the image data under the target neural network model.
[0074] Optionally, the interpretability device based on the dynamic behavior of the model determines, from the difference dynamic boundary values corresponding to the multiple pixel points respectively, the pixel points whose difference dynamic boundary values exceed a second threshold as the set of target pixel points of the image data. By normalizing the difference between the left and right dynamic boundaries of the target pixel points, a relatively large difference value is processed into a smaller value, which is convenient for arithmetic processing and makes the difference in the influence of each pixel point on the classification result more obvious. For example, it is processed within the range of 0-1. Among them, the larger the difference dynamic boundary value of the target pixel point, the more the classification result of the image data in the target neural network model depends on the features of the target pixel point. For example, when the difference dynamic boundary is normalized within [0,1], a larger value means that the predicted class tends to have this feature for the input data (that is, if this feature value becomes smaller, the category will change). Otherwise, the predicted class tends to not have this feature.
[0075] Step S504: Obtain a target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries corresponding to multiple pixel points and the image data.
[0076] Specifically, the interpretability device based on the dynamic behavior of the model obtains the target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries corresponding to the multiple pixel points and the image data. The visual perception brightness corresponding to the third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to the fourth pixel point, where the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point. Implementing the embodiments of the present application, the decision-making mechanism of the target neural network model can be directly reflected by the difference between the left and right dynamic boundaries. Secondly, the extracted features are relatively simple and easy to understand. Generating the target feature map based on the magnitude of the difference between the left and right dynamic boundaries can more accurately discover the differential influence of each pixel point on the classification result. For example, since the visual perception brightness corresponding to the third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to the fourth pixel point, the embodiments of the present application can be applied to intelligent medical diagnosis to help doctors quickly analyze the medical image data of patients, that is, quickly determine the set of target pixel points (the visual perception brightness of the dynamic key features is greater) in the medical image data of patients, achieving the effect of accurately and quickly locating the lesion, and at the same time can also help establish a trust relationship between the patient and the intelligent medical diagnosis device.
[0077] Optionally, the interpretability device based on the dynamic behavior of the model obtains the target feature map corresponding to the image data according to the differential dynamic boundary values corresponding to the multiple pixel points and the image data. The visual perception brightness corresponding to the first pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to the second pixel point, where the differential dynamic boundary value of the first pixel point is greater than the differential dynamic boundary value of the second pixel point. Therefore, the decision-making mechanism of the target neural network model can be reflected after normalizing the differential dynamic boundary values. Secondly, generating the interpretation map, that is, the target feature map, based on the numerical magnitude of the differential dynamic boundary can more intuitively discover the differential influence of each pixel point on the classification result.
[0078] Optionally, the interpretability device based on the dynamic behavior of the model obtains the dynamic boundary range of the target pixel point set according to the difference between the left and right dynamic boundaries corresponding to the multiple pixel points respectively; according to the dynamic boundary range, the robustness boundary range of the target neural network model is obtained, and the robustness boundary range is within the dynamic boundary range and is used to determine the classification result after inputting into the target neural network model. Therefore, the left dynamic boundary is the range limit where reducing the pixel point value will not change the classification result, and the right dynamic boundary is the range limit where increasing the pixel point value will not change the classification result. The robustness boundary range is within the left and right dynamic boundary ranges. From the perspective of the robustness of the target neural network model, the robustness boundary range is equivalent to the dynamic boundary range corresponding to the image data, and both are used to determine the classification result after inputting into the target neural network model, and then the robustness boundary of the neural network model can be verified, thereby evaluating the security performance of the model. For example: it is used to analyze, verify, and evaluate the reliability and security of the convolutional neural network, thereby reducing the potential risks in the actual deployment of the convolutional neural network (such as: detecting adversarial samples from the inconsistency between human cognition and the model decision-making mechanism). Further, the embodiments of the present application can analyze the credibility of the neural network model's judgment and prompt the risks of the neural network model's judgment.
[0079] Aiming at the problems that the convolutional neural network model lacks transparency and interpretability, and the existing interpretability methods cannot accurately interpret the convolutional neural network model and reflect the dynamic behavior of the model regarding input features, etc., in the embodiments of the present application, for a given neural network model and image data (i.e., the input data of the neural network model), the asymmetric dynamic boundary of the image data under the neural network model can be obtained. The asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary, where the left dynamic boundary is the range limit where reducing the pixel point value will not change the classification result, and the right dynamic boundary is the range limit where increasing the pixel point value will not change the classification result. Therefore, a difference operation is performed on the left dynamic boundary and the right dynamic boundary to obtain the difference between the left and right dynamic boundaries. Finally, an analysis is performed according to the difference between the left and right dynamic boundaries to obtain the target pixel point set of the image data. Among them, the input data is image data. This method for each input data of the neural network model uses the dynamic behavior of the neural network model regarding the target pixel points in the input data (i.e., increasing or decreasing the pixel point value of the target pixel point without changing its classification result) to interpret the decision-making mechanism of the neural network model, thereby increasing the transparency of the neural network and establishing a trust relationship with users. In addition, using the method proposed in the present application, by deeply studying the internal working mechanism of the convolutional neural network, the robustness boundary of the neural network model can be verified, thereby evaluating the security performance of the model.
[0080] The method of the embodiments of the present application is elaborated in detail above, and the related devices of the embodiments of the present application are provided below.
[0081] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of another interpretability device based on the dynamic behavior of a model provided by an embodiment of the present application. The interpretability device based on the dynamic behavior of the model may include a first acquisition unit 701, a difference unit 702, and a pixel point unit 703, and may further include a normalization unit 704, a first feature map unit 705, a second feature map unit 706, and a second acquisition unit 707. The detailed descriptions of each unit are as follows.
[0082] The first acquisition unit 701 is configured to acquire the asymmetric dynamic boundary of a target pixel point in the image data. The image data is the input data of the target neural network model. The target pixel point is any one of multiple pixel points in the image data. The asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary. The left dynamic boundary is the lower boundary of the pixel value of the target pixel point, and the right dynamic boundary is the upper boundary of the pixel value of the target pixel point. When the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged;
[0083] The difference unit 702 is configured to perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point;
[0084] The pixel point unit 703 is configured to obtain a target pixel point set of the image data according to the differences between the left and right dynamic boundaries respectively corresponding to the multiple pixel points. The target pixel point set includes one or more pixel points among the multiple pixel points whose differences between the left and right dynamic boundaries exceed a first threshold.
[0085] In a possible implementation manner, the target neural network model includes a convolutional layer and / or a fully connected layer.
[0086] In a possible implementation manner, the device further includes: a normalization unit 704, configured to perform normalization processing on the difference between the left and right dynamic boundaries of the target pixel point to obtain a difference dynamic boundary value of the target pixel point; the pixel point unit 703 is specifically configured to: determine, from the difference dynamic boundary values respectively corresponding to the multiple pixel points, the pixel points whose difference dynamic boundary values exceed a second threshold as the target pixel point set of the image data.
[0087] In a possible implementation manner, the device further includes: a first feature map unit 705, configured to obtain a target feature map corresponding to the image data according to the differential dynamic boundary values respectively corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a first pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a second pixel point, and where the differential dynamic boundary value of the first pixel point is greater than the differential dynamic boundary value of the second pixel point.
[0088] In a possible implementation manner, the device further includes: a second feature map unit 706, configured to obtain a target feature map corresponding to the image data according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a fourth pixel point, and where the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point.
[0089] In a possible implementation manner, the device further includes: a second obtaining unit 707, configured to obtain a dynamic boundary range of the target pixel point set according to the difference between the left and right dynamic boundaries respectively corresponding to the multiple pixel points; a robustness unit 708, configured to obtain a robustness boundary range of the target neural network model according to the dynamic boundary range, where the robustness boundary range is within the dynamic boundary range and is used to determine a classification result after inputting into the target neural network model.
[0090] It should be noted that for the functions of the functional units in the interpretability device 70 based on the dynamic behavior of the model described in the embodiments of the present application, reference may be made to the relevant descriptions of steps S501 - S504 in the method embodiments described above Figure 5 and will not be elaborated here.
[0091] As Figure 8 shown, Figure 8 is a schematic structural diagram of another interpretability device based on the dynamic behavior of the model provided by the embodiments of the present application. The device 20 includes at least one processor 201, at least one memory 202, and at least one communication interface 203. In addition, the device may further include general components such as an antenna, which will not be elaborated here.
[0092] The processor 201 may be a general - purpose central processing unit (CPU), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the above - mentioned programs.
[0093] A communication interface 203 for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), core network, wireless local area networks (WLAN), etc.
[0094] The memory 202 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.
[0095] Among them, the memory 202 is used to store the application program code for executing the above solutions, and is controlled by the processor 201 for execution. The processor 201 is used to execute the application program code stored in the memory 202.
[0096] The code stored in the memory 202 can execute the above Figure 5The provided interpretability method based on the dynamic behavior of the model, for example: obtaining the asymmetric dynamic boundary of the target pixel point in the image data, where the image data is the input data of the target neural network model, the target pixel point is any one of multiple pixel points in the image data, the asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary, the left dynamic boundary is the lower boundary of the pixel value of the target pixel point, the right dynamic boundary is the upper boundary of the pixel value of the target pixel point, and when the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged; performing a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point; according to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively, obtaining a set of target pixel points of the image data, where the set of target pixel points includes one or more pixel points whose differences between the left and right dynamic boundaries exceed a first threshold.
[0097] It should be noted that for the interpretability device 20 based on the dynamic behavior of the model described in the embodiments of the present application, the functions of each functional unit can be referred to the relevant descriptions of steps S501 - S504 in the method embodiment described above Figure 5 and will not be elaborated here.
[0098] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0099] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be implemented in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0100] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0101] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0102] In addition, each functional unit in the embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0103] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute all or part of the steps of the above-mentioned methods in the various embodiments of the present application. Among them, the aforementioned storage medium may include: various media that can store program codes such as USB flash drives, mobile hard disks, magnetic disks, optical disks, read-only memory (ROM), or random access memory (RAM).
[0104] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. An interpretability method based on the dynamic behavior of a model, characterized in that, Including: Obtain the asymmetric dynamic boundary of the target pixel point in the image data. The image data is the input data of the target neural network model. The target pixel point is any one of multiple pixel points in the image data. The asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary. The left dynamic boundary is the lower boundary of the pixel value of the target pixel point, and the right dynamic boundary is the upper boundary of the pixel value of the target pixel point. When the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged; Perform a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point; According to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively, obtain the target pixel point set of the image data. The target pixel point set includes one or more pixel points among the multiple pixel points whose differences between the left and right dynamic boundaries exceed a first threshold; According to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively, obtain the dynamic boundary range of the target pixel point set; According to the dynamic boundary range, obtain the robustness boundary range of the target neural network model. The robustness boundary range is within the dynamic boundary range and is used to determine the classification result after inputting into the target neural network model.
2. The method according to claim 1, characterized in that, The target neural network model includes a convolutional layer and / or a fully connected layer.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Normalize the difference between the left and right dynamic boundaries of the target pixel point to obtain the differential dynamic boundary value of the target pixel point; The obtaining the target pixel point set of the image data according to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points in the image data includes: Determine, from the differential dynamic boundary values corresponding to the multiple pixel points respectively, that the pixel points whose differential dynamic boundary values exceed a second threshold are the target pixel point set of the image data.
4. The method according to claim 3, characterized in that, The method further includes: According to the differential dynamic boundary values corresponding to the multiple pixel points respectively and the image data, obtain the target feature map corresponding to the image data. Among the multiple pixel points in the target feature map, the visual perception brightness corresponding to the first pixel point is greater than the visual perception brightness corresponding to the second pixel point, where the differential dynamic boundary value of the first pixel point is greater than the differential dynamic boundary value of the second pixel point.
5. The method according to claim 1 or 2, characterized in that, The method further includes: According to the differences between the left and right dynamic boundaries corresponding to the multiple pixel points respectively and the image data, obtain the target feature map corresponding to the image data. Among the multiple pixel points in the target feature map, the visual perception brightness corresponding to the third pixel point is greater than the visual perception brightness corresponding to the fourth pixel point, where the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point.
6. An interpretability device based on the dynamic behavior of a model, characterized in that, Including: A first acquisition unit for acquiring an asymmetric dynamic boundary of a target pixel point in image data, where the image data is the input data of a target neural network model, the target pixel point is any one of multiple pixel points in the image data, the asymmetric dynamic boundary includes a left dynamic boundary and a right dynamic boundary, the left dynamic boundary is the lower boundary of the pixel value of the target pixel point, the right dynamic boundary is the upper boundary of the pixel value of the target pixel point, and when the pixel value of the target pixel point changes between the left dynamic boundary and the right dynamic boundary, the classification result of the image data in the target neural network model remains unchanged; A difference unit for performing a difference operation on the left dynamic boundary and the right dynamic boundary of the target pixel point to obtain the difference between the left and right dynamic boundaries of the target pixel point; A pixel point unit for obtaining a set of target pixel points of the image data according to the differences between the left and right dynamic boundaries respectively corresponding to the multiple pixel points, where the set of target pixel points includes one or more pixel points among the multiple pixel points whose differences between the left and right dynamic boundaries exceed a first threshold; A second acquisition unit for obtaining the dynamic boundary range of the set of target pixel points according to the differences between the left and right dynamic boundaries respectively corresponding to the multiple pixel points; A robustness unit for obtaining a robustness boundary range of the target neural network model according to the dynamic boundary range, where the robustness boundary range is within the dynamic boundary range and is used to determine the classification result after inputting into the target neural network model.
7. The device according to claim 6, characterized in that, The target neural network model includes a convolutional layer and / or a fully connected layer.
8. The device according to claim 6 or 7, characterized in that, The device further includes: A normalization unit for normalizing the difference between the left and right dynamic boundaries of the target pixel point to obtain a difference dynamic boundary value of the target pixel point; The pixel point unit is specifically used for: Determining, from the difference dynamic boundary values respectively corresponding to the multiple pixel points, the pixel points whose difference dynamic boundary values exceed a second threshold as the set of target pixel points of the image data.
9. The device according to claim 8, characterized in that, The device further includes: A first feature map unit for obtaining a target feature map corresponding to the image data according to the difference dynamic boundary values respectively corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a first pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a second pixel point, and among them, the difference dynamic boundary value of the first pixel point is greater than the difference dynamic boundary value of the second pixel point.
10. The device according to claim 6 or 7, characterized in that, The device further includes: A second feature map unit for obtaining a target feature map corresponding to the image data according to the differences between the left and right dynamic boundaries respectively corresponding to the multiple pixel points and the image data, where the visual perception brightness corresponding to a third pixel point among the multiple pixel points in the target feature map is greater than the visual perception brightness corresponding to a fourth pixel point, and among them, the difference between the left and right dynamic boundaries of the third pixel point is greater than the difference between the left and right dynamic boundaries of the fourth pixel point.
11. A chip system, characterized in that, The chip system includes at least one processor, a memory, and an interface circuit. The memory, the interface circuit, and the at least one processor are interconnected by lines, and instructions are stored in the at least one memory. When the instructions are executed by the processor, the method described in any one of claims 1-5 is implemented.
12. A computer storage medium, characterized in that, The computer storage medium stores a computer program which, when executed by a processor, implements the method described in any one of claims 1-5 above.
13. A computer program, characterized in that, The computer program includes instructions which, when executed by a computer, cause the computer to execute the method described in any one of claims 1-5.