Methods, devices, and storage media for interpreting model inference results
By performing cluster partitioning and saliency graph processing on the inference results of AI models, the problems of long processing time and poor interpretation effect in existing technologies are solved, realizing automatic, fast and accurate interpretation of inference results, and supporting model improvement and visualization.
Patent Information
- Application Number
- CN202110586245.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-05-27
AI Technical Summary
Existing technologies are time-consuming and produce poor results when interpreting the inference results of AI models, failing to meet user needs.
By grouping image samples into clusters based on their labels and automatically identifying explanatory objects using saliency maps and contrast saliency maps, comparative interpretation is achieved, common features are eliminated, and dissimilar features are highlighted for visualization and quantitative evaluation.
It enables automatic, fast, and accurate interpretation of reasoning results, improves the richness and visualization of interpretations, supports model improvement and debugging, and enhances the value and accuracy of interpretations.
Smart Images

Figure CN115482424B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for interpreting model reasoning results. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0003] AI models are typically black-box models, meaning users often know what happens but not why. Currently, related technologies are time-consuming and produce poor explanations when interpreting the inference results of AI models, which fails to meet relevant needs. Summary of the Invention
[0004] In view of this, a method, apparatus and storage medium for interpreting model reasoning results are proposed.
[0005] In a first aspect, embodiments of this application provide a method for interpreting model inference results. The method includes: obtaining at least one first label, where the first label represents the inference result of an image sample; dividing the at least one first label into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels within the same cluster have a higher probability of appearing in the same image sample compared to the first labels within other clusters; determining a first contrast saliency map corresponding to each cluster based on each first saliency map corresponding to each cluster, or determining a second contrast saliency map corresponding to each first label within a cluster based on each first saliency map corresponding to each cluster and each first saliency map corresponding to each first label within a cluster, wherein the first saliency map represents each cluster or each first label within a cluster as the basis for the inference result, the first contrast saliency map represents each cluster as the basis for the inference result relative to other clusters, and the second contrast saliency map represents each first label within a cluster as the basis for the inference result relative to other first labels within the cluster.
[0006] According to the embodiments of this application, by obtaining at least one first label, dividing the at least one first label into at least one cluster, and determining the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster, or by determining the second contrast saliency map corresponding to each first label within a cluster based on the first saliency map corresponding to each cluster and the first saliency map corresponding to each first label within a cluster, the objects for contrast interpretation can be automatically determined without manual selection, saving a lot of resources. Moreover, all first labels can be contrast interpreted, making the interpretation more complete, which is more helpful for improving and debugging the inference model, making the inference results more accurate. The contrast interpretation between clusters can solve the problem of interpretation of multiple subject objects and can be applied to multi-label image scenarios. Through the contrast interpretation of first labels within a cluster, and since the first labels within the same cluster have a higher probability of appearing in the same image sample compared to the first labels within other clusters, more refined distinction of key features of labels within a cluster can be achieved, making the contrast interpretation results more valuable. By displaying the contrast saliency map, the results of the contrast interpretation can be systematically and concisely visualized.
[0007] According to the first aspect, in a first possible implementation of the method for interpreting the model inference results, the highlighted areas of the first contrast saliency map represent each cluster as the basis for the inference results relative to other clusters, and the highlighted areas of the second contrast saliency map represent each first label within a cluster as the basis for the inference results relative to other first labels within the cluster.
[0008] According to the embodiments of this application, by using the highlighted areas of the first contrast saliency map to represent each cluster relative to other clusters as the basis for the inference result, and the highlighted areas of the second contrast saliency map to represent each first label within a cluster relative to other first labels within the cluster as the basis for the inference result, a more intuitive interpretation of the inference result can be achieved, and the interpretation content can be made richer, which helps to improve and debug the inference model and obtain a more accurate inference model.
[0009] According to the first aspect or the first possible implementation of the first aspect, in the second possible implementation of the method for interpreting the model inference results, determining the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster includes: eliminating common features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, determining the differential features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, and determining the first contrast saliency map corresponding to each cluster based on the differential features; determining the second contrast saliency map corresponding to each first label within a cluster based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within a cluster includes: determining the image region corresponding to the cluster indicated by the first saliency map of the cluster corresponding to each first label, eliminating common features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster within the image region, determining the differential features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster, and determining the second contrast saliency map corresponding to each first label within a cluster based on the differential features.
[0010] According to the embodiments of this application, by eliminating common features and determining the saliency map based on the difference features, the comparison interpretation results can be displayed with fine-grained differences between the comparison subject and the comparison object, making the interpretation more accurate and specific. By determining the corresponding image region of the cluster indicated by the first saliency map of each first label, the comparison interpretation results of the labels within the same cluster can be limited to the same area in the image, making the comparison interpretation more accurate.
[0011] According to the first aspect or the first or second possible implementation of the first aspect, in the third possible implementation of the method for interpreting the model inference results, dividing the at least one first label into at least one cluster includes: dividing the at least one second label into at least one cluster, wherein the second label represents the inference result of all image samples, each cluster includes at least one of the second labels, and the second labels in the same cluster have a higher probability of appearing in the same image sample than the second labels in other clusters; and dividing the first labels in the at least one first label that are in the same cluster as the corresponding second labels into one cluster according to the correspondence between the at least one first label and the at least one second label.
[0012] According to the embodiments of this application, by dividing at least one second label into at least one cluster, and according to the correspondence between the at least one first label and the at least one second label, the first labels in the at least one first label that correspond to the second label in the same cluster are divided into one cluster, the objects of comparison and interpretation can be automatically determined without manual selection, saving a lot of resources, and making easily confused labels into one cluster, thereby making the results of comparison and interpretation more referential and valuable.
[0013] According to the first aspect or the first, second or third possible implementation of the first aspect, in a fourth possible implementation of the method for interpreting the model inference result, the method further includes: determining a first order corresponding to pixels in the image sample based on the first contrast saliency map; sequentially erasing a predetermined number of pixels in the image sample based on the first order; determining an evaluation corresponding to the first contrast saliency map based on at least one image sample obtained during the pixel erasure process; and / or determining a second order corresponding to pixels in the image sample based on the second contrast saliency map; sequentially erasing a predetermined number of pixels in the image sample based on the second order; and determining an evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process.
[0014] According to the embodiments of this application, by determining the order of pixels in the image sample based on the saliency map, and then sequentially erasing a predetermined number of pixels in the image sample according to the order, and determining the evaluation corresponding to the saliency map based on at least one image sample obtained during the pixel erasure process, the results of the comparative interpretation can be quantitatively evaluated, thereby intuitively demonstrating the value of the comparative interpretation content.
[0015] According to the fourth possible implementation of the first aspect, in the fifth possible implementation of the method for interpreting the model inference result, the first sorting is determined based on the brightness of each pixel in the first contrast saliency map, and the second sorting is determined based on the brightness of each pixel in the second contrast saliency map.
[0016] According to the embodiments of this application, a more accurate evaluation of the comparative interpretation results indicated by the saliency map can be achieved.
[0017] According to the fourth or fifth possible implementation of the first aspect, in the sixth possible implementation of the method for interpreting the model inference result, determining the evaluation corresponding to the first contrast saliency map based on at least one image sample obtained during the pixel erasure process includes: determining a first prediction probability of the cluster corresponding to the first contrast saliency map and a second prediction probability of other clusters besides the first cluster based on at least one image sample obtained during the pixel erasure process, wherein the first prediction probability represents the probability that the inference result of the image sample is the corresponding cluster, and the second prediction probability represents the probability that the inference result of the image sample is other clusters besides the first cluster; determining the evaluation corresponding to the first contrast saliency map based on the first prediction probability and the second prediction probability. Evaluation of the first contrast saliency map; determining the evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process, including: determining the third prediction probability of the first label within the cluster corresponding to the second contrast saliency map and the fourth prediction probability of other first labels within the cluster other than the first label within the cluster based on at least one image sample obtained during the pixel erasure process, wherein the third prediction probability represents the probability that the inference result of the image sample is the corresponding first label within the cluster, and the fourth prediction probability represents the probability that the inference result of the image sample is other first labels other than the first label within the cluster; determining the evaluation of the second contrast saliency map based on the third prediction probability and the fourth prediction probability.
[0018] According to the embodiments of this application, when conducting evaluations, the degree of support of the explanatory basis presented in the comparative explanation results for the comparative subject and the degree of denial of the comparative object can be taken into account, thereby comprehensively measuring the loyalty of the comparative explanation results and making the evaluation results more accurate.
[0019] Secondly, embodiments of this application provide an apparatus for interpreting model inference results. The apparatus includes: an acquisition module for acquiring at least one first label, where the first label represents the inference result of an image sample; a first determination module for dividing the at least one first label into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels within the same cluster have a higher probability of appearing in the same image sample compared to the first labels within other clusters; and a second determination module for determining a first contrast saliency map corresponding to each cluster based on the first saliency maps corresponding to each cluster, or determining a second contrast saliency map corresponding to each first label within a cluster based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within the cluster, wherein the first saliency map represents each cluster or each first label within a cluster as the basis for the inference result, the first contrast saliency map represents each cluster as the basis for the inference result relative to other clusters, and the second contrast saliency map represents each first label within a cluster as the basis for the inference result relative to other first labels within the cluster.
[0020] According to the second aspect, in a first possible implementation of the device for interpreting the model inference results, the highlighted areas of the first contrast saliency map represent each cluster as the basis for the inference result relative to other clusters, and the highlighted areas of the second contrast saliency map represent each first label within a cluster as the basis for the inference result relative to other first labels within the cluster.
[0021] According to the second aspect or the first possible implementation of the second aspect, in the second possible implementation of the device for interpreting the model inference results, determining the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster includes: eliminating common features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, determining the differential features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, and determining the first contrast saliency map corresponding to each cluster based on the differential features; determining the second contrast saliency map corresponding to each first label within a cluster based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within a cluster includes: determining the image region corresponding to the cluster indicated by the first saliency map of the cluster corresponding to each first label, eliminating common features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster within the image region, determining the differential features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster, and determining the second contrast saliency map corresponding to each first label within a cluster based on the differential features.
[0022] According to the second aspect or the first or second possible implementation of the second aspect, in a third possible implementation of the device for interpreting the model inference results, dividing the at least one first label into at least one cluster includes: dividing the at least one second label into at least one cluster, wherein the second label represents the inference result of all image samples, each cluster includes at least one of the second labels, and the second labels in the same cluster have a higher probability of appearing in the same image sample than the second labels in other clusters; and dividing the first labels in the at least one first label that correspond to the at least one second label into a cluster according to the correspondence between the at least one first label and the at least one second label.
[0023] According to the second aspect or the first, second or third possible implementation of the second aspect, in a fourth possible implementation of the device for interpreting the model inference result, the device further includes: a third determining module, configured to determine a first order corresponding to pixels in the image sample based on the first contrast saliency map; a first pixel erasure module, configured to sequentially erasure a predetermined number of pixels in the image sample based on the first order; a fourth determining module, configured to determine an evaluation corresponding to the first contrast saliency map based on at least one image sample obtained during the pixel erasure process; and / or a fifth determining module, configured to determine a second order corresponding to pixels in the image sample based on the second contrast saliency map; a second pixel erasure module, configured to sequentially erasure a predetermined number of pixels in the image sample based on the second order; and a sixth determining module, configured to determine an evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process.
[0024] According to the fourth possible implementation of the second aspect, in the fifth possible implementation of the interpretation device for the model inference result, the first sorting is determined based on the brightness of each pixel in the first contrast saliency map, and the second sorting is determined based on the brightness of each pixel in the second contrast saliency map.
[0025] According to the fourth or fifth possible implementation of the second aspect, in the sixth possible implementation of the device for interpreting the model inference result, determining the evaluation corresponding to the first contrast saliency map based on at least one image sample obtained during the pixel erasure process includes: determining a first prediction probability of the cluster corresponding to the first contrast saliency map and a second prediction probability of other clusters besides the first cluster based on at least one image sample obtained during the pixel erasure process, wherein the first prediction probability represents the probability that the inference result of the image sample is the corresponding cluster, and the second prediction probability represents the probability that the inference result of the image sample is other clusters besides the first cluster; determining the evaluation corresponding to the first contrast saliency map based on the first prediction probability and the second prediction probability. Evaluation of the first contrast saliency map; determining the evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process, including: determining the third prediction probability of the first label within the cluster corresponding to the second contrast saliency map and the fourth prediction probability of other first labels within the cluster other than the first label within the cluster based on at least one image sample obtained during the pixel erasure process, wherein the third prediction probability represents the probability that the inference result of the image sample is the corresponding first label within the cluster, and the fourth prediction probability represents the probability that the inference result of the image sample is other first labels other than the first label within the cluster; determining the evaluation of the second contrast saliency map based on the third prediction probability and the fourth prediction probability.
[0026] Thirdly, embodiments of this application provide an apparatus for interpreting model inference results, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to, when executing the instructions, implement one or more of the model inference result interpretation methods described in the first aspect or multiple possible implementations of the first aspect.
[0027] Fourthly, embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement a method for interpreting model reasoning results from one or more of the first aspect or various possible implementations of the first aspect.
[0028] Fifthly, embodiments of this application provide a terminal device that can execute one or more of the model reasoning result interpretation methods described in the first aspect or various possible implementations of the first aspect.
[0029] Sixthly, embodiments of this application provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in an electronic device, the processor in the electronic device executes one or more of the model reasoning result interpretation methods of the first aspect or various possible implementations of the first aspect.
[0030] These and other aspects of this application will become more apparent in the description of the following embodiments(s). Attached Figure Description
[0031] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this application together with the specification and serve to explain the principles of this application.
[0032] Figure 1 This diagram illustrates an application scenario of a method for interpreting model reasoning results according to an embodiment of this application.
[0033] Figure 2 A structural diagram of an apparatus for interpreting model reasoning results according to an embodiment of this application is shown.
[0034] Figure 3 A flowchart illustrating the comparative interpretation stage of a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0035] Figure 4 A schematic diagram illustrating the clustering of tags according to an embodiment of this application is shown.
[0036] Figure 5A schematic diagram of a hierarchical tree according to an embodiment of this application is shown.
[0037] Figure 6 This diagram illustrates the determination of the cluster corresponding to the label in an image sample according to an embodiment of the present application.
[0038] Figure 7 A schematic diagram showing a contrast saliency diagram according to an embodiment of this application is provided.
[0039] Figure 8 A flowchart is shown of the evaluation phase of a method for interpreting model reasoning results according to an embodiment of this application.
[0040] Figure 9 A schematic diagram of erasing pixels from an image sample according to an embodiment of this application is shown.
[0041] Figure 10 A schematic diagram illustrating the determination of comparative loyalty according to an embodiment of this application is shown.
[0042] Figure 11 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0043] Figure 12 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0044] Figure 13 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0045] Figure 14 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0046] Figure 15 A structural diagram of an apparatus for interpreting model reasoning results according to an embodiment of this application is shown.
[0047] Figure 16 A structural diagram of an apparatus for interpreting model reasoning results according to an embodiment of this application is shown.
[0048] Figure Labels
[0049] Highlighted area H Detailed Implementation
[0050] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0051] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0052] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.
[0053] Figure 1 This diagram illustrates an application scenario of a method for interpreting model inference results according to an embodiment of this application. The method for interpreting model inference results according to this application can be applied to data annotation systems, such as... Figure 1 As shown, the data annotation system may include an interpretation device and a model inference device. The model inference device, based on a deep learning model, is used to perform inference and prediction on the data, obtaining the annotation results (i.e., inference results). The interpretation device can be used to implement the interpretation method of the model inference results in this embodiment, and can be used to compare and interpret the annotation results of the data. In one possible implementation, the interpretation device can also be used to quantitatively evaluate the interpretation results. The data annotation system can interact with a database, which can be used to store data and provide the necessary data to the model inference device. As shown, through the data annotation system, the model in the model inference device can be trained. Based on the annotation results of the data by the model inference device, recommendation services can be provided to users. Through the interpretation device, the annotation results of the model inference can be interpreted, and the interpretation results can also be visualized.
[0054] For example, in scenarios where images are annotated, the model inference device can infer the corresponding labels for the images, and the interpretation device can compare and interpret the inferred labels. For instance, for a given image, it can explain the basis for determining that the image is labeled A rather than B / C, or it can explain that the image is labeled B rather than A / C. Furthermore, it can display a saliency map to the user based on the comparison and interpretation results, allowing for a more intuitive explanation of the annotation results. The saliency map can display the interpretation result of a certain label in the image in a specific form, such as a highlighted area. For example, the highlighted area in the saliency map can indicate the basis for the image being labeled A. Similarly, the comparison saliency map can display the basis for the comparison and interpretation results of a certain label relative to other labels in the image in a specific form, such as a highlighted area. For example, the highlighted area in the comparison saliency map can indicate the basis for the image being labeled A rather than B / C.
[0055] Through the embodiments of this application, cognitively friendly explanations can be provided to users for the annotation results of the data, so as to meet the legal requirements for explainable AI. At the same time, it can help users understand the model. Based on the explanation results, users can also optimize the model in the model inference device to achieve more accurate reasoning results.
[0056] It should be noted that the above-described scenario of annotating images is only one example of this application. The embodiments of this application can also be used in scenarios other than those described above, and this application does not impose any limitations on them.
[0057] Figure 2 A structural diagram of an apparatus for interpreting model reasoning results according to an embodiment of this application is shown. Figure 2 As shown, the interpretation device can interact with the database and the model inference device to interpret the model inference results. The interpretation device may include a user interface (UI), a comparative interpretation module, a basic interpreter, and an evaluation module. The basic interpreter can be used to interpret the inference results using a saliency map. The comparative interpretation module can be used to perform comparative interpretation of the inference results based on the saliency map interpretation results. The UI can be used to visualize the comparative interpretation results. The evaluation module can be used to perform quantitative evaluation of the comparative interpretation results.
[0058] It should be noted that the structure of the above-described explanatory device is only an example. The explanatory device may include only some of its modules or may include more modules. This application does not limit this.
[0059] The following combination Figures 3-10 Taking the comparison and interpretation of the annotation results of image samples using the above-mentioned device as an example, the flow of the method for interpreting model inference results provided in this application embodiment is explained.
[0060] The process of interpreting the model reasoning results in this application embodiment may include a comparison and interpretation stage and an evaluation stage.
[0061] Figure 3 A flowchart illustrating the comparative interpretation stage of a method for interpreting model reasoning results according to an embodiment of this application is shown.
[0062] like Figure 3 As shown, the process of the comparative explanation phase may include:
[0063] Step S301: The comparison and interpretation module clusters all labels in the label set based on the probability co-occurrence relationship to obtain cluster classification information.
[0064] The co-occurrence relationship can be the probability relationship between all labels in the label set that they co-occur in the same image. The label set can represent the set of labels corresponding to all image samples in the dataset (i.e., the inference results of all image samples). For example, the model inference device can infer the corresponding label for each image sample in the dataset, and the comparison and interpretation module can determine L*N high-probability labels in the label set. The L high-probability labels can be the L labels with the highest predicted probability among the labels corresponding to each inferred image sample, or labels with predicted probabilities exceeding a threshold (e.g., L labels). N can represent the number of image samples in the dataset. For N image samples in the dataset, L*N high-probability labels can be obtained. The comparative interpretation module can cluster all labels in the label set based on the output distribution of these L*N labels under the model inference device (i.e., in which image samples the labels appear). Specifically, if two labels frequently appear in the same image sample (or have a high probability of appearing in the same image sample), they can be considered indistinguishable and grouped into the same cluster; if two labels do not frequently appear in the same image sample, they can be considered easily distinguishable and grouped into different clusters. For example, the output distribution of the L*N labels under the model inference device can be used to determine whether two labels in the label set frequently appear in the same image sample through learning and fitting methods.
[0065] The size of L and the threshold used to determine whether a label is of high probability can be determined as needed, and this application does not impose any restrictions on them. The learning and fitting method can be, for example, learning and fitting all labels in the label set based on a hierarchical latent tree model (HLTM), or it can utilize other models, and this application does not impose any restrictions on them.
[0066] The following is Figure 4 , Figure 5 The process of S301 will be illustrated using an example. Figure 4 A schematic diagram illustrating the clustering of tags according to an embodiment of this application is shown. For example... Figure 4 As shown in (a), taking a dataset with 9 image samples as an example, the inference model in the model inference device can be used to infer the corresponding label for each image sample. The figure uses ResNet50 as the inference model, but other models, such as GoogleNet, can also be used; this application does not limit this. Figure 4(a) The labels for each image are listed in descending order of predicted probability. For image 1, the labels obtained through the inference model can include cello, robin, and violin (the predicted probabilities decrease sequentially; the same applies to images 2-9). For image 2, the labels obtained through the inference model can include cello, violin, and oboe. For image 3, the labels obtained through the inference model can include cello, violin, robe, and trombone. For image 4, the labels obtained through the inference model can include acoustic guitar and electric guitar. For image 5, the labels obtained through the inference model can include acoustic guitar and electric guitar. For image 6, the labels obtained through the inference model can include acoustic guitar, electric guitar, and banjo. For image 7, the labels obtained through the inference model can include electric guitar and acoustic guitar. For image 8, the labels obtained through the inference model can include electric guitar, acoustic guitar, and banjo. For image 9, the labels obtained through the inference model can include electric guitar and acoustic guitar. The top L high-probability labels in each image sample can be determined. When L is set to 2, 18 labels in the inferred sample can be determined. In one possible implementation, a probability threshold for the predicted probability can also be set as a condition for selecting labels, such as using labels with a corresponding predicted probability greater than 95% as the selected labels. This application does not limit this.
[0067] The contrast interpretation module can learn and fit all the labels in the label set using HLTM based on the distribution relationship of these 18 labels in the image samples (i.e., the output distribution under the model inference device), to obtain the following result: Figure 4 (b) shows a hierarchical tree based on probabilistic co-occurrence relationships. Among them, Figure 4 (b) This example only shows the probability co-occurrence relationship of some labels in the label set. The hierarchical tree can also include other labels in the label set besides violin, cello, acoustic guitar, electric guitar, and banjo, such as robin, trombone, robe, oboe, etc. In the hierarchical tree, labels under the same node can represent labels with high probability co-occurrence (i.e., high probability of appearing in the same image sample). Z66, Z1396, Z2138, and Z1284 can represent implicit node names. Clustering can be performed based on the probability co-occurrence relationship represented by the hierarchical tree. For example, a certain level can be selected as the basis for clustering. Labels under the same parent node can be labels within the same cluster. The shorter the path between two nodes, the higher the probability that the labels corresponding to the two nodes appear in the same image sample. Therefore, the labels obtained within the same cluster can be easily confused labels.
[0068] In this application, it is possible to determine which level in the hierarchical tree to use as the basis for clustering, and this application does not impose any restrictions on this.
[0069] For example, the labels corresponding to Z1396 can be considered as one cluster, and the labels within this cluster can include "cello" and "violin," indicating that the labels "cello" and "violin" are highly probable to appear in the same image sample. Similarly, the labels corresponding to Z2138 can be considered as another cluster, and the labels within this cluster can include "electric guitar," "acoustic guitar," and "banjo," indicating that the labels "electric guitar," "acoustic guitar," and "banjo" are highly probable to appear in the same image sample. Z2138 can also include two child nodes: Z1284 and "banjo," indicating that within Z2138, the labels "electric guitar" and "acoustic guitar" have a higher probability of appearing in the same image sample compared to the label "banjo."
[0070] The following is an example of using image samples from the ImageNet label set, inferring labels from the image samples using ResNet50, and determining the hierarchical tree corresponding to all labels in the label set based on the output distribution of the top L*N high-probability labels under the model inference device. Figure 5 A schematic diagram of a hierarchical tree according to an embodiment of this application is shown.
[0071] like Figure 5 As shown, the image samples used can come from the ImageNet label set, and the inference model can be ResNet50. Figure 5 This example uses only the cello, violin, flute, clarinet, electric guitar, acoustic guitar, banjo, accordion, bassoon, saxophone, trombone, and cornet. The hierarchical tree can also include other tags from the ImageNet tag set. For example... Figure 5 In the hierarchical tree shown, each node has labels that are highly co-occurring, and the corresponding image samples are below the labels. For example, cello and violin are in the same node, and flute and clarinet are in the same node.
[0072] See again Figure 3 After obtaining the cluster classification information, in step S302, the model inference device determines the top K high-probability labels corresponding to the input image sample, and the comparison and interpretation module divides these K high-probability labels into m clusters according to the cluster classification information.
[0073] Figure 6 This diagram illustrates the determination of clusters corresponding to labels in an image sample according to an embodiment of this application. Figure 6 As shown, Figure 6 The image samples in (a) can be, for example... Figure 6 (b) The leftmost image can be inferred using a ResNet50 model. The first K (K=5) labels determined based on the inference results of the inference model can be, for example... Figure 6(b) The cello, violin, acoustic guitar, five-stringed guitar, and electric guitar shown in the middle and middle images can be clustered by the comparison and interpretation module based on the cluster classification information obtained in S301. For example, if the cello and violin are grouped into one cluster and the guitar, five-stringed guitar, and electric guitar are grouped into another cluster in S301, the five labels can be considered to correspond to each other based on the correspondence between the five labels and the labels used for clustering in S301. If the content of the five labels and the labels used for clustering in S301 are the same, they can be considered to correspond. For example, if the content of the cello in the five labels is the same as the content of the cello in the labels used for clustering in S301, they can be considered to correspond. Similarly, based on the clustering situation in S301, the five labels corresponding to this image sample can be divided into two clusters, m1 and m2. m1 includes cello and violin, and m2 includes acoustic guitar, five-stringed guitar, and electric guitar.
[0074] In this application, K can be set as a quantity threshold to determine the top K tags with the highest probability among the tags inferred by the model inference device as the tags used for clustering. Alternatively, a probability threshold can be set to select tags with a probability higher than the probability threshold (e.g., K tags) as the tags used for clustering. This application does not impose any restrictions on this.
[0075] In step S303, the basic interpreter performs saliency map interpretation on the K classes and m clusters in the determined image samples to obtain the corresponding saliency maps.
[0076] Here, the K classes can correspond to the K labels determined in the image samples, and the m clusters can be the clusters determined based on the K labels in S302. The basic interpreter can utilize gradient-weighted class activation mapping (GradCAM) or interpretation based on randomized input sampling. The RISE (Signal Analysis and Interpretation) method interprets K classes and m clusters respectively, resulting in K+m saliency maps. For the saliency maps corresponding to the K classes, the highlighted areas in each saliency map explain why the inference result belongs to that class. That is, the highlighted areas in the saliency map are different from other areas, which can represent the basis for the model inference device to determine that the inference result belongs to that class. For example, in the case that class K1 is cello, the highlighted areas in the saliency map corresponding to K1 explain why the inference result is cello. For the saliency maps corresponding to the m clusters, the highlighted areas in each saliency map explain why the inference result belongs to that cluster. That is, the highlighted areas in the saliency map are different from other areas, which can represent the basis for the model inference device to determine that the inference result belongs to that cluster. For example, in the case that cluster m1 includes cello and violin, the highlighted areas in the saliency map corresponding to m1 explain why the inference result is cello and violin.
[0077] It should be noted that, in addition to GradCAM and RISE, other methods can be used to obtain saliency maps, and this application does not impose any restrictions on them.
[0078] Step S304: The comparison and interpretation module performs inter-cluster comparison and inter-class comparison within the same cluster based on the saliency map, and obtains the comparison and interpretation results between clusters and between classes within the same cluster.
[0079] Among them, inter-cluster comparison and inter-class comparison within the same cluster can be performed by eliminating common features in the saliency maps between clusters or between classes within the same cluster and highlighting the differences. For example, the inter-cluster comparison interpretation result can be obtained by the following formula (1), and the inter-class comparison interpretation result within the same cluster can be obtained by the following formulas (2) and (3):
[0080]
[0081] in, This can represent the contrastive interpretation result corresponding to the i-th cluster, where m can represent the number of clusters, ReLU can represent the corrected linear unit, and H... i H can represent the saliency map corresponding to the i-th cluster. i′ This can represent the saliency map corresponding to the i′-th cluster, where i and i′ correspond to different clusters among m clusters, and H i and H i′ It can be obtained from S303. Where H... i -H i′ The operation can eliminate H i and H i′ The common features between them determine H i and H i′ The differences between them.
[0082]
[0083] in, The support function can be obtained by formula (3) to represent the comparative interpretation result corresponding to the j-th class within the i-th cluster. It represents the support function, where δ is a predetermined parameter, and J... i H can represent the number of classes within the i-th cluster. ij H can represent the saliency map corresponding to the j-th class within the i-th cluster. ij′ This can represent the saliency map corresponding to the j′ class within the i-th cluster, where j and j′ correspond to different classes within the i-th cluster, and H... ij and H ij′ It can be obtained from S303. Where H... ij -H ij′ The operation can eliminate Hij and H ij′ The common features between them determine H ij and H ij′ The differences between them.
[0084]
[0085] in The support function can be represented by the coordinates (a, b), where, This can represent the value of the pixel at coordinates (a, b) in the saliency map corresponding to the i-th cluster's contrast interpretation result. By multiplying the linear unit by the support function in formula (2), the contrast interpretation results between labels within the same cluster can be limited to the same region in the image sample, that is, it can make... Limited to the range (image region) within the image samples corresponding to the i-th cluster. The predefined parameter δ in the support function can be used to filter pixels with values lower than [a certain value] in the corresponding saliency map. Pixels.
[0086] It should be noted that the comparison interpretation results between clusters and between classes within the same cluster can also be determined in other ways besides those shown in formulas (1)-(3), and this application does not impose any restrictions on this.
[0087] In step S305, the UI displays saliency maps based on the comparison interpretation results between clusters and between classes within the same cluster.
[0088] The calculation result can be obtained from formula (1). The determined pixel values are displayed to obtain the saliency map corresponding to the i-th cluster (an example of the first saliency map below), which can be obtained by formula (2). The determined pixel values are displayed to obtain the contrast saliency map corresponding to the j-th class within the i-th cluster (an example of the second contrast saliency map below).
[0089] Figure 7 A schematic diagram showing a contrast saliency diagram according to an embodiment of this application is provided. Figure 7 As shown, with Figure 6 Taking the leftmost image in (b) as an example, we can obtain 5 labels (cello, violin, acoustic guitar, banjo, electric guitar). These 5 labels can correspond to 5 classes respectively, and the 5 labels are divided into 2 clusters. The first cluster can include cello and violin, and the second cluster can include acoustic guitar, banjo, electric guitar. Figure 7 (a) and Figure 7 (b) can represent the comparative interpretation results between clusters. Figure 7 (a.1) Figure 7 (a.2) Figure 7 (b.1) Figure 7 (b.2) Figure 7 (b.3) can represent the comparative interpretation results within the same cluster and between classes.
[0090] Figure 7 (a) shows a saliency map of the first cluster. The highlighted area H on the right half of the saliency map can represent the range of image samples corresponding to the first cluster in the image samples. Furthermore, the highlighted area can be used to indicate that the model inference device uses cello and violin instead of acoustic guitar, banjo, and electric guitar as the basis for the inference result. That is, the model inference device can infer from the highlighted area that the label corresponding to the image is cello and violin instead of acoustic guitar, banjo, and electric guitar, thus explaining why the inference result is cello and violin instead of acoustic guitar, banjo, and electric guitar. Figure 7 (b) shows the contrast saliency map of the second cluster. The highlighted area H on the left half of the contrast saliency map can represent the range of image samples corresponding to the second cluster in the image samples. Furthermore, the highlighted area can be used to indicate that the model inference device uses acoustic guitar, banjo, and electric guitar instead of cello and violin as the basis for the inference result. That is, the model inference device can infer from the highlighted area that the label corresponding to the image is acoustic guitar, banjo, and electric guitar instead of cello and violin, thus explaining why the inference result is acoustic guitar, banjo, and electric guitar instead of cello and violin.
[0091] Figure 7 (a.1) shows the contrast saliency map within the first cluster and the first class. The contrast interpretation result corresponding to this contrast saliency map can be limited to the range of image samples corresponding to the first cluster in the image samples. That is, the highlighted region H of this contrast saliency map can be limited to... Figure 7 (a) The highlighted area on the right half of the corresponding saliency map, and the highlighted area of the saliency map can be used to indicate that the model inference device uses the cello instead of the violin as the basis for the inference result, that is, the model inference device can infer from the highlighted area that the label corresponding to the image is cello instead of violin, thus explaining the reason why the inference result is cello instead of violin. Figure 7 (a.2) shows the contrast saliency map of the second class within the first cluster. The contrast interpretation result corresponding to the contrast saliency map can also be limited to the range of image samples corresponding to the first cluster in the image samples. Furthermore, the highlighted area H of the contrast saliency map can be used to indicate that the model inference device uses a violin instead of a cello as the basis for the inference result, thereby explaining why the inference result can be represented as a violin instead of a cello. Figure 7(b.1) shows the contrast interpretation results within the second cluster and the first class. The contrast interpretation results corresponding to this contrast saliency map can be limited to the range of image samples corresponding to the second cluster in the image samples, that is, the highlighted region H of the contrast saliency map can be limited to... Figure 7 (b) The highlighted area on the left half of the corresponding contrast saliency map, and the highlighted area of the contrast saliency map can be used to indicate that the model inference device uses an acoustic guitar instead of a banjo or electric guitar as the basis for the inference result, thereby explaining why the inference result is an acoustic guitar instead of a banjo or electric guitar; Figure 7 (b.2) shows the contrast interpretation results within the second cluster and the second class. The contrast interpretation results corresponding to this contrast saliency map can also be limited to the range of image samples corresponding to the second cluster in the image samples. Furthermore, the highlighted area H of this contrast saliency map can be used to indicate that the model inference device uses the banjo instead of the acoustic guitar or electric guitar as the basis for the inference result, thereby explaining why the inference result is the banjo instead of the acoustic guitar or electric guitar. Figure 7 (b.3) shows the contrast interpretation results of the third class within the second cluster. The contrast interpretation results corresponding to the contrast saliency map can also be limited to the range of image samples corresponding to the second cluster in the image samples. Furthermore, the highlighted area H of the contrast saliency map can be used to indicate that the model inference device uses an electric guitar instead of an acoustic guitar or banjo as the basis for the inference result, thereby explaining why the inference result is an electric guitar instead of an acoustic guitar or banjo.
[0092] Therefore, the comparative explanation results can be presented more comprehensively and richly. By explaining why it is A instead of B, rather than just explaining why it is A, users can better understand the reasoning results of the model reasoning device and improve the model in the model reasoning device based on the comparative explanation results.
[0093] In the above saliency diagram, the highlighted portion H can represent evidence for the comparative interpretation results, for example, through... Figure 7 (a.1) The highlighted area indicates that the reasoning result is a cello rather than a violin, based on the features under the instrument. Therefore, it can be assumed that the reasoning of the model reasoning device is based on the fact that violins usually have protective pads under their bodies, while cellos usually do not.
[0094] Figure 8 A flowchart illustrating the evaluation phase of a method for interpreting model reasoning results according to an embodiment of this application is shown. Figure 8 As shown, the evaluation phase process may include:
[0095] Step S801: Based on pixel importance, the evaluation module sorts the pixels in the corresponding image samples.
[0096] Among them, pixel importance can be determined according to Figure 3 In the corresponding saliency map in S305, the brightness of each pixel is determined. Higher pixel brightness indicates greater importance of the pixel. The corresponding saliency map can be determined based on the comparison subject and object being evaluated, such as... Figure 7 In the saliency diagram shown, when the evaluation module evaluates the cello as the subject of comparison and the violin as the object of comparison (i.e., the corresponding comparative explanation explains why it is a cello and not a violin), the evaluation module's evaluation object can be the comparative explanation result of "why it is a cello and not a violin," and the module can select the corresponding comparative explanation result. Figure 7 (a.1) As a saliency map. When the evaluation module evaluates cluster 1 as the subject of comparison and cluster 2 as the object of comparison, meaning the corresponding comparative explanation explains "why cluster 1 is used instead of cluster 2," the evaluation module's evaluation object can be the comparative explanation result of "why cluster 1 is used instead of cluster 2," and the module can select the comparison explanation result that corresponds to this comparative explanation result. Figure 7 (a) As a saliency diagram.
[0097] In step S802, according to the sorting order, the evaluation module sequentially erases pixels in the image sample until the number of erased pixels reaches a predetermined threshold. Each time pixels are erased, the model inference device obtains the predicted probabilities of the comparison subject and the comparison object based on the input image sample after the pixels are erased.
[0098] The number of pixels to be erased each time can be determined as needed. The pixels to be erased each time can be the more important pixels among all the pixels determined by sorting, or pixels with higher brightness. The method of erasing pixels can be, for example, setting the pixels to be erased in the image sample to white.
[0099] The predicted probabilities of the comparison subject and the comparison object can be represented as the probabilities of the corresponding comparison subject and the comparison object in the inference results inferred by the model inference device, respectively.
[0100] Step S803: The evaluation module determines the comparative fidelity of the corresponding comparative interpretation results based on the predicted probabilities of the comparative subject and the comparative object.
[0101] The method for calculating loyalty can be shown in formula (4):
[0102]
[0103] CAUC can represent the area under the comparison line and can be used to measure the fidelity of the comparison. The smaller the CAUC value, the higher the fidelity of the corresponding comparison interpretation result, that is, the more accurately the comparison interpretation result reflects the evidence that distinguishes the comparison subject from the comparison object. P M (c|X[r,n] P can represent the predicted probability of the subject when the r-th pixel is erased during the pixel erasure process. M (C|X [r,n] ) can represent the predicted probability of the comparison object when the r-th pixel is erased during the pixel erasure process, H can represent the saliency map of the comparison subject, M can represent the corresponding inference model in the model inference device (e.g., ResNet50, GoogleNet). X can represent the image sample, c can represent the comparison subject, C can represent the comparison object, and n can represent the total number of times pixels are erased.
[0104] Figure 9 A schematic diagram illustrating the removal of pixels from an image sample according to an embodiment of this application is shown. Figure 9 As shown, image samples can be, for example... Figure 9 (a) When the subject of comparison is a cello and the object of comparison is a violin within the same cluster, the saliency map of the corresponding image sample can be, for example... Figure 9 (b) Referring to S801, the pixels in the image sample are sorted according to the importance of the pixels in the contrast saliency map, and then referring to S802, the pixels in the image sample are erased one by one. Figure 9 (c.1)- Figure 9 (c.3) can represent image samples after a predetermined number of pixels are erased each time, and the number of pixels erased each time can be as follows: Figure 9 (c.1)- Figure 9 The white area is shown in (c.3).
[0105] Figure 10 A schematic diagram illustrating the determination of comparative loyalty according to an embodiment of this application is shown. Figure 10 As shown in (a), the horizontal axis represents the percentage of erased pixels in the image sample relative to the total number of pixels, and the vertical axis represents the prediction probability. Figure 10 (a) shows the change in the predicted probabilities of the subject and object of comparison during the process of erasing pixels from an image sample.
[0106] Among them, with Figure 9 Taking the process shown in the figure as an example, the comparison subject is a cello and the comparison object is a violin. In the process of erasing pixels from the image sample one by one, the prediction probability (i.e., P(c)) corresponding to the comparison subject (i.e., the cello) decreases, while the prediction probability (i.e., P(C)) corresponding to the comparison object (i.e., the violin) increases. This is because, when the erased pixels can truly reflect the evidence that distinguishes the comparison subject and the comparison object, the model inference device will have more difficulty distinguishing the comparison subject and the comparison object in the process of erasing pixels.
[0107] like Figure 10(b) The horizontal axis represents the percentage of the total number of pixels in the image sample that have been erased, while the vertical axis represents the product of the predicted probability of the subject and the difference between the predicted probabilities of the subject and the object (i.e., P(c)*(1-P(C)). As pixels are erased from the image sample, the product of the predicted probabilities of the subject and the object gradually decreases. The smaller the product, the more accurately the erased pixels in the current image sample reflect the evidence that distinguishes the subject from the object.
[0108] Based on the product determined during the pixel erasure process, the contrast fidelity of the contrast interpretation results regarding the comparison subject and object can be determined. For example, CAUC can be calculated referring to the example in S803; the smaller the CAUC value, the higher the contrast fidelity.
[0109] Figure 11 A flowchart illustrating a method for interpreting model inference results according to an embodiment of this application is shown. This method can be executed by a server or a terminal device. For example, the server can execute the method described below and control the terminal device or display device to display visualized execution results, such as saliency maps, contrast saliency maps, etc. Figure 11 As shown, the method includes:
[0110] Step S1101: Obtain at least one first label, where the first label represents the inference result of the image sample;
[0111] Step S1102: Divide the at least one first label into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels in the same cluster have a higher probability of appearing in the same image sample compared with the first labels in other clusters;
[0112] Step S1103: Determine the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster, or determine the second contrast saliency map corresponding to each first label within the cluster based on the first saliency map corresponding to each cluster and the first saliency map corresponding to each first label within the cluster. The first saliency map represents each cluster or each first label within the cluster as the basis for the inference result, the first contrast saliency map represents each cluster as the basis for the inference result relative to other clusters, and the second contrast saliency map represents each first label within the cluster as the basis for the inference result relative to other first labels within the cluster.
[0113] According to the embodiments of this application, by obtaining at least one first label, dividing the at least one first label into at least one cluster, and determining the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster, or by determining the second contrast saliency map corresponding to each first label within a cluster based on the first saliency map corresponding to each cluster and the first saliency map corresponding to each first label within a cluster, the objects for contrast interpretation can be automatically determined without manual selection, saving a lot of resources. Moreover, all first labels can be contrast interpreted, making the interpretation more complete, which is more helpful for improving and debugging the inference model, making the inference results more accurate. The contrast interpretation between clusters can solve the problem of interpretation of multiple subject objects and can be applied to multi-label image scenarios. Through the contrast interpretation of first labels within a cluster, and since the first labels within the same cluster have a higher probability of appearing in the same image sample compared to the first labels within other clusters, more refined distinction of key features of labels within a cluster can be achieved, making the contrast interpretation results more valuable. By displaying the contrast saliency map, the results of the contrast interpretation can be systematically and concisely visualized.
[0114] The first tag can be, for example Figure 3 In step S302, there are K labels (which can also correspond to the K classes mentioned above), and the clusters can be, for example, m clusters in S302. This application does not limit the number of first labels and clusters. The first saliency map can be, for example, the saliency map obtained in S303, and the contrast saliency map obtained in S305 is the first contrast saliency map and the second contrast saliency map.
[0115] For example, the first label corresponding to a certain image sample could be, for instance, cello, violin, electric guitar, or acoustic guitar. Dividing the at least one first label into at least one cluster, for example, grouping cello and violin into one cluster and electric guitar and acoustic guitar into another. Based on the first saliency maps corresponding to each cluster, a first contrast saliency map corresponding to each cluster is determined. For example, contrast saliency maps corresponding to the first and second clusters can be determined separately. The contrast saliency map corresponding to the first cluster can represent the basis for the inference result of cello and violin relative to electric guitar and acoustic guitar as image samples, and the contrast saliency map corresponding to the second cluster can represent the basis for the inference result of electric guitar and acoustic guitar relative to cello and violin as image samples. Based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within each cluster, the second contrast saliency maps corresponding to each first label within each cluster are determined. For example, in the first cluster, the contrast saliency maps corresponding to the cello and violin are determined respectively, and in the second cluster, the second contrast saliency maps corresponding to the electric guitar and acoustic guitar are determined respectively. Taking the contrast saliency map corresponding to the cello as an example, it can represent the basis for the inference result of the cello relative to the violin as an image sample in the first cluster.
[0116] Examples of steps S1101-S1103 can be found in [reference needed]. Figure 3 S302-S305.
[0117] In one possible implementation, the highlighted areas of the first saliency map represent the basis for inference results of each cluster relative to other clusters, and the highlighted areas of the second saliency map represent the basis for inference results of each first label within a cluster relative to other first labels within the cluster.
[0118] According to the embodiments of this application, by using the highlighted areas of the first contrast saliency map to represent each cluster relative to other clusters as the basis for the inference result, and the highlighted areas of the second contrast saliency map to represent each first label within a cluster relative to other first labels within the cluster as the basis for the inference result, a more intuitive interpretation of the inference result can be achieved, and the interpretation content can be made richer, which helps to improve and debug the inference model and obtain a more accurate inference model.
[0119] The highlighted areas in the first contrast saliency map can be referenced. Figure 7 (a) Figure 7 (b) An example of the highlighted area; the highlighted areas in the second contrast saliency map can be referenced. Figure 7 (a.1) Figure 7 (a.2) Figure 7 (b.1) Figure 7 (b.2) Figure 7 Example of the highlighted area in (b.3).
[0120] In one possible implementation, determining a first contrast saliency map corresponding to each cluster based on each first saliency map corresponding to each cluster includes: eliminating common features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, determining the differential features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, and determining the first contrast saliency map corresponding to each cluster based on the differential features; determining a second contrast saliency map corresponding to each first label within a cluster based on each first saliency map corresponding to each cluster and each first saliency map corresponding to each first label within a cluster includes: determining the image region corresponding to the cluster indicated by the first saliency map of the cluster corresponding to each first label, eliminating common features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster within the image region, determining the differential features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster, and determining the second contrast saliency map corresponding to each first label within a cluster based on the differential features.
[0121] According to the embodiments of this application, by eliminating common features and determining the saliency map based on the difference features, the comparison interpretation results can be displayed with fine-grained differences between the comparison subject and the comparison object, making the interpretation more accurate and specific. By determining the corresponding image region of the cluster indicated by the first saliency map of each first label, the comparison interpretation results of the labels within the same cluster can be limited to the same area in the image, making the comparison interpretation more accurate.
[0122] Common features can be, for example, regions with the same pixel brightness in the first saliency map, while differential features can be, for example, regions with different pixel brightness in the first saliency map. Common features are eliminated, differential features are determined, and based on these differential features, the first contrast saliency map corresponding to each cluster is determined, for example, according to formula (1). The image region corresponding to the cluster indicated by the first saliency map of each first label can be determined, for example, according to formula (3). Within this image region, common features are eliminated, differential features are determined, and based on these differential features, the second contrast saliency map corresponding to each first label within the cluster is determined, for example, according to formula (2).
[0123] It should be noted that the above process can also be determined in other ways besides those shown in formulas (1)-(3), and this application does not impose any restrictions on this.
[0124] An example of the above process can be found in [reference]. Figure 3 The relevant process of step S304 in the middle.
[0125] Figure 12 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown. Figure 12 As shown, dividing the at least one first tag into at least one cluster includes:
[0126] Step S1201: Divide at least one second label into at least one cluster. The second label represents the reasoning result of all image samples. Each cluster includes at least one of the second labels. The second label in the same cluster has a higher probability of appearing in the same image sample than the second label in other clusters.
[0127] Step S1202: Based on the correspondence between the at least one first tag and the at least one second tag, the first tags that belong to the same cluster as the corresponding second tags among the at least one first tags are grouped into one cluster.
[0128] According to the embodiments of this application, by dividing at least one second label into at least one cluster, and according to the correspondence between the at least one first label and the at least one second label, the first labels in the at least one first label that correspond to the second label in the same cluster are divided into one cluster, the objects of comparison and interpretation can be automatically determined without manual selection, saving a lot of resources, and making easily confused labels into one cluster, thereby making the results of comparison and interpretation more referential and valuable.
[0129] The second tag can be, for example... Figure 3 In step S301, all tags in the tag set are included, but this application does not limit this. The correspondence between the at least one first tag and the at least one second tag can be set as needed. For example, if the content of the first tag and the second tag are the same, it indicates that the first tag corresponds to the second tag.
[0130] For example, the second label can be, for example, violin, cello, acoustic guitar, electric guitar, banjo. At least one second label can be divided into at least one cluster. For example, violin and cello can be divided into one cluster, and acoustic guitar, electric guitar, banjo can be divided into another cluster. According to the correspondence between the at least one first label and the at least one second label, the first labels in the at least one first label that belong to the same cluster as the corresponding second label can be divided into one cluster. For example, when the first label includes violin, cello, acoustic guitar, electric guitar, according to the above clustering situation, violin and cello in the first label can be divided into one cluster, and acoustic guitar and electric guitar in the first label can be divided into another cluster.
[0131] Examples of steps S1201-S1202 can be found by referring to Figure 3 In step S301 and Figure 4 .
[0132] Figure 13 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown. Figure 13 As shown, the method also includes:
[0133] Step S1301: Determine the first sorting of pixels in the image sample based on the first comparison saliency map;
[0134] Step S1302: According to the first sorting, a predetermined number of pixels in the image sample are erased sequentially;
[0135] Step S1303: Determine the evaluation corresponding to the first saliency map based on at least one image sample obtained during the pixel erasure process; and / or
[0136] Step S1304: Determine the second sorting of pixels in the image sample based on the second contrast saliency map;
[0137] Step S1305: According to the second sorting, a predetermined number of pixels in the image sample are erased sequentially;
[0138] Step S1306: Determine the evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process.
[0139] According to the embodiments of this application, by determining the order of pixels in the image sample based on the saliency map, and then sequentially erasing a predetermined number of pixels in the image sample according to the order, and determining the evaluation corresponding to the saliency map based on at least one image sample obtained during the pixel erasure process, the results of the comparative interpretation can be quantitatively evaluated, thereby intuitively demonstrating the value of the comparative interpretation content.
[0140] The number of pixels can be determined as needed. The number of pixels erased each time can be the same or different. This application does not limit this. The corresponding evaluation can be, for example, the comparison loyalty mentioned above.
[0141] Examples of steps S1301 and S1304 can be found in [reference]. Figure 8 Examples of steps S801, S1302, and S1305 can be found in [reference needed]. Figure 8 Examples of steps S802, S1303, and S1306 can be found in [reference needed]. Figure 8 S803.
[0142] In one possible implementation, the first sorting is determined based on the brightness of each pixel in the first contrast saliency map, and the second sorting is determined based on the brightness of each pixel in the second contrast saliency map.
[0143] According to the embodiments of this application, a more accurate evaluation of the comparative interpretation results indicated by the saliency map can be achieved.
[0144] The first and second sorting can be, for example, sorted by pixel brightness from high to low.
[0145] An example of the above process can be found in [reference]. Figure 8 The relevant description of S801.
[0146] Figure 14 A flowchart illustrating a method for interpreting model reasoning results according to an embodiment of this application is shown. Figure 14 As shown, based on at least one image sample obtained during the pixel erasure process, the evaluation corresponding to the first saliency map is determined, including:
[0147] Step S1401: Based on at least one image sample obtained during the pixel erasure process, determine the first prediction probability of the cluster corresponding to the first contrast saliency map and the second prediction probability of other clusters outside the first cluster. The first prediction probability represents the probability that the inference result of the image sample is the corresponding cluster, and the second prediction probability represents the probability that the inference result of the image sample is other clusters outside the first cluster.
[0148] Step S1402: Determine the evaluation of the first contrast saliency map based on the first predicted probability and the second predicted probability;
[0149] Based on at least one image sample obtained during the pixel erasure process, determine the evaluation corresponding to the second saliency map, including:
[0150] Step S1403: Based on at least one image sample obtained during the pixel erasure process, determine the third prediction probability of the first label within the cluster corresponding to the second contrast saliency map, and the fourth prediction probability of other first labels within the cluster other than the first label within the cluster. The third prediction probability represents the probability that the inference result of the image sample is the corresponding first label within the cluster, and the fourth prediction probability represents the probability that the inference result of the image sample is other first labels other than the first label within the cluster.
[0151] Step S1404: Determine the evaluation of the second contrast saliency map based on the third prediction probability and the fourth prediction probability.
[0152] According to the embodiments of this application, when conducting evaluations, the degree of support of the explanatory basis presented in the comparative explanation results for the comparative subject and the degree of denial of the comparative object can be taken into account, thereby comprehensively measuring the loyalty of the comparative explanation results and making the evaluation results more accurate.
[0153] The first and third prediction probabilities can be determined from P in formula (4) above. M (c|X [r,n] The second and fourth prediction probabilities can be determined based on P in formula (4) above. M (C|X [r,n] The first prediction probability, the second prediction probability, the third prediction probability, and the fourth prediction probability can also be determined in other ways. This application does not limit this. The evaluation of the first contrast saliency map is determined based on the first prediction probability and the second prediction probability, and the evaluation of the second contrast saliency map is determined based on the third prediction probability and the fourth prediction probability. For example, the above formula (4) can be referred to, or it can be determined in other ways besides formula (4). This application does not limit this.
[0154] Examples of steps S1401-S1404 can be found here. Figure 8 The relevant description of S803.
[0155] Figure 15 A structural diagram of an apparatus for interpreting model reasoning results according to an embodiment of this application is shown. Figure 15 As shown, the device 1500 includes:
[0156] The acquisition module 1501 is used to acquire at least one first label, wherein the first label represents the inference result of the image sample;
[0157] The first determining module 1502 is used to divide the at least one first label into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels in the same cluster have a higher probability of appearing in the same image sample than the first labels in other clusters;
[0158] The second determining module 1503 is used to determine a first contrast saliency map corresponding to each cluster based on each first saliency map corresponding to each cluster, or to determine a second contrast saliency map corresponding to each first label within a cluster based on each first saliency map corresponding to each cluster and each first saliency map corresponding to each first label within a cluster, wherein the first saliency map represents each cluster or each first label within a cluster as the basis for the inference result, the first contrast saliency map represents each cluster as the basis for the inference result relative to other clusters, and the second contrast saliency map represents each first label within a cluster as the basis for the inference result relative to other first labels within the cluster.
[0159] According to the embodiments of this application, by obtaining at least one first label, dividing the at least one first label into at least one cluster, and determining the first contrast saliency map corresponding to each cluster based on the first saliency map corresponding to each cluster, or by determining the second contrast saliency map corresponding to each first label within a cluster based on the first saliency map corresponding to each cluster and the first saliency map corresponding to each first label within a cluster, the objects for contrast interpretation can be automatically determined without manual selection, saving a lot of resources. Moreover, all first labels can be contrast interpreted, making the interpretation more complete, which is more helpful for improving and debugging the inference model, making the inference results more accurate. The contrast interpretation between clusters can solve the problem of interpretation of multiple subject objects and can be applied to multi-label image scenarios. Through the contrast interpretation of first labels within a cluster, and since the first labels within the same cluster have a higher probability of appearing in the same image sample compared to the first labels within other clusters, more refined distinction of key features of labels within a cluster can be achieved, making the contrast interpretation results more valuable. By displaying the contrast saliency map, the results of the contrast interpretation can be systematically and concisely visualized.
[0160] In one possible implementation, the highlighted areas of the first saliency map represent the basis for inference results of each cluster relative to other clusters, and the highlighted areas of the second saliency map represent the basis for inference results of each first label within a cluster relative to other first labels within the cluster.
[0161] According to the embodiments of this application, by using the highlighted areas of the first contrast saliency map to represent each cluster relative to other clusters as the basis for the inference result, and the highlighted areas of the second contrast saliency map to represent each first label within a cluster relative to other first labels within the cluster as the basis for the inference result, a more intuitive interpretation of the inference result can be achieved, and the interpretation content can be made richer, which helps to improve and debug the inference model and obtain a more accurate inference model.
[0162] In one possible implementation, determining a first contrast saliency map corresponding to each cluster based on each first saliency map corresponding to each cluster includes: eliminating common features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, determining the differential features in the first saliency maps corresponding to each cluster relative to the first saliency maps of other clusters, and determining the first contrast saliency map corresponding to each cluster based on the differential features; determining a second contrast saliency map corresponding to each first label within a cluster based on each first saliency map corresponding to each cluster and each first saliency map corresponding to each first label within a cluster includes: determining the image region corresponding to the cluster indicated by the first saliency map of the cluster corresponding to each first label, eliminating common features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster within the image region, determining the differential features in the first saliency maps corresponding to each first label within a cluster relative to the first saliency maps of other first labels within a cluster, and determining the second contrast saliency map corresponding to each first label within a cluster based on the differential features.
[0163] According to the embodiments of this application, by eliminating common features and determining the saliency map based on the difference features, the comparison interpretation results can be displayed with fine-grained differences between the comparison subject and the comparison object, making the interpretation more accurate and specific. By determining the corresponding image region of the cluster indicated by the first saliency map of each first label, the comparison interpretation results of the labels within the same cluster can be limited to the same area in the image, making the comparison interpretation more accurate.
[0164] In one possible implementation, dividing the at least one first label into at least one cluster includes: dividing at least one second label into at least one cluster, wherein the second label represents the inference result of all image samples, each cluster includes at least one of the second labels, and the second labels in the same cluster have a higher probability of appearing in the same image sample than the second labels in other clusters; and dividing the first labels in the at least one first label that are in the same cluster as the corresponding second labels into one cluster according to the correspondence between the at least one first label and the at least one second label.
[0165] According to the embodiments of this application, by dividing at least one second label into at least one cluster, and according to the correspondence between the at least one first label and the at least one second label, the first labels in the at least one first label that correspond to the second label in the same cluster are divided into one cluster, the objects of comparison and interpretation can be automatically determined without manual selection, saving a lot of resources, and making easily confused labels into one cluster, thereby making the results of comparison and interpretation more referential and valuable.
[0166] In one possible implementation, the device further includes: a third determining module, configured to determine a first sorting of pixels in the image sample based on the first saliency map; a first pixel erasure module, configured to sequentially erasure a predetermined number of pixels in the image sample based on the first sorting; a fourth determining module, configured to determine an evaluation corresponding to the first saliency map based on at least one image sample obtained during the pixel erasure process; and / or a fifth determining module, configured to determine a second sorting of pixels in the image sample based on the second saliency map; a second pixel erasure module, configured to sequentially erasure a predetermined number of pixels in the image sample based on the second sorting; and a sixth determining module, configured to determine an evaluation corresponding to the second saliency map based on at least one image sample obtained during the pixel erasure process.
[0167] According to the embodiments of this application, by determining the order of pixels in the image sample based on the saliency map, and then sequentially erasing a predetermined number of pixels in the image sample according to the order, and determining the evaluation corresponding to the saliency map based on at least one image sample obtained during the pixel erasure process, the results of the comparative interpretation can be quantitatively evaluated, thereby intuitively demonstrating the value of the comparative interpretation content.
[0168] In one possible implementation, the first sorting is determined based on the brightness of each pixel in the first contrast saliency map, and the second sorting is determined based on the brightness of each pixel in the second contrast saliency map.
[0169] According to the embodiments of this application, a more accurate evaluation of the comparative interpretation results indicated by the saliency map can be achieved.
[0170] In one possible implementation, determining the evaluation corresponding to the first contrast saliency map based on at least one image sample obtained during the pixel erasure process includes: determining a first predicted probability of the cluster corresponding to the first contrast saliency map and a second predicted probability of other clusters outside the first cluster based on at least one image sample obtained during the pixel erasure process, wherein the first predicted probability represents the probability that the inference result of the image sample is the corresponding cluster, and the second predicted probability represents the probability that the inference result of the image sample is other clusters outside the first cluster; determining the evaluation of the first contrast saliency map based on the first predicted probability and the second predicted probability; determining the evaluation corresponding to the second contrast saliency map based on at least one image sample obtained during the pixel erasure process includes: determining a third predicted probability of the first label within the cluster corresponding to the second contrast saliency map and a fourth predicted probability of other first labels within the cluster outside the first label within the cluster based on at least one image sample obtained during the pixel erasure process, wherein the third predicted probability represents the probability that the inference result of the image sample is the corresponding first label within the cluster, and the fourth predicted probability represents the probability that the inference result of the image sample is other first labels outside the first label within the cluster; determining the evaluation of the second contrast saliency map based on the third predicted probability and the fourth predicted probability.
[0171] According to the embodiments of this application, when conducting evaluations, the degree of support of the explanatory basis presented in the comparative explanation results for the comparative subject and the degree of denial of the comparative object can be taken into account, thereby comprehensively measuring the loyalty of the comparative explanation results and making the evaluation results more accurate.
[0172] Figure 16 A structural diagram of an interpretation apparatus for model reasoning results according to an embodiment of this application is shown. This interpretation apparatus is applicable to… Figure 1 In the data annotation system shown, the above is performed. Figures 3-14 Any of the methods described herein illustrates an interpretation method for the model inference results. For example, the interpretation device may be a server, or a chip (system) or other component or part that can be installed inside the server. As another example, the interpretation device may also be the aforementioned interpretation device 1500. This application does not limit this aspect.
[0173] like Figure 16 As shown, the interpretation device 700 may include a processor 701 and a transceiver 702. Optionally, the interpretation device 700 may include a memory 703. The processor 701 is coupled to the transceiver 702 and the memory 703, for example, they can be connected via a communication bus.
[0174] The following combination Figure 16 The various components of the explanatory device 700 will be described in detail.
[0175] The processor 701 described above is the control center of the interpreter 700. It can be a single processor or a collective term for multiple processing elements. For example, the processor 701 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0176] Optionally, the processor 701 can perform various functions of the interpretation processing device 700 by running or executing software programs stored in the memory 703 and calling data stored in the memory 703.
[0177] In a specific implementation, as one example, the processor 701 may include one or more CPUs, for example... Figure 16 CPU0 and CPU1 are shown in the diagram.
[0178] In one possible implementation, the interpreter 700 may also include multiple processors, for example... Figure 16 The processors 701 and 704 are shown. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, "processor" can refer to one or more communication devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0179] The transceiver 702 described above is used for communication with other interpreting devices. For example, if the interpreting device 700 is a server, the transceiver 702 can be used to communicate with another server.
[0180] Optionally, transceiver 702 may include a receiver and a transmitter. Figure 16 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0181] Optionally, the transceiver 702 can be integrated with the processor 701 or exist independently, and can be connected via the input / output ports of the interpreter 700. Figure 16(Not shown in the image) is coupled to the processor 701, but this application embodiment does not limit this.
[0182] The memory 703 described above can be used to store software programs that execute the scheme of this application, and the processor 701 controls the execution. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.
[0183] The memory 703 can be a read-only memory (ROM) or other type of static storage communication device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage communication device capable of storing information and instructions, or it can be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage communication device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited to these. It should be noted that the memory 703 can be integrated with the processor 701 or exist independently, and can be accessed through the input / output ports of the interpreter 700. Figure 16 (Not shown in the image) is coupled to the processor 701, but this application embodiment does not limit this.
[0184] It should be noted that, Figure 16 The structure of the explanatory apparatus 700 shown does not constitute a limitation on the implementation of the explanatory apparatus. The actual explanatory apparatus may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0185] Embodiments of this application provide an apparatus for interpreting model inference results, including: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions.
[0186] An embodiment of this application provides a terminal device that can perform the above-described method.
[0187] Embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the above-described method.
[0188] Embodiments of this application provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0189] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital video disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing.
[0190] The computer-readable program instructions or code described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0191] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this application.
[0192] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0193] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0194] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0195] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
[0196] It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented using hardware (such as circuits or ASICs (Application Specific Integrated Circuits)) that performs the corresponding function or action, or using a combination of hardware and software, such as firmware.
[0197] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0198] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for interpreting model reasoning results, characterized in that, The method includes: Obtain at least one first label, where the first label represents the inference result of the image sample; The at least one first label is divided into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels in the same cluster have a higher probability of appearing in the same image sample than the first labels in other clusters; Based on each first saliency map corresponding to each cluster, determine the first contrast saliency map corresponding to each cluster; or, based on each first saliency map corresponding to each cluster and each first label within the cluster, determine the second contrast saliency map corresponding to each first label within the cluster. Here, the first saliency map represents each cluster or each first label within the cluster as the basis for the inference result, the first contrast saliency map represents each cluster relative to other clusters as the basis for the inference result, and the second contrast saliency map represents each first label within the cluster relative to other first labels within the cluster as the basis for the inference result. Based on the first saliency maps corresponding to each cluster, determine the first contrastive saliency map corresponding to each cluster, including: Eliminate the common features in the first saliency map corresponding to each cluster and the first saliency map relative to other clusters, determine the difference features in the first saliency map corresponding to each cluster and the first saliency map relative to other clusters, and determine the first contrast saliency map corresponding to each cluster based on the difference features. Based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within each cluster, determine the second contrastive saliency maps corresponding to each first label within each cluster, including: The image region corresponding to the first saliency map of each first label is determined. Within the image region, the common features of the first saliency maps corresponding to each first label in the cluster and the first saliency maps relative to other first labels in the cluster are eliminated. The difference features of the first saliency maps corresponding to each first label in the cluster and the first saliency maps relative to other first labels in the cluster are determined. Based on the difference features, the second contrast saliency map corresponding to each first label in the cluster is determined.
2. The method according to claim 1, characterized in that, The highlighted areas of the first saliency map represent the basis for inference results of each cluster relative to other clusters, and the highlighted areas of the second saliency map represent the basis for inference results of each first label within a cluster relative to other first labels within the cluster.
3. The method according to claim 1, characterized in that, Dividing the at least one first tag into at least one cluster includes: At least one second label is divided into at least one cluster, where the second label represents the inference result of all image samples. Each cluster includes at least one of the second labels. The second label in the same cluster has a higher probability of appearing in the same image sample than the second label in other clusters. Based on the correspondence between the at least one first tag and the at least one second tag, the first tags that belong to the same cluster as the corresponding second tags among the at least one first tags are grouped into one cluster.
4. The method according to claim 1, characterized in that, The method further includes: Based on the first contrast saliency map, determine the first sorting of pixels in the image sample; According to the first sorting, a predetermined number of pixels in the image sample are erased sequentially. Based on at least one image sample obtained during the pixel erasure process, determine the evaluation corresponding to the first saliency map; and / or Based on the second saliency map, determine the second sorting of pixels in the image sample; According to the second sorting, a predetermined number of pixels in the image sample are erased sequentially. The evaluation corresponding to the second saliency map is determined based on at least one image sample obtained during the pixel erasure process.
5. The method according to claim 4, characterized in that, The first sorting is determined based on the brightness of each pixel in the first contrast saliency map, and the second sorting is determined based on the brightness of each pixel in the second contrast saliency map.
6. The method according to claim 4, characterized in that, Based on at least one image sample obtained during the pixel erasure process, determine the evaluation corresponding to the first saliency map, including: Based on at least one image sample obtained during the pixel erasure process, a first predicted probability of the cluster corresponding to the first contrast saliency map and a second predicted probability of other clusters outside the first cluster are determined. The first predicted probability represents the probability that the inference result of the image sample is the corresponding cluster, and the second predicted probability represents the probability that the inference result of the image sample is other clusters outside the first cluster. The evaluation of the first contrast saliency map is determined based on the first predicted probability and the second predicted probability. Based on at least one image sample obtained during the pixel erasure process, determine the evaluation corresponding to the second saliency map, including: Based on at least one image sample obtained during the pixel erasure process, a third predicted probability of the first label within the cluster corresponding to the second contrast saliency map and a fourth predicted probability of other first labels within the cluster other than the first label within the cluster are determined. The third predicted probability represents the probability that the inference result of the image sample is the corresponding first label within the cluster, and the fourth predicted probability represents the probability that the inference result of the image sample is other first labels other than the first label within the cluster. The evaluation of the second contrast saliency map is determined based on the third and fourth prediction probabilities.
7. An apparatus for interpreting the results of model reasoning, characterized in that, The device includes: An acquisition module is used to acquire at least one first label, wherein the first label represents the inference result of the image sample; The first determining module is used to divide the at least one first label into at least one cluster, wherein each cluster includes at least one of the first labels, and the first labels in the same cluster have a higher probability of appearing in the same image sample than the first labels in other clusters; The second determining module is used to determine a first contrast saliency map corresponding to each cluster based on each first saliency map corresponding to each cluster, or to determine a second contrast saliency map corresponding to each first label within a cluster based on each first saliency map corresponding to each cluster and each first saliency map corresponding to each first label within a cluster, wherein the first saliency map represents each cluster or each first label within a cluster as the basis for the inference result, the first contrast saliency map represents each cluster as the basis for the inference result relative to other clusters, and the second contrast saliency map represents each first label within a cluster as the basis for the inference result relative to other first labels within the cluster; Based on the first saliency maps corresponding to each cluster, determine the first contrastive saliency map corresponding to each cluster, including: Eliminate the common features in the first saliency map corresponding to each cluster and the first saliency map relative to other clusters, determine the difference features in the first saliency map corresponding to each cluster and the first saliency map relative to other clusters, and determine the first contrast saliency map corresponding to each cluster based on the difference features. Based on the first saliency maps corresponding to each cluster and the first saliency maps corresponding to each first label within each cluster, determine the second contrastive saliency maps corresponding to each first label within each cluster, including: The image region corresponding to the first saliency map of each first label is determined. Within the image region, the common features of the first saliency maps corresponding to each first label in the cluster and the first saliency maps relative to other first labels in the cluster are eliminated. The difference features of the first saliency maps corresponding to each first label in the cluster and the first saliency maps relative to other first labels in the cluster are determined. Based on the difference features, the second contrast saliency map corresponding to each first label in the cluster is determined.
8. An apparatus for interpreting the results of model reasoning, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1-6 when executing the instructions.
9. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-6.
10. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is executed in an electronic device, a processor in the electronic device performs the method of any one of claims 1-6.
Citation Information
Patent Citations
Image scene recognition method and device
CN105809146A
Visual analytics system for convolutional neural network based classifiers
CN110892414A