Information processing method, information processing system, and storage medium

By employing computer-based information processing methods, the influence of inference results from multiple inference engines is acquired and displayed, solving the problem of low efficiency in analyzing behavioral differences in inference models in existing technologies and achieving highly efficient analysis of inference results.

CN114788265BActive Publication Date: 2026-05-08PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2020-09-18
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to evaluate the behavioral differences of multiple inference models in object detection processing in parallel, resulting in low efficiency in the analysis of detection results. In particular, when the number, location, or range of detected objects is inconsistent, users need to spend time and effort comparing and aligning the detection results.

Method used

By using computer-executed information processing methods, the inference results of multiple inference engines are obtained, and the influence of input data on each inference result is displayed in parallel or overlapping. Based on these results, a combination is determined, and the influence of the combination is displayed through a prompting device, thereby realizing the parallel evaluation of multiple inference results.

Benefits of technology

It improves the efficiency of users in analyzing multiple inference results, reduces the workload of users in aligning and comparing detection results, and achieves efficient inference result analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114788265B_ABST
    Figure CN114788265B_ABST
Patent Text Reader

Abstract

As an information processing method executed by a computer, a plurality of inference results as results of inferences for the same input data respectively performed by a plurality of inference machines is acquired (S11), an influence of the input data on the acquired plurality of inference results is acquired for each of the inference results (S12), one or more combinations of the plurality of inference results are decided based on the plurality of inference results, the one or more combinations being combinations of the plurality of inference results with each other (S14), and the influences acquired for the inference results included in the same combination in the decided plurality of inference results are prompted to each other in parallel or in overlap via a prompting device (S15).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to information processing methods performed by computers, etc. Background Technology

[0002] Techniques for indicating the results of object detection captured in an image are proposed (e.g., see Patent Document 1 and Non-Patent Document 1).

[0003] Prior art literature

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2018-181273

[0006] Non-patent literature

[0007] Non-patent literature 1: Erik Bochinski et al, “High-Speed ​​tracking-by-detection without using image information”, 14 th IEEE International Conference onAdvanced Video and Signal Based Surveillance(AVSS), August 2017 Summary of the Invention

[0008] The problem the invention aims to solve

[0009] In previous proposed techniques, it was difficult to evaluate in parallel the behavior of multiple different inference engines that performed inference processes such as object detection.

[0010] This disclosure provides information processing methods, etc., that can evaluate the behavior of multiple different inference engines in parallel.

[0011] Problem-solving methods

[0012] One aspect of the information processing method disclosed herein is a method executed by a computer, which obtains multiple inference results as the result of inference performed by multiple inference engines on the same input data, obtains the effect of the input data on each of the multiple inference results for each of the multiple inference results, determines one or more combinations based on the multiple inference results, said one or more combinations being combinations of the multiple inference results with each other, and prompts, via a prompting device, prompts in parallel or overlapping the effects obtained for each of the multiple inference results included in the same combination.

[0013] Furthermore, one aspect of the information processing system disclosed herein includes: a reasoning result acquisition unit that acquires multiple reasoning results, which are the results of reasoning performed by multiple inference engines for the same input data; an input data influence acquisition unit that acquires the influence of the input data on each of the multiple reasoning results; a decision unit that determines one or more combinations based on the multiple reasoning results, the one or more combinations being combinations of the multiple reasoning results with each other; and an influence prompting unit that, via a prompting device, prompts in parallel or overlapping the influences acquired for each of the multiple reasoning results included in the same combination.

[0014] In addition to the methods and systems described above, these general or specific forms can also be implemented by devices, integrated circuits, or computer-readable CD-ROMs and other recording media, or by any combination of devices, systems, integrated circuits, methods, computer programs and recording media.

[0015] The effects of the invention

[0016] The information processing methods disclosed herein enable the parallel evaluation of the behavior of multiple different inference engines. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of an information processing method that can be implemented.

[0018] Figure 2 This is a schematic diagram used to illustrate the differences in object detection results among different object detection models.

[0019] Figure 3A This is a schematic diagram representing the object detection results based on three object detection models for a common input image.

[0020] Figure 3B This is a diagram representing an example of the data from the object detection result set output by the object detection model.

[0021] Figure 4A This is an illustrative representation of applying the information processing method of the embodiment to... Figure 3A The image shows an example of the state after the object detection results are displayed.

[0022] Figure 4B This represents an example of data from an object detection result set after applying the information processing method of the implementation method.

[0023] Figure 5A This is a diagram illustrating the result of the insertion of alternative data performed in the information processing method of the embodiment.

[0024] Figure 5B It means and Figure 5A The figure shows an example of the object detection result set data corresponding to the processing results.

[0025] Figure 6 This is used to illustrate the insertion of replacement data performed by the information processing method implemented in the embodiments, and... Figure 5A The examples shown are diagrams in different forms.

[0026] Figure 7A This shows a display example in a UI screen where the information processing method of the implementation method is applied.

[0027] Figure 7B This shows a display example in a UI screen where the information processing method of the implementation method is applied.

[0028] Figure 7C This shows a display example in a UI screen where the information processing method of the implementation method is applied.

[0029] Figure 7D This shows a display example in a UI screen where the information processing method of the implementation method is applied.

[0030] Figure 8 This is a block diagram illustrating an example of the structure of a computer that performs an information processing method according to an implementation method.

[0031] Figure 9A This is a schematic diagram illustrating the analysis frame obtained through the information processing method described above.

[0032] Figure 9B This is an example of data representing an object detection result set that includes data related to the analysis box described above.

[0033] Figure 10 This is a flowchart illustrating the steps of the information processing method of this embodiment. Detailed Implementation

[0034] (The insights that form the basis of this disclosure)

[0035] The inventors have discovered the following problems with the prior art described above.

[0036] For example, in object detection performed by multiple inference models on the same image, there may be cases where the number of detected subjects (hereinafter referred to as detected objects) differs. In such cases, when the creator of the inference model wants to analyze the detection results for each detected object across the inference models, simply finding the results for the same detected object from the outputs of each inference model is tedious, inconvenient, and inefficient. Furthermore, even if the number of detected objects is the same, sometimes the estimated position or the captured range of a detected object may differ between the inference models. Establishing a correspondence visually is tedious, error-prone, and inefficient.

[0037] In view of the problem of such inefficiency, one aspect of the information processing method of this disclosure is a method executed by a computer, which obtains multiple inference results as the result of inference performed by multiple inference engines on the same input data, obtains the effect of the input data on each of the multiple inference results for each of the multiple inference results, determines one or more combinations based on the multiple inference results, the more than one combination being a combination of the multiple inference results with each other, and prompts, via a prompting device, the effects obtained for each of the multiple inference results and the inference results included in the same combination are prompted side by side or overlapping.

[0038] Therefore, reasoning results from multiple inference models (inference engines) can be selected and combined, for example, to consider inference results about a common object. Furthermore, the influence of input data on the inference results included in this combination can be temporarily presented in parallel or overlapping. Thus, the behavior (i.e., the influence) of multiple different inference engines can be evaluated in parallel. As a result, the user's effort in analyzing inference results is reduced compared to the past. Therefore, users can analyze inference results more efficiently than before.

[0039] Alternatively, the plurality of inference engines may be object detectors, the plurality of inference results may be object detection result sets containing multiple object detection results, the influence obtained for each of the plurality of inference results may be the influence of the input data on the plurality of object detection results contained in the object detection result set, the plurality of combinations may contain the plurality of object detection results contained in the plurality of object detection result sets that are different from each other, and the influence of being prompted in parallel or overlapping manner may be the influence obtained for the plurality of object detection results contained in the same combination.

[0040] Therefore, the influence of the input data on the object detection results included in this combination is temporarily displayed side-by-side or overlapping. As a result, the time required for the user to analyze the object detection results is reduced compared to the past, thus enabling the user to perform object detection result analysis more efficiently than before. This is particularly significant in object detection, where the analysis time becomes enormous due to detecting multiple objects in each object detector.

[0041] Alternatively, each of the plurality of object detection results may include a category based on the detected object, and the categories included in each of the plurality of object detection results in each of the more than one combination are common.

[0042] Therefore, when performing object detection for multiple categories, even if the estimated positions of the detected objects are close, temporary prompts to the user can be avoided if they belong to different categories. As a result, the time required for the user to analyze the object detection results is reduced compared to the past, allowing for more efficient analysis of object detection results.

[0043] Here, for example, each of the plurality of object detection results may include a detection box, and the combination of the plurality of results may be determined based on the overlap or positional relationship of the detection boxes contained in the plurality of object detection results contained in mutually different sets of the plurality of object detection results. Alternatively, for example, each of the plurality of object detection results may include a detection likelihood, and the combination of the plurality of results may be further determined based on the proximity of the detection likelihoods contained in the plurality of object detection results contained in mutually different sets of the plurality of object detection results.

[0044] Therefore, it is possible to make the combination of detection results for a common object more reliable from the object detection results of each object detector.

[0045] Alternatively, in the indication of effects, if the same combination does not contain the multiple object detection results in the set of object detection results generated by one of the multiple object detectors, the effects and alternative data obtained for each of the multiple object detection results included in the same combination can be displayed side by side. Furthermore, if there are isolated object detection results among the multiple object detection results that are not included in any of the more than one combination, the effects and alternative data obtained for the isolated object detection results can be displayed side by side.

[0046] Therefore, when the number of objects detected by the object detectors differs, the user can easily determine which object detector did not detect which object, or which object detector detected it.

[0047] Alternatively, the impact prompt may indicate the impact on the reference object detection result and the impact on each of the plurality of object detection results included in the combination containing the reference object detection result, wherein the reference object detection result is one of the plurality of object detection results included in a reference object detection result set as one of the plurality of object detection result sets. Furthermore, it may also involve accepting an operation to select the reference object detection result set, switching the reference object detection result set to the object detection result set selected through the operation.

[0048] This allows users to focus on specific object detectors and analyze object detection results more efficiently than before. Furthermore, it enables more efficient analysis when the object detector being analyzed is changed.

[0049] Alternatively, the operation of selecting the input data can also be accepted, switching the input data to the input data selected by the operation, and in the effect prompt, indicating the effect obtained by each of the plurality of inference engines for the selected input data, respectively, on the plurality of inference results.

[0050] Therefore, even with multiple input data, users can analyze the object detection results of each object detector more efficiently than before.

[0051] Alternatively, it may also accept the operation of selecting a group of object detection results, switching the object detection results opposite to the prompted influence or the prompted object detection results to the object detection results of the group selected by the operation.

[0052] As a result, users can select object detection results related to the prompts in the UI based on the attributes, and can analyze the object detection results more efficiently than before.

[0053] Alternatively, the prompt regarding the effect may also include information about the inference engine among the plurality of inference engines that outputs the inference results respectively related to the prompted effect.

[0054] Therefore, users can determine which of the multiple inference engines is responsible for the notification.

[0055] Furthermore, one aspect of the information processing method disclosed herein is a computer-executed method that obtains multiple object detection result sets, each of which contains multiple object detection results performed by multiple object detectors on the same input data. Based on the multiple object detection results contained in each of the multiple object detection result sets, one or more combinations are determined. The one or more combinations are combinations of one of the multiple object detection results contained in each of the multiple object detection result sets. The object detection results contained in the same combination are presented side-by-side or overlappingly in the multiple object detection results contained in each of the multiple object detection result sets via a prompting device.

[0056] This allows for the parallel evaluation of the behavior (i.e., object detection results) of multiple different object detectors. As a result, users can more efficiently compare the object detection results of each object detector compared to previous methods.

[0057] Furthermore, one aspect of the information processing system disclosed herein includes: a reasoning result acquisition unit that acquires multiple reasoning results, which are the results of reasoning performed by multiple inference engines for the same input data; an input data influence acquisition unit that acquires the influence of the input data on each of the multiple reasoning results; a decision unit that determines one or more combinations based on the multiple reasoning results, the one or more combinations being combinations of the multiple reasoning results with each other; and an influence prompting unit that, via a prompting device, prompts in parallel or overlapping the influences acquired for each of the multiple reasoning results and for the reasoning results included in the same combination.

[0058] Therefore, reasoning results from multiple reasoning models (inference engines) are selected and combined, for example, those concerning a common object. Furthermore, the influence of input data on the reasoning results included in this combination is temporarily presented, either in parallel or overlapping. Thus, the behavior of multiple different inference engines can be evaluated in parallel. As a result, the user's effort in analyzing reasoning results is reduced compared to the past. Therefore, users can analyze reasoning results more efficiently than before.

[0059] In addition to the methods and systems described above, these general or specific forms can also be implemented by devices, integrated circuits, or computer-readable CD-ROMs and other recording media, or by any combination of devices, systems, integrated circuits, methods, computer programs and recording media.

[0060] Hereinafter, embodiments of an information processing method and information processing system according to one aspect of the present disclosure will be described with reference to the accompanying drawings. The embodiments shown herein represent a specific example of the present disclosure. Therefore, the numerical values, shapes, constituent elements, configurations and connection methods of the constituent elements, as well as the steps (processes) and the order of the steps shown in the following embodiments are examples and do not limit the present disclosure. Furthermore, among the constituent elements in the following embodiments, constituent elements not described in the independent technical solutions are constituent elements that can be arbitrarily added. Additionally, the figures are schematic diagrams and are not necessarily rigorously illustrated.

[0061] (Implementation Method)

[0062] Figure 1 This is a schematic diagram of an information processing method applicable to various implementations. This screen functions as a user interface (hereinafter referred to as UI) for application software used in the analysis of the results of inference performed by an inference engine. This application software can be a web application that displays this screen (hereinafter referred to as the UI screen) on a web browser, or it can be a native application or a hybrid application. This application software is executed, for example, by a processor found in various business or personal computers such as tower, desktop, and tablet computers, and displays such a UI screen on a monitor used by the user.

[0063] UI screen 10 accepts user operations and displays the results of inference performed by the inference engine or information related to those results, according to the user's operation. Furthermore, this embodiment uses the case where the inference engine is an object detector that detects objects captured in an input image shown by the input data as an example for explanation. The information related to the inference results in this example refers to, for example, whether each part of the input data affects the results of the inference performed by the object detector, and if so, whether the effect is related to its magnitude and directionality (whether it is a positive or negative effect), or both. In other words, the effect is, for example, the reaction of the inference engine in the inference processing that outputs the inference results to the input data. Alternatively, the effect may also be represented as the analysis results of that reaction (i.e., analysis boxes or analysis values).

[0064] The UI screen 10 is divided into three parts from left to right: the input data bar, the model bar, and the result bar. The input data bar includes an input data selection section 20A, the model bar includes a model selection section 20B and a result display selection section 20C, and the result bar includes a result display section 40, a display data switching section 50A, a unified analysis box switching section 50B, and an individual analysis box switching section 50C.

[0065] The input data selection unit 20A prompts the user to select from candidates of input data that are overlaid in the results bar, displaying either the inference results or the analysis results (hereinafter, when referring to inference results and analysis results interchangeably, they are also called result information). In this example, the input data is image data, and the input data is prompted as a thumbnail of the image. The image selected by the user is displayed in... Figure 1 The example shown is a thumbnail of the entire image, surrounded by a thick frame. Hereafter, the input data will also be referred to as the input image.

[0066] The model selection unit 20B displays candidate inference engines with result information, allowing the user to select one. Here, "inference engine" refers to a machine learning inference model (hereinafter, for convenience, it will be simply referred to as "model"). Figure 1 In the example shown, the user selects "Model A" and "Model E" from the candidate inference models after narrowing down the inference models used for object detection. Furthermore, even if the multiple models selected by the model selection unit 20B share a common purpose, they differ in type, training dataset, or training volume due to differences in training methods or network structures.

[0067] The result selection unit 20C allows the user to select the items of result information to be displayed. In the case of object detection, the user selects, for example, the type (category) of the detected object as an item of result information. Figure 1 In the example shown, the available categories were cars, bicycles, pedestrians, and motorcycles; pedestrians were selected. Additionally, in Figure 1 In the example shown, other items in the displayed results information also allow the user to select the type of result (TP: True Positive, FP: False Positive, FN: False Negative). Figure 1 In the example shown, TP and FN were selected.

[0068] When a user finishes making selections in the input data selection section 20A, model selection section 20B, and display result selection section 20C, and clicks or touches the display button below the display result selection section 20C, the result information is displayed in the result display section 40 of the result bar according to the selected content. Figure 1 In the example, in the result display unit 40, the input image selected by the input data selection unit 20A, information related to the object detection result as the reasoning result for the input image, and the name of the object detection model that outputs the object detection result are displayed together.

[0069] exist Figure 1In the example shown, two display images, one for each model (Model A, Model E) selected by the model selection unit 20B and the other for each model, are displayed horizontally on the result display unit 40, respectively. Alternatively, the display images can also be displayed vertically.

[0070] The result display unit 40 displays which input image was selected by the input data selection unit 20A, and the display data switching unit 50A can be used to switch between them. Figure 1 In the example shown, the display data switching unit 50A is a slider that switches the input image displayed on the result display unit 40 according to the position of the slider moved by the user. Along with the switching of the input image, the result information superimposed on the input image is also switched. Furthermore, in this example, the name of the input image currently displayed on the result display unit 40 and its order among the selected input images are displayed on the display data switching unit 50A.

[0071] The unified analysis box switching unit 50B and the individual analysis box switching unit 50C allow the user to switch analysis boxes. Analysis boxes are displayed in the result display unit 40 as part of the display image, superimposed on the input image, and are a form that influences the inference results. The analysis box to be switched is superimposed on the input image to generate the display image. For example, when each model performs object detection on an input image and detects multiple objects, an analysis box is generated for each of the multiple detection results (i.e., detection boxes). However, in the result display unit 40, what is temporarily superimposed on each input image is one of the multiple analysis boxes generated from the multiple detection boxes. Figure 1 In the example, when the user clicks or taps the up-and-down pointing triangle (hereinafter, both click and tap operations are referred to as "press") of the analysis box unification switching unit 50B, which has only one component, in the results bar, the analysis boxes superimposed on the input image displayed on the results display unit 40 are switched sequentially in the analysis boxes generated for the detection results of multiple models. Furthermore, the individual analysis box switching unit 50C, which functions as a slider, is set in the results display unit 40 to display images for each input image superimposed with an analysis box related to each model. When the user moves the slider, the analysis boxes superimposed on the input image can be switched individually for each image.

[0072] By applying the information processing method of this embodiment to the UI screen 10, the UI screen 10 can be used to perform more efficient analysis of inference results. Alternatively, users can use a UI screen that allows them to operate in the same way as before, and perform more efficient analysis of inference results.

[0073] It is not uncommon for discrepancies like those described above to exist between inference results from multiple models based on the same input data. For example, in object detection, the number of detected subjects (hereinafter also referred to as detected objects), the estimated location within the image, or the captured range may differ between models. Figure 2 This is a schematic diagram used to illustrate the differences in object detection results between such models. Figure 2 The image shows bounding boxes enclosing pedestrians or vehicles as detected objects in the input image Image1, based on two models, Model A and Model B, respectively. When comparing the two object detection results, the number of detected pedestrians differs, although the bounding boxes for pedestrians considered to be common are similar in size and extent but not identical. Furthermore, the third bounding box from the top in Model B encloses a traffic cone on the road that was mistakenly detected as a pedestrian or vehicle, an example of result type FP. Moreover, a vehicle detected in Model A and located at the right end of the image was not detected in Model B. Additionally, these bounding boxes are arranged from top to bottom according to the order in which they appear in the detection results output by each model, but their order is as follows... Figure 2 As shown, there are sometimes differences between models.

[0074] The user uses UI screen 10 to compare and analyze the object detection results output by multiple models for each detected object. At this time, a request may arise to arrange and temporarily display the results for common detected objects on UI screen 10. In the prior art, the user uses the individual analysis frame switching unit 50C to display the results found from the object detection results output by each model on the results display unit 40. These found results are considered to be the detection results for common detected objects, such as pedestrians captured near the left end of an image. Alternatively, although no operation unit for this is shown in UI screen 10, the unified analysis frame switching unit 50B is used after rearranging and aligning the object detection results output by each model in the order desired by the user.

[0075] The information processing method of this embodiment can be said to save users time in rearranging. The following uses examples to describe in detail the rearrangement based on this method.

[0076] Figure 3AThis diagram schematically illustrates the object detection results based on three object detection models for a common input image. Solid rectangles represent the entire region of the input image used for object detection. Furthermore, the dashed and dotted-line boxes within the solid rectangles are the detection boxes obtained by each model as a result of the object detection process, and the differences in line style indicate, for example, detection boxes for different categories of detected objects such as pedestrians and vehicles. The vertical arrangement of these detection boxes corresponds to the order of the data related to the detection boxes within the object detection results output by each object detection model. Figure 3B This is an example of object detection data output from Model A, which is the model itself. Rows 2 through 6 of this data contain output values ​​related to the detection boxes. The content of these output values ​​indicates... Figure 3A The object detection results from Model A contain the position and size (range) of each bounding box within the entire image region, as well as the category of the detected object. The same data is also output from Models B and C, which are other object detection models. Furthermore, when multiple object detection results output by each object detection model are aggregated, this is hereinafter referred to as the object detection result set. That is, Figure 3B The illustrated data is a set of object detection results consisting of 5 bounding boxes, obtained by Model A performing object detection on one input image. However, Figure 3B This example illustrates the structure and values ​​of such data; the format and values ​​of the object detection result set are not limited to this example. Furthermore, the items included in each object detection result are not limited to this example. For instance, the types of results described above can also be included in the object detection results. For the purposes of subsequent explanations... Figure 4B , Figure 5B and Figure 9B The same applies.

[0077] When the input image is selected as the display object in the input data selection unit 20A of the UI screen 10, the object detection results (detection boxes) contained in the display image are switched in the vertical order of the image through the operation of the analysis box unified switching unit 50B and the analysis box individual switching unit 50C in the result display unit 40.

[0078] Figure 4A This is an illustrative representation of applying the information processing method of this embodiment to... Figure 3A The image shows an example of the state after object detection. Figure 4AIn the example shown, the bounding boxes, which are the object detection results based on each object detection model, are sorted according to the category of the detected object. This reflects how, through this information processing method, within the data of the object detection result set output by each object detection model, the output values ​​of each bounding box are rearranged according to the category of the detected object ("vehicle", "pedestrian"). For example, in Figure 3B The data shown is from the object detection result set output by Model A, such as... Figure 4B As shown, the output values ​​of the detection boxes in rows 2 through 6 are rearranged in order of the detected object's category ("vehicle", "pedestrian"). The same rearrangement is performed on the object detection result sets output from Models B and C.

[0079] Figure 5A This diagram illustrates the result of processing performed after the sorting based on the category of the detected object in the information processing method of this embodiment. For Figure 5A The results shown are consistent with Figure 4A The differences between the states shown will be explained in detail.

[0080] When reference Figure 5A In this case, among the object detection result sets output by the object detection models of Model A, B, and C, the detection results for common objects (or, the detection results with a high probability of being common objects) are placed in the same position (order) in the vertical direction.

[0081] In this arrangement of object detection results, the object detection results in the object detection result sets output by any one model are rearranged based on their order. Then, the order of the object detection results in the result set output by the model that is not the reference for the order can be determined based on the overlap of detection boxes of the same category within the object detection result sets output by the two models. This overlap can be expressed as IoU (Intersection over Union), which is the ratio of the intersection to the union of detection boxes. Alternatively, more simply, the order can be determined based on the size of the overlapping portions of detection boxes of the same category.

[0082] Furthermore, as another example of a method for determining the order of object detection results within an object detection result set, the order can be determined based on the positional relationship of detection boxes of the same category among the object detection result sets output by multiple object detection models. As a specific example of based on the positional relationship of detection boxes, the distance between corresponding vertices or corresponding edges of two detection boxes, or the distance between the geometric centroids of two detection boxes, can be used.

[0083] Furthermore, as mentioned above, the combination of object detection results (or inference results) in this disclosure can be determined by placing the detection results for each detected object in the same order within the object detection result sets output by multiple object detection models. Alternatively, it could be expressed as follows: Figure 5A Two or three object detection results contained in the same row are included in the same group. Object detection results included in the same group are temporarily displayed in the results display section 40, even if the user does not search or rearrange them. Figure 5A In the example shown, the detection boxes that are the object detection results based on each model and placed in the same row are detection boxes of the same category, and have a larger IoU or are closer in position within the image compared to detection boxes in other rows.

[0084] Whether to include detection results from different models in the same combination can be determined by comparing overlap or distance with a predetermined threshold, based on factors such as IoU (Interchange of Value) and the positional relationship of the detection boxes. Even detection results from boxes of the same level with an IoU greater than those in other rows are excluded from the same combination if the IoU is below a predetermined threshold. Alternatively, the overall optimization can be achieved using algorithms such as the Hungarian Algorithm to determine this.

[0085] Furthermore, the method by which the computer executing the information processing method of this embodiment selects the object detection result set (hereinafter also referred to as the reference object detection result set) as the basis for rearranging the object detection results within each object detection result set is not particularly limited. For example, a user can select the reference object detection result set. Furthermore, for example, the object detection result set with the fewest number of detection results (detection boxes) or the object detection result set with the most number of detection results included in the object detection results can be selected as the reference object detection result set. Furthermore, for example, random selection can be performed during rearrangement. Figure 5A The example shown is a process performed in the information processing method of this embodiment, in which the order of object detection results in the object detection result set of Model A is used as a reference.

[0086] also, Figure 5A The results shown include Figure 4A The rectangle with attached dots is not shown in the diagram. This represents data (hereinafter referred to as substitute data) that represents alternative information inserted into the object detection result set based on the model that did not detect the detected object, in cases where there is a detected object in only a portion of the models selected by the model selection unit 20B. Figure 5BThis is a diagram representing an insertion example of the object detection result set output to Model A, which represents the alternative information. Figure 5B The "{blank}" in lines 4 and 5 are examples of inserted alternative information. This information corresponds to... Figure 5A The two rectangles of the sub-dots seen in the column of the object detection results based on Model A. Alternative data is displayed side-by-side with the display images of object detection results based on other models in the results display unit 40, providing the user with information that the detected object was not detected in that model. Such alternative data in the results display unit 40 may be displayed as, for example, a box with a blank center or a simple blank area without a box, or as characters, symbols, patterns, or graphics conveying the message that the object was not detected.

[0087] also, Figure 6 It is used to explain and Figure 5A Examples of different insertions replace the form of the data in the graph. In this example, the order of object detection results is based on the order of object detection results within the object detection result set based on Model A. However, in Figure 6 In this context, the replacement data is not inserted into the object detection result set of Model A, and is maintained in... Figure 4A The sorted state is shown below. Detection results of objects not detected in Model A, which are included in the object detection result sets of Model B and Model C, are grouped at the end (bottom) of the sort order.

[0088] For example, when a user presses the display button, the object detection results in the object detection result set can be automatically rearranged (sorted). Alternatively, the UI screen 10 may also have an operation component (not shown) that indicates the user to perform such a rearrangement, and the rearrangement is performed based on the operation of the operation component.

[0089] also, Figure 5A The 4th and 6th from above Model A and Figure 6 The fourth one from above Model A corresponds to the detection result of the object detected by only one of the three models (hereinafter referred to as the isolated object detection result). The isolated object detection result is not included in any combination of object detection results determined as described above. The display image based on such isolated object detection results can also be displayed in the result display unit 40 alongside the alternative data.

[0090] An example of the rearranged UI screen 10 is shown below. Figures 7A to 7D .exist Figures 7A to 7DThe image shown here is a portion of the results display section 40 in the results bar of the UI screen 10 and the unified switching section 50B of the analysis box. For the image of the object being detected, refer to... Figure 2 Image 1. In this example, it is assumed that three models, Model A, Model B, and Model C, perform object detection on image 1.

[0091] When reference Figure 7A In this state, a display image showing the analysis results detected in one of the three models, including the detection results of a pedestrian mapped from slightly to the left of the center, is shown side by side. In the examples of this disclosure, the analysis results (magnitude and directionality of influence) referred to as analysis boxes are schematically represented by patterned regions that overlap with the image. Parts without overlapping analysis boxes indicate that this part of the input image has no or minimal influence on the object detection results.

[0092] Figure 7B Indicates in Figure 7A In the state shown, the result display unit 40 is in the state after pressing the downward-pointing triangle contained in the analysis box unification switching unit 50B once. A display image is shown, which contains the analysis results of the detection results of a pedestrian near the left end of either input image, from the object detection result sets output by Model A and Model B respectively. Additionally, in Model C, since the pedestrian was not detected, a box containing a graphic as substitute data is displayed alongside other display images.

[0093] Figure 7C Indicates in Figure 7B In the state shown, the result display unit 40 is activated by pressing the downward-pointing triangle included in the analysis box unification switching unit 50B once further. In this state, a display image containing the analysis results for pedestrian detection results only included in the object detection result set output by Model B is shown side by side, along with boxes containing graphs as alternative data for Model A and Model C.

[0094] Figure 7D Indicates in Figure 7C In the state shown, the result display unit 40 is activated by pressing the downward-pointing triangle contained in the analysis box unification switching unit 50B once more. In this state, two display images containing the analysis results of the detection results of the car, which are included in both the object detection result set based on Model A and the object detection result set based on Model C, are shown side by side, along with a box containing a graph as alternative data for Model B.

[0095] Thus, in the information processing method of this embodiment, the overlap or positional relationship of the detection boxes of object detection results included in the object detection result set is determined based on the combinations of each other contained in the inference results of multiple models. In the object detection processing based on each model in such determined combinations, the user is prompted in parallel with the effects of the input image on the object detection results included in the same combination. As a result, the user can view the detection results for the same object simultaneously for analysis without having to search for detection results for the same object or rearrange the searched detection results.

[0096] Next, the structure of the computer that performs such a rearrangement and the information processing method of this embodiment for the computer to perform the rearrangement will be described. Figure 8 This is a block diagram illustrating a structural example of a computer 100 that performs the information processing method of this embodiment.

[0097] Computer 100 is any of the aforementioned computers, and includes an information processing device 80 and a display device 60.

[0098] The information processing apparatus 80 comprises a storage unit storing the aforementioned application software and a processing unit for reading and executing the application software. The information processing apparatus 80 includes an analysis unit 81, a synthesis unit 82, a prompting object selection unit 83, and a prompting unit 84, which are functional components provided by executing the application software.

[0099] The analysis unit 81 calculates the presence (and in some cases, magnitude and orientation) of the influence caused by each part of the image (hereinafter referred to as the input image) of the object being processed as an object in the object detection result set output by the object detection model for each of the multiple object detection results (i.e., each detection box). Such influence is calculated using various methods. For example, the influence can be calculated using methods disclosed in the following documents. The analysis unit 81 is an example of the inference result acquisition unit and the input data influence acquisition unit in this embodiment. Document: Denis Gudovskiy, Alec Hodgkinson, Takuya Yamaguchi, Yasunori Ishii, and Sotaro Tsukizawa; "Explain to Fix: A Framework to Interpret and Correct DNN Object Detector Predictions"; arXiv: 1811.08011v1; November 19, 2018.

[0100] Figure 9AThis is a schematic diagram illustrating the calculations of the effects performed by the analysis unit 81. Figure 9A In the input image, the dashed boxes represent the detection boxes that constitute the object detection results. The left side of the input image is an extract of the regions containing the detection boxes, and the squares represent the pixels that make up the input image (some omitted). The values ​​within the squares are values ​​calculated by the analysis unit 81, representing the influence of each pixel's value on the object detection results (hereinafter also referred to as analysis values). Furthermore, in... Figure 9A In the above text, for ease of explanation, analysis values ​​are input for each pixel, but this does not mean that such data is appended to the values ​​that each pixel already possesses. Furthermore, the above describes the case of calculating analysis values ​​per pixel, but it is not a limitation. Based on grouping multiple pixels and setting them as superpixels, analysis values ​​can be calculated per superpixel.

[0101] The range of pixels for calculating the analysis values ​​varies depending on the design values ​​of the object detection model and the size of the detection box. Figure 9A In the example shown, the analysis value is calculated by extending 2 pixels upwards and downwards and 4 pixels horizontally from the pixel overlapping with the detection box. The analysis box described above corresponds to the set of pixels whose influence is calculated as a whole, or the outline of the set of pixels within a specific range whose influence is calculated.

[0102] Figure 9B This is a diagram illustrating an example of object detection result set data to which the analysis values ​​calculated by the analysis unit 81 are appended to the object detection results. In this example, in Figure 3B The object detection result set output by Model A, as shown in the figure, has been inserted with data such as... Figure 9A The data of the analytical values ​​calculated by the analysis unit 81. Figure 9B In the data set, the second line contains data indicating the position and size of the detection box and the category of the object being detected with that detection box attached. The third line contains data of the analysis value calculated by the analysis unit 81 for that detection box (omitted in the middle). The data following the fourth line contains data of the position and size of other detection boxes included in the object detection result set, the category of the object being detected, and the analysis value calculated for that detection box.

[0103] The synthesis unit 82 synthesizes and outputs an image, i.e., a display image, which is an image in which the analysis box is superimposed on the input image, based on the analysis values ​​calculated by the analysis unit 81. In this disclosure, the analysis box is composed of, for example... Figure 1 and Figures 7A to 7D The regions with patterns, as illustrated, are used to represent the input image. In the case of multiple detection boxes in a single input image, the display image is synthesized for each detection box. This display image summarizes a single input image and the analysis results calculated by the analysis unit 81, which represent the influence of the input image on the object detection results performed on that input image.

[0104] The prompting object selection unit 83 selects a display image to be displayed on the results display unit 40 via the prompting unit 84 based on input from the operation unit 2050 included in the UI screen 10 displayed on the display device 60 such as a monitor. The operation unit 2050 is a component on the UI screen 10 that accepts the user's operation of selecting a group of object detection results, which includes the analysis results in the display image displayed on the results display unit 40. This group is defined by attributes of the object detection results, such as the category of the detected object and the result type. In this embodiment, the operation unit 2050 includes an input data selection unit 20A, a model selection unit 20B, a display result selection unit 20C, a display data switching unit 50A, an analysis frame unification switching unit 50B, and an analysis frame individual switching unit 50C. The input from the operation unit 2050 includes information about the input image and model selected by the user, as well as information about the group of object detection results, that is, information about the category of the detected object or the result type. The prompting object selection unit 83 selects an image to be displayed on the results display unit 40 from the display images output by the synthesis unit 82 according to this information. Furthermore, the object selection unit 83 rearranges the selected display images within each object detection result set and determines the combination. Additionally, in object detection result sets lacking detection results for objects common to other models, substitute information is inserted according to settings or specifications. The object selection unit 83 is an example of the decision unit in this embodiment.

[0105] The prompting unit 84 outputs the display image included in the same combination determined by the prompting object selection unit 83 to the display device 60 for side-by-side display on the result display unit 40. Alternatively, depending on the settings or specifications, alternative data is output in a manner that displays it side-by-side with the display image on the result display unit 40. The prompting unit 84 and the result display unit 40 of the display device 60 are examples of influence prompting units in this embodiment. Furthermore, the display device 60 is an example of a prompting device in this embodiment.

[0106] Figure 10 This is a flowchart illustrating the steps of the information processing method of this embodiment executed by a computer 100 having this structure. Furthermore, the processing performed by each functional component described above will be explained below as processing performed by the computer 100.

[0107] Computer 100 obtains an object detection result set, which is obtained by different object detection models performing object detection on a common input image and outputting the results (step S11). This object detection result set includes output values ​​related to each detection box that serves as the object detection result (refer to...). Figure 3BFurthermore, the components that actually acquire the input image data and perform object detection based on the object detection model on that data can be a computer 100 or a device other than a computer 100. For example, it can be an information processing device in an information processing terminal such as a smartphone or a camera that performs object detection on images captured by the imaging elements of each device.

[0108] Next, the computer 100 calculates the influence of each part of the input image on the calculation of the output value related to the detection box, which is included in the object detection result set. Specifically, the calculation of the influence refers to calculating an analytical value representing the magnitude (including presence or absence) and directionality of the influence, or both of them.

[0109] Next, the computer 100 synthesizes a display image that overlays the effects calculated in step S12 onto the input image (step S13). Specifically, based on the calculated analysis values, a display image obtained by overlaying the analysis box onto the input image is synthesized.

[0110] Next, the computer 100 determines the combination of detection results (detection boxes) included in the object detection result sets output by different object detection models (step S14). Here, the detection results included in the same combination are those that are determined to be detection results for a common object based on the overlap or positional relationship between the detection boxes, or those that are determined to be detection results for a common object with a probability higher than a predetermined benchmark. In addition, during the process of determining this combination, alternative information may be inserted into the object detection result set according to settings or specifications.

[0111] Finally, the computer 100 prompts for the selection of display images according to the settings of the groups of object detection results defined by attributes such as category and result type (step S15). Since the selection is performed in units of the groups determined in step S14, display images containing analysis results of detection results for common objects (or objects that are highly likely to be common objects) are presented side by side.

[0112] Furthermore, the above-described process is just one example, and the process of the information processing method in this embodiment is not limited to this. For example, the composition of the display image (S13) or the acquisition of the effect (S12) and the composition of the display image (S13) can be performed after the combination decision (S14) and before the display image is actually prompted. In this case, the input image included in the display image displayed on the result display unit 40 is first selected based on the object detection result. Then, the acquisition of the effect and the composition of the display image can be performed by combining only the object detection result for the selected input image with the object. In addition, user operations related to the group setting on the UI screen 10 can be accepted at any stage before step S15. According to this setting, the process that becomes the display object becomes the processing object in the unexecuted process. Then, if the user operation related to the group setting is accepted again after step S15, the display image in the result display unit 40 is switched to the display image based on the object detection result selected according to the latest group setting. Alternatively, the combination decision (step S14) can be performed and the display image can be temporarily displayed directly. Then, step S14 can be performed, for example, after receiving a rearrangement request from a user, and the display content on the result display unit 40 can be updated according to the result. Alternatively, the display content reflecting the result of step S14 and the display content not reflecting it can be reversibly switched.

[0113] Furthermore, in the above description, the representation of the combination of object detection results (detection boxes) is used. However, in this disclosure, the representation also includes the combination of input images or display images based on the overlap or positional relationship of the detection boxes.

[0114] (Variations and other supplementary matters)

[0115] The information processing methods of one or more forms disclosed herein are not limited to the embodiments described above. Various modifications conceived by those skilled in the art to the above embodiments, without departing from the spirit of this disclosure, can also be included in the forms of this disclosure. Examples of such modifications and other supplementary matters to the description of the embodiments are given below.

[0116] (1) In the above embodiment, the calculation of the effect is performed by the computer 100, but it is not limited thereto. The calculation of the effect may also be performed in a device other than the computer 100 equipped with an information processing device, and the computer 100 may also directly or indirectly receive and obtain the input of the effect output from the device.

[0117] Furthermore, in the above embodiments, the representation of the influence is only illustrated by a box (analysis box) along the region or its outline, but is not limited to this. For example, it can also be represented by a point or a cross or other graphic placed at the center or geometric center of the region.

[0118] (2) In the description of the above embodiments and the accompanying drawings referred to therein, multiple display images that overlap multiple object detection results contained in the same combination and the effects obtained on these object detection results are displayed side by side in a row on the left and right sides of the result display unit 40, but the display method is not limited to this. For example, they can be displayed side by side in upper and lower columns, or they can be arranged in a matrix composed of multiple rows and multiple columns, such as in a grid pattern. Furthermore, for example, detection boxes that are object detection results of multiple models, or analysis boxes obtained for detection boxes, can be summarized and overlaid on a single input image. In this case, the analysis box may not be... Figures 7A to 7D The areas shown in the example are not patterned, but are represented only by outlines, or by areas painted with higher transparency.

[0119] (3) In the description of the above embodiments and the accompanying drawings referred to therein, an example is shown in which a display image is shown alongside the input image, indicating that the effect obtained for the object detection results included in the same determined combination is superimposed on the input image. However, the display image shown in the result display unit 40 is not limited to such a display image. For example, in the result display unit 40, the input image superimposed as the object detection result may be shown instead of the effect or in addition to the effect. Furthermore, the information superimposed on the input image may be switchable according to the user's operation.

[0120] (4) In the description of the above embodiments and the accompanying drawings referred to therein, for ease of explanation, only the object detection results contained in one combination and the effect obtained on the object detection results are temporarily displayed on the UI screen 10 in the result display unit 40, but it is not limited to this. As long as the display images based on the object detection results contained in the same combination are arranged together and surrounded by a frame, and the images are easy for the user to grasp, the display images corresponding to multiple combinations can also be temporarily displayed in the result display unit 40.

[0121] (5) In the description of the above embodiments, it was explained that the combination is determined by rearranging the contents of the object detection result set data and placing them in the same order, but the implementation of the combination is not limited to this. For example, instead of rearranging the contents of the object detection result set data, an identifier indicating the combination to which each object detection result belongs may be added, or a table that maintains combination information separately from the object detection result set may be generated or updated. Alternatively, the combination of display images of the quantities prompted in the result display unit 40 may be determined by performing calculations each time based on the user's operation.

[0122] (6) Not only when comparing inference results across multiple models, but also in order to organize inference results based on a single model, it is possible to perform the following: Figure 3A and Figure 4A The sorting described corresponds to the category of the object being tested.

[0123] On the other hand, even if the objects displayed are based on the inference results of multiple models or their analysis results, sorting according to categories is not necessary. That is, sorting according to categories can be omitted, and decisions can be made based on the overlap or positional relationship of the detection boxes.

[0124] (7) In use Figures 3A to 6 In the above description of the example embodiment shown, the combination of object detection results is determined based on the overlap or positional relationship between the detection boxes, but the method of determining the combination is not limited to this. For example, when the influence, i.e., the analysis box, has been calculated, the combination can also be determined based on the similarity between the influences. The similarity of the influences mentioned here can be determined based on the comparison of the size of the vectors or the similarity between these vectors (e.g., cosine distance), and the vectors are, for example, based on the size (vertical × horizontal) of the pixel region from which the analysis value is calculated by the analysis unit 81 for the object detection results output by the two object detection models respectively.

[0125] (8) In addition, the combination of object detection results can be determined not only based on the overlap or position of the detection boxes, but also based on the proximity of the likelihood of the detection results. For example, a higher score can be assigned to cases where the overlap of the detection boxes is greater and the likelihood of the detection results is closer, and the combination can be determined based on the score.

[0126] (9) Some or all of the functional components of the aforementioned information processing systems can be constituted by a single system LSI (Large Scale Integration). A system LSI is a multifunctional LSI manufactured by integrating multiple components onto a single chip. Specifically, it is a computer system comprising a microprocessor, ROM (Read-Only Memory), RAM (Random Access Memory), etc. The ROM stores the computer program. The microprocessor operates according to the computer program, thereby enabling the system LSI to realize the functions of each component.

[0127] Furthermore, while this is referred to as a system LSI, it is sometimes also called an IC, LSI, super LSI, or very large-scale LSI, depending on the level of integration. Additionally, the method of integrated circuitization is not limited to LSI; it can also be implemented using dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI fabrication, or reconfigurable processors that can reconfigure the connections or settings of the internal circuitry of the LSI, can also be used.

[0128] Furthermore, if advancements in semiconductor technology or other derived technologies lead to the development of integrated circuits that replace LSIs, then this technology can also be used for the integration of functional blocks. Applications such as biotechnology are also possible.

[0129] (10) One form of this disclosure is not limited to use Figure 10 The flowchart described above illustrates that the information processing method can also be used by a program executed by computer 100, or by an information processing system containing computer 100. Alternatively, one embodiment of this disclosure can be a computer-readable, non-transitory recording medium on which such a computer program is recorded.

[0130] Industrial availability

[0131] The information processing methods disclosed herein can be used in a user interface for comparing the results of computer reasoning processing.

[0132] Explanation of reference numerals in the attached figures

[0133] 10 UI Screens

[0134] 20A Input Data Selection Section

[0135] 20B Model Selection Department

[0136] 20C Display Results Selection Section

[0137] 40 Results Display Department

[0138] 50A Display Data Switching Unit

[0139] 50B Analysis Frame Unified Switching Section

[0140] 50C Analysis Frame Individual Switching Section

[0141] 60 display devices

[0142] 80 Information processing devices

[0143] 81 Analysis Department

[0144] 82 Synthesis Department

[0145] 83. Prompt for object selection

[0146] 84. Reminder Department

[0147] 100 computers

[0148] 2050 Operations Department

Claims

1. An information processing method, executed by a computer, wherein, For each of the multiple inference engines, obtain the detection results of multiple objects for the same image by that inference engine. For each of the plurality of object detection results from each of the inference engines, the following influence information is obtained: the influence information indicates whether the image has an influence on each of the plurality of object detection results, and, if there is an influence, indicates one or both of the magnitude and directionality of the influence. Based on the categories of the detected objects from the multiple object detection results of each of the inference engines, one or more combinations are determined. These combinations consist of object detection results of the same category from the multiple object detection results of each of the inference engines. The influence information obtained from the object detection results included in the same combination of the multiple object detection results of each inference engine is superimposed on the object detection results included in the same combination, and the object detection results included in the same combination with the superimposed influence information are presented side by side.

2. The information processing method according to claim 1, wherein, Each of the plurality of inference engines is an object detector.

3. The information processing method according to claim 1, wherein, The detection results for each of the multiple objects each contain a detection bounding box. The combination of one or more is further determined based on the overlap or positional relationship between the detection frames contained in the object detection results of each of the inference engines.

4. The information processing method according to claim 3, wherein, The detection results for each of the multiple objects include the likelihood of the detection. The combination of one or more objects is further determined based on the proximity of the likelihoods of the detections contained in the object detection results of each of the various inference engines to each other.

5. The information processing method according to claim 2, wherein, In the prompt of the impact information, if the multiple object detection results generated by one of the multiple object detectors do not include object detection results belonging to any category of any of the more than one combination, alternative data to replace the object detection results are prompted side by side.

6. The information processing method according to claim 2, wherein, Furthermore, if there is an isolated object detection result among the multiple object detection results that is not included in any of the more than one combinations, the impact information and alternative data obtained for the isolated object detection result are displayed side by side.

7. The information processing method according to claim 1, wherein, It also accepts the operation of selecting the image. Switch the image to the image selected through the operation.

8. The information processing method according to claim 2, wherein, It also accepts the option to select a combination of the object detection results. Switch the object detection result relative to the prompted impact information or the prompted object detection result to the combination of object detection results selected by the operation.

9. The information processing method according to any one of claims 1 to 8, wherein, The prompt in the impact information also indicates information about the inference engine among the plurality of inference engines that outputs the object detection results respectively related to the impact information being prompted.

10. An information processing system, wherein, have: The reasoning result acquisition unit acquires, for each of the multiple inference engines, the object detection results of that inference engine for the same image. The input data influence acquisition unit acquires influence information for each of the plurality of object detection results of each of the inference engines. The influence information is information indicating whether the image has an influence on each of the plurality of object detection results, and if there is an influence, information indicating one or both of the magnitude and directionality of the influence. The decision unit determines one or more combinations based on the categories of the detected objects in the multiple object detection results of each of the inference engines. The more than one combination is a combination of object detection results of the same category in the multiple object detection results of each of the inference engines. as well as The impact prompting unit, via a prompting device, overlays the impact information obtained from the object detection results included in the same combination of the plurality of object detection results of each of the inference engines with the object detection results included in the same combination, and then prompts the object detection results included in the same combination with the overlaid impact information side by side.

11. A computer-readable storage medium storing a program, said program causing the processor to: In an information processing apparatus having a processor, the processor, by execution by said processor, to: For each of the multiple inference engines, obtain the object detection results for the same image as that inference engine. For each of the plurality of object detection results from each of the inference engines, the following influence information is obtained: the influence information indicates whether the image has an influence on each of the plurality of object detection results, and, if there is an influence, indicates one or both of the magnitude and directionality of the influence. Based on the categories of the detected objects from the multiple object detection results of each of the inference engines, one or more combinations are determined. These combinations consist of object detection results of the same category from the multiple object detection results of each of the inference engines. The influence information obtained from the object detection results in the same combination of the multiple object detection results of each inference engine is superimposed on the object detection results in the same combination, and the object detection results in the same combination with the superimposed influence information are presented side by side.

Citation Information

Patent Citations

  • Image processing apparatus, method thereof, and program

    JP2018181273A

  • Methods and systems for performing sleeping object detection in video analytics

    US20180268563A1