Information processing device, information processing method, and program

The information processing apparatus improves object detection accuracy in images with multiple subjects by estimating detection difficulty and generating a display image, addressing the limitations of existing techniques.

WO2025134178A1PCT designated stage expired Publication Date: 2025-06-26NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/045224
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing techniques for improving object detection accuracy in images are limited when the image contains both the target object and other objects, as they do not effectively utilize images with multiple subjects.

Method used

An information processing apparatus and method that acquire images, estimate the difficulty of object detection for multiple types of objects in each region, and generate a display image indicating detection difficulty, using a machine-learned difficulty estimation model.

Benefits of technology

This approach enhances the accuracy of identifying target objects in images with multiple subjects by providing a visual representation of detection difficulty, thereby improving object detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023045224_26062025_PF_FP_ABST
    Figure JP2023045224_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises: an acquisition means for acquiring one or a plurality of images; an estimation means for estimating the degree of difficulty of object detection for a plurality of types of objects in each region included in an image acquired by the acquisition means; and a generation means that refers to the estimation results of the estimation means and generates a display image indicating the degree of difficulty of detecting the plurality of types of objects for each region. The information processing device assists a user in decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] There are known techniques for improving the accuracy of identifying a target object included as a subject in an image. For example, Patent Literature 1 discloses a technique for improving the accuracy of object detection in an image by using an image in which a background image and a detection target image are combined.

[0003] Japanese Patent Application Publication No. 2019-159630

[0004] The technology described in Patent Document 1 does not use images that show objects other than the target object, and therefore has the problem that in images that include both the target object and objects other than the target object as subjects, the degree to which the technology contributes to improving the accuracy of identifying the target object is limited.

[0005] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology for improving the accuracy of identifying a target object in an image that includes the target object and objects other than the target object as subjects.

[0006] An information processing device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring one or more images, an estimation means for estimating the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition means, and a generation means for generating a display image showing the difficulty of detection of the multiple types of objects for each region by referring to the estimation results by the estimation means.

[0007] An information processing device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring a training image, and a learning means for machine learning a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image, by referring to the training image.

[0008] An information processing method according to an exemplary aspect of the present disclosure includes an acquisition process for acquiring one or more images, an estimation process for estimating the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition process, and a generation process for generating a display image showing the difficulty of detection of the multiple types of objects for each region by referring to the estimation results of the estimation process.

[0009] An information processing method according to an exemplary aspect of the present disclosure includes an acquisition process for acquiring a training image, and a learning process for machine learning a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image by referring to the training image.

[0010] A program relating to an exemplary aspect of the present disclosure causes a computer to perform an acquisition process to acquire one or more images, an estimation process to estimate the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition process, and a generation process to generate a display image showing the difficulty of detection of the multiple types of objects for each region by referring to the estimation results of the estimation process.

[0011] A program relating to an exemplary aspect of the present disclosure causes a computer to perform an acquisition process to acquire a training image and a learning process to machine-learn a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image by referring to the training image.

[0012] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technology can be provided that improves the accuracy of identifying a target object in an image that includes the target object and an object other than the target object as subjects.

[0013] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing the configuration of an information processing device according to the present disclosure. FIG. 4 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 5 is a block diagram showing the configuration of an information processing device according to the present disclosure. FIG. 6 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 7 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 8 is a diagram showing an example configuration of learning data creation and learning processing according to the present disclosure. FIG. 9 is a diagram showing an example configuration of processing related to learning and inference according to the present disclosure. FIG. 10 is a diagram showing an example of overlaying an object detection difficulty image according to the present disclosure. FIG. 11 is a diagram showing an example presentation of an object detection difficulty image according to the present disclosure. FIG. 12 is a diagram showing an example presentation of an object detection difficulty image according to the present disclosure. FIG. 13 is a block diagram showing the configuration of a computer functioning as an information processing device according to the present disclosure.

[0014] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0015] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0016] (Overview of Information Processing Device 1) First, an overview of the information processing device 1 according to this exemplary embodiment will be described. The information processing device 1 is, as an example, a device that estimates the difficulty of object detection for multiple types of objects in each region included in an image.

[0017] Here, the object may be, for example, an object included as a subject in the image data, and may be classified by type of object.

[0018] (Configuration of information processing device 1) The configuration of the information processing device 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1. As shown in Fig. 1, the information processing device 1 includes an acquisition unit 11, an estimation unit 12, and a generation unit 13.

[0019] (Acquisition Unit 11) The acquisition unit 11 acquires one or more images. For example, the acquisition unit 11 may acquire one or more background images that serve as backgrounds for objects. Here, the background images may be included in a known background image dataset, for example.

[0020] (Estimation unit 12) The estimation unit 12 estimates the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition unit 11. The difficulty of object detection may be indicated by, for example, accuracy information of the object detection.

[0021] (Generation Unit 13) The generation unit 13 refers to the estimation result by the estimation unit 12 and generates a display image that indicates the detection difficulty of multiple types of objects for each region.

[0022] (Effects of Information Processing Device 1) As described above, the information processing device 1 employs a configuration including an acquisition means that acquires one or more images, an estimation means that estimates the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition means, and a generation means that generates a display image indicating the detection difficulty of the multiple types of objects for each region by referring to the estimation results by the estimation means. Therefore, the information processing device 1 has the effect of improving the accuracy of identifying a target object in an image that includes both the target object and an object other than the target object as subjects.

[0023] (Flow of Information Processing Method S1) The flow of information processing method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of information processing method S1. As shown in Fig. 2, information processing method S1 includes an acquisition process (step) S11, an estimation process (step) S12, and a generation process (step) S13.

[0024] (Step S11) In step S11, the acquisition unit 11 acquires one or more images.

[0025] (Step S12) In step S12, the estimation unit 12 estimates the difficulty of object detection for a plurality of types of objects in each region included in the image acquired by the acquisition unit 11.

[0026] (Step S13) In step S13, the generating unit 13 refers to the estimation result by the estimating unit 12 and generates a display image indicating the detection difficulty of a plurality of types of objects for each region.

[0027] (Effects of Information Processing Method S1) As described above, information processing method S1 employs a configuration including an acquisition process that acquires one or more images, an estimation process that estimates the difficulty of object detection for multiple types of objects in each region included in the images acquired by the acquisition process, and a generation process that generates a display image indicating the detection difficulty of the multiple types of objects for each region by referring to the estimation results of the estimation process. Therefore, information processing method S1 has the effect of improving the accuracy of identifying a target object in an image that includes both the target object and an object other than the target object as subjects.

[0028] [Second Exemplary Embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0029] (Overview of Information Processing Device 2) First, an overview of the information processing device 2 according to this exemplary embodiment will be described. As an example, the information processing device 2 is a device that performs machine learning to generate a model that estimates the difficulty of object detection for multiple types of objects in each region included in an image.

[0030] (Configuration of information processing device 2) The configuration of the information processing device 2 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 2. As shown in Fig. 3, the information processing device 2 includes an acquisition unit 21 and a learning unit 22.

[0031] (Acquisition Unit 21) The acquisition unit 21 acquires learning images.

[0032] (Learning Unit 22) The learning unit 22 performs machine learning on a difficulty estimation model that estimates the difficulty of object detection for a plurality of types of objects in each region included in the target image, by referring to a learning image.

[0033] (Effects of Information Processing Device 2) As described above, the information processing device 2 employs a configuration including an acquisition unit that acquires training images and a learning unit that performs machine learning, by referring to the training images, on a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image. Therefore, the information processing device 2 has the effect of improving the accuracy of identifying a target object in an image that includes a target object and an object other than the target object as subjects.

[0034] (Flow of Information Processing Method S2) The flow of information processing method S2 will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing the flow of information processing method S2. As shown in Fig. 4, information processing method S2 includes an acquisition process (step) S21 and a learning process (step) S22.

[0035] (Step S21) In step S21, the acquisition unit 21 acquires learning images.

[0036] (Step S22) In step S22, the learning unit 22 performs machine learning on a difficulty estimation model that estimates the difficulty of object detection for a plurality of types of objects in each region included in the target image, by referring to the learning image.

[0037] (Effects of Information Processing Method S2) As described above, information processing method S2 employs a configuration including an acquisition process for acquiring training images, and a learning process for machine learning a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image, by referring to the training images. Therefore, information processing method S1 has the effect of improving the accuracy of identifying a target object in an image that includes a target object and an object other than the target object as subjects.

[0038] [Third Exemplary Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0039] (Overview of Information Processing Device 1A) Information processing device 1A is, for example, a device that estimates the difficulty of object detection for multiple types of objects in each region included in an image. Specific examples of object types include people, suitcases, doctors in medical settings, patients, and IV stands.

[0040] For example, consider a case where it is desired to improve the detection rate of an object in a background image. In this case, for example, in the past, a model was first trained using images in which an object is placed in each region included in the background image, and the difficulty of detecting the object for each region included in the background image was estimated. Then, in the past, for example, the model was retrained using images in which the object is placed in a region where the detection difficulty exceeds a predetermined threshold, thereby improving the detection rate of the object in the background image.

[0041] However, in the past, for example, when an image containing an object different from the object for which the detection difficulty is to be estimated is used, it can be difficult to estimate the difficulty of detecting the object.

[0042] As a specific example, consider a case where an abandoned bag is to be detected in a public space where many people pass by, such as a train station or an airport. Here, for example, the object to be detected may be a person and baggage. In this case, in the past, even if an attempt was made to improve the baggage detection rate and estimate the baggage detection difficulty for each area included in an image where many people pass by, it was difficult to estimate the baggage detection difficulty when an object such as a person was captured in the foreground of the baggage.

[0043] For example, the information processing device 1A is a device that can estimate the difficulty of object detection for multiple types of objects in each area contained in an image, even if the image contains multiple types of objects to be detected.

[0044] (Configuration of information processing device 1A) The configuration of information processing device 1A will be described with reference to Fig. 5. Fig. 5 is a block diagram showing the configuration of information processing device 1A. Information processing device 1A includes a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.

[0045] (Control unit 10) The control unit 10 controls each unit of the information processing device 1A in an integrated manner. The control unit 10 includes an acquisition unit 11, an estimation unit 12, and a generation unit 13 that are included in the information processing device 1, and a learning unit 22 that is included in the information processing device 2. The acquisition unit 11 included in the control unit 10 also functions as the acquisition unit 21 that is included in the information processing device 2.

[0046] (Acquisition Unit 11) The acquisition unit 11 acquires one or more images.

[0047] (Estimation Unit 12) The estimation unit 12 estimates the difficulty of object detection for a plurality of types of objects in each region included in the image acquired by the acquisition unit 11.

[0048] For example, if the types of objects to be detected are a person and a suitcase, the estimation unit 12 may estimate the difficulty of object detection for each of the person and the suitcase in each area included in the image acquired by the acquisition unit 11.

[0049] For example, the estimation unit 12 may generate a difficulty level image for each type of object, the difficulty level indicating the estimated difficulty level for each region. For example, if the types of objects to be detected are a person and a suitcase, the estimation unit 12 may generate a difficulty level image for each of the person and the suitcase, the difficulty level indicating the estimated difficulty level for each region.

[0050] Furthermore, the estimation unit 12 may estimate the difficulty of object detection for multiple types of objects using, for example, one or more difficulty estimation models. For example, if the types of objects to be detected are a person and a suitcase, the estimation unit 12 may estimate the difficulty of object detection for each of the person and the suitcase using one or more difficulty estimation models. Note that the difficulty estimation model may be, for example, a machine learning model using a CNN (Convolution Neural Network).

[0051] Here, the difficulty estimation model may be, for example, a model trained by machine learning using training images including a different type of object from the type of object that the difficulty estimation model is intended to detect. For example, if the difficulty estimation model is intended to detect a person, the difficulty estimation model may be a model trained by machine learning using training images including a suitcase.

[0052] (Example of Training Image) The training image may be, for example, an image in which different types of objects are arranged in positions in a background image according to the result of a segmentation process applied to the background image. The segmentation process may be, for example, a process of determining the ground or the sky in the background image using a known algorithm. For example, if the types of objects to be detected are a person and a suitcase, since a person or a suitcase is usually located on the ground, the training image may be an image in which a person or a suitcase is arranged on the ground in the background image.

[0053] Furthermore, the learning images may be images in which objects of a type different from the type of object that the difficulty estimation model is intended to detect are arranged, depending on whether or not an image of the interior of a room is included in the background image. For example, whether or not an image of the interior of a room is included in the background image may be determined using a known algorithm.

[0054] Furthermore, for example, the object may be a three-dimensional model, and the learning image may be an image in which a type of object different from the type of object targeted for detection by the difficulty estimation model is placed in a background image based on the result of estimating the depth of a subject included in the background image. Here, the depth of the subject included in the background image may be estimated using, for example, a known algorithm. In this case, for example, the size or angle of the object to be placed in the background image may be set based on the result of estimating the depth of the subject included in the background image.

[0055] Furthermore, the learning image may be, for example, an image in which, in a background image, an object of a type different from the type of object to be detected by the difficulty estimation model is arranged according to information regarding its positional relationship with other objects. As a specific example, since it is unlikely in real life that a person would be placed directly below a vehicle, except for exceptional cases such as maintenance, the placement of such an object may be avoided.

[0056] Furthermore, the training image may be, for example, an image in which information regarding the boundary between a background image and an object placed on the background image is referenced, and the boundary is processed using a predetermined method. Here, for example, the training image may be an image in which the boundary between the background image and the object image placed on the background image is processed using a predetermined method to approximate the appearance of the background image and the object image placed on the background image, so that the boundary between the images appears natural to the user. The predetermined method may be, for example, a known technique. Specific examples of the predetermined method include: Feathering; Poisson blending; and Deep image blending. Here, for example, Feathering may be a method of blurring the boundary between object images. For example, Poisson blending may be a method of combining images while maintaining an image gradient. For example, Deep image blending may be a method of using deep learning to perform processing so that the boundary between the images appears natural to the user.

[0057] (Outline of Difficulty Estimation Model Used by Estimation Unit 12) The difficulty estimation model used by the estimation unit 12 may be, for example, a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label. The layers included in the difficulty estimation model will be described later in the configuration example of FIG. 9. Note that the object detection model may be, for example, a machine learning model that uses a CNN (Convolution Neural Network).

[0058] (Generation Unit 13) The generation unit 13 refers to the estimation result by the estimation unit 12 and generates a display image that indicates the detection difficulty of multiple types of objects for each region.

[0059] The generation unit 13 may generate a display image by, for example, integrating difficulty level images for each type of object.

[0060] (Acquisition Unit 21) As described above, the acquisition unit 11 also functions as the acquisition unit 21. The acquisition unit 21 acquires learning images. For example, the acquisition unit 21 may generate learning images by applying a segmentation process to a background image and arranging different types of objects in the background image at positions according to the results of the segmentation process. The segmentation process may be, for example, a process of determining the ground or the sky in the background image using a known algorithm. For example, if the types of objects to be detected are a person and a suitcase, since a person or a suitcase is usually located on the ground, the acquisition unit 21 may generate learning images by arranging the person or the suitcase on the ground in the background image.

[0061] (Learning Unit 22) The learning unit 22 performs machine learning on a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in the target image, by referring to the training images. For example, if the types of objects to be detected are people and suitcases, the learning unit 22 may perform machine learning on a difficulty estimation model that estimates the difficulty of object detection for each of the people and the suitcases, by referring to the training images. Note that the difficulty estimation model may be, for example, a machine learning model using a CNN (Convolution Neural Network).

[0062] For example, the learning unit 22 may train the difficulty estimation model by machine learning using training images that include a type of object different from the type of object that the difficulty estimation model is intended to detect. For example, if the difficulty estimation model is intended to detect a person, the learning unit 22 may train the difficulty estimation model by machine learning using training images that include a suitcase.

[0063] (Outline of Difficulty Estimation Model Trained by Machine Learning by Learning Unit 22) The difficulty estimation model trained by machine learning by the learning unit 22 may be, for example, a model including at least one layer included in an object detection model trained by machine learning in advance, and one or more layers including one or more parameters updated in a learning process in which the accuracy information output by the object detection model is used as a correct answer label. The layers included in the difficulty estimation model will be described later in the configuration example of FIG. 9. Note that the object detection model may be, for example, a machine learning model using a CNN (Convolution Neural Network).

[0064] (Storage Unit 20) The storage unit 20 stores various types of data referenced by the control unit 10 and various types of data generated by the control unit 10. As an example, the storage unit 20 stores the following: Input image data AD Object information OB Detection difficulty information DD Display image data OD Learning image data TD Accuracy information PR Difficulty estimation model DM

[0065] The input image data AD may be, for example, one or more images acquired by the acquisition unit 11. The input image data AD may also be, for example, data including one or more background images that serve as backgrounds for the object. Here, the background images may be, for example, images included in a known background image data set.

[0066] The object information OB may be, for example, information about the type of object to be detected by the information processing device 1 A. Specific examples of the object information OB include people, suitcases, vehicles such as bicycles, doctors, patients, and IV stands.

[0067] The detection difficulty information DD may be, for example, information indicating the difficulty of object detection estimated for multiple types of objects in each region included in the input image data AD.

[0068] The display image data OD may be, for example, data including a display image indicating the difficulty of object detection for each region, generated by the generation unit 13. For example, the display image data OD may be an object detection difficulty image DDM, which will be described later in the configuration example of FIG.

[0069] The training image data TD may be, for example, a training image acquired by the acquisition unit 21. Furthermore, the training image data TD may be, for example, a training image that is provided as a reference for the learning unit 22 when the learning unit 22 performs machine learning on the difficulty estimation model DM.

[0070] The accuracy information PR may be, for example, the degree to which the object to be detected is detected by the object detection model in each region included in the input image data AD. Here, the accuracy information PR may be, for example, the detection rate at which the object is detected in each region included in the input image data AD.

[0071] The difficulty estimation model DM may be, for example, a machine learning model used when estimating the difficulty of object detection for multiple types of objects in each region included in an image. Furthermore, for example, the learning unit 22 may perform machine learning on the difficulty estimation model DM by referring to the training image data TD. Note that the difficulty estimation model DM may be, for example, a machine learning model using a CNN (Convolution Neural Network).

[0072] (Communication Unit 30) The communication unit 30 communicates with devices external to the information processing device 1A. As an example, the communication unit 30 communicates with external devices connected to the information processing device 1A via a network N (not shown). The communication unit 30 transmits data supplied from the control unit 10 to the outside, and supplies data received from the outside to the control unit 10. Note that the specific configuration of the network N does not limit this exemplary embodiment, and as an example, a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public line network, a mobile data communication network, or a combination of these networks can be used.

[0073] (Input / Output Unit 40) The input / output unit 40 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. Alternatively, the input / output unit 40 may be configured to be connected to input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. In this configuration, the input / output unit 40 accepts various types of information input to the information processing device 1A from the connected input devices. Furthermore, the input / output unit 40 outputs various types of information to connected output devices under the control of the control unit 10. An example of the input / output unit 40 is an interface such as a USB (Universal Serial Bus).

[0074] (Flow of information processing method S1A executed by information processing device 1A) The flow of information processing method S1A executed by information processing device 1A will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the flow of information processing method S1A. Note that step S12 in information processing method S1 corresponds to steps S121 to S122 in information processing method S1A.

[0075] (Step S11) In step S11, the acquisition unit 11 acquires one or more images.

[0076] (Step S121) In step S121, the estimation unit 12 may estimate the difficulty of object detection for a plurality of types of objects using, for example, one or a plurality of difficulty estimation models.

[0077] (Step S122) In step S122, the estimation unit 12 may generate, for example, a difficulty level image indicating the estimated difficulty level for each region, for each type of object.

[0078] (Step S13) In step S13, the generating unit 13 refers to the estimation result by the estimating unit 12 and generates a display image indicating the detection difficulty of a plurality of types of objects for each region.

[0079] (Flow of information processing method S2A executed by information processing device 1A) The flow of information processing method S2A executed by information processing device 1A will be described with reference to Fig. 7. Fig. 7 is a flow chart showing the flow of information processing method S2A.

[0080] (Step S21) In step S21, the acquisition unit 21 acquires learning images.

[0081] (Step S22) In step S22, the learning unit 22 may perform machine learning of a difficulty estimation model using, for example, learning images that include a type of object different from the detection target.

[0082] (Creation of Learning Data and Learning Process) Fig. 8 is a diagram showing an example of the configuration of creation of learning data and learning process. As will be described with reference to Fig. 8 , the difficulty estimation model according to this exemplary embodiment is a model that has been machine-learned using learning images that include objects of a type different from the type of object that the difficulty estimation model is to detect. In the example of Fig. 8 , the objects to be detected are a person and a suitcase.

[0083] In the example of Figure 8, images AD1 and AD2 are images in which people and suitcases are randomly arranged in a background image. Also, in the example of Figure 8, the pre-trained model LM is an object detection model that has been machine-learned in advance to detect people and suitcases from images. Here, in the example of Figure 8, the estimation unit 12 detects objects included in images AD1 and AD2 using the pre-trained model LM and records the detection rate of each object in each image. In the example of Figure 8, the detection rate of image 1 may correspond to image AD1, and the detection rate of image 2 may correspond to image AD2, respectively.

[0084] In the example of Fig. 8, the person detection difficulty estimation model DMH and the suitcase detection difficulty estimation model DMS are obtained by extracting at least one layer included in the intermediate layer of the pre-trained model LM. Here, the person detection difficulty estimation model DMH and the suitcase detection difficulty estimation model DMS may each be, for example, an object detection difficulty estimation model DM. The difficulty estimation model DM will be described later in the configuration example of Fig. 9.

[0085] In the learning example of FIG. 8 , the acquisition unit 21 generates training images by placing objects other than the detection target in a background image. Here, the estimation unit 12 may, for example, train each difficulty estimation model DM by machine learning to regress, i.e., estimate, the detection rate of the detection target object by referring to the training images. For example, the estimation unit 12 may train the person detection difficulty estimation model DMH by machine learning to regress the detection rate of the person in the training image TDH by referring to a training image TDH, which is an image in which a suitcase, an object other than a person, is placed in the background image. Furthermore, for example, the estimation unit 12 may train the suitcase detection difficulty estimation model DMS by machine learning to regress the detection rate of the suitcase in the training image TDS by referring to a training image TDS, which is an image in which a person, an object other than a suitcase, is placed in the background image. In other words, the difficulty estimation models according to this exemplary embodiment, the person detection difficulty estimation model DMH and the suitcase detection difficulty estimation model DMS, may be models trained by machine learning using training images including objects of a type different from the type of object to be detected by the difficulty estimation model.

[0086] The training images TDH and TDS may be, for example, images in which different types of objects are arranged in positions in a background image according to the results of a segmentation process applied to the background image. The segmentation process may be, for example, a process of determining the ground or sky in the background image using a known algorithm. For example, if the types of objects to be detected are a person and a suitcase, the person or suitcase is usually located on the ground, so the training images may be images in which the person or suitcase is arranged on the ground in the background image.

[0087] In the example of FIG. 8, the learning unit 22 may perform the above processes instead of the estimation unit 12.

[0088] 9 is a diagram showing an example of the configuration of processing related to learning and inference. As will be described with reference to Fig. 9, the difficulty estimation model according to this exemplary embodiment is a model that includes at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers that include one or more parameters that are updated in a learning process that uses accuracy information output by the object detection model as a correct answer label.

[0089] (Calculation of Object Detection Rate in Background Image) In the learning example of FIG. 9 , the estimation unit 12 extracts features from input images to the estimation unit 12 using the pre-trained model LD. The input images to the estimation unit 12 may be, for example, multiple images in which objects are arranged in a background image. Furthermore, the input images to the estimation unit 12 may be, for example, images in which objects are arranged in each region included in the background image. Here, for example, there may be multiple types of background images in which objects are arranged. Furthermore, for example, in multiple input images to the estimation unit 12 that have the same background image, the color or shape of the object arranged in the background image may differ between the multiple input images. Furthermore, the backbone included in the pre-trained model LD may be, for example, a group of layers used by the estimation unit 12 to extract features from images using the pre-trained model LD.

[0090] The images ADH1 and ADH2 in FIG. 9 are examples of input images to the estimation unit 12, and although they have the same background image, the areas in which the rectangular objects are arranged are different.

[0091] 9 , the estimation unit 12 uses the backbone and detection head included in the pre-trained model LD to output a detection result DR of an object in an input image to the estimation unit 12. Here, for example, the detection result DR may be a detection result of an object in an input image to the estimation unit 12 that is different from the images ADH1 and ADH2. In the example of the detection result DR, the detected object is surrounded by a rectangle.

[0092] Furthermore, for example, when multiple input images to the estimation unit 12 have the same background image and the colors or shapes of objects located in the same region of the background image differ between the multiple input images, whether or not the object can be detected using the pre-trained model LD may depend on the color or shape of the object. In this case, the estimation unit 12 may calculate, for example, the object detection rate using the pre-trained model LD for each region included in the same background image in the multiple input images.

[0093] (Learning of a Model for Estimating an Object Detection Rate from a Background Image) In the example of Figure 9, the estimation unit 12 extracts at least one layer from the intermediate layer of the backbone in the pre-trained model LD. Here, the difficulty estimation model DM according to this exemplary embodiment may include at least one layer included in the pre-trained model LD, which is an object detection model that has been machine-learned in advance. Note that, for example, the object detection process using the pre-trained model LD may not be completed in the extracted layer.

[0094] 9, the estimation unit 12 performs processing using the layer in the CNN after layer extraction. In the learning example of FIG. 9, the estimation unit 12 performs flattening and full connection processing on DDL, which is the layer at the time when processing using the layer in the CNN is completed and is the layer immediately before flattening. In the example of FIG. 9, the layer at the time of extraction is an array of size 1280 x 20 x 20, while the array ED at the time when processing up to full connection is completed has a size of 1 x 1.

[0095] Here, in the example of Figure 9, the difficulty estimation model DM may be a model that includes at least one layer included in the pre-trained model LD and one or more layers that include one or more parameters that are updated in a learning process in which the detection rate output by the pre-trained model LD is used as the correct label.

[0096] In the learning example of FIG. 9 , when the difficulty estimation model DM includes layers up to just before flattening, the estimation unit 12 may train parameters included in each layer so as to estimate the detection rate, i.e., accuracy information, of an object for each region included in the background image using the values ​​included in the array ED. Here, the difficulty estimation model DM according to this exemplary embodiment may include one or more layers including one or more parameters updated in a learning process in which the detection rate output by the pre-trained model LD is used as the correct label. For example, the estimation unit 12 may calculate the loss between the detection rate of the object and the estimated value included in the array ED for multiple input images including the same background image, update the parameters included in each layer so as to minimize the loss, and train the difficulty estimation model DM. In this case, the estimated value included in the array ED may be, for example, a value between 0 and 1, or between 0% and 100%.

[0097] In the example of learning in FIG. 9, the learning unit 22 may perform the above processes instead of the estimation unit 12.

[0098] (Configuration for Inferring Detection Rate from Background Image) In the example of inference in Fig. 9 , the estimation unit 12 uses a backbone to extract feature amounts from a background image BD, which is an input image to the estimation unit 12. For example, the estimation unit 12 may extract feature amounts from the background image BD using a backbone included in the pre-trained model LD in the example of training in Fig. 9 . Here, the backbone may include, for example, the difficulty estimation model DM that has undergone the training process in the example of training in Fig. 9 .

[0099] In the inference example of FIG. 9 , similar to the learning example of FIG. 9 , the estimation unit 12 extracts at least one layer included in the intermediate layer in the backbone and performs layer processing in the CNN for the extracted layer. Here, in the inference example of FIG. 9 , the layer DDE immediately before flattening is output as is. The layers DDL and DDE immediately before flattening may be, for example, images. Furthermore, for example, the layer DDE may include information regarding the result of estimating the object detection rate in the background image BD.

[0100] 9 , the generation unit 13 may refer to the layer image DDM to generate a difficulty level image DDM indicating the difficulty of object detection for each region included in the background image BD. Here, the difficulty level image DDM may indicate, for example, an estimated value of the object detection rate for each region included in the background image BD. Furthermore, the generation unit 13 may, for example, color-code, such as by using a shade of gray, each region included in the background image BD according to the difficulty of object detection in the difficulty level image DDM.

[0101] (Integration Process by Generator 13) For example, the integration process by the generator 13 may include a process of superimposing difficulty level images for each type of object.

[0102] FIG. 10 illustrates an example of overlaying object detection difficulty images. In the example of FIG. 10 , the estimation unit 12 regresses the detection rate of objects in each region included in the background image BD using the person detection difficulty estimation model DMH and the suitcase detection difficulty estimation model DMS. Also, in the example of FIG. 10 , the generation unit 13 references the processing results of the estimation unit 12 and outputs a person detection difficulty image DDMH and a suitcase detection difficulty image DDMS, respectively. Here, the generation unit 13 may, for example, overlay the person detection difficulty image DDMH and the suitcase detection difficulty image DDMS to generate an integrated detection difficulty image DDM. Here, specific examples of the process of overlaying difficulty images for each object type in the integration process by the generation unit 13 include a process of taking the average, product, or maximum value or logical product of the detection difficulty for each region included in each difficulty image, such as by comparatively brightening the image. Also, for example, the generation unit 13 may sequentially estimate the detection difficulty for each object type for images obtained by frame-decomposing video data, and overlay them in the time direction.

[0103] Furthermore, for example, the generation unit 13 may weight each type of object and then overlay difficulty level images for each type of object.

[0104] Furthermore, for example, the integration process by the generation unit 13 may include a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or on the integrated image.

[0105] 11 is a diagram showing an example in which an icon representing an object is superimposed on an object detection difficulty image. For example, as shown in FIG. 11 , the generation unit 13 may superimpose a bicycle icon in an area of ​​the display image data OD where the detection difficulty of the bicycle is high, or may superimpose a person icon in an area where the detection difficulty of the person is high. Furthermore, the generation unit 13 may, for example, color-code the icons representing the objects in the display image data OD, using different shades of color, depending on the detection difficulty of each type of object.

[0106] Furthermore, for example, the integration process by the generation unit 13 may include a process of highlighting at least one of areas with a relatively high level of difficulty and areas with a relatively low level of difficulty in the difficulty image for each type of object, or in the image after integration.

[0107] 12 is a diagram showing an example of a presentation of a difficulty level image that highlights areas where the difficulty of object detection is high. As shown in the example of Fig. 12, the generation unit 13 may display areas in the display image data OD where the difficulty of object detection exceeds a predetermined threshold as rectangles.

[0108] 13 is a diagram showing an example of a presentation in which an area to be focused on is emphasized in the object detection difficulty image. As shown in the example of FIG. 13, the generation unit 13 may change the color or brightness between the area to be focused on and other areas in the display image data OD, or may display the areas other than the area to be focused on in a blurred manner.

[0109] (Example of Use of Object Detection Difficulty Level Image) For example, the position or angle of the imaging device may be set with reference to the object detection difficulty level image generated by the generation unit 13. As a specific example, when a surveillance camera is installed in a medical facility or the like for the purpose of issuing a warning or an instruction to leave to an unauthorized intruder, the position or angle of the surveillance camera may be set with reference to the object detection difficulty level image generated by the generation unit 13 so that the intruder can be easily detected.

[0110] (Advantages of Information Processing Device 1A) As described above, in information processing device 1A, the estimation means generates, for each type of object, a difficulty level image indicating the estimated difficulty level for each region, and the generation means generates a display image by integrating the difficulty levels for each type of object. Therefore, in addition to the advantages of information processing device 1, information processing device 1A has the advantage of being able to display a difficulty level image that reflects the difficulty levels of detection for multiple types of objects.

[0111] Furthermore, in the information processing device 1A, the integration process by the generating means includes a process of overlaying difficulty images for each type of object. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A has the effect of being able to highlight and display areas on an image where multiple types of objects are difficult to detect.

[0112] Furthermore, in the information processing device 1A, the integration process by the generating means includes a process of superimposing at least one of specific text and images on the difficulty level image for each object type or on the integrated image. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A has the effect of being able to present the difficulty level of object detection in each area on the image for each object type.

[0113] Furthermore, in the information processing device 1A, the integration process by the generating means includes a process of highlighting at least one of areas with a relatively high level of difficulty and areas with a relatively low level of difficulty in the difficulty level image for each type of object or in the integrated image. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A has the effect of being able to quickly identify areas on the image that are highlighted according to the difficulty level of object detection.

[0114] Furthermore, in the information processing device 1A, the estimation means is configured to estimate the difficulty of object detection for multiple types of objects using one or multiple difficulty estimation models. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can also achieve the effect of being able to estimate the difficulty of object detection for each of multiple types of objects.

[0115] Furthermore, the information processing device 1A employs a configuration in which the difficulty estimation model is a model that has been machine-learned using training images that include a type of object that is different from the type of object that the difficulty estimation model is intended to detect. Therefore, in addition to the effects achieved by the information processing device 1, the information processing device 1A has the effect of being able to machine-learn a difficulty estimation model that takes into account interactions between objects.

[0116] Furthermore, the information processing device 1A employs a configuration in which the learning images are images in which different types of objects are arranged in positions in a background image according to the results of segmentation processing applied to the background image. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can also achieve the effect of being able to arrange objects in positions that are actually possible in the background image.

[0117] Furthermore, the information processing device 1A employs a configuration in which the difficulty estimation model includes at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers that include one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct answer label. Thus, in addition to the effects of the information processing device 1, the information processing device 1A achieves the effect of being able to generate a difficulty estimation model based on components included in an object detection model that has been machine-learned in advance.

[0118] Furthermore, in the information processing device 1A, the learning means is configured to perform machine learning of the difficulty estimation model using training images that include a type of object different from the type of object that the difficulty estimation model is intended to detect. Therefore, in addition to the effects achieved by the information processing device 2, the information processing device 1A has the effect of being able to machine learn the difficulty estimation model by taking into account interactions between objects.

[0119] Furthermore, in the information processing device 1A, the acquisition means applies a segmentation process to the background image and places different types of objects in the background image at positions according to the results of the segmentation process, thereby generating a learning image. Therefore, in addition to the effects of the information processing device 2, the information processing device 1A has the effect of being able to place objects in positions that are actually possible in the background image.

[0120] Furthermore, the information processing device 1A employs a configuration in which the difficulty estimation model includes at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers that include one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct answer label. Thus, in addition to the effects of the information processing device 2, the information processing device 1A achieves the effect of being able to generate a difficulty estimation model based on components included in an object detection model that has been machine-learned in advance.

[0121] [Software Implementation Example] Some or all of the functions of the information processing devices 1, 2, 1A (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0122] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 14. Figure 14 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0123] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0124] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0125] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0126] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0127] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0128] (Appendix A1) An information processing device comprising: an acquisition means for acquiring one or more images; an estimation means for estimating the difficulty of object detection for multiple types of objects in each region included in the images acquired by the acquisition means; and a generation means for generating a display image indicating the difficulty of detection of the multiple types of objects for each region by referring to the estimation results by the estimation means.

[0129] (Appendix A2) The information processing device described in Appendix A1, wherein the estimation means generates a difficulty level image for each type of object, the difficulty level image indicating the estimated difficulty level for each area, and the generation means generates the display image by integrating the difficulty level images for each type of object.

[0130] (Supplementary Note A3) The information processing device according to Supplementary Note A2, wherein the integration process by the generating means includes a process of superimposing difficulty level images for each type of object.

[0131] (Supplementary Note A4) The information processing device according to Supplementary Note A2 or A3, wherein the integration process by the generating means includes a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or on the integrated image.

[0132] (Appendix A5) The information processing device described in any one of Appendices A2 to A4, wherein the integration process by the generation means includes a process of highlighting at least one of the areas where the difficulty level is relatively high and the areas where the difficulty level is relatively low in the difficulty level image for each type of object or in the image after integration.

[0133] (Supplementary Note A6) The information processing device according to any one of Supplementary Notes A1 to A5, wherein the estimation means estimates the difficulty of object detection for the plurality of types of objects using one or a plurality of difficulty estimation models.

[0134] (Supplementary Note A7) The information processing device according to Supplementary Note A6, wherein the difficulty estimation model is a model that is machine-learned using learning images that include a type of object different from a type of object that the difficulty estimation model is intended to detect.

[0135] (Supplementary Note A8) The information processing device according to Supplementary Note A7, wherein the learning image is an image in which the different types of objects are arranged in positions in a background image according to a result of a segmentation process applied to the background image.

[0136] (Appendix A9) The information processing device described in any one of Appendices A6 to A8, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0137] (Appendix A10) An information processing device comprising: an acquisition means for acquiring a learning image; and a learning means for machine learning a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image, by referring to the learning image.

[0138] (Supplementary Note A11) The information processing device according to Supplementary Note A10, wherein the learning means performs machine learning on the difficulty estimation model using learning images including a type of object different from a type of object that is a detection target of the difficulty estimation model.

[0139] (Supplementary Note A12) The information processing device according to Supplementary Note A11, wherein the acquisition means applies a segmentation process to a background image, and generates the learning image by arranging the different types of objects in the background image at positions according to a result of the segmentation process.

[0140] (Appendix A13) The information processing device described in any one of Appendices A10 to A12, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0141] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0142] (Appendix B1) An information processing method including: an acquisition process in which at least one processor acquires one or more images; an estimation process in which the at least one processor estimates the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition process; and a generation process in which the at least one processor generates a display image indicating the difficulty of detection of the multiple types of objects for each region by referring to the estimation results of the estimation process.

[0143] (Appendix B2) An information processing method described in Appendix B1, wherein in the estimation process, the at least one processor generates a difficulty image for each type of object, the difficulty image indicating the estimated difficulty for each region, and in the generation process, the at least one processor generates the display image by integrating the difficulty images for each type of object.

[0144] (Supplementary Note B3) The information processing method according to Supplementary Note B2, wherein the integration process in the generation process includes a process of superimposing difficulty level images for each type of object.

[0145] (Appendix B4) The information processing method according to Appendix B2 or B3, wherein the integration process by the generation process includes a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or on the integrated image.

[0146] (Appendix B5) The information processing method described in any one of Appendices B2 to B4, wherein the integration process by the generation process includes a process of highlighting at least one of the areas with a relatively high level of difficulty and the areas with a relatively low level of difficulty in the difficulty image for each type of object or in the image after integration.

[0147] (Supplementary Note B6) The information processing method according to any one of Supplementary Notes B1 to B5, wherein in the estimation process, the at least one processor estimates the difficulty of object detection for the plurality of types of objects using one or more difficulty estimation models.

[0148] (Supplementary Note B7) The information processing method according to Supplementary Note B6, wherein the difficulty estimation model is a model that is machine-learned using training images that include a type of object different from the type of object that the difficulty estimation model is intended to detect.

[0149] (Supplementary Note B8) The information processing method according to Supplementary Note B7, wherein the learning image is an image in which the different types of objects are arranged in positions in a background image according to a result of a segmentation process applied to the background image.

[0150] (Appendix B9) The information processing method described in any one of Appendices B6 to B8, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0151] (Appendix B10) An information processing method including: an acquisition process in which the at least one processor acquires a learning image; and a learning process in which the at least one processor performs machine learning on a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image, by referring to the learning image.

[0152] (Supplementary Note B11) The information processing method according to Supplementary Note B10, wherein in the learning process, the at least one processor performs machine learning on the difficulty estimation model using learning images including a type of object different from a type of object that the difficulty estimation model is to detect.

[0153] (Supplementary Note B12) The information processing method according to Supplementary Note B11, wherein in the acquisition process, the at least one processor applies a segmentation process to a background image, and generates the learning image by arranging the different types of objects in the background image at positions according to a result of the segmentation process.

[0154] (Appendix B13) The information processing method described in any one of Appendices B10 to B12, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0155] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0156] (Appendix C1) A program that causes a computer to function as an information processing device, the program causing the computer to function as: an acquisition means that acquires one or more images; an estimation means that estimates the difficulty of object detection for multiple types of objects in each area included in the image acquired by the acquisition means; and a generation means that generates a display image that indicates the difficulty of detection for each area of ​​the multiple types of objects by referring to the estimation results by the estimation means.

[0157] (Appendix C2) The program described in Appendix C1, wherein the estimation means generates a difficulty level image for each type of object, the difficulty level image indicating the estimated difficulty level for each area, and the generation means generates the display image by integrating the difficulty level images for each type of object.

[0158] (Supplementary Note C3) The program according to Supplementary Note C2, wherein the integration process by the generating means includes a process of superimposing difficulty level images for each type of object.

[0159] (Appendix C4) The program according to appendix C2 or C3, wherein the integration process by the generating means includes a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or on the integrated image.

[0160] (Appendix C5) The program described in any one of Appendices C2 to C4, wherein the integration process by the generating means includes a process of highlighting at least one of the areas where the difficulty level is relatively high and the areas where the difficulty level is relatively low in the difficulty level image for each type of object or in the image after integration.

[0161] (Supplementary Note C6) The program according to any one of Supplementary Notes C1 to C5, wherein the estimation means estimates the difficulty of object detection for the plurality of types of objects using one or a plurality of difficulty estimation models.

[0162] (Supplementary Note C7) The program according to Supplementary Note C6, wherein the difficulty estimation model is a model that is machine-learned using training images that include a type of object different from a type of object that the difficulty estimation model is intended to detect.

[0163] (Supplementary Note C8) The program according to Supplementary Note C7, wherein the learning image is an image in which the different types of objects are arranged in positions in a background image according to a result of a segmentation process applied to the background image.

[0164] (Appendix C9) The program described in any one of Appendices C6 to C8, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process that uses accuracy information output by the object detection model as a correct label.

[0165] (Appendix C10) A program that causes the computer to function as: an acquisition means that acquires learning images; and a learning process that performs machine learning to develop a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each area included in a target image by referring to the learning images.

[0166] (Supplementary Note C11) The program according to Supplementary Note C10, wherein the learning means trains the difficulty estimation model by machine learning using training images including a type of object different from a type of object that the difficulty estimation model is to detect.

[0167] (Appendix C12) The program according to Appendix C11, wherein the acquisition means applies a segmentation process to a background image, and generates the learning image by arranging the different types of objects in the background image at positions according to a result of the segmentation process.

[0168] (Appendix C13) The program described in any one of Appendices C10 to C12, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process that uses accuracy information output by the object detection model as a correct label.

[0169] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0170] (Appendix D1) An information processing device comprising at least one processor, the at least one processor executing: an acquisition process for acquiring one or more images; an estimation process for estimating the difficulty of object detection for multiple types of objects in each region included in the image acquired by the acquisition process; and a generation process for generating a display image indicating the difficulty of detection of the multiple types of objects for each region by referring to the estimation results of the estimation process.

[0171] (Appendix D2) The information processing device described in Appendix D1, wherein in the estimation process, the at least one processor generates a difficulty image for each type of object, the difficulty image indicating the estimated difficulty for each region, and in the generation process, the at least one processor generates the display image by integrating the difficulty images for each type of object.

[0172] (Supplementary Note D3) The information processing device according to Supplementary Note D2, wherein the integration process in the generation process includes a process of superimposing difficulty level images for each type of object.

[0173] (Supplementary Note D4) The information processing device according to Supplementary Note D2 or D3, wherein the integration process by the generation process includes a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or on the integrated image.

[0174] (Appendix D5) The information processing device described in any one of Appendices D2 to D4, wherein the integration process by the generation process includes a process of highlighting at least one of the areas where the difficulty level is relatively high and the areas where the difficulty level is relatively low in the difficulty level image for each type of object or in the image after integration.

[0175] (Supplementary Note D6) The information processing device according to any one of Supplementary Notes D1 to D5, wherein in the estimation process, the at least one processor estimates the difficulty of object detection for the plurality of types of objects using one or a plurality of difficulty estimation models.

[0176] (Supplementary Note D7) The information processing device according to Supplementary Note D6, wherein the difficulty estimation model is a model that is machine-learned using learning images that include a type of object different from a type of object that the difficulty estimation model is intended to detect.

[0177] (Supplementary Note D8) The information processing device according to Supplementary Note D7, wherein the learning image is an image in which the different types of objects are arranged in positions in a background image according to a result of a segmentation process applied to the background image.

[0178] (Appendix D9) The information processing device described in any one of Appendices D6 to D8, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0179] (Appendix D10) The information processing device, wherein the at least one processor executes an acquisition process of acquiring a learning image, and a learning process of machine learning a difficulty estimation model that estimates the difficulty of object detection for multiple types of objects in each region included in a target image, by referring to the learning image.

[0180] (Supplementary Note D11) The information processing device described in Supplementary Note D10, wherein in the learning process, the at least one processor performs machine learning on the difficulty estimation model using learning images including a type of object different from a type of object that the difficulty estimation model is to detect.

[0181] (Supplementary Note D12) The information processing device according to Supplementary Note D11, wherein in the acquisition process, the at least one processor applies a segmentation process to a background image, and generates the learning image by arranging the different types of objects in the background image at positions according to a result of the segmentation process.

[0182] (Appendix D13) The information processing device described in any one of Appendices D10 to D12, wherein the difficulty estimation model is a model including at least one layer included in an object detection model that has been machine-learned in advance, and one or more layers including one or more parameters that are updated in a learning process in which accuracy information output by the object detection model is used as a correct label.

[0183] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0184] (Appendix E1) A non-transitory recording medium having recorded thereon a program that causes a computer to function as an information processing device, the program causing the computer to execute: an acquisition process that acquires one or more images; an estimation process that estimates the difficulty of object detection for multiple types of objects in each area included in the image acquired by the acquisition process; and a generation process that generates a display image that indicates the difficulty of detection for each area of ​​the multiple types of objects by referring to the estimation results of the estimation process.

[0185] 1, 1A, 1A, 2 Information processing device 10 Control unit 11, 21 Acquisition unit 12 Estimation unit 13 Generation unit 20 Storage unit 22 Learning unit 30 Communication unit 40 Input / output unit C1 Processor C2 Memory

Claims

1. An information processing apparatus comprising: an acquisition unit that acquires one or more images; an estimation unit that estimates the difficulty level of object detection for a plurality of types of objects in each region included in the images acquired by the acquisition unit; and a generation unit that generates a display image indicating the detection difficulty level of the plurality of types of objects for each region with reference to the estimation result by the estimation unit.

2. The information processing apparatus according to claim 1, wherein the estimation unit generates a difficulty level image indicating the estimated difficulty level for each region for each type of object, and the generation unit generates the display image by integrating the difficulty level images for each type of object.

3. The information processing apparatus according to claim 2, wherein the integration process by the generation unit includes a process of overlapping the difficulty level images for each type of object.

4. The information processing apparatus according to claim 2 or 3, wherein the integration process by the generation unit includes a process of superimposing at least one of specific text and an image on the difficulty level image for each type of object or the integrated image.

5. The information processing apparatus according to any one of claims 2 to 4, wherein the integration process by the generation unit includes a process of highlighting at least one of a region where the difficulty level is relatively high and a region where the difficulty level is relatively low in the difficulty level image for each type of object or the integrated image.

6. The information processing apparatus according to any one of claims 1 to 5, wherein the estimation unit estimates the difficulty level of object detection for the plurality of types of objects using one or more difficulty level estimation models.

7. The information processing apparatus according to claim 6, wherein the difficulty level estimation model is a model machine-learned using learning images including types of objects different from the types of objects to be detected by the difficulty level estimation model.

8. The information processing apparatus according to claim 7, wherein the learning image is an image in which the different types of objects are arranged at positions according to the result of a segmentation process applied to a background image in the background image.

9. The information processing apparatus according to any one of claims 6 to 8, wherein the difficulty estimation model includes at least one layer included in a pre-trained object detection model, and one or more layers including one or more parameters updated in a learning process using the probability information output by the object detection model as a correct label.

10. An information processing apparatus comprising: an acquisition unit that acquires learning images; and a learning unit that performs machine learning on a difficulty estimation model that estimates the difficulty of object detection for a plurality of types of objects in each region included in a target image, with reference to the learning images.

11. The information processing apparatus according to claim 10, wherein the learning unit performs machine learning on the difficulty estimation model using learning images including types of objects different from the types of objects to be detected by the difficulty estimation model.

12. The information processing apparatus according to claim 11, wherein the acquisition unit applies a segmentation process to a background image, and generates the learning images by arranging the different types of objects at positions corresponding to the results of the segmentation process in the background image.

13. The information processing apparatus according to any one of claims 10 to 12, wherein the difficulty estimation model includes at least one layer included in a pre-trained object detection model, and one or more layers including one or more parameters updated in a learning process using the probability information output by the object detection model as a correct label.

14. An information processing method including: an acquisition process of acquiring one or more images; an estimation process of estimating the difficulty of object detection for a plurality of types of objects in each region included in the image acquired in the acquisition process; and a generation process of generating a display image showing the detection difficulty of the plurality of types of objects for each region, with reference to the estimation result of the estimation process.

15. An information processing method including: an acquisition process of acquiring learning images; and a learning process of performing machine learning on a difficulty estimation model that estimates the difficulty of object detection for a plurality of types of objects in each region included in a target image, with reference to the learning images.

16. A program for causing a computer to execute an acquisition process of acquiring one or more images, an estimation process of estimating the difficulty level of object detection for a plurality of types of objects in each region included in the images acquired by the acquisition process, and a generation process of generating a display image indicating the detection difficulty level of the plurality of types of objects for each region with reference to the estimation result of the estimation process.

17. A program for causing a computer to execute an acquisition process of acquiring learning images, and a learning process of machine learning a difficulty level estimation model for estimating the difficulty level of object detection for a plurality of types of objects in each region included in a target image with reference to the learning images.

Citation Information

Patent Citations

  • Difficulty map creation device

    JP2022079164A

  • Object recognition device and object recognition system

    WO2016199244A1