Trained model, image processing device, imaging device, and program
A trained model and image processing device accurately separate and focus on multiple subjects in a camera by assigning distance-based identification values, effectively addressing misfocus issues in autofocus systems.
Patent Information
- Application Number
- PCT/JP2025/022711
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2025-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing autofocus technologies in cameras struggle with accurately identifying and separating multiple subjects at different distances and orientations in a captured image, particularly when they are close to each other, leading to potential misfocus on unintended subjects.
A trained model that assigns uniform identification values based on distance to subject areas in training images, generating information for subject separation, and an image processing device that identifies subjects using these values, along with depth and defocus information, to accurately track and focus on intended subjects.
Enables high-speed and low-computational-cost subject separation, preventing autofocus from shifting to unintended subjects by distinguishing subjects based on distance and orientation, even when they are close together.
Smart Images

Figure JP2025022711_02012026_PF_FP_ABST
Abstract
Description
Trained model, image processing device, imaging device, and program
[0001] The present invention relates to a trained model, an image processing device, an imaging device, and a program. This application claims priority to Japanese Patent Application No. 2024-101408, filed on June 24, 2024, the contents of which are incorporated herein by reference.
[0002] Conventionally, in cameras equipped with an autofocus (AF) function, there is a technology that prevents the autofocus target from shifting to a subject unintended by the user by identifying each subject being captured (see, for example, Patent Document 1).
[0003] JP 2024-042614 A
[0004] One aspect of the present invention is a trained model that uses training data in which each pixel in a subject area of a subject included in a training image is assigned a uniform first identification value that corresponds to the distance from the imaging position to the position of the subject, and generates, for an input image containing multiple subjects, information in which different second identification values are estimated for each of the subject areas of multiple subjects that are located at different distances from the imaging position of the input image.
[0005] One aspect of the present invention is an image processing device that includes an image supply unit that supplies an input image containing multiple subjects to a trained model that generates information that estimates different second identification values for each of the multiple subjects' subject regions that are located at different distances from the imaging position of the input image; a result acquisition unit that acquires results generated by the trained model; and an identification unit that, based on the distribution of identification values for each position in the captured image shown by the results, identifies each of the subjects represented by subject shadows that have different identification values from each other, among each of the subject shadows that are collections of uniform identification values corresponding to the subjects, as separate subjects.
[0006] One aspect of the present invention is an imaging device including the image processing device described above.
[0007] One aspect of the present invention is a program that causes a computer to execute the following steps: an image supply step of supplying an input image containing multiple subjects to a trained model that generates information that estimates different identification values for each of the multiple subjects' subject regions that are located at different distances from the imaging position of the input image; a result acquisition step of acquiring results generated by the trained model; and an identification step of identifying, as separate subjects, each of the subjects represented by subject shadows that have different identification values from each other, among each of the subject shadows, which are collections of uniform identification values corresponding to the subjects, based on the distribution of the identification values for each position in the captured image shown by the results.
[0008] FIG. 1 is a block diagram for explaining an example of the configuration of an imaging device according to the present embodiment. FIG. 2 is a block diagram for explaining an example of the functional configuration of an image processing unit according to the first embodiment. FIG. 3 is a diagram for explaining input data and teacher data used by a trained model during learning. FIG. 4 is a diagram for explaining the correspondence between distance and area. FIG. 5 is a diagram for explaining an example of generating teacher data to which an area-based identification value is assigned. FIG. 6 is a diagram for explaining an example of data estimated by a trained model. FIG. 7 is a flowchart for explaining an example of the processing flow of an image processing unit according to the first embodiment. FIG. 8 is a block diagram for explaining an example of the functional configuration of an image processing unit according to the second embodiment. FIG. 9 is a flowchart for explaining an example of the processing flow of an image processing unit according to the second embodiment. FIG. 10 is a block diagram for explaining an example of the internal configuration of an image processing unit according to the present embodiment.
[0009] One method for identifying each subject (subject separation) is to perform instance segmentation on a captured image. However, instance segmentation requires high computational costs and may not be suitable for autofocus functions that require real-time and high speed.
[0010] There is also a method for identifying each subject by performing object detection on a captured image to detect each subject. However, with methods using object detection, the closer the distance between detected subjects in the captured image, the more difficult it tends to be to identify each subject, making it difficult to properly identify subjects that are close to each other in the captured image.
[0011] There is also a method for identifying each subject using a defocus map acquired when implementing an autofocus function. However, this identification method using a defocus map has a problem in that it is difficult to identify two or more subjects that are at the same distance from the imaging device because the identification is performed using the distance to the imaging device. Furthermore, if the number of AF measuring points used to acquire the defocus map is small, the resolution of the defocus map becomes low, and it may be impossible to properly identify each subject.
[0012] [Embodiment 1] A preferred embodiment of a trained model, image processing device, imaging device, and program according to the present embodiment, which can perform suitable object separation, will be described in detail below with reference to the accompanying drawings. In the drawings, identical or similar parts are designated by identical or similar reference numerals. Note that the present embodiment is not limited to these embodiments and includes various modifications or improvements. In other words, the components described below include those that can be easily imagined by a person skilled in the art and those that are substantially identical, and the components described below can be combined as appropriate. Furthermore, in the present embodiment, various omissions, substitutions, or modifications of components may be made without departing from the spirit of the present invention.
[0013] 1 is a block diagram illustrating an example of the configuration of an imaging device 1 according to this embodiment. The imaging device 1 includes an imaging unit 11, an image processing unit 12, and a control unit 13. The imaging device 1 is, for example, a camera. The imaging device 1 may also be, for example, an information processing device such as a smartphone, a tablet terminal, or a personal computer.
[0014] The imaging unit 11 has multiple lenses, a motor for driving at least some of the multiple lenses, and an imaging element. Light from an object (subject) (not shown) is collected by the multiple lenses to form a subject image on the imaging surface of the imaging unit 11. The imaging unit 11 photoelectrically converts the light that forms the subject image using the imaging element to generate pixel signals. The imaging unit 11 outputs the pixel signals to the image processing unit 12.
[0015] The image processing unit 12 acquires an image signal from the imaging unit 11. The image processing unit 12 processes the image signal to generate image (captured image) data showing a subject. The image processing unit 12 also processes the image signal to identify each subject shown in the captured image. In other words, the image processing unit 12 performs subject separation.
[0016] The control unit 13 controls each unit in the imaging device 1 based on a control program stored in a storage unit (not shown). The control unit 13 may display image data on a display unit such as a viewfinder provided in the imaging device 1, or may store the image data in a storage medium (not shown). This allows a user operating the imaging device 1 to capture an image of a subject. The control unit 13 also controls the drive of a motor provided in the imaging unit 11 based on the result of subject separation by the image processing unit 12. For example, the control unit 13 controls the drive of a motor to autofocus on a specific subject. This makes it possible to track and focus on the specific subject. The control unit 13 also controls the image processing unit 12 to perform image processing on the captured image.
[0017] FIG. 2 is a block diagram illustrating an example of the functional configuration of the image processing unit 12 according to the first embodiment. The image processing unit 12 includes an image supply unit 121, a trained model 122, a result acquisition unit 123, an identification unit 124, a depth information acquisition unit 125, a defocus information acquisition unit 126, and an orientation estimation unit 127. Each of these functional units is implemented using, for example, a computer including a central processing unit (CPU) and memory, and software. Each functional unit may also be implemented using electronic circuits as necessary. Furthermore, each functional unit does not necessarily have to be included in a single device, and the image processing unit 12 may be configured from multiple devices. The image processing unit 12 may also be referred to as an image processing device.
[0018] The image supply unit 121 acquires a captured image from the imaging unit 11. The image supply unit 121 supplies the acquired captured image to the trained model 122.
[0019] The trained model 122 generates information required to identify each subject depicted in the captured image based on the supplied captured image. Hereinafter, the information generated by the trained model 122 and the training method of the trained model 122 will be described with reference to FIGS.
[0020] 3A and 3B are diagrams for explaining input data and training data used during training by the trained model 122. In Fig. 3A, captured images of multiple people are shown as input data for the trained model 122.
[0021] 3(B) shows, as training data for the trained model 122, data in which the captured image is divided into regions for each subject, and each pixel in the region is labeled with the same identification value for each subject. That is, the training data is data in which each subject is represented by a subject shadow, which is a collection of uniform identification values. The subject shadow has a contour similar to that of the corresponding subject. The identification value is a value required to identify each subject, and corresponds to the distance from the imaging position (the position of the imaging device 1) to the position of the subject. The identification value may be, for example, a value obtained by actually measuring the distance from the imaging position to the position of the subject.
[0022] Furthermore, the identification value may be a value based on a value obtained by estimating the distance from the imaging position to each position shown in the captured image using, for example, an existing monocular depth estimation method. When the distance (depth) estimated by monocular depth estimation is used as the identification value, the distance may differ for each part of the subject (for a person, the face, hands, feet, etc.), so the identification value is a representative value of the distance for each part of the subject. A representative value represents a distribution of values with a single value, such as the mode, average, or median.
[0023] The identification value may also be a value determined based on the area within the subject's shadow. The area within the subject's shadow may be the number of pixels that make up the subject's shadow. FIG. 4 is a diagram illustrating the correspondence between distance and area. FIG. 4A shows a subject imaged 5 m away from the imaging device 1. FIG. 4B shows a subject imaged 10 m away from the imaging device 1. Compared to the subject's shadow imaged 5 m away, the subject's shadow imaged 10 m away is smaller and has a smaller area. Therefore, the area of the subject's shadow corresponds to the distance from the imaging device 1 to the subject.
[0024] Furthermore, the value determined based on the area within the subject shadow may be the number of pixels that make up the contour of the subject. Similar to the number of pixels that make up the subject shadow, the number of pixels that make up the contour of the subject increases as the distance from the imaging device 1 to the subject decreases, and decreases as the distance from the imaging device 1 to the subject increases. Specifically, contour extraction may be performed on the subject in the image, and the number of pixels that make up the extracted contour of the subject may be used as the identification value. Note that the identification value may also be the number of pixels that make up the contour of the subject shadow.
[0025] The identification value may be the value obtained by the various methods described above, or may be a value obtained by performing a predetermined calculation on the obtained value. In particular, when the area of the subject shadow is used as the identification value, the area value calculated to be within a predetermined range, such as 0 to 100, may be used as the identification value, since the obtained area value varies depending on the resolution of the captured image. Specifically, the identification value may be a value obtained by multiplying the number of pixels of the subject shadow relative to the number of pixels of the captured image by 100. The identification value may also be a value based on a power of the area value. By using a power of the area value as the identification value, the difference in identification values between multiple subject shadows that appear small and far away and therefore have a small area difference and are difficult to distinguish can be increased, making them easier to distinguish.
[0026] By using the input data and training data described above, the trained model 122 is trained to generate output data for the input data in which a value corresponding to the distance from the imaging position to the object position is labeled as an identification value for the object region. Below, a description will be given of a case in which the trained model 122 has learned a value based on the area of the object shadow as the identification value. Note that the object region is the region in the captured image in which the object is represented. A separate object region is defined for each object. That is, there is one object region for one object, and two object regions for two objects. Note that the object region is approximately the same as the region indicated by the object shadow.
[0027] 5A and 5B are diagrams illustrating an example of generating training data to which area-based classification values are assigned. The training data used in the trained model 122 is generated using data obtained by performing instance segmentation on captured images, which are input data. In FIG. 5A, captured images of five people are shown as input data.
[0028] FIG. 5B shows the result of instance segmentation performed on the captured image shown in FIG. 5A. In the data after instance segmentation, each subject is divided into regions, and pixels within the regions are labeled with an ID that identifies each person. The IDs that identify each person are, for example, "Person 1" to "Person 5." Note that, although a person is shown as an example of a subject in the above description, this embodiment is not limited to this example. In the following description, a subject in this embodiment is an object that is detected by the imaging device. The subject may be, for example, an animal such as a dog, cat, bird, or insect, or a vehicle such as a car, airplane, or train.
[0029] 5C shows training data created using the data after instance segmentation. The training data is created by labeling each pixel in each object region of the data after instance segmentation with a uniform classification value. This reduces the effort required to determine the region of each object when creating the training data.
[0030] Specifically, when creating training data using a value based on the area of the subject's shadow as the identification value, the identification value of the subject indicated by "Person 1" may be determined based on the number of pixels in the area labeled "Person 1." Also, when creating training data using a value based on the area of the subject's shadow as the identification value, contour extraction may be performed on the area labeled "Person 1," and the number of pixels that make up the extracted contour may be used to determine the identification value of the subject indicated by "Person 1." The identification values of "Person 2" to "Person 5" are determined in the same manner as "Person 1."
[0031] FIG. 6 is a diagram illustrating an example of training data for the trained model 122. FIG. 6 illustrates an example of training data in which images of two people and two cars are supplied to the trained model 122 as input data. The trained model 122 may have multiple channels for each type of subject to be autofocused. The type of subject is classified based on the relationship between the distance of the subject to the imaging device and the area of the subject's shadow. For example, generally different types of subjects such as people and cars, and vehicles of the same type such as cars and trucks that have different relationships between distance and area, are classified as different types. Note that the subject to be autofocused may be, for example, animals such as people, dogs, cats, birds, and insects, or vehicles such as cars, airplanes, and trains.
[0032] In FIG. 6 , the trained model 122 includes a “Person” channel and a “car” channel. The ratio of area to distance for cars and people is significantly different. If the ratio of identification value to area of the subject shadow is the same for people and cars, the “Person” channel, which is learning the characteristics of the person's outline shape, estimates the identification value of a person at a distance of 5 m as 50 and the identification value of a person at a distance of 10 m as 20. Meanwhile, the “car” channel, which is learning the characteristics of the car's outline shape, estimates the identification value of a car at a distance of 10 m as 60 and the identification value of a car at a distance of 15 m as 25. To ensure that different types of subjects are assigned the same identification value at the same distance, the trained model 122 uses the identification value of one type of subject in the training data as a reference and adjusts the identification values of other types of subjects so that they have the same identification value at the same distance. That is, in the training data for the trained model 122, the ratio of identification value to area of the subject shadow may differ for each type of subject. 6 shows an example in which the ratio of the identification value to the area in the "car" channel is adjusted to one-third of the ratio of the identification value to the area in the reference "Person" channel. Note that the identification values shown in FIG. 6 are values added for the purpose of explanation, and the embodiment is not limited to these values.
[0033] Returning to FIG. 2, the functional units of the image processing unit 12 will be further described.
[0034] The result acquisition unit 123 acquires the results estimated by the trained model 122. Specifically, the result acquisition unit 123 acquires data in which the identification values of the respective pixels of the captured image are estimated based on the captured image acquired by the image supply unit 121.
[0035] The identification unit 124 separates and identifies each subject based on the information output from the result acquisition unit 123. Specifically, the identification unit 124 identifies subjects whose subject shadows have different identification values as separate subjects based on the distribution of identification values for each pixel in the captured image indicated by the estimation result. The identification unit 124 outputs the identification result to the control unit 13.
[0036] The depth information acquisition unit 125 acquires depth information, which estimates the distance from the image capturing position to the subject position for each pixel in the captured image, from a depth estimation unit (not shown). The depth information may be, for example, a depth map that is the result of monocular depth estimation performed on the captured image acquired by the image supply unit 121. The depth information may also be, for example, information acquired by a TOF (Time of Flight) sensor.
[0037] The defocus information acquisition unit 126 acquires the defocus amount for each detection area in the captured image from a defocus map generation unit (not shown). The autofocus method may be any method, such as phase difference AF using a separator lens and a phase difference AF sensor, image plane phase difference AF using an image sensor equipped with image plane phase difference pixels, laser AF using reflected waves of infrared rays or ultrasonic waves, or contrast AF using contrast between multiple captured images.
[0038] The orientation estimation unit 127 acquires the results of the identification of each subject by the identification unit 124. The orientation estimation unit 127 also acquires depth information from the depth information acquisition unit 125. The orientation estimation unit 127 also acquires the defocus amount from the defocus information acquisition unit 126. The subject shadow, which is the result estimated by the trained model 122, contains information on the subject's area and position, but does not contain information on the distance to the position of each part of the subject, such as the subject's nose. In other words, the subject shadow does not contain three-dimensional information about the subject. For a subject identified by the identification unit 124, the orientation estimation unit 127 estimates the orientation of the subject relative to the imaging direction based on at least one of the difference in distance to each position in the subject shadow from the imaging position and the defocus amount, which are obtained from the depth information. Specifically, the orientation estimation unit 127 may estimate the orientation of the subject based on the position of a characteristic part, such as the position of the subject's nose, or a part with a specific shape, such as the bending direction of a joint. The orientation estimation unit 127 outputs information indicating the estimated orientation of the subject to the control unit 13. The orientation estimation unit 127 may estimate the orientation of the subject based on both the depth information acquired by the depth information acquisition unit 125 and the defocus amount acquired by the defocus information acquisition unit 126. Using the two pieces of information, the depth information and the defocus amount, the orientation estimation unit 127 can estimate the orientation of the subject with higher accuracy. This can improve the accuracy of autofocusing and tracking of the subject by the imaging device 1. The orientation estimation unit 127 may also estimate the orientation of the subject using existing posture estimation from an image, without relying on the depth information acquisition unit 125 and the defocus information acquisition unit 126.
[0039] The control unit 13 tracks the subject based on the information acquired from the identification unit 124, thereby distinguishing and tracking multiple subjects that are located at different or the same distance from the imaging unit 11 and are located close to each other in the captured image. By distinguishing between subjects in such positional relationships, the control unit 13 prevents the autofocus target from shifting to an unintended subject. Furthermore, the control unit 13 predicts the direction in which the subject will move based on the information acquired from the orientation estimation unit 127, and tracks and appropriately focuses on the subject.
[0040] The control unit 13 may also distinguish between multiple subjects based on information acquired from the identification unit 124 and information acquired from the orientation estimation unit 127. For example, if the trained model 122 is trained using a value based on the area of the subject's shadow as the identification value, the same identification value may be estimated for subjects located at the same distance from the imaging unit 11 and close to each other in the captured image. If the subjects are located close to each other in the captured image and have the same identification value, it is difficult to distinguish between the subjects based solely on the information acquired from the identification unit 124. The control unit 13 according to this embodiment may distinguish between the subjects based on the differences in the orientations of the subjects indicated by the information acquired from the orientation estimation unit 127 in addition to the information acquired from the identification unit 124. By distinguishing between the subjects in the above-described positional relationship, the control unit 13 prevents the autofocus target from shifting to an unintended subject.
[0041] The control unit 13 may cause a display unit such as a viewfinder included in the imaging device 1 to display a warning message for an occluded subject among multiple subjects that are located at different distances from the imaging unit 11 and are located close to each other in the captured image (subjects that are displayed overlapping each other in the captured image). The warning message may be, for example, an icon or text displayed near the occluded subject to notify the subject that it is occluded. The warning message may also be, for example, a frame surrounding the occluded subject that is different in color or shape from a frame surrounding an unoccluded subject. This allows the user operating the imaging device 1 to realize that the subject is being captured in an occluded state.
[0042] Furthermore, the control unit 13 may identify the subject using a known subject detection method such as CV (Computer Vision), and when a subject registered as a subject for which autofocus (AF), auto exposure (AE), or the like in the imaging device 1 is prioritized is occluded, the control unit 13 may display a warning on the display unit to warn that the subject is occluded. This allows the user operating the imaging device 1 to realize that an image is being captured while the subject is occluded. This allows the user operating the imaging device 1 to realize that an image of a subject to be captured via a wired connection is being captured while the subject is occluded.
[0043] FIG. 7 is a flowchart illustrating an example of the processing flow of the image processing unit 12 according to the first embodiment. The image processing unit 12 supplies the captured image acquired from the imaging unit 11 to the trained model 122 (step S101). The image processing unit 12 acquires information generated by the trained model 122 (step S102). The image processing unit 12 classifies subjects with different identification values as distinct subjects based on the distribution of identification values estimated by the trained model 122 (step S103). The image processing unit 12 acquires depth information and a defocus amount from a functional unit (not shown) (step S104). The image processing unit 12 estimates the orientation of the subject based on the depth information and the difference in distance between each part of the subject indicated by the defocus amount (step S105). Note that in FIG. 7, the image processing unit 12 does not necessarily acquire both the depth information and the defocus amount; it may acquire only one of them.
[0044] In the above description, an example is shown in which the trained model 122 and the depth estimation unit (not shown) estimate a classification value and a depth for each pixel in a captured image. However, this embodiment is not limited to this example, and estimation is not necessarily required for each pixel. Specifically, the trained model 122 and the depth estimation unit may estimate a classification value and a depth at a resolution different from that of the captured image acquired by the image supply unit 121 (for example, one-fourth).
[0045] In the above description, the trained model 122 does not necessarily have to be provided in the image processing unit 12. In this case, the image processing unit 12 may read out and use the trained model 122 stored in a storage unit (not shown).
[0046] In the above description, an example is shown in which the image processing unit 12 includes the depth information acquisition unit 125 and the defocus information acquisition unit 126. However, the present embodiment is not limited to this example, and the image processing unit 12 may include either the depth information acquisition unit 125 or the defocus information acquisition unit 126. Furthermore, the image processing unit 12 does not necessarily have to include the depth information acquisition unit 125 or the defocus information acquisition unit 126.
[0047] Summary of First Embodiment According to the above-described embodiment, the trained model 122 uses training data in which each pixel in a subject region of a subject included in a captured image is assigned a uniform identification value corresponding to the distance from the imaging position to the subject position, to learn relationships. The trained model 122 generates information from the captured image that estimates the identification values of at least some of the pixel positions in the captured image. The trained model 122 according to the embodiment estimates the subject region while generating information (distribution of identification values) that can identify each subject. By estimating the subject region, the trained model 122 can generate information that distinguishes between subjects that are the same distance from the imaging position to the subject position. Furthermore, the trained model 122 can generate information in which different identification values are assigned to subjects that are close to each other or overlapping each other in the captured image by estimating an identification value according to distance. Furthermore, unlike instance segmentation, the trained model 122 does not perform object detection, and therefore can generate information that can identify each subject with computational cost similar to that of semantic segmentation. Therefore, by using the trained model 122, which generates information that can perform appropriate subject separation at low computational cost, it is possible to prevent the autofocus target from shifting to an unintended subject.
[0048] Furthermore, according to the above-described embodiment, the training data is generated using data obtained as a result of performing instance segmentation on the training images, in which each pixel in the subject region (subject shadow) included in the training images is identically labeled. By using the data obtained as a result of instance segmentation, the creator of the training data does not need to manually segment the captured images into regions for each subject, thereby reducing the effort required to create the training data.
[0049] Furthermore, according to the above-described embodiment, the identification value used for learning is a value based on the area of the subject shadow, which is a collection of uniform identification values. The area of the subject shadow can be easily determined by counting the number of pixels assigned with the ID of the target subject in the instance-segmented data. The trained model 122 according to the embodiment does not require the preparation of new data (values corresponding to distance) to create training data, making it easy to create training data.
[0050] Furthermore, according to the above-described embodiment, during learning, the ratio of the identification value to the value indicating the area differs for each type of subject. The area value according to the distance differs for each type of subject. When a large subject is captured at a distance and a small subject is captured at a close distance, the size of the subject shadows of the two subjects may be the same. By varying the ratio of the identification value to the area for each type of subject, it is possible to estimate the same identification value for different types of subjects at the same distance. Therefore, the trained model 122 according to the embodiment can estimate an identification value according to the distance of the subject.
[0051] According to the above-described embodiment, the image processing unit 12 (image processing device) includes an image supply unit 121 that acquires a captured image and supplies it to the trained model 122; a result acquisition unit 123 that acquires results generated by the trained model 122; and an identification unit 124 that, based on the distribution of identification values for each position in the captured image indicated by the result, identifies each of the objects represented by subject shadows with different identification values as separate objects, among the object shadows, which are collections of uniform identification values corresponding to the objects. The image processing unit 12 according to the embodiment estimates each object region and an identification value corresponding to the distance, thereby enabling identification of objects located close to each other but at different depths in the captured image. Furthermore, the image processing unit 12 uses the trained model 122 that does not perform object detection using a bounding box or the like, enabling identification of each object with low computational cost. The image processing unit 12 according to the embodiment enables high-speed identification of each object and prevents the autofocus target from moving onto an unintended object.
[0052] Furthermore, according to the above-described embodiment, the image processing unit 12 includes a depth information acquisition unit 125 that acquires depth information indicating the distance from the imaging position to the position of the subject for each position in the captured image, and an orientation estimation unit 127 that estimates the orientation of the identified subject relative to the imaging direction based on the difference in distance from the imaging position to each position in the subject indicated by the depth information. The trained model 122 estimates a uniform identification value within each subject region and the subject's subject shadow. Therefore, the information generated by the trained model 122 does not include information on the position of each part within the subject from the imaging device 1. Using the depth map acquired by the depth information acquisition unit 125, the image processing unit 12 can three-dimensionally grasp the subject and estimate the orientation of the subject based on the difference in distance from the imaging device 1 to the position of each part within the subject, i.e., the unevenness of the subject. By grasping the orientation of the subject, the control unit 13 can appropriately track the subject.
[0053] According to the above-described embodiment, the image processing unit 12 further includes a defocus information acquisition unit 126 that acquires a defocus amount for each position in the captured image, and an orientation estimation unit 127 that estimates the orientation of the identified subject relative to the imaging direction based on the difference in distance from the imaging position to each position in the subject area, as indicated by the defocus amount. Using the defocus map acquired by the defocus information acquisition unit 126, the image processing unit 12 can estimate the orientation of the subject by grasping the subject three-dimensionally based on the unevenness of the subject, similar to the case where depth information is used. By grasping the orientation of the subject, the control unit 13 can appropriately track the subject.
[0054] [Embodiment 2] Next, an image processing unit 12A according to embodiment 2 will be described with reference to Figures 8 and 9. In the image processing unit 12 according to embodiment 1, the trained model 122 corrected the identification value for each type of subject at the stage of generating training data. In contrast, the image processing unit 12A according to the embodiment differs from the image processing unit 12 according to embodiment 1 in that the trained model 122 corrects the identification value for each type of subject after generating information. In embodiment 2, the trained model 122 does not have a different ratio of the identification value to the area for each type of subject. Because the trained model 122 corrects the identification value for each type of subject after generating information, when there are many types of subjects in a captured image, each subject can be identified with higher accuracy.
[0055] 8 is a block diagram illustrating an example of the functional configuration of the image processing unit 12A according to embodiment 2. The image processing unit 12A differs from the image processing unit 12 in that it further includes a type identification unit 221 and a correction unit 222.
[0056] The type identification unit 221 acquires information generated by the trained model 122 from the result acquisition unit 123. The type identification unit 221 may also acquire captured images from the image supply unit 121. The type identification unit 221 identifies the type of each subject based on the feature values of each subject indicated in the captured image or the information generated by the trained model 122. The feature values of the subject represent the characteristics of the subject, and may be, for example, the color, shape, texture, size, or contour shape of the subject's shadow. For example, the type identification unit 221 may identify different types of subjects and identify the type of each subject by performing semantic segmentation on the acquired captured image. The type identification unit 221 may also identify the type of each subject from the captured image using a known subject detection method such as CV. The type identification unit 221 can also identify the type of subject indicated by each subject shadow from the shape of the subject shadow indicated in the generated information. The type identification unit 221 outputs the information generated by the trained model 122 and information indicating the type of each subject to the correction unit 222.
[0057] The correction unit 222 acquires information generated by the trained model 122 and information indicating the type of each subject. The correction unit 222 corrects the identification value in each subject shadow according to the type identified by the type identification unit 221. Specifically, the correction unit 222 corrects the identification value of a subject shadow indicating another type of subject based on the ratio of the identification value for one type of subject to the distance from the imaging position to the position of the subject. As a result, even if the types of subjects are different, the same identification value will be obtained when the distance is the same. The correction unit 222 outputs information obtained by correcting the identification value of each subject shadow indicated by the information generated by the trained model 122 to the identification unit 124.
[0058] 9 is a diagram illustrating an example of the processing flow of the image processing unit 12A according to the second embodiment. The image processing unit 12A performs processing similar to steps S101 and S102 to acquire information generated by the trained model 122. The image processing unit 12A identifies the type of each subject from the feature quantities of the subject indicated in at least one of the information generated by the trained model 122 or the captured image (step S201). The image processing unit 12A corrects the identification value of each subject shadow indicated in the information generated by the trained model 122 according to the identified type of subject (step S202). The image processing unit 12A performs processing similar to steps S103 to S105 using the corrected information.
[0059] Summary of Embodiment 2 According to the above-described embodiment, the trained model 122 performs training using a value based on the area of the subject shadow as an identification value. The image processing unit 12A further includes a type identification unit 221 that identifies the type of subject and a correction unit 222 that corrects the identification value of the subject shadow according to the identified type of subject. The identification unit 124 identifies each subject using the identification value corrected by the correction unit 222. When a large subject is captured at a distance and a small subject is captured at a close distance, the size of the subject shadows of the two subjects may be the same. By varying the ratio of the identification value to the area for each type of subject, the same identification value can be estimated even for different types of subjects at the same distance. Therefore, the image processing unit 12A according to embodiment 2 can estimate an identification value according to the distance of the subject.
[0060] FIG. 10 is a block diagram showing an example of the internal configuration of an image processing unit according to this embodiment. At least some of the functions of the image processing unit 12 or 12A can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions contained in a program read from the RAM 902 or the like. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. RAM stands for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from and to the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906. Furthermore, all or part of the functional units provided in the image processing unit 12 or the image processing unit 12A may be realized using hardware such as an ASIC, a PLD, or an FPGA. Furthermore, all or part of the functional units may be realized by a combination of software and hardware.
[0061] Note that all or part of the functions of each unit of the image processing unit 12 or image processing unit 12A in the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.
[0062] Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as recording units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks such as the Internet or over communication lines such as telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above-mentioned programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system.
[0063] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design modifications can be made without departing from the spirit of the present invention. Furthermore, the configurations described in the above-described embodiments and examples can be combined.
[0064] According to the present invention, an effect of enabling suitable subject separation can be obtained.
[0065] 1...imaging device, 11...imaging unit, 12...image processing unit, 13...control unit, 121...image supply unit, 122...trained model, 123...result acquisition unit, 124...identification unit, 125...depth information acquisition unit, 126...defocus information acquisition unit, 127...orientation estimation unit, 221...type identification unit, 222...correction unit
Claims
1. A trained model that uses training data in which a uniform first identification value that corresponds to the distance from the imaging position to the position of the subject is assigned to each pixel in the subject area of a subject included in a training image, and generates information in which different second identification values are estimated for each subject area of multiple subjects that are located at different distances from the imaging position of the input image for an input image that includes multiple subjects.
2. The trained model according to claim 1, wherein the first identification value is a value that increases or decreases depending on the distance from the imaging position to the position of the subject.
3. The trained model according to claim 1 or claim 2, wherein the first discrimination value used for training is a value based on the area of the subject shadow, which is a collection of uniform first discrimination values.
4. A trained model according to any one of claims 1 to 3, wherein the trained model estimates a different second identification value for each type of the plurality of subjects when the distances from the imaging position of the input image to the plurality of subjects are equal.
5. The trained model according to claim 2, wherein during training, the ratio of the first identification value to a value indicating the area of the subject shadow, which is a uniform collection of the first identification values, differs for each type of subject.
6. The trained model according to any one of claims 1 to 5, wherein the training data is generated using data obtained as a result of performing instance segmentation on a training image, in which each pixel in the subject region included in the training image is identically labeled.
7. An image processing device comprising: an image supply unit that supplies an input image including a plurality of subjects to a trained model that generates information estimating different second identification values for each of the plurality of subjects located at different distances from the imaging position of the input image; a result acquisition unit that acquires the results generated by the trained model; and an identification unit that, based on the distribution of the identification values for each position in the captured image shown by the results, identifies each of the subjects represented by subject shadows that are collections of uniform identification values corresponding to the subjects and that have different identification values as separate subjects.
8. An image processing device comprising: an image supply unit that acquires a captured image and supplies it to a trained model described in any one of claims 1 to 6; a result acquisition unit that acquires results generated by the trained model; and an identification unit that, based on the distribution of identification values for each position in the captured image shown by the results, identifies, as separate subjects, each of the subjects represented by subject shadows that have different identification values from each other, among each of the subject shadows that are collections of uniform identification values corresponding to the subject.
9. The image processing device described in claim 7 or claim 8, wherein the trained model learns a value based on the area of the subject shadow as a first identification value that is a uniform value corresponding to the distance from the imaging position to the position of the subject, and further comprises: a type identification unit that identifies the type of the subject; and a correction unit that corrects the second identification value of the subject shadow indicated by the result according to the identified type of the subject, and the identification unit identifies each of the subjects using the second identification value corrected by the correction unit.
10. An image processing device as described in any one of claims 7 to 9, further comprising: a depth information acquisition unit that acquires depth information indicating the distance from the imaging position to the position of the subject for each position in the captured image; and an orientation estimation unit that estimates the orientation of the identified subject relative to the imaging direction based on the difference in distance from the imaging position to the positions of each part constituting the subject, as indicated by the depth information.
11. An image processing device as described in any one of claims 6 to 10, further comprising: a defocus information acquisition unit that acquires a defocus amount for each position of a part that constitutes a captured image; and an orientation estimation unit that estimates the orientation of the identified subject relative to the imaging direction based on the difference in distance from the imaging position to each position within the subject, as indicated by the defocus amount.
12. An imaging device comprising the image processing device according to any one of claims 6 to 11.
13. A program that causes a computer to execute the following steps: an image supplying step of supplying an input image containing multiple subjects to a trained model that generates information estimating different identification values for each of the multiple subjects' subject regions that are located at different distances from the imaging position of the input image; a result obtaining step of obtaining results generated by the trained model; and an identification step of identifying, as separate subjects, each of the subjects represented by subject shadows that are collections of uniform identification values corresponding to the subjects, with different identification values, based on the distribution of identification values for each position in the captured image shown by the results.
Citation Information
Patent Citations
Focus adjustment device and imaging device
JP2009192774A
Image communication device
JP2017022600A
Application device, application method, and application program
JP2018045517A
Remote operation auxiliary system, remote operation auxiliary method, and program
JP2024033189A