Image processing device, image processing method, and program
The image processing apparatus uses subject detection and classification methods to accurately identify subjects of interest by analyzing overlap and correction areas, enhancing autofocus and tracking control in image processing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-01
AI Technical Summary
Existing image processing technologies struggle to accurately distinguish between athletes or performers and the surrounding audience in images, leading to misidentification of subjects during autofocus and tracking control, and fail to estimate the position of subjects not in contact with the ground.
An image processing apparatus that includes subject detection, region detection, and classification means to identify subjects of interest by analyzing the overlap between subject regions and correction areas, using machine learning and rule-based methods to determine if subjects are within or outside the correction areas.
Enhances the accuracy of identifying subjects of interest in images, improving autofocus and tracking control by correctly distinguishing between subjects and background elements.
Smart Images

Figure 2026089495000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for detecting a subject in an image.
Background Art
[0002] When a user takes a picture or automatically distributes a picture of a stadium or a venue in sports or an event, in order to make an athlete or a performer the target of autofocus control (AF) or tracking control as a target subject, it is necessary to distinguish the athlete or the performer from the surrounding audience. However, when an athlete or a performer and the audience in the image are adjacent to each other, there is a possibility that the audience may be misrecognized as the target subject.
[0003] Patent Document 1 describes a technique for continuing to focus on a past subject when it is not possible to consider that the subject at the previous focus detection (past subject) and the subject at the current focus detection (new subject) are the same, and when the new subject is farther away than the past subject. Patent Document 2 describes a technique for estimating the position of a subject on a road surface from the ground contact position where the subject contacts the road surface.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] In Patent Document 1, when an audience as a new subject is in front of an athlete as a past subject, the focus is set on the audience. In Patent Document 2, it is necessary that the subject is in contact with the ground, and the position of a subject at a position higher than the ground or the position of a subject cut off from the image cannot be estimated.
[0006] This invention has been made in view of the above problems, and its purpose is to realize a technology that can easily identify a subject of interest in an image and improve the processing accuracy for that subject of interest. [Means for solving the problem]
[0007] To solve the above problems and achieve the objective, the image processing apparatus of the present invention includes acquisition means for acquiring image data, subject detection means for detecting the region of a subject in the image data, region detection means for detecting a first region in the image data that is different from the region of the subject, and classification means for classifying subjects that are located within the first region and subjects that are located outside the first region based on the overlap between the first portion of the region of the subject and the first region. [Effects of the Invention]
[0008] According to the present invention, it is possible to easily identify a subject of interest in an image and improve the processing accuracy for that subject of interest. [Brief explanation of the drawing]
[0009] [Figure 1] A block diagram illustrating the device configuration of Embodiment 1. [Figure 2] A flowchart illustrating the control process of Embodiment 1. [Figure 3] A diagram illustrating the determination target area of Embodiment 1. [Figure 4] A diagram illustrating the determination result of whether the area is inside or outside the region in Embodiment 2. [Figure 5] A diagram illustrating the results of detecting the competition area in Embodiment 2. [Figure 6] A flowchart illustrating the competition area detection process of Embodiment 3. [Figure 7] A block diagram illustrating the device configuration of Embodiment 4. [Modes for carrying out the invention]
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0011] In this embodiment, an example will be described in which when photographing or automatically distributing the state of an arena or a venue in sports, events, etc., a subject of interest existing in a field, a stage, etc. in an image is discriminated, and autofocus control (AF) or tracking control is performed on the subject of interest, or automatic editing and automatic distribution are performed.
[0012] [Embodiment 1] First, Embodiment 1 will be described with reference to FIG. 1.
[0013] In Embodiment 1, an example will be described in which an image processing apparatus is applied to an imaging apparatus, a subject of interest existing in a field, a stage, etc. in an image is discriminated, and AF control or tracking control is performed on the subject of interest.
[0014] Note that the imaging apparatus of this embodiment is applied to a digital still camera, a digital video camera, a smartphone, a tablet computer, etc.
[0015] [Device Configuration] First, with reference to FIG. 1, the configuration and functions of the imaging apparatus of this embodiment will be described.
[0016] FIG. 1 is a block diagram illustrating the configuration of the imaging apparatus of this embodiment.
[0017] The imaging apparatus 100 of this embodiment includes a lens unit 101 controlled by a main control unit 140. The lens unit 101 constitutes a photographing optical system that forms an optical image of a subject, which is reflected light of the subject, on the imaging unit 131 in accordance with the control of the main control unit 140.
[0018] The lens unit 101 includes a fixed first-group lens 102, a zoom lens 103 driven by a zoom lens driving unit 104, an aperture 105 driven by an aperture driving unit 106, a fixed third-group lens 107, and a focus lens 109 driven by a focus lens driving unit 110.
[0019] The zoom lens 103 moves in the optical axis direction to change the focal length and perform a zoom operation. The aperture 105 changes the aperture diameter to adjust the amount of light of the subject image formed on the imaging surface of the imaging unit 131. The focus lens 109 has a focus lens function for correcting the movement of the focal plane associated with the zoom operation and a compensator lens function for adjusting the focus state.
[0020] The zoom control unit 121 performs zoom control to change the focal length by controlling the motor of the zoom lens driving unit 104 to drive the zoom lens 103 in accordance with the control of the main control unit 140. The aperture control unit 122 performs exposure control to adjust the amount of light during shooting by controlling the motor of the aperture driving unit 106 to drive the aperture 105 in accordance with the control of the main control unit 140 and adjusting the aperture diameter of the aperture 103. The focus control unit 124 performs AF to adjust the focus state of the subject by controlling the motor of the focus lens driving unit 110 to drive the focus lens 109 in accordance with the control of the main control unit 140.
[0021] Each lens of the lens unit 101 is usually composed of a plurality of lenses, but in FIG. 1, it is shown simplified as only one lens.
[0022] The subject image formed on the imaging surface of the imaging unit 131 by the lens unit 101 is converted into an electrical signal by the imaging unit 131. The imaging unit 131 is an image sensor equipped with photoelectric conversion elements such as a CCD or CMOS that photoelectrically convert the subject image (optical image) into an electrical signal. The imaging unit 131 has m pixels of photoelectric conversion elements arranged horizontally and n pixels of photoelectric conversion elements arranged vertically. The image signal generated by the imaging unit 131 is subjected to predetermined signal processing by the imaging signal processing unit 132 and output as image data. This allows the image of the imaging surface to be acquired. For example, with NTSC and FHD60p settings, image data equivalent to 1920 pixels × 1080 pixels will be acquired for each frame (1 / 60 second).
[0023] Image data processed by the imaging signal processing unit 132 is output to the imaging control unit 133 and temporarily stored in the volatile memory 143. The image data stored in the volatile memory 143 is subjected to various image processing by the image processing unit 141, and after compression processing is performed by the compression / decompression unit 142, it is recorded on a recording medium 147 such as a memory card.
[0024] The compression / decompression unit 142 compresses and encodes the image data output from the image processing unit 141 using a video or still image compression method and records it as an image file on the recording medium 147, or decodes the image file read from the recording medium 147. The recording medium 147 can be a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The recording medium 147 may be configured to be removable from the imaging device 100, or it may be configured not to be easily removable from the imaging device 100.
[0025] The image processing unit 141 applies predetermined image processing to the image data stored in the volatile memory 143. These predetermined image processing steps include resizing to an optimal size (such as scaling), calculating the similarity of images between frames, and performing gamma correction and white balance processing based on the subject area. The image processing unit 141 also generates display data based on the processed image data and sends it to the display unit 145, thereby displaying a preview image or live view image on the display unit 145. Furthermore, the image processing unit 141 generates display data by superimposing the subject detection results from the subject detection unit 150 onto the processed image data and sends it to the display unit 145, thereby displaying an image including the subject detection results on the display unit 145.
[0026] The subject detection unit 150 performs subject detection processing on the image data, detects the area of the subject in the image, and stores the subject detection results (information such as the subject's posture, center of gravity, and the position of the face and eyes) in the volatile memory 143. In this embodiment, the subject is a person, such as a competitor in a sports competition, performer at an event, or spectator. However, the subject detected by the subject detection unit 150 is not limited to a person; it may also be a vehicle or an animal.
[0027] The region detection unit 151 performs region detection processing on the image data and detects a specific region in the image that is different from the subject region. The specific region is, for example, a field, court, goal where a sports competition is held, or a stage at an event venue. In this embodiment, the region detection unit 151 detects the competition area where a sports competition is held, but it is not limited to this. The region detection unit 151 may detect the competition area based on the image data, or it may acquire positional information about the competition area from an external source. If the competition area has a rectangular outline, the region can be identified from positional information at four points: the upper left, lower left, upper right, and lower right. As an example of a region other than a rectangle, in the case of track and field, the competition area (track) is elliptical, so the competition area can be identified from the positional information of the outline.
[0028] The subject classification unit 152 generates a judgment target area corrected for the competition area based on the area detection results of the subject detection unit 150 and the area detection unit 151, and classifies subjects that are within the judgment target area from subjects that are outside the judgment target area (in-area / out-of-area determination). The image processing unit 141 can determine whether a subject is a competitor or a spectator based on the in-area / out-of-area determination result of the subject classification unit 152.
[0029] The tracking control unit 153 works in cooperation with the focus control unit 124 to perform tracking control that keeps the focus on the subject of interest based on the subject classification unit 152's determination of whether the subject is inside or outside the area. In this embodiment, the subject of interest is a competitor who is present within the competition area, and may be one or more people.
[0030] The vibration detection unit 154 includes a gyro sensor, an accelerometer, an electronic compass, etc., and detects vibrations of the imaging device 100. The vibration detection unit 154 detects the amount of vibration in the three orthogonal axis directions of the imaging device 100 and detects the amount of change in the position and orientation of the imaging device 100.
[0031] The main control unit 140 uses the volatile memory 143 as a ring buffer to buffer image data of multiple frames captured within a predetermined period, detection results from the subject detection unit 150 based on the image data, area detection results from the area detection unit 151, area inside / outside determination results from the subject classification unit 152, and detection results from the shake detection unit 154.
[0032] The display unit 145 displays images being captured (live view), captured still images, videos being recorded, detected subjects and subjects of interest in the displayed image, and a GUI for interactive operation. The display unit 145 is a display device such as a liquid crystal display or an organic EL display. The display unit 145 may be an integrated component of the imaging device 100, or it may be an external device connected to the imaging device 100.
[0033] The operation unit 146 consists of operating elements such as switches, buttons, rings, and levers that accept user input and output an operation signal corresponding to the operating element operated by the user to the main control unit 140. The main control unit 140 controls each component of the imaging device 100, including the lens unit 101, by outputting control signals based on the operation signals. The operating elements include, for example, a touch panel integrated into the display unit 145. The photographer, as the user, can perform various operations on the imaging device 100 by operating the operation unit 146. In addition, the photographer can make various settings for the imaging device 100 by operating the GUI (Graphical User Interface) displayed on the display unit 145 using the operation unit 146.
[0034] The control unit 146 includes at least a still image capture button, a video recording button, a mode dial, and a power switch. The still image capture button is an operating member for instructing the main control unit 140 to capture a still image. The video recording button is an operating member for instructing the main control unit 140 to capture a video. The mode dial is an operating member for switching the operating mode of the imaging device 100. The mode dial makes it possible to switch the operating mode of the imaging device 100 to one of still image capture mode, video recording mode, or playback mode. The power switch is an operating member for switching the power of the imaging device 100 on or off.
[0035] The power control unit 148 controls the supply of power from the battery 149 to each component of the imaging device 100 according to the state of the imaging device 100, in accordance with the control of the main control unit 140. The battery 149 is a secondary battery or the like that capable of supplying power to operate the imaging device 100.
[0036] In still image shooting mode, the main control unit 140 starts automatic exposure (AE) control and AF control when the still image shooting button is half-pressed. Furthermore, when the still image shooting button is fully pressed, the main control unit 140 executes still image shooting processing, which records the image data captured by the imaging unit 131 onto the recording medium 147.
[0037] Furthermore, in video recording mode, the main control unit 140 performs AE control and AF control on the image data (frames) captured by the imaging unit 131 in response to the first pressing of the video recording button, and continues the video recording process to record a video of a predetermined duration onto the recording medium 147. The video recording process stops in response to the video recording button being pressed again.
[0038] The volatile memory 143 is, for example, DRAM and is used as a buffer memory for temporarily holding image data captured by the imaging unit 131, an image display memory for the display unit 145, and a work area for the main control unit 140.
[0039] The non-volatile memory 144 is, for example, a flash ROM, and stores control programs executed by the main control unit 140. When the power is turned on by user operation and the imaging device 100 is started, the control programs stored in the non-volatile memory 144 are read (loaded) into a portion of the volatile memory 143. The main control unit 140 controls the operation of the imaging device 100 according to the control programs loaded into the volatile memory 143.
[0040] The main control unit 140 performs calculation processing to control the imaging device 100, including the lens unit 101. The main control unit 140 includes hardware processors such as a CPU and MPU that control each component of the imaging device 100. The main control unit 140 controls each component of the imaging device 100 and realizes the functions of the imaging device 100 by, for example, loading a program stored in the non-volatile memory 144 into the volatile memory 143 and executing it. Alternatively, instead of the main control unit 140 controlling the entire imaging device 100, the entire imaging device 100 may be controlled by multiple hardware components (e.g., multiple processors and circuits) sharing the processing.
[0041] The main control unit 140 controls the focus control unit 124 to drive the focus lens 109 based on the focus detection result obtained by the phase difference detection method or the TV-AF method, and performs AF control.
[0042] Furthermore, the main control unit 140 performs automatic exposure (AE) processing to automatically determine exposure conditions (shutter speed or storage time, aperture value, and sensitivity) based on the brightness information of the subject. The brightness information of the subject can be obtained, for example, by the image processing unit 141. The main control unit 140 can also determine exposure conditions based on a predetermined area, such as a person's face.
[0043] Each component of the imaging device 100 is connected via a bus 160 to enable data exchange and is controlled by the main control unit 140.
[0044] The subject detection unit 150 performs subject detection processing using inference processing with machine learning such as deep learning. The trained model used in machine learning is composed of a neural network, and in this embodiment, it is composed of a Convolutional Neural Network (CNN). However, the inference model in this embodiment is not limited to a CNN and may be composed of a neural network such as a Transformer. Furthermore, subject detection processing may be performed using rule-based methods other than machine learning.
[0045] Deep learning-based inference processing can be performed using a GPU (Graphics Processing Unit) or a DSP (Digital Signal Processor). A GPU or DSP is a processor capable of performing large amounts of multiply-accumulate operations, bias addition, and nonlinear processing, and has the computational processing power to perform matrix operations of neural networks in a short amount of time. The inference processing may be performed collaboratively by the CPU of the main control unit 140 and the GPU or DSP of the subject detection unit 150, or it may be performed by either the CPU of the main control unit 140 or the GPU or DSP of the subject detection unit 150.
[0046] The subject detection unit 150 detects the coordinates of a rectangular region circumscribing the detected subject from the image data, using these coordinates as the subject's position and size information. Furthermore, based on the subject's position and size information, the subject detection unit 150 calculates a confidence level (probability value) for each subject, representing its likelihood of being the subject of interest. The confidence level is expressed as an integer value from 0 to 255, with a higher confidence level indicating a lower probability of false detection.
[0047] Furthermore, when the region detection unit 151 detects a competition area, it performs region detection processing using machine learning-based inference, similar to the subject detection unit 150. Alternatively, it may also perform region detection processing using rule-based methods other than machine learning, similar to the subject detection unit 150. The region detection unit 151 may output a rectangular region containing the field, or, if the field is an ellipse or the like, it may output the outline of the field.
[0048] The functions (functional units) of each component of the imaging device 100 in this embodiment are realized by the hardware shown in Figure 1 and / or by software programs executed by the control unit that operates as each of the functional units shown in Figure 1. Alternatively, if the functional units shown in Figure 1 are implemented using hardware instead of software, it is sufficient to have the circuit configurations corresponding to each functional unit shown in Figure 1.
[0049] <Control process of Embodiment 1> Next, with reference to Figure 2, the control process of the imaging device 100 of Embodiment 1 will be described.
[0050] The processing shown in Figure 2 is achieved when the main control unit 140 utilizes a pre-trained model and executes a program stored in the non-volatile memory 144 to control each component shown in Figure 1.
[0051] In step S201, the imaging control unit 133 controls the imaging unit 131 to capture an image, and the imaging signal processing unit 132 processes the image signal obtained by the imaging unit 131 to acquire image data.
[0052] In step S202, the region detection unit 161 detects the competition area in the image data acquired in step S201.
[0053] In step S203, the subject classification unit 152 corrects the competition area detected in step S202 and generates a target area for judgment. Details of the processing in step S203 will be described later in Figure 3.
[0054] In step S204, the subject detection unit 150 detects the subject area in the image data acquired in step S201.
[0055] In step S205, the subject classification unit 152 performs an inside / outside determination to classify subjects into those located within the competition area and those located outside the competition area, based on the determination target area generated in step S203 and the subject area detected in step S204. Details of the inside / outside determination method in step S205 will be described later in Figure 3. The inside / outside determination result in step S205 is stored in the volatile memory 143 along with subject attribute information such as the subject's position.
[0056] In step S206, the focus control unit 124 performs AF control on subjects located within the competition area, based on the area in / out determination result from step S205. Note that this is not limited to AF control; any processing that can be performed using the area in / out determination result, such as tracking control or subject attribute determination, may be executed.
[0057] <Processing to generate areas to be judged> Next, with reference to Figure 3, the process of generating the area to be judged in step S203 of Figure 2 will be explained.
[0058] Figure 3(a) is an example of the arrangement of athletes, spectators, and the field in a sports competition.
[0059] In Figure 3(a), the running figures are athletes, and the standing figures are spectators. Figures 301 and 302 are athletes, and figures 303 and 304 are spectators. For the sake of clarity, the symbols for spectators other than figures 303 and 304 are omitted. Rectangular frames 311 to 314 are subject detection frames corresponding to the regions of figures 301 to 304, which are subjects detected from the image. Region 320 is the competition area (field), and region 321 is a correction area calculated based on the height to which athletes can move, such as by jumping. For example, when taking a picture at approximately the same eye level as an athlete, if the height to which an athlete can physically jump is h [mm], the distance from the imaging device 100 to the edge of the field is z [mm], the focal length of the imaging device 100 is f [mm], and the pixel pitch of the image sensor is Δ [mm], then the jump width in the image is calculated as hf / (zΔ) [pix]. In this embodiment, assuming that the image is taken from outside the field, the distance to the field is substituted with a known field size. If the distance to the field can be obtained by another method, that value may be used. The width 322 of the correction area 321 is set by multiplying that value by a coefficient of 1 or more (e.g., 1.5). This is because the width 322 is a value calculated at the edge of the field, and it is a margin to absorb variations in the height of the competitors' jumps.
[0060] If the imaging device 100 is fixed, or if the depression angle information of the imaging device 100 is known, the width 322 may be calculated by projecting the jump height from z[mm] and h[mm] onto the image using the depression angle information.
[0061] The subject classification unit 152 sets the area comprising the competition area 320 and the correction area 321 as the area to be judged. The subject classification unit 152 then determines whether an object is inside or outside the area based on the degree of overlap between a portion of the subject detection frames 311 to 314 and the area to be judged. In this embodiment, if a portion of the subject detection frames 311 to 314 overlaps with the area to be judged, for example, if the midpoint of the base of the subject detection frames 311 to 314 is within the area to be judged, the subject classification unit 152 determines that the object corresponding to that subject detection frame is within the area to be judged. Conversely, if a portion of the subject detection frames 311 to 314 does not overlap with the area to be judged, for example, if the midpoint of the base of the subject detection frames 311 to 314 is not within the area to be judged, the subject classification unit 152 determines that the object corresponding to that subject detection frame is outside the area to be judged.
[0062] In the example in Figure 3, the midpoints of the bases of the subject detection frames 311 and 312 for athletes 301 and 302 are within the detection area, but the subject detection frames 313 and 314 for spectators 303 and 304 are outside the detection area. By setting a correction area 321, it becomes possible to determine that athletes 301 and 302 are within the detection area even when they jump. Furthermore, by using the base of the subject detection frame, it becomes possible to determine that athletes are within the detection area even when their bodies are inverted, such as in gymnastics.
[0063] Furthermore, the determination of whether an object is inside or outside the area is not limited to the midpoint of the bottom edge of the object detection frame. It may also be determined by using the ratio of the length of the bottom edge included in the area to be determined to the bottom edge, or by the ratio of the overlap between the area of the bottom of the object detection frame and the area to be determined. For example, the bottom 1 / 3 of the object detection frame can be used as the bottom edge. When using the overlap ratio, it is possible to perform not only two classifications (inside or outside the competition area) but also three classifications. For example, if the overlap ratio is less than the first threshold Th_1, it can be determined to be outside the competition area; if the overlap ratio is greater than or equal to the first threshold Th_1 but less than the second threshold Th_2 (>Th_1), it can be determined to be unknown; and if the overlap ratio is greater than or equal to the second threshold Th_2, it can be determined to be inside the competition area. It is also possible to assign a confidence level to the overlap ratio.
[0064] In the example in Figure 3, the explanation assumes that the competition area 320 is at the same height as the ground, but as shown in Figure 3(b), it is also possible to set a height-adjusted area such as a goal as a correction area for the competition area. Figure 3(b) is a diagram illustrating an example of setting a height-adjusted area such as a goal in the competition area. In Figure 3(b), for example, in basketball, there is a high possibility that players will jump near the goal, and when shooting on the court, the accuracy of the determination may be improved by changing the conditions for determining whether an area includes the goal or not. In the example in Figure 3(b), the area 330 is the goal area including the goal attached to the competition area 320, and the subject classification unit 152 can determine that if the midpoint of the upper edge of the subject detection frame 311 is inside the goal area 330, then the subject corresponding to the subject detection frame 311 is inside the goal area 330, and if the midpoint of the upper edge of the subject detection frame 311 is outside the goal area 330, then the subject corresponding to the subject detection frame 311 is outside the goal area 330.
[0065] As described above, according to Embodiment 1, by determining whether or not a subject exists within the determination target area based on the degree of overlap between a part of the subject detection frame and the determination target area, the subject of interest can be easily identified in the image, and the processing accuracy for the subject of interest can be improved.
[0066] Although this embodiment describes an example where the subject is a person, this embodiment can also be applied when the subject is a vehicle, such as at a racetrack, or when the subject is an animal, such as at a horse racing track.
[0067] [Embodiment 2] Next, Embodiment 2 will be described.
[0068] Embodiment 1 described an example in which it is determined whether or not a subject is within the area to be determined based on the degree of overlap between a part of the subject detection frame and the area to be determined. In contrast, Embodiment 2 performs the determination of whether or not a subject is inside or outside the area using time-series information of the subject detection frame.
[0069] In this embodiment, an example is described in which a determination of whether a subject is inside or outside the region is made for each frame of image data as time-series information for the subject detection frame. However, the embodiment is not limited to this, and any time information may be used.
[0070] The following describes an example where there is one or more subjects of interest, the midpoint coordinates of the base of the i-th subject detection frame in the image of frame n are (xi(n), yi(n)), and a determination of whether the subject is inside or outside the region is performed for all frames, followed by tracking control of the subject of interest. Alternatively, instead of all frames, a determination of whether the subject is inside or outside the region may be performed for predetermined frames, and matching to determine whether they are the same subject may be performed to control tracking of the subject of interest.
[0071] Now, referring to Figure 4, we will explain how to determine whether an area is inside or outside a region for each frame of image data.
[0072] In the example in Figure 4, the row shows the ID "i" of the subject being tracked, and the column shows the frame number. The "in" and "out" in each cell indicate whether the i-th subject detection frame (xi(k), yi(k)) is inside or outside the target area in frame k. If the frame of interest is frame n, the subject classification unit 152 counts the number of "in" occurrences m for each subject in the image of frame n within a predetermined number of frames M prior to frame n. If m / M exceeds a predetermined threshold Th_j, it is determined that the subject is inside the target area. In this case, m / M can be considered as the confidence level that subject i is inside the target area.
[0073] Furthermore, if the region detection unit 151 detects the competition area based on image data and updates it at predetermined time intervals, the detection accuracy may deteriorate due to changes in the situation. In that case, it is possible to set a confidence level for each detected area within the competition area using the results of the determination of whether each subject is inside or outside the area up to the (n-1)th frame.
[0074] Next, referring to Figure 5, we will explain how the region detection unit 151 calculates the reliability of its detection when it detects a competition area based on image data.
[0075] In Figure 5, region 501 represents the competition area detected by the region detection unit 151, and region 502 represents the area where spectators are located. Other parts common to Figure 3 are denoted by the same reference numerals and their explanations are omitted.
[0076] The numbers next to each subject indicate the probability of them being within the detection area. In this embodiment, for the sake of explanation, all spectators within area 502 are assumed to have a probability of 20%, but normally the values would differ for each subject. The area detection unit 151 sets the confidence level within the detected competition area 501 as the probability value for each subject detection frame. This makes it possible to suppress the deterioration of the detection accuracy of the competition area 501.
[0077] Furthermore, if the imaging device 100 is fixed, the movement speed of the subject detection frame can also be used. For example, if the difference between the midpoint coordinate yi(n) of the base of the i-th subject detection frame in the n-th frame image and the midpoint coordinate yi(n-1) of the base of the i-th subject detection frame in the (n-1)-th frame image is greater than or equal to a threshold, the subject is excluded from the area in / out of the area determination. This makes it possible to exclude subjects that are highly likely to have jumped from the determination target and to set the determination target area to be approximately the same as the competition area 320.
[0078] As described above, according to Embodiment 2, by using the time-series information of the subject detection frame to determine whether an object is inside or outside the area, it is possible to improve the accuracy of identifying the subject of interest in the image.
[0079] [Embodiment 3] Next, Embodiment 3 will be described.
[0080] Embodiment 3 describes an example in which competition area detection processing is performed according to the type of sports competition or event.
[0081] Figure 6 is a flowchart illustrating the competition area detection process of Embodiment 3 in step S202 of Figure 2.
[0082] The device configuration of Embodiment 3 is the same as that shown in Figure 1 of Embodiment 1, and other configurations identical to those of Embodiment 1 will not be described. The following description will focus on the differences from Embodiment 1.
[0083] In step S601, the region detection unit 151 determines the type of sport, such as soccer or basketball, from the image data. The method for determining the type of sport can be, for example, obtaining information about a sport specified by the user via a GUI, or it can be automatically determined using a dictionary learned through machine learning. The same applies to the type of event, not just sports. The following describes an example of automatically determining the type of sport using a region detection dictionary.
[0084] Furthermore, a dictionary trained using machine learning groups together words and phrases that share some common characteristics.
[0085] In step S602, the region detection unit 151 reads region detection dictionaries specific to courts and goals for each sport, such as soccer and basketball, from the volatile memory 143.
[0086] In step S603, the region detection unit 151 detects the competition area using the region detection dictionary acquired in step S602, and proceeds to the process shown in step S203 in Figure 2.
[0087] The process described above makes it possible to set the target area for judgment with greater accuracy.
[0088] Alternatively, a dictionary that has been trained to allow steps S601 and S603 to be performed simultaneously may be used.
[0089] The subject classification unit 152 may change the conditions for determining whether an object is inside or outside the area based on the competition judgment result in step S601. For example, in competitions where there is little possibility of spectators in the foreground, the determination of whether an object has gone outside the area from below the area to be judged may not be performed. This makes it possible to reduce erroneous determinations regarding whether an object is inside or outside the area when there is a false detection within the area to be judged.
[0090] As described above, Embodiment 3 makes it possible to set the area to be judged with higher accuracy.
[0091] [Embodiment 4] Next, Embodiment 4 will be described.
[0092] Embodiment 4 describes an example in which an image processing device 700 and an imaging device 750 are connected in a communicative manner, the imaging device 750 automatically captures images of a stadium or venue during sports or events, the image processing device 700 performs area detection on the images acquired from the imaging device 750, and automatically edits and / or distributes them.
[0093] The image processing device of this embodiment is applicable to smartphones, tablet computers, desktop computers, etc., that can communicate with an imaging device.
[0094] Figure 7 is a block diagram illustrating the configuration of the image processing device 700 of Embodiment 4.
[0095] In Embodiment 4, multiple imaging devices 750 are connected to an image processing device 700, and the image processing device 700 acquires image data from each imaging device 750 and performs region in / out determination. In this embodiment, if the relative positions of the imaging devices 750 are known, it is also possible to share the region in / out determination results with each imaging device 750.
[0096] The system memory unit 701 is, for example, a flash ROM, and stores programs and operating constants for each function unit of the system control unit 710. The system memory 720 is a volatile memory such as DRAM, and loads constants, variables, and data read from the system memory 720 during the operation of the system control unit 710. The image memory unit 702 is, for example, a flash ROM, and stores image data acquired from the imaging device 750. Figure 7 illustrates a configuration in which two imaging devices 750 are connected, but it may also be one or three or more.
[0097] The system control unit 710 performs calculations to control the image processing device 700. The system control unit 710 includes hardware processors such as a CPU and MPU that control each component of the image processing device 700. For example, the system control unit 710 controls each component of the image processing device 700 and realizes the functions of the image processing device 700 by loading a program stored in the system memory 720 into the system memory 720 and executing it. Alternatively, instead of the system control unit 710 controlling the entire image processing device 700, the entire image processing device 700 may be controlled by multiple hardware components (e.g., multiple processors and circuits) sharing the processing.
[0098] The system control unit 710 includes a subject detection unit 703, an area detection unit 704, a subject classification unit 705, and an image editing unit 706. The functions of the subject detection unit 703, area detection unit 704, and subject classification unit 705 are the same as those of the subject detection unit 150, area detection unit 151, and subject classification unit 152 in Figure 1. In this embodiment, the area in / out determination result of the subject classification unit 705 is fed back to the imaging device 750, enabling each imaging device 750 to perform AF control and tracking control for subjects (subjects of interest) that are present in the competition area based on the area in / out determination result.
[0099] The image editing unit 706 performs at least one of the following: editing and distribution of image data, based on the subject classification unit 705's determination of whether an area is inside or outside the designated area. The image editing unit 706 edits the image data based on the subject classification unit 705's determination of whether an area is inside or outside the designated area, and stores the edited image data in the system storage unit 701. Image editing includes cropping videos focusing on a specific athlete, and automatically extracting highlight scenes of play. The image editing unit 706 also uploads the edited images to a server device 730 that provides cloud services via the internet 740.
[0100] As described above, according to Embodiment 4, it is possible to perform at least one of automatic image editing and automatic distribution based on the determination result of whether an image captured by the imaging device 750 is inside or outside a region.
[0101] The functions (functional units) of each component of the image processing apparatus 700 in this embodiment are realized by the hardware shown in Figure 7 and / or by software programs executed by the control unit that operates as each of the functional units shown in Figure 7. Alternatively, if the functional units shown in Figure 7 are implemented using hardware instead of software, it is sufficient to have the circuit configurations corresponding to each functional unit shown in Figure 7.
[0102] [Other embodiments] This embodiment can also be implemented by supplying a program that implements one or more of the functions of the above-described embodiment to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.
[0103] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention.
[0104] The disclosures herein include the following image processing apparatus, image processing methods, and programs. [Item 1] Means for acquiring image data, A subject detection means for detecting the region of the subject in the aforementioned image data, A region detection means for detecting a first region different from the region of the subject in the image data, An image processing apparatus comprising: classification means for classifying subjects located within the first region and subjects located outside the first region based on the overlap between a first portion of the region of the subject and the first region. [Item 2] The image processing apparatus according to item 1, characterized in that the classification means determines that the subject is present in the first region if the first portion of the subject's region overlaps with the first region, and that the subject is not present in the first region if the first portion of the subject's region does not overlap with the first region. [Item 3] The image processing apparatus according to item 1 or 2, characterized in that the first portion is the bottom edge of the region of the subject. [Item 4] The image processing apparatus according to any one of items 1 to 3, characterized in that the first region includes a region set based on the movable height of the subject. [Item 5] The region detection means detects a second region attached to the first region, The image processing apparatus according to any one of items 1 to 4, characterized in that the classification means classifies subjects located within the second region and subjects located outside the second region based on the overlap between the second portion of the subject region and the second region. [Item 6] The image processing apparatus according to item 5, characterized in that the second portion is the upper edge of the region of the subject. [Item 7] The first area mentioned above is an area that includes either a field, court, or goal in a sports competition, or a stage in an event venue. The image processing apparatus according to any one of items 1 to 6, characterized in that the subject is a person participating in the sports competition or a person appearing at the event venue. [Item 8] The image processing apparatus according to any one of items 1 to 7, characterized in that the classification means classifies each frame of the image data into subjects located within the first region and subjects located outside the first region. [Item 9] The image processing apparatus according to any one of items 1 to 7, characterized in that the classification means classifies subjects located within the first region and subjects located outside the first region based on the speed at which the subjects move. [Item 10] The image processing apparatus according to item 8 or 9, characterized in that the region detection means obtains the confidence level of the result of classifying the subject for each frame of the image data. [Item 11] The image processing apparatus according to item 7, characterized in that the region detection means detects the first region based on the image data and a learned dictionary. [Item 12] The image processing apparatus according to item 11, further comprising a second acquisition means for acquiring the type of sports competition or event venue. [Item 13] The image processing apparatus according to item 12, characterized in that the region detection means switches the dictionary based on the type acquired by the second acquisition means. [Item 14] The image processing apparatus according to item 12, characterized in that the classification means switches the conditions for classifying the subject based on the type acquired by the second acquisition means. [Item 15] An editing means for editing the image data based on the results of classification by the classification means, An image processing apparatus according to any one of items 1 to 14, characterized by having at least one of a storage means for storing an image edited by the editing means and a distribution means for distributing an image edited by the editing means. [Item 16] An imaging means for capturing an image and generating image data, The image processing apparatus according to any one of items 1 to 15, further comprising control means for performing focus control or tracking control with respect to a subject located within the first region. [Item 17] An image processing method performed by an image processing device, Steps to acquire image data, The steps include detecting the region of the subject in the image data, A step of detecting a first region in the image data that is different from the region of the subject, An image processing method characterized by having a step for classifying subjects located within the first region and subjects located outside the first region based on the overlap between a first portion of the region of the subject and the first region. [Item 18] A program to cause a computer to function as an image processing device as described in any of items 1 through 16. [Explanation of Symbols]
[0105] 100, 750…Imaging device, 140…Main control unit, 150, 703…Subject detection unit, 151, 704…Region detection unit, 152, 705…Subject classification unit, 700…Image processing device, 706…Image editing unit
Claims
1. Means for acquiring image data, A subject detection means for detecting the region of the subject in the aforementioned image data, A region detection means for detecting a first region in the image data that is different from the region of the subject, An image processing apparatus characterized by having a classification means for classifying subjects located within the first region and subjects located outside the first region based on the overlap between a first portion of the region of the subject and the first region.
2. The image processing apparatus according to claim 1, characterized in that the classification means determines that the subject is present in the first region if the first portion of the subject's region overlaps with the first region, and that the subject is not present in the first region if the first portion of the subject's region does not overlap with the first region.
3. The image processing apparatus according to claim 1, characterized in that the first portion is the bottom edge of the region of the subject.
4. The image processing apparatus according to claim 1, characterized in that the first region includes a region set based on the movable height of the subject.
5. The region detection means detects a second region attached to the first region, The image processing apparatus according to claim 1, characterized in that the classification means classifies subjects located within the second region and subjects located outside the second region based on the overlap between the second portion of the subject region and the second region.
6. The image processing apparatus according to claim 5, characterized in that the second portion is the upper edge of the region of the subject.
7. The first area is an area that includes either a field, court, or goal in a sports competition, or a stage in an event venue. The image processing apparatus according to claim 1, characterized in that the subject is a person participating in the sports competition or a person appearing at the event venue.
8. The image processing apparatus according to claim 1, characterized in that the classification means classifies each frame of the image data into subjects located within the first region and subjects located outside the first region.
9. The image processing apparatus according to claim 1, characterized in that the classification means classifies subjects located within the first region and subjects located outside the first region based on the speed at which the subjects are moving.
10. The image processing apparatus according to claim 8, characterized in that the region detection means obtains the confidence level of the result of classifying the subject for each frame of the image data.
11. The image processing apparatus according to claim 7, characterized in that the region detection means detects the first region based on the image data and a learned dictionary.
12. The image processing apparatus according to claim 11, further comprising a second acquisition means for acquiring the type of sports competition or event venue.
13. The image processing apparatus according to claim 12, characterized in that the region detection means switches the dictionary based on the type acquired by the second acquisition means.
14. The image processing apparatus according to claim 12, characterized in that the classification means switches the conditions for classifying the subject based on the type acquired by the second acquisition means.
15. An editing means for editing the image data based on the results of classification by the classification means, The image processing apparatus according to claim 1, further comprising at least one of the following: a storage means for storing an image edited by the editing means and a distribution means for distributing an image edited by the editing means.
16. An imaging means for capturing an image and generating image data, The image processing apparatus according to claim 1, further comprising control means for performing focus control or tracking control with respect to a subject located within the first region.
17. An image processing method performed by an image processing device, Steps to acquire image data, The steps include detecting the region of the subject in the image data, A step of detecting a first region in the image data that is different from the region of the subject, An image processing method characterized by having a step for classifying subjects that are located within the first region and subjects that are located outside the first region, based on the overlap between a first portion of the region of the subject and the first region.
18. A program for causing a computer to function as an image processing device according to any one of claims 1 to 16.