Image processing apparatus, image processing method, computer-readable storage medium, and computer program product
By using region detection and classification technology in the image processing device, the problem of spectators being mistakenly identified as subjects of interest was solved, improving the accuracy of distinguishing between contestants and spectators and ensuring the accuracy of focusing and tracking control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2025-11-12
- Publication Date
- 2026-05-22
Smart Images

Figure CN122073005A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image processing apparatus, image processing method, computer-readable storage medium, and computer program product, and particularly to image processing for detecting a subject in an image. Background Technology
[0002] When users photograph or automatically assign scenes such as sports stadiums and event venues, it is necessary to distinguish athletes and performers from the surrounding audience in order to set the athlete or performer as the subject of focus as the autofocus (AF) control target or tracking control target. However, if the athlete or performer is adjacent to the audience in the image, the audience may be mistakenly identified as the subject of focus.
[0003] Japanese Patent Application Publication No. 2006-330567 discloses a technique for continuously focusing on a past subject when the subject at a previous focus detection moment (past subject) and the subject at the current focus detection moment (new subject) cannot be considered the same. Japanese Patent Application Publication No. 2011-065338 discloses a technique for estimating the position of a subject on the road surface based on the contact position between the subject and the road surface.
[0004] According to Japanese Patent Application Publication No. 2006-330567, if the viewer, as a new subject, is closer than the contestant, who was a previous subject, the viewer will be unexpectedly focused on. According to Japanese Patent Application Publication No. 2011-065338, the subject needs to be in contact with the ground, and the position of a subject located above the ground or a subject cropped from the image cannot be estimated. Summary of the Invention
[0005] This disclosure is made in consideration of the aforementioned problems and provides the technical advantage of being able to easily identify subjects of interest in an image and improve the processing accuracy of subjects of interest.
[0006] To address the aforementioned problems, this disclosure relates to an image processing apparatus, comprising: an acquisition unit configured to acquire image data; a subject detection unit configured to detect a subject region in the image data; a region detection unit configured to detect a first region in the image data that is different from the subject region; and a classification unit configured to classify subjects existing within the first region and subjects existing outside the first region based on the overlap between a first portion of the subject region and the first region.
[0007] To address the aforementioned problems, this disclosure relates to an image processing method executed by an image processing apparatus, comprising the steps of: acquiring image data; detecting a subject region in the image data; detecting a first region in the image data that is different from the subject region; and classifying subjects existing within the first region and subjects existing outside the first region based on the overlap between a first portion of the subject region and the first region.
[0008] In order to solve the aforementioned problems, this disclosure relates to a computer-readable storage medium storing a program for enabling a computer to be used as the image processing apparatus specified above.
[0009] In order to solve the aforementioned problems, this disclosure relates to a computer program product that includes a program for enabling a computer to be used as the image processing apparatus specified above.
[0010] According to this disclosure, subjects of interest in an image can be easily identified and the processing accuracy of subjects of interest can be improved.
[0011] The features of this disclosure will become clear from the following description of embodiments with reference to the accompanying drawings. The following description of the embodiments is by way of example. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the embodiments.
[0013] Figure 1 This is a block diagram illustrating the device configuration according to the first embodiment;
[0014] Figure 2 The flowchart illustrates the control process according to the first embodiment;
[0015] Figure 3A and Figure 3B These are views used to interpret and determine the target area according to the first embodiment;
[0016] Figure 4 This is a table illustrating the results of determining the interior / exterior of the region according to the second embodiment;
[0017] Figure 5 This is a view used to explain the motion region detection results according to the second embodiment;
[0018] Figure 6 The flowchart illustrating the motion region detection process is based on the third embodiment; and
[0019] Figure 7 This is a block diagram illustrating the device configuration according to the fourth embodiment. Detailed Implementation
[0020] In the following, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claims. Several features are described in the embodiments, but not all such features are required, and several such features can be appropriately combined. Furthermore, in the drawings, the same or similar configurations are given the same reference numerals, and repeated descriptions thereof are omitted.
[0021] This embodiment will describe an example in which, when shooting or automatically assigning a scene in a sports field or on a stage, an event, or other similar setting, an object of interest present in the image is identified, and then automatic focus control (AF) or tracking control, or automatic editing or assignment is performed.
[0022] [First Embodiment]
[0023] Reference Figure 1 The first embodiment is described.
[0024] The first embodiment describes an example of applying an image processing apparatus to a camera device, identifying a subject of interest present in an image on a site or stage, and performing AF control or tracking control on the subject of interest.
[0025] Note that the camera device in this embodiment is applied to digital cameras, digital camcorders, smartphones, tablets, etc.
[0026] <Device Configuration>
[0027] First, refer to Figure 1 The configuration and function of the camera device according to this embodiment are described.
[0028] Figure 1 This is a block diagram illustrating the configuration of the camera device according to this embodiment.
[0029] The imaging device 100 according to this embodiment includes a lens unit 101 controlled by a main control unit 140. The lens unit 101 forms an imaging optical system, which, under the control of the main control unit 140, causes the imaging unit 131 to form an optical image of the subject as reflected light from the subject.
[0030] The lens unit 101 includes a fixed first lens group 102, a zoom lens 103 driven by a zoom lens driving unit 104, an aperture 105 driven by an aperture driving unit 106, a fixed third lens group 107, and a focusing lens 109 driven by a focusing lens driving unit 110.
[0031] The zoom lens 103 moves along the optical axis to change the focal length, thereby performing a zoom operation. The aperture 105 changes its diameter to adjust the amount of light forming on the image of the subject on the imaging plane of the imaging unit 131. The focusing lens 109 has a focusing lens function that corrects the movement of the focal plane during zooming and a compensating lens function that adjusts the focus state.
[0032] Under the control of the main control unit 140, the zoom control unit 121 drives the zoom lens 103 by controlling the motor of the zoom lens drive unit 104, thereby performing zoom control to change the focal length. Under the control of the main control unit 140, the aperture control unit 122 drives the aperture 105 by controlling the motor of the aperture drive unit 106, thereby performing exposure control to adjust the aperture diameter of the aperture 105 and adjust the amount of light during shooting. Under the control of the main control unit 140, the focus control unit 124 drives the focus lens 109 by controlling the motor of the focus lens drive unit 110, thereby performing AF control to adjust the focus state of the subject.
[0033] Each lens in lens unit 101 is typically formed by multiple lenses, but... Figure 1 The middle part is represented in a simplified way by a lens.
[0034] The subject image formed by the lens unit 101 on the imaging plane of the imaging unit 131 is converted into an electrical signal by the imaging unit 131. The imaging unit 131 is an image sensor including photoelectric conversion elements, such as a CCD or CMOS sensor that photoelectrically converts the subject image (optical image) into an electrical signal. Photoelectric conversion elements with m pixels in the horizontal direction and n pixels in the vertical direction are arranged in the imaging unit 131. The image signal generated by the imaging unit 131 undergoes predetermined signal processing by the image capture signal processing unit 132 and is output as image data. This allows an image to be obtained on the imaging plane. For example, in NTSC and FHD / 60p settings, image data corresponding to 1920 pixels x 1080 pixels is obtained for each frame (1 / 60 second).
[0035] Image data processed by the image capture signal processing unit 132 is output to the imaging control unit 133 and temporarily stored in volatile memory 143. The image data stored in volatile memory 143 is subjected to various image processing by the image processing unit 141, compressed by the image compression / decompression unit 142, and then recorded in a recording medium 147 such as a memory card.
[0036] The image compression / decompression unit 142 compresses and encodes the image data output from the image processing unit 141 using a moving image or still image compression method, so as to record the obtained data as an image file in the recording medium 147, and decodes the image file read from the recording medium 147. The recording medium 147 is a hard disk drive (HDD), a solid-state drive (SSD), a memory card, etc. The recording medium 147 can be configured to be removable from the camera device 100, or not easily removable from the camera device 100.
[0037] Image processing unit 141 applies predetermined image processing to image data stored in volatile memory 143. Predetermined image processing includes scaling (enlarging / reducing to an optimal size), calculating image similarity between frames, and gamma correction and white balance processing based on the subject area. Furthermore, image processing unit 141 generates display data based on the predetermined image processing data and sends the display data to display unit 145, thereby displaying a preview image or live view image on display unit 145. Image processing unit 141 generates display data by superimposing the subject detection result from subject detection unit 150 onto the predetermined image processing data and sends the display data to display unit 145, thereby displaying an image including the subject detection result on display unit 145.
[0038] The subject detection unit 150 performs subject detection processing on the image data to detect subject regions in the image and stores the subject detection results (such as the subject's posture, center of gravity, and information on the position of the face and eyes) in the volatile memory 143. Note that in this embodiment, the subject is a person, specifically an athlete participating in a sport, a performer in an activity, or a spectator. Note that the subject detected by the subject detection unit 150 is not limited to people; it can also be a vehicle or an animal.
[0039] The region detection unit 151 performs region detection processing on the image data to detect specific regions in the image that differ from the subject area. For example, the specific region might be a sports field, a ball court, a goal, or a stage in an activity area. This embodiment assumes that the region detection unit 151 detects a moving region, but this disclosure is not limited to this. The region detection unit 151 can detect the moving region based on the image data or obtain positional information about the moving region from external sources. When the outline of the moving region is rectangular, the region can be specified based on positional information from four locations: upper left, lower left, upper right, and lower right. As an example other than a rectangular region, in the case of sports, the moving region (track) is elliptical, so the moving region can be specified based on the positional information of the outline.
[0040] Subject classification unit 152 generates a defined target region by correcting the motion region based on the region detection results of subject detection unit 150 and region detection unit 151, and classifies subjects existing within the defined target region and subjects existing outside the defined target region (region interior / exterior determination). Based on the region interior / exterior determination results of subject classification unit 152, image processing unit 141 can determine whether the subject is a contestant or a spectator.
[0041] Based on the region-internal / external determination result of the subject classification unit 152, the tracking control unit 153 works in conjunction with the focus control unit 124 to continuously focus on the subject of interest for tracking control. This embodiment assumes that the subject of interest is an athlete existing within the movement area, and that there are one or more subjects of interest.
[0042] The jitter detection unit 154 includes a gyroscope sensor, an accelerometer, and an electromagnetic compass, and detects the jitter of the camera device 100. The jitter detection unit 154 detects the amount of jitter of the camera device 100 in three mutually perpendicular axes, and also detects the amount of change in the position and orientation of the camera device 100.
[0043] By using volatile memory 143 as a circular buffer, the main control unit 140 can buffer multiple frames of image data captured within a predetermined time period, as well as data such as the detection results of the subject detection unit 150 based on the image data, the region detection results of the region detection unit 151, the region interior / exterior determination results of the subject classification unit 152, and the detection results of the shake detection unit 154.
[0044] Display unit 145 displays images being captured (live view) or still images taken, moving images being recorded, subjects detected in the displayed images and subjects of interest, GUIs for interactive operation, etc. Display unit 145 is a display device such as a liquid crystal display or an organic EL display. Display unit 145 can be integrated with camera device 100, or it can be an external device connected to camera device 100.
[0045] The operation unit 146 includes an operation component comprising a switch, a button, a ring, and an operation lever for receiving user operation, and outputs operation signals corresponding to the operation components operated by the user to the main control unit 140. The main control unit 140 performs control by outputting control signals to each component of the imaging device 100, including the lens unit 101, based on the operation signals. The operation component includes, for example, a touch panel integrated with the display unit 145. The user, acting as the photographer, can perform various operations on the imaging device 100 by operating the operation unit 146. The photographer can also make various settings in the imaging device 100 by using the operation unit 146 to operate the graphical user interface (GUI) displayed on the display unit 145.
[0046] The operation unit 146 includes at least a still image capture button, a moving image capture button, a mode dial, and a power switch. The still image capture button is an operation component used to instruct the main control unit 140 to perform still image capture processing. The moving image capture button is an operation component used to instruct the main control unit 140 to perform moving image capture processing. The mode dial is an operation component used to switch the operating mode of the camera device 100. The mode dial can be used to switch the operating mode of the camera device 100 to any one of the still image capture mode, moving image capture mode, and playback mode. The power switch is an operation component used to switch the power on / off of the camera device 100.
[0047] Under the control of the main control unit 140, the power control unit 148 controls the supply of power from the battery 149 to each component of the camera device 100 according to the state of the camera device 100. The battery 149 is a secondary battery capable of supplying power to operate the camera device 100.
[0048] When the still image capture button is half-pressed in still image capture mode, the main control unit 140 initiates automatic exposure (AE) control and AF control. When the still image capture button is fully pressed, the main control unit 140 performs still image capture processing to record the image data captured by the imaging unit 131 onto the recording medium 147.
[0049] When the main control unit 140 presses the motion image capture button for the first time in motion image capture mode, it performs AE control and AF control on the image data (frames) captured by the imaging unit 131, continues motion image capture processing to record motion images for a predetermined time in the recording medium 147, and stops motion image capture processing when the motion image capture button is pressed again.
[0050] The volatile memory 143 is, for example, DRAM, and is used as a buffer memory for temporarily storing image data captured by the imaging unit 131, an image display memory for the display unit 145, the working area of the main control unit 140, etc.
[0051] The non-volatile memory 144 is, for example, a flash ROM, and stores control programs executed by the main control unit 140. When the camera device 100 is powered on and activated by user operation, the control program stored in the non-volatile memory 144 is read (loaded) into a portion of the volatile memory 143. The main control unit 140 controls the operation of the camera device 100 according to the control program loaded into the volatile memory 143.
[0052] The main control unit 140 performs arithmetic processing for controlling the imaging device 100, including the lens unit 101. The main control unit 140 includes a hardware processor, such as a CPU or MPU, which controls the various components of the imaging device 100. The main control unit 140 controls the various components of the imaging device 100 by loading a program stored in non-volatile memory 144 into volatile memory 143 and executing the program, thereby enabling the functionality of the imaging device 100. Note that the entire imaging device 100 can be controlled by having multiple hardware components (e.g., multiple processors or circuits) share processing, rather than by the main control unit 140.
[0053] The main control unit 140 performs AF control based on the focus detection result of the phase difference detection method or the TV-AF method to control the focus control unit 124 to drive the focusing lens 109.
[0054] Furthermore, the main control unit 140 performs automatic exposure (AE) processing, which automatically determines exposure conditions (shutter speed or cumulative time, f-number, and sensitivity) based on the brightness information of the subject. For example, the image processing unit 141 can obtain the brightness information of the subject. The main control unit 140 can determine the exposure conditions by referring to a predetermined area such as a human face.
[0055] The various components of the camera device 100 are connected to exchange data via bus 160 and are controlled by the main control unit 140.
[0056] The subject detection unit 150 performs subject detection processing using inference processing through machine learning, such as deep learning. The learning model used for machine learning is formed by a neural network, which in this embodiment is formed by a convolutional neural network (CNN). Note that the inference model according to this embodiment is not limited to CNN, and can also be formed by a neural network such as a transformer. Rule-based methods other than machine learning can be used for subject detection processing.
[0057] Inference processing in deep learning can be performed by a graphics processing unit (GPU) or a digital signal processor (DSP). A GPU or DSP is a processor capable of performing extremely large product sum operations, bias addition operations, and nonlinear processing, and has arithmetic processing capabilities for performing matrix operations such as those on neural networks in a short time. Note that in inference processing, the CPU of the main control unit 140 and the GPU or DSP of the subject detection unit 150 can cooperate to perform arithmetic processing, or one of the CPU of the main control unit 140 and the GPU or DSP of the subject detection unit 150 can perform arithmetic processing.
[0058] The subject detection unit 150 detects the coordinates of the rectangular region surrounding the subject detected from the image data, using these coordinates as subject position and size information. Furthermore, based on the subject position and size information, the subject detection unit 150 calculates a reliability (probability value) for each subject, representing the likelihood of it being a subject of interest. The reliability is represented by an integer value from 0 to 255; a higher reliability value indicates a lower probability of detection error.
[0059] Even when the region detection unit 151 detects a moving region, similar to the subject detection unit 150, region detection processing is performed using machine learning inference. Similar to the subject detection unit 150, rule-based methods other than machine learning can be used for region detection processing. The region detection unit 151 can output a rectangular region including the field, and can output the outline of the field if the field is elliptical.
[0060] Note that the function (functional unit) of each component of the camera device 100 in this embodiment is determined by... Figure 1 The hardware shown and / or the hardware provided as Figure 1 The operation of each functional unit shown is implemented by a software program executed by the control unit. Furthermore, in Figure 1 In the case where each functional unit shown is formed in hardware rather than implemented in software, it provides the same functionality as... Figure 1 The circuit configuration corresponding to each functional unit shown is illustrated.
[0061] <Control Processing of the First Embodiment>
[0062] The following will refer to Figure 2 The control processing of the camera device 100 according to the first embodiment is described.
[0063] When the main control unit 140 controls the system by using the learning model and executing the program stored in the non-volatile memory 144... Figure 1 When implementing the components shown, Figure 2 The processing shown.
[0064] In step S201, the imaging control unit 133 controls the imaging unit 131 to capture an image, and causes the captured image signal processing unit 132 to process the image signal obtained by the imaging unit 131, thereby obtaining image data.
[0065] In step S202, the region detection unit 151 detects moving regions from the image data obtained in step S201.
[0066] In step S203, the subject classification unit 152 corrects the motion region detected in step S202 to generate a defined target region. (See below for further details.) Figure 3A and Figure 3B Describe the processing details in step S203.
[0067] In step S204, the subject detection unit 150 detects the subject region from the image data obtained in step S201.
[0068] In step S205, the subject classification unit 152, based on the target region generated in step S203 and the subject region detected in step S204, determines the interior / exterior of the region and classifies subjects existing within the motion region and subjects existing outside the motion region. (See below for further details.) Figure 3A and Figure 3B The details of the method for determining the interior / exterior of the region in step S205 are described. The result of determining the interior / exterior of the region in step S205, together with the attribute information of the subject (such as the position of the subject), is stored in volatile memory 143.
[0069] In step S206, based on the region interior / exterior determination result in step S205, the focus control unit 124 performs AF control by setting the subject present in the motion region as the subject of interest. Note that this disclosure is not limited to AF control, and any processing such as tracking control or subject attribute determination processing can be performed, as long as the region interior / exterior determination result can be used to perform the processing.
[0070] <Target Area Determination and Generation Processing>
[0071] Next, we will refer to Figure 3A and Figure 3B describe Figure 2 The process of determining the target region in step S203.
[0072] Figure 3A It is a view illustrating the arrangement of athletes, spectators, and venues in a sporting event.
[0073] Reference Figure 3AAssume the runners are the competitors and the standing people are the spectators. Personnel 301 and 302 are competitors, and personnel 303 and 304 are spectators. For ease of description, the reference numerals for other spectators are omitted except for personnel 303 and 304. Rectangles 311 to 314 are subject detection boxes, corresponding to the areas of personnel 301 to 304, which are detected as subjects from the image. Area 320 is the motion area (field), and area 321 is a corrected area calculated based on the height that the competitor can move (e.g., jump). For example, when shooting from almost the same line of sight as the competitor, the jump width in the image can be calculated using the formula hf / (zΔ) [pixels], where h [mm] represents the actual jump height that the competitor can jump, z [mm] represents the distance from the camera device 100 to the edge of the field, f [mm] represents the focal length of the camera device 100, and Δ [mm] represents the pixel pitch of the image sensor. This embodiment assumes the shooting position is outside the field and uses known field dimensions instead of the distance to the field. If the distance to the court can be obtained by other methods, this value can be used. The width 322 of zone 321 is set by multiplying this value by a factor of 1 or greater (e.g., 1.5). Considering that the width 322 is a value calculated from the edge of the court, it serves as a margin for adjusting for differences in jump height among athletes.
[0074] Given the known tilt angle information of the camera device 100, for example, if the camera device 100 is fixed, the width 322 can be calculated based on z[mm] and h[mm] by projecting the height that the athlete can jump onto the image and using the tilt angle information.
[0075] Subject classification unit 152 sets the region obtained by combining motion region 320 and correction region 321 as the target region. Then, subject classification unit 152 determines the region's interior / exterior based on the degree of overlap between a portion of each of subject detection frames 311 to 314 and the target region. In this embodiment, if a portion of each of subject detection frames 311 to 314 overlaps with the target region—for example, if the midpoint of the bottom side of each of subject detection frames 311 to 314 falls within the target region—then subject classification unit 152 determines that the subject corresponding to that subject detection frame exists within the target region. Alternatively, if a portion of each of subject detection frames 311 to 314 does not overlap with the target region—for example, if the midpoint of the bottom side of each of subject detection frames 311 to 314 falls outside the target region—then subject classification unit 152 determines that the subject corresponding to that subject detection frame exists outside the target region.
[0076] exist Figure 3A and Figure 3BIn the example shown, the midpoints of the bottom sides of the subject detection frames 311 and 312 for athletes 301 and 302 fall within the defined target area, but the subject detection frames 313 and 314 for spectators 303 and 304 fall outside the defined target area. By setting the correction area 321, it can be determined that athletes 301 and 302 exist within the defined target area even when they are jumping. By using the bottom side of the subject detection frames, it can be determined that athletes exist within the defined target area even when their bodies are flipping up and down, for example, in gymnastics.
[0077] Note that instead of using the midpoint of the bottom side of the subject detection frame, the ratio of the bottom length included in the target region to the bottom length, or the overlap ratio between the area of the lower part of the subject detection frame and the target region, can be used to determine whether the region is inside or outside. For example, the lower 1 / 3 of the region can be used as the lower part of the subject detection frame. If the overlap ratio is used, the motion region can be divided not only into two categories (inside and outside) but also into three categories. For example, if the overlap ratio is below the first threshold Th_1, the subject can be determined to exist outside the motion region. If the overlap ratio falls within the range of the first threshold Th_1 (including the endpoints) to the second threshold Th_2 (> Th_1) (excluding the endpoints), it can be determined to be unknown. If the overlap ratio is equal to or higher than the second threshold Th_2, the subject can be determined to exist within the motion region. The overlap ratio can also be assigned a reliability value.
[0078] exist Figure 3A and Figure 3B In the example shown, it is assumed that the motion area 320 is at the same height as the ground. However, as... Figure 3B The diagram shows that areas with height (such as goalposts) can be set as correction areas for the movement area. Figure 3B This is a view used to illustrate an example of setting a correction zone with height (such as a goal) for a movement area. (See reference...) Figure 3B For example, in basketball, where a player is highly likely to jump near the goal and shoot on the court, the accuracy of the determination can be improved by changing the determination conditions inside / outside the area between the area including the goal and the area excluding the goal. Figure 3B In the example shown, region 330 is the goal region, which includes the goal belonging to movement region 320. In goal region 330, if the midpoint of the top side of the subject detection frame 311 falls within goal region 330, the subject classification unit 152 can determine that the subject corresponding to the subject detection frame 311 exists within goal region 330; if the midpoint of the top side of the subject detection frame 311 falls outside goal region 330, the subject classification unit 152 can determine that the subject corresponding to the subject detection frame 311 exists outside goal region 330.
[0079] As described above, according to the first embodiment, by determining whether a subject exists within a defined target region based on the degree of overlap between a portion of the subject detection box and the defined target region, the subject of interest in the image can be easily identified, thereby improving the processing accuracy of the subject of interest.
[0080] Note that this embodiment has explained the example of a human subject. However, this embodiment is also applicable when the subject is similar to a vehicle on a racetrack, or when the subject is similar to an animal in a racetrack.
[0081] [Second Embodiment]
[0082] The second embodiment will now be described.
[0083] The first embodiment illustrates an example of determining the presence of a subject within a defined target region based on the degree of overlap between a portion of the subject detection frame and the defined target region. In contrast, the second embodiment uses time-series information of the subject detection frame to determine whether a region is inside or outside.
[0084] This embodiment describes an example of using the time-series information of the subject detection box to determine the interior / exterior of each frame of image data, but this disclosure is not limited thereto and any time information can be used.
[0085] The following describes an example of the presence of one or more subjects of interest. The midpoint coordinates of the bottom side of the i-th subject detection box in frame n are represented by (xi(n), yi(n)), and region-inside / outside determination is performed on all frames for tracking control of the subjects of interest. Note that instead of all frames, region-inside / outside determination can be performed on each predetermined frame, allowing for matching whether the subjects of interest are the same, and enabling tracking control of the subjects of interest.
[0086] Now refer to Figure 4 Describe a method for determining the interior / exterior of a region for each frame of image data.
[0087] exist Figure 4In the example shown, the row indicates the ID "i" of the tracked subject, and the column indicates the frame number. Furthermore, the terms "inside" and "outside" described in each unit indicate whether the i-th subject detection box (xi(k), yi(k)) exists within or outside the defined target region in frame k. When the frame of interest is frame n, the subject classification unit 152 counts the number of times "inside" m in a predetermined number M frames prior to frame n for each subject in the image of frame n, and if m / M exceeds a predetermined threshold Th_j, it determines that the subject exists within the defined target region. In this case, m / M can be considered as the reliability of subject i existing within the defined target region.
[0088] Furthermore, when the region detection unit 151 detects moving regions based on image data and updates them at predetermined time intervals, the detection accuracy may deteriorate due to state changes. In this case, the results of determining the interior / exterior of the region of each subject up to the (n-1)th frame can be used to provide reliability for each detected region in the moving region.
[0089] The following will refer to Figure 5 This describes a method for calculating detection reliability when the region detection unit 151 detects a moving region based on image data.
[0090] exist Figure 5 In the diagram, region 501 indicates the motion region detected by region detection unit 151, and region 502 indicates the region including the audience. Figure 3A and Figure 3B The common remainder is indicated by the same reference numerals, and its description is omitted.
[0091] The numerical value next to each subject indicates the probability that the subject exists within a defined target area. In this embodiment, for ease of description, all probability values for the audience in area 502 are 20%, but this value is typically different for each subject. The area detection unit 151 sets the reliability of the detected motion area 501 to the probability value of each subject detection frame. This can suppress the degradation of the detection accuracy of the motion area 501.
[0092] Furthermore, if the camera device 100 is fixed, the movement speed of the subject detection box can be used. For example, if the difference between the midpoint coordinate yi(n) of the bottom side of the i-th subject detection box in the n-th frame and the midpoint coordinate yi(n-1) of the bottom side of the i-th subject detection box in the (n-1)-th frame is equal to or greater than a threshold, the subject is excluded from the targets determined inside / outside the region. This can exclude subjects that are highly likely to jump from the determined targets and set the determined target region to be substantially equal to the motion region 320.
[0093] As described above, according to the second embodiment, by using the time-series information of the subject detection box to determine the inside / outside of the region, the recognition accuracy of the subject of interest in the image can be improved.
[0094] [Third Embodiment]
[0095] The third embodiment will now be described.
[0096] The third embodiment will describe an example of performing motion region detection processing based on the type of motion or event.
[0097] Figure 6 This is an example Figure 2 The flowchart of the motion region detection process in the third embodiment of step S202.
[0098] Note that the device configuration according to the third embodiment is different from that according to the first embodiment. Figure 1 The apparatus shown has the same configuration, and descriptions of other components that are the same as in the first embodiment will be omitted. The following will mainly describe the parts that differ from the first embodiment.
[0099] In step S601, the region detection unit 151 determines the type of sport, such as soccer or basketball, based on the image data. As a method for determining the sport type, for example, information about the sport specified by the user through a GUI can be obtained, or the sport type can be automatically determined using a dictionary learned through machine learning. Note that this disclosure is not limited to sport; it also applies to the type of event. An example of automatically determining the sport type using a region detection dictionary will be described below.
[0100] Note that the dictionary learned by machine learning is obtained by grouping words and phrases that have commonalities.
[0101] In step S602, the area detection unit 151 reads from the volatile memory 143 an area detection dictionary specifically for each sport, such as football or basketball, for the field or goal.
[0102] In step S603, the region detection unit 151 uses the region detection dictionary obtained in step S602 to detect the moving region and proceeds to... Figure 2 The processing in step S203.
[0103] Through the above processing, a target area can be defined with higher precision.
[0104] Note that a learned dictionary can be used to enable the simultaneous execution of steps S601 and S603.
[0105] The subject classification unit 152 can modify the determination conditions for the interior / exterior of the region based on the motion determination results in step S601. For example, in cases where it is unlikely that the viewer is moving in front, a method that is uncertain whether the subject has moved out from below the determined target region can be considered. Therefore, in the case of erroneously detecting the determined target region, the determination error of the interior / exterior of the region can be reduced.
[0106] As described above, according to the third embodiment, a target area with higher precision can be defined.
[0107] [Fourth Embodiment]
[0108] The fourth embodiment will now be described.
[0109] The fourth embodiment will describe an example in which, in a system in which the image processing device 700 and the camera device 750 are communicatively connected to each other, the camera device 750 is used to automatically capture scenes in a sports field or sports venue, event, etc., and the image processing device 700 determines the interior / exterior of the region in the image obtained from the camera device 750 and automatically edits and / or assigns the image.
[0110] The image processing apparatus according to this embodiment is applied to smartphones, tablets, desktop computers, etc., that can communicate with a camera device.
[0111] Figure 7 This is a block diagram illustrating the configuration of the image processing apparatus 700 according to the fourth embodiment.
[0112] In the fourth embodiment, multiple camera devices 750 are connected to the image processing device 700, and the image processing device 700 obtains image data from each camera device 750 and performs region interior / exterior determination. Note that in the system according to this embodiment, as long as the positional relationship between the camera devices 750 is known, the region interior / exterior determination results can be shared among the camera devices 750.
[0113] System storage unit 701 is, for example, a flash ROM, and stores programs, operating constants, etc., for each functional unit of system control unit 710. System memory 720 is a volatile memory such as DRAM, which is loaded with constants and variables for the operation of system control unit 710, data read from system storage unit 701, etc. Image storage unit 702 is, for example, a flash ROM, and stores image data acquired from imaging device 750. Figure 7 An example configuration is shown in which two camera devices 750 are connected, but one, three or more camera devices 750 can be connected.
[0114] The system control unit 710 performs arithmetic processing for controlling the image processing apparatus 700. The system control unit 710 includes a hardware processor, such as a CPU or MPU, that controls the various components of the image processing apparatus 700. The system control unit 710 controls the various components of the image processing apparatus 700 by loading a program stored in the system storage unit 701 into the system memory 720 and executing the program, thereby realizing the function of the image processing apparatus 700. Note that the entire image processing apparatus 700 can be controlled by having multiple hardware components (e.g., multiple processors or circuits) share processing, rather than by the system control unit 710.
[0115] The system control unit 710 includes a subject detection unit 703, a region detection unit 704, a subject classification unit 705, and an image editing unit 706. The functions of the subject detection unit 703, the region detection unit 704, and the subject classification unit 705 are similar to... Figure 1 The subject detection unit 150, region detection unit 151, and subject classification unit 152 in the embodiment have the same function. In this embodiment, the region interior / exterior determination result of the subject classification unit 705 is fed back to each camera device 750, and each camera device 750 can perform AF control or tracking control on the subject (subject of interest) existing in the motion area based on the region interior / exterior determination result.
[0116] Image editing unit 706 performs at least one of image data editing and allocation based on the region interior / exterior determination result of subject classification unit 705. Image editing unit 706 edits image data based on the region interior / exterior determination result of subject classification unit 705 and stores the edited image data in system storage unit 701. Image editing includes automatically extracting cropped motion images by focusing on a particular athlete or exciting scene from a play. Image editing unit 706 uploads the edited image to server device 730 providing cloud services, etc., via network 740.
[0117] As described above, according to the fourth embodiment, at least one of automatic image editing and automatic allocation can be performed based on the result of determining the interior / exterior region of the image captured by the camera device 750.
[0118] Note that the function (functional unit) of each component of the image processing apparatus 700 in this embodiment is determined by... Figure 7 The hardware shown and / or the hardware provided are as follows: Figure 7 The operation of each functional unit shown is implemented by a software program executed by the control unit. Furthermore, in Figure 7 In the case where each functional unit shown is formed in hardware rather than implemented in software, it provides the same functionality as... Figure 7The circuit configuration corresponding to each functional unit shown is illustrated.
[0119] Other embodiments
[0120] The embodiments of this disclosure can also be implemented by providing software (including computer program products of computer programs) that performs the functions of the above embodiments to a system or device via a network or various storage media, and the computer (central processing unit (CPU) or microprocessor unit (MPU) of the system or device) reads and executes the computer program.
[0121] While this disclosure has been described with reference to exemplary embodiments, it is to be understood that this disclosure is not limited to the disclosed exemplary embodiments. The scope of the appended claims is to be given the broadest interpretation in order to cover all such modifications and equivalent structures and functions.
Claims
1. An image processing apparatus, the apparatus comprising: An acquisition unit configured to acquire image data; A subject detection unit, configured to detect a subject region in image data; A region detection unit is configured to detect a first region in the image data that is different from the subject region; as well as A classification unit is configured to classify subjects present within the first region and subjects present outside the first region based on the overlap between a first portion of the subject region and the first region.
2. The apparatus according to claim 1, wherein When the first part of the subject region overlaps with the first region, the classification unit determines that the subject exists within the first region, and If the first part of the subject region does not overlap with the first region, the classification unit determines that the subject exists outside the first region.
3. The apparatus according to claim 1, wherein the first portion is the bottom side of the subject area.
4. The apparatus of claim 1, wherein the first region comprises a region set based on the height at which the subject can move.
5. The apparatus according to claim 1, wherein The region detection unit detects the second region that belongs to the first region, and The classification unit classifies subjects that exist within the second region and subjects that exist outside the second region based on the overlap between the second part of the subject region and the second region.
6. The apparatus of claim 5, wherein the second portion is the top side of the subject region.
7. The apparatus according to claim 1, wherein The first zone includes the playing field, the court and goalposts, and one of the stages in the playing field. The subject is one of the people participating in the sport or appearing at the event venue.
8. The apparatus of claim 1, wherein the classification unit classifies, for each frame of image data, subjects present in the first region and subjects present outside the first region.
9. The apparatus of claim 1, wherein the classification unit classifies subjects present in the first region and subjects present outside the first region based on the speed at which the subjects move.
10. The apparatus of claim 8, wherein the region detection unit obtains the reliability of the result of classifying the subject for each frame of image data.
11. The apparatus of claim 7, wherein the region detection unit detects the first region based on image data and a learned dictionary.
12. The apparatus of claim 11, further comprising a second obtaining unit configured to obtain the type of the sport or the activity venue.
13. The apparatus of claim 12, wherein the region detection unit switches the dictionary based on the type obtained by the second acquisition unit.
14. The apparatus of claim 12, wherein the classification unit switches the conditions for classifying the subject based on the type obtained by the second obtaining unit.
15. The apparatus according to any one of claims 1 to 14, further comprising: An editing unit is configured to edit image data based on the results of classification by a classification unit; as well as At least one of a storage unit and an allocation unit, wherein the storage unit is configured to store an image edited by the editing unit, and the allocation unit is configured to allocate an image edited by the editing unit.
16. The apparatus according to any one of claims 1 to 14, further comprising: An imaging unit configured to generate image data by capturing images; as well as The control unit is configured to perform either focus control or tracking control on a subject present in the first region.
17. An image processing method executed by an image processing apparatus, comprising the following steps: Obtain image data; Detect the subject region in image data; Detect the first region in the image data that differs from the subject region; as well as Subjects located within the first region and subjects located outside the first region are classified based on the overlap between the first part of the subject region and the first region.
18. A computer-readable storage medium storing a program for enabling a computer to function as an image processing apparatus according to any one of claims 1 to 16.
19. A computer program product comprising a program for enabling a computer to function as an image processing apparatus according to any one of claims 1 to 16.