Image processing apparatus, image capturing apparatus, control method, and storage medium

The image processing device uses multidimensional time series data to map and infer subject roles, addressing misidentification issues in scenes with non-distinctive postures by accurately determining the main subject.

JP2026012292APending Publication Date: 2026-01-23CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025181977
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing image processing systems struggle to accurately identify the main subject in scenes where the posture of subjects is not distinctive, leading to potential misidentification.

Method used

An image processing device that acquires multiple captured images, detects subject roles using multidimensional time series data, maps this data into a multidimensional space, and infers the subject's role based on positional relationships within this space.

Benefits of technology

Enables accurate identification of subject roles in complex scenes, ensuring the main subject is correctly identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012292000001_ABST
    Figure 2026012292000001_ABST
Patent Text Reader

Abstract

To select a main subject suitable for a scene.SOLUTION: A first acquisition unit configured to acquire a plurality of captured images obtained by intermittent imaging, a second acquisition unit configured to acquire a role of each of at least some subjects included in the plurality of captured images acquired by the first acquisition unit, a generation unit configured to detect a state of the subject for each of the plurality of captured images and generate multidimensional time-series data indicating a plurality of types of changes in the state of each subject, and a mapping unit configured to map the multidimensional time-series data of each subject generated by the generation unit in a multidimensional space in association with the role of the subject; The information processing apparatus includes a configuration unit configured to configure a plurality of subspaces formed by multidimensional time-series data in the space, an input unit configured to receive multidimensional time-series data of an inference target object, and an inference unit configured to determine and output a role of the inference target object based on a positional relationship with the plurality of subspaces when the multidimensional time-series data of the inference target object is mapped to the multidimensional space.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device, an imaging device, a control method, and a program, and more particularly to an image processing technique for recognizing a subject. [Background technology]

[0002] When there are multiple subjects within an imaging range, there are imaging devices that identify one of the subjects as a main subject and control focus and tracking. Patent Document 1 discloses a method of acquiring posture information of the multiple subjects and setting the subject with the most different posture as the main subject based on the distance between feature vectors extracted from the posture information. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-105850 Summary of the Invention [Problem to be solved by the invention]

[0004] However, since the method of Patent Document 1 detects subjects with unusual postures, there is a possibility that a subject that should be the main subject cannot be identified in a scene where the posture of the subject is not distinctive.

[0005] The present invention has been made in consideration of the above-mentioned problems, and has as its object to provide an image processing device, an imaging device, a control method, and a program that suitably identify the role of a subject in a scene. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, the image processing device of the present invention is characterized by having: a first acquisition means for acquiring a plurality of captured images obtained by intermittent imaging; a second acquisition means for acquiring a role for each of at least some of the subjects included in the plurality of captured images acquired by the first acquisition means; a generation means for detecting the state of the subject in each of the plurality of captured images and generating multidimensional time series data showing multiple types of changes in the state of each subject; a construction means for mapping the multidimensional time series data of each subject generated by the generation means into a multidimensional space while associating it with the role of the subject, and constructing a plurality of subspaces formed by the multidimensional time series data within the space; an input means for accepting the multidimensional time series data of the subject to be inferred; and an inference means for determining and outputting the role of the subject to be inferred based on the positional relationship with the plurality of subspaces when the multidimensional time series data of the subject to be inferred is mapped into the multidimensional space. [Effects of the Invention]

[0007] With this configuration, the present invention makes it possible to suitably identify the role of the subject in the scene. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a digital camera system according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating the structure of an image sensor 22 according to an embodiment of the present invention; [Figure 3] FIG. 1 is a diagram illustrating detection of the head orientation according to an embodiment of the present invention. [Figure 4] FIG. 1 is a diagram illustrating a trajectory of head movement according to an embodiment of the present invention. [Figure 5] FIG. 1 is a diagram illustrating the direction of a person according to an embodiment of the present invention; [Figure 6] FIG. 1 is a diagram illustrating attention levels according to an embodiment of the present invention; [Figure 7] FIG. 1 is a diagram showing an example of time-series data according to an embodiment of the present invention; [Figure 8]FIG. 10 is another diagram showing an example of time-series data according to an embodiment of the present invention. [Figure 9] FIG. 10 is yet another diagram illustrating an example of time-series data according to an embodiment of the present invention. [Figure 10] FIG. 1 is a diagram illustrating a partial space according to an embodiment of the present invention; [Figure 11] FIG. 1 is a block diagram illustrating a functional configuration involved in main subject selection according to an embodiment of the present invention. [Figure 12] 1 is a flowchart illustrating a selection process executed by a system control unit 50 according to an embodiment of the present invention. [Figure 13] FIG. 1 is a diagram illustrating a line of sight direction according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0009] [Embodiment] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] In the embodiment described below, the present invention is applied to a digital camera system (digital camera with interchangeable lenses) as an example of an imaging device, which includes an image processing device that identifies the role of a subject in a scene and selects a main subject based on the identification result. However, the present invention can be applied to any device that can identify the role of a subject based on a captured image. Such devices include video cameras, computer devices (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, drones, drive recorders, etc. These are merely examples, and the present invention can also be applied to other electronic devices.

[0011] <<Digital camera system configuration>> 1 is a block diagram showing an example of the configuration of a digital camera system according to this embodiment. The digital camera system includes a body 100 of a digital camera with an interchangeable lens and a lens unit 150 that is detachable from the body 100. Note that this embodiment will be described assuming that the imaging device is of an interchangeable lens type, but this is not an essential element for implementing the present invention.

[0012] The lens unit 150 has a communication terminal 6 that comes into contact with a communication terminal 10 provided on the main body 100 when the lens unit 150 is attached to the main body 100. Power is supplied from the main body 100 to the lens unit 150 through the communication terminal 10 and the communication terminal 6. Furthermore, the lens system control circuit 4 of the lens unit 150 and the system control unit 50 of the main body 100 can communicate bidirectionally through the communication terminal 10 and the communication terminal 6.

[0013] In the lens unit 150, the lens group 103 is an imaging optical system composed of multiple lenses, including a movable lens. The movable lenses include at least a focus lens. Depending on the lens unit 150, one or more lenses with different characteristics, such as a variable magnification lens or a blur correction lens, may also be included. The AF drive circuit 3 includes a motor, actuator, etc. that drive the focus lens. The focus lens is driven by the lens system control circuit 4 controlling the AF drive circuit 3. The iris drive circuit 2 includes a motor actuator, etc. that drive the iris 102. The opening amount of the iris 102 is adjusted by the lens system control circuit 4 controlling the iris drive circuit 2.

[0014] The mechanical shutter 101 is driven by the system control unit 50 to adjust the exposure time of the image sensor 22. The mechanical shutter 101 is kept fully open during video shooting.

[0015] The image sensor 22 is, for example, a CCD image sensor or a CMOS image sensor. The image sensor 22 has a plurality of pixels arranged two-dimensionally, and each pixel has one microlens, one color filter, and one or more photoelectric conversion units. In this embodiment, each pixel has multiple photoelectric conversion units, and each pixel is configured to be able to read out a signal for each photoelectric conversion unit. By configuring the pixels in this way, it is possible to generate a captured image, a parallax image pair, and an image signal for phase-difference AF from signals read out from the image sensor 22.

[0016] FIG. 2(a) is a diagram schematically showing the correspondence between the exit pupil of the lens unit 150 and each photoelectric conversion unit when a pixel of the image sensor 22 has two photoelectric conversion units.

[0017] Two photoelectric conversion units 201a and 201b provided in a pixel share one color filter 252 and one microlens 251. Light that has passed through a partial region 253a of the exit pupil (region 253) is incident on the photoelectric conversion unit 201a, and light that has passed through a partial region 253b of the exit pupil is incident on the photoelectric conversion unit 201b.

[0018] Therefore, for a pixel included in an arbitrary pixel region, an image formed by a signal read from the photoelectric conversion unit 201a and an image formed by a signal read from the photoelectric conversion unit 201b form a parallax image pair. The parallax image pair can be used as image signals (image A signal and image B signal) for phase-difference AF. Furthermore, a normal image signal (captured image) can be obtained by adding the signal read from the photoelectric conversion unit 201a and the signal read from the photoelectric conversion unit 201b for each pixel.

[0019] In this embodiment, each pixel of the image sensor 22 is configured to function both as a pixel for generating a signal for phase difference AF (focus detection pixel) and as a pixel for generating a normal image signal (image capturing pixel). However, it is not necessary for all pixels of the image sensor 22 to be configured to be able to realize both of these functions; some pixels of the image sensor 22 may be configured as focus detection pixels and the remaining pixels as image capturing pixels.

[0020] Furthermore, the focus detection pixels do not necessarily have to be configured with multiple photoelectric conversion units in one pixel as shown in FIG. 2(a). They can also be configured with one photoelectric conversion unit in one pixel as shown in FIG. 2(b). FIG. 2(b) shows an example of the correspondence between the focus detection pixels and a partial region 253c of the exit pupil 253 through which incident light passes. In the focus detection pixels shown in FIG. 2(b), the photoelectric conversion unit 201 functions similarly to the photoelectric conversion unit 201b in FIG. 2(a) due to the opening 254. Therefore, by distributing the focus detection pixels shown in FIG. 2(b) and other types of focus detection pixels that function similarly to the photoelectric conversion unit 201a in FIG. 2(a) throughout the image sensor 22, it becomes possible to set a focus detection area of ​​virtually any location and size.

[0021] 2(a) and 2(b) is a configuration in which an image sensor for obtaining an image to be recorded is used as a sensor for phase-difference AF, but the present invention can be implemented regardless of the AF method, such as other AF that can set a focus detection area of ​​any size and position. For example, the present invention can also be implemented with a configuration that uses contrast AF. When only contrast AF is used, each pixel has one photoelectric conversion unit.

[0022] 1 converts into a digital image signal (image data) an analog image signal output from the imaging element 22. The A / D converter 23 may be provided in the imaging element 22.

[0023] The image data (RAW image data) output by the A / D converter 23 is processed by the image processing unit 24 as necessary, and then stored in the memory 32 via the memory control unit 15. The memory 32 is a storage device that is used as a buffer memory for temporarily storing image data and audio data, and as a video memory for the display unit 28.

[0024] The image processing unit 24 applies predetermined image processing to the image data to generate signals and image data, and acquire and / or generate various types of information. The image processing unit 24 may be a dedicated hardware circuit such as an ASIC designed to realize a specific function, or may be configured to realize a specific function by a processor such as a DSP executing software.

[0025] The image processing applied by the image processing unit 24 includes preprocessing, color interpolation, correction, detection, data processing, and evaluation value derivation. Preprocessing includes signal amplification, reference level adjustment, and defective pixel correction. Color interpolation, also known as demosaicing, interpolates values ​​of color components not included in image data. Correction includes white balance adjustment, image brightness correction, correction of optical aberrations of the lens unit 150, and color correction. Detection includes detection and tracking of feature regions (e.g., facial regions and human body regions), person recognition, and the like. Data processing includes scaling, encoding and decoding, and header information generation. Evaluation value derivation includes derivation of a pair of image signals for phase-difference AF, evaluation values ​​for contrast AF, and evaluation values ​​used for automatic exposure control. Note that these are examples of image processing that the image processing unit 24 can perform and do not limit the image processing that the image processing unit 24 performs. Furthermore, some of the processing, such as the evaluation value derivation processing, may be executed by the system control unit 50, or may be executed by the image processing unit 24 and the system control unit 50 working together.

[0026] The D / A converter 19 generates an analog signal suitable for display on the display unit 28 from the image data for display stored in the memory 32, and supplies the generated analog signal to the display unit 28. The display unit 28 has, for example, a liquid crystal display device, and presents a display based on the analog signal from the D / A converter 19 on its display surface.

[0027] By continuously capturing video (capture control) and displaying the captured video (display control), display unit 28 can function as an electronic viewfinder (EVF). The video displayed to enable display unit 28 to function as an EVF is called a live view image. Display unit 28 may be provided inside main body 100 so that it can be viewed through an eyepiece, or may be provided on the surface of the housing of main body 100 so that it can be viewed without using an eyepiece. Display unit 28 may be provided both inside main body 100 and on the surface of the housing.

[0028] The system control unit 50 is, for example, a CPU (also called an MPU: microprocessor). The system control unit 50 reads out programs stored in the nonvolatile memory 56, expands them into the system memory 52, and executes them to control the operation of the main body 100 and the lens unit 150 and realize various functions of the digital camera system. The system control unit 50 also controls the operation of the lens unit 150 by sending various commands to the lens system control circuit 4 via communication terminals 10 and 6.

[0029] The nonvolatile memory 56 stores programs executed by the system control unit 50, various setting values ​​of the digital camera system, image data for a GUI (Graphical User Interface), etc. The system memory 52 is a main memory used when the system control unit 50 executes programs. The data (information) stored in the nonvolatile memory 56 may be rewritable.

[0030] As part of its operation, the system control unit 50 performs automatic exposure control (AE) processing based on evaluation values ​​derived by the image processing unit 24 or itself, and determines the shooting conditions (exposure conditions). Shooting conditions, for example, for still image shooting are shutter speed, aperture value, and sensitivity. The system control unit 50 determines one or more of the shutter speed, aperture value, and sensitivity according to the AE mode that is set. The system control unit 50 controls the aperture value (opening size) of the aperture mechanism of the lens unit 150. The system control unit 50 also controls the operation of the mechanical shutter 101.

[0031] In addition, the system control unit 50 drives the focus lens of the lens unit 150 based on the evaluation value or defocus amount derived by the image processing unit 24 or itself, and performs autofocus detection (AF) processing to focus the lens group 103 on a subject within the focus detection area.

[0032] The system timer 53 is an internal clock that is used by the system control unit 50 .

[0033] The operation unit 70 is a user interface provided in the main body 100, and has multiple input devices (buttons, switches, dials, etc.) that can be operated by the user. Some of the input devices provided in the operation unit 70 have names corresponding to the functions assigned to them. For convenience, the shutter button 61, mode selector switch 60, and power switch 72 are illustrated separately from the operation unit 70, but are included in the operation unit 70. If the display unit 28 is a touch display equipped with a touch panel, the touch panel is also included in the operation unit 70. Operations of the input devices included in the operation unit 70 are monitored by the system control unit 50. More specifically, when an operation input is made to an input device, the operation unit 70 outputs a corresponding control signal to the system control unit 50, and the system control unit 50 executes processing corresponding to the operation input made based on the control signal.

[0034] The shutter button 61 has a first shutter switch 62 that is turned on when half-pressed and outputs a signal SW1, and a second shutter switch 64 that is turned on when fully pressed and outputs a signal SW2. When the system controller 50 detects the signal SW1 (ON of the first shutter switch 62), it executes preparatory operations for still image shooting. The preparatory operations include AE ​​processing and AF processing. When the system controller 50 detects the signal SW2 (ON of the second shutter switch 64), it executes the still image shooting operation (image capturing and recording operations) according to the shooting conditions determined by the AE processing.

[0035] The power supply control unit 80 is composed of a battery detection circuit, a DC-DC converter, a switch circuit for switching between powered blocks, etc., and detects whether a battery is installed, the type of battery, and the remaining battery power. The power supply control unit 80 also controls the DC-DC converter based on the detection results and instructions from the system control unit 50, and supplies the required voltage to each unit, including the recording medium 200, for the required period of time.

[0036] The power supply unit 30 is composed of a battery, an AC adapter, etc. The I / F 18 is an interface used for inputting and outputting data to and from the recording medium 200. The recording medium 200 is a storage device such as a memory card or a hard disk that is detachably attached to the main body 100, and records data files such as captured images and audio. The data files recorded on the recording medium 200 can be read out via the I / F 18 and played back via the image processing unit 24 and the system control unit 50.

[0037] The communication unit 54 realizes communication with external devices by at least one of wireless communication and wired communication. Images captured by the image sensor 22 (captured images (including live view images)) and images recorded on the recording medium 200 can be transmitted to external devices via the communication unit 54. In addition, image data and various other information can be received from external devices via the communication unit 54.

[0038] The camera body orientation detection unit 55 detects the orientation of the camera body of the main body 100 relative to the direction of gravity. The camera body orientation detection unit 55 may be an acceleration sensor or an angular velocity sensor. The system control unit 50 can record orientation information corresponding to the camera body orientation detected by the camera body orientation detection unit 55 during shooting in a data file that stores image data obtained during that shooting. The orientation information can be used, for example, to display recorded images in the same orientation as when they were shot. In this specification, the orientation of the camera body (main body 100) and the orientation of the subject will be described separately, and the former will be referred to as the camera body orientation and the latter as the subject orientation.

[0039] Overview of main subject selection Next, a detailed description will be given of main subject selection realized by the main body 100 of this embodiment. In the following description, the main body 100 detects an image of a person (subject) in a series of captured images obtained for live view as a feature region, infers the role of the person based on time-series information indicating the state of the person, and selects the person to be the main subject based on the inference result.

[0040] The detection of characteristic regions may be realized, for example, by applying a known method to the live view image and identifying regions corresponding to predetermined characteristics. Characteristic regions relating to a person may include regions in which images of specific parts of a person appear, such as a face region, a head region, a torso region, and even eyeball regions corresponding to the eyes in the face region. In this embodiment, the detection of characteristic regions is performed by the image processing unit 24, and information such as the position, size, and reliability of each characteristic region in the captured image is transmitted to the system control unit 50 as the detection result.

[0041] In response to the detection of a characteristic region, the system control unit 50 can perform various controls to make the image of the characteristic region into an appropriate state. For example, the system control unit 50 may perform autofocus detection (AF) to focus on the characteristic region, or autoexposure control (AE) to properly expose the characteristic region. The system control unit 50 may also perform automatic white balance to properly adjust the white balance of the characteristic region, or automatic flash light intensity adjustment to properly adjust the brightness of the characteristic region, for example.

[0042] <Constructing an inference model> Here, we will explain the inference model constructed in the main body 100 of this embodiment to identify the role of a person. The inference model is constructed prior to the image capture in which the main subject is selected, and the necessary information is stored in the non-volatile memory 56 or the like. Therefore, the inference model is constructed as a result of learning based on video images (a series of captured images: multiple captured images) captured in advance. In the learning, the role of each person (subject) in the captured images is assigned as a label.

[0043] Below, an example will be described in which learning is performed based on video images of ball game scenes (ball game scenes) such as soccer and basketball, and an inference model is constructed.

[0044] In video footage of a ball game scene, the subjects within the imaging range (area of ​​interest) can be classified into limited roles (action types), such as a player holding the ball, a player on the team receiving a pass, or a player on the opposing team playing defense. Each of these players performs different actions, and differences in their characteristics can be seen in the changes in the person's state over time.

[0045] For example, a player in possession of the ball will look around in search of open space on the field (court) to get closer to the goal, and so will likely exhibit a tendency to change his or her head direction and gaze direction more frequently than other players. Furthermore, a player in possession of the ball must avoid the opposing team's defense, so his or her movement trajectory may exhibit a meandering pattern. Furthermore, a player in possession of the ball is the center of the game, so he or she will likely be the focus of attention from other players, i.e., the gaze direction of multiple other players may be concentrated on that player.

[0046] For example, a defensive player may exhibit the characteristic of frequently looking at the player in possession of the ball, and a defensive player may exhibit the characteristic of moving in the shortest distance toward the player in possession of the ball.

[0047] In the main body 100 of this embodiment, in order to extract and learn such features, the head of a person in a video is detected as a feature region, and the head orientation, the direction of the person (player), and the attention level indicating whether the person is being watched by other people are derived. Then, based on these derived parameters, time-series data (features) for learning are constructed. The detection of a person's head and the derivation of each parameter may be realized, for example, by incorporating a circuit configuration or software configuration of a detector for detecting the features in the image processing unit 24. The detector may perform a predetermined analysis process, or may be obtained by, for example, labeling a head image with a manually set head orientation and performing machine learning using a forward propagation neural network.

[0048] Here, the head direction is estimated based on the position of the nose (triangle in the figure) included in the person's head, as shown in Figure 3. Figures 3(a) and (b) show images of the same subject taken at different times, and the head direction transitions from direction 311 to direction 312 over time.

[0049] Furthermore, the direction of travel of a person is estimated based on the trajectory of head movement between frames of a moving image, as shown in Fig. 4, for example. Fig. 4 shows the movement of the image of the same subject in an image that appears when three captured images (three consecutive frame images) obtained by capturing an image of a fixed imaging range are superimposed, with images 401, 402, and 403 each representing the subject at a different time. As shown in the figure, the position of the subject changes over time in the order of image 401 → image 402 → image 403, so the direction of travel of the person can be estimated based on the trajectory 410 of head movement.

[0050] 4, for convenience, the direction of travel of a person is estimated based on the transition of the position at which the image of the head is captured in the captured image obtained by capturing an image of a fixed capturing range, but it should be readily understood that the invention is not limited to this. When capturing a ball game scene, the orientation of the main body 100 may change. Therefore, when framing occurs during the capture of a moving image, the framing angle can be derived from the change in the camera body orientation detected by the camera body orientation detection unit 55, and the amount of movement corresponding to the framing angle can be added to obtain a trajectory.

[0051] The direction of the head and the direction of travel of a person may be expressed as directions in a plane parallel to a field 501, as shown in Fig. 5, for example. Fig. 5 shows a schematic diagram of a bird's-eye view of the field 501, and eight directions are shown for a subject 502: rear, right rear, right, right front, front, left front, left, and left rear, but the direction of the head and the direction of travel may be directions not shown.

[0052] The attention level is also derived according to the number of other subjects that simultaneously capture the subject within their field of view, as shown in FIG. 6, for example. The field of view is estimated as a range of a certain angle (within a plane parallel to field 501) centered on the direction of the head. In FIG. 6, images 601 to 603 each correspond to a different subject, and field of view ranges 611 to 613 based on the direction of the head are shown. In the example shown, the subject in image 601 is included in field of view range 612 of the subject in image 602 and field of view range 613 of the subject in image 603, and the attention level is derived as 2. On the other hand, the subjects in images 602 and 603 are not included in the field of view ranges of other subjects, and therefore the attention level is derived as 0.

[0053] When these parameters are detected for each frame of the captured image included in the video to be learned, time-series data for learning, which are feature quantities, are generated. In this embodiment, the following three types of time-series data are used as feature quantities for learning.

[0054] The first feature is time-series data of changes in head direction relative to the direction of travel. This feature represents the difference between the subject's head direction and the direction of travel of the subject at each frame time in a time series. Figure 7(a) shows an example of time-series data of a player in possession of the ball, and Figure 7(b) shows an example of time-series data of a defensive player, distributed over time. As shown in Figure 7(a), the changes in the head direction of the player in possession of the ball relative to the direction of travel show the characteristic of the player looking back and forth and moving forward in the direction of travel. On the other hand, as shown in Figure 7(b), the changes in the head direction of the defensive player relative to the direction of travel show the characteristic of the head direction fluctuating less than that of the player in possession of the ball, and moving forward in the direction of travel. This corresponds to the behavior unique to defensive players, such as when a player tackles a player in possession of the ball while looking diagonally forward to the right.

[0055] The second feature is time-series data of changes in the subject's direction of travel. This feature represents the subject's movement between frames (for example, the difference in movement within a plane corresponding to the field 501) over time. FIG. 8(a) shows an example of five frames of overlapping time-series data for a player in possession of the ball, and FIG. 8(b) shows an example of five frames of overlapping time-series data for a defensive player. Each figure is a two-dimensional distribution with axes indicating the player's front and back directions and axes indicating the player's rightward and leftward directions, and the distance of each point from the intersection of the axes corresponds to the amount of movement in one frame. In other words, the closer a point is to the intersection of the axes, the less the player moves between frames, and the further away it is from the intersection, the more the player moves.

[0056] As shown in Figure 8(a), the change in the direction of the player in possession of the ball shows that the player not only moves forward but also changes direction to face right along the way. This corresponds to, for example, a situation where the player in possession of the ball temporarily moves to the right to avoid a defensive player. On the other hand, as shown in Figure 8(b), the change in the direction of the defensive player shows that the player continues to move forward.

[0057] The third feature is time-series data of changes in the attention level of a subject. This feature indicates the number of other subjects paying attention to the subject at each frame time in a time series. Figure 9(a) shows an example of the time-series data of a player in possession of the ball, and Figure 9(b) shows an example of the time-series data of a defensive player, distributed over time. As shown in Figure 9(a), the change in the attention level of a player in possession of the ball shows a characteristic in which the attention level fluctuates between 2 and 4 players, since the player is more likely to be noticed by other players. On the other hand, as shown in Figure 9(b), the change in the attention level of a defensive player shows a characteristic in which the attention level fluctuates between 0 and 1 player, since the defensive player is less likely to be noticed by other players than the player in possession of the ball.

[0058] Next, based on the three types of time-series data for learning obtained in this way, multidimensional time-series data, i.e., in the above example, three-dimensional vector data that combines "changes in head orientation relative to the direction of travel," "changes in subject travel direction," and "changes in attention level," is generated. In other words, the multidimensional time-series data used for learning is data sampled for each subject from a video with a length of a predetermined number of frames, and is data associated with roles associated with the subjects as labels.

[0059] Learning in this embodiment is performed by mapping sampled multidimensional time-series data (hereinafter simply referred to as time-series data) into a multidimensional space (having the same number of dimensions as the time-series data) and clustering the data. Here, clustering refers to the process of grouping time-series data exhibiting similar characteristics regardless of the role associated with the time-series data. Clustering makes it possible to group time-series data that are close to each other in the multidimensional space, and within each group, a subspace formed by time-series data exhibiting a common feature can be configured in the multidimensional space. Therefore, multiple subspaces are configured in the multidimensional space by time-series data sampled for subjects with various roles.

[0060] Fig. 10 is a diagram illustrating a plurality of subspaces configured in a multidimensional space. In Fig. 10, to facilitate understanding of the invention, the space is shown in two dimensions, and the explanation will be given assuming that time-series data is associated with "role 1 (player holding the ball)" or "role 2 (defensive player)." In the example shown, four types of subspaces 1001 to 1004 are configured in the space through learning, and the roles associated with the time-series data forming each subspace are shown in correspondence with the space.

[0061] As described above, time-series data forming one subspace have similar characteristics, but are not necessarily associated with the same role. For example, the action of running facing forward is performed by both a player in possession of the ball and a defensive player, so time-series data associated with roles 1 and 2 belong to subspace 1001 related to the characteristic of moving forward.

[0062] On the other hand, movements that attract a lot of attention from other players, such as frequent changes in head direction and gaze direction, are characteristics that are primarily characteristic of the behavior of players in possession of the ball. Therefore, only time-series data associated with role 1 belong to subspace 1002, which relates to such characteristics. Similarly, subspaces 1003 and 1004 are related to characteristics that appear in the behavior unique to defensive players, and only time-series data associated with role 2 belong to subspaces 1003 and 1004.

[0063] In the latter case, for a subspace formed by time-series data associated with one type of role, if time-series data showing similar characteristics is obtained for the subject to be inferred, the role of the subject can be immediately identified. In this embodiment, when such a subspace is formed as a result of learning, a representative value of the time-series data related to the subspace is stored so that the role related to the subspace can be inferred, and can be referenced when the inference model is used. The representative value of the time-series data related to the subspace is set as the average value of the time-series data forming the subspace.

[0064] Therefore, the inference model constructed as a result of learning includes information on representative values ​​of at least a subspace (definite space) that can uniquely identify the role of the subject related to the time-series data to which it belongs, among multiple subspaces configured in a multidimensional space. For example, in the embodiment shown in FIG. 10, subspaces 1002 to 1004 are definite spaces for role 1 or 2, and information on their representative values ​​is included in the inference model. In the following description, information on representative values ​​of definite spaces may be referred to as a "template."

[0065] According to this information, for example, when input time series data related to a subject to be inferred is mapped onto a multidimensional space, it becomes possible to determine to which definite space within the space the time series data related to the subject to be inferred belongs. Here, whether or not the time series data related to the subject to be inferred belongs to a definite space is determined based on the distance between the time series data and a representative value of each definite space (Euclidean distance in a multidimensional space). For example, when the time series data related to the subject to be inferred is closer to the representative value of one of the definite spaces by more than a predetermined threshold, it is determined that the time series data related to the definite space belongs to that definite space, and it is inferred that the subject to be inferred has a role related to that definite space. The threshold may be determined according to the size of a subspace, or may be fixed and set regardless of the subspace.

[0066] <Main subject selection> Based on the inference model constructed by such pre-learning, the main body 100 of this embodiment infers the role of each person in captured images (live view images) obtained by intermittent imaging, and selects a main subject based on the inference results. More specifically, multidimensional time-series data related to feature amounts is constructed for the head of each person extracted as a feature region in the live view images, and the inference model infers the role of each person based on the time-series data. Based on the inference results from the inference model, the people in the captured images are classified into classes according to their roles.

[0067] In this embodiment, a player holding the ball is selected as the main subject, and if there is a person classified into the class of "player holding the ball," system control unit 50 selects that person as the main subject. Note that the control performed by selecting the main subject is not limited to imaging control and image processing control such as the automatic AF control and automatic AE control described above, but may also include operation control of the automatic camera platform to which main body 100 is attached (subject tracking control).

[0068] <Functional configuration> 11 is a block diagram showing the functional configuration related to the above-described main subject selection, which is realized in main body 100 of this embodiment. Each of the functional configurations shown in the figure is realized by at least one of system control unit 50 and image processing unit 24 executing the corresponding operation.

[0069] The body part detection unit 1101 detects the body parts of a subject included in a captured image. In this embodiment, the body part detection unit 1101 detects the head of a person as described above and outputs the detection result. The body part detection unit 1101 detects the body parts of a subject included in the captured image both during learning related to inference model construction and during inference related to imaging for main subject selection.

[0070] The feature generation unit 1102 generates multidimensional time-series data related to features based on the detection results by the body part detection unit 1101 for multiple captured images obtained by intermittent imaging. In this embodiment, the feature generation unit 1102 generates three-dimensional time-series data that combines "changes in head orientation relative to the direction of travel," "changes in subject travel direction," and "changes in attention level" over a time length corresponding to a predetermined number of frames, as described above. Like the body part detection unit 1101, the generation of time-series data by the feature generation unit 1102 is performed both during learning related to inference model construction and during inference related to imaging in which a main subject is selected.

[0071] During learning related to building an inference model, the role setting unit 1103 associates the role of each subject with the time-series data of the subject. The role of each subject is set by the user via the operation unit 70.

[0072] The learning unit 1104 performs learning based on the time-series data associated with roles and constructs an inference model. More specifically, the learning unit 1104 maps multiple time-series data associated with roles into a multidimensional space and performs clustering, thereby constructing multiple subspaces formed by time-series data distributed without being separated by more than a predetermined distance, for example. For each of the multiple constructed subspaces, the learning unit 1104 also references role information associated with the time-series data forming the subspace and identifies a deterministic section formed only by time-series data of one role. The learning unit 1104 then averages the time-series data forming the deterministic space to derive a representative value (template) of the time-series data related to the deterministic space, and constructs an inference model including this representative value.

[0073] During imaging to select a main subject, the inference unit 1105 infers the role of the subject to be inferred based on the time-series data of the subject to be inferred generated by the feature generation unit 1102. More specifically, the inference unit 1105 inputs the time-series data of the subject to be inferred to the inference model constructed by the learning unit 1104, and obtains an inference result of the role of the subject. At this time, the inference model determines whether or not there exists a deterministic space in which the distance between the input time-series data and a representative value is closer than a predetermined threshold, and if there exists, determines the role of the subject to be inferred to be the role corresponding to the deterministic space, and outputs it as the inference result.

[0074] During imaging for selecting a main subject, the selection unit 1106 selects a subject to be the main subject based on the inference result by the inference unit 1105. The role of the subject to be the main subject may be set based on the set shooting mode, for example, and in the example of the shooting mode for shooting a ball game scene described above, the role of "player holding the ball" is set for the main subject.

[0075] <<Selection Process>> The selection process executed for selecting a main subject by the main body 100 of this embodiment having such a configuration will be described in detail using the flowchart in Fig. 12. The process corresponding to this flowchart can be realized by the system control unit 50 reading out a corresponding processing program stored in, for example, nonvolatile memory 56, and expanding and executing it in system memory 52. ​​This selection process will be described as being started, for example, when the main body 100 is started up in shooting mode. Furthermore, since the acquisition of a predetermined number of captured images is required to generate time-series data used for inference, the selection process may be executed for the first time after the predetermined number of frames have elapsed after start-up, and then executed every predetermined number of frames thereafter.

[0076] In S1201, the body part detection unit 1101 detects the head of a person (subject to be inferred) for each of a predetermined number of captured images obtained by imaging. The body part detection unit 1101 stores the detection results of the person's head in the memory 32. At this time, the body part detection unit 1101 associates the head detection results with each person based on the continuity between frames.

[0077] In S1202, the feature generation unit 1102 generates time series data related to the feature of each person to be inferred based on the person head detection results stored in S1201. More specifically, the feature generation unit 1102 derives time series data for each of "change in head orientation relative to the traveling direction," "change in subject traveling direction," and "change in attention level," and generates three-dimensional time series data that combines these.

[0078] In S1203, the inference unit 1105 uses the inference model to obtain inference results for the role of each person based on the time series data for each person to be inferred generated in S1202, and classifies the people to be inferred by role.

[0079] In S1204, the selection unit 1106 determines whether or not there is a person to be inferred who is classified into a class related to the role of "player in possession of the ball." If the selection unit 1106 determines that there is a person to be inferred who is classified into a class related to the role of "player in possession of the ball," the process proceeds to S1205, and if it determines that there is no person to be inferred, the process proceeds to S1206.

[0080] In S1205, the selection unit 1106 selects the person to be inferred who is classified into a class relating to the role of "player holding the ball" as the main subject, stores the information of the selection in the memory 32, and completes this selection process.

[0081] On the other hand, if it is determined in S1204 that there is no person to be inferred that is classified into a class related to the role of "player holding the ball," the selection unit 1106 completes this selection process by, for example, retaining the most recent selection result in S1206. That is, the selection unit 1106 maintains the person being tracked as the main subject immediately before executing this selection process as the main subject. Note that if the person selected as the most recent main subject is not within the imaging range or it is not possible to determine where in the imaging range the person is, for example, a person present near the position of the person selected as the most recent main subject may be selected as the main subject.

[0082] In this way, the image processing device of this embodiment can suitably identify the role of a subject in a scene. Specifically, when learning multiple captured images obtained by intermittent imaging of a ball game scene, etc., it is possible to cluster the images based on features appearing in the time-series data regardless of the assigned role labels, and prepare a template for a deterministic space in which a role can be uniquely identified. Therefore, the inference model obtained by learning can output more reliable inference results regarding the role of the subject to be inferred. Furthermore, compared to the technology described in Patent Document 1, the present invention limits the parts of the subject to be detected and does not require processing such as joint connection. Therefore, it is possible to stably select a subject with a desired role as the main subject with simple processing, even in scenes where multiple people are mixed together.

[0083] In this embodiment, as an example, learning is performed by assigning roles such as "player holding the ball" and "defensive player" based on video of ball game scenes such as soccer or basketball. However, the present invention is not limited to this. That is, the inference model is not limited to the above-mentioned ball games, and it goes without saying that it can be used for any scene, as long as learning is performed by labeling the role of each subject using video obtained from a scene in which multiple subjects are included in the imaging range. Furthermore, since the types of roles assigned to subjects vary depending on the scene, the construction (learning) of the inference model and the setting of the role selected as the main subject are performed for each scene, and the detector and inference model used during imaging to select the subject can be switched.

[0084] The switching may be performed based on the user's selection of the sport to be photographed, or on settings of photographing conditions such as the installation position of the digital camera system, the photographing angle of view, and the focal length. Alternatively, for example, when a person's head is detected, the photographing conditions may be determined based on information on the focal length, the subject distance, the subject size, the camera body posture, etc., and the switching may be performed without user operation. In addition, the sport to be photographed may be identified based on information on the location of the photographing location obtained by GPS or the like, or information on the clothing (uniform) and type of equipment (ball) used by the subject in the captured image, and control may be performed to use the corresponding detector or inference model.

[0085] Furthermore, in the above-described learning, the direction of travel is estimated simply based on the trajectory of the subject's head movement. However, the following method may be used to identify the type of movement depending on the situation of the scene. For example, in a ball game scene, the players' movements tend to be centered on the ball, and the main body 100 is often framed to capture the ball. Therefore, the center position of the game (ball) can be identified by averaging the trajectories of the heads of all players within the imaging range, and this can be subtracted from the trajectory of each player's head to obtain the trajectory of each player's head relative to the center of the game. In this way, it is possible to learn the characteristics of each player that are more suited to the scene.

[0086] In the example described above using the figures, the learning video was captured from a bird's-eye view of the subject, but the present invention is not limited to this. The video may be, for example, a horizontal shot of the field, i.e., a shot of the player from directly to the side. In this case, since it is difficult to detect the player's movement in the depth direction from the movement of the player's image in the captured image, the player's movement in the depth direction may be estimated using a depth map generated from a phase-difference AF signal in addition to the captured image.

[0087] 6, a method for deriving an attention level for a situation in which only the image of the player appears in the image capture range has been described, but the present invention is not limited to this. For example, depending on the angle of view at which the image is captured, the image capture range may include images of people who are not directly involved in the competition, such as spectators. If the attention level is derived based on the field of view of the images of these people, it may be difficult to extract features that are in line with the content of the competition. Therefore, for example, the detected head size or depth map may be further used to separate the image of the person in question, and then the attention level may be derived.

[0088] Furthermore, in the above-described example, the field of view of each subject is determined assuming that the subject is looking forward, but the present invention is not limited to this. That is, the present invention is not limited to determining the field of view assuming that the gaze direction 1321 of the person 1301 coincides with the forward direction 1311 of the person, as shown in FIG. 13(a). For example, the subject's eyeball region may be detected, and the field of view of the subject may be determined based on the gaze direction determined based on the position of the pupil in that region, and the attention level may be derived based on that. In this case, when the gaze direction 1322 of the person 1301 does not coincide with the forward direction 1311, as shown in FIG. 13(b), the field of view can be more accurately included in the feature. Furthermore, gaze direction information based on pupil images can also be learned by constructing time-series data of gaze movement that is not reflected in head rotation, thereby further improving the accuracy of inferring a person's role.

[0089] Furthermore, the parts of the subject detected in generating time-series data indicating the characteristics of changes in the state of the subject are not limited to the person's head or eyes, but may include other parts. For example, the direction of the subject's torso, the trajectory of movement, the overall state of the person, the direction of the feet, etc. may be detected, and the direction of the subject's movement may be derived based on this. Furthermore, the time-series data generated for learning based on these detected parts is not limited to the above examples. For example, time-series data of changes in the up-and-down movement of the head, time-series data of changes in the direction of movement relative to the longitudinal direction of the field (court), etc. may be added or changed.

[0090] In the above-described aspect, the inference model is described as being constructed by forming a subspace by mapping time-series data to a multidimensional space and performing clustering, and storing template information only for a deterministic space in which a role can be uniquely identified. However, the implementation of the present invention is not limited to this, and the inference model may include information on a subspace formed by including time-series data associated with different roles. In this case, if the time-series data of the subject to be inferred belongs to the subspace, processing may be performed, for example, by presenting multiple types of roles as candidates and allowing the user to select one, or by weighting based on the most recent inference result to determine one of the roles.

[0091] Furthermore, the information about the subspace included in the inference model does not need to be time-series data of representative values, but may be, for example, information indicating a multidimensional range representing the subspace. That is, the inference may be performed by determining whether the input time-series data belongs to the subspace based on the positional relationship between the input time-series data of the subject to be inferred and the subspace. Alternatively, for example, information about multiple feature points defining the subspace may be included in the inference model, and the inference may be performed by determining whether the input time-series data belongs to the subspace based on the distance between the feature points and the input time-series data. Furthermore, the inference result does not need to be output on the condition that the input time-series data belongs to any subspace. For example, even if the input time-series data does not belong to any subspace, the role associated with the closest subspace may be output. In this case, the inference result may be associated with information indicating a low likelihood.

[0092] Furthermore, in the above-described embodiment, learning is performed using clustering, but learning may be performed using other techniques, such as using a recurrent neural network.

[0093] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0094] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0095] 100: Main body, 24: Image processing unit, 32: Memory, 50: System control unit, 52: System memory, 55: Camera body orientation detection unit, 56: Non-volatile memory, 70: Operation unit, 1101: Body part detection unit, 1102: Feature generation unit, 1103: Role setting unit, 1104: Learning unit, 1105: Inference unit, 1106: Selection unit

Claims

1. a first acquisition means for acquiring a plurality of captured images obtained by intermittent imaging; a second acquisition means for acquiring a role for each of at least some of the subjects included in the plurality of captured images acquired by the first acquisition means; a generating means for detecting a state of the subject for each of the plurality of captured images and generating multidimensional time series data indicating a plurality of types of changes in the state of each subject; a constructing means for mapping the multidimensional time-series data of each of the subjects generated by the generating means onto a multidimensional space in association with the role of the subject, and constructing a plurality of subspaces formed by the multidimensional time-series data within the space; an input means for receiving the multidimensional time series data of a subject to be inferred; an inference means for determining and outputting a role of the subject to be inferred based on a positional relationship with the plurality of subspaces when the multidimensional time-series data of the subject to be inferred is mapped onto the multidimensional space; 1. An image processing device comprising:

2. 2. The image processing device according to claim 1, wherein the inference means determines a role of the subject to be inferred based on the subspace to which the multidimensional time-series data of the subject to be inferred belongs in the multidimensional space.

3. the constructing means determines, for each of the plurality of subspaces, a representative value of the multidimensional time series data forming the subspace; The inference means identifies the subspace to which the multidimensional time series data of the subject to be inferred belongs, based on a distance between the representative value of each of the plurality of subspaces in the multidimensional space and the multidimensional time series data of the subject to be inferred.

3. The image processing device according to claim 2.

4. 4. The image processing device according to claim 2, wherein the inference means identifies the subspace that is closest to the multidimensional time series data of the subject to be inferred as the subspace to which the multidimensional time series data of the subject to be inferred belongs.

5. The image processing device according to any one of claims 2 to 4, characterized in that the inference means identifies the subspace that is closer to the multidimensional time series data of the subject to be inferred by more than a predetermined threshold as the subspace to which the multidimensional time series data of the subject to be inferred belongs.

6. the configuration means sets a role for each of the plurality of subspaces based on a role associated with the multidimensional time-series data forming the subspace; The inference means determines a role of the subject to be inferred based on a role set in the subspace to which the multidimensional time-series data of the subject to be inferred belongs.

6. The image processing device according to claim 2, wherein the image processing device is a computer.

7. The image processing device according to claim 6, characterized in that the inference means determines the role of the subject to be inferred when there is one type of role set in the subspace to which the multidimensional time-series data of the subject to be inferred belongs.

8. 8. The image processing device according to claim 1, wherein the state of the subject includes at least one of a direction of travel of the subject, a direction of the subject's head, a trajectory of the subject's head, a direction of the subject's torso, a trajectory of the subject's torso, a direction of the subject's line of sight, and a level of attention of the subject.

9. An imaging device, An imaging means; An image processing device according to any one of claims 1 to 8; a control unit that controls the operation of the imaging device based on the role of the subject to be inferred output by the inference unit of the image processing device; An imaging device having the above configuration.

10. 10. The imaging apparatus according to claim 9, wherein the control means determines a main subject based on a role of the subject to be inferred.

11. a first acquisition step of acquiring a plurality of captured images obtained by intermittent imaging; a second acquisition step of acquiring a role for each of at least some of the subjects included in the plurality of captured images acquired in the first acquisition step; a generating step of detecting a state of the subject for each of the plurality of captured images and generating multidimensional time-series data indicating a plurality of types of changes in the state of each subject; a construction step of mapping the multidimensional time-series data of each subject generated in the generation step into a multidimensional space in association with the role of the subject, and constructing a plurality of subspaces formed by the multidimensional time-series data within the space; an input step of receiving the multidimensional time series data of a subject to be inferred; an inference step of determining and outputting a role of the subject to be inferred based on a positional relationship with the plurality of subspaces when the multidimensional time-series data of the subject to be inferred is mapped onto the multidimensional space; 1. A method for controlling an image processing apparatus, comprising:

12. A method for controlling an imaging device having an imaging unit and the image processing device according to any one of claims 1 to 8, comprising: A control method comprising a control step of controlling an operation of the imaging device based on the role of the subject to be inferred output by the inference means of the image processing device.

13. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing device and method, and imaging device

    JP2021105850A