Image processing apparatus and image processing method

The image processing device uses a machine learning model to infer the ranking of subjects in a single image frame, effectively addressing the inefficiency of existing methods and enabling real-time tracking of the leading subject.

JP2025132822APending Publication Date: 2025-09-10CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024030637
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-09-10

AI Technical Summary

Technical Problem

Existing image processing methods require multiple images or complex calculations to identify the leading subject among multiple subjects moving in the same direction, making them inefficient for real-time analysis.

Method used

An image processing device and method that utilize a machine learning model to infer the ranking of subjects in a single image frame, outputting an evaluation value for each region to determine the leading subject.

Benefits of technology

Accurately estimates the leading subject among multiple subjects in a single image frame, enabling efficient and real-time tracking and photography.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025132822000001_ABST
    Figure 2025132822000001_ABST
Patent Text Reader

Abstract

To provide an image processing apparatus and an image processing method that can accurately estimate, from an image of one frame, the top subject of a plurality of subjects included in the image.SOLUTION: An image processing apparatus acquires image data of one frame including a plurality of subjects moving in a predetermined direction. The image processing apparatus has a machine learning model for inferring the ranking in the predetermined direction of the plurality of subjects, and the machine learning model outputs, for each area of the plurality of subjects, an evaluation value having a larger value or a smaller value as the ranking is closer to a specific ranking.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device and an image processing method, and more particularly to a technique for estimating the positional relationship between a plurality of subjects. [Background technology]

[0002] When shooting scenes in which multiple subjects move in the same direction, such as track and field events or races, it may be desirable to track not a specific subject but a subject at a specific position (for example, the lead subject). Patent Document 1 proposes a method for identifying the relative positions of the subjects by detecting the three-dimensional positions of the subject's characteristic parts (head, joints, etc.) using images of the same subject captured from multiple different directions. Patent Document 2 also proposes a method for detecting the lead subject from multiple subjects moving in the same direction based on the direction of movement. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] US Patent Application Publication No. 2020 / 0401793 [Patent Document 2] Japanese Patent Publication No. 2022-022767 Summary of the Invention [Problem to be solved by the invention]

[0004] The method of Patent Document 1 requires detecting the joints of a subject from images of the same human subject captured by multiple image capture devices. Therefore, it is not possible to detect the leading subject based on one frame of an image captured by a single image capture device. Patent Document 2 also requires multiple images to detect the subject's movement direction.

[0005] In one aspect, the present invention provides an image processing device and an image processing method that can accurately estimate, from an image of one frame, a leading subject among a plurality of subjects included in the image. [Means for solving the problem]

[0006] In one aspect, the present invention provides an image processing device having an acquisition means for acquiring image data of one frame including a plurality of subjects moving in a predetermined direction, and a machine learning model for inferring the ranking of the plurality of subjects in the predetermined direction from the image data, wherein the machine learning model outputs an evaluation value for each region of the plurality of subjects, the larger or smaller the value the closer the ranking is to a specific ranking. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide an image processing device and an image processing method that can accurately estimate the leading subject among multiple subjects included in an image from an image of one frame. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of a digital camera as an example of an image processing apparatus according to an embodiment; [Figure 2] Flowchart regarding the operation of the object detection unit according to the embodiment [Figure 3] Flowchart regarding the operation of the object detection unit according to the embodiment [Figure 4] FIG. 1 is a block diagram showing an example of the configuration of an object detection unit according to an embodiment; [Figure 5] FIG. 10 is a block diagram showing another example of the configuration of the object detection unit according to the embodiment; [Figure 6] FIG. 1 is a diagram showing an example of an input image in an embodiment. [Figure 7] FIG. 1 is a diagram showing an example of an input image to a NN and a heat map output by the NN in an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] The present invention will be described in detail below based on exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Furthermore, although multiple features are described in the embodiments, not all of them are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] In the following, the present invention will be described in terms of an embodiment using a digital camera as an example of an image processing device. However, imaging functionality is not essential to the present invention, and the present invention can be implemented in any electronic device having one or more arithmetic circuits or processors. Such electronic devices include video cameras, computer devices (personal computers, tablet computers, media players, PDAs, etc.), smartphones, smart watches, game consoles, robots, drones, and drive recorders. These are merely examples, and the present invention can also be implemented in other electronic devices.

[0011] (Digital camera configuration) FIG. 1 is a block diagram showing an example of the functional configuration of a digital camera 100 (hereinafter simply referred to as camera 100) according to an embodiment of the invention. Camera 100 is capable of capturing and recording moving and still images. Functional blocks within camera 100 are communicably connected to one another via bus 160. The operation of camera 100 is realized by one or more programmable processors in main control unit 151 loading programs stored in ROM 155, for example, into RAM 154, executing the programs, and controlling the functional blocks.

[0012] Each functional block of the camera 100 can be implemented by software or a combination of software and hardware, except for parts that can clearly only be realized by hardware (e.g., the photographic lens 101, the image sensor 141, etc.). For example, a functional block may be realized by dedicated hardware such as an ASIC (Application Specific IC). A functional block may also be realized by a processor such as a CPU executing a program stored in memory. Multiple functional blocks may also be realized by a common configuration (e.g., one ASIC). Hardware that realizes part of the functions of one functional block may also be included in hardware that realizes another functional block.

[0013] The photographing lens 101 (lens unit) has a first lens group 102, a zoom lens 111, an aperture 103, a third lens group 121, a focus lens 131, a zoom motor (ZM) 112, an aperture motor (AM) 104, and a focus motor (FM) 132. The first lens group 102, the zoom lens 111, the aperture 103, the third lens group 121, and the focus lens 131 constitute a photographing optical system. The first lens group 102 and the third lens group 121 are fixed lenses, while the zoom lens 111 and the focus lens 131 are movable lenses. Although the lenses 102, 111, 121, and 131 are illustrated as single lenses, each may be composed of multiple lenses. The photographing lens 101 may also be configured as an interchangeable lens that can be removed from the camera 100.

[0014] The aperture control unit 105 controls the operation of the aperture motor 104 in accordance with commands from the main control unit 151 , and changes the aperture diameter of the aperture 103 . The zoom control unit 113 controls the operation of the zoom motor 112 in accordance with commands from the main control unit 151 , and changes the focal length (angle of view) of the photographic lens 101 .

[0015] The focus control unit 133 converts the defocus amount calculated by the depth calculation unit 163 into a drive amount and drive direction of the focus motor 132. Based on this drive amount and drive direction, the focus control unit 133 controls the operation of the focus motor 132 and drives the focus lens 131, thereby controlling the focus state of the photographing lens 101.

[0016] Here, the focus control unit 133 performs autofocus (AF) using a phase difference detection method, but the focus control unit 133 may also perform AF using a contrast detection method based on a contrast evaluation value of an image signal obtained from the image sensor 141.

[0017] The image sensor 141 may be, for example, a known CCD or CMOS color image sensor having a primary-color Bayer array color filter. The image sensor 141 has a pixel array in which multiple pixels are arranged two-dimensionally, and peripheral circuits for reading out signals from each pixel. Each pixel accumulates charge according to the amount of incident light through photoelectric conversion. By reading out signals from each pixel having a voltage according to the amount of charge accumulated during the exposure period, a group of pixel signals (analog image signals) representing a subject image formed on the imaging surface of the image sensor 141 by the photographing lens 101 can be obtained. The sensor control unit 143 controls the reading out of the analog image signals from the image sensor 141 in accordance with instructions from the main control unit 151.

[0018] The image signal read out from the imaging element 141 is supplied to the signal processing unit 142. The signal processing unit 142 applies signal processing such as noise reduction processing, A / D conversion processing, and automatic gain control processing to the analog image signal read out from the imaging element 141, and outputs the result as image data to the sensor control unit 143. The sensor control unit 143 stores the image data received from the signal processing unit 142 in the RAM 154.

[0019] When recording image data stored in RAM 154, main control unit 151 generates a data file according to the recording format by, for example, adding a predetermined header to the image data. At this time, main control unit 151 encodes the image data using compression / decompression unit 153 as necessary and stores the encoded data in a data file. Main control unit 151 records the generated data file on recording medium 157, such as a memory card.

[0020] Furthermore, when displaying image data stored in RAM 154, main control unit 151 uses image processing unit 152 to scale the image data so that it fits the display size of display unit 150, and then writes the scaled image data to an area of ​​RAM 154 used as a video memory (VRAM area). Display unit 150 reads the image data for display from the VRAM area of ​​RAM 154 and displays it on a display device such as an LCD or organic EL display. Display unit 150 also displays the detection result of the main subject detected by object detection unit 162 (such as a frame indicating the main subject area).

[0021] When shooting video (when in shooting standby mode or while recording video), camera 100 causes display unit 150 to function as an electronic viewfinder (EVF) by instantly displaying the shot video on display unit 150. The video and its frame images displayed when display unit 150 functions as an EVF are called live view images or through images. Furthermore, when shooting a still image, camera 100 displays the most recently shot still image on display unit 150 for a certain period of time so that the user can check the shooting results. These display operations are also realized under the control of main control unit 151.

[0022] The compression / decompression unit 153 encodes and decodes image data. For example, when recording still images or moving images, the image data and audio data are encoded using a predetermined encoding method. When playing back still image data files or moving image data files recorded on the recording medium 157, the compression / decompression unit 153 decodes the encoded data and stores it in the RAM 154.

[0023] The RAM 154 is used as a system memory for executing programs, a video memory, a buffer memory, and the like. The ROM 155 stores programs executable by the processor of the main control unit 151, various setting values, information specific to the camera 100, GUI data, etc. The ROM 155 may be electrically rewritable.

[0024] Operation unit 156 is a collective term for a group of input devices such as switches, buttons, keys, and touch panels that allow the user to input instructions to camera 100. Input via operation unit 156 is detected by main control unit 151 via bus 160, and main control unit 151 controls each unit to realize operations according to the input.

[0025] The main control unit 151 has one or more programmable processors such as a CPU or MPU, and controls each unit by loading a program stored in, for example, ROM 155 into RAM 154 and executing it, thereby realizing the functions of the camera 100. The main control unit 151 also executes AE processing to automatically determine exposure conditions (shutter speed or accumulation time, aperture value, sensitivity) based on information about the brightness of the subject. Information about the brightness of the subject can be obtained, for example, from the image processing unit 152. The main control unit 151 can also determine the exposure conditions based on brightness information about the area of ​​a specific subject, such as a person's face.

[0026] During video shooting, the main control unit 151 fixes the aperture 103 and controls exposure by the electronic shutter speed (accumulation time) and the magnitude of the gain. The main control unit 151 notifies the sensor control unit 143 of the determined accumulation time and the magnitude of the gain. The sensor control unit 143 controls the operation of the image sensor 141 so that shooting is performed in accordance with the notified exposure conditions.

[0027] The image processing unit 152 applies predetermined image processing to the image data stored in the RAM 154. The image processing applied by the image processing unit 152 may include, for example, color interpolation processing, correction processing, detection processing, data processing, evaluation value calculation processing, special effect processing, and the like. Color interpolation, also known as demosaicing, is performed when the image sensor is equipped with a color filter, and is a process of interpolating the values ​​of color components that are not included in the individual pixel data that make up the image data. The correction processing may include white balance adjustment, tone correction, correction of image degradation caused by optical aberration of the imaging optical system 101 (image restoration), correction of the effects of vignetting in the imaging optical system 101, color correction, and the like. The detection process may include detection of characteristic regions (for example, face or head regions or human body regions) and their movements, person recognition processing, and the like. Data processing can include processes such as cutting out an area (trimming), combining, scaling, etc. Data processing also includes generating image data for display or image data for recording. The evaluation value calculation process may include processing for generating an evaluation value used in automatic exposure control (AE), etc. Instead of the depth calculation unit 163, the image processing unit 152 may calculate the defocus amount. Special effect processing can include adding a blur effect, changing color tones, relighting, and the like. It should be noted that these are examples of processes that can be applied by the image processing unit 152, and do not limit the processes that can be applied by the image processing unit 152.

[0028] Information about the area of ​​the main subject selected by main control unit 151 may be used for image processing (such as white balance adjustment processing or processing to generate luminance information about the subject) in image processing unit 152. When focus control unit 133 performs contrast detection AF, image processing unit 152 can generate a contrast evaluation value and supply it to focus control unit 133. Image processing unit 152 stores the processed image data, information about the detected characteristic area, and the like in RAM 154.

[0029] The position and orientation change acquisition unit 161 is configured with an orientation sensor such as a gyroscope, an acceleration sensor, or an electronic compass, and measures changes in the position and orientation of the camera 100. In this embodiment, as an example, the optical axis of the photographing lens 101 is defined as the roll axis, an axis perpendicular to the roll axis and parallel to the longitudinal direction of the image sensor is defined as the pitch axis, and an axis perpendicular to the roll axis and pitch axis is defined as the yaw axis. The position and orientation change acquisition unit 161 stores information indicating the detected changes in position and orientation in the RAM 154. The image processing unit 152 and the like can refer to the information indicating the changes in position and orientation detected by the position and orientation change acquisition unit 161.

[0030] Object detection unit 162 infers the order of a predetermined object (here, a human head) present in an image from one frame of image data (for example, display image data). Then, object detection unit 162 selects one area from the image as the area of ​​the main subject based on the inference result. Object detection unit 162 stores information on the selected area of ​​the main subject in RAM 154.

[0031] In this embodiment, the object detection unit 162 performs inference using a trained machine learning model that uses a neural network (NN). The object detection unit 162 may have a hardware circuit for performing neural network calculations at high speed. Examples of such hardware circuits include a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), an FPGA (Field Programmable Gate Array), and an ASIC. A parameter set for realizing the trained machine learning model can be stored in, for example, the ROM 155. Examples of neural networks include a convolutional neural network (CNN) and a recurrent neural network.

[0032] As will be described later, the form of the output result of the NN used in the object detection unit 162 may differ depending on the learning method. However, the NN outputs an evaluation value (also called priority, likelihood, or score) that has a larger or smaller value for each region of a predetermined object in the input two-dimensional image, the greater the likelihood that the region is located at the top.

[0033] Based on the output of the object detection section 162, the main control section 151 can set the focus detection area to the area with the largest likelihood or score. This enables tracking shooting so that the leading subject is in focus even if the leading person changes. Note that if the focus detection area is set to the area with the nth priority (n is an integer of 2 or more), it is possible to track and shoot the nth subject from the beginning. The main control section 151 saves information about the area in which the focus detection area is set as information about the main subject area in RAM 154. The information about the main subject area is used for processing by the image processing section 152, etc.

[0034] In addition, the main control unit 151 may control the related functional blocks so that the captured image data, the detection results of the object detection unit 162, and the detection results of the position and orientation change acquisition unit 161 during the most recent predetermined period are stored in the RAM 154.

[0035] The main control unit 151 can also display the detection results by the object detection unit 162. For example, based on the output result of the object detection unit 162, the main control unit 151 may display a frame-shaped indicator indicating the area of ​​a predetermined object and the corresponding ranking superimposed on the image data used for the detection.

[0036] The depth calculation unit 163 calculates the defocus amount of an arbitrary region in an image using any known method. The depth calculation unit 163 can calculate the defocus amount based on the phase difference between a pair of focus detection signals obtained from the image sensor 141, for example. Alternatively, the depth calculation unit 163 may calculate the defocus amount based on the phase difference between a pair of focus detection signals obtained from an AF sensor provided separately from the image sensor 141. The defocus direction is indicated by the sign of the defocus amount.

[0037] The depth calculation unit 163 may calculate the defocus amount for one focus detection area set by the main control unit 151, or may calculate the defocus amount for each area obtained by dividing the entire image in the horizontal and vertical directions. Information representing the distribution of defocus amounts in the entire image is also called a defocus map. The defocus amounts calculated by the depth calculation unit 163 and information about the corresponding areas are stored in RAM 154 and can be referenced by the focus control unit 133, image processing unit 152, etc.

[0038] (Operation of the object detection unit 162) The operation of the object detection unit 162 will be described with reference to the flowchart of FIG. In S200, the object detection unit 162 inputs image data to a trained machine learning model. The image data input to the machine learning model is two-dimensional image data, and may be, for example, image data for display generated by the image processing unit 152. Preprocessing such as reduction processing may be applied to the image data for display before inputting it to the machine learning model.

[0039] In addition to the two-dimensional image data, information on a rectangular area in which a specific object exists within the two-dimensional image data can also be input to the machine learning model. In this case, information on the area of ​​the specific object detected as a feature area by the image processing unit 152 can be used.

[0040] In S201, the object detection unit 162 acquires a detection result from a machine learning model. As described above, the form of the detection result differs depending on the learning method of the machine learning model. However, for each region where the presence of a predetermined object is estimated, an evaluation value (also called likelihood or score) is acquired, with the leading region having a larger or smaller value. The object detection unit 162 stores the acquired detection result in RAM 154.

[0041] A typical NN learning method will be explained using Figure 3. First, a training dataset is prepared, consisting of a set of input data to the NN and teacher data that represents the desired output result for the input data. Here, the input data is assumed to be 2D image data, or a combination of 2D image data and information indicating the area of ​​a specific object. The information indicating the area of ​​the specific object may be, for example, but is not limited to, the coordinates of the two diagonal vertices of a rectangular area inscribed by the area of ​​the specific object, or the coordinate of one vertex of the rectangular area and the horizontal and vertical sizes of the area.

[0042] The input data may be a real-life image or a CG (Computer Graphics) image. The training data is prepared in the same format as the output format of the NN. For example, if a two-dimensional evaluation value or score map is to be output from the NN, the correct evaluation value or score map is prepared. Also, if an evaluation value or score for each region of a specific object is to be output, the correct evaluation value or score is prepared.

[0043] Note that the learning of the NN may be performed by the main control unit 151 or the object detection unit 162, but since the learning requires many calculations, it is more practical to perform it on an external device with higher processing power. For convenience, the following description will be given assuming that the learning is mainly performed by the object detection unit 162.

[0044] In S300, the object detection unit 162 inputs the input data of the training data set to the NN.

[0045] In S301, the object detection unit 162 executes the NN calculation and stores the result in the RAM 154, for example.

[0046] In S302, the object detection unit 162 inputs the result obtained in S301 and the training data corresponding to the input data into a predetermined evaluation function.

[0047] In S303, the object detection unit 162 executes calculation of the evaluation function and stores the result in the RAM 154, for example.

[0048] In S304, the object detection unit 162 changes the parameters of the NN so that the absolute value of the calculation result of the evaluation function becomes smaller. This change can be performed using a known optimization algorithm such as a gradient variation method.

[0049] In S305, the object detection unit 162 determines whether the learning is sufficient, and if it is determined that the learning is sufficient, ends the learning, but if it is not determined that the learning is sufficient, repeats the process from S300. The object detection unit 162 determines that the learning is sufficient when a predetermined convergence condition is met, for example, when the absolute value of the change in the calculation result of the evaluation function obtained in S303 is less than a threshold, or when the accuracy rate of the NN output is equal to or greater than a threshold.

[0050] Although FIG. 3 shows that the parameters are changed for each sample of training data (each set of input data and teacher data), the parameters may be changed based on the calculation results of the evaluation function for multiple samples.

[0051] Based on the general learning method described above, the learning method of the NN in this embodiment will be described. In this embodiment, the input data included in the training dataset is image data of a competition (sprint, middle, and long distance running, relay, race walking, marathon, long-distance relay, etc.) in which multiple runners (subjects) run along a course. The training data is the position and ranking of the runners' head regions included in the input data. The position of the head region is defined as the image coordinates of the center (intersection of the diagonals) of a rectangle inscribed by the head region, and the horizontal and vertical sizes (number of pixels) of the rectangle.

[0052] The NN outputs candidates for the head region and a priority or score for each candidate. The priority or score has a value ranging from 0 to 1, with a higher or lower value the closer the head region is to the top. The candidates for the head region and the priority or score may be obtained from the same layer or different layers among the multiple layers that make up the NN.

[0053] The NN can be configured to output the priorities or scores corresponding to the head regions in any format. For example, it can be configured to output a two-dimensional distribution of scores corresponding to the input image as a heat map, or to output one score for each head region. When configuring it to output scores as a heat map, training data can be used in which the scores for each head region are Gaussian-distributed, with the center of gravity as the peak. In this case, the score corresponding to the peak can be a fixed value according to the ranking. Figures 7(a) and (b) show examples of input images and heat maps of scores output by the NN.

[0054] When a heat map of scores is output from the NN, the main control unit 151 selects a head region including a pixel position with the highest value in the heat map, and sets, for example, a focus detection region. Also, when information about the head region and the corresponding score are output separately from the NN, the control unit 151 selects the head region with the highest score, and sets, for example, a focus detection region.

[0055] FIG. 4 shows an example of a configuration in which, for input image data 400, candidates for the head region and a priority or score 402 for each candidate are obtained from the same layer (for example, the output layer) of a neural network 401.

[0056] 5 shows an example of a configuration in which a first NN 501 detects head region candidates from input image data, and a second NN 503 obtains a priority or score for each head region candidate using the detection results of the first NN 501. When using the configuration of FIG. 5, information on the head region can be obtained from the first NN 501, and the corresponding score can be obtained from the second NN 503 in the form of a numerical value or heat map.

[0057] In this embodiment, the output of the evaluation function used in learning the neural network is set to a value that takes into account the ranking of the subject, thereby achieving learning that improves the accuracy of detecting subjects of a specific ranking. Specifically, for each candidate head region, a priority or score is calculated, with the higher or lower value corresponding to the position closer to the top, and the greater the difference from the value in the training data, the greater the loss. In this case, the closer the head region is to the top of the training data, the greater the loss relative to the magnitude of the error.

[0058] Consider the scene shown in Figure 6(a), in which three runners, A, B, and C, are running from right to left in the image. Runner A is in the lead (1st place), runner B is in 3rd place, and runner C is in 2nd place.

[0059] The evaluation function l_k for rank k (k is an integer between 1 and 3) is expressed as the difference between the training data T_k and the output value y_k of the NN. Generally, it can be expressed as a function f of the following equation (1-1) with the rank T_k and the output value y_k as arguments. The absolute value function or cross-entropy function can be used for the function f. As an example, when using the absolute value function of the difference, the function f becomes -|y_k-T_k|. Note that the absolute value is negative because it is common in this field to make the value of the evaluation function negative.

[0060] The evaluation function L for the entire image is a function obtained by multiplying the evaluation function f of rank k by a weighting coefficient c_k and adding the result, as shown in the following equation (1-2). In NN training, the parameters of the NN are optimized so that the absolute value of the evaluation function L becomes small.

[0061] When using NN for tracking and photographing a leading subject, the higher the ranking, the more accurate the inference (higher the accuracy rate) is required. Therefore, in the evaluation function L, the higher the ranking, the greater the weight is used for the evaluation function l_k, as shown in the following equation (1-3), so that the contribution of the evaluation function l_k with a higher ranking is greater. l_k=f(y_k,T_k) …Formula (1-1) L=Σc_k*f(y_k,T_k) …Formula (1-2) c_i ≧ c_j (when i has higher priority than j) ... Equation (1-3)

[0062] If the weight c_k of the evaluation function l_k of the ranking is constant regardless of the ranking k, the NN is trained to have a similar accuracy rate regardless of the ranking. On the other hand, in this embodiment, the weight c_k is made larger as the ranking becomes higher, so the NN is trained to have a higher accuracy rate for subjects with higher rankings. Therefore, the accuracy of detecting the leading or first-place runner using the NN can be improved.

[0063] When using weights according to rank, it is possible to achieve learning with a high accuracy rate for subjects in a specific rank by increasing the weight of a specific rank other than the top or 1st place. Therefore, by storing learning results (parameters) with different ranks with high accuracy rates in, for example, ROM 155 and applying parameters according to the desired rank to the NN, it is possible to accurately detect subjects in the desired rank.

[0064] Furthermore, when using weights according to rank, the ranks of the subjects to be detected can be narrowed down by setting the weights of ranks that meet the conditions, such as 4th place or lower, to 0. Furthermore, when there are multiple subjects with the same rank, the weights according to rank may be the same value, or one of the weights may be increased according to a predetermined condition. For example, a weight can be increased for a subject with a large head area or a subject whose in-image distance from the head area of ​​the leading subject is short. These are merely examples, and other conditions may also be used. Furthermore, multiple conditions may be used in combination.

[0065] According to this embodiment, a machine learning model that has been trained to detect a subject of a specific rank from among multiple subjects contained in an image is used, and the machine learning model is trained to have a higher detection accuracy for subjects of a specific rank than for subjects of other ranks. This makes it possible to detect subjects of a specific rank with high accuracy, and improves the accuracy of processing using the detection results of the machine learning model, such as tracking photography.

[0066] ●(Second embodiment) Next, a second embodiment of the present invention will be described. This embodiment differs from the first embodiment in the evaluation function used during NN training.

[0067] In this embodiment, an evaluation function is used that is based on the difference between the evaluation value for each region of a specific object detected by the NN in the input image data of the training dataset and the evaluation value of the corresponding teacher data for each of a plurality of combinations of orders. In other words, while the first embodiment calculates the difference for each rank, this embodiment calculates the difference for each combination of ranks.

[0068] The evaluation function in this embodiment will be further explained below using an example. As in the first embodiment, assume a scene in which three runners A, B, and C are running from right to left in the image shown in Fig. 6(a). Runner A is in the lead (first place), runner B is in third place, and runner C is in second place.

[0069] The evaluation function l_k,j for rank k,j (1≦k,j≦3) is expressed as the deviation of the scores y_k,y_j output by the NN for the subject of rank k,j from the training data T_k,T_j for rank k,j. Generally, it can be expressed as a function f of the following equation (2-1) with the rank T_k,j and the output values ​​y_k,y_j as arguments. The function f can be an absolute value function or a cross-entropy function. As an example, when using the absolute value function of the difference, the function f is -|(y_k-y_j)-(T_k-T_j)|.

[0070] The evaluation function L for the entire image is a function that multiplies the evaluation function f for rank k,j by the weighting coefficient c_k,j and sums them, as shown in the following equation (2-2). In NN training, the parameters of the NN are optimized so that the absolute value of the evaluation function L becomes small.

[0071] When using NN for tracking and photographing a leading subject, the higher the ranking, the more accurate the inference (higher the accuracy rate) is required. Therefore, in the evaluation function L, the contribution of the evaluation function l_k with a higher ranking is increased, so that the evaluation function for the combination with a higher ranking uses a larger weight, as shown in the following equation (2-3). l_k,j=f(y_k,y_j,T_k,T_j)...Equation (2-1) L=Σc_k,j*f(y_k,y_j,T_k,T_j)...Equation (2-2) c_k,j≧c_m,n (when k is smaller than m, or when k=m, the difference between k and j is greater than the difference between m and n)... Equation (2-3)

[0072] As an example, in this embodiment, it is possible to set c_1,2 = 3, c_1,3 = 5, and c_2,3 = 1. If the weight c_k,j of the evaluation function l_k,j for the combination of rankings is kept constant regardless of the rankings k, j, the NN is trained to have a similar accuracy rate regardless of the rankings.

[0073] On the other hand, in this embodiment, the weight c_k,j is increased as the difference in the combined rankings increases and as the rankings closer to the top are included. Therefore, the NN is trained so that the accuracy rate increases for combinations of subjects with higher rankings. This makes it possible to improve the accuracy of detecting the leading or first-place runner using the NN, just like in the first embodiment.

[0074] The evaluation function described in this embodiment can be used in combination with the evaluation function described in the first embodiment. In this case, the NN is trained to update the parameters so that the sum of the outputs of the two evaluation functions becomes smaller.

[0075] ●(Third embodiment) Next, a second embodiment of the present invention will be described. This embodiment differs from the first and second embodiments in the training data set used when training the NN.

[0076] Specifically, the teacher data in the training dataset is the distance to a reference position or a value based on the distance. The reference position can be predetermined, such as a goal or a boundary in the direction of movement of the shooting range. The distance-based value may be a score or ranking converted from the distance. The distance may be an image distance or a real distance. When the real distance is used, an actually measured distance may be used, or a distance obtained by simulating the shooting scene using 3D modeling may be used. The distance may also be estimated from an image of an indicator in real space that indicates the distance to the reference position. When input image data for the training dataset is generated using 3D modeling, it is possible to generate input image data suitable for learning, thereby achieving efficient learning.

[0077] The training of the NN can be performed using the evaluation function described in the first or second embodiment, in which the argument is a distance or a difference in a value based on the distance. The training is completed by optimizing the parameters so that the output of the evaluation function is smaller than a predetermined threshold.

[0078] ●(Fourth embodiment) Next, a second embodiment of the present invention will be described. This embodiment differs from the first and second embodiments in the training data set used when training the NN.

[0079] In this embodiment, data augmentation is applied to a training dataset. Data augmentation is the process of applying processing such as rotation, inversion, masking, and cropping to other input image data included in a previously prepared training dataset to generate new input image data. By adding this new input image data to the training dataset along with the corresponding teacher data, the generalization performance of the NN can be improved.

[0080] When adding a training dataset through data augmentation, if part of a specific object's region is lost due to processing of the input image data, the corresponding training data is modified to match the processed input image data. For example, if the specific object's region is the head region, the training data for the head region that is missing a certain percentage (e.g., 30%) or more is deleted, and the training data for the remaining head region is modified. Modification of the training data can be done, for example, by reassigning the ranking or score for the head region.

[0081] FIG. 6(b) shows an example in which data augmentation is applied, in which the input image data shown in FIG. 6(a) is moved to the center (top is cropped). As shown in FIG. 6(b), the head region of runner A is outside the cropping range due to data augmentation (more than 30% is lost). Therefore, the object detection unit 162 deletes the data related to the head region of runner A from the training data for the input image data shown in FIG. 6(a) and reclassifies runners B and C to second and first place. The object detection unit 162 then adds the corrected training data to the training dataset together with the input image data as training data corresponding to the input image data shown in FIG. 6(b). Note that the black area at the bottom of FIG. 6(b) is an area that does not exist in the original input image data. Areas hidden by masking are also black areas.

[0082] Data augmentation can reduce the time and effort required to prepare input image data to be used as a training dataset. Furthermore, training data corresponding to input image data generated by data augmentation is generated by modifying training data corresponding to the original input image data. Therefore, even if at least a portion of the region of a specific object to be detected is lost due to data augmentation, learning can be achieved using appropriate training data.

[0083] The addition of a training data set according to this embodiment can be implemented in combination with the first to third embodiments.

[0084] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0085] The disclosure of the present embodiment includes the following image processing device, imaging device, image processing method, and program. (Item 1) an acquisition means for acquiring image data of one frame including a plurality of subjects moving in a predetermined direction; a machine learning model that infers a ranking of the plurality of subjects in the predetermined direction from the image data, the machine learning model outputs an evaluation value for each region of the plurality of subjects, the evaluation value having a larger or smaller value as the ranking approaches a specific ranking; 1. An image processing device comprising: (Item 2) 2. The image processing device according to item 1, wherein the machine learning model is trained so that the closer the ranking is to a specific ranking, the higher the accuracy rate. (Item 3) 3. The image processing device according to item 1 or 2, characterized in that the machine learning model is a machine learning model using a neural network, and is trained to correct parameters of the neural network so that the output of the neural network for input image data included in a training dataset and the output of an evaluation function that takes as arguments teacher data corresponding to the input image data satisfy a convergence condition. (Item 4) Item 3. The image processing device according to item 3, characterized in that the evaluation function is expressed by a weighted sum of the evaluation functions of the rankings, and the weight of the evaluation function of the ranking corresponding to the specific ranking is greater than the weight of the evaluation function of the ranking corresponding to another ranking. (Item 5) 5. The image processing device according to item 3 or 4, wherein the evaluation function for the ranking is an evaluation function for each ranking. (Item 6) 5. The image processing device according to item 3 or 4, wherein the evaluation function of the rankings is an evaluation function for each combination of rankings. (Item 7) 7. The image processing device according to any one of items 1 to 6, wherein a training dataset used to learn the machine learning model is composed of a set of input image data and corresponding teacher data, and the teacher data has the form of a heat map representing the distribution of the evaluation values. (Item 8) 7. The image processing device according to any one of items 1 to 6, wherein the training dataset used to learn the machine learning model is composed of a set of input image data and corresponding teacher data, and the teacher data has the form of a distance to a reference position or a value based on the distance as the evaluation value. (Item 9) 9. The image processing device according to item 8, wherein the input image data is image data generated by 3D modeling of a photographed scene. (Item 10) The image processing device according to any one of items 7 to 9, characterized in that the input image data includes input image data generated by processing other input image data, and training data for the input image data generated by processing the other input image data is generated based on the training data for the other input image data. (Item 11) Item 11. The image processing device according to item 10, characterized in that when the processing causes at least a portion of the area of ​​the subject contained in the other input image to no longer be included, the image processing device generates training data for the input image data generated by processing the other input image data by correcting the training data for the other input image data in accordance with the processing. (Item 12) 12. The image processing device according to any one of items 1 to 11, wherein the specific ranking is first. (Item 13) The image processing device according to any one of items 1 to 12, An imaging element; generating means for generating one frame of image data using the imaging element; a detection means for inputting the image data into a machine learning model included in the image processing device and detecting the subject of the specific rank from a plurality of subjects included in the image data; and a setting unit that sets a focus detection area to the subject of the specific rank detected by the detection unit. (Item 14) An image processing method executed by an image processing device, acquiring one frame of image data including a plurality of subjects moving in a predetermined direction; and inferring, from the image data, a ranking of the plurality of objects in the predetermined direction using a machine learning model; the machine learning model outputs an evaluation value for each region of the plurality of subjects, the evaluation value having a larger or smaller value as the ranking is closer to a specific ranking; An image processing method comprising: (Item 15) 13. A program for causing a computer to function as the image processing device according to any one of items 1 to 12.

[0086] The present invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Therefore, the following claims are appended to clarify the scope of the invention. [Explanation of symbols]

[0087] 100...imaging device, 101...lens unit, 141...imaging element, 151...main control unit, 152...image processing unit, 162...object detection unit

Claims

1. an acquisition means for acquiring image data of one frame including a plurality of subjects moving in a predetermined direction; a machine learning model that infers a ranking of the plurality of subjects in the predetermined direction from the image data, the machine learning model outputs an evaluation value for each region of the plurality of subjects, the evaluation value having a larger or smaller value as the ranking approaches a specific ranking; 1. An image processing device comprising:

2. The image processing device according to claim 1 , wherein the machine learning model is trained so that the closer the ranking is to a specific ranking, the higher the accuracy rate.

3. The image processing device described in claim 1, characterized in that the machine learning model is a machine learning model using a neural network, and is trained to modify the parameters of the neural network so that the output of the neural network for input image data included in a training dataset and the output of an evaluation function that takes as arguments teacher data corresponding to the input image data satisfy a convergence condition.

4. 4. The image processing device according to claim 3, wherein the evaluation function is expressed by a weighted sum of evaluation functions of ranks, and the weight of the evaluation function of the rank corresponding to the specific rank is greater than the weight of the evaluation function of the rank corresponding to another rank.

5. 4. The image processing apparatus according to claim 3, wherein the evaluation function for the ranking is an evaluation function for each ranking.

6. 4. The image processing apparatus according to claim 3, wherein the evaluation function of the rankings is an evaluation function for each combination of rankings.

7. The image processing device according to claim 1, characterized in that the training dataset used to learn the machine learning model is composed of a set of input image data and corresponding teacher data, and the teacher data has the form of a heat map representing the distribution of the evaluation values.

8. The image processing device described in claim 1, characterized in that the training dataset used to learn the machine learning model is composed of a set of input image data and corresponding teacher data, and the teacher data has the form of a distance to a reference position or a value based on the distance as the evaluation value.

9. 9. The image processing apparatus according to claim 8, wherein the input image data is image data generated by 3D modeling of a photographed scene.

10. The image processing device described in claim 7, characterized in that the input image data includes input image data generated by processing other input image data, and training data for the input image data generated by processing the other input image data is generated based on the training data for the other input image data.

11. The image processing device described in claim 10, characterized in that when the processing causes at least a portion of the subject area contained in the other input image to no longer be included, the training data for the other input image data is corrected in accordance with the processing, thereby generating training data for the input image data generated by processing the other input image data.

12. 2. The image processing apparatus according to claim 1, wherein the specific ranking is first place.

13. An image processing device according to any one of claims 1 to 12; An imaging element; generating means for generating one frame of image data using the imaging element; a detection means for inputting the image data into a machine learning model included in the image processing device and detecting the subject of the specific rank from a plurality of subjects included in the image data; and a setting unit that sets a focus detection area to the subject of the specific rank detected by the detection unit.

14. An image processing method executed by an image processing device, acquiring one frame of image data including a plurality of subjects moving in a predetermined direction; and inferring, from the image data, a ranking of the plurality of objects in the predetermined direction using a machine learning model; the machine learning model outputs an evaluation value for each region of the plurality of subjects, the evaluation value having a larger or smaller value as the ranking approaches a specific ranking; An image processing method comprising:

15. A program for causing a computer to function as the image processing device according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Subject tracking device and subject tracking method, and imaging apparatus

    JP2022022767A

  • Apparatus and methods for determining multi-subject performance metrics in a three-dimensional space

    US20200401793A1