Image processing apparatus, imaging apparatus, and image processing method
The image processing device uses machine learning-based detection methods to reference multiple frames for joint and object detection, addressing accuracy issues in human body and object detection by interpolating from past frames, thereby maintaining detection precision.
Patent Information
- Application Number
- JP2024103360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2044-06-26
AI Technical Summary
Existing methods for detecting human bodies and associated objects in images, such as those used for autofocus in imaging devices, suffer from accuracy issues when interpolation from past images is not sufficient due to dynamic shape changes.
An image processing device utilizing first and second detection means based on machine learning to detect objects and infer behavior, allowing reference to up to n and m frames for joint and object detection respectively, with m > n, to minimize accuracy loss from occlusions.
The method effectively maintains detection accuracy by utilizing past information when parts of the human body or objects are occluded, reducing the impact of dynamic movements and occlusions.
Smart Images

Figure 2026005118000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device having a subject detection mechanism, an imaging device, and an image processing method. [Background technology]
[0002] 2. Description of the Related Art Conventionally, various techniques have been proposed for detecting a subject to be controlled in order to perform imaging control such as autofocus (AF) in imaging devices such as digital cameras.
[0003] Patent Document 1 discloses a technology for recognizing road markings, and when the reliability of recognition is low, interpolation is performed from past images to improve accuracy. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2018-88144 Summary of the Invention [Problem to be solved by the invention]
[0005] Since the interpolated road does not change shape dynamically, the disclosed method described in Patent Document 1 is thought to be sufficient for recognizing road markings, but when applied to recognizing human bodies, it is not always possible to obtain high accuracy by interpolating from past images. [Means for solving the problem]
[0006] In order to achieve the above-mentioned object, the image processing device of the present invention comprises a first detection means for detecting a first object from an image using a first learning model based on machine learning, a second detection means for detecting a second object different from the first object from an image using a second learning model based on machine learning, and a behavior detection means for detecting a specific behavior of a subject based on the detection results by the first and second detection means in a moving image consisting of multiple input frames, wherein the first detection means is capable of referring to the detection results of up to n frames (n≧0) before the target frame in order to output the detection result of the first object in the target frame, and the second detection means is capable of referring to the detection results of up to m frames (m>n) before the target frame in order to output the detection result of the second object in the target frame. [Effects of the Invention]
[0007] According to the present invention, when a part of a human body, an object associated with the human body, belongings, or sporting goods used in a particular sport is occluded or cannot be detected, past information can be used to minimize the decrease in detection accuracy. [Brief explanation of the drawings]
[0008] [Figure 1] Interchangeable lens camera system block diagram [Figure 2] Flowchart diagram showing the procedure for detecting human behavior [Figure 3] A diagram showing the operations of the joint detection unit 251 and the object detection unit 252. [Figure 4] A diagram showing the operations of the joint detection unit 251 and the object detection unit 252. [Figure 5] FIG. 10 shows an example of joint detection and object detection in multiple frames. DETAILED DESCRIPTION OF THE INVENTION
[0009] The best mode for carrying out the present invention will be described in detail below with reference to the accompanying drawings. Note that the embodiment described below is an example of a means for realizing the present invention, and should be appropriately modified or changed depending on the configuration of the device to which the present invention is applied and various conditions, and the present invention is not limited to the following embodiment.
[0010] [First embodiment] In this embodiment, when photographing a group sport, a specific movement to be photographed is detected, and one of multiple players is selected as the main subject for photography control such as autofocus (AF). To determine the main subject, the positions of the player's joints are detected to detect the specific movement. Depending on the sport, the ball is detected, and the relationship between the ball and the position of the player's joints is detected to detect the specific movement.
[0011] However, there are cases where the opposite arm or ball cannot be detected sufficiently or continuously because it is hidden by the human body or torso. When detecting a person's posture or the movement of the ball they are holding, the occlusion of these elements hinders movement detection.
[0012] For example, if you keep the ball's coordinates before it gets hidden, and the ball gets hidden in the next frame, the ball won't be moving very fast, so using the position from the previous frame will still provide a reasonable hint.
[0013] However, if the previous position of the human body is used just because the opposite arm cannot be detected because it is hidden by the torso, the human body will move while changing its posture, resulting in an unnatural posture. When trying to estimate movement from a person's posture, an unnatural posture will hinder detection.
[0014] Therefore, in this embodiment, when joint points or the ball cannot be detected in the target frame, interpolation is performed by referencing past frames. At this time, a feature of this embodiment is that it is configured so that more past frames can be referenced for the ball (second subject), which is more stable, than for the more unstable joint (first subject).
[0015] The camera 20 of this embodiment detects and tracks specific actions of people and supports photography.
[0016] FIG. 1 is a block diagram showing an example of the functional configuration of an interchangeable lens camera as an example of an imaging apparatus according to a first embodiment of the present invention.
[0017] The imaging device of this embodiment is composed of an interchangeable lens unit 10 and a camera body 20. A lens control unit 106, which controls the overall operation of the lens, and a camera control unit 250, which controls the overall operation of the camera system including the lens unit 10, can communicate with each other via a terminal provided on the lens mount. The camera body 20 is a digital still camera or video camera that captures a subject and records video and still image data on various media such as tape, solid-state memory, optical disks, and magnetic disks, but is not limited to these. Examples include mobile phones (smartphones), personal computers (laptop, desktop, tablet, etc.), game consoles, in-vehicle sensors, FA (factory automation) devices, drones, and medical devices. The present invention is applicable to any device that has an imaging device built in or externally connected. Therefore, the term "imaging device" in this specification is intended to encompass any electronic device equipped with an imaging function. Furthermore, the term "image processing device" in this specification is intended to encompass any electronic device that determines a main subject based on an image captured by the imaging device.
[0018] First, the configuration of lens unit 10 will be described. Fixed lens 101, aperture 102, and focus lens 103 constitute the imaging optical system. Aperture 102 is driven by aperture drive unit 104 and controls the amount of light incident on image sensor 201, which will be described later. Focus lens 103 is driven by focus lens drive unit 105, and the focal length of the imaging optical system changes depending on the position of focus lens 103. Aperture drive unit 104 and focus lens drive unit 105 are controlled by lens control unit 106, which determines the opening size of aperture 102 and the position of focus lens 103.
[0019] The lens operation unit 107 is a group of input devices that allow the user to make settings related to the operation of the lens unit 10. Specific operations of the lens unit 10 include switching between AF (autofocus) and MF (manual focus) modes, adjusting the position of the focus lens using MF, setting the operating range of the focus lens, and setting the image stabilization mode. When the lens operation unit 107 is operated, the lens control unit 106 performs control according to the operation.
[0020] The lens control unit 106 controls the aperture driving unit 104 and the focus lens driving unit 105 in accordance with control commands and control information received from the camera control unit 250 (described later), and also transmits lens control information to the camera control unit 250.
[0021] Next, a description will be given of the configuration of the camera body 20. The camera body 20 is configured so that it can acquire an image signal from a light beam that has passed through the imaging optical system of the lens unit .
[0022] The image sensor 201 is composed of a CCD or CMOS sensor. A light beam incident from the photographing optical system of the lens unit 10 forms an image on the light receiving surface of the image sensor 201 and is converted into a signal charge corresponding to the amount of incident light by photodiodes provided in pixels arranged on the image sensor 201. The signal charge accumulated in each photodiode is sequentially read out from the image sensor 201 as a voltage signal corresponding to the signal charge, in response to a drive pulse output by a timing generator 214 in accordance with a command from a camera control unit 250.
[0023] The CDS / AGC / AD converter 202 performs correlated double sampling to remove reset noise, gain adjustment, and AD conversion on the imaging signal read out from the image sensor 201. The converter 202 outputs the imaging signal and AF signal that have been subjected to these processes to an image input controller 203 and an AF signal processing unit 204, respectively.
[0024] The image input controller 203 stores the imaging signal output from the converter 202 as an image signal in the SDRAM 209 via the bus 21. The image signal stored in the SDRAM 209 is read out by the display control unit 205 via the bus 21 and displayed on the display unit 206. In a recording mode in which the image signal is recorded, the image signal stored in the SDRAM 209 is recorded by the recording medium control unit 207 in a recording medium 208 such as a semiconductor memory.
[0025] The ROM 210 stores control programs and processing programs executed by the camera control unit 250, as well as various data required for executing these programs. The flash ROM 211 stores various setting information related to the operation of the camera 20 set by the user.
[0026] The joint detection unit 251 in the camera control unit 250 detects the coordinates of each joint of the target person in the imaging signal as shown in Fig. 4(a) input from the image input controller 203, as shown in Fig. 4(b) 310, estimates the connection relationships, and groups them as one person. The detection and connection method is a well-known method using a neural network or other machine learning model, and will not be described in detail, but will be partially described below.
[0027] The object detection unit 252 detects and stores a specific object in an imaging signal such as that shown in Fig. 4(a) input from the image input controller 203 as a bounding box with coordinates and length and width as shown in Fig. 4(b) 411. The detection method is a well-known method using a neural network or other machine learning model, and will not be described in detail, but will be partially described later.
[0028] The behavior detection unit 253 inputs detection information such as the coordinates and size in the image of each joint and object output from the joint detection unit 251 and object detection unit 252, and infers the behavior of the target person. Examples of behavior include suspicious behavior that is desired to be detected by a surveillance camera, or the behavior of kicking a ball in soccer filming. Detection methods use neural networks, decision trees, other machine learning models, or simple rule-based Algorithms. These learning and implementation methods use known methods and will not be described in detail.
[0029] Furthermore, there may be a plurality of behavior detection units 253 to correspond to a plurality of modes for photographing a plurality of sporting events, respectively, or to correspond to a plurality of photographing environments.
[0030] The camera control unit 250 controls each unit within the camera body 20 while exchanging information with them. Furthermore, in response to input from the camera operation unit 213 based on user operation, the camera control unit 250 performs various processes corresponding to user operation, such as turning the power on / off, changing various settings, capturing an image, AF processing, and playing back a recorded image. Furthermore, the camera control unit 250 transmits control commands for the lens unit 10 (lens control unit 106) and information about the camera body 20 to the lens control unit 106, and obtains information about the lens unit 10 from the lens control unit 106. The camera control unit 250 is configured by a microcomputer, and controls the entire camera system including the interchangeable lens 10 by executing a computer program stored in the ROM 210.
[0031] The following describes the processing performed by the camera body 20. The camera control unit 250 performs the following processing in accordance with an imaging processing program, which is a computer program. Figure 2 is a flowchart showing the procedure for detecting the behavior of a person on the camera body 20. "S" represents a step.
[0032] In S201, the joints and their positions in the human body are detected by the joint detection unit 251, which includes a learning model trained by machine learning. In S202, the process from S204 to S205 is repeated for the number of joint connections in the human body. In this embodiment, the process includes 13 joints: the neck, head, left and right shoulders, left and right elbows, left and right wrists, left and right hip joints, left and right knees, and left and right ankles.
[0033] In S203, the joints are connected and the result of person detection is output to the behavior detection unit 253. In S205, the object detection unit 252, which includes a learning model created by machine learning, detects an object and outputs the result of detection to the behavior detection unit 253. In S206, the joints of the person detected up to S204 and the coordinates of the object detected in S205, etc. are input to the behavior detection unit 253, and the probability of a specific behavior is inferred.
[0034] In S207, it is confirmed whether the behavior detection of all people on the screen has been completed. For example, if two heads are detected, it is assumed that there are two people on the screen and the process is repeated twice. If there are people whose behavior detection has not yet been completed, the process proceeds to S202; if not, the process proceeds to S208.
[0035] In S208, the behavior probabilities of all people on the screen are displayed. Depending on the purpose of use, the person with the highest behavior probability may be selected from all people, or people with a behavior probability above a threshold may be listed.
[0036] In this case, if different behavior detection units 253 are selected for different people and the scale and bias of the behavior probabilities are different, normalization or the like is used to enable direct comparison.
[0037] This completes the process of detecting behavior.
[0038] The operations of the joint detection unit 251 and the object detection unit 252 will be described with reference to FIGS.
[0039] For posture estimation, 14 points are detected, as shown in Figure 3(a): head, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles.
[0040] In this embodiment, the illustration becomes too complicated and difficult to see, so as shown in FIG. 3(b), six joints are used: the head, center of gravity, left and right hands, and left and right feet. However, in practice, the number of joints may be increased to 14 or the like, or may be reduced in consideration of the amount of calculation required for detection and connection.
[0041] For example, suppose an image like that shown in Figure 3(c) is input to the joint detection unit 251. Then, the head response will appear in the head peak map as shown in Figure 3(d). Similarly, the center of gravity will appear as shown in Figure 3(e), the right hand as shown in Figure 3(f), the left hand as shown in Figure 3(g), the right foot as shown in Figure 3(h), and the left foot as shown in Figure 3(i).
[0042] The paths connecting the joints require five connections, as shown in FIG. 3(b), between the head and center of gravity 301, the right hand and center of gravity 302, the left hand and center of gravity 303, the right foot and center of gravity 253, and the left foot and center of gravity 305.
[0043] For example, if there is a two-point head response as shown in FIG. 3(d), in this case there will be two people on the screen, so the loop of S202 to S204 in FIG. 2 will be processed twice.
[0044] In the first loop, when connecting the head 311 to the center of gravity, there are two candidates, 312 and 313 in Figure 3(e). By selecting 312 using some known algorithm, such as the joint with the shorter distance, one connection between the head and the center of gravity is completed. Similarly, in the case of six joints including the center of gravity, five connections are repeated and the connection is completed as shown in Figure 3(b).
[0045] 3(d) to 3(i). The joints detected by the joint detection unit 251 may be detected based on whether or not a peak is detected, as shown in FIG. 3(d) 311, or a certain threshold may be set so that if the peak value is equal to or greater than the threshold (satisfies the criteria), it is detected, and if it is equal to or less than the threshold, it is not detected (missing, does not satisfy the criteria).
[0046] Also, assume that one person has six joints. Assume that both feet of the person in Figure 4(a) are detected, and that the left foot 401 is detected as 402 in the peak map for detecting the left foot, Figure 4(e). Assume that the left foot 404 of the person in the next frame, Figure 4(c), is hidden by clothing, and the peak value is low below the threshold, as shown in Figure 4(f) 403 of the peak map for detecting the left foot, and therefore it cannot be detected. In this case, as shown in Figure 4(d), five joints may be connected without the coordinate of the left foot, and the left foot coordinate may be left as a missing value, and joint information and ball information may be input to the action detection unit 253. However, action detection may also be performed by connecting six joints using the coordinate of the left foot 402 that was detected in the previous frame, Figure 4(a). The number of frames to go back will be described later.
[0047] In this case, although the example in Figure 4 shows one person on the screen, if there are multiple people, care must be taken not to connect the joints of different people in Figure 4(a) and (c). For example, if the number of frames per second is sufficiently large and the time difference between frames is small, the head movement distance between frames should not be very large, so after confirming that it is the same person between frames, the coordinates of the joints that were not detected can be used.
[0048] As with joints, the bounding boxes detected by the object detection unit 252 in Figure 4(a) are detected as shown in Figure 4(b) 411, but these may also be detected or not, or a certain threshold may be set and if the detection value is above the threshold it is detected, and if it is below the threshold it is not detected.
[0049] Similarly, suppose that the coordinates and size of the artist ball 404 in Figure 4(a) are detected using a bounding box as shown in Figure 4(b). Assume that the ball 405 in the next frame, Figure 4(c), is half outside the frame and cannot be detected. In this case, the joint information and ball information may be input to the behavior detection unit 253 as missing values without the ball, but behavior may also be detected using the coordinates and size 411 of the ball detected in the previous frame, Figure 4(a). How many frames back can be referenced will be described later.
[0050] In most ball games, only one ball is used during a game, so there is no need for frame-to-frame matching like with human joints. However, in practice, there may be multiple balls lying around, so you can perform frame-to-frame matching in the same way and then use the coordinates of the previous frame.
[0051] The process when no joints or objects are detected in the first embodiment will be described using the table in FIG. 5(a).
[0052] Naturally, it is advantageous for the behavior detection unit 253 to have all the information when inferring the behavior of the target person from the information on each joint and object output from the joint detection unit 251 and the object detection unit 252. For example, having all the joint information on the entire body makes it easier to infer the behavior the person is about to take, and the positions of balls, shuttlecocks, rackets, etc. in sports provide important clues.
[0053] However, detection may not be possible due to the accuracy of image analysis or occlusion. In such cases, it is possible to use information from past frames.
[0054] In the case of a person, if the coordinate information of a joint that could not be detected by the joint detection unit 251 is interpolated from a past frame, and if the human body is moving significantly, referencing a past frame that is too far in time will result in a distorted and unnatural posture for the human body. In other words, when trying to infer behavior from a person's posture, detection accuracy is likely to decrease.
[0055] In the case of an object, the coordinate and size information of an object that could not be detected by the object detection unit 252 is interpolated from past frames. If the object is a ball and the action detection target is a sports action, even if the position is slightly off, there is no problem such as the human body becoming in an unnatural posture. Therefore, there is no significant problem even if past frames from a certain distance in time are referenced.
[0056] In the table of Figure 5(a), frame number 501 is the number of each frame in the video, and is assigned in increments over time. 502 indicates the detection state of one joint A (for simplicity) among the joints detected by joint detection unit 251 of a specific person in the image. The processing is similar even if there are multiple joints. 503 indicates the object detection state output from object detector 252.
[0057] In this case, the joint detection unit 251 uses the coordinates of the joints that are missing up to n frames before for interpolation, where n = 1. Also, the object detection unit 252 uses the coordinates and sizes of the objects that are missing up to m frames before for interpolation, where m = 2. Here, m > n.
[0058] An example will be explained using Figure 5(a).
[0059] In frame number 1, both joint A and the object have been detected, so the joint and object information is input to the behavior detection unit 253 as is, and the behavior is inferred.
[0060] Neither joint A nor the object is detected in frame number 2. Looking back at the frames where detection was possible, both joint A and the object were detected in the previous frame, frame number 1. The coordinates of joint A and the coordinates and size of the object from frame number 1 are input to the behavior detection unit 253, and behavior is inferred.
[0061] In frame number 3, neither joint A nor the object is detected. Looking back at the frames where detection was possible, both joint A and the object were detected in frame number 1, two frames earlier. Since joint A has n=1, this is not adopted and is treated as a missing value. Since the coordinates and size of the object are m=2, those from frame number 1 are adopted and both are input into the behavior detection unit 253, and behavior is inferred.
[0062] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0063] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention.
[0064] As described above, according to the first embodiment of the present invention, when a part of a human body, an associated object, belongings, or sporting goods used in a particular sport is occluded or cannot be detected, it is possible to effectively utilize past detection information while minimizing the deterioration of detection accuracy.
[0065] [Second embodiment] (If the ball coordinate / size difference is greater than the threshold, it will not be adopted) Except for the processing described below, the same configuration and control as in the first embodiment are used. In this embodiment, when detecting human behavior, if a person's joints (points) cannot be detected, past frames can be referenced, and a feature of this embodiment is that it determines whether the past frames can be reliably referenced at this time. Specifically, if there is a change between past frames that exceeds a reference standard, those past frames are not referenced.
[0066] The process when no joints or objects are detected in the second embodiment will be described using the table in FIG. 5(b).
[0067] In the table of Figure 5(b), frame number 511 is the number of each frame in the video, and is assigned in increments over time. 512 indicates (for simplicity) the X coordinate or detection state of a specific object detected by object detection unit 252. This is also true for the Y coordinate and detected size of the object. 513 is the difference between the X coordinate of 512 and the previous frame.
[0068] In this case, the object detection unit 252 uses the coordinates and sizes of the objects that are missing up to m frames before for interpolation, and m = 2. Here again, it is assumed that m>n.
[0069] Here, the specific object is a ball in a ball game. The speed of ball games is known depending on the type of game. The speed of a volleyball spike is said to be 150 km / h. By taking this into account the focal length of the camera lens, the shooting distance, the number of pixels in the image, the frame rate, and leaving some margin, we can calculate the maximum possible moving speed in the image. Here, we will use 100 pixels. A ball moving faster than this may be a false detection.
[0070] An example will be explained using Figure 5(b).
[0071] In frame number 1, an object is detected and the X coordinate is also output. Since this is the first frame, there is no difference. As with other embodiments, the joint coordinates are also detected by the joint detection unit 251, which will not be described in detail. The joint and object information is input as is to the behavior detection unit 253, and behavior is inferred.
[0072] In frame number 2, an object is detected and the X coordinate is also output. The difference of 10 between frames 1 and 2 is recorded as the difference in the object's X coordinate. For this frame as well, the joint and object information is input as is to the behavior detection unit 253, and behavior is inferred.
[0073] In frame number 3, an object is detected and the X coordinate is also output. The difference 490 between frames 2 and 3 is recorded as the difference in the object's X coordinate. For this frame as well, the joint and object information is input as is to the behavior detection unit 253, and behavior is inferred.
[0074] In frame number 4, no object is detected and no X coordinate is output. The object's X coordinate difference is also not calculated. Therefore, when the X coordinate difference of the object in frame number 3 is checked, it is 490. As mentioned above, the maximum movement speed of the ball is used as the difference threshold, so the X coordinate difference threshold is 100, and therefore the X coordinate difference in frame number 3 is greater than the X coordinate difference threshold. Therefore, the X coordinate of 600 in frame number 3 is not used as the position coordinate of the ball in frame number 4. Therefore, in frame number 4, the coordinates of the object (ball) are output to the behavior detection unit 253 as missing values, and the behavior detection unit 253 infers the person's behavior using the joint point information.
[0075] The above process checks the possibility of false detection of an object based on the moving speed. If there is a possibility of false detection, it is not used to interpolate the missing value when the object is not detected. Ultimately, a value that may be falsely detected in frame number 3 is used, but this has the effect of preventing the influence of false detection from spreading to multiple frames. This is even more effective if the behavior detection unit 253 includes filtering processing in the time direction, etc.
[0076] Furthermore, in this embodiment, the difference from the past frame is simply compared with a threshold value, but the position of the current frame may be predicted from the past frame, and the difference between the predicted position and the detected coordinates may be used instead of the X coordinate difference 513 of the object for processing.
[0077] In this embodiment, the constants are m, n, the object's X coordinate difference threshold, Y coordinate difference threshold, and diameter difference threshold. Multiple sets of these constants may be provided depending on the camera environment, the action to be inferred, the type of sport, etc.
[0078] As described above, according to the second embodiment of the present invention, when a part of a human body, an object associated with the part, belongings, or sporting goods used in a particular sport is occluded or cannot be detected, it is possible to effectively utilize past detection information while suppressing a decrease in detection accuracy. Furthermore, in this embodiment, it is possible to further suppress a decrease in detection accuracy even when there is a high possibility of false detection of an object.
[0079] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0080] (Other embodiments) The object of the present invention can also be achieved as follows: A storage medium storing software program code describing procedures for realizing the functions of each of the above-described embodiments is supplied to a system or device, and the computer (or CPU, MPU, etc.) of the system or device reads and executes the program code stored in the storage medium.
[0081] In this case, the program code itself read from the storage medium will realize the novel functions of the present invention, and the storage medium storing the program code and the program will constitute the present invention.
[0082] Furthermore, examples of storage media for supplying the program code include flexible disks, hard disks, optical disks, magneto-optical disks, etc. Also usable are CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-Rs, magnetic tapes, non-volatile memory cards, ROMs, etc.
[0083] The functions of the above-described embodiments are realized by making the computer executable the read program code. Furthermore, the functions of the above-described embodiments may be realized by an operating system (OS) or the like running on the computer performing some or all of the actual processing based on the instructions of the program code.
[0084] The following case is also included: First, program code is read from a storage medium and written into memory on an expansion board inserted into a computer or on an expansion unit connected to the computer. Then, based on the instructions of the program code, a CPU or other device on the expansion board or unit performs some or all of the actual processing.
[0085] The disclosure of this embodiment includes the following configurations, methods, and programs.
[0086] (Configuration 1) a first detection means for detecting a first object from an image using a first learning model based on machine learning; a second detection means for detecting a second object different from the first object from an image using a second learning model based on machine learning; a behavior detection means for detecting a specific behavior of a subject based on the detection results of the first and second detection means in a moving image made up of a plurality of input frames, the first detection means is capable of referring to detection results up to n frames (n≧0) before the target frame in order to output the detection result of the first object in the target frame; The image processing device is characterized in that the second detection means is capable of referring to detection results up to m frames (m>n) before the target frame in order to output the detection results of the second object in the target frame.
[0087] (Configuration 2) 2. The image processing device according to claim 1, wherein the first detection means and the second detection means detect specific behavior of the subject in the target frame by interpolating with reference to frames prior to the target frame when the inference result for the input of the target frame does not satisfy a criterion.
[0088] (Configuration 3) 2. The image processing device according to claim 1, wherein the first object is a joint of a person, and the second object is a ball.
[0089] (Configuration 4) The image processing device according to claim 1, characterized in that the second detection means does not refer to a past frame if at least one of the difference in coordinates or size of the detected second object between past frames prior to the target frame is greater than a reference value.
[0090] (Configuration 5) 2. The image processing device according to claim 1, wherein the first and second detection means have a plurality of modes suited to a plurality of different sports, and the values of n and m are switched for each mode.
[0091] (Configuration 6) an image processing device according to any one of claims 1 to 5; imaging means for capturing the image; An imaging device comprising:
[0092] (Method 1) a first detection step of detecting a first object from an image using a first learning model based on machine learning; a second detection step of detecting a second object different from the first object from the image using a second learning model based on machine learning; a behavior detection step of detecting a specific behavior of a subject based on the detection results of the first and second detection means in a moving image made up of a plurality of input frames, In the first detection step, in order to output the detection result of the first object in the target frame, it is possible to refer to the detection result of up to n frames (n≧0) before the target frame, An image processing method characterized in that in the second detection step, the detection results of the second object in the target frame can be referenced up to m frames before (m>n) in the target frame in order to output the detection results.
[0093] (Program 1) A computer-executable program in which the steps of the image processing method according to claim 7 are described. [Explanation of symbols]
[0094] 10 Lens unit 106 Lens control unit 107 Lens operation section 20 Camera 201 Image sensor 203 Image Input Controller 204 AF signal processing unit 205 Display control unit 206 Display section 207 Recording medium control unit 208 Recording Media 250 Camera control unit 251 Joint detection unit 252 Object detection unit 253 Behavior Detection Unit
Claims
1. a first detection means for detecting a first object from an image using a first learning model based on machine learning; a second detection means for detecting a second object different from the first object from an image using a second learning model based on machine learning; a behavior detection means for detecting a specific behavior of a subject based on the detection results of the first and second detection means in a moving image made up of a plurality of input frames, the first detection means is capable of referring to detection results up to n frames (n≧0) before the target frame in order to output the detection result of the first object in the target frame, An image processing device characterized in that the second detection means is capable of referring to detection results up to m frames (m>n) before the target frame in order to output the detection results of the second object in the target frame.
2. The image processing device according to claim 1, characterized in that, when the detection result for the input of the target frame does not satisfy a criterion, the first detection means and the second detection means detect specific behavior of the subject in the target frame by interpolating by referring to frames prior to the target frame.
3. 2. The image processing apparatus according to claim 1, wherein the first object is a joint of a person, and the second object is a ball.
4. 2. The image processing device according to claim 1, wherein the second detection means does not refer to a past frame if at least one of the difference in coordinates and the difference in size of the detected second object between past frames prior to the target frame is greater than a reference value.
5. 2. The image processing device according to claim 1, wherein the first and second detection means have a plurality of modes suited to a plurality of different sports, and the values of n and m are switched for each mode.
6. An image processing device according to any one of claims 1 to 5; imaging means for capturing the image; An imaging device comprising:
7. a first detection step of detecting a first object from an image using a first learning model based on machine learning; a second detection step of detecting a second object different from the first object from the image using a second learning model based on machine learning; a behavior detection step of detecting a specific behavior of a subject based on the detection results of the first and second detection means in a moving image made up of a plurality of input frames, In the first detection step, in order to output a detection result of the first object in the target frame, it is possible to refer to a detection result up to n frames before the target frame (n≧0), An image processing method characterized in that in the second detection step, the detection results of the second object in the target frame can be referenced up to m frames (m>n) before the target frame in order to output the detection results.
8. A computer-executable program in which the steps of the image processing method according to claim 7 are described.
Citation Information
Patent Citations
Image processing apparatus, method thereof, and program
JP2018181273A
Processing device and program
JP2019179288A
Determination system and determination method
JP2020031406A
Ball game image analysis device, ball game image analysis system, ball game image analysis method, and computer program
JP2021145702A
Tactile providing system and tactile providing device
JP2022094597A