Image processing program, image processing apparatus, and image processing method
The image processing program uses spatial and shape-based thresholding to extract bounding boxes around target objects in dynamic scenes, enhancing object identification and skeleton recognition accuracy in crowded environments.
Patent Information
- Application Number
- JP2023550811
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing image processing methods struggle to accurately identify the area related to a moving object in an image when multiple objects are present, particularly in dynamic environments like sports arenas where athletes perform acrobatic movements, as they often fail to distinguish between the target object and background objects.
An image processing program that calculates threshold values based on three-dimensional spatial relationships and object shapes to selectively extract bounding boxes around target objects, using a multi-viewpoint image processing system with synchronized cameras to capture and process images, thereby identifying the target object by determining appropriate size ranges for bounding boxes.
This approach effectively identifies and tracks the target object within a crowded scene, improving the accuracy of subsequent skeleton recognition and enabling efficient online processing by automatically adjusting to changes in camera distance and object posture.
Smart Images

Figure 0007700866000001 
Figure 0007700866000002 
Figure 0007700866000003
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology.
Background Art
[0002] In relation to image processing, a technique for extracting a region that surrounds an object such as a person shown in an image from an image captured by a camera is known. The region surrounding the object may be called a bounding box.
[0003] When the bounding box is set as a region for performing human body skeleton recognition, as an application field thereof, a gymnastics scoring support system composed of skeleton recognition and technique recognition has attracted attention (see, for example, Non-Patent Document 1). As technologies related to sports, an Adaptive Appearance Model for three-dimensional robust object tracking, a camera pose estimation technique for a sports arena, and a technique for specifying the position of a sports arena from an image are also known (see, for example, Non-Patent Document 2, Non-Patent Document 3, and Non-Patent Document 4). In addition, a technique for specifying the position of a person using a panoramic video is also known (see, for example, Non-Patent Document 5).
[0004] On the other hand, as another application example of the bounding box, an image monitoring device that detects a target object from a monitoring image obtained by imaging a monitoring space is also known (see, for example, Patent Document 1). There is also known an object recognition system that can stably obtain a highly accurate background under any circumstances and does not fail in detection even under the influence of illumination variation, shielding, etc. (see, for example, Patent Document 2). A system for object identification and behavior characterization using video analysis is also known (see, for example, Patent Document 3).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Document
[0006]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0007] When photographing an athlete performing in a sports arena with a camera and scoring the athlete's performance using the captured image, after identifying the area in the image where the athlete to be scored appears, the method of recognizing the athlete's skeleton is effective when analyzing human movements with high precision such as scoring. However, when a plurality of people including the athlete to be scored appear in the image, it is difficult to identify the area where the athlete to be scored appears.
[0008] Note that such a problem occurs not only when scoring an athlete's performance using an image, but also when identifying an area related to various objects appearing in the image.
[0009] In one aspect, the present invention aims to identify an area related to a moving object in an image in which a plurality of objects appear. The movement here includes not only a general standing posture but also an acrobatic posture such as a backflip in gymnastics.
Means for Solving the Problems
[0010] In one proposal, the image processing program causes a computer to execute the following processing.
[0011] The computer From an image captured by an imaging device arranged in a predetermined space, extract regions surrounding each of a plurality of objects, which are a plurality of persons within the predetermined space or a plurality of predetermined tools used in a predetermined competition. Based on the distance between the three-dimensional coordinates indicating the vertices of a target space, which is a space preset within the predetermined space, and the three-dimensional coordinates indicating the position of the imaging device, and information representing the three-dimensional shape of a specific target object, which is a person or a predetermined tool existing within the target space among the plurality of objects, calculate a first threshold value indicating the lower limit of the size of the region surrounding the specific target object and a second threshold value indicating the upper limit of the size of the region surrounding the specific target object. From among the regions surrounding each of the plurality of objects, select a region that is larger than the first threshold value and smaller than the second threshold value as the region surrounding the specific target object. Extract.
Effects of the Invention
[0012] According to one aspect, in an image in which a plurality of objects appear, an area related to a moving object can be identified.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Embodiments for Carrying Out the Invention
[0014] Hereinafter, embodiments will be described in detail with reference to the drawings.
[0015] When recognizing the skeleton of an athlete using an image taken in a sports arena and scoring the performance from the result, it is effective to identify the athlete performing the performance within the target space from among the many people shown in the image. The target space is a predetermined space in a three-dimensional space where the athlete to be scored performs the performance.
[0016] In this case, among the many people shown in the image, the people shown in the background are excluded, and the athlete performing the performance within the target space is identified. The people shown in the background include other athletes, coaches, spectators, referees, cameramen, etc. who are not performing the performance within the target space. Hereinafter, the athlete performing the performance within the target space may be referred to as the athlete within the target space.
[0017] When identifying the athlete within the target space using an information processing device (computer), the information processing device extracts the bounding box surrounding each person from the image. Then, the information processing device extracts the bounding box of the athlete within the target space by excluding the bounding boxes of the people shown in the background or foreground among the extracted bounding boxes. Thereby, the accuracy of subsequent skeleton recognition can be improved.
[0018] As a method for excluding the bounding boxes of the people shown in the background, for example, the methods of Comparative Example 1 and Comparative Example 2 can be mentioned.
[0019] FIG. 1 is a flowchart showing the method of Comparative Example 1. First, the user selects a plurality of bounding boxes within the target area from among the many bounding boxes extracted from the image, and estimates the bounding box range from the size of the selected bounding boxes (step 101). The target area represents the area within the image corresponding to the target space, and the bounding box range represents the range of the size of the bounding box surrounding the athlete within the target space.
[0020] Next, the information processing device excludes the bounding boxes having sizes that do not belong to the estimated bounding box range among the bounding boxes extracted from the image (step 102).
[0021] However, the closer the camera approaches the target space in three-dimensional space, the larger the target area in the image and the bounding box of the player in the target space become, and the farther the camera moves away from the target space, the smaller the target area and the bounding box of the player in the target space become. Therefore, in the method of Comparative Example 1, every time the distance between the camera and the target space is changed, the manual selection of the bounding box and the estimation of the bounding box range are repeated. For this reason, the method of Comparative Example 1 is not efficient and is not suitable for online processing.
[0022] FIG. 2 is a flowchart showing the method of Comparative Example 2. In the method of Comparative Example 2, as shown in Non-Patent Documents 2 to 4, it is premised that a person moves on the ground and the ground is a plane.
[0023] First, the information processing apparatus estimates the position of each person's foot by back-projecting the midpoint of the bottom side of each bounding box extracted from the image into three-dimensional space (step 201). Next, the information processing apparatus excludes the bounding boxes among the bounding boxes extracted from the image for which the estimated foot position does not exist on the ground of the sports arena (step 202).
[0024] According to the method of Comparative Example 2, in sports such as soccer and basketball, it is possible to exclude the bounding boxes of spectators and the like shown in the background. However, in a sport such as gymnastics where a player performs an act in a space away from the floor of the sports arena, the premise that a person moves on the ground does not hold. When the method of Comparative Example 2 is applied to an image of a gymnast, since the position of the gymnast's foot in the target space does not exist on the floor, the bounding box of that gymnast is erroneously excluded.
[0025] FIG. 3 shows a functional configuration example of the image processing apparatus according to the embodiment. The image processing apparatus 301 in FIG. 3 includes an object extraction unit 311, a determination unit 312, and a target object extraction unit 313.
[0026] FIG. 4 is a flowchart showing an example of the image processing performed by the image processing apparatus 301 in FIG. 3. First, the object extraction unit 311 extracts regions (bounding boxes) related to each of a plurality of objects from the image captured by the imaging device (step 401).
[0027] Next, the determination unit 312 determines a threshold value based on the range of the space in which the target object moves and the position of the imaging device (step 402). Then, the target object extraction unit 313 extracts the region (bounding box) related to the target object from among the regions related to each of the plurality of objects based on the comparison result of comparing the size of the region related to each of the plurality of objects with the threshold value (step 403).
[0028] According to the image processing apparatus 301 in FIG. 3, in an image in which a plurality of objects are shown, the region related to the moving target object can be specified.
[0029] FIG. 5 shows a configuration example of a multi-viewpoint image processing system including the image processing apparatus 301 in FIG. 3. The multi-viewpoint image processing system in FIG. 5 includes imaging devices 501-1 to 501-N (N is an integer of 2 or more indicating the number of imaging devices), a synchronization signal generator 502, a capture device 503, and an image processing device 504.
[0030] The multi-viewpoint image processing system captures an image of a gymnastics arena and performs processing to assist in scoring the performance of gymnasts in the target space. The imaging devices 501-1 to 501-N are installed at predetermined locations in the gymnastics arena so as to surround the gymnasts in the target space. The people shown in the image correspond to the objects, and the gymnasts in the target space correspond to the target objects. The target space corresponds to the range of the space in which the target object moves.
[0031] The imaging devices 501-1 to 501-N, the synchronization signal generator 502, the capture device 503, and the image processing device 504 are hardware. The image processing device 504 corresponds to the image processing device 301 in FIG. 3.
[0032] The synchronization signal generator 502 outputs a synchronization signal to the imaging devices 501-1 to 501-N via a signal cable.
[0033] The imaging device 501-i (i = 1 to N) is, for example, a camera having an imaging element such as a CCD (Charge-Coupled Device) or a CMOS (Complementary Metal-Oxide-Semiconductor). Each imaging device 501-i captures an image of the target space in the gymnastics arena in synchronization with the synchronization signal output from the synchronization signal generator 502, and outputs the image to the capture device 503 via a signal cable.
[0034] The capture device 503 aggregates the digital signals of the images output from each of the imaging devices 501-1 to 501-N and outputs them to the image processing device 504. The image output from the capture device 503 includes images at a plurality of synchronized times.
[0035] FIG. 6 shows a functional configuration example of the image processing device 504 in FIG. 5. The image processing device 504 in FIG. 6 includes an acquisition unit 611, an object extraction unit 612, a determination unit 613, a target object extraction unit 614, a skeleton recognition unit 615, an output unit 616, and a storage unit 617. The object extraction unit 612, the determination unit 613, and the target object extraction unit 614 respectively correspond to the object extraction unit 311, the determination unit 312, and the target object extraction unit 313 in FIG. 3.
[0036] The storage unit 617 stores target space information 622 and imaging device information 623. The target space information 622 is information indicating the target space. The target space is set in advance according to the type of the item to be scored (such as pommel horse, rings, parallel bars, etc., and there are 10 events for both men and women) in the case of gymnastics. As the shape of the target space, a rectangular parallelepiped, a combination of a plurality of rectangular parallelepipeds, a prism, a cylinder, etc. can be used. For example, when the shape of the target space is represented by a solid having m vertices, the target space information 622 may include the three-dimensional coordinates (xj, yj, zj) (j = 1 to m) of the m vertices.
[0037] FIG. 7 shows an example of a target space. FIG. 7(a) shows an example of a rectangular parallelepiped target space. When the shape of the target space is a rectangular parallelepiped 701, the target space information 622 includes the three-dimensional coordinates (xj, yj, zj) (j = 1 to 8) of the eight vertices of the rectangular parallelepiped 701.
[0038] FIG. 7(b) shows an example of a target space formed by combining two rectangular parallelepipeds. The shape of the target space in FIG. 7(b) is a shape formed by combining a rectangular parallelepiped 701 and a rectangular parallelepiped 702. In this case, the target space information 622 includes the three-dimensional coordinates (xj, yj, zj) (j = 1 to 8) of the eight vertices of the rectangular parallelepiped 701 and the three-dimensional coordinates (xj, yj, zj) (j = 9 to 16) of the eight vertices of the rectangular parallelepiped 702.
[0039] FIG. 7(c) shows an example of a hexagonal prism target space. When the shape of the target space is a hexagonal prism 703, the target space information 622 includes the three-dimensional coordinates (xj, yj, zj) (j = 1 to 12) of the twelve vertices of the hexagonal prism 703.
[0040] The imaging device information 623 includes the position coordinates (Xi, Yi, Zi) (i = 1 to N) of each imaging device 501-i in the three-dimensional space, the focal length f, and the effective scale (three-dimensional distance per pixel) s. In this example, for simplicity, it is assumed that the focal lengths f of the imaging devices 501-1 to 501-N are the same and the aspect ratio is 1.
[0041] The acquisition unit 611 acquires the video output from the capture device 503, extracts the image 621 at each time from the acquired video, and stores it in the storage unit 617.
[0042] The object extraction unit 612 performs a person detection process to extract the bounding box of each of the plurality of persons shown in the image 621 from the image 621, generates region information 626 indicating the bounding box of each person, and stores it in the storage unit 617. The bounding box of a person corresponds to the region related to the object. The region information 626 includes, for example, the coordinates (xp, yp) of the upper left vertex of the bounding box in the image 621, the width wp of the bounding box, and the height hp of the bounding box.
[0043] FIG. 8 shows an example of a bounding box extracted from the image 621. The bounding box 801 surrounds a gymnast performing a pommel horse routine. In this case, the coordinates of the lower right vertex 812 of the bounding box 801 are represented by (xp + wp, yp + hp) using the coordinates (xp, yp) of the upper left vertex 811 of the bounding box 801, the width wp of the bounding box 801, and the height hp of the bounding box 801.
[0044] FIG. 9 shows an example of a target space and a bounding box in the image 621. FIG. 9(a) shows an example of a target space and a bounding box in a rings routine. The target space 901 is a space in a three-dimensional space where a gymnast performs a rings routine. The bounding box 911 is a bounding box of the gymnast within the target space 901, and the bounding box 912 is a bounding box of a background person.
[0045] FIG. 9(b) shows an example of a target space and a bounding box in a pommel horse routine. The target space 902 is a space in a three-dimensional space where a gymnast performs a pommel horse routine. The bounding box 913 is a bounding box of the gymnast within the target space 902, and the bounding box 914 is a bounding box of a background person.
[0046] The determination unit 613 estimates the size of the gymnast within the target space, generates object information 624 indicating the estimated size of the gymnast, and stores it in the storage unit 617. The size of the gymnast is represented, for example, by a solid that surrounds the gymnast's body without gaps in a three-dimensional space. As the solid representing the size of the gymnast, a rectangular parallelepiped, a combination of a plurality of rectangular parallelepipeds, a prism, a cylinder, etc. can be used. Since the posture of the gymnast performing the routine changes as the gymnast moves, the shape of the solid representing the gymnast also changes as the gymnast moves.
[0047] FIG. 10 shows an example of a three-dimensional object representing a gymnast. In this example, a rectangular parallelepiped is used as the three-dimensional object representing the gymnast. FIGS. 10(a) to 10(e) show examples of rectangular parallelepipeds corresponding to various postures of the gymnast. The height H represents the vertical length of the rectangular parallelepiped, the width Wmax represents the maximum value of the horizontal length of the rectangular parallelepiped as viewed from a plurality of cameras, and the width Wmin represents the minimum value of the horizontal length of the rectangular parallelepiped as viewed from a plurality of cameras.
[0048] As described above, since the shape of the three-dimensional object representing the gymnast varies depending on the posture of the gymnast, the determination unit 613 estimates the size of the gymnast based on the statistical information regarding the occurrence frequency of each of the plurality of shapes of the gymnast's body.
[0049] The statistical information regarding the occurrence frequency of the body shape can be obtained, for example, from motion capture, skeleton recognition by three-dimensional sensing (Non-Patent Document 1), an open dataset, or the like. As the object information 624, for example, a combination of a threshold value T1 indicating the lower limit of Smin and a threshold value T2 indicating the upper limit of Smax can be used. Smin represents the sum of Wmin and H, and Smax represents the sum of Wmax and H.
[0050] FIG. 11 shows an example of the statistical information regarding the occurrence frequency of the body shape. FIG. 11(a) shows an example of the distribution of Smin of the rectangular parallelepipeds representing each of the M postures. The horizontal axis represents the value of Smin (mm), and the vertical axis represents the number of postures represented by the rectangular parallelepipeds corresponding to the value of Smin.
[0051] The determination unit 613 calculates T1 such that in the distribution of FIG. 11(a), the number of postures having Smin less than or equal to T1 is M*(α / 100), and the number of postures having Smin greater than T1 is M*((100-α) / 100). As a result, α% of the postures having Smin close to the minimum value among the M postures are excluded as outliers. α is determined based on experimental results or the like. α may be a numerical value in the range of 0 to 20.
[0052] FIG. 11(b) shows an example of the distribution of Smax of rectangular parallelepipeds representing each of the M postures. The horizontal axis represents the value of Smax (mm), and the vertical axis represents the number of postures represented by the rectangular parallelepiped corresponding to the value of Smax.
[0053] The determination unit 613 calculates T2 such that in the distribution of FIG. 11(b), the number of postures having Smax of T2 or more is M*(α / 100), and the number of postures having Smax smaller than T2 is M*((100 - α) / 100). As a result, α% of the postures having Smax close to the maximum value among the M postures are excluded as outliers.
[0054] By determining T1 and T2 using the statistical information regarding the occurrence frequency of the body shape, even when the posture of the gymnast varies, appropriate object information 624 reflecting the change in the posture can be generated.
[0055] Next, the determination unit 613 determines a threshold indicating the range of the size of the bounding box TB surrounding the gymnast in the target space using the target space information 622, the imaging device information 623, and the object information 624. The bounding box TB corresponds to the region related to the object.
[0056] The determination unit 613 determines, for example, a threshold B1 indicating the lower limit of the size of the bounding box TB and a threshold B2 indicating the upper limit of the size of the bounding box TB in the image 621, and stores the combination of B1 and B2 in the storage unit 617 as the range information 625. B1 is an example of the first threshold, and B2 is an example of the second threshold.
[0057] The determination unit 613 can calculate B1 and B2 in pixel units, for example, by the following equation based on perspective transformation.
[0058] B1 = (f / s)*(T1 / Dmax) (1) B2 = (f / s)*(T2 / Dmin) (2)
[0059] f represents the focal length included in the imaging device information 623, and s represents the three-dimensional distance per pixel included in the imaging device information 623. Therefore, f / s corresponds to the focal length in pixel units. T1 and T2 represent the threshold values included in the object information 624.
[0060] Dmax represents the maximum value of the three-dimensional distance D(i,j) from the imaging device 501-i to the j-th vertex (j = 1 to m) of the target space, and Dmin represents the minimum value of D(i,j). The determination unit 613 can calculate D(i,j) using the position coordinates (Xi, Yi, Zi) of the imaging device 501-i included in the imaging device information 623 and the three-dimensional coordinates (xj, yj, zj) of the j-th vertex included in the target space information 622.
[0061] FIG. 12 shows examples of Dmax and Dmin. FIG. 12(a) shows an example of Dmax and Dmin for the target space 901 in FIG. 9(a). In this case, among the three-dimensional distances D(i,j) (j = 1 to 8) from the imaging device 501-i to the eight vertices of the target space 901, the maximum value is used as Dmax and the minimum value is used as Dmin.
[0062] FIG. 12(b) shows an example of Dmax and Dmin for the target space 902 in FIG. 9(b). In this case, among the three-dimensional distances D(i,j) (j = 1 to 8) from the imaging device 501-i to the eight vertices of the target space 902, the maximum value is used as Dmax and the minimum value is used as Dmin.
[0063] By determining B1 and B2 using the target space information 622, the imaging device information 623, and the object information 624, appropriate range information 625 can be generated according to the three-dimensional distance from the imaging device 501-i to the target space and the size of the gymnast.
[0064] The object extraction unit 614 calculates the size BS of each bounding box using the width wp and height hp included in the region information 626 by the following equation.
[0065] BS = wp + hp (3)
[0066] Next, the object extraction unit 614 compares the BS of each bounding box with B1 and B2 included in the range information 625, respectively. Then, the object extraction unit 614 extracts, as the bounding box TB, one or a plurality of bounding boxes having a BS that is larger than B1 and smaller than B2 from among the plurality of bounding boxes indicated by the region information 626. The object extraction unit 614 stores the region information 626 of the extracted bounding box TB in the storage unit 617 as the object region information 627.
[0067] As an example, the case where the following four bounding boxes are extracted from the image 621 will be described.
[0068] BOX1 wp=53 hp=97 BOX2 wp=46 hp=128 BOX3 wp=475 hp=598 BOX4 wp=102 hp=421
[0069] In this case, according to Equation (3), the BS of each bounding box is calculated as follows.
[0070] BOX1 BS=150 BOX2 BS=174 BOX3 BS=1073 BOX4 BS=523
[0071] When B1 = 245 and B2 = 847, the BS of BOX1 and BOX2 is smaller than B1, and the BS of BOX3 is larger than B2. The BS of BOX4 is larger than B1 and smaller than B2. Therefore, BOX1 to BOX3 are excluded, and only BOX4 is extracted as the bounding box TB.
[0072] By extracting a bounding box having a BS that is larger than B1 and smaller than B2 as the bounding box TB, it is possible to exclude the background and foreground people from among the many people shown in the image and identify the gymnast in the target space.
[0073] Figure 13 shows an example of object extraction processing for extracting the boundary box TB. Figure 13(a) shows an example of a plurality of boundary boxes extracted from the image 621 capturing the performance of the suspension ring. The boundary box 1311 is the boundary box of the gymnast within the target space 1301, and the boundary box 1312 is the boundary box of the assistant who assists the gymnast in the action of hanging on the suspension ring. The other 14 small boundary boxes are the boundary boxes of the background people.
[0074] Figure 13(b) shows an example of the boundary box TB extracted by the object extraction processing for the 16 boundary boxes in Figure 13(a). In this example, the 14 boundary boxes of the background are excluded, and the boundary box 1311 and the boundary box 1312 are extracted as the boundary box TB. Among them, since the boundary box 1312 of the assistant is no longer extracted from the image 621 after the start of the performance, it can be excluded from the targets of subsequent processing by performing person tracking processing.
[0075] Figure 14 shows an example of a plurality of boundary boxes extracted from four images 621 captured by four imaging devices 501-i (i = 1 to 4) at the same time during the performance of the suspension ring.
[0076] Figure 14(a) shows an example of a plurality of boundary boxes extracted from the image 621 captured by the imaging device 501-1. Figure 14(b) shows an example of a plurality of boundary boxes extracted from the image 621 captured by the imaging device 501-2. Figure 14(c) shows an example of a plurality of boundary boxes extracted from the image 621 captured by the imaging device 501-3. Figure 14(d) shows an example of a plurality of boundary boxes extracted from the image 621 captured by the imaging device 501-4.
[0077] The boundary boxes in Figure 14(b) and Figure 14(d) include not only the boundary boxes of each person but also large boundary boxes containing a plurality of people.
[0078] FIG. 15 shows an example of a boundary box TB extracted by the object extraction process for the boundary box in FIG. 14.
[0079] FIG. 15(a) shows an example of a boundary box TB extracted by the object extraction process for the boundary box in FIG. 14(a). FIG. 15(b) shows an example of a boundary box TB extracted by the object extraction process for the boundary box in FIG. 14(b). FIG. 15(c) shows an example of a boundary box TB extracted by the object extraction process for the boundary box in FIG. 14(c). FIG. 15(d) shows an example of a boundary box TB extracted by the object extraction process for the boundary box in FIG. 14(d).
[0080] In any of the images 621, the boundary boxes of the background people are excluded, and the boundary boxes of the gymnasts and assistants in the target space are extracted as the boundary box TB.
[0081] When the type of the performance to be scored is changed, the target space information 622 and the object information 624 are changed according to the type of the performance. When the installation location of the imaging device 501-i is changed, the imaging device information 623 is changed according to the installation location. When the target space information 622, the imaging device information 623, or the object information 624 is changed, the determination unit 613 updates the range information 625 by recalculating the threshold B1 and the threshold B2. The object extraction unit 614 extracts the boundary box TB using the updated range information 625.
[0082] The skeleton recognition unit 615 performs skeleton recognition processing using the images of the people included in the boundary box TB indicated by the object region information 627, thereby recognizing the skeletons of the gymnasts in the target space and generating the three-dimensional coordinates of the recognized skeletons. The skeleton recognition processing includes person tracking processing, two-dimensional pose estimation processing, three-dimensional pose estimation processing, and smoothing processing. As the skeleton recognition processing, for example, the processing described in Non-Patent Document 1 can be used.
[0083] The output unit 616 outputs the three-dimensional coordinates of the skeleton of the gymnast. By recognizing the techniques of the performance from the temporal changes in the three-dimensional coordinates of the skeleton, it is possible to assist in scoring the performance performed by the gymnast.
[0084] According to the multi-viewpoint image processing system of FIG. 5, in an image in which a plurality of people in the gymnastics arena are shown, it is possible to specify the bounding box of the gymnast performing the performance within the target space.
[0085] Even when the distance between the imaging device 501-i and the target space is changed, since the determination of the threshold value B1 and the threshold value B2 and the object extraction process are automatically performed, it is possible to efficiently specify the bounding box of the gymnast within the target space. Therefore, image processing suitable for online processing is realized.
[0086] In addition, since the premise that the person moves on the ground is not used, even when the position of the foot of the gymnast in the target space does not exist on the floor surface, the bounding box of the gymnast is not excluded.
[0087] FIG. 16 is a flowchart showing an example of image processing performed by the image processing device 504 of FIG. 6. First, the acquisition unit 611 acquires the video output from the capture device 503 and extracts the image 621 at each time from the acquired video (step 1601).
[0088] Next, the object extraction unit 612 performs a person detection process to extract the bounding box of each of the plurality of people from the image 621 and generates region information 626 indicating the bounding box of each person (step 1602).
[0089] Next, the determination unit 613 estimates the size of the gymnast in the target space and generates object information 624 indicating the estimated size of the gymnast (step 1603). Next, the determination unit 613 uses the target space information 622, the imaging device information 623, and the object information 624 to determine a threshold value B1 and a threshold value B2 indicating the range of the size of the bounding box TB surrounding the gymnast in the target space. Then, the determination unit 613 generates range information 625 including B1 and B2 (step 1604).
[0090] Next, the object extraction unit 614 compares the size BS of each bounding box with B1 and B2 included in the range information 625 respectively, and extracts a bounding box having a BS larger than B1 and smaller than B2 as the bounding box TB. Then, the object extraction unit 614 generates object region information 627 indicating the bounding box TB (step 1605).
[0091] Next, the skeleton recognition unit 615 performs skeleton recognition processing using the image of the person included in the bounding box TB to generate the three-dimensional coordinates of the skeleton of the gymnast in the target space (step 1606). Then, the output unit 616 outputs the three-dimensional coordinates of the skeleton of the gymnast (step 1607).
[0092] The multi-viewpoint image processing system in FIG. 5 can be applied not only to the process of specifying the bounding box of the gymnast shown in the image of the gymnastics arena, but also to the process of specifying the bounding box of the object shown in various images.
[0093] The applicable fields may be support for scoring performances in figure skating, detection of the postures of performers in various events, or support for scoring in basketball practice.
[0094] FIG. 17 shows an example of object extraction processing for extracting the bounding boxes of performers from an image 621 that captures the stage of an event. The people shown in the image 621 correspond to objects, and the performers who perform on the stage correspond to the objects. In this example, using the image included in the extracted bounding boxes of the performers, the pose of the performer is detected, and a 3D avatar of the performer is displayed on the screen on the stage.
[0095] FIG. 17(a) shows an example of the bounding boxes of each of a plurality of people extracted from the image 621 that captures the stage. The bounding boxes 1714 to 1719 are the bounding boxes of the performers who perform on the stage. The bounding boxes 1711 to 1713 and the bounding boxes 1720 to 1723 are the bounding boxes of the background audience.
[0096] FIG. 17(b) shows an example of the bounding boxes of the performers extracted by the object extraction processing for the 13 bounding boxes in FIG. 17(a). In this example, the seven bounding boxes of the background are excluded, and the bounding boxes 1714 to 1719 are extracted as the bounding boxes of the performers.
[0097] FIG. 18 shows an example of object extraction processing for extracting the bounding box of a basketball that a player is dribbling from an image 621 that captures a basketball practice scene. The basketball shown in the image 621 corresponds to an object, and the basketball that the player is dribbling corresponds to the object. In this example, using the image included in the extracted bounding box of the basketball, scoring of the dribble is supported.
[0098] FIG. 18(a) shows an example of the bounding boxes of each of a plurality of basketballs extracted from the image 621 that captures a basketball practice scene. The bounding boxes 1811 and 1812 are the bounding boxes of the target basketball that a player in the target space 1801 is dribbling. The bounding boxes 1813 and 1814 are the bounding boxes of the basketballs held by other players.
[0099] In this example, a decagonal prism space including the player dribbling the target basketball is used as the target space 1801. Therefore, the three-dimensional coordinates of each vertex of the target space 1801 are dynamically updated as the player moves. The position of the player dribbling the target basketball can be estimated, for example, by the method described in Non-Patent Document 5. The size of the target basketball is known.
[0100] FIG. 18(b) shows an example of the bounding box of the target basketball extracted by the object extraction process for the four bounding boxes in FIG. 18(a). In this example, the bounding box 1813 and the bounding box 1814 are excluded, and the bounding box 1811 and the bounding box 1812 are extracted as the bounding boxes of the target basketball.
[0101] The configuration of the image processing apparatus 301 in FIG. 3 is merely an example, and some components may be omitted or changed according to the use or conditions of the image processing apparatus 301.
[0102] The configuration of the multi-viewpoint image processing system in FIG. 5 is merely an example, and some components may be omitted or changed according to the use or conditions of the multi-viewpoint image processing system. The configuration of the image processing apparatus 504 in FIG. 6 is merely an example, and some components may be omitted or changed according to the use or conditions of the multi-viewpoint image processing system.
[0103] For example, in the image processing apparatus 504 in FIG. 6, if the image 621 is previously stored in the storage unit 617, the acquisition unit 611 can be omitted. If there is no need to perform the skeleton recognition process, the skeleton recognition unit 615 can be omitted.
[0104] The flowcharts of FIGS. 1, 2, 4, and 16 are merely examples, and some processes may be omitted or changed according to the configuration or conditions of the information processing apparatus or the image processing apparatus. For example, in the image processing of FIG. 16, if the image 621 is previously stored in the storage unit 617, the process of step 1601 can be omitted. If there is no need to perform the skeleton recognition process, the process of step 1606 can be omitted.
[0105] The target spaces shown in FIGS. 7, 9, 12, 13, and 18 are merely examples, and the image processing apparatus 504 may perform image processing using a target space of another shape. The bounding boxes shown in FIGS. 8, 9, 13 to 15, 17, and 18 are merely examples, and the position and size of the bounding box change according to the image 621. The image processing apparatus 504 may perform image processing using a region of another shape instead of the rectangular bounding box.
[0106] The solid shown in FIG. 10 is merely an example, and the solid representing the gymnast changes according to the type of performance. The statistical information shown in FIG. 11 is merely an example, and the statistical information regarding the occurrence frequency of the body shape changes according to the type of performance.
[0107] Smin does not necessarily have to be the sum of Wmin and H, and Smax does not necessarily have to be the sum of Wmax and H. The image processing apparatus 504 may calculate T1 and T2 using other indicators as Smin and Smax.
[0108] Equations (1) to (3) are merely examples, and the image processing apparatus 504 may calculate B1, B2, and BS using other calculation formulas.
[0109] FIG. 19 shows a hardware configuration example of an information processing apparatus used as the image processing apparatus 301 in FIG. 3 and the image processing apparatus 504 in FIG. 6. The information processing apparatus in FIG. 19 includes a CPU (Central Processing Unit) 1901, a memory 1902, an input device 1903, an output device 1904, an auxiliary storage device 1905, a media drive device 1906, and a network connection device 1907. These components are hardware and are connected to each other by a bus 1908. The capture device 503 in FIG. 5 may be connected to the bus 1908.
[0110] The memory 1902 is, for example, a semiconductor memory such as a ROM (Read Only Memory), a RAM (Random Access Memory), or a flash memory, and stores programs and data used for processing. The memory 1902 may operate as the storage unit 617 in FIG. 6.
[0111] The CPU 1901 (processor) operates as the object extraction unit 311, the determination unit 312, and the target object extraction unit 313 in FIG. 3, for example, by executing a program using the memory 1902. The CPU 1901 also operates as the acquisition unit 611, the object extraction unit 612, the determination unit 613, the target object extraction unit 614, and the skeleton recognition unit 615 in FIG. 6 by executing a program using the memory 1902.
[0112] The input device 1903 is, for example, a keyboard, a pointing device, etc., and is used for inputting instructions or information from a user or an operator. The output device 1904 is, for example, a display device, a printer, a speaker, etc., and is used for outputting inquiries or processing results to a user or an operator. The output device 1904 may operate as the output unit 616 in FIG. 6.
[0113] The auxiliary storage device 1905 is, for example, a magnetic disk device, an optical disk device, a magneto-optical disk device, a tape device, or the like. The auxiliary storage device 1905 may be a hard disk drive or an SSD (Solid State Drive). The information processing device can store programs and data in the auxiliary storage device 1905 and load them into the memory 1902 for use. The auxiliary storage device 1905 may operate as the storage unit 617 in FIG. 6.
[0114] The medium drive device 1906 drives the portable recording medium 1909 and accesses the recorded content thereon. The portable recording medium 1909 is a memory device, a flexible disk, an optical disk, a magneto-optical disk, or the like. The portable recording medium 1909 may be a CD-ROM (Compact Disk Read Only Memory), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, or the like. A user or an operator can store programs and data in the portable recording medium 1909 and load them into the memory 1902 for use.
[0115] As described above, the computer-readable recording medium for storing the programs and data used in the processing is a physical (non-transitory) recording medium such as the memory 1902, the auxiliary storage device 1905, or the portable recording medium 1909.
[0116] The network connection device 1907 is a communication interface circuit that is connected to a communication network such as a LAN (Local Area Network) or a WAN (Wide Area Network) and performs data conversion associated with communication. The information processing device can receive programs and data from an external device via the network connection device 1907 and load them into the memory 1902 for use. The network connection device 1907 may operate as the output unit 616 in FIG. 6.
[0117] Note that the information processing apparatus does not necessarily need to include all the components shown in FIG. 19, and it is also possible to omit or change some of the components according to the application or conditions. For example, if the interface with the user or operator is unnecessary, the input device 1903 and the output device 1904 may be omitted. If the information processing apparatus does not use the portable recording medium 1909 or the communication network, the medium drive device 1906 or the network connection device 1907 may be omitted.
[0118] Although the disclosed embodiments and their advantages have been described in detail, those skilled in the art will be able to make various changes, additions, and omissions without departing from the scope of the invention clearly described in the claims.
Claims
Extract, from an image captured by an imaging device arranged in a predetermined space, regions surrounding respective ones of a plurality of objects, which are a plurality of persons or a plurality of predetermined tools used in a predetermined competition, within the predetermined space. Based on the distance between the three-dimensional coordinates indicating the vertices of a target space, which is a space preset within the predetermined space, and the three-dimensional coordinates indicating the position of the imaging device, and information representing the shape of a solid representing a specific target object, which is a person or a predetermined tool existing within the target space among the plurality of objects, calculate a first threshold value indicating the lower limit of the size of the region surrounding the specific target object and a second threshold value indicating the upper limit of the size of the region surrounding the specific target object. Extract, as the region surrounding the specific target object, a region that is larger than the first threshold value and smaller than the second threshold value from among the regions surrounding the respective ones of the plurality of objects. An image processing program for causing a computer to execute the processing. The image processing program according to claim 1, wherein the information representing the shape of the solid is a threshold value of the dimensions of the solid.
3. The shape representing the specific target object changes according to the shape of the specific target object. The image processing program further causes the computer to execute a process of calculating the threshold value of the dimensions of the solid using statistical information regarding the occurrence frequency of each of a plurality of shapes of the specific target object, the image processing program according to claim 2. The image processing program according to any one of claims 1 to 3, wherein the process of calculating the first threshold value and the second threshold value is based on perspective transformation. Object extraction for extracting, from an image captured by an imaging device arranged in a predetermined space, regions surrounding respective ones of a plurality of objects, which are a plurality of persons or a plurality of predetermined tools used in a predetermined competition, within the predetermined space. A determination unit that calculates a first threshold value indicating the lower limit of the size of the region surrounding the specific target object and a second threshold value indicating the upper limit of the size of the region surrounding the specific target object based on the distance between the three-dimensional coordinates indicating the vertices of a target space, which is a space preset within the predetermined space, and the three-dimensional coordinates indicating the position of the imaging device, and information representing the shape of a solid representing a specific target object, which is a person or a predetermined tool existing within the target space among the plurality of objects. An image processing apparatus, comprising: an object extraction unit that extracts, as a region surrounding the object of interest, a region that is larger than the first threshold and smaller than the second threshold from among the regions surrounding each of the plurality of objects.
6. Extract regions surrounding each of a plurality of objects, which are a plurality of persons or a plurality of predetermined tools used in a predetermined competition, within a predetermined space, from an image captured by an imaging device disposed in the predetermined space. Based on the distance between the three-dimensional coordinates indicating the vertices of a target space, which is a space preset within the predetermined space, and the three-dimensional coordinates indicating the position of the imaging device, and information representing the shape of a solid representing an object of interest, which is a person or a predetermined tool existing within the target space among the plurality of objects, calculate a first threshold indicating the lower limit of the size of the region surrounding the object of interest and a second threshold indicating the upper limit of the size of the region surrounding the object of interest. Extract, as a region surrounding the object of interest, a region that is larger than the first threshold and smaller than the second threshold from among the regions surrounding each of the plurality of objects. An image processing method, characterized in that a computer executes the processing.
Citation Information
Patent Citations
Object detection system
JP2005135014A
Image monitoring device
JP2010045501A
Object recognition system, monitoring system using the same, and watching system
JP2011209794A
Image processing device, image processing method, and program
JP2019159739A
Region extraction device and program
JP2020160812A