Image processing apparatus, image processing method, and program
The described method addresses the challenge of accurately detecting a person's fallen state in images by using a combination of person detection and fall determination modules, which effectively handle reflections and reduce computational load, thereby enhancing detection accuracy and efficiency.
Patent Information
- Application Number
- JP2023203957
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-12
AI Technical Summary
Existing technologies face challenges in accurately detecting the fallen state of a person in images, particularly when the person is reflected on a floor surface or when shadows cause erroneous detection, leading to increased computational load.
The proposed solution involves a person detection module to identify objects representing people in images and a fall determination module that assesses the fall state of detected objects based on similarity calculations using intermediate feature maps from deep neural networks. This approach selectively removes objects likely to be reflections from the detection target, thereby reducing computational load.
This method effectively suppresses erroneous detections of reflected images and reduces the computational burden, enabling accurate detection of a person's fallen state in images.
Smart Images

Figure 2025089029000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.
Background Art
[0002] In recent years, technologies for detecting the fallen state of a person from the video of a surveillance camera have been proposed and are being applied in the fields of customer security in stores and urban surveillance. For example, there is a need to confirm the presence or absence of the store's fault when a customer falls and is injured and sues the store side, and to discover and notify the sitting-in and overnight staying of vagrants at ATMs and the like.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0005] When a person standing in front of a camera is reflected on the floor surface and appears to be lying on the floor, there may be a case where the reflected image of the person on the floor surface is erroneously detected. In such a case, there may be an erroneous detection due to the shadow of the person. Patent Document 1 discloses a method for suppressing the erroneous detection due to the shadow of a person based on the luminance value. However, since the luminance value of the reflected image of the person on the floor surface has little difference from that of the actual body of the person, the erroneous detection due to the reflection of the person cannot be suppressed. In addition, Non-Patent Document 1 discloses a method for detecting a reflection region by a model using a Deep Neural Network (DNN). However, in the method disclosed in Non-Patent Document 1, the amount of calculation increases.
[0006] An object of the present invention is to suppress the erroneous detection of a reflected image generated when a person is reflected on a floor surface or the like while suppressing an increase in the amount of calculation when detecting the falling state of a person in an image.
Means for Solving the Problems
[0007] The present invention includes a person detection means for detecting an object representing a person from an image, and a fall determination means for determining a fall state of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected by the person detection means and the images of the respective regions of the plurality of objects.
Effect of the Invention
[0008] According to the present invention, when detecting the fall state of a person in an image, it is possible to suppress an increase in the amount of calculation and suppress the misdetection of a reflected image generated by the person being reflected on the floor or the like.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Mode for Carrying Out the Invention
[0010] Hereinafter, with reference to the accompanying drawings, embodiments for carrying out the present invention will be described in detail. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.
[0011] <Embodiment 1> The image processing apparatus according to this embodiment detects the fallen state of a person in an image. In this embodiment, a plurality of objects representing a person are detected by a Deep Neural Network (hereinafter referred to as DNN) for human detection. Further, the fallen state of the person corresponding to the plurality of objects is determined by a DNN for fall detection. Then, using the intermediate feature map of the DNN for human detection and the intermediate feature map of the DNN for fall detection, a feature vector corresponding to each of the plurality of objects is calculated, and the similarity between the plurality of objects is calculated. Then, based on the similarity between the plurality of objects, an object highly likely to represent a reflection image generated by the person being reflected on the floor surface is removed from the detection target. The procedure will be described below.
[0012] FIG. 1 shows an example of the hardware configuration of the image processing apparatus. The image processing apparatus 100 includes a CPU 101, a ROM 102, a RAM 103, a secondary storage device 104, an imaging device 105, an input device 106, a display device 107, and a network I / F 108. These components are interconnected via a bus 109.
[0013] The CPU 101 executes various processes using programs and data stored in the ROM 102 and the secondary storage device 104. The CPU 101 controls the entire image processing apparatus 100. The ROM 102 is a non-volatile memory and stores various programs and data. The RAM 103 is a volatile memory and stores temporary data such as images. The RAM 103 also stores programs and data loaded from the ROM 102 and the secondary storage device 104.
[0014] The secondary storage device 104 is a rewritable large-capacity storage device such as a hard disk drive or a flash memory, and stores various programs, image data, various setting data, etc. The data stored in the secondary storage device 104 also includes model parameters related to the learned deep learning models (DNN for human detection, DNN for fall determination) used in this embodiment. The programs and data stored in the secondary storage device 104 are appropriately loaded into the RAM 103 according to the control by the CPU 101 and become the processing targets by the CPU 101.
[0015] The imaging device 105 is composed of an imaging lens, an imaging sensor such as a CCD or a CMOS, a video signal processing unit, etc., and captures images. Note that the imaging device 105 may be integrally configured with the image processing device 100, or may be connected to the image processing device 100 via a network I / F 108 or the like. The input device 106 is an input device such as a keyboard, a mouse, a touch panel, etc., and enables input from the user. The display device 107 is a display device such as a liquid crystal display, and displays processing results, etc. to the user. The network I / F 108 is an interface for connecting to a network such as the Internet or an intranet.
[0016] FIG. 2 shows a functional configuration example of the image processing device 100 according to this embodiment. The image processing device 100 includes an image acquisition unit 201, a human detection unit 202, an image cutout unit 203, a fall determination unit 204, a similarity calculation unit 205, and a removal unit 206. By the CPU 101 reading out the programs stored in the ROM 102 and the secondary storage device 104 and executing them in the RAM 103, each function shown in FIG. 2 is realized.
[0017] The image acquisition unit 201 sequentially acquires images from the imaging device 105 at predetermined time intervals and provides the images to the human body detection unit 202. Note that the image acquisition unit 201 is not limited to the input of captured images from the imaging device 105. For example, images may be input by streaming input via a network and reading video data (recorded video) from the secondary storage device 104.
[0018] The human body detection unit 202 detects an object representing a person from the images acquired by the image acquisition unit 201. The detection result of the object includes information on the circumscribed rectangle of the object (hereinafter referred to as the human body rectangle) and information on its likelihood (hereinafter referred to as the human body likelihood). Further, the human body detection unit 202 generates a feature map corresponding to the object. The image cutting-out unit 203 acquires, as a cut-out image, a region corresponding to the human body rectangle detected by the human body detection unit 202 in the image acquired by the image acquisition unit 201.
[0019] The fall determination unit 204 generates score information (hereinafter referred to as the fall score) representing the fall state of the person corresponding to the object from the cut-out image acquired by the image cutting-out unit 203. Further, the fall determination unit 204 generates a feature map corresponding to the object. The similarity calculation unit 205 generates a feature vector corresponding to the object from the feature map generated by the human body detection unit 202 and the feature map generated by the fall determination unit 204. Then, the similarity between the objects is calculated. The removal unit 206 removes, from the detection targets, an object highly likely to represent a reflection image caused by a person being reflected on the floor surface, based on the similarity between the objects and the fall states of the respective objects.
[0020] FIG. 3 is a flowchart showing the processing of the image processing apparatus 100 according to the present embodiment. In the following description of the flowchart, in the description of each step, the step is denoted by prefixing S to omit the description of the step. Software implementing the processing corresponding to each step of the flowchart is read from the ROM 102 or the secondary storage device 104 and executed by the CPU 101.
[0021] In S101, the image acquisition unit 201 acquires an image from the imaging device 105. The image 400 shown in FIG. 4 is an example of the acquired image. In the image 400, there are person images 401 to 403 corresponding to three persons, and person images (reflection images) 404 to 406 generated by each person being reflected on the floor surface. The number of person images in the image is arbitrary, and may be 0, may be 3, and is not particularly limited.
[0022] In S102, the human body detection unit 202 applies a DNN for human body detection to the image 400 acquired in S101 to detect an object representing a person. As the DNN for human body detection, the DNN for detection described in Non-Patent Document 2 is used. FIG. 5 shows the detection result of the object with respect to the image 400. The objects 501 to 506 correspond to the person images 401 to 406. Note that the method is not limited to using the DNN for human body detection as long as it can detect an object representing a person. The human body detection unit 202 is an example of a person detection means.
[0023] The objects 501 to 506 each include information on a human body rectangle and a human body likelihood. The human body rectangle includes the upper left x coordinate, the upper left y coordinate, the width, and the height of the rectangle. The human body likelihood is a value from 0 to 1, and the closer it is to 1, the higher the possibility that a human body exists in the human body rectangle.
[0024] Subsequently, the human body detection unit 202 generates a feature map corresponding to each object 501 to 506 using the feature map corresponding to the entire image obtained from the DNN for human body detection and the objects 501 to 506. A method for generating a feature map corresponding to each object will be described with reference to FIG. 6.
[0025] The human body detection unit 202 performs RoIAlign described in Non-Patent Document 3 using the intermediate feature map 601 corresponding to the entire image obtained from the fourth convolutional layer of the DNN for human body detection and the objects 501 to 506. Specifically, by sampling the values of the regions corresponding to each human body rectangle on the intermediate feature map by bilinear interpolation, intermediate feature maps 602 to 607 corresponding to each object are generated. The size of the intermediate feature maps 602 to 607 is H Di ×W Di ×C D as represented. Here, i is the index of the corresponding human figure, H Di is the height of the intermediate feature map, W Di is the width of the intermediate feature map, and C D is the number of channels of the intermediate feature map (here, equal to the number of output channels of the fourth convolutional layer of the DNN for human body detection). Note that the method for generating the feature map corresponding to each object representing a human body is not limited to the above method. Also, the layer and the number of layers for obtaining the intermediate feature map are arbitrary and not particularly limited. In this way, the human body detection unit 202 obtains the feature maps corresponding to each of the objects 501 to 506.
[0026] In S103, the image cutting-out unit 203 generates cut-out images by cutting out the regions within the human body rectangles in the image 400 using the image 400 obtained in S101 and the objects 501 to 506 obtained in S102. FIG. 7 shows the cut-out images 701 to 706. The cut-out images 701 to 706 correspond to the human figures 401 to 406.
[0027] In S104, the fall determination unit 204 applies a DNN for fall determination to the cut-out images 701 to 706 generated in S103 to perform fall determination for each object. FIG. 8 schematically shows the DNN 800 for fall determination used in this embodiment. The DNN 800 for fall determination has a feature extraction layer 801 and a discrimination layer 802. By inputting the cut-out images 701 to 706 to the DNN 800 for fall determination, a fall determination result 900 and an intermediate feature map are output. Note that the method of using the DNN for fall determination is not limited as long as it can determine the fall state of the person corresponding to the object. The fall determination unit 204 is an example of a fall determination means.
[0028] FIG. 9 shows an example of the fall determination result 900. The fall determination result 900 includes score information (fall score) in which the value corresponding to the classification of the fall state such as "fallen" or "not fallen" increases as the possibility that the person is in the corresponding state is higher. The fall determination result 900 is set data of the objects 501 to 506 and their fall scores. The intermediate feature map is obtained from the last convolutional layer of the feature extraction layer 801 of the DNN 800 for fall determination. The size of the intermediate feature map is H Ci ×W Ci ×C C represented by. Here, i is the index of the corresponding person image, H Ci is the height of the intermediate feature map, W Ci is the width of the intermediate feature map, and C C is the number of channels of the intermediate feature map (here, equal to the output channel number of the last convolutional layer of the feature extraction layer 801). Note that as long as it is an intermediate feature map obtained in the process of outputting the fall determination result, the layer and the number of layers for obtaining the intermediate feature map are arbitrary and not particularly limited. In this way, the fall determination unit 204 obtains a feature map corresponding to each of the objects 501 to 506 by inputting the cut-out images 701 to 706 of the objects 501 to 506 to the DNN for fall determination.
[0029] In S105, the similarity calculation unit 205 calculates the similarity between objects using the intermediate feature maps 602 to 607 obtained by the human body detection unit 202 and the intermediate feature map obtained by the fall determination unit 204. The similarity calculation unit 205 is an example of a similarity calculation means.
[0030] FIG. 10 shows an example of the similarity 1000 between each object. The similarity 1000 between each object is the similarity corresponding to each combination of the objects 501 to 506. First, the similarity calculation unit 205 generates an intermediate feature vector by averaging and concatenating in the spatial direction the intermediate feature maps 602 to 607 obtained by the human body detection unit 202 corresponding to each object and the intermediate feature map obtained by the fall determination unit 204. The intermediate feature vector generated here is a vector of C D + C C dimensions. Note that the method of generating the intermediate feature vector is not limited to the above method, and may be concatenated using basic statistics such as the maximum, minimum, and median in the spatial direction of the intermediate feature map. Also, when C D and C C are equal, instead of concatenating the intermediate feature maps, the sum or difference may be used. In that case, the intermediate feature vector is a vector of C D or C C dimensions. Any method of generating an intermediate feature vector using the intermediate feature maps obtained by each of the human body detection unit 202 and the fall determination unit 204 corresponding to each object so that the number of dimensions is equal may be used. Also, the similarity calculation unit 205 may generate an intermediate feature vector using only one of the intermediate feature maps obtained by the human body detection unit 202 and the intermediate feature map obtained by the fall determination unit 204.
[0031] Next, the similarity calculation unit 205 calculates the cosine similarity between the intermediate feature vectors as the similarity between objects. Note that the correlation coefficient between the intermediate feature vectors may be calculated as the similarity between objects. Any method that can calculate the similarity between objects using the intermediate feature vectors is not particularly limited.
[0032] In S106, the removal unit 206 uses the fall determination result 900 obtained by the fall determination unit 204 and the similarity 1000 between each object obtained by the similarity calculation unit 205 to remove an object that satisfies the condition of being a reflected image of a person from the detection target. In the present embodiment, the removal unit 206 determines for each object whether the following two conditions are satisfied, and if it is determined that either one of the two conditions is satisfied, the object is removed. The removal unit 206 is an example of a removal means. After that, a series of processes in this flowchart are completed.
[0033] <Condition 1>: There exists an object whose fall score is equal to or higher than a first threshold value and whose similarity is equal to or higher than a second threshold value. Also, the fall score of the person (similar person) corresponding to the object whose similarity is equal to or higher than the second threshold value is less than the first threshold value, and the position on the image is below the position of the object corresponding to the similar person. <Condition 2>: There exists an object whose fall score is equal to or higher than a first threshold value and whose similarity is equal to or higher than a second threshold value. Also, the fall score of the person (similar person) corresponding to the object whose similarity is equal to or higher than the second threshold value is equal to or higher than the first threshold value, and the position on the image is below the position of the object corresponding to the similar person.
[0034] Here, since the objects 504 and 506 satisfy Condition 1 and the object 505 satisfies Condition 2, the removal unit 206 removes the objects 504 to 506. Note that Condition 1 and Condition 2 are examples of the conditions for a reflected image of a person, and the threshold value of the fall score (first threshold value) and the threshold value of the similarity (second threshold value) are set to 0.6 and 0.5 respectively, but arbitrary values such as 0.3 or 0.7 may be set.
[0035] Under the above Conditions 1 and 2, in order to remove the reflected image that appears on the floor surface, a condition is added that the position on the image of the first object is below the position of the second object whose similarity to the first object is equal to or greater than the second threshold value. On the other hand, when a person falls near a wall surface such as glass or a mirror in a room or a show window on the street, the reflected image reflected on the wall surface may also appear to be falling. Therefore, considering the positional relationship with other objects that are likely to reflect light such as the floor surface and wall surface in the image, lighting conditions, the orientation and settings of the imaging device 105, etc., other conditions regarding the positional relationship between the first object and the second object may be set.
[0036] In the above embodiment, after determining the falling state (S104), the similarity between objects is calculated (S105). However, after calculating the similarity between objects, the falling state may be determined. In this case, the similarity calculation unit 205 calculates the similarity between objects using, for example, the intermediate feature map obtained by the human body detection unit 202. Then, the fall determination unit 204 performs a fall determination for each object based on the cutout images 701 to 706 of the objects 501 to 506. In this case, when there is a second object whose similarity to the first object is equal to or greater than the second threshold value, the fall determination unit 204 may not perform a fall determination for the first object or the second object. Specifically, the fall determination unit 204 determines whether to perform a fall determination based on the images of the respective regions of the first object and the second object. For example, when the position of the first object on the image is below the position of the second object, the fall determination unit 204 performs a fall determination for the second object but does not perform a fall determination for the first object. Thereby, an object that is likely to be a reflected image of a person can be excluded from the target of the fall determination.
[0037] In addition, when, for a combination of objects with a similarity equal to or higher than the second threshold value, one indicates a state of falling and the other indicates a state of not falling, the removing unit 206 may remove the former one.
[0038] According to the present embodiment, when detecting the falling state of a person in an image, it is possible to suppress an increase in the amount of calculation and suppress the erroneous detection of a reflected image generated when the person is reflected on the floor or the like.
[0039] <Embodiment 2> In the present embodiment, a pair with a high possibility of combination between an object representing the entity of a person and an object due to the reflection of the person on the floor is selected, and the similarity is calculated only between the selected objects. Hereinafter, the procedure will be described. Note that the description of the same parts as those in Embodiment 1 will be omitted, and the description will be centered on the differences.
[0040] FIG. 11 shows a functional configuration example of the image processing apparatus 100 according to the present embodiment. The image processing apparatus 100 in FIG. 2 is different in that it has a selection unit 1101. The selection unit 1101 selects, for each object, an object that is highly likely to be a reflected image generated when the person corresponding to the object is reflected on the floor.
[0041] FIG. 12 is a flowchart showing the processing of the image processing apparatus 100 shown in the present embodiment. First, the image processing apparatus 100 executes the processing shown in S201 to S204. The processing shown in S201 to S204 is the same as that in S101 to S104 in FIG. 3.
[0042] In S205, the selection unit 1101 selects, as a pair, an object that is highly likely to be a reflection image formed by the person corresponding to the target object being reflected on the floor surface, for the target object. First, as shown in FIG. 13, the selection unit 1101 calculates coordinates 1302 that are symmetric with respect to the bottom side of the human body rectangle, for each center coordinate 1301 of the human body rectangles obtained by the human body detection unit 202. Next, the selection unit 1101 selects, as a pair, an object having a center coordinate at a position close to the coordinates 1302. Specifically, the Euclidean distance between the center coordinate of another human body rectangle and the coordinates 1302 is calculated, and an object with a Euclidean distance equal to or less than a predetermined threshold value is selected as a pair. Here, object 504 is selected as a pair for object 501, object 505 is selected as a pair for object 502, and object 506 is selected as a pair for object 503. Note that the method is not limited to the above method, as long as it is a method capable of selecting, as a pair, an object that is highly likely to be a reflection image formed by the person corresponding to the target object being reflected on the floor surface or the like, based on the positional relationship with the target object. The selection unit 1101 is an example of selection means.
[0043] In S206, the similarity calculation unit 205 calculates the similarity between the objects selected as a pair in S205, using the intermediate feature maps 602 to 607 obtained by the human body detection unit 202 and the intermediate feature map obtained by the fall determination unit 204. In Embodiment 1, the similarity was calculated for all combinations of objects, but in this embodiment, the similarity is calculated only for the objects selected as a pair by the selection unit 1101.
[0044] FIG. 14 shows an example of the similarity 1400 between each object. In this embodiment, the similarity calculation unit 205 calculates the similarity only for the combinations of object 501 and object 504, object 502 and object 505, and object 503 and object 506.
[0045] The process shown in S207 to be executed next is the same as the process shown in S106, so the description thereof is omitted. After that, a series of processes in this flowchart are completed.
[0046] According to this embodiment, the number of pairs for calculating the similarity is limited, and an increase in the amount of calculation can be suppressed.
[0047] <Embodiment 3> In this embodiment, a DNN for joint detection is used to detect the joints of a person, and based on the detection result of the joints, the similarity between objects is calculated. The procedure will be described below. Note that the parts similar to those in Embodiment 2 will be omitted from the description, and the description will focus on the differences.
[0048] FIG. 15 shows a functional configuration example of the image processing apparatus 100 according to this embodiment. The image processing apparatus 100 in FIG. 11 is different in that it has a joint detection unit 1501. The joint detection unit 1501 detects the joints of an object from the cut-out image acquired by the image cut-out unit 203.
[0049] FIG. 16 is a flowchart showing the processing of the image processing apparatus 100 shown in this embodiment. First, the image processing apparatus 100 executes the processes shown in S301 to S303. The processes shown in S301 to S303 are the same as those in S201 to S203 in FIG. 12.
[0050] In S304, the joint detection unit 1501 applies a DNN for joint detection to the cut-out images 701 to 706 generated in S303 to detect the joints of each object. As the DNN for joint detection, the detection DNN described in Non-Patent Document 4 is used. By inputting the cut-out images 701 to 706 to the DNN for joint detection, a joint likelihood map and an intermediate feature map are output. The joint likelihood map is a heat map of the position information of the joints. The joints in this embodiment consist of 17 points including the nose, both eyes, both ears, both shoulders, both elbows, both waists, both knees, and both ankles. The joint detection unit 1501 is an example of a joint detection means.
[0051] FIG. 17(a) schematically shows a joint likelihood map corresponding to the cropped image 701 as an example of the joint likelihood map. The joint likelihood map is heatmap information of 17 sheets corresponding to each of the 17 joints, and holds a higher value at a position where there is a higher possibility of the joint existing. The joint detection unit 1501 calculates joint coordinates for each joint by obtaining a peak position from the joint likelihood map. FIG. 17(b) schematically shows joint coordinates corresponding to the cropped image 701 as an example of the joint coordinates. In FIG. 17(b), the joint coordinates of 17 points are shown for the human body in the cropped image 701.
[0052] The intermediate feature map is obtained from the fifth convolutional layer of the DNN for joint detection. The size of the intermediate feature map is represented by H Pi ×W Pi ×C P Here, i is the index of the corresponding person image, H Pi is the height of the intermediate feature map, W Pi is the width of the intermediate feature map, and C P is the number of channels of the intermediate feature map (here, equal to the number of output channels of the fifth convolutional layer of the DNN for joint detection). Note that as long as it is an intermediate feature map obtained in the process of outputting the joint likelihood map, the layer and the number of layers for obtaining the intermediate feature map are arbitrary and not particularly limited. In this way, the joint detection unit 1501 obtains a feature map corresponding to each of the objects 501 to 506 by inputting the cropped images 701 to 706 of the objects 501 to 506 into the DNN for joint detection.
[0053] In S305, the fall determination unit 204 applies a DNN for fall determination to the cut-out images 701 to 706 generated in S303 and the joint point likelihood map obtained in S304 to perform a fall determination for each object. FIG. 18 schematically shows the DNN 1800 for fall determination used in this embodiment. The DNN 1800 for fall determination includes feature extraction layers 1801 and 1802, a combination layer 1803, and an identification layer 1804. By inputting the cut-out images 701 to 706 and the joint point likelihood map to the DNN 1800 for fall determination, a fall determination result 1805 and an intermediate feature map are output. The feature extraction layers 1801 and 1802 respectively take the cut-out image and the joint point likelihood map as inputs. The feature extraction layer 1801 and the feature extraction layer 1802 have a structure that inputs to the combination layer 1803 with the spatial information matched by making the output size a numerical vector of M×N dimensions. The identification layer 1804 has a structure that outputs a fall determination result 1805 based on the feature amounts combined by the combination layer 1803.
[0054] The process of S306 to be executed next is the same as the process shown in S205, and thus the description thereof is omitted. In S307, the similarity calculation unit 205 calculates the similarity between the objects selected as a pair in S306. Here, in addition to the intermediate feature maps 602 to 607 obtained by the human body detection unit 202 and the intermediate feature map obtained by the fall determination unit 204, the joint point likelihood map obtained by the joint point detection unit 1501 and the joint point coordinates are used. In this embodiment, similarly to Embodiment 2, the similarity is calculated only for the objects selected as a pair by the selection unit 1101.
[0055] First, the similarity calculation unit 205 averages and concatenates in the spatial direction the intermediate feature maps obtained by the human body detection unit 202, the intermediate feature maps obtained by the joint point detection unit 1501, and the intermediate feature maps obtained by the fall determination unit 204 corresponding to each object. Thereby, an intermediate feature vector is generated. The intermediate feature vector generated here is C D +C P +C CIt is a vector of dimensions. Note that it may be concatenated using basic statistics such as the maximum, minimum, and median in the spatial direction of the intermediate feature map. Also, C D C P C C When they are equal, instead of concatenating the intermediate feature maps, sums or differences may be used. In that case, the intermediate feature vector is C D C P or C C dimensional vector. A method of generating an intermediate feature vector such that the number of dimensions is equal may be used, using the intermediate feature maps obtained by each of the human body detection unit 202, the joint point detection unit 1501, and the fall determination unit 204 corresponding to each object. Also, the similarity calculation unit 205 may generate an intermediate feature vector using only at least any one of the intermediate feature maps obtained by the human body detection unit 202, the intermediate feature map obtained by the joint point detection unit 1501, and the intermediate feature map obtained by the fall determination unit 204.
[0056] Next, the similarity calculation unit 205 calculates the cosine similarity between the intermediate feature vectors for the selected object pairs. Also, the similarity calculation unit 205 inverts the joint point likelihood map of one of the selected object pairs in the vertical direction and calculates the similarity between the joint point likelihood maps based on the average of the absolute values of the differences from the joint point likelihood map of the other object. The smaller the average of the absolute values of the differences from the inverted joint point likelihood map, the greater the similarity. Note that the similarity may be calculated based on the average of the squares of the differences between the joint point likelihood maps, and any method that can calculate the similarity between the joint point likelihood maps may be used.
[0057] Furthermore, the similarity calculation unit 205 may calculate the similarity of the joint point coordinates based on the average of the absolute values of the differences between the joint point coordinates of one object and the joint point coordinates of the other object after vertically inverting the y coordinate of the joint point coordinates of one object among the selected object pairs. The smaller the average of the absolute values of the differences from the inverted joint positions, the greater the similarity. Note that the similarity may be calculated based on the average of the squares of the differences in joint point coordinates, and any method that can calculate the similarity between joint point coordinates may be used.
[0058] Finally, the similarity calculation unit 205 calculates the average of the cosine similarity between the intermediate feature vectors, the similarity between the joint likelihood maps, and the similarity between the joint point coordinates as the similarity between objects. Note that the maximum of each similarity may be calculated as the similarity between objects. Any method that can calculate the similarity between objects is not particularly limited.
[0059] The process shown in S308 to be executed next is the same as the process shown in S106, so the description is omitted. After that, the series of processes in this flowchart ends.
[0060] According to this embodiment, by using the detection result of the joint points of a person, it is possible to suppress the misdetection of the reflected image generated when the person is reflected on the floor or the like.
[0061] <Embodiment 4> In this embodiment, a human body ratio representing the ratio of the sizes of each part of the human body is calculated from the joint points of the person, and the similarity between objects is calculated based on the human body ratio. The procedure will be described below. Note that the description of the same parts as in Embodiment 3 will be omitted, and the description will focus on the differences.
[0062] FIG. 19 shows a functional configuration example of the image processing apparatus 100 according to this embodiment. The image processing apparatus 100 in FIG. 15 is different in that it has a human body ratio calculation unit 1901 instead of the selection unit 1101. The human body ratio calculation unit 1901 calculates a human body ratio representing the ratio of the sizes of each part of the human body from the joint point coordinates obtained by the joint point detection unit 1501.
[0063] Figure 20 is a flowchart showing the processing of the image processing apparatus 100 shown in this embodiment. First, the image processing apparatus 100 executes the processing shown in S401 to S405. The processing shown in S401 to S405 is the same as that in S301 to S305 in FIG. 16.
[0064] In S406, the human body ratio calculation unit 1901 calculates the human body ratio of each object using the joint point coordinates obtained by the joint point detection unit 1501. Specifically, the area of the minimum circumscribed rectangle that encloses all the joint point positions of the nose, both eyes, and both ears is calculated as the head size, and the area of the human body rectangle is calculated as the whole body size. Then, the ratio of the head size to the whole body size is output as the human body ratio. Note that the ratio of the width between both shoulders to the width between both waists may be calculated as the human body ratio, or a plurality of ratios may be calculated. Any method that can calculate the ratio between the lengths of each part of the human body or the ratio to the length of the whole body is acceptable. The human body ratio calculation unit 1901 is an example of the human body ratio calculation means.
[0065] In S407, the similarity calculation unit 205 calculates the similarity between objects using the human body ratio of each object obtained in S406. Specifically, the smaller the absolute value of the difference in the human body ratio between objects, the greater the similarity between the objects. Any method that can calculate the similarity between objects using the human body ratio is not particularly limited. When the human body ratio consists of a plurality of ratios, the cosine similarity may be used. Note that in this embodiment, since the selection unit 1101 is not provided, the similarity using the human body ratio is calculated for all combinations of objects. Note that, similar to Embodiment 3, the similarity may be calculated only for the objects as a pair selected by the selection unit 1101.
[0066] Next, the processing shown in S408 to be executed is the same as the processing shown in S106, so the description is omitted. After that, a series of processes in this flowchart ends.
[0067] According to this embodiment, by using the body ratio of a person, the calculation accuracy of the similarity between objects representing the person can be improved.
[0068] <Embodiment 5> In this embodiment, a predicted image that pseudo-reproduces the reflected image of a person on the floor surface is generated by estimating the three-dimensional model of the person, and the similarity between objects is calculated based on the comparison result between the actual image and the predicted image. The procedure will be described below. Note that the description of the same parts as in Embodiment 1 will be omitted, and the description will focus on the differences.
[0069] FIG. 21 shows a functional configuration example of the image processing apparatus 100 according to this embodiment. The image processing apparatus 100 in FIG. 2 is different in that it has a three-dimensional model estimation unit 2101. The three-dimensional model estimation unit 2101 estimates the three-dimensional model of a person, and generates a predicted image (hereinafter referred to as a pseudo-reflection image) that pseudo-reproduces the reflected image of the person on the floor surface using the estimated three-dimensional model.
[0070] FIG. 22 is a flowchart showing the processing of the image processing apparatus 100 shown in this embodiment. First, the image processing apparatus 100 executes the processing shown in S501 to S504. The processing shown in S501 to S504 is the same as that in S101 to S104 in FIG. 3.
[0071] In S505, the three-dimensional model estimation unit 2101 generates a pseudo-reflection image by performing three-dimensional model estimation on the cutout images 701 to 706 generated in S503. As a result, a pseudo-reflection image is generated for each of the cutout images 701 to 706 generated in S503. For the three-dimensional model estimation, the method described in Non-Patent Document 5 is used. The three-dimensional model consists of polygon data. The three-dimensional model estimation unit 2101 generates a pseudo-reflection image by inverting the y coordinate of each vertex of the polygon data and performing rendering. Note that any method that can generate a pseudo-reflection image from the cutout image may be used, and the method for generating the pseudo-reflection image is not limited to this.
[0072] In S506, the similarity calculation unit 205 calculates the similarity between the cut-out images 701 to 706 generated in S503 and the pseudo-reflection image obtained in S505. The similarity here is calculated for each combination of the objects 501 to 506 and the pseudo-reflection images corresponding to the objects 501 to 506. As a method for calculating the similarity, there is a method of calculating the cosine similarity of the pixel value histograms between the cut-out image and the pseudo-reflection image. Note that any method that can calculate the similarity between images may be used.
[0073] The process shown in S507, which is executed next, is the same as the process shown in S106, so the description thereof is omitted. After that, a series of processes in this flowchart ends.
[0074] According to the present embodiment, by using the reflection image pseudo-reproduced by estimating the three-dimensional model of a person, it is possible to improve the calculation accuracy of the similarity between objects representing the person.
[0075] <Other Embodiments> Although the embodiment has been described in detail above, the present invention can be implemented in embodiments such as a system, an apparatus, a method, a program, or a recording medium (storage medium), for example. Specifically, it may be applied to a system composed of a plurality of devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus composed of a single device.
[0076] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and having one or more processors in the computer of the system or apparatus read and execute the program. It can also be realized by a circuit (for example, an ASIC) that realizes one or more functions.
[0077] The disclosure of each of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) A person detection means for detecting an object representing a person from an image, Based on the similarity between a plurality of objects detected by the person detection means and the images of the respective regions of the plurality of objects, a fall determination means for determining the fall states of a plurality of persons corresponding to the plurality of objects, An image processing apparatus characterized by comprising: (Configuration 2) The fall determination means does not determine the fall state of the first object or the second object when there is a second object whose similarity to the first object is equal to or greater than a threshold value. The image processing apparatus according to Configuration 1. (Configuration 3) The fall determination means does not determine the fall state of the first object or the second object based on the positional relationship between the first object and the second object when there is a second object whose similarity to the first object is equal to or greater than a threshold value. The image processing apparatus according to Configuration 1. (Configuration 4) Removing means for removing the first object from the detection target when the first object whose fall state is determined by the fall determination means satisfies a predetermined condition, The image processing apparatus according to any one of Configurations 1 to 3, further comprising: (Configuration 5) The predetermined condition relates to the fall state of the first object and a second object whose similarity to the first object is equal to or greater than a threshold value. The image processing apparatus according to Configuration 4. (Configuration 6) The predetermined condition relates to the positional relationship between the first object and a second object whose similarity to the first object is equal to or greater than a threshold value. The image processing apparatus according to Configuration 4 or 5. (Configuration 7) The person detection means detects an object using a first learning model, and acquires a feature map of the object from the first learning model, Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the person detection means The image processing apparatus according to any one of Configurations 1 to 6, further comprising the above (Configuration 8) Acquisition means for acquiring a feature map corresponding to each of the plurality of objects by inputting an image of each region of the plurality of objects into a second learning model for determining a person's falling state Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the acquisition means The image processing apparatus according to any one of Configurations 1 to 7, further comprising the above (Configuration 9) Selection means for selecting a pair that is likely to be a combination of a person and a reflected image of the person from among the plurality of objects detected by the person detection means Similarity calculation means for calculating the similarity between the objects corresponding to the pair selected by the selection means The image processing apparatus according to any one of Configurations 1 to 8, further comprising the above (Configuration 10) The image processing apparatus according to Configuration 9, wherein the selection means selects the pair based on the positional relationship of the plurality of objects (Configuration 11) Joint point detection means for detecting joint points of a person from an image of each region of the plurality of objects Similarity calculation means for calculating the similarity between the plurality of objects based on the detection result by the joint point detection means The image processing apparatus according to any one of Configurations 1 to 10, further comprising the above (Configuration 12) The image processing apparatus according to Configuration 11, wherein the fall determination means determines the fall states of a plurality of persons corresponding to the plurality of objects based on the detection result by the joint point detection means (Configuration 13) The joint point detection means detects joint points using a third learning model, and obtains a feature map of the object from the third learning model. Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the joint point detection means The image processing apparatus according to Configuration 11 or 12, further comprising the above. (Configuration 14) The joint point detection means obtains a joint point likelihood map of the object from the third learning model. The similarity calculation means vertically flips the joint point likelihood map of one of the two objects, calculates the difference from the other joint point likelihood map, and calculates the similarity between the two objects based on the calculated difference. The image processing apparatus according to Configuration 13. (Configuration 15) The joint point detection means calculates joint point coordinates from the joint point likelihood map obtained from the third learning model. The similarity calculation means vertically flips the joint point position of one of the two objects, calculates the difference from the joint point position of the other object, and calculates the similarity between the two objects based on the calculated difference. The image processing apparatus according to Configuration 13 or 14. (Configuration 16) Further comprising a human body ratio calculation means for calculating a human body ratio representing the ratio of the sizes of each part of the human body based on the detection result by the joint point detection means. The similarity calculation means calculates the similarity based on the human body ratio. The image processing apparatus according to any one of Configurations 11 to 15. (Configuration 17) Generating means for estimating a three-dimensional model of a plurality of persons corresponding to the plurality of objects from images of the respective regions of the plurality of objects, and generating a predicted image that reproduces a reflected image generated by the person being reflected in another object based on the estimated three-dimensional model. Similarity calculation means for calculating the similarity between the plurality of objects based on the image of the area of each of the plurality of objects and the predicted image corresponding to each of the plurality of objects. The image processing apparatus according to any one of Configurations 1 to 16, further comprising the above. (Method) A person detection step of detecting an object representing a person from an image, A fall determination step of determining the fall states of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected in the person detection step and the image of the area of each of the plurality of objects. An image processing method characterized by including the above. (Program) A program for causing a computer of an image processing apparatus to function as person detection means for detecting an object representing a person from an image, and fall determination means for determining the fall states of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected by the person detection means and the image of the area of each of the plurality of objects.
Claims
1. A person detection means for detecting an object representing a person from an image, A fall determination means for determining the fall states of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected by the person detection means and the images of the respective regions of the plurality of objects, An image processing apparatus characterized by comprising the above.
2. The image processing apparatus according to claim 1, wherein when there is a second object whose similarity to the first object is equal to or greater than a threshold value, the fall determination means does not determine the fall state of the first object or the second object.
3. The image processing apparatus according to claim 1, wherein when there is a second object whose similarity to the first object is equal to or greater than a threshold value, the fall determination means does not determine the fall state of the first object or the second object based on the positional relationship between the first object and the second object.
4. Removing means for removing the first object from the detection target when the first object whose fall state is determined by the fall determination means satisfies a predetermined condition, The image processing apparatus according to claim 1, further comprising the above.
5. The image processing apparatus according to claim 4, wherein the predetermined condition relates to the fall states of the first object and a second object whose similarity to the first object is equal to or greater than a threshold value.
6. The image processing apparatus according to claim 4, wherein the predetermined condition relates to the positional relationship between the first object and a second object whose similarity to the first object is equal to or greater than a threshold value.
7. The person detection means detects an object using a first learning model and obtains a feature map of the object from the first learning model, Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the person detection means, The image processing apparatus according to claim 1, further comprising the above.
8. An acquisition means for acquiring a feature map corresponding to each of the plurality of objects by inputting the image of each region of the plurality of objects into a second learning model for determining the fall state of a person. Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the acquisition means; The image processing apparatus according to claim 1, further comprising the above.
9. Selection means for selecting a pair that is likely to be a combination of a person and a reflected image of the person from among the plurality of objects detected by the person detection means; Similarity calculation means for calculating the similarity between the objects corresponding to the pair selected by the selection means; The image processing apparatus according to claim 1, further comprising the above.
10. The image processing apparatus according to claim 9, wherein the selection means selects the pair based on the positional relationship of the plurality of objects.
11. Joint point detection means for detecting joint points of a person from images of the respective regions of the plurality of objects; Similarity calculation means for calculating the similarity between the plurality of objects based on the detection result by the joint point detection means; The image processing apparatus according to claim 1, further comprising the above.
12. The image processing apparatus according to claim 11, wherein the fall determination means determines the fall states of a plurality of persons corresponding to the plurality of objects based on the detection result by the joint point detection means.
13. The joint point detection means detects joint points using a third learning model, and obtains a feature map of the object from the third learning model. Similarity calculation means for calculating the similarity between the plurality of objects using the feature map obtained by the joint point detection means The image processing apparatus according to claim 11, further comprising the above.
14. The joint point detection means obtains a joint point likelihood map of the object from the third learning model. The similarity calculation means vertically flips the joint point likelihood map of one of the two objects, calculates the difference from the other joint point likelihood map, and calculates the similarity between the two objects based on the calculated difference. The image processing apparatus according to claim 13.
15. The joint point detection means calculates joint point coordinates from the joint point likelihood map obtained from the third learning model. The similarity calculation means in claim 13 is characterized in that it calculates the similarity between the two objects by inverting the joint point positions of one of the two objects in the vertical direction, calculating the difference from the joint point positions of the other object, and calculating the similarity based on the calculated difference.
16. The image processing apparatus further includes a human body ratio calculation means for calculating a human body ratio representing the ratio of the sizes of the respective parts of the human body based on the detection result by the joint point detection means. The similarity calculation means in claim 11 is characterized in that it calculates the similarity based on the human body ratio.
17. Generating means for estimating a three-dimensional model of a plurality of persons corresponding to the plurality of objects from the images of the respective regions of the plurality of objects, and generating a predicted image that reproduces a reflected image generated by the person being reflected in another object based on the estimated three-dimensional model. Similarity calculation means for calculating the similarity between the plurality of objects based on the images of the respective regions of the plurality of objects and the predicted images corresponding to the respective plurality of objects. The image processing apparatus according to claim 1, further comprising the above.
18. A person detection step of detecting an object representing a person from an image. A fall determination step of determining the fall states of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected in the person detection step and the images of the respective regions of the plurality of objects. An image processing method characterized by including the above.
19. A computer of an image processing apparatus. Person detection means for detecting an object representing a person from an image. Fall determination means for determining the fall states of a plurality of persons corresponding to the plurality of objects based on the similarity between the plurality of objects detected by the person detection means and the images of the respective regions of the plurality of objects. A program for causing the above to function.
Citation Information
Patent Citations
Elevator user detection system
JP2021147227A