Object Tracking Device, Object Tracking Method, and Program
The object tracking device addresses the challenge of accurately associating objects between frames by using numerical values to analyze the spatial relationships and trajectories of objects across multiple frames, ensuring precise object tracking.
Patent Information
- Application Number
- JP2021085259
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing object tracking techniques struggle to accurately associate the same object between consecutive frames, often misidentifying it with other objects, especially when objects are in adjacent lanes or following each other.
An object tracking device that acquires and processes three consecutive captured images to extract possible combinations of object paths and associate objects based on calculated numerical values representing the distance and positional relationships between objects in different frames.
This approach enables accurate tracking of objects without misidentification by considering the spatial relationships and trajectories of objects across multiple frames, improving the precision of object association.
Smart Images

Figure 0007685873000001 
Figure 0007685873000002 
Figure 0007685873000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for associating objects between captured images.
Background Art
[0002] Conventionally, various techniques have been proposed for tracking objects in captured moving images. For example, a tracking frame group composed of a plurality of frames for tracking an object is set, and in each frame F(i) in the tracking frame group, an object region R(i - 1) where the object in the previous frame F(i - 1) exists is set, and a region having the same size as the object region R(i - 1) is set on the frame F(i) as a comparison region. Then, a plurality of the comparison regions are set by moving the comparison region vertically, horizontally, and within a predetermined range by (+X, +Y) with respect to the center position of the object region R(i - 1), and for each comparison region, the total sum d(X, Y) of the differences between corresponding pixels with the object region R(i - 1) is calculated, and the comparison region where the total sum d(X, Y) is minimized is set as the object region R(i) of the frame F(i). Such a technique is known (Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When tracking an object (such as a vehicle or a pedestrian), the association between two consecutive frames as described in Patent Document 1 was performed. However, since the association between two consecutive frames was a method of linking those that are close between frames based on the coordinates of the detected object, for example, vehicles in adjacent lanes or following vehicles may become candidates for connection, and there was a problem that the same object could not be accurately tracked.
[0005] In view of the above circumstances, an object of the present invention is to provide a technique capable of tracking an object without misidentifying it with another object.
Means for Solving the Problems
[0006] To achieve the above object, an object tracking device according to an aspect of the present invention includes an acquisition unit, a detection unit, an extraction unit, and a specification unit. The acquisition unit acquires a first captured image, a second captured image, and a third captured image in chronological order. The detection unit specifies the coordinate positions of one or more objects in each captured image. The extraction unit extracts possible combinations of the paths of the objects between the captured images. The specification unit associates the objects between the captured images based on the extraction result of the extraction unit.
[0007] According to the object tracking device, possible combinations of the paths of the objects are extracted from the coordinate positions of the objects between three captured images, and the same objects are associated based on the extracted result. Therefore, the objects can be tracked without misidentifying them with other objects.
[0008] The extraction unit has a first processing unit that calculates a first numerical value, which is the distance from the coordinate position of the object in the second captured image to the coordinate position of the midpoint of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image. The specification unit may associate the objects between the captured images based on the first numerical value.
[0009] Thereby, since the relationship of the coordinate positions of the second captured image on the line segment connecting the first captured image and the third captured image is calculated, the same objects can be associated from the coordinate positions of the objects between the three captured images.
[0010] The extraction unit further has a second processing unit that calculates a second numerical value which is the absolute value of the difference between half the distance of the line segment connecting the coordinate position of the object in the first captured image of the object and the coordinate position of the object in the third captured image of the object, and the distance of the line segment connecting the coordinate position of the object in the first captured image of the object and the coordinate position of the object in the second captured image of the object. The specifying unit may associate the object between the captured images based on the first numerical value and the second numerical value.
[0011] Thereby, in addition to the relationship of the coordinate positions of the respective captured images based on the first numerical value, the distance of the line segment connecting the coordinate position of the first captured image and the coordinate position of the second captured image, and the distance of the line segment connecting the coordinate position of the first captured image and the coordinate position of the third captured image can be calculated. Therefore, it is possible to associate the same object from the coordinate positions of the object among the three captured images.
[0012] The extraction unit further has a third processing unit that calculates a third numerical value which is the absolute value of the difference between half the distance of the line segment connecting the coordinate position of the object in the first captured image of the object and the coordinate position of the object in the third captured image of the object, and the distance of the line segment connecting the coordinate position of the object in the second captured image of the object and the coordinate position of the object in the third captured image of the object. The specifying unit may associate the object between the captured images based on the first numerical value and the third numerical value.
[0013] Thereby, in addition to the relationship of the coordinate positions of the respective captured images based on the first numerical value, the distance of the line segment connecting the coordinate position of the second captured image and the coordinate position of the third captured image, and the distance of the line segment connecting the coordinate position of the first captured image and the coordinate position of the third captured image can be calculated. Therefore, it is possible to associate the same object from the coordinate positions of the object among the three captured images.
[0014] The extraction unit further includes a third processing unit that calculates a third numerical value, which is the absolute value of the difference between half the distance of the line segment connecting the coordinate position of the object in the first captured image of the object and the coordinate position of the object in the third captured image of the object, and the distance of the line segment connecting the coordinate position of the object in the second captured image of the object and the coordinate position of the object in the third captured image of the object. The specifying unit may associate the object between the captured images based on the first numerical value, the second numerical value, and the third numerical value.
[0015] Thereby, since the relationship between the coordinate positions of each captured image based on the first numerical value, the second numerical value, and the third numerical value can be calculated, it is possible to associate the same object from the coordinate positions of the object among the three captured images.
[0016] For each combination extracted by the extraction unit, the specifying unit calculates the first numerical value, the second numerical value, and the third numerical value, and may calculate a combination in which the first numerical value, the second numerical value, and the third numerical value are each the minimum value.
[0017] Thereby, based on the first numerical value, the second numerical value, and the third numerical value, it is possible to associate the same object from the coordinate positions of the object among the three captured images.
[0018] The specifying unit may associate the object between the captured images based on the sum of the first numerical value, the second numerical value, and the third numerical value based on the combination that becomes the minimum value.
[0019] Thereby, it is possible to associate the same object from the coordinate positions of the object among the three captured images.
[0020] When the minimum value is less than a predetermined threshold value, the specifying unit may associate the object between the captured images based on the combination that becomes the minimum value.
[0021] This makes it possible to track the same object without mistakenly associating other objects that exceed a predetermined range.
[0022] An object tracking method according to one embodiment of the present invention acquires a first captured image, a second captured image, and a third captured image in chronological order, specifies the coordinate positions of one or more objects within each captured image, extracts possible combinations of the paths of the objects between the respective captured images, and associates the objects between the respective captured images based on the extraction result of the extraction unit.
[0023] A program according to one embodiment of the present invention causes execution of steps of acquiring a first captured image, a second captured image, and a third captured image in chronological order, specifying the coordinate positions of one or more objects within each captured image, extracting possible combinations of the paths of the objects between the respective captured images, and associating the objects between the respective captured images based on the extraction result of the extraction unit.
Advantages of the Invention
[0024] As described above, according to the present invention, it is possible to track an object without misidentifying it as another object.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0026] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0027] <Configuration of Control Device 100 and Configuration of Each Part> FIG. 1 is a diagram showing a control device 100 according to an embodiment of the present invention installed on a road. FIG. 2 is a block diagram showing a configuration including the control device 100 according to an embodiment of the present invention.
[0028] In the example shown in FIG. 1, a road extending in the north-south direction is shown. The road extending in the north-south direction has two lanes on one side and four lanes on both sides. Here, in the description of this specification, the road with the traveling direction being the north direction is called road D1, the road with the traveling direction being the south direction is called road D2, the left lane of each road traveling in the north-south direction is called lane R1, and the right lane is called lane R2.
[0029] As shown in FIG. 1, the control device 100 is provided near roads 1 and 2. In this embodiment, the object is described as vehicle 5, but it is not limited to this, and it may be a person or the like.
[0030] The control device 100 is installed along lanes R1 and R2. In FIG. 1, the interval at which the control device 100 is installed is approximately 1 km, but it is not limited to this, and of course other values can also be taken.
[0031] As shown in FIG. 2, the control device 100 includes an object tracking device 10, an imaging unit 20, a storage unit 30, and a communication unit 40.
[0032] (Imaging unit) In the control device 100 of the present embodiment, the imaging unit 20 is installed on the lane boundary line S1 of the road D1 and the lane boundary line S2 of the road D2 in FIG. 1. By installing the imaging unit 20 on the lane boundary line S2, it becomes easy to simultaneously image objects (for example, vehicle 5) existing on both lanes R1 and R2. However, of course, it is not limited to this, and any installation position where the object can be imaged is acceptable. These imaging units 20 are installed at a height, position, and angle such that vehicles 5 traveling on the lanes R1 and R2 of both roads can be imaged. In the present embodiment, the control device 100 is installed on a two-lane road, but it is not limited to this, and it may be installed on a four-lane road.
[0033] The frame rate when the imaging unit 20 images an object can be arbitrarily set, for example, 30 fps.
[0034] (Storage unit) The storage unit 30 includes a non-volatile memory that stores various programs necessary for the processing of the object tracking device 10, captured images acquired from the imaging unit 20, and various information. The above-mentioned various programs may be read from a portable recording medium such as an optical disk or a semiconductor memory, or may be downloaded from a server device on a network.
[0035] (Communication unit) The communication unit 40 communicates between the object tracking device 10 and the traffic control center 200 by wire or wirelessly. Specifically, the communication unit 40 transmits information such as the number of vehicles and the vehicle speed calculated by the object tracking device 10, which will be described later, to the traffic control center 200.
[0036] (Object tracking device) The object tracking device 10 is composed of a CPU (Central Processing Unit) or the like, and as shown in FIG. 2, includes an acquisition unit 11, an analysis unit 12, a detection unit 13, an extraction unit 14, and a specification unit 15.
[0037] The acquisition unit 11 acquires the first captured image, the second captured image, and the third captured image in chronological order. In this embodiment, the first captured image is the oldest in chronological order, and the third captured image is the newest captured image in chronological order. Regarding the captured image, the acquisition unit 11 may directly acquire it from the imaging unit 20, or may acquire the captured image stored in the storage unit 30.
[0038] The analysis unit 12 identifies the objects existing in each captured image acquired by the acquisition unit 11. The objects in each captured image can be detected by adopting a known image processing technique for detecting objects from a still image.
[0039] The detection unit 13 identifies the coordinate positions of the objects in each captured image specified by the analysis unit 12. The method for identifying the coordinate positions of the objects by the detection unit 13 is to take the left end in the captured image as the origin, and identify the four values of the maximum x coordinate, the minimum x coordinate, the maximum y coordinate, and the minimum y coordinate among the coordinate positions of the objects in the captured image, and calculate the center of gravity point from the frame surrounded by the four values of the maximum x coordinate, the minimum x coordinate, the maximum y coordinate, and the minimum y coordinate of the specified object, and identify the center of gravity point as the coordinate position. In this embodiment, the frame is an arbitrarily shaped frame circumscribing the object, and is, for example, approximately rectangular in shape.
[0040] The extraction unit 14 extracts possible combinations of the paths of the objects between the captured images. The extraction unit 14 includes a first processing unit 141, a second processing unit 142, and a third processing unit 143.
[0041] The first processing unit 141 calculates a first numerical value that is the distance between the coordinate position of the object in the second captured image, the coordinate position of the midpoint of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image.
[0042] Here, the first numerical value will be described with reference to FIG. 8(A). FIG. 8(A) is a diagram for explaining the calculation method of the first numerical value. As a premise, in FIG. 8(A), X is the object in the first captured image, Y is the object in the second captured image, and Z is the object in the third captured image. The coordinate position of the object X is (x, y) = (X1, X2), the coordinate position of the object Y is (x, y) = (Y1, Y2), and the coordinate position of the object Z is (x, y) = (Z1, Z2). The calculation method of the first numerical value is to calculate the coordinate position T ((X1 + Z1) ÷ 2, (X2 + Z2) ÷ 2) of the midpoint of the line segment connecting the coordinate position (X1, X2) of the object X in the first captured image and the coordinate position (Z1, Z2) of the object Z in the third captured image. The first numerical value is derived by calculating the distance L1 between the coordinate position T and the coordinate position (Y1, Y2) of the object Y in the second captured image.
[0043] The association of the object based on the first numerical value will be described. First, considering that the same object has moved among the three captured images, it is presumed that the movement trajectory of the object is linear. Therefore, when the coordinate position (X1, X2) of the object X in the first captured image and the coordinate position (Z1, Z2) of the object Z in the third captured image are connected by a line segment, the coordinate position (Y1, Y2) of the object Y in the second captured image therebetween is presumed to be located near the coordinate position T of the midpoint of the above line segment. Therefore, in the calculation method of the first numerical value, the distance L1 between the coordinate position T of the midpoint and the coordinate position (Y1, Y2) of the object Y in the second captured image is calculated, and based on the result, the association of the object among the three captured images can be performed.
[0044] Also, when the distance L1 is the shortest (the first numerical value is the minimum value), that combination is inferred as the movement of the object among the three captured images. Therefore, the association of the object among the three captured images can be performed based on the combination that becomes the minimum value according to the first numerical value.
[0045] The second processing unit 142 calculates a second numerical value that is the absolute value of the difference between half of the distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image, and the distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the second captured image.
[0046] Here, the second numerical value will be described with reference to FIG. 8(B). FIG. 8(B) is a diagram for explaining the calculation method of the second numerical value. Similar to FIG. 8(A), in FIG. 8(B), X is the object in the first captured image, Y is the object in the second captured image, and Z is the object in the third captured image. The coordinate position of the object X is (x, y) = (X1, X2), the coordinate position of the object Y is (x, y) = (Y1, Y2), and the coordinate position of the object Z is (x, y) = (Z1, Z2). The calculation method of the second numerical value is to calculate half of the distance L2 of the line segment connecting the coordinate position (X1, X2) of the object X in the first captured image and the coordinate position (Z1, Z2) of the object Z in the third captured image. Next, the distance L3 of the line segment connecting the coordinate position (X1, X2) of the object X in the first captured image and the coordinate position (Y1, Y2) of the object Y in the second captured image is calculated. The second numerical value is derived by calculating the absolute value |L2 - L3| of the difference between the distance L2 and the distance L3.
[0047] The association of an object based on the second numerical value will be described. First, consider the case where the same object has moved among three captured images. Since it is assumed that the trajectory of the object's movement is linear, the distance L2, which is half of the distance obtained by connecting the coordinate positions (X1, X2) of the object X in the first captured image and the coordinate positions (Z1, Z2) of the object Z in the third captured image with a line segment, and the distance L3 between the coordinate positions (X1, X2) of the object X in the first captured image and the coordinate positions (Y1, Y2) of the object Y in the second captured image are presumed to take similar values. Therefore, in the calculation method of the second numerical value, the absolute value of the difference between the above half distance L2 and the distance L3 is calculated, and based on the result, the association of the object among the three captured images can be performed.
[0048] Also, when the absolute value of the above difference is the shortest (the second numerical value is the minimum value), that combination is presumed to be the movement of the object among the three captured images. Therefore, the association of the object among the three captured images can be performed based on the combination where the second numerical value is the minimum value.
[0049] The third processing unit 143 calculates a third numerical value, which is the absolute value of the difference between half of the distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image, and the distance of the line segment connecting the coordinate position of the object in the second captured image and the coordinate position of the object in the third captured image.
[0050] Here, the third numerical value will be described with reference to FIG. 8(C). FIG. 8(C) is a diagram for explaining the calculation method of the third numerical value. Similar to FIGS. 8(A) and 8(B), in FIG. 8(C), X is the object in the first captured image, Y is the object in the second captured image, and Z is the object in the third captured image. The coordinate position of the object X is (x, y) = (X1, X2), the coordinate position of the object Y is (x, y) = (Y1, Y2), and the coordinate position of the object Z is (x, y) = (Z1, Z2). The calculation method of the third numerical value calculates the half distance L2 of the line segment connecting the coordinate positions (X1, X2) of the object X in the first captured image and the coordinate positions (Z1, Z2) of the object Z in the third captured image. Next, the distance L4 of the line segment connecting the coordinate positions (Y1, Y2) of the object Y in the second captured image and the coordinate positions (Z1, Z2) of the object Z in the third captured image is calculated. The third numerical value is derived by calculating the absolute value |L2 - L4| of the difference in distance between the distance L2 and the distance L4.
[0051] Regarding the association of objects based on the third numerical value, it is the same as the second numerical value. First, it is considered that the same object has moved among the three captured images. Since the movement trajectory of the object is presumed to be linear, the half distance L2 of the distance obtained by connecting the coordinate positions (X1, X2) of the object X in the first captured image and the coordinate positions (Z1, Z2) of the object Z in the third captured image with a line segment, and the distance L4 between the coordinate positions (Y1, Y2) of the object Y in the second captured image and the coordinate positions (Z1, Z2) of the object Z in the third captured image, are presumed to take similar values. Therefore, in the calculation method of the third numerical value, the absolute value of the difference between the half distance L2 and the distance L4 is calculated, and based on the result, the association of objects among the three captured images can be performed.
[0052] Also, when the absolute value of the difference is the shortest (the third numerical value is the minimum value), that combination is presumed to be the movement of the object among the three captured images. Therefore, the association of objects among the three captured images can be performed based on the combination where the third numerical value is the minimum value.
[0053] The specifying unit 15 associates the objects between the captured images based on the extraction result of the extraction unit 14. The extraction result of the extraction unit 14 refers to the first numerical value, the second numerical value, and the third numerical value.
[0054] The specifying unit 15 associates the objects between the captured images based on the first numerical value, the second numerical value, and the third numerical value.
[0055] Processing is performed based on the three numerical values of the first numerical value, the second numerical value, and the third numerical value, so that the accuracy of object association among the three captured images is improved compared to the processing by each individual numerical value.
[0056] Specifically, for each possible combination of the object paths among the captured images extracted by the extraction unit 14, the specific unit 15 calculates the first numerical value, the second numerical value, and the third numerical value, and calculates a combination in which the first numerical value, the second numerical value, and the third numerical value respectively become the minimum values.
[0057] By performing object association based on the combination in which each of the first numerical value, the second numerical value, and the third numerical value becomes the minimum value, the accuracy of object association is further improved.
[0058] More specifically, the specific unit 15 associates the objects among the captured images based on the sum of the first numerical value, the second numerical value, and the third numerical value based on the combination that becomes the above-mentioned minimum value.
[0059] By performing object association based on the sum based on the combination in which each of the first numerical value, the second numerical value, and the third numerical value becomes the minimum value, the accuracy of object association is further improved.
[0060] According to the object tracking device 10, a possible combination of the object paths is extracted from the coordinate positions of the object among the three captured images, and based on the extracted result, the same object is associated, so that the object can be tracked without misidentifying it as another object.
[0061] In the case of two captured images, only the positional relationship with the coordinate positions of two objects is processed. In contrast, in the case of three captured images, the positional relationship between the coordinate position of the first captured image and the coordinate position of the second captured image, the positional relationship between the coordinate position of the second captured image and the coordinate position of the third captured image, and the positional relationship between the coordinate position of the first captured image and the coordinate position of the third captured image are used to perform the association of objects between each captured image. Therefore, the accuracy of the association is further improved.
[0062] In this embodiment, the object tracking device 10 executes a process of associating the same vehicle from the captured images, and calculates the number of vehicles and the speed of the vehicle based on the above-mentioned associating process. The specific process of the object tracking device 10 will be described in detail later.
[0063] [Operation of Object Tracking Device] FIGS. 3 and 4 are flowcharts showing an example of the processing procedure of the object tracking device 10, and will be described with reference to FIGS. 5 to 8 along the flowcharts of FIGS. 3 and 4. FIG. 5 is a diagram showing the captured images from the first frame to the fourth frame. The first frame is the first captured image at the start of the operation of the object tracking device 10. Also, the first frame to the third frame shown in FIG. 5 will be described as the first round, and the second frame to the fourth frame will be described as the second round.
[0064] <First round> First, the first round will be described. First, the acquisition unit 11 acquires the first captured image from the imaging unit 20 or the storage unit 30 (step 1). In the first round, the first captured image corresponds to one frame.
[0065] Subsequently, after the acquisition unit 11 acquires the first captured image, the analysis unit 12 identifies the objects in the first captured image (step 2). For the detection of objects in the first captured image, a known technique for detecting objects from a still image can be adopted. The object 5a in the first frame of FIG. 5 is identified by the analysis unit 12.
[0066] Subsequently, the detection unit 13 detects the coordinate position of the object in the first captured image specified by the analysis unit 12 (step 3).
[0067] FIG. 6 is a diagram showing the center of gravity of the object in FIG. 5, and (A) is a diagram showing the first to third frames.
[0068] Here, an explanation of FIG. 6(A) will be given. t represents the time of the latest captured image when three captured images are acquired, t - 2 indicates the oldest time in the three acquired captured images in terms of time series, and t - 1 indicates that it is between t and t - 2 in terms of time series. In the present embodiment, t - 2 corresponds to the first frame, t - 1 corresponds to the second frame, and t corresponds to the third frame.
[0069] Also, the object existing in lane R1 is denoted as object 5a, and the object existing in lane R2 is denoted as object 5b. That is, the coordinate position of the object detected in the first frame is denoted as 5a1, the coordinate position of the object detected in the second frame is denoted as 5a2, 5b2, and the coordinate position of the object detected in the second frame is denoted as 5a3, 5b3. Here, in the following description, although denoted as objects 5a and 5b, it is for convenience of explanation that symbols are attached, and the objects are not already distinguished at the stages of step 2 and step 3.
[0070] That is, as shown in FIG. 6(A), the detection unit 13 detects that the object 5a in the first frame is G(t - 2, 5a1) with respect to the center of gravity position.
[0071] Then, the accumulation number (the number of captured images whose coordinate positions are detected by the detection unit 13) is counted (step 4). When the accumulation number reaches 3, it proceeds to step 5, and when it is less than 3, steps 1 to 3 are repeated until it becomes 3. Hereinafter, steps 1 to 3 are executed for the second frame and the third frame in the same manner as the first frame until the accumulation number reaches 3.
[0072] Next, in the same way as the first frame, first, the acquisition unit 11 acquires a second captured image from the imaging unit 20 or the storage unit 30 (step 1). In the first round, the second captured image corresponds to two frames.
[0073] Subsequently, when the acquisition unit 11 has acquired the second captured image, the analysis unit 12 identifies the objects in the second captured image (step 2). The objects 5a and 5b in the second frame in FIG. 5 are identified by the analysis unit 12.
[0074] Subsequently, the detection unit 13 detects the coordinate positions of the objects in the second captured image identified by the analysis unit 12 (step 3). As shown in FIG. 6(A), with respect to the center-of-gravity position, the object 5a in the second frame is detected as G(t - 1, 5a2), and the object 5b in the second frame is detected as G(t - 1, 5b2).
[0075] Next, in the same way as the first frame and the second frame, first, the acquisition unit 11 acquires a third captured image from the imaging unit 20 or the storage unit 30 (step 1). In the first round, the third captured image corresponds to three frames.
[0076] Subsequently, when the acquisition unit 11 has acquired the third captured image, the analysis unit 12 identifies the objects in the third captured image (step 2). The objects 5a and 5b in the third frame in FIG. 5 are identified by the analysis unit 12.
[0077] Subsequently, the detection unit 13 detects the coordinate positions of the objects in the third captured image identified by the analysis unit 12 (step 3). As shown in FIG. 6(A), with respect to the center-of-gravity position, the objects in the third frame are detected as G(t, 5a3) and G(t, 5b3).
[0078] Then, when the coordinate positions are detected by the detection unit 13 from the first frame to the third frame, since the accumulation count reaches 3, the process proceeds to step 5.
[0079] Here, first, the processing of steps 5 to 7 will be described. In Steps 5 to 7, if an object that has not been extracted as a possible combination of the paths of the object among the captured images is detected, it is determined that detection has occurred, and the process proceeds to the next step. At this time, in Steps 5 to 7, if an object is detected in each captured image, the next step is executed each time one object is detected. That is, Steps 5 to 7 are executed for each possible combination of the paths of the object among the captured images. A more specific explanation will be described later.
[0080] First, the extraction unit 14 extracts the object in the first captured image (Step 5). The extraction unit 14 determines that detection has occurred if the object in the detected first captured image (the first frame) is an object that has not been detected as a possible combination of the paths of the object among the captured images. In the case of FIG. 6(A), the object G(t - 2, 5a1) in the first frame is determined to have been detected.
[0081] Subsequently, if it is determined that detection has occurred in Step 5, the extraction unit 14 detects the object in the second captured image (the second frame) (Step 6). In the case of FIG. 6(A), in the second frame, the objects G(t - 1, 5a2) and G(t - 1, 5b2) are detected. Here, if either object is detected, the process proceeds to Step 7. Here, it is assumed that G(t - 1, 5a2) is selected. The selection method at this time is not particularly limited. For example, an object at a distance closer to the origin in the captured image may be detected first.
[0082] Subsequently, if it is determined that detection has occurred in Step 6, the extraction unit 14 extracts the object in the third captured image (the third frame) (Step 7). In the case of FIG. 6(A), in the third frame, the objects G(t, 5a3) and G(t, 5b3) are detected. Here, if either object is detected, the process proceeds to Step 8. Here, it is assumed that G(t, 5a3) is selected.
[0083] According to the above steps 5 to 7, the first set (G(t - 2, 5a1), G(t - 1, 5a2), G(t, 5a3)) of possible combinations of the object's path among the captured images is extracted.
[0084] Subsequently, for the initial value setting in step 8, it is determined whether to assign an initial value if the combination of the object is a combination contrary to the result of the previous process of the currently executed process, for example, if the second - round process is being executed, it is contrary to the result of the first - round process. In the first round, since there is no previous process, there is no contrary combination, and the initial value cannot be assigned.
[0085] After step 8, the extraction unit 14 calculates whether the combination has an initial value assigned (step 9). Regarding the first set in the first round, since no initial value is assigned, it proceeds to step 10.
[0086] Then, for the first set, the first numerical value, the second numerical value, and the third numerical value are calculated (step 10). FIG. 7 is a table showing an example of the first numerical value, the second numerical value, the third numerical value, and their total value. (A) is a table recording from frame 1 to frame 3.
[0087] Explanation of the first numerical value of the first set. The first numerical value of the first set is calculated as the coordinate position T((5a1 + 5a3)÷2) of the mid - point of the line segment connecting the coordinate position 5a1 of the object 5a within 1 frame and the coordinate position 5a3 of the object 5a within 3 frames. The distance L1, which is the distance between the above - mentioned coordinate position T and the coordinate position 5a2 of the object 5a within 2 frames, is calculated and derived.
[0088] Explanation of the second numerical value of the first set. The second numerical value of the first set calculates the half distance L2 of the line segment connecting the coordinate position 5a1 of the object 5a within one frame and 5a3 of the object 5a within three frames. Next, the half distance L3 of the line segment connecting the coordinate position 5a1 of the object 5a within one frame and the coordinate position 5a2 of the object 5a within two frames is calculated. It is derived by calculating the absolute value |L2 - L3| of the difference in distance between the distance L2 and the distance L3.
[0089] An explanation will be given for the third numerical value of the first set. The third numerical value of the first set calculates the half distance L2 of the line segment connecting the coordinate position 5a of the object 5a within one frame and the coordinate position 5a of the object 5a within three frames. Next, the half distance L4 of the line segment connecting the coordinate position 5a2 of the object 5a within two frames and the coordinate position 5a3 of the object 5a within three frames is calculated. It is derived by calculating the absolute value |L2 - L4| of the difference in distance between the distance L2 and the distance L4.
[0090] After calculating the first numerical value, the second numerical value, and the third numerical value in step 10, the specifying unit 15 calculates the total value of the first numerical value, the second numerical value, and the third numerical value (step 11).
[0091] In step 11, after calculating the total value of the first set, return to step 7 and perform the process of step 7. In this step 7, objects that have not been detected as possible combinations of the paths of the objects between each captured image in the third frame are extracted. That is, the objects of G(t, 5b3) that have not been extracted as a combination in the first set are extracted. In this case, the combination of the second set becomes (G(t - 2, 5a1), G(t - 1, 5a2), G(t, 5b3)). Hereinafter, the processes from step 8 to step 11 are performed in the same manner as in the first set.
[0092] After the end of step 11, return to step 7 again. Here, in the third frame, since there are no objects that have not been detected as possible combinations of the paths of the objects between the captured images, return to step 6. In step 6, objects that have not been detected as possible combinations of the paths of the objects between the captured images in the second frame are extracted. That is, the objects of G(t-1, 5b2) that have not been extracted as combinations in the first and second sets are extracted and proceed to step 7.
[0093] In step 7, the objects of G(t, 5a3) and G(t, 5b3) are cited as the objects in the third frame. Here, for example, the object of G(t, 5a3) is selected. In this case, the combination of the third set is (G(t-2, 5a1), G(t-1, 5b2), G(t, 5a3)). Hereinafter, the processes from step 8 to step 11 are performed in the same manner as in the first and second sets.
[0094] After the end of step 11, return to step 7 again. In this step 7, objects that have not been detected as possible combinations of the paths of the objects between the captured images in the third frame are extracted. That is, the objects of G(t, 5b3) that have not been extracted as combinations in the first to third sets are extracted. In this case, the combination of the fourth set is (G(t-2, 5a1), G(t-1, 5b2), G(t, 5b3)). Hereinafter, the processes from step 8 to step 11 are performed in the same manner as in the first to third sets.
[0095] After the end of step 11 of the fourth set, return to step 7. Here, in the third frame, since there are no objects that have not been detected as possible combinations of the paths of the objects between the captured images, return to step 6. Also, in the second frame, since there are no objects that have not been detected as possible combinations of the paths of the objects between the captured images, return to step 5. Furthermore, in the first frame, since there are no objects that have not been detected as possible combinations of the paths of the objects between the captured images, shift to A and perform step 12 described in FIG. 4.
[0096] Subsequently, when the total value is calculated for each combination of the possible paths of the object in each captured image, the specifying unit 15 calculates whether there are any remaining combinations (step 12).
[0097] Step 12 for the combinations will be described with reference to FIG. 7(A). FIG. 7 is a table showing an example of the first numerical value, the second numerical value, the third numerical value, and their total value, and (A) is a table showing the first to third frames. The vertical axis in FIG. 7(A) indicates the coordinate position of the object in the third frame, and the horizontal axis indicates the combinations of the object in the first frame and the object in the second frame. The combinations are the combinations including G(t, 5a3) on the vertical axis and the combinations including G(t, 5b3). When there is a remaining combination for which steps 13 to 15 described later have not been performed, it is processed as having a remaining combination.
[0098] When there is a remaining combination, the specifying unit 15 calculates the combination with the lowest value (step 13). For example, the combination at (t, 5a3) is the first set (G(t - 2, 5a1), G(t - 1, 5a2), G(t, 5a3)) and the third set (G(t - 2, 5a1), G(t - 1, 5b2), G(t, 5a3)). Comparing the first set and the third set, since the total value of the first set is lower, the first set is calculated. Here, a known technique may be used to calculate the combination with the lowest value. A known technique is, for example, the Hungarian method.
[0099] Subsequently, the specifying unit 15 calculates whether the combination with the lowest value is equal to or greater than a predetermined threshold value (a value lower than the initial value, 8 in this embodiment) or less than it (step 14). The threshold value is not limited to 8 and is appropriately determined according to the angle of view of the captured image, the frame rate, etc. Here, it is calculated whether the combination with the lowest value is the first set and whether the first set is less than the threshold value. In this embodiment, since the first set is less than the threshold value, the process proceeds to step 15.
[0100] Subsequently, the combination is set as a combination candidate on the assumption that the object is associated between the captured images (step 15).
[0101] Subsequently, return to step 12. Regarding the combination, since there are remaining combinations including G(t, 5b3) that have not been examined, there are remaining combinations, and proceed to step 13.
[0102] In step 13, the combination at (t, 5b3) is the combination of the second set (G(t - 2, 5a1), G(t - 1, 5a2), G(t, 5b3)) and the fourth set (G(t - 2, 5a1), G(t - 1, 5b2), G(t, 5b3)). Comparing the second set and the fourth set, since the total value of the second set is lower, the second set is calculated.
[0103] Subsequently, the specifying unit 15 calculates whether the combination with the lowest value is equal to or greater than a predetermined threshold value (a value lower than the initial value, 8 in this embodiment) (step 14). Here, the combination with the lowest value is calculated as the second set, and it is calculated whether the second set is less than the threshold value. In this embodiment, since the second set is equal to or greater than the threshold value, it does not become a combination candidate and returns to step 12.
[0104] By setting the threshold value, it is possible to exclude combinations that are not correct as the association of the object but become the combination with the lowest value.
[0105] In step 12, since the number of remaining combinations with the lowest value has become zero, the combination is determined (step 16). The determined combination becomes the first set of combinations.
[0106] Thereafter, based on the determined combination, the specifying unit 15 calculates the speed and number of objects in each captured image (step 17).
[0107] Regarding the calculated speed and number of objects, as described above, they are transmitted from the communication unit 40 to the traffic control center 200.
[0108] In addition, in the present embodiment, although the control device 100 is illustrated as a road without traffic lights such as highways for roads 1 and 2, it is not limited thereto, and it may be installed on a road having intersections such as general roads. Thereby, the number and speed of vehicles calculated by the object tracking device 10 can be known, and the traffic situation can be grasped. Therefore, the display time of the traffic light can be set to an appropriate length according to the traffic situation, and for example, traffic congestion can be alleviated. Further, the control device 100 may be provided with a device for controlling the display time of the traffic light in order to control the display time of the traffic light described above.
[0109] <Second round> Next, the process of the second round will be described. First, similar to the first round, the acquisition unit 11 acquires the first captured image from the imaging unit 20 or the storage unit 30 (step 1). Here, in the second round, the first captured image is the second frame, the second captured image is the third frame, and the third captured image is the fourth frame.
[0110] Subsequently, similar to the first round, the analysis unit 12 identifies the objects in each captured image (step 2). Objects 5a and 5b are identified as the objects in the second frame of FIG. 5.
[0111] Subsequently, the detection unit 13 detects the coordinate positions of the objects in the captured image identified by the analysis unit 12 (step 3).
[0112] FIG. 6 is a diagram showing the center of gravity of the objects in FIG. 5, and (B) is a diagram showing from the second frame to the fourth frame.
[0113] Here, an explanation of FIG. 6(B) will be given. In the present embodiment, t-2 corresponds to the second frame, t-1 corresponds to the third frame, and t corresponds to the fourth frame.
[0114] That is, as shown in FIG. 6(B), the detection unit 13 detects that the object 5a in the second frame is G(t-2, 5a2) and the object 5b in the second frame is G(t-2, 5b2) with respect to the center of gravity position.
[0115] Then, in the same manner as in the first round, when the coordinate positions are detected by the detection unit 13 from the second frame to the fourth frame, since the accumulation number reaches 3, the process proceeds to step 5.
[0116] The extraction unit 14 extracts the object in the first captured image (step 5). When the object in the detected first captured image (second frame) is not detected as a possible combination of the object's path among the captured images, it is determined that detection is successful. In the case of FIG. 6(B), G(t-1, 5a1) and G(t-2, 5b2) of the object in the second frame are detected. Here, either object is detected, and the process proceeds to step 7. Here, it is assumed that G(t-2, 5a2) is selected.
[0117] Subsequently, when detection is successful in step 5, the extraction unit 14 detects the object in the second captured image (third frame) (step 6). In the case of FIG. 6(B), in the third frame, G(t-1, 5a3) and G(t-1, 5b3) of the object are detected. Here, either object is detected, and the process proceeds to step 7. Here, it is assumed that G(t-1, 5a3) is selected.
[0118] Subsequently, when detection is successful in step 6, the extraction unit 14 extracts the object in the third captured image (fourth frame) (step 7). In the case of FIG. 6(B), in the fourth frame, G(t, 5a4), G(t, 5b4), and G(t, 5a5) of the object are detected. Here, either object is detected, and the process proceeds to step 8. Here, it is assumed that G(t, 5a4) is selected.
[0119] According to the above steps 5 to 7, the first set in the second round as a possible combination of the object's path among the captured images is (G(t-2, 5a2), G(t-1, 5a3), G(t, 5a4)). Similarly, the same process is performed for other combinations.
[0120] Subsequently, for the initial value setting in step 8, in the second round, since the process one round before is the first round, initial values are assigned to combinations contrary to the first round.
[0121] Here, combinations that are contrary to the first round will be described. A combination contrary to the first round refers to a combination that is contrary to the combination determined in the first round.
[0122] The combination of the first round and the combination of the second round have the second frame and the third frame overlapping. Therefore, an initial value is assigned to the combination of the second round that is contrary to the second frame and the third frame of the determined combination of the first round. The determined combination of the first round is the first set of (G(t - 2, 5a1), G(t - 1, 5a2), G(t, 5a3)). That is, an initial value is assigned to a combination where the second frame and the third frame of the combination of the second round are (G(t - 2, 5a2), G(t - 1, 5b3)), (G(t - 1, 5b2), G(t, 5a3)). For example, a combination of (G(t - 2, 5a2), G(t - 1, 5b3), G(t, 5a4)) can be cited.
[0123] Here, the initial value only needs to be a value larger than the numerical value that can be calculated by the combination. For example, it is 1000.
[0124] After step 8, the extraction unit 14 calculates whether the combination has an initial value assigned (step 9). Since the combination of (G(t - 2, 5a2), G(t - 1, 5b3), G(t, 5a4)) in the second round has an initial value assigned in step 8, it does not proceed to step 10, but returns to step 7, and objects that have not been detected as possible combinations of the path of the object between each captured image in the fourth frame are extracted, and the processes of step 8 and step 9 are performed again. Hereinafter, regarding steps 10 to 11, the same processing as in the first round is performed to calculate the total value.
[0125] Since the initial value is larger than the numerical value that can be calculated by the combination, it becomes easy to determine whether the initial value is assigned in step 10, and it is possible to move to the next combination without having to specifically calculate the first numerical value, the second numerical value, and the third numerical value.
[0126] Regarding the combination in step 12, it will be described with reference to FIG. 7(B). FIG. 7 is a table showing an example of the first numerical value, the second numerical value, the third numerical value, and their total value, and (B) is a table showing the second frame to the fourth frame. The vertical axis of FIG. 7(b) indicates the coordinate position of the object in the fourth frame, and the horizontal axis indicates the combination of the object in the second frame and the object in the third frame. The combination refers to the combination including G(t, 5a4) in the third frame which is the vertical axis, the combination including G(t, 5b4), and the combination including G(t, 5a5). When there is a remaining combination for which steps 13 to 15 described later are not performed, it is processed as having a remaining combination.
[0127] And when there is a remaining combination, the specifying unit 15 calculates the combination with the lowest value (step 13). For example, when calculating the combination at (t, 5a4), since the total value of (G(t - 2, 5a2), G(t - 1, 5a3), G(t, 5a4)) is low, this combination is calculated.
[0128] Subsequently, the specifying unit 15 calculates whether the combination with the lowest value is equal to or greater than a predetermined threshold value (a value lower than the initial value, 8 in this embodiment) or less than it (step 14). Here, it is calculated whether the combination with the lowest value is less than the threshold value. Since it is less than the threshold value in this embodiment, it proceeds to step 15.
[0129] Subsequently, that combination is set as a combination candidate on the assumption that the objects are associated between the respective captured images (step 15).
[0130] Subsequently, it returns to step 12. Regarding the combinations, for the combinations including G(t, 5b4) and the combinations including G(t, 5a5) that have not been considered, steps 13 and subsequent steps are similarly performed. Based on the determined combinations, the specifying unit 15 calculates the speed and the number of objects in each captured image (step 17).
[0131] <Modification Example> In the above description, the initial value setting (step 8) was provided between step 7 and step 9, but the initial value setting (step 8) may also be provided between step 4 and step 5.
[0132] In this case, the coordinate positions of the object by the detection unit 13 from the second frame to the fourth frame are detected, and when the accumulation count reaches 3 (step 4), the initial value is set (step 8). Here, a value is assigned in advance to all possible combinations for the initial value. For example, the initial value is 1000 as described above.
[0133] Subsequently, after executing steps 5 to 7 described above, the extraction unit 14 determines whether each of the combinations extracted in steps 5 to 7 violates the first round (step 9).
[0134] Subsequently, if the above-described combination is not a combination that violates the first round, the first numerical value, the second numerical value, and the third numerical value are calculated, and a process of overwriting the above-described initial value with the first numerical value, the second numerical value, and the third numerical value is performed (step 10). On the contrary, if the above-described combination is a combination that violates the first round, it returns to step 7 with the above-described initial value unchanged.
[0135] Although the embodiments of the present invention have been described above, the present invention is not limited to only the above-described embodiments, and it goes without saying that various modifications can be made.
[0136] In this embodiment, the control device 100 is installed on the road, but it is not limited thereto. For example, the imaging unit 20 and the communication unit 40 may be installed on the road, and the object tracking device 10 and the storage unit 30 may be arranged in, for example, a traffic control center 200, which is different from the road.
[0137] For example, in the present embodiment, the first numerical value, the second numerical value, and the third numerical value are calculated, and based on these, the specifying unit 15 performs the association of the object between the captured images. However, the present invention is not limited to this, and the association of the object between the captured images may be performed based on only the first numerical value, or a combination of the first numerical value and the second numerical value, or a combination of the first numerical value and the third numerical value. By executing the processing according to the flowcharts of FIGS. 3 and 4 also by the above combination, the association of the object between the captured images can be performed.
Explanation of Signs
[0138] 10… Object tracking device 11… Acquisition unit 12… Analysis unit 13… Detection unit 14… Extraction unit 141… First processing unit 142… Second processing unit 143… Third processing unit 15… Specifying unit 20… Imaging unit 30… Storage unit 40… Communication unit 100… Control device 200… Traffic control center
Claims
1. An acquisition unit that acquires a first captured image, a second captured image, and a third captured image in chronological order; An object position detection unit that specifies the coordinate positions of one or more objects within each captured image; An extraction unit that extracts possible combinations of the paths of the objects between the respective captured images; An identification unit that associates the objects between the respective captured images based on the extraction result of the extraction unit and comprising: The extraction unit has a first processing unit that calculates a first numerical value which is the distance between the coordinate position of the object in the second captured image and the coordinate position of the midpoint of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image; The identification unit associates the objects between the respective captured images based on the first numerical value Object tracking device.
2. The object tracking device according to claim 1, wherein the extraction unit further has a second processing unit that calculates a second numerical value which is the absolute value of the difference between the half distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image and the distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the second captured image; The identification unit associates the objects between the respective captured images based on the first numerical value and the second numerical value Object tracking device.
3. The object tracking device according to claim 1, wherein the extraction unit further has a third processing unit that calculates a third numerical value which is the absolute value of the difference between the half distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image and the distance of the line segment connecting the coordinate position of the object in the second captured image and the coordinate position of the object in the third captured image; The identification unit associates the objects between the respective captured images based on the first numerical value and the third numerical value Object tracking device.
4. The object tracking device according to claim 2, wherein the extraction unit further has a third processing unit that calculates a third numerical value which is the absolute value of the difference between the half distance of the line segment connecting the coordinate position of the object in the first captured image and the coordinate position of the object in the third captured image and the distance of the line segment connecting the coordinate position of the object in the second captured image and the coordinate position of the object in the third captured image (the "existence" in the original seems to be a typo and should be "second" here); The specific unit associates the object between the captured images based on the first numerical value, the second numerical value, and the third numerical value. Object tracking device.
5. The object tracking device according to claim 4, For each combination extracted by the extraction unit, the specific unit calculates the first numerical value, the second numerical value, and the third numerical value, and calculates a combination in which the first numerical value, the second numerical value, and the third numerical value are each the minimum value. Object tracking device.
6. The object tracking device according to claim 5, The specific unit associates the object between the captured images based on the sum of the first numerical value, the second numerical value, and the third numerical value based on the combination that becomes the minimum value. Object tracking device.
7. The object tracking device according to claim 5 or 6, When the minimum value is less than a predetermined threshold value, the specific unit associates the object between the captured images based on the combination that becomes the minimum value. Object tracking device.
8. Acquire a first captured image, a second captured image, and a third captured image in chronological order, Specify the coordinate positions of one or more objects within each captured image, Extract, by an extraction unit, possible combinations of the paths of the object between the captured images, Based on the extraction result of the extraction unit, associate the object between the captured images, Calculate a first numerical value that is the distance from the coordinate position of the midpoint of the line segment connecting the coordinate position of the object in the second captured image, the coordinate position of the object in the first captured image, and the coordinate position of the object in the third captured image of the object, Associate the object between the captured images based on the first numerical value. Object tracking method.
9. A step of acquiring a first captured image, a second captured image, and a third captured image in chronological order, A step of specifying the coordinate positions of one or more objects within each captured image, A step of extracting, by an extraction unit, possible combinations of the paths of the object between the captured images, A step of associating the object between the captured images based on the extraction result of the extraction unit, A step of calculating a first numerical value that is the distance from the coordinate position of the midpoint of the line segment connecting the coordinate position of the object in the second captured image, the coordinate position of the object in the first captured image, and the coordinate position of the object in the third captured image of the object. A step of associating the object between the captured images based on the first numerical value A program for causing the execution.
Citation Information
Patent Citations
Device and method for making areas correspond to each other
JP1997161071A
Object tracking device
JP2010218232A
Device and method for tracking mobile object and program
JP2020057424A