Object detection device, object detection method, and program
The object detection device addresses collision prevention by estimating distances and speeds between objects, issuing warnings to ensure safe interactions, enhancing operational safety in environments such as warehouses.
Patent Information
- Application Number
- JP2023223466
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing methods for estimating distances between objects do not prevent collisions, and there is a need for systems that can issue warnings and control object interactions to avoid potential collisions.
An object detection device that inputs captured images, synthesizes bounding boxes around objects, estimates distances and speeds, and issues warnings when objects approach too closely or exceed safe speeds/angles, using a learning model to track and control interactions.
Effectively prevents collisions by providing visual and intuitive warnings for potential dangers, ensuring safe operation of objects in environments like warehouses.
Smart Images

Figure 2025105137000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object detection device, an object detection method, and a program.
Background Art
[0002] Conventionally, a method of estimating the distance between objects without using a ranging sensor or the like has been studied. For example, in "Problems to be Solved by the Invention" of Patent Document 1 (Japanese Unexamined Patent Application Publication No. 2021-056717), a technique of estimating the distance between objects by converting two-dimensional coordinates into three-dimensional coordinates with the coordinates of the center of the lower end of the bounding box as the position of the object is disclosed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As described in Patent Document 1, the distance between objects can be estimated, but collisions between objects cannot be prevented.
Means for Solving the Problems
[0005] The object detection device according to the first aspect inputs, using a learning image in which an object is captured, an image including a region in which the object is captured from an arbitrary moving image, stores an object identified by a plurality of the learning images, outputs an image in which a bounding box corresponding to the region in which the object is captured is synthesized, estimates the distance between a plurality of objects between the bounding boxes, and issues a warning and controls when the plurality of objects approach each other at a certain distance.
[0006] The object detection device from the second perspective inputs, using a learning image in which an object is captured, an image including a region in which the object is captured from an arbitrary moving image, stores the objects identified by a plurality of the learning images, outputs an image in which a bounding box corresponding to the region in which the object is captured is synthesized, estimates the object speed using the bounding box of each frame of the moving image, and issues a warning and controls when a certain speed is exceeded.
[0007] The object detection device from the third perspective inputs, using a learning image in which an object is captured, an image including a region in which the object is captured from an arbitrary moving image, stores the objects identified by a plurality of the learning images, outputs an image in which a bounding box corresponding to the region in which the object is captured is synthesized, estimates the rotation angle of the object using the bounding box of each frame of the moving image, and issues a warning and controls when a certain angle is exceeded.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0009] <1> Configuration of the Object Detection Device Hereinafter, the configuration of the object detection device according to an embodiment of the present disclosure will be described with reference to the drawings. FIG. 1 is a schematic diagram showing the configuration of the object detection device 20 according to the present embodiment. The object detection device 20 is a device for avoiding collisions between objects O.
[0010] As shown in FIG. 2, the region of the object O is defined by the coordinate information b1 to b4 corresponding to the four vertices of the bounding box B synthesized in the image.
[0011] In the example shown in FIG. 2, a "forklift" and a "person" are shown as the object O, but the object O is not limited to this. Any object can be adopted as the object O. In addition, the object O can be set by distinguishing not only the type of the object but also the state and the like.
[0012] The object detection device 20 can be realized by any computer and includes a storage unit 21, an input unit 22, an output unit 23, and a processing unit 24. Note that the object detection device 20 may also be realized as hardware using an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.
[0013] The storage unit 21 stores various information and is realized by any storage device such as a memory and a hard disk.
[0014] The input unit 22 is realized by any input device such as a keyboard, a mouse, and a touch panel, and inputs various information to the computer.
[0015] The output unit 23 is realized by any output device such as a display, a touch panel, and a speaker, and outputs various information from the computer.
[0016] The processing unit 24 executes various information processes, and is realized by a processor such as a CPU or a GPU, and a memory. Here, when one or more programs stored in the storage unit 21 are read into the CPU, GPU, etc. of the computer, the processing unit 24 functions as a generation unit 24A, a synthesis unit 24B, a setting unit 24C, and a control unit 24D. Hereinafter, each function of the processing unit 24 will be described.
[0017] The generation unit 24A generates coordinate information of the area where the object O is imaged. Further, the generation unit 24A generates a file in which the coordinate information b1 to b4 of the vertices of the bounding box B corresponding to the area where the object O is imaged is described. Note that the coordinate information b1 to b4 can be defined by the two-dimensional coordinates of each vertex. However, not limited thereto, the coordinate information b1 to b4 can also be defined by the two-dimensional coordinates of one vertex and the width and height from that vertex on the premise that the bounding box B is a square or a rectangle. In the former case, a file in which eight values corresponding to four vertices on the two-dimensional coordinates are described is generated. In the latter case, a file in which a total of four values, two values corresponding to one vertex on the two-dimensional coordinates and two values indicating the width and height therefrom, are described is generated.
[0018] The synthesis unit 24B synthesizes the bounding box B and displays it on the display of the output unit 23.
[0019] The setting unit 24C generates an arbitrary bounding box B in the image by specifying coordinates, and sets that the object image in which the object O is imaged using the bounding box B.
[0020] The control unit 24D estimates the distance between the respective bounding boxes B so that the objects O do not collide with each other, and issues a warning and controls when a collision is likely to occur. In addition, a warning is issued also at high speed and high-speed rotation.
[0021] <2>Operation of the object detection device The operation of the object detection device 20 according to this embodiment will be described.
[0022] (1) Learning of Visual Detection Model: 1. Data Preparation: Data focused on forklift operations within the warehouse environment was collected. This collection included forklifts in various directions and lighting conditions. During the labeling process, each image was annotated to emphasize the forklift. This annotation is necessary for the model to recognize and distinguish the forklift within the environment.
[0023] 2. Model Learning: An advanced YoloX visual detection model, known for its efficiency and accuracy in object detection models, was selected. To significantly shorten the learning time and improve the detection accuracy, learning was carried out by leveraging a pre-trained model. Here, the labeled data was used as the model. The pre-trained YoloX model was fine-tuned to accurately identify and distinguish forklifts and humans on the data. This is very important as it prevents overfitting while optimizing the generalization function of the model, ensuring reliable performance in actual applications.
[0024] 3. Results and Visualization: The detection results are visualized through a bounding box B that encloses the object O identified within the frame.
[0025] (2) Visual Tracking Model: 1. Model Selection: The optimized OC-SORT tracking model, an algorithm known for its effectiveness in multiple object tracking environments, is utilized. This selection is based on performance characteristics suitable for warehouse monitoring.
[0026] 2. Identification and Labeling of Objects: During frame processing, an identification number (ID) is assigned to each detected object O, and consistent tracking is performed throughout the sequence. This is to maintain the continuity of tracking data between frames. This system records a label that specifies the type of the identified object O, specifically whether it is a "person" or a "forklift".
[0027] 3. Tracking Continuity and Data Management: Monitor the assigned ID and label across consecutive frames. Adopt the OC-SORT tracking model to achieve visual tracking within the warehouse.
[0028] (3) Estimation of Warehouse Measurements from Camera Input Only: 1. Conversion to Perspective View: Here, a technique known as "perspective transformation" is used to associate two-dimensional video image coordinates (measured in pixels) with the actual three-dimensional space of the warehouse (usually in meters).
[0029] 2. Definition of the Ground Surface: In the video image, identify the ground surface area and depict its contour as a trapezoid. This distorted shape is due to the camera's perspective. This step is important for marking the actual monitored target area. While the trapezoid is distorted by perspective, recognize that the actual warehouse floor is square or rectangular. This is the area for mapping the observations.
[0030] 3. Calculation of Homographic Transformation: This coordinate transformation involves the calculation of a "homography matrix". This matrix is a 3×3 matrix between the distorted (trapezoidal) coordinates and the actual (square) coordinates.
[0031] 4. Application of the Transformation: Apply the homographic transformation to all points of interest within the trapezoid. Convert the coordinates from the video's pixel-based system to the actual coordinates of the warehouse floor. This conversion is very important for accurate activity monitoring, distance measurement, and spatial recognition during the automatic video surveillance process.
[0032] (4) Proximity estimation: Calculate the distance between a person and a forklift, and between forklifts. Refer to Figure 3. To avoid proximity risks between a person and a forklift, and between forklifts, a systematic distance calculation is adopted. 1. Identification of reference points: Extract the central lower part of the bounding box B that encloses a person or a forklift within the frame. This is set as the reference position of each object O.
[0033] 2. Distance estimation: Draw line segments between each pair of the identified reference points, especially between a person and a forklift, and between two forklifts. The length of this line segment calculated using actual coordinate measurements represents the actual distance separating the two objects.
[0034] (5) Estimation of forklift speed (refer to Figure 4): To estimate the speed of a forklift within the video, a series of procedures including coordinate transformation and distance and speed calculations are performed. 1. Distance calculation: The central lower part of the bounding box B of the forklift within the frame is set as the reference position of the moving forklift. Calculate the displacement of the forklift by measuring the distance this point moves between two consecutive frames (frame n - 1 and frame n). Here, calculate the difference in the positions of the corresponding actual points within these frames in terms of coordinates (in meters).
[0035] 2. Speed calculation: Since the frames are consecutive, this interval is the reciprocal of the frame rate of the video, and calculate the time interval between two frames. Specifically, if the video is recorded at "f" frames per second, the time between frames is 1 / f seconds. Divide the calculated distance by the time interval. Since the time interval is 1 / f, the formula becomes as follows. Speed = Distance * f This formula calculates the speed of the forklift in meters per second, assuming that the distance is measured in meters and the frame rate is the number of frames per second.
[0036] (6) Estimation of the moving direction of the forklift (see Fig. 5): 1. Evaluation of displacement: Focus on the bounding box B of the forklift within the frame. Specifically, pay attention to the center point of the bounding box B of the forklift. Note the displacement of this center point between two specific frames (frame n and frame n+1). This displacement is the data necessary to determine the operating direction of the forklift. This is because using the center point of the bounding box B is less affected by changes in the shape of the bounding box B, and the angle is easy to measure.
[0037] 2. Calculation of the angle: Extract the change in the X coordinate (dx) and the change in the Y coordinate (dy) of the center point of the bounding box B of the forklift between frame n and frame n+1. Use the arctangent function considering the ratio of the vertical displacement to the horizontal displacement to calculate the moving angle α. The formula is as follows. α = arctan(dy / dx) Here, "dy" represents the change in the vertical coordinate, and "dx" represents the change in the horizontal coordinate. Next, the arctangent function converts this ratio into the angle α, representing the moving direction of the forklift on the warehouse floor.
[0038] 3. Periodic calculation: Instead of calculating the direction for each consecutive frame, optimize by performing this calculation at regular intervals. The actual interval is every 5 frames, which can balance the calculation efficiency and sufficient data resolution for accurate direction estimation.
[0039] (7) Warning indication system: 1. Proximity warning: Serious danger: When the distance between the forklift and a person (or another forklift) is less than 3 m, the system warns of serious danger. Here, the bounding box B surrounding the objects O turns red. A red line segment connecting the center of the lower end of the bounding box B is also displayed, visually emphasizing the proximity. Moderate danger: When the measured distance is less than 5 m and more than 3 m, the system indicates a moderate danger level. The bounding box B turns yellow to indicate caution, and a yellow line segment connecting the center of the lower end of the bounding box B is displayed.
[0040] 2. High-speed warning: It activates when the speed of the forklift exceeds 3 m / s. This threshold is considered to be operating beyond the safe speed limit within the warehouse environment. When detected, the system highlights the bounding box B of the forklift in magenta to indicate high-speed operation and immediately displays a "High-speed" warning to draw attention.
[0041] 3. High-speed rotation warning: In the case of operations under more special conditions where the forklift not only moves at a speed exceeding 3 m / s but also turns at an angle exceeding 30 degrees, such maneuvers significantly increase the risk of accidents. In response, the system colors the bounding box B green and issues a "High-speed rotation" warning to inform of this combined danger.
[0042] These visual and intuitive step-by-step warning systems are designed to function as an immediate information reference for supervisors and staff, enabling a prompt response to potential collisions in warehouse operations.
[0043] (8) Others When two objects O and O' intersect and the rear object O' is blocked from the image, its bounding box B is reduced or eliminated. When the bounding box B is shrunk, it means that the object O' is partially blocked from the video. At this time, since the position information of the object O' cannot be obtained, the position information of the object O' before occlusion is recorded, and the bounding box B' is displayed on the display in white based on the last image data before occlusion. When the object O' is displayed on the display again, the position of the object O' is recalibrated using the new data.
[0044] When the bounding box B is erased, it means that the object O' is completely blocked. After that, when the object O' moves and is displayed on the display, the last visible position shown on the display is recorded so that the ID attached to the object O' does not change from before. The bounding box B' is synthesized and displayed on the display in white. In this way, the occluded object O' can be continuously tracked. When the object is displayed on the display again, the system updates the data using the new video.
[0045] <Other Embodiments> The present disclosure is not limited to the above-described embodiments as they are. The present disclosure can be embodied by modifying the components without departing from the gist thereof at the implementation stage. Also, the present disclosure can form various disclosures by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components may be appropriately combined from different embodiments.
Description of Reference Numerals
[0046] 20 Object Detection Device 21 Storage Unit 22 Input Unit 23 Output Unit 24 Processing Unit 24A Generation Unit 24B Composition Unit 24C Setting Unit 24D Control Unit
Claims
1. An input unit that inputs, from an arbitrary moving image, an image including a region in which the object is shown, using a learning image in which the object is shown; A storage unit that stores objects identified by a plurality of the learning images; An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized; A control unit that estimates the distance between a plurality of objects by a line segment between bounding boxes, and issues a warning when the plurality of objects approach a certain distance; An object detection device, characterized by comprising the above.
2. An input unit that inputs, from an arbitrary moving image, an image including a region in which the object is shown, using a learning image in which the object is shown; A storage unit that stores objects identified by a plurality of the learning images; An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized; A control unit that estimates the speed of the object from the moving distance and time of the line segment of the bounding box between frames of the moving image, and issues a warning when the speed exceeds a certain speed; An object detection device, characterized by comprising the above.
3. An input unit that inputs, from an arbitrary moving image, an image including a region in which the object is shown, using a learning image in which the object is shown; A storage unit that stores objects identified by a plurality of the learning images; An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized; A control unit that estimates the rotation angle of the object from the moving direction of the center point of the bounding box between frames of the moving image, and issues a warning when the angle exceeds a certain angle; An object detection device, characterized by comprising the above.
4. When two objects intersect and the object behind in the moving image is partially blocked, the position information of the object behind cannot be obtained. Therefore, record the position information of the object behind before occlusion, synthesize a new bounding box based on the last moving image data before occlusion, and display it on the display. When the object behind is displayed on the display again, recalibrate the position of the object behind using the new data. The object detection device according to Claims 1 to 3.
5. When two objects intersect and the object behind the moving image is completely blocked, if the object behind moves and is redisplayed on the display, record the last position shown on the display so that the identification number attached to the object behind does not change from before, form a new bounding box, display it on the display, continuously track the blocked object behind, and when the object behind is displayed on the display again, the system updates the data using the new video. The object detection device according to claims 1 to 3.
6. A computer, An input unit that inputs, using a learning image in which an object is shown, an image including a region in which the object is shown from an arbitrary moving image, A storage unit that stores objects identified by a plurality of the learning images, An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized, A control unit that estimates the distance between a plurality of objects by a line segment between bounding boxes and issues a warning when the plurality of objects approach a certain distance from each other, A program that functions as.
7. A computer, An input unit that inputs, using a learning image in which an object is shown, an image including a region in which the object is shown from an arbitrary moving image, A storage unit that stores objects identified by a plurality of the learning images, An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized, A control unit that estimates the object speed based on the moving distance and time of the line segment of the bounding box between frames of the moving image and issues a warning when a certain speed is exceeded, A program that functions as.
8. A computer, An input unit that inputs, using a learning image in which an object is shown, an image including a region in which the object is shown from an arbitrary moving image, A storage unit that stores objects identified by a plurality of the learning images, An output unit that outputs an image in which a bounding box corresponding to the region in which the object is shown is synthesized, A control unit that estimates the rotation angle of the object based on the moving direction of the center point of the bounding box between frames of the moving image and issues a warning when a certain angle is exceeded, A program that functions as.
9. An object detection method using a computer, Inputting, using a learning image in which an object is shown, an image including a region in which the object is shown from an arbitrary moving image, Store the objects identified by the plurality of the learning images. Output an image in which a bounding box corresponding to the region in which the object is depicted is synthesized. Estimate the distances between a plurality of objects by line segments between the bounding boxes, and when the plurality of objects approach a certain distance, issue a warning and perform control. Object detection method.
10. An object detection method using a computer, comprising: Input, using a learning image in which an object is depicted, an image including the region in which the object is depicted from an arbitrary moving image. Store the objects identified by the plurality of the learning images. Output an image in which a bounding box corresponding to the region in which the object is depicted is synthesized. Estimate the speed of the object based on the moving distance and time of the line segments of the bounding boxes between the frames of the moving image, and when the speed exceeds a certain speed, issue a warning and perform control. Object detection method.
11. An object detection method using a computer, comprising: Input, using a learning image in which an object is depicted, an image including the region in which the object is depicted from an arbitrary moving image. Store the objects identified by the plurality of the learning images. Output an image in which a bounding box corresponding to the region in which the object is depicted is synthesized. Estimate the rotation angle of the object based on the moving direction of the center points of the bounding boxes between the frames of the moving image, and when the angle exceeds a certain angle, issue a warning and perform control. Object detection method.
Citation Information
Patent Citations
Object detection device
JP2021056717A