Target detection methods, devices, electronic equipment and storage media
By predicting the motion position of the target object in a multi-view camera and using the motion velocity vector and time information to mark the target position box in the next shooting frame, the problem of jump and lag in target detection in multi-view cameras is solved, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202311083240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-08-25
AI Technical Summary
Multi-view cameras suffer from target location bounding box jumps and lags during target detection, affecting detection performance.
By determining the target object's motion velocity vector information and time consumption information, the target object's motion position in the next captured frame is predicted, and the target position box is marked in that frame, thus solving the problem of target position box jump and lag.
It improves the accuracy of target detection, reduces jumps and lags in target location boxes, and enhances detection performance.
Smart Images

Figure CN119520966B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a target detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of the video security industry, camera models have become increasingly diversified, sophisticated, and customized. To meet the market demand for a single device to cover a wide area, multi-camera products, such as binocular and quad-camera systems, have emerged. Multi-camera systems offer the advantage of using fewer devices in the same monitoring scenario, satisfying aesthetic requirements, ease of operation, and cost savings. However, the fusion of images from multiple imaging sensors in multi-camera products can cause frequent left-right jumps and lag in the target detection process, affecting detection accuracy. Summary of the Invention
[0003] This invention provides a target detection method, apparatus, electronic device, and storage medium to solve the problems of target position bounding box jumps and lags in target detection under multi-view camera shooting scenarios.
[0004] According to one aspect of the present invention, a target detection method is provided, the method comprising:
[0005] Determine the target motion velocity vector information of the target object in the first captured frame;
[0006] Determine the target time information from the start of the first captured image to the formation of the second captured image. The first captured image is formed by stitching together multiple images from a multi-view camera, and the second captured image is formed by stitching together multiple images from a multi-view camera after the first captured image.
[0007] Based on the target motion velocity vector information and the target time information, the predicted motion position of the target object in the upcoming second shooting frame is determined;
[0008] When forming the second shooting frame, a target position box for tracking the target object is determined in the formed second shooting frame based on the predicted motion position.
[0009] According to another aspect of the present invention, a target detection apparatus is provided, the apparatus comprising:
[0010] The first determining module is used to determine the target motion velocity vector information of the target object in the first captured image;
[0011] The second determining module is used to determine the target time information from the start of the first captured image to the formation of the second captured image. The first captured image is formed by stitching together multiple imaging images from a multi-view camera, and the second captured image is formed by stitching together multiple imaging images from a multi-view camera after the first captured image.
[0012] The third determining module is used to determine the predicted motion position of the target object in the upcoming second shooting frame based on the target motion velocity vector information and the target time information;
[0013] The fourth determining module, when forming the second shooting frame, determines the target position box for tracking the target object in the formed second shooting frame based on the predicted motion position.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the target detection method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the target detection method according to any embodiment of the present invention.
[0019] The technical solution of this invention determines the target motion velocity vector information of the target object in the first shooting frame and the target time information from the start of the first shooting frame to the formation of the second shooting frame. Then, during the period from the start of the first shooting frame to the formation of the second shooting frame, the motion position of the target object detected in the first shooting frame can be corrected according to the target motion velocity vector information and the target time information, so as to obtain as accurately as possible the predicted motion position that the target object may reach when the second shooting frame is formed. Then, when the second shooting frame is formed, the target position box of the tracked target object can be marked in the second shooting frame, which solves the problem of target position box jump and lag in target detection in multi-view camera shooting scenarios and improves the detection accuracy of target objects in the shooting frame.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a target detection method provided according to an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of the position frame markings of a captured image generated by stitching and fusing multiple imaging frames from a multi-view camera in the prior art applicable to the present invention.
[0024] Figure 3a This is a schematic diagram illustrating the marking of the motion position of a target object in the captured image of a multi-camera, according to the prior art applicable to embodiments of the present invention.
[0025] Figure 3b This is a schematic diagram illustrating another method of marking the movement position of a target object in the captured image of a multi-camera, according to the prior art applicable to embodiments of the present invention.
[0026] Figure 3c This is a schematic diagram illustrating another method of marking the movement position of a target object in the captured image of a multi-camera, according to the prior art applicable to embodiments of the present invention.
[0027] Figure 3d This is a schematic diagram illustrating another method of marking the movement position of a target object in the captured image of a multi-camera, according to the prior art applicable to embodiments of the present invention.
[0028] Figure 4a This is a schematic diagram illustrating the principle of marking the position of a target object in the image captured by a multi-view camera according to an embodiment of the present invention;
[0029] Figure 4b This is a schematic diagram illustrating another principle for marking the position of a target object in the image captured by a multi-camera, applicable according to an embodiment of the present invention.
[0030] Figure 5 This is a flowchart of another target detection method provided according to an embodiment of the present invention;
[0031] Figure 6 This is a schematic diagram of the structure of a target detection device according to an embodiment of the present invention;
[0032] Figure 7 This is a schematic diagram of the structure of an electronic device that implements the target detection method of this invention. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] Figure 1 The present invention provides a flowchart of a target detection method. This embodiment is applicable to the tracking and detection of target objects in the footage captured by a multi-view camera. The method can be executed by a target detection device, which can be implemented in hardware and / or software. The target detection device can be configured in any electronic device with network communication function, such as a multi-view camera.
[0036] like Figure 1 As shown, the target detection method in this embodiment may include the following process:
[0037] S110. Determine the target motion velocity vector information of the target object in the first captured image. The first captured image is formed by stitching together multiple imaging images from a multi-view camera.
[0038] A multi-lens camera is equipped with multiple lenses that can be arranged sequentially along the circumference of the camera. Each lens can be independently adjusted on three axes to achieve wide-angle or specific angle image capture. Multi-lens cameras can be multi-lens network cameras, etc.
[0039] For multi-view cameras, the images from each imaging sensor can be received via video input devices VIF Dev0 and VIF Dev2. An image processing engine then stitches and merges these images into a single, complete image, which becomes the captured footage from the multi-view camera. The image processing engine supports initial image quality adjustments to an input image, including but not limited to noise reduction, sharpening, and brightness adjustment, before scaling each image to a specific resolution and outputting it through the respective output interfaces.
[0040] The target object can refer to any object in motion, including a moving car, a pedestrian, or other people or objects in motion. The first captured image can be formed by stitching together the corresponding images from multiple imaging sensors of a multi-view camera, and the first captured image includes the target object. The target motion velocity vector information can be the speed and direction of motion of the target object in the first captured image.
[0041] S120. Determine the target time information from the start of the first shooting frame to the formation of the second shooting frame. The second shooting frame is formed by stitching together multiple imaging frames from the multi-view camera after the first shooting frame.
[0042] Multi-view cameras continuously capture images in real time. The second image is formed after the first image is captured by the multi-view cameras by stitching together the corresponding images from the multiple imaging sensors. The process from capturing the first image to capturing the second image involves detecting the movement of the target object within the multi-view camera's images and stitching together the corresponding images from the multiple imaging sensors. These operations take time. Therefore, the time taken from capturing the first image to capturing the second image is calculated.
[0043] S130. Based on the target motion velocity vector information and the target time information, determine the predicted motion position of the target object in the upcoming second shooting frame.
[0044] See Figure 2Regarding the issue of target location bounding box jumps and lags, traditional solutions involve stitching together the images from the two imaging sensors of a multi-view camera to form the captured image. Motion position detection is typically performed sequentially from left to right to obtain the target location bounding box for the target object in the captured image. However, detection modes that prioritize left over right or right over left can cause left-right jumps in the target location bounding box due to differences in the detection order. This is detrimental to the observation of the live effect, reduces the detection rate, and takes a long time to perform sequential left-right detection. Furthermore, the detected moving object bounding box may differ from the object's current position after detection, and this difference increases with the speed of the target object's movement. It may be impossible to mark the target location bounding box on the captured image in time, resulting in target location bounding box lag. (See [link to relevant documentation]). Figures 3a-3d This illustrates the target location bounding box marking in a traditional scheme involving multiple consecutive frames of footage.
[0045] Furthermore, the method of stitching and fusing the images from two imaging sensors of a multi-view camera to form a unified image for detecting the motion position of targets within that image places high demands on CPU performance, potentially leading to high computational costs. Moreover, as the number of cameras increases and the resolution rises, the unified detection rate after stitching and fusing multiple frames decreases, resulting in more severe target position frame lag issues and even excessive CPU load on the camera.
[0046] Based on the above, for motion position detection of target objects in the first captured frame of a multi-camera system, the long detection time may prevent timely position detection, thus hindering the marking of the target position bounding box in the first captured frame. Furthermore, considering the practical application scenarios of multi-camera systems, it's easy to see that most moving target objects do not experience sudden changes in speed and direction within a short period. Therefore, after determining the target motion velocity vector information and target time information from the first captured frame, this solution allows for the prediction of the target object's position in the second captured frame before the second captured frame is formed. This enables timely motion position detection of the target object in the second captured frame.
[0047] As an optional but not limited implementation, the target time information is determined based on the total detection time of motion position detection of the target object in the multi-camera's captured image and the fusion time required to stitch and merge multiple images from the multi-camera to form the second captured image after the first captured image. The total detection time is the sum of the detection time of sequentially performing motion position detection on each sub-image region in the first captured image.
[0048] See Figure 4aThe time taken from the start of the first shot to the formation of the second shot can be determined as the target time. In this way, the motion position predicted by using the target time and the target motion velocity vector is kept as close as possible to the motion position obtained by real-time detection of the target object's motion position during the second shot.
[0049] See Figure 4a The calculation of the target's motion detection time can be combined with the time required to detect the motion position of the target object in the multi-camera's captured image and to stitch together the corresponding images from multiple imaging sensors. The total detection time for the target object's motion position can be the sum of the detection times for sequentially detecting the motion position of each sub-image region in the first captured image from left to right.
[0050] Optionally, the target time information indication from the start of the first captured image to the formation of the second captured image is greater than or equal to the total detection time for motion position detection of the target object in the multi-camera captured images and the fusion time required to stitch and merge multiple imaging images from the multi-cameras to form the second captured image after the first captured image. Understandably, the target time should not be too long to prevent inaccurate judgment of the shooting environment due to an excessively long time period.
[0051] As an optional but not limited implementation, determining the predicted motion position of the target object in the upcoming second captured image, based on the target motion velocity vector information and the target time information, may include the following steps A1-A2:
[0052] Step A1: Determine the target movement position of the target object in the first captured frame.
[0053] Step A2: Based on the target motion velocity vector information and the target time information, determine the first predicted motion distance of the target object in the multi-view camera's shooting frame during the reference motion period. The reference motion period is the time period from the start of the first shooting frame to the formation of the second shooting frame.
[0054] Step A3: Based on the target's motion position and the first predicted motion distance, determine the predicted motion position of the target object in the upcoming second shooting frame.
[0055] The target motion position can be the location of the target object in the first captured frame. For example, the target motion position of the target object can be marked in the first captured frame using a target position bounding box. The motion position of the target object in the first captured frame can be detected using methods including but not limited to: clustering theory-based methods, fuzzy theory-based methods, statistical theory-based methods, background modeling-based methods, neural network-based methods, optical flow methods, frame difference methods, etc.
[0056] See Figure 4a and Figure 4b The distance the target object moves within the multi-camera view during the time interval from the start of the first shot to the formation of the second shot can be calculated using the target's velocity vector information and the target's time consumption information. Once the target object's distance within the multi-camera view is obtained, it can be superimposed on the target object's position in the first shot. This allows for prediction of the target object's position in the second shot, using the position in the first shot preceding the second shot.
[0057] Optionally, the first and second captured frames can be two adjacent frames formed by stitching together multiple images from a multi-camera system, or two frames spaced a preset number of frames apart. It is necessary to ensure that the time taken from the first captured frame to the formation of the second captured frame is greater than or equal to the total detection time for motion position detection of the target object in the captured frames of the multi-camera system and the fusion time required to stitch together multiple images from the multi-camera system after the first captured frame to form the second captured frame.
[0058] In some implementations, the positions of the individual cameras that make up a multi-camera system are relatively fixed (such as an integrated multi-camera / binocular camera), and the stitching algorithm is based on cropping, correction, and stitching based on fixed coordinates. Therefore, the fusion time for each frame is basically the same. In other implementations, the individual cameras that make up a multi-camera system are relatively independent. The stitching algorithm is based on the automatic identification and matching of feature points before cropping, correction, and stitching. Therefore, the fusion time is affected by the difficulty of feature point matching. In the above embodiments, when predicting the motion position of the target object in the second shooting frame, the target time can be adaptively adjusted in conjunction with the actual scene. For example, when the scene has few or too cluttered details (such as a water scene), the target time can be extended by 20% compared to other scenes.
[0059] By using the above method, since the target object's movement position in the upcoming second shot is predicted by using the target object's movement position in the first shot before the second shot is formed, the detection efficiency can be improved as the number of "eyes" of the multi-eye camera increases, which will not cause the moving object to lag or jump phenomenon to become more obvious as the "eyes" increase.
[0060] As an optional but not limited implementation, determining the first predicted motion distance of the target object in the multi-view camera's captured image during the reference movement period, based on the target motion velocity vector information and the target time information, may include the following steps B1-B2:
[0061] Step B1: Based on the target motion velocity vector information and the target time information, determine the second predicted motion distance of the target object in the real scene area corresponding to the multi-view camera's captured image during the reference movement period.
[0062] Step B2: Based on the second predicted motion distance and the preset motion distance mapping information, determine the first predicted motion distance of the target object in the multi-camera's shooting image during the reference movement. The preset motion distance mapping information is used to describe the distances mapped from the multi-camera's shooting image to the real scene area corresponding to the multi-camera's shooting image, which are pre-calibrated based on the multi-camera's focal length.
[0063] See Figure 4b The target motion velocity vector information indicates the target object's position in the first captured image and its motion direction relative to a preset reference object in the multi-camera image, respectively. The motion velocity and direction indicated in the target motion velocity vector information describe the target object's speed and direction of motion within the real-world scene area corresponding to the multi-camera image. Furthermore, the target motion velocity vector information and target time information can be used to calculate the second predicted motion distance of the target object within the real-world scene area corresponding to the multi-camera image during its reference movement.
[0064] Since the real scene area corresponding to the multi-camera's captured image is not proportional to the multi-camera's captured image, it is necessary to determine the mapping relationship between different distances in the real scene area corresponding to the multi-camera's captured image and the distances in the multi-camera's captured image. Then, using the mapping relationship, the second predicted motion distance is mapped from the real scene area corresponding to the multi-camera's captured image to the multi-camera's captured image, thus obtaining the first predicted motion distance of the target object in the multi-camera's captured image during the reference movement.
[0065] For example, assuming the target motion velocity vector information indicates the target object's motion velocity in the direction of motion as v, and the target time information indicates the time from the start of the first shooting frame to the formation of the second shooting frame as t seconds, then the target object has moved s(m) = v*t during the period from the start of the first shooting frame to the formation of the second shooting frame. Then, based on the focal length of the multi-view camera, the preset motion distance mapping information is queried. Using the ratio of the distance of the shooting frame corresponding to the real scene area indicated by the preset motion distance mapping information to the distance of the shooting frame, this value is set as k. It can be deduced that the first predicted motion distance of the target object in the multi-view camera's shooting frame during the reference movement period is S = k*s. Then, the predicted motion position of the target object in the upcoming second shooting frame can be obtained according to the first predicted motion distance.
[0066] S140. When forming the second shooting frame, determine the target position box for tracking the target object in the formed second shooting frame based on the predicted motion position.
[0067] When determining the target motion position of the target object in the first captured image, the size of the target object can also be determined. After predicting the predicted motion position of the target object in the second captured image, a target position box matching the size of the target object can be marked at the corresponding predicted position based on the predicted motion position and the size of the target object. In this way, when displaying the second captured image, the target object in the second captured image can be tracked using the target position box. The target position box can be a rectangle.
[0068] The technical solution of this invention determines the target motion velocity vector information of the target object in the first shooting frame and the target time information from the start of the first shooting frame to the formation of the second shooting frame. Then, during the period from the start of the first shooting frame to the formation of the second shooting frame, the motion position of the target object detected in the first shooting frame can be corrected according to the target motion velocity vector information and the target time information, so as to obtain as accurately as possible the predicted motion position that the target object may reach when the second shooting frame is formed. Then, when the second shooting frame is formed, the target position box of the tracked target object can be marked in the second shooting frame, which solves the problem of target position box jump and lag in target detection in multi-view camera shooting scenarios and improves the detection accuracy of target objects in the shooting frame.
[0069] Figure 5 The present invention provides a flowchart of another target detection method. The technical solution of this embodiment further optimizes the process of determining the target motion velocity vector information of the target object in the first shooting frame in the above embodiment based on the above embodiment. This embodiment can be combined with various optional solutions in one or more of the above embodiments.
[0070] like Figure 5 As shown, the target detection method in this embodiment may include the following process:
[0071] S510. Determine the target motion state information of the target object in the first shooting frame. The target motion state information is either a first motion state or a second motion state. When the target object is in the first motion state, the motion distance within the reference time period is not less than the reference motion distance. When the target object is in the second motion state, the motion distance within the reference time period is less than the reference motion distance.
[0072] For multi-view cameras, the motion state of a target object entering the camera's frame is not unique. The motion state of the target object can be described by its speed. For example, some target objects move very fast, and these objects will move a large distance in a few frames. If the detection time for the target object's motion position in the frame is long, there will not be enough time to mark the target position bounding box in the frame. On the other hand, some target objects move very slowly, and these objects will not move a large distance in a few frames. It can be considered that the motion position of the target object is consistent across multiple consecutive frames.
[0073] Based on the above analysis, when using a multi-view camera, it is not always necessary to use the movement position and speed of the target object in the first captured image to predict the movement position of the target object in the upcoming second captured image. Instead, it is necessary to consider the target object's movement state information in the first captured image, judge the speed of the target object's movement in the captured image, and thus decide whether to start using the movement position and speed of the target object in the first captured image to predict the movement position of the target object in the upcoming second captured image.
[0074] As an optional but not limited implementation, determining the target motion state information of the target object in the first captured frame may include the following steps C1-C2:
[0075] Step C1: After detecting that the target object has entered the shooting screen of the multi-view camera, determine at least two third shooting frames. The at least two third shooting frames are shooting frames formed by stitching together multiple imaging frames from the multi-view camera and including a preset number of frames of the first shooting frame.
[0076] Based on the images captured by the multi-camera fusion, a third shooting image is formed by stitching together multiple images of the target object continuously passing through the multi-camera. The first 5-10 frames of the multi-camera image when the target object first enters the multi-camera are used as the third shooting image, and the target object's movement direction and speed are analyzed and recorded. The target object is then judged to be in the first or second movement state based on its movement direction and speed.
[0077] Step C2: Determine the target motion position of the target object in each frame of the third shooting frame, and determine the target motion state information of the target object in the third shooting frame based on each target motion position.
[0078] Optionally, determining the target motion state information of the target object in the third captured frame based on each target motion position may include: if the difference in position distance between the target motion positions corresponding to two adjacent frames of the third captured frame is greater than or equal to a reference motion distance, then the target motion state information is determined to be a first motion state; if the difference in position distance between the target motion positions corresponding to two adjacent frames of the third captured frame is less than the reference motion distance, then the target motion state information is determined to be a second motion state, wherein the reference duration is the interval duration between two adjacent frames of the third captured frame.
[0079] S520. If the target motion state information is the first motion state, then determine the target motion speed vector information of the target object in the first shooting frame. The motion speed vector information includes the motion speed and the motion direction.
[0080] Based on the above embodiments, optionally, after determining the target motion state information of the target object in the first captured image, the following steps D1-D2 may also be included:
[0081] Step D1: If the target motion state information is the second motion state, then determine the target motion position of the target object in the first captured image.
[0082] Step D2: When forming the second shooting frame, determine the target position box for tracking the target object in the formed second shooting frame based on the target's motion position.
[0083] For example, since the movement of the target object in the multi-camera's captured images may be uncertain, the target object is collected in each frame of the third captured image during the movement process. The influence of environmental changes in each frame of the third captured image is eliminated, and each pair of adjacent frames of the third captured image is compared 5 times per second. When the target object moves less than the reference movement distance within a reference time for 5 consecutive times, the target object is considered to be in a "slow movement state". At this time, the target motion velocity vector information and target time information are not used to predict the target object's position in the second captured image, but instead the target position box is marked in real time. At this time, the multi-camera fusion time is negligible in the "slow movement state". When the target object moves more than or equal to the reference movement distance within a reference time for 5 consecutive times, the target object is considered to be in a "fast movement state". At this time, the target motion velocity vector information and target time information are used to predict the target object's position in the second captured image.
[0084] S530. Determine the target time information from the start of the first shooting frame to the formation of the second shooting frame. The first shooting frame is formed by stitching together multiple imaging frames from the multi-view camera, and the second shooting frame is formed by stitching together multiple imaging frames from the multi-view camera after the first shooting frame.
[0085] S540. Based on the target motion velocity vector information and the target time information, determine the predicted motion position of the target object in the upcoming second shooting frame.
[0086] S550: When forming the second shooting frame, a target position box for tracking the target object is determined in the formed second shooting frame based on the predicted motion position.
[0087] The technical solution of this invention determines the target motion velocity vector information of the target object in the first shooting frame and the target time information from the start of the first shooting frame to the formation of the second shooting frame. Then, during the period from the start of the first shooting frame to the formation of the second shooting frame, the motion position of the target object detected in the first shooting frame can be corrected according to the target motion velocity vector information and the target time information, so as to obtain as accurately as possible the predicted motion position that the target object may reach when the second shooting frame is formed. Then, when the second shooting frame is formed, the target position box of the tracked target object can be marked in the second shooting frame, which solves the problem of target position box jump and lag in target detection in multi-view camera shooting scenarios and improves the detection accuracy of target objects in the shooting frame.
[0088] Figure 6This invention provides a structural block diagram of a target detection device. This embodiment is applicable to the tracking and detection of target objects in the footage captured by a multi-view camera. The target detection device can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities, such as a multi-view camera.
[0089] like Figure 6 As shown, the target detection device in this embodiment may include the following process:
[0090] The first determining module 610 is used to determine the target motion velocity vector information of the target object in the first captured image;
[0091] The second determining module 620 is used to determine the target time information from the start of the first shooting frame to the formation of the second shooting frame. The first shooting frame is formed by stitching together multiple imaging frames from a multi-view camera, and the second shooting frame is formed by stitching together multiple imaging frames from a multi-view camera after the first shooting frame.
[0092] The third determining module 630 is used to determine the predicted motion position of the target object in the upcoming second shooting frame based on the target motion velocity vector information and the target time information;
[0093] The fourth determining module 640, when forming the second shooting frame, determines the target position box for tracking the target object in the formed second shooting frame based on the predicted motion position.
[0094] Based on the above embodiments, optionally, the target time information is determined based on the total detection time for detecting the motion position of the target object in the captured image and the fusion time required to stitch and fuse multiple imaging images from a multi-view camera to form a second captured image. The total detection time is the sum of the detection time for sequentially detecting the motion position of each sub-image region in the first captured image.
[0095] Based on the above embodiments, optionally, determining the target motion velocity vector information of the target object in the first captured image includes:
[0096] Determine the target motion state information of the target object in the first shooting frame. The target motion state information is either a first motion state or a second motion state. When the target object is in the first motion state, the motion distance within the reference time is not less than the reference motion distance. When the target object is in the second motion state, the motion distance within the reference time is less than the reference motion distance.
[0097] If the target motion state information is the first motion state, then the target motion velocity vector information of the target object in the first captured image is determined, and the motion velocity vector information includes the motion velocity and the motion direction.
[0098] Based on the above embodiments, optionally, determining the target motion state information of the target object in the first captured image includes:
[0099] After detecting that the target object enters the shooting screen of the multi-view camera, at least two third shooting frames are determined. The at least two third shooting frames are shooting frames formed by stitching together multiple imaging frames from the multi-view camera and including a preset number of frames of the first shooting frame.
[0100] Determine the target object's motion position in each frame of the third-shot image, and determine the target object's motion state information in the third-shot image based on each target motion position.
[0101] Based on the above embodiments, optionally, the target motion state information of the target object in the third captured frame is determined according to the motion positions of each target, including:
[0102] If the difference in positional distance between the target motion positions corresponding to two adjacent third-shot frames is greater than or equal to the reference motion distance, then the target motion state information is determined to be the first motion state.
[0103] If the difference in positional distance between the target motion positions corresponding to two adjacent third-shot frames is less than the reference motion distance, then the target motion state information is determined to be the second motion state, and the reference duration is the interval duration between two adjacent third-shot frames.
[0104] Optionally, based on the above embodiments, after determining the target motion state information of the target object in the first captured image, the method further includes:
[0105] If the target motion state information is the second motion state, then the target motion position of the target object in the first captured image is determined;
[0106] When forming the second shooting frame, a target position box for tracking the target object is determined in the formed second shooting frame based on the target's motion position.
[0107] Based on the above embodiments, optionally, determining the predicted motion position of the target object in the upcoming second captured image based on the target motion velocity vector information and the target time consumption information includes:
[0108] Determine the target's movement position in the first captured frame;
[0109] Based on the target motion velocity vector information and the target time information, a first predicted motion distance of the target object in the multi-view camera's capture image during the reference motion period is determined, wherein the reference motion period is the time period from the start of the first capture image to the formation of the second capture image;
[0110] Based on the target's movement position and the first predicted movement distance, the predicted movement position of the target object in the upcoming second shooting frame is determined.
[0111] Based on the above embodiments, optionally, determining the first predicted motion distance of the target object in the multi-view camera's captured image during the reference movement period, based on the target motion velocity vector information and the target time information, includes:
[0112] Based on the target motion velocity vector information and the target time information, a second predicted motion distance is determined in the real scene area corresponding to the multi-view camera's captured image during the target object's reference movement.
[0113] Based on the second predicted motion distance and the preset motion distance mapping information, the first predicted motion distance of the target object in the multi-camera's shooting frame during the reference movement is determined. The preset motion distance mapping information is used to describe the distances of different distances of the multi-camera's shooting frame corresponding to the real scene area, which are pre-calibrated based on the focal length of the multi-camera, mapped to the multi-camera's shooting frame.
[0114] The technical solution of this invention determines the target motion velocity vector information of the target object in the first shooting frame and the target time information from the start of the first shooting frame to the formation of the second shooting frame. Then, during the period from the start of the first shooting frame to the formation of the second shooting frame, the motion position of the target object detected in the first shooting frame can be corrected according to the target motion velocity vector information and the target time information, so as to obtain as accurately as possible the predicted motion position that the target object may reach when the second shooting frame is formed. Then, when the second shooting frame is formed, the target position box of the tracked target object can be marked in the second shooting frame, which solves the problem of target position box jump and lag in target detection in multi-view camera shooting scenarios and improves the detection accuracy of target objects in the shooting frame.
[0115] The target detection device provided in the embodiments of the present invention can execute the target detection method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the target detection method. For details, please refer to the relevant operations of the target detection method in the foregoing embodiments.
[0116] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0117] Figure 7 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0118] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0119] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0120] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as object detection methods.
[0121] In some embodiments, the target detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the target detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the target detection method by any other suitable means (e.g., by means of firmware).
[0122] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0126] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0127] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0128] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0129] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A target detection method, characterized in that, The method includes: Determine the target motion velocity vector information of the target object in the first captured frame; Determine the target time information from the start of the first captured image to the formation of the second captured image. The first captured image is formed by stitching together multiple images from a multi-view camera, and the second captured image is formed by stitching together multiple images from a multi-view camera after the first captured image. Based on the target motion velocity vector information and the target time information, the predicted motion position of the target object in the upcoming second shooting frame is determined; When forming the second shooting frame, a target position box for tracking the target object is determined in the formed second shooting frame based on the predicted motion position.
2. The method according to claim 1, characterized in that, The target time information is determined based on the total detection time for detecting the motion position of the target object in the captured image and the fusion time required to stitch together and fuse multiple images from a multi-view camera to form the second captured image. The total detection time is the sum of the detection time for sequentially detecting the motion position of each sub-image region in the first captured image.
3. The method according to claim 1, characterized in that, Determine the target motion velocity vector information of the target object in the first captured frame, including: Determine the target motion state information of the target object in the first shooting frame. The target motion state information is either a first motion state or a second motion state. When the target object is in the first motion state, the motion distance within the reference time is not less than the reference motion distance. When the target object is in the second motion state, the motion distance within the reference time is less than the reference motion distance. If the target motion state information is the first motion state, then the target motion velocity vector information of the target object in the first captured image is determined, and the motion velocity vector information includes the motion velocity and the motion direction.
4. The method according to claim 3, characterized in that, Determine the target motion state information of the target object in the first captured frame, including: After detecting that the target object enters the shooting screen of the multi-view camera, at least two third shooting frames are determined. The at least two third shooting frames are shooting frames formed by stitching together multiple imaging frames from the multi-view camera and including a preset number of frames of the first shooting frame. Determine the target object's motion position in each frame of the third-shot image, and determine the target object's motion state information in the third-shot image based on each target motion position.
5. The method according to claim 4, characterized in that, Based on the movement positions of each target, the target motion state information in the third-shot frame is determined, including: If the difference in positional distance between the target motion positions corresponding to two adjacent third-shot frames is greater than or equal to the reference motion distance, then the target motion state information is determined to be the first motion state. If the difference in positional distance between the target motion positions corresponding to two adjacent third-shot frames is less than the reference motion distance, then the target motion state information is determined to be the second motion state, and the reference duration is the interval duration between two adjacent third-shot frames.
6. The method according to claim 3, characterized in that, After determining the target object's motion state information in the first captured frame, the process also includes: If the target motion state information is the second motion state, then the target motion position of the target object in the first captured image is determined; When forming the second shooting frame, a target position box for tracking the target object is determined in the formed second shooting frame based on the target's motion position.
7. The method according to claim 1, characterized in that, Based on the target motion velocity vector information and the target time information, the predicted motion position of the target object in the upcoming second captured image is determined, including: Determine the target's movement position in the first captured frame; Based on the target motion velocity vector information and the target time information, a first predicted motion distance of the target object in the multi-view camera's capture image during the reference motion period is determined, wherein the reference motion period is the time period from the start of the first capture image to the formation of the second capture image; Based on the target's movement position and the first predicted movement distance, the predicted movement position of the target object in the upcoming second shooting frame is determined.
8. The method according to claim 7, characterized in that, Based on the target motion velocity vector information and the target time information, determining the first predicted motion distance of the target object in the multi-view camera's captured image during the reference movement includes: Based on the target motion velocity vector information and the target time information, a second predicted motion distance is determined in the real scene area corresponding to the multi-view camera's captured image during the target object's reference movement. Based on the second predicted motion distance and the preset motion distance mapping information, the first predicted motion distance of the target object in the multi-camera's shooting frame during the reference movement is determined. The preset motion distance mapping information is used to describe the distances of different distances of the multi-camera's shooting frame corresponding to the real scene area, which are pre-calibrated based on the focal length of the multi-camera, mapped to the multi-camera's shooting frame.
9. A target detection device, characterized in that, The device includes: The first determining module is used to determine the target motion velocity vector information of the target object in the first captured image; The second determining module is used to determine the target time information from the start of the first captured image to the formation of the second captured image. The first captured image is formed by stitching together multiple imaging images from a multi-view camera, and the second captured image is formed by stitching together multiple imaging images from a multi-view camera after the first captured image. The third determining module is used to determine the predicted motion position of the target object in the upcoming second shooting frame based on the target motion velocity vector information and the target time information; The fourth determining module, when forming the second shooting frame, determines the target position box for tracking the target object in the formed second shooting frame based on the predicted motion position.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the target detection method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the target detection method according to any one of claims 1-8.
Citation Information
Patent Citations
Athlete tracking method and system
CN112070795A
Image processing method and device
CN116134484A