Moving object tracking method and device, computer-readable medium and electronic device
By using virtual camera technology to obtain and crop video frames in monitoring, online meetings and large-scale competition director scenarios, the problems of high costs and poor results in the existing technology are solved, and low-cost and high-quality tracking and shooting of moving objects are achieved.
Patent Information
- Application Number
- CN202210841788.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-07-18
AI Technical Summary
In the prior art, in monitoring, online meetings and large-scale competition director scenarios, tracking and shooting of sports objects requires a large labor and hardware cost, and the effect is affected by the photographer's level and gimbal/motor quality, resulting in poor tracking results.
By obtaining the original video frame, determining the moving object, and creating a virtual camera based on the property parameters of the real camera, performing cropping processing, generating target video frames for the screen to move with the moving object, avoiding manual operations and gimbal/motor driving.
It reduces the manpower and hardware cost of tracking and shooting of moving objects, improves the stability and smoothness of target video frames, and achieves better tracking and shooting effects independently of the photographer level and gimbal/motor stability.
Smart Images

Figure CN115240107B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a moving object tracking method, a moving object tracking device, a computer-readable medium, and an electronic device. Background Art
[0002] In scenarios such as surveillance, online conferencing, or directing large-scale competitions, the scene is large and the camera's field of view (FOV) is limited. Therefore, it is necessary to rotate the camera's shooting angle to track moving objects in the scene.
[0003] Currently, related technical solutions either require manually rotating the camera to achieve tracking shooting of a moving object, or require setting a pan / tilt head or a motor on the camera and driving the pan / tilt head or the motor to rotate the camera to achieve tracking shooting of a moving object. However, these solutions require high manpower and hardware costs, and the shooting process is affected by the photographer's photography skills and the quality of the pan / tilt head or the motor, resulting in poor tracking shooting effects. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a moving object tracking method, a moving object tracking device, a computer-readable medium and an electronic device, thereby reducing, at least to a certain extent, the manpower and hardware costs required for tracking and photographing moving objects in a scene in the related art.
[0005] According to a first aspect of the present disclosure, a moving object tracking method is provided, comprising:
[0006] Obtaining an original video frame, wherein the original video frame is captured by a real camera with a fixed shooting angle;
[0007] determining a moving object in the original video frame;
[0008] Creating a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame;
[0009] The original video frame is cropped based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object.
[0010] According to a second aspect of the present disclosure, there is provided a moving object tracking apparatus, comprising:
[0011] A video frame acquisition module is used to acquire original video frames, where the original video frames are captured by a real camera with a fixed shooting angle;
[0012] A moving object determination module, configured to determine a moving object in the original video frame;
[0013] A virtual camera creation module, configured to create a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame;
[0014] The video frame cropping module is used to crop the original video frame based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object.
[0015] According to a third aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0016] According to a fourth aspect of the present disclosure, there is provided an electronic device, comprising:
[0017] processor; and
[0018] The memory is used to store one or more programs. When the one or more programs are executed by one or more processors, the one or more processors implement the above method.
[0019] An embodiment of the present disclosure provides a method for tracking a moving object. The method can obtain a raw video frame captured by a real camera with a fixed shooting angle, then determine the moving object in the raw video frame. A virtual camera can then be created based on the real camera's attribute parameters and the moving object's position in the raw video frame. The raw video frame can then be cropped based on the virtual camera to obtain a target video frame whose image follows the moving object. On the one hand, a real camera with a fixed shooting angle captures a raw video frame with a large field of view. Then, a virtual camera is created to crop the raw video frame to obtain a target video frame with a smaller field of view, but the view angle tracks the movement of the original moving object. This method does not require manual operation or rotation of the real camera using a pan / tilt or motor, effectively reducing the labor and hardware costs associated with tracking and capturing moving objects. On the other hand, the virtual camera is created based on the real camera's attribute parameters and the moving object's position in the raw video frame. It is not affected by the manual camera operator's skill or the stability of the pan / tilt or motor, effectively improving the capture quality of the target video frame and ensuring its stability and smoothness.
[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0022] Figure 1 A schematic diagram showing an exemplary system architecture to which embodiments of the present disclosure may be applied;
[0023] Figure 2 A schematic diagram schematically illustrates a flow chart of a moving object tracking method in an exemplary embodiment of the present disclosure;
[0024] Figure 3 A schematic diagram schematically illustrates a principle of moving object tracking in an exemplary embodiment of the present disclosure;
[0025] Figure 4 A schematic diagram of a process for determining a rotation matrix of a virtual camera in an exemplary embodiment of the present disclosure is schematically shown;
[0026] Figure 5 A schematic diagram schematically illustrates a principle of a virtual camera rotation angle in an exemplary embodiment of the present disclosure;
[0027] Figure 6 A schematic diagram of a process for determining an intrinsic parameter matrix of a virtual camera in an exemplary embodiment of the present disclosure is schematically shown;
[0028] Figure 7 The following schematically illustrates a process flow of cropping an original video frame in an exemplary embodiment of the present disclosure;
[0029] Figure 8 A schematic diagram of a process for smoothing a virtual camera tracking and shooting process in an exemplary embodiment of the present disclosure is shown;
[0030] Figure 9 A schematic diagram schematically illustrates a principle of smoothing a virtual camera tracking and shooting process in an exemplary embodiment of the present disclosure;
[0031] Figure 10 Schematically illustrating a flowchart of tracking shooting in a protagonist mode in an exemplary embodiment of the present disclosure;
[0032] Figure 11 A schematic diagram schematically illustrates the composition of a moving object tracking device in an exemplary embodiment of the present disclosure;
[0033] Figure 12 A schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0035] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0036] Figure 1 A schematic diagram shows a system architecture of an exemplary application environment in which a moving object tracking method and apparatus according to an embodiment of the present disclosure can be applied.
[0037] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with image acquisition functions, including but not limited to desktop computers, portable computers, smart phones, and tablet computers, etc. It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as needed. For example, the server 105 may be a server cluster consisting of multiple servers.
[0038] The moving object tracking method provided in the embodiments of the present disclosure is generally executed by the terminal devices 101, 102, and 103, and accordingly, the moving object tracking apparatus is generally disposed in the terminal devices 101, 102, and 103. However, it is readily understood by those skilled in the art that the moving object tracking method provided in the embodiments of the present disclosure may also be executed by the server 105, and accordingly, the moving object tracking apparatus may also be disposed in the server 105, and this is not particularly limited in the present exemplary embodiment.
[0039] For example, in an exemplary embodiment, a user may capture original video frames through a real camera associated with terminal devices 101, 102, and 103, and then upload the original video frames to server 105. After the server generates a target video frame through the motion object tracking method provided by the embodiment of the present disclosure, the target video frame is transmitted to terminal devices 101, 102, 103, etc.
[0040] In one technical solution, a camera can be mounted on a pan-tilt head or motor in the surveillance field. The periodic rotation of the pan-tilt head or motor drives the camera to change its viewing angle and scan the environment. When a moving person is detected entering the field of view, the person's position is determined through face detection, and the pan-tilt head or motor is controlled based on this positional information to keep the person within the frame, thereby capturing the person's trajectory as they enter the surveillance area. Throughout this process, an algorithm detects and provides control signals to drive the motor or pan-tilt head, completing the video tracking process. While this solution can expand the surveillance camera's field of view and automatically follow potentially valuable objects, improving surveillance efficiency and value, the additional pan-tilt head or motor required to drive the camera's rotation increases the camera's manufacturing cost.
[0041] One technical solution, in the field of conference systems, provides a camera with a pan / tilt system and, in combination with a microphone array, automatically tracks participants in a conference setting, for example, automatically tracking the area where a speaker is speaking. The microphone array locates the area where the speaker is speaking, and controls the pan / tilt system to rotate to that area, enabling camera tracking of a single person. While this solution enables real-time switching and changing of camera perspectives, enhancing participant engagement and focus, the camera pan / tilt system and motors incur additional costs, increasing the manufacturing cost of the conference camera.
[0042] In one technical solution, in the field of event directing, for example, live broadcasts of football matches, multiple cameras are typically deployed to record footage from different areas and ranges. The footage is then transmitted back to the director, who then decides which camera to play for the audience. Each camera's live broadcast requires a camera operator operating a camera with a stable tripod. This solution's professional camera operators and directors can present a comprehensive and professional match footage to the audience. However, the need for multiple cameras and camera operators for a single match incurs significant labor and hardware costs. Furthermore, the camera's performance is affected by the camera operator, resulting in poor smoothness and stability.
[0043] Based on one or more problems in the related art, the present disclosure first provides a moving object tracking method. The moving object tracking method of the exemplary embodiment of the present disclosure is specifically described below by taking a terminal device executing the method as an example.
[0044] Figure 2 A schematic flow chart of a moving object tracking method in this exemplary embodiment is shown, including the following steps S210 to S240:
[0045] In step S210 , an original video frame is obtained, where the original video frame is captured by a real camera with a fixed shooting angle.
[0046] In an exemplary embodiment, the original video frame refers to a video frame whose picture contains the entire field of view of the scene to be monitored. For example, in a football game directing scenario, the original video frame may be a video frame whose picture contains the entire football field or the movement range of all game players; in a security monitoring scenario, the original video frame may be a video frame containing the field of view of all monitoring areas; in an online conference field scenario, the original video frame may be a video frame containing the conference scene and all conference participants. This exemplary embodiment does not specifically limit this.
[0047] A real camera refers to an image acquisition device used to capture original video frames in a real scene. For example, a real camera can be an image acquisition device such as a surveillance camera, a conference camera, a professional camera, or an electronic device with image acquisition function, such as a smart phone. Of course, a real camera can also be an input device that realizes data transmission through a communication connection (such as a wired connection, a wireless connection, or other communication connection methods) with a computing device. This example embodiment does not impose any special restrictions on the form of expression of the real camera.
[0048] The real camera can be a standard camera or an ultra-wide-angle camera, depending on the size of the scene being captured or the coverage area. For example, in a conference room, due to the small coverage area required, a camera with a standard field of view can be used. For scenes such as a football field or a large surveillance area, an ultra-wide-angle camera can be selected as needed, such as one with a diagonal field of view between 110° and 150°. Of course, the image quality of the real camera can also be customized based on the scene requirements, such as the image resolution, pixel size, image signal processing (ISP) performance, and lens structure and process. This example embodiment does not impose any specific restrictions on the parameter requirements of the real camera.
[0049] After determining the real camera according to the application scenario, the real camera can be fixed in the shooting scene, so that the real camera can capture all areas and all moving objects in the shooting scene. Then, the original video frames corresponding to the shooting scene can be collected by the real camera with a fixed shooting angle.
[0050] In step S220, a moving object in the original video frame is determined.
[0051] In an exemplary embodiment, a moving object refers to an object in an original video frame that needs to be tracked and photographed. For example, a moving object can be a person walking in a monitored area or a moving vehicle in a monitored scene, or a player in a game scene or a game ball in a ball game, etc. Of course, a moving object can also be other types of objects that need to be tracked and photographed, and this embodiment does not specifically limit this.
[0052] The moving object in the original video frame can be determined through target detection and target tracking algorithms, or through auxiliary tools (such as microphone arrays) and other external factors of the image (such as the sound emitted by the moving object). This embodiment does not impose any special restrictions on the method of determining the moving object.
[0053] In step S230 , a virtual camera is created according to the attribute parameters of the real camera and the position of the moving object in the original video frame.
[0054] In an exemplary embodiment, the attribute parameters of a real camera may include the optical center coordinates and optical axis data of the real camera, the output image resolution of the real camera, the intrinsic parameter matrix of the real camera, etc. The required attribute parameters are pre-acquired or calibrated according to the actual application scenario. This exemplary embodiment does not specifically limit the type of attribute parameters of the real camera.
[0055] The position of the moving object in the original video frame can be the position coordinates of the four corner points of the detection frame corresponding to the moving object, or the position coordinates of the geometric center point of the detection frame corresponding to the moving object. Of course, it can be the position coordinates of the pre-defined weighted center point of the moving object. This example embodiment does not make any special restrictions on this.
[0056] A virtual camera refers to a virtual image acquisition device created based on the attribute parameters of a real camera, whose shooting field of view follows the position of the moving object in the original video frame. To a certain extent, the virtual camera can be considered as a bunch of parameters. Through this virtual camera, the partial image area corresponding to the moving object in the original video frame can be output to the display device, presenting the effect of the video screen tracking the movement of the moving object. That is, the virtual camera replaces the photographer or the gimbal, motor, etc. associated with the real camera to track the moving object in the original video frame and achieve a "moving camera" effect.
[0057] In step S240 , the original video frame is cropped based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object.
[0058] In an exemplary embodiment, after creating a virtual camera, the operation mode of a real camera can be simulated according to the relevant parameters of the virtual camera, and a "shooting" operation can be performed on the basis of the original video frame to capture the partial image corresponding to the moving object in the original video frame, that is, to achieve the cropping of the original video frame, and finally obtain the target video frame in which the picture follows the movement of the moving object.
[0059] Figure 3 A schematic diagram schematically illustrates a principle of moving object tracking in an exemplary embodiment of the present disclosure.
[0060] refer to Figure 3 As shown, a target scene can be captured by a real camera 310 with a fixed shooting angle to obtain an original video frame 320, which includes a moving object 330. The moving object 330 in the original video frame 320 can be determined by using an object detection and tracking algorithm, or by using external factors (such as the sound emitted by the moving object) such as auxiliary tools (such as a microphone array) to determine the moving object 330 in the original video frame 320.
[0061] Specifically, a virtual camera 340 can be created based on the attribute parameters of the real camera 310 and the position of the moving object 330 in the original video frame 320, and the original video frame 320 can be cropped by the virtual camera 340 to obtain a target video frame 350 in which the picture follows the movement of the moving object 330.
[0062] The original video frame can be obtained based on a real camera with a fixed shooting angle, and then a virtual camera can be created according to the attribute parameters of the real camera and the position of the moving object in the original video frame. Then, the target video frame whose picture follows the movement of the moving object can be output through the virtual camera based on the original video frame. In addition, the virtual camera does not require a professional photographer, nor does it require a gimbal or motor to drive the camera to rotate, which effectively reduces the labor cost and hardware cost of tracking and shooting moving objects. In addition, the tracking and shooting of moving objects through the virtual camera is not only simple to operate, but also data-driven, and the tracking and shooting effect of the target video frame obtained is more stable and smooth, which effectively improves the quality of tracking and shooting of moving objects.
[0063] Steps S210 to S240 are described in detail below.
[0064] In an exemplary embodiment, extrinsic camera parameters and intrinsic camera parameters may be determined according to property parameters of a real camera and the position of a moving object in an original video frame, and then a virtual camera may be created based on the extrinsic camera parameters and the intrinsic camera parameters.
[0065] Among them, the external camera parameters refer to the parameters of the virtual camera in the world coordinate system. For example, the external camera parameters corresponding to the virtual camera can be the optical center position of the virtual camera in the world coordinate system, the optical axis of the virtual camera, or the rotation direction of the virtual camera in the world coordinate system. This example embodiment does not specifically limit the type of the external camera parameters of the virtual camera.
[0066] Intrinsic camera parameters refer to parameters related to the characteristics of the virtual camera itself. For example, the intrinsic camera parameters of the virtual camera can be the focal length of the virtual camera, the aspect ratio of the output image of the virtual camera, the pixel size of the virtual camera, and other parameters. This example embodiment does not specifically limit the type of the intrinsic camera parameters of the virtual camera.
[0067] The optical axis of the virtual camera can be determined based on the position of the moving object in the original video frame, and then the optical axis angle is calculated based on the optical axis in the attribute parameters of the real camera to determine the rotation direction of the virtual camera, that is, the external camera parameters of the virtual camera; the intrinsic camera parameters of the virtual camera can also be obtained based on the intrinsic parameter matrix transformation of the real camera, but this example embodiment is not limited to this.
[0068] Optionally, the external camera parameters of the virtual camera may include a rotation matrix, the optical center of the virtual camera and the optical center of the real camera may be optical centers at the same position, and the attribute parameters of the real camera may include a first optical axis.
[0069] Can be achieved through Figure 4 The steps in the above are used to determine the external camera parameters of the virtual camera. Figure 3 Specifically, it may include:
[0070] Step S410, determining a second optical axis of the virtual camera according to the position of the moving object in the original video frame and the optical center;
[0071] Step S420: Determine a rotation matrix corresponding to the virtual camera based on the angle between the first optical axis and the second optical axis.
[0072] The position of the moving object in the original video frame can be determined by determining key points. For example, target detection can be performed on the original video frame to determine the minimum bounding box, connected domain, or detection frame corresponding to the moving object in the original video frame. Then, the geometric center point corresponding to the minimum bounding box, connected domain, or detection frame is determined, and the position coordinates corresponding to the geometric center point are used as the position of the moving object in the original video frame. Alternatively, key feature points corresponding to the moving object can be specified. For example, when the moving object is a person, the feature key point can be a feature point at any position on the face of the person, and the position coordinates corresponding to the feature key point are used as the position of the moving object in the original video frame. Of course, the position of the moving object in the original video frame can also be determined by other methods, which are not limited in this example embodiment.
[0073] The optical center refers to the center of the lens in the camera, the optical axis refers to the line connecting the centers of the two spheres that make up the lens in the camera, the first optical axis refers to the optical axis corresponding to the real camera, and the second optical axis refers to the optical axis corresponding to the hypothetical virtual camera.
[0074] It should be noted that the "first" and "second" in the "first optical axis" and "second optical axis" of this embodiment are only used to distinguish the different optical axes corresponding to the real camera and the virtual camera, and do not have any special meaning and should not impose any special limitations on this example embodiment.
[0075] In this embodiment, the optical center of the real camera can be used as the optical center of the virtual camera. That is, the real camera and the virtual camera are co-optical. Thus, a line can be constructed based on the optical center of the real camera and the position of the moving object in the original video frame. This line can be considered the second optical axis of the virtual camera. After determining the second optical axis of the virtual camera, the first optical axis of the real camera can be used as the three-dimensional rotation axis of the second optical axis of the virtual camera. Furthermore, the angle between the first optical axis of the real camera and the second optical axis of the virtual camera can be determined, i.e., the rotation angle of the virtual camera, and thus the rotation matrix of the virtual camera in the world coordinate system can be determined.
[0076] For example, in the three-dimensional coordinate system xyz established by the imaging plane corresponding to the original video frame, the rotation angle of the second optical axis of the virtual camera around the first optical axis of the real camera can be calculated by the relationship group (1):
[0077]
[0078] Among them, θ x It can represent the angle component of the rotation angle on the x-axis, θ y It can represent the angle component of the rotation angle on the y-axis, θ z It can represent the angle component of the rotation angle on the z coordinate axis, and the position coordinate (u, v) can represent the position of the moving object in the original video frame.
[0079] After calculating the angle between the first and second optical axes, or the rotation angle, the virtual camera's corresponding rotation vector can be calculated based on this angle. This rotation vector can then be converted into a rotation matrix using the Rodriguez formula. Since the real and virtual cameras are assumed to be concentric, the z-axis component of the rotation angle doesn't need to be considered during the calculation, effectively reducing the computational effort.
[0080] Those skilled in the art will understand that, although the present embodiment is described as assuming that the real camera and the virtual camera have the same optical center, the present embodiment may also include one or more virtual cameras, and these virtual cameras may not have the same optical center as the real camera, and the present embodiment does not impose any special limitation on this.
[0081] Figure 5 The following is a schematic diagram schematically illustrating a principle of a virtual camera rotation angle in an exemplary embodiment of the present disclosure.
[0082] refer to Figure 5 As shown, during calculation, it can be assumed that the optical center of the virtual camera and the optical center of the real camera can be the optical center 510 at the same position. By setting the optical center of the virtual camera to overlap with the optical center of the real camera, the amount of calculation can be effectively reduced and the system performance can be improved.
[0083] Specifically, the attribute parameters of the real camera may include a first optical axis 520. The first optical axis 520 can be obtained by pre-calibrating the real camera. The second optical axis 530 of the virtual camera can be obtained by connecting the position of the moving object in the original video frame and the optical center; determine the angle between the second optical axis 530 and the first optical axis 520, that is, use the optical center 510 as the rotation center and the first optical axis 520 as the rotation axis to determine the rotation angle θ of the second optical axis 530, and then determine the rotation matrix corresponding to the virtual camera according to the rotation angle θ.
[0084] In an exemplary embodiment, the intrinsic camera parameters may include an intrinsic parameter matrix, which may be obtained by Figure 6 The steps in the above example determine the intrinsic parameter matrix of the virtual camera. Figure 6 Specifically, it may include:
[0085] Step S610, determining focal length data corresponding to the virtual camera according to attribute parameters of the real camera;
[0086] Step S620, obtaining a preset image cropping ratio;
[0087] Step S630 : determining an intrinsic parameter matrix corresponding to the virtual camera based on the focal length data and the image cropping ratio.
[0088] Among them, the image cropping ratio refers to the pre-set length and width parameters of the virtual camera output image. For example, assuming that the output image size of the real camera is 1920*1080, then the image cropping ratio of the virtual camera can be 100*100, or 100*50, etc. The specific setting can be customized according to the actual application situation, and this embodiment does not impose any special limitation on this.
[0089] The focal length data corresponding to the virtual camera is the only parameter that determines the size of the virtual camera's field of view. Since the virtual camera needs to perform cropping based on the original video frames captured by the real camera, the focal length data of the virtual camera can be restricted according to the attribute parameters of the real camera. For example, the attribute parameters of the real camera can be the length and width data of the real camera's output image.
[0090] Optionally, the focal length of the virtual camera can be determined by conditions 1 and 2: Condition 1: The boundaries of the virtual camera cannot exceed the boundaries of the real camera, otherwise the image mapping will have empty areas without valid color information to fill. That is, for any point (x, y) in the virtual camera coordinates, it is converted to the coordinates of the real camera (x', y') through the established mapping relationship, where the values of x' and y' are both in the output image of the real camera; Condition 2: The field of view of the virtual camera must be able to include every valid moving object detected in the original video frame.
[0091] For example, the intrinsic parameter matrix corresponding to the virtual camera can be determined by equation (2):
[0092]
[0093] Among them, K virtual It can represent the intrinsic parameter matrix corresponding to the virtual camera, f virtual It can represent the focal length data corresponding to the virtual camera, w virtual It can represent the length of the cropped image output by the virtual camera, h virtual It can represent the width of the cropped image output by the virtual camera.
[0094] In an exemplary embodiment, the Figure 7 The steps in the above code are used to crop the original video frame. Figure 7 Specifically, it may include:
[0095] Step S710, obtaining the intrinsic parameter matrix of the real camera obtained through calibration;
[0096] Step S720: creating a projection transformation relationship from the virtual camera to the real camera according to the intrinsic parameter matrix of the real camera, the rotation matrix corresponding to the virtual camera, and the intrinsic parameter matrix corresponding to the virtual camera;
[0097] Step S730 : determining a cropping area based on the projection transformation relationship, and performing cropping processing on the original video frame using the cropping area to obtain a target video frame whose picture follows the movement of the moving object.
[0098] Among them, the projection transformation relationship refers to the transformation relationship of converting points in the real camera pixel coordinate system to points in the virtual camera pixel coordinate system. For example, the projection transformation relationship from the virtual camera to the real camera can be represented by a homography matrix (Homograph). Of course, it can also be represented by other constraint matrices. This example embodiment does not make any special restrictions on this.
[0099] For example, the projection transformation relationship from the virtual camera to the real camera can be expressed by equation (3):
[0100]
[0101] Among them, H can represent the projection transformation relationship from the virtual camera to the real camera, that is, the homography matrix, K original It can represent the intrinsic parameter matrix of the real camera, K virtual It can represent the intrinsic parameter matrix of the virtual camera, and R can represent the rotation matrix from the virtual camera to the real camera. The rotation matrix can be obtained by the Rodriguez formula: R = cosθI + (1-cosθ)nn T +sinθn ^ , where I can represent the identity matrix, n can represent the rotation axis of the real camera, and the symbol “^” is the conversion symbol from vector to antisymmetric. virtual and x original The homogeneous coordinates of the points in the virtual camera pixel coordinate system and the real camera pixel coordinate system can be expressed respectively. Specifically, the homogeneous coordinates of the points in the virtual camera pixel coordinate system can be expressed as x virtual =[u virtual v virtual 1] T , the homogeneous coordinates of the point in the real camera pixel coordinate system can be expressed as x origial =[u origial v origial 1] T , the unit is pixel.
[0102] By determining the homography matrix, that is, the projection transformation relationship from the virtual camera to the real camera, through the intrinsic parameter matrix of the real camera, the rotation matrix corresponding to the virtual camera, and the intrinsic parameter matrix corresponding to the virtual camera, the cropping area can be accurately determined, and the original video frame can be cropped through the cropping area to obtain the target video frame that follows the movement of the moving object, thereby achieving a "moving camera" effect. In addition, the virtual camera does not require a professional photographer, nor does it require a gimbal or motor to drive the camera rotation, effectively reducing the labor and hardware costs of tracking and shooting moving objects. In addition, the tracking and shooting of moving objects through the virtual camera is not only simple to operate, but also data-driven, and the tracking and shooting effect of the target video frame obtained is more stable and smooth, effectively improving the quality of tracking and shooting moving objects.
[0103] In an exemplary embodiment, the target video frame can be smoothed during the cropping process of the original video frame based on the virtual camera. By smoothing the target video frame, the smoothness and stability of the target video frame can be effectively improved, and the tracking effect of the moving object can be improved.
[0104] Specifically, you can Figure 8 The steps in the smoothing process of the target video frame are implemented, refer to Figure 8 Specifically, it may include:
[0105] Step S810, determining an estimated position to which the optical axis of the virtual camera of the current frame in the target video frame is rotated, and determining an estimated focal length of the virtual camera in the current frame;
[0106] Step S820, obtaining a historical position to which the optical axis of the virtual camera rotates in the previous frame of the target video frame, and a historical focal length of the virtual camera;
[0107] Step S830, determining a target position of the optical axis of the virtual camera based on a preset smoothing weight coefficient, the estimated position, and the historical position;
[0108] Step S840, determining a target focal length of the virtual camera based on the smoothing weight coefficient, the estimated focal length, and the historical focal length;
[0109] Step S850 : determining a smoothed projection transformation relationship according to the target position and the target focal length, and performing cropping processing on the original video frame based on the smoothed projection transformation relationship to obtain a smoothed target video frame.
[0110] Among them, the estimated position refers to the position to which the optical axis of the virtual camera in the current frame needs to be rotated, that is, the theoretical position of the optical axis of the virtual camera at the current moment; the estimated focal length refers to the focal length that needs to be set when the optical axis of the virtual camera in the current frame rotates to the estimated position, that is, the theoretical focal length of the virtual camera at the current moment.
[0111] The smoothing weight coefficient refers to a pre-set parameter used to adjust the rotation and focal length change of the virtual camera. The smoothing weight coefficient can generally include two weight values, and the sum of these two weight values is 1. Of course, the smoothing weight coefficient can also be set to multiple weight values, which are specifically set according to the actual application situation. This example embodiment is not limited to this.
[0112] The estimated position of the virtual camera at the current moment and the historical position at the previous moment can be weighted by a smoothing weight coefficient, so that the change between the target position and the historical position is smoother; the estimated focal length of the virtual camera at the current moment and the historical position at the previous moment can be weighted by a smoothing weight coefficient, so that the change between the target position and the historical position is smoother; by smoothing the rotation position of the optical axis and smoothing the focal length change of the virtual camera, the smoothness between each video frame in the target video frame is effectively improved.
[0113] Optionally, the position threshold and focal length threshold can be set in advance so that the difference between the calculated target position and the historical position is less than the position threshold, and the difference between the calculated target focal length and the historical focal length is less than the focal length threshold. The position threshold and focal length threshold can ensure that the optical axis rotation change and focal length change of the virtual camera between adjacent moments are continuous, further ensuring the smooth change between each video frame in the target video frame, and improving the smoothness of the target video frame.
[0114] For example, the target video frame can be smoothed using the relationship group (4):
[0115]
[0116] Among them, C current and They can represent the estimated position of the optical axis of the current frame virtual camera and the estimated focal length of the current frame virtual camera, respectively. previous and They can represent the historical position to which the optical axis of the virtual camera of the video frame at the previous moment rotates and the historical focal length value of the virtual camera of the video frame at the previous moment respectively; α and β can represent smoothing weight coefficients. Specifically, the larger the α setting is, the stronger the tracking ability of the video frame output by the virtual camera for the moving object is, and the larger the β setting is, the smoother the video frame output by the virtual camera is; T c It can be expressed as the position threshold, T fIt can represent the focal length threshold, that is, the maximum value of the optical axis rotation position and focal length of the previous and next frames, and is used to limit the output of the change amount of the virtual camera of the previous and next frames, effectively ensuring the smoothness of the target video frame.
[0117] Figure 9 The following schematically illustrates a principle diagram of smoothing a virtual camera tracking and shooting process in an exemplary embodiment of the present disclosure.
[0118] refer to Figure 9 As shown, the estimated position 910 to which the optical axis of the virtual camera of the current frame in the target video frame 900 rotates, and the estimated focal length 920 of the virtual camera in the current frame can be determined first. Then, the historical position 930 to which the optical axis of the virtual camera of the previous frame in the target video frame 900 rotates, and the historical focal length 940 of the virtual camera can be obtained; a preset smoothing weight coefficient can be obtained, and the estimated position 910 and the historical position 930 can be smoothed based on the smoothing weight coefficient to obtain the target position 950 of the optical axis of the virtual camera that is finally output; the estimated focal length 920 and the historical focal length 940 can be smoothed based on the smoothing weight coefficient to obtain the target focal length 960 of the virtual camera that is finally output; and the original video frame can be cropped according to the target position 950 and the target focal length 960 to obtain the smoothed target video frame, thereby effectively improving the stability and smoothness of the output target video frame and improving the shooting and tracking quality of the moving object.
[0119] In an exemplary embodiment, target tracking can be performed on the original video frame based on a pre-trained target tracking model to determine the moving object in the original video frame. The target tracking model is constructed based on the kernel correlation filter algorithm (KCF). The target tracking model can output the coordinate position of the detection frame corresponding to the moving object in the original video frame. For example, the face or human body in the original video frame can be detected, the coordinate position of the human body frame or the face frame can be output, and the detected human body or face can be numbered to facilitate subsequent tracking and shooting.
[0120] Alternatively, a moving object in the original video frame can be determined based on a microphone array associated with a real camera. The microphone array may include multiple microphones, and the source of the sound can be determined based on the sound data received by the microphone array, thereby determining the position of the moving object to be tracked in the original video frame.
[0121] It can be understood that the target tracking model can be combined with the microphone array to jointly determine the moving objects that need to be tracked and photographed. For example, in an online meeting scenario, the target tracking model can be used to detect all moving objects (i.e., participants) in the original video frame, and then the microphone array can be used to determine the moving object that is currently speaking. The position of the speaking moving object is used as the position that needs to be tracked and photographed, and the virtual camera is driven to shoot.
[0122] In an exemplary embodiment, the motion object may include an object number, which may be Figure 9 The steps in this paper are used to determine the second optical axis of the virtual camera. Figure 9 Specifically, it may include:
[0123] Step S1010, in response to setting the moving object with the target object number as the protagonist object, determining the number of the moving objects in the original video frame;
[0124] Step S1020: determining the second optical axis of the virtual camera based on the position of the protagonist object, a preset optical axis weight coefficient, and the number of the moving objects, so that the position of the second optical axis is within a preset range around the protagonist object.
[0125] Among them, the target object number refers to the object number of the protagonist object determined by various means. For example, the target object number can be determined by user pre-specification or real-time selection, or it can be determined by a microphone array. Of course, it can also be determined by other means, and this example embodiment does not make any special limitations on this.
[0126] A protagonist object refers to a moving object that the virtual camera focuses on for tracking and capturing. For example, a moving object currently speaking in an online meeting scenario can be the protagonist object, or a player currently dribbling in a football game scenario can be the protagonist object. This example embodiment does not specifically limit the type of protagonist object. After determining the protagonist object, the virtual camera can be driven to focus the output image on or around the protagonist object.
[0127] When two or more moving objects are detected in the original video frame, the terminal device can be controlled to enter the protagonist mode. At this time, the protagonist object can be determined according to the predetermined target object number, and then the second optical axis of the virtual camera can be determined based on the position of the protagonist object, the preset optical axis weight coefficient, and the number of moving objects. Then, by determining the position of the second optical axis, the field of view angle of the virtual camera is driven to be within a preset range around the protagonist object. This example embodiment does not impose any special restrictions on the preset range around the protagonist object, as long as the protagonist object is in the output image of the virtual camera, that is, the target video frame.
[0128] For example, when two or more moving objects are detected in the original video frame, the second optical axis of the virtual camera is focused more on the main object through weighted average calculation. Specifically, the position of the second optical axis of the virtual camera can be determined by the relationship group (5):
[0129]
[0130] Among them, F lead It can represent the position coordinates of the main object (such as the geometric center point). It can represent the position coordinates (such as the geometric center point) of the moving objects other than the protagonist object in the original video frame. γ1 and γ2 can represent the optical axis weight coefficients. The larger the γ1 setting is, the more the optical axis of the virtual camera is focused on the protagonist object. n can represent the number of all moving objects in the original video frame.
[0131] The protagonist mode can better control the virtual camera's tracking and shooting of specific moving objects. The operation is simple and can be achieved without the user adjusting too many parameters, effectively improving the shooting efficiency of the protagonist object.
[0132] In summary, in this exemplary embodiment, a raw video frame captured by a real camera with a fixed shooting angle can be obtained, and a moving object in the raw video frame can be determined. A virtual camera can then be created based on the properties of the real camera and the position of the moving object in the raw video frame. The raw video frame can then be cropped based on the virtual camera to obtain a target video frame whose image follows the movement of the moving object. On the one hand, a raw video frame with a large field of view is captured by a real camera with a fixed shooting angle. Then, by creating a virtual camera, a target video frame with a smaller field of view is obtained based on the raw video frame, but the view angle tracks the movement of the original moving object. This does not require manual operation or the use of a pan / tilt or motor to drive the real camera to rotate, effectively reducing the labor and hardware costs associated with tracking and capturing moving objects. On the other hand, the virtual camera is created based on the properties of the real camera and the position of the moving object in the raw video frame. It is not affected by the manual camera operator's skill or the stability of the pan / tilt or motor, effectively improving the capture quality of the target video frame and ensuring its stability and smoothness.
[0133] This application uses an ultra-wide-angle real camera with a fixed shooting angle to obtain and track moving objects, such as human bodies (faces), through a target detection algorithm, to create a virtual camera, so that the optical axis of the virtual camera points to the center of a single human body (face) or the center of mass of multiple human bodies (faces), rotates with the human body (face) in real time, and can zoom in and out in real time according to the proportion of the person in the picture, achieving the effect of "intelligent camera movement" of the virtual camera.
[0134] It should be noted that the above figures are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0135] For further reference, Figure 11 As shown, in the embodiment of this example, a moving object tracking device 1100 is further provided, comprising a video frame acquisition module 1110, a moving object determination module 1120, a virtual camera creation module 1130, and a video frame cropping module 1140.
[0136] The video frame acquisition module 1110 is used to acquire original video frames, where the original video frames are captured by a real camera with a fixed shooting angle;
[0137] The moving object determination module 1120 is used to determine the moving object in the original video frame;
[0138] The virtual camera creation module 1130 is used to create a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame;
[0139] The video frame cropping module 1140 is configured to crop the original video frame based on the virtual camera to obtain a target video frame in which the image follows the movement of the moving object.
[0140] In an exemplary embodiment, the virtual camera creation module 1130 may include:
[0141] an extrinsic parameter determination unit, configured to determine extrinsic camera parameters and intrinsic camera parameters according to attribute parameters of the real camera and the position of the moving object in the original video frame;
[0142] A creating unit is configured to create a virtual camera based on the extrinsic camera parameters and the intrinsic camera parameters.
[0143] In an exemplary embodiment, the extrinsic camera parameters may include a rotation matrix, the optical centers of the virtual camera and the real camera may be the same, and the attribute parameters of the real camera may include a first optical axis; and the extrinsic parameter determination unit may be configured to:
[0144] determining a second optical axis of the virtual camera according to the position of the moving object in the original video frame and the optical center;
[0145] A rotation matrix corresponding to the virtual camera is determined based on an angle between the first optical axis and the second optical axis.
[0146] In an exemplary embodiment, the intrinsic camera parameters may include an intrinsic parameter matrix, and the intrinsic and extrinsic parameter determination unit may be configured to:
[0147] Determining focal length data corresponding to the virtual camera according to attribute parameters of the real camera;
[0148] Get the preset image cropping ratio;
[0149] An intrinsic parameter matrix corresponding to the virtual camera is determined based on the focal length data and the image cropping ratio.
[0150] In an exemplary embodiment, the video frame cropping module 1140 may be configured to:
[0151] Obtaining the intrinsic parameter matrix of the real camera obtained by calibration;
[0152] Creating a projection transformation relationship from the virtual camera to the real camera according to the intrinsic parameter matrix of the real camera, the rotation matrix corresponding to the virtual camera, and the intrinsic parameter matrix corresponding to the virtual camera;
[0153] A cropping area is determined based on the projection transformation relationship, and the original video frame is cropped using the cropping area to obtain a target video frame whose picture follows the movement of the moving object.
[0154] In an exemplary embodiment, the moving object tracking apparatus 1100 may further include a target video frame smoothing module, which may be configured to:
[0155] Performing smoothing on the target video frame;
[0156] The smoothing process on the target video frame includes:
[0157] Determining an estimated position to which the optical axis of the virtual camera of a current frame in the target video frame is rotated, and determining an estimated focal length of the virtual camera in the current frame;
[0158] Obtaining a historical position to which the optical axis of the virtual camera rotates in a previous frame of the target video frame, and a historical focal length of the virtual camera;
[0159] Determining a target position of the optical axis of the virtual camera based on a preset smoothing weight coefficient, the estimated position, and the historical position;
[0160] determining a target focal length of the virtual camera based on the smoothing weight coefficient, the estimated focal length, and the historical focal length;
[0161] A smoothed projection transformation relationship is determined according to the target position and the target focal length, and the original video frame is cropped based on the smoothed projection transformation relationship to obtain a smoothed target video frame.
[0162] In an exemplary embodiment, the difference between the target position and the historical position may be smaller than a position threshold, and the difference between the target focal length and the historical focal length may be smaller than a focal length threshold.
[0163] In an exemplary embodiment, the moving object determination module 1120 may be configured to:
[0164] Tracking a target in the original video frame based on a pre-trained target tracking model to determine a moving object in the original video frame, wherein the target tracking model is constructed based on a kernel correlation filtering algorithm; and / or
[0165] A moving object in the original video frame is determined based on a microphone array associated with the real camera.
[0166] In an exemplary embodiment, the moving object may include an object number, and the internal and external parameter determination unit may further be configured to:
[0167] In response to setting the moving object with the target object number as the protagonist object, determining the number of the moving objects in the original video frame;
[0168] Based on the position of the protagonist object, a preset optical axis weight coefficient, and the number of the moving objects, a second optical axis of the virtual camera is determined so that the position of the second optical axis is within a preset range around the protagonist object.
[0169] The specific details of each module in the above device have been described in detail in the implementation method part. The undisclosed details can be found in the implementation method part, so they will not be repeated here.
[0170] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0171] The exemplary embodiments of the present disclosure further provide an electronic device. The electronic device may be the aforementioned terminal devices 101, 102, 103, and server 105. Generally, the electronic device may include a processor and a memory, the memory being configured to store executable instructions of the processor, and the processor being configured to execute the aforementioned moving object tracking method by executing the executable instructions.
[0172] Below Figure 12 The structure of the electronic device is exemplarily described by taking the mobile terminal 1200 in FIG. 1 as an example. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 12 The construction in can also be applied to fixed type equipment.
[0173] like Figure 12 As shown, the mobile terminal 1200 may specifically include: a processor 1201, a memory 1202, a bus 1203, a mobile communication module 1204, an antenna 1, a wireless communication module 1205, an antenna 2, a display screen 1206, a camera module 1207, an audio module 1208, a power module 1209 and a sensor module 1210.
[0174] The processor 1201 may include one or more processing units, for example, the processor 1201 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit). The moving object tracking method in this exemplary embodiment may be executed by an AP, a GPU, or a DSP. When the method involves neural network-related processing, it may be executed by an NPU. For example, the NPU may load neural network parameters and execute neural network-related algorithm instructions.
[0175] The encoder can encode (i.e., compress) an image or video to reduce the data size for easy storage or transmission. The decoder can decode (i.e., decompress) the encoded data of the image or video to restore the image or video data. The mobile terminal 1200 can support one or more encoders and decoders, such as: image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0176] The processor 1201 may be connected to the memory 1202 or other components via a bus 1203 .
[0177] Memory 1202 can be used to store computer-executable program code, which includes instructions. Processor 1201 executes various functional applications and data processing of mobile terminal 1200 by running the instructions stored in memory 1202. Memory 1202 can also store application data, such as images, videos, and other files.
[0178] The communication functions of mobile terminal 1200 are implemented through mobile communication module 1204, antenna 1, wireless communication module 1205, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1204 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1200. Wireless communication module 1205 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1200.
[0179] The display screen 1206 is used to implement display functions, such as displaying a user interface, images, and videos. The camera module 1207 is used to implement shooting functions, such as capturing images and videos. The audio module 1208 is used to implement audio functions, such as playing audio and capturing voice. The power module 1209 is used to implement power management functions, such as charging the battery, powering the device, and monitoring battery status.
[0180] The sensor module 1210 may include one or more sensors for implementing corresponding sensing detection functions. For example, the sensor module 1210 may include an inertial sensor for detecting the motion posture of the mobile terminal 1200 and outputting inertial sensing data.
[0181] The exemplary embodiments of the present disclosure further provide a computer-readable storage medium having stored thereon a program product capable of implementing the methods described above in this specification. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of the present disclosure.
[0182] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0183] In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the foregoing.
[0184] In addition, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0185] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
[0186] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for tracking a moving object, characterized in that: include: Obtaining an original video frame, wherein the original video frame is captured by a real camera with a fixed shooting angle; Determining a moving object in the original video frame; Creating a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame; Performing cropping processing on the original video frame based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object, and performing smoothing processing on the target video frame; The smoothing process on the target video frame includes: Determining an estimated position to which the optical axis of the virtual camera of a current frame in the target video frame is rotated, and determining an estimated focal length of the virtual camera in the current frame; Obtaining a historical position to which the optical axis of the virtual camera rotates in a previous frame of the target video frame, and a historical focal length of the virtual camera; Determining a target position of the optical axis of the virtual camera based on a preset smoothing weight coefficient, the estimated position, and the historical position; determining a target focal length of the virtual camera based on the smoothing weight coefficient, the estimated focal length, and the historical focal length; A smoothed projection transformation relationship is determined according to the target position and the target focal length, and the original video frame is cropped based on the smoothed projection transformation relationship to obtain a smoothed target video frame.
2. The method according to claim 1, characterized in that The step of creating a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame includes: Determining extrinsic camera parameters and intrinsic camera parameters according to the property parameters of the real camera and the position of the moving object in the original video frame; A virtual camera is created based on the extrinsic camera parameters and the intrinsic camera parameters.
3. The method according to claim 2, characterized in that The external camera parameters include a rotation matrix, the optical center of the virtual camera is the same as that of the real camera, and the attribute parameters of the real camera include a first optical axis; The determining of the external camera parameters includes: determining a second optical axis of the virtual camera according to the position of the moving object in the original video frame and the optical center; A rotation matrix corresponding to the virtual camera is determined based on an angle between the first optical axis and the second optical axis.
4. The method according to claim 2, characterized in that The intrinsic camera parameters include an intrinsic parameter matrix, and determining the intrinsic camera parameters includes: Determining focal length data corresponding to the virtual camera according to attribute parameters of the real camera; Get the preset image cropping ratio; An intrinsic parameter matrix corresponding to the virtual camera is determined based on the focal length data and the image cropping ratio.
5. The method according to any one of claims 1 to 4, characterized in that The step of cropping the original video frame based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object includes: Obtaining the intrinsic parameter matrix of the real camera obtained by calibration; Creating a projection transformation relationship from the virtual camera to the real camera according to the intrinsic parameter matrix of the real camera, the rotation matrix corresponding to the virtual camera, and the intrinsic parameter matrix corresponding to the virtual camera; A cropping area is determined based on the projection transformation relationship, and the original video frame is cropped using the cropping area to obtain a target video frame whose picture follows the movement of the moving object.
6. The method according to claim 1, characterized in that The difference between the target position and the historical position is smaller than a position threshold, and the difference between the target focal length and the historical focal length is smaller than a focal length threshold.
7. The method according to claim 1, characterized in that The determining of the moving object in the original video frame includes: Tracking a target in the original video frame based on a pre-trained target tracking model to determine a moving object in the original video frame, wherein the target tracking model is constructed based on a kernel correlation filtering algorithm; and / or A moving object in the original video frame is determined based on a microphone array associated with the real camera.
8. The method according to claim 3 or 6, characterized in that The moving object includes an object number, and determining the second optical axis of the virtual camera includes: In response to setting the moving object with the target object number as the protagonist object, determining the number of the moving objects in the original video frame; Based on the position of the protagonist object, a preset optical axis weight coefficient, and the number of the moving objects, a second optical axis of the virtual camera is determined so that the position of the second optical axis is within a preset range around the protagonist object.
9. A moving object tracking device, characterized in that: include: A video frame acquisition module is used to acquire original video frames, where the original video frames are captured by a real camera with a fixed shooting angle; A moving object determination module, configured to determine a moving object in the original video frame; A virtual camera creation module, configured to create a virtual camera according to the attribute parameters of the real camera and the position of the moving object in the original video frame; A video frame cropping module, configured to crop the original video frame based on the virtual camera to obtain a target video frame whose picture follows the movement of the moving object; A target video frame smoothing module, configured to perform smoothing processing on the target video frame; Wherein, the smoothing processing of the target video frame includes: determining the estimated position to which the optical axis of the virtual camera of the current frame in the target video frame is rotated, and determining the estimated focal length of the virtual camera in the current frame; obtaining the historical position to which the optical axis of the virtual camera of the previous frame in the target video frame is rotated, and the historical focal length of the virtual camera; determining the target position of the optical axis of the virtual camera based on a preset smoothing weight coefficient, the estimated position and the historical position; determining the target focal length of the virtual camera based on the smoothing weight coefficient, the estimated focal length and the historical focal length; determining the smoothed projection transformation relationship according to the target position and the target focal length, and cropping the original video frame based on the smoothed projection transformation relationship to obtain a smoothed target video frame.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
11. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 8 by executing the executable instructions.
Citation Information
Patent Citations
Video processing method and electronic equipment
CN113014793A