Video stitching method and device, electronic equipment and computer readable storage medium
By acquiring the real-time status parameters of the camera and dynamically selecting and optimizing the viewing angle and pose parameters, the high latency problem caused by camera heterogeneity in the existing technology is solved, and high-quality real-time video stitching is achieved.
Patent Information
- Application Number
- CN202511131300.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-11
AI Technical Summary
Existing video stitching technologies fail to effectively consider the heterogeneity of cameras, resulting in high processing latency, inability to meet real-time video stitching requirements, and a lack of real-time adaptability to changes in device status.
By acquiring real-time status parameters from multiple cameras, the target camera is dynamically selected, and video stitching with minimal stitching error and processing latency is achieved through optimization based on viewpoint and pose parameters.
It achieves high-quality, real-time video stitching while taking into account camera heterogeneity, reducing stitching errors and processing latency, and meeting the requirements of real-time video stitching.
Smart Images

Figure CN120935470A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, specifically to a video splicing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Video stitching technology is used to stitch together videos captured simultaneously by multiple cameras into a panoramic video. However, this technology treats multiple cameras as identical devices for scheduling, failing to consider the heterogeneity of parameters such as resolution and bandwidth among the cameras. Because large-scale video processing is computationally complex, this technology often relies on centralized processing methods, resulting in high processing latency and failing to meet the requirements of real-time video stitching. Furthermore, this technology lacks real-time adaptability to changes in device status, particularly in the event of temporary device failure or addition, where dynamic scheduling is impossible. Summary of the Invention
[0003] This application provides a video stitching method, apparatus, electronic device, and computer-readable storage medium, which can perform video stitching in real time, dynamically, and with high quality, taking into account the heterogeneity of cameras.
[0004] In a first aspect, embodiments of this application provide a video stitching method, including:
[0005] Obtain real-time status parameters from multiple cameras;
[0006] Under the constraints of satisfying the real-time status parameters of multiple cameras, with the goal of minimizing stitching error and processing delay, the target camera and the viewing angle parameters of the target camera are determined.
[0007] Based on the target camera and its viewing angle parameters, the pose parameters of multiple target cameras are optimized.
[0008] Based on the optimized pose parameters, the videos captured by the multiple target cameras are stitched together.
[0009] In one embodiment, the real-time status parameters of the camera include viewing angle parameters; the stitching error is determined through the following steps:
[0010] Acquire multiple video frames captured by the multiple cameras;
[0011] Obtain a panoramic video frame stitched together from multiple video frames;
[0012] Determine the first coordinate information of the scale-invariant feature transform (SIFT) feature points in multiple video frames;
[0013] Determine the second coordinate information of the target feature point corresponding to the SIFT feature point in the panoramic video frame;
[0014] Based on the viewpoint parameters of the multiple cameras, the pose transformation matrix is determined;
[0015] Based on the pose transformation matrix, the first coordinate information is transformed into third coordinate information in the panoramic space;
[0016] The third coordinate information is projected to obtain the fourth coordinate information;
[0017] The splicing error is determined based on the second coordinate information and the fourth coordinate information.
[0018] In one embodiment, the processing delay is determined by the following steps:
[0019] Acquire pose recalculation delay;
[0020] Get the frame buffer latency caused by caching multiple video frames;
[0021] The sum of the pose recalculation delay and the frame buffer delay is determined as the processing delay.
[0022] In one embodiment, determining the target camera and its viewing angle parameters with the goal of minimizing stitching errors and processing latency includes:
[0023] Obtain the error coefficient and delay coefficient;
[0024] The splicing error is weighted based on the error coefficient to obtain the weighted splicing error;
[0025] The processing delay is weighted based on the delay coefficient to obtain the weighted processing delay.
[0026] The target camera and its viewing angle parameters are determined with the goal of minimizing the sum of the weighted stitching error and the weighted processing delay.
[0027] In one embodiment, the real-time status parameters of the camera include resolution;
[0028] The acquisition of error coefficients includes:
[0029] Determine the resolution requirements;
[0030] Obtain the average resolution of the multiple cameras;
[0031] The error coefficient is determined based on the average resolution of the multiple cameras and the resolution requirement.
[0032] In one embodiment, the real-time status parameters of the camera include bandwidth;
[0033] The acquisition of the delay coefficient includes:
[0034] Obtain the average bandwidth of the multiple cameras;
[0035] The reciprocal of the average bandwidth is determined as the delay coefficient.
[0036] In one embodiment, the method further includes:
[0037] Obtain the learning rate;
[0038] Obtain the partial derivative of the splicing error with respect to the error coefficient;
[0039] The adjustment amount is determined based on the learning rate and the partial derivative.
[0040] The error coefficient is adjusted based on the adjustment amount to obtain the adjusted error coefficient; the adjusted error coefficient is used to redetermine the target camera and the viewing angle parameters.
[0041] In one embodiment, the real-time status parameters of the camera include computing power; satisfying the constraints of the real-time status parameters of multiple cameras includes:
[0042] Obtain the resolution and frame rate of the video streams processed by each of the multiple cameras;
[0043] Obtain the computing power conversion coefficient;
[0044] The computing load of each of the multiple cameras is determined based on the computing power conversion coefficient, the resolution and frame rate of the video streams processed by each of the multiple cameras;
[0045] When the computational load of each camera does not exceed its own computational capability, the computational capability constraints of multiple cameras are satisfied.
[0046] In one embodiment, the real-time status parameters of the camera include a bandwidth limit; satisfying the constraints of the real-time status parameters of multiple cameras includes:
[0047] Obtain the sum of the bandwidth occupied by the frame rates of one or more video streams output by the multiple cameras;
[0048] If the sum of the bandwidths corresponding to each of the multiple cameras does not exceed the upper limit of the bandwidth of the camera, then the bandwidth constraint of the camera is determined to be satisfied.
[0049] Secondly, embodiments of this application provide a video stitching device, including:
[0050] The acquisition module is used to acquire real-time status parameters from multiple cameras;
[0051] The determination module is used to determine the target camera and the viewing angle parameters of the target camera under the constraints of the real-time status parameters of multiple cameras, with the goal of minimizing stitching error and processing delay;
[0052] An optimization module is used to optimize the pose parameters of multiple target cameras based on the target camera and the viewpoint parameters of the target camera;
[0053] The stitching module is used to stitch together videos captured by multiple target cameras based on optimized pose parameters.
[0054] Thirdly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the video stitching method described above.
[0055] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the video stitching method described above.
[0056] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.
[0057] The embodiments of this application have the following beneficial effects:
[0058] The target camera is selected from multiple cameras, taking into account their heterogeneity. The target camera and its viewing angle parameters are determined by considering the real-time state parameters of multiple cameras, allowing for real-time adaptation and adjustment to changes in their states. The target camera and its viewing angle parameters are determined with the goal of minimizing stitching error and processing latency, resulting in smaller stitching error, higher stitching quality, and lower processing latency, meeting the requirements for real-time video stitching. Furthermore, the pose parameters of the multiple target cameras have been optimized to improve the video stitching quality. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 This is a schematic diagram illustrating the steps of a video stitching method provided in an embodiment of this application;
[0061] Figure 2 This is a schematic diagram of the connection of multiple cameras provided in one embodiment of this application;
[0062] Figure 3 This is a dependency relationship diagram provided in an embodiment of this application;
[0063] Figure 4 This is a schematic diagram of the structure of a video splicing device provided in one embodiment of this application;
[0064] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0065] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0066] In one embodiment, such as Figure 1 As shown, a video stitching method is provided. Although the logical order is illustrated in the step diagram, in some cases, the steps shown or described can be performed in a different order than that shown in the diagram. Specifically, this video stitching method can be applied to electronic devices such as terminals or servers. The terminal can be, but is not limited to, one or more of smartphones, tablets, laptops, desktop computers, and in-vehicle computers. The server can be a physical server or a cloud server providing various cloud services. It is worth noting that this application does not limit the number of terminals or servers; any number of terminals or servers can be used depending on implementation needs. For example, a server can be a single server or a server cluster composed of multiple servers, etc.
[0067] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0068] according to Figure 1The video stitching method shown includes at least steps S110 to S140, which are described in detail below:
[0069] In step S110, real-time status parameters of multiple cameras are obtained.
[0070] A camera can be a standalone device or a component on an electronic device used to capture video.
[0071] Figure 2 This is a schematic diagram illustrating the connection of multiple cameras according to an embodiment of this application. The multiple cameras are connected via a software bus, enabling unified management of their resources. A software bus is a virtual communication architecture that enables data transmission and resource sharing between devices through a software-defined approach. It can break hardware interface limitations and build flexible cross-device connections. Multiple cameras connected via the software bus can form a resource pool. This resource pool can interconnect dispersed cameras through a network, integrating them into a unified, on-demand camera cluster to achieve efficient scheduling and collaborative video stitching.
[0072] The real-time status parameters of a camera are used to characterize the camera's current operating status and performance attributes. Optionally, the real-time status parameters of a camera may include, but are not limited to, any one or more of the following: computing power, available bandwidth, resolution, and viewing angle parameters.
[0073] The computing power of a camera refers to the computational performance of the camera's built-in chip or associated processing unit in performing real-time analysis and processing of image data.
[0074] The available bandwidth of a camera refers to the maximum data transmission rate that can be stably used in the current network environment when the camera transmits data with external systems (such as servers or terminal devices). It is used to reflect the ability and stability of real-time data transmission.
[0075] The resolution of a camera refers to the number of pixels in the image captured and / or processed by the camera, reflecting the image's clarity and ability to capture details.
[0076] A camera's field of view parameters can include its spatial position and orientation. The camera's spatial position refers to its specific coordinates within the physical environment, identifying its geographical or scene location. The camera's orientation refers to the real-time pointing direction of the camera lens (including pitch and yaw angles), which determines its field of view and the direction of the captured image. The camera's orientation can be acquired by an inertial measurement unit (IMU).
[0077] The real-time status parameters of the cameras can be acquired in real time and represent the current status parameters of the cameras. Multiple cameras report their own real-time status parameters in real time. The electronic device constructs a resource pool based on the real-time status parameters reported by multiple cameras, dynamically selects and schedules the cameras in the resource pool, and then stitches the videos captured by the cameras together.
[0078] In step S120, under the constraints of the real-time status parameters of multiple cameras, the target camera and its viewing angle parameters are determined with the goal of minimizing stitching error and processing delay.
[0079] Meeting the constraints of the real-time status parameters of multiple cameras means that the real-time status parameters of multiple cameras, after dynamic scheduling and task allocation, do not exceed the constraints of the real-time status parameters of multiple cameras.
[0080] The stitching error is the sum of reprojection errors generated during feature point matching and pose optimization in the geometric modeling module. Stitching error reflects the consistency of feature point projections from video frames onto the panoramic space; a smaller stitching error indicates higher stitching quality.
[0081] Processing latency refers to the time interval from when multiple cameras acquire video frames to when the stitching process is completed and the complete stitched image is output, reflecting the system's real-time processing efficiency of video data.
[0082] To avoid performance degradation, malfunctions, hardware overload, and / or other security risks associated with cameras, it is necessary to constrain the tasks performed by the cameras to ensure they meet the constraints of the cameras' real-time status parameters. For example, this means ensuring that the current computational load of a camera does not exceed its current computing capacity, and / or that the bandwidth occupied by the data being transmitted by a camera does not exceed the available bandwidth of the camera.
[0083] To ensure stitching quality and improve video stitching efficiency, multiple cameras can be selected with the goal of minimizing stitching errors and processing latency. This results in a camera selection vector and viewing angle parameters for multiple target cameras. The camera selection vector indicates whether each camera is selected as a target camera. Target cameras are those selected for video capture and transmission. Multiple target cameras capture video and transmit it to electronic equipment. The electronic equipment then stitches the videos captured by the multiple target cameras to obtain a panoramic video. Alternatively, multiple target cameras can capture video and transmit it to a switch, which then forwards the video to the electronic equipment used for video stitching.
[0084] In step S130, the pose parameters of multiple target cameras are optimized based on the target camera and its viewpoint parameters.
[0085] The viewing angle parameters of the target camera are its initial parameters; however, these initial parameters are unreliable. For example, the camera may shift due to wind and vibration; tilted bracket installation may cause pitch angle errors; and changes in the lens flange focal length at high temperatures may affect the equivalent optical center position. Actual measurements in the stadium show that wind and vibration cause camera displacement greater than 5cm, and tilted bracket installation causes pitch angle errors greater than 3°.
[0086] To facilitate understanding, the camera's viewpoint and pose parameters will be explained first. Please refer to Table 1 for details.
[0087] Table 1
[0088]
[0089]
[0090] Table 1 compares the camera's viewpoint and pose parameters. In the mathematical representation of the viewpoint parameters, Θ... i The t represents the viewpoint parameter of the i-th camera. i The translation vector, r i The rotation parameters are represented by T, which represents the transpose of the matrix. T allows the translation and rotation parameters to be combined into a single column vector, facilitating processing by optimization algorithms. The translation vector can be determined using GPS coordinates, while the rotation parameters can be Euler angles, determined by an IMU. In the mathematical representation of the pose parameters, T... i R represents the pose parameters of the i-th camera. i It is a rotation matrix used to characterize the camera's orientation, t i It is a translation vector used to represent the position of the camera.
[0091] The camera's viewpoint parameters, used as the initial input to the optimization problem, are directly reported by the hardware (e.g., GPS, IMU) and may contain errors. The camera's pose parameters, optimized through a geometry module, represent the precise pose and are used as the optimized result for video stitching. These pose parameters are optimized extrinsic parameters, characterizing the camera's precise position and orientation in the global coordinate system. Compared to the viewpoint parameters, the pose parameters eliminate initial measurement errors through an algorithm.
[0092] The viewpoint parameters in this embodiment can be considered as initial, unoptimized pose parameters. These viewpoint parameters only provide coarse positioning for the camera and cannot be directly used for high-precision stitching. Optimization is required to improve stitching accuracy and thus enhance stitching quality. Table 2 compares the stitching results of video stitching using the camera's viewpoint parameters directly versus using optimized pose parameters. Table 2 confirms that to ensure stitching quality, the camera's pose parameters need to be optimized for video stitching.
[0093] Table 2
[0094]
[0095] The pose parameters of the target camera can be optimized using Lie groups and Lie algebras, based on the target camera and its viewpoint parameters, with the goal of minimizing the stitching error. The viewpoint parameters of the target camera are used to provide initial values for the Lie group optimization.
[0096] In one embodiment, the optimized pose parameters are determined according to the following formula:
[0097]
[0098] The meaning and mathematical form of each character in this formula can be found in Table 3.
[0099] Table 3
[0100]
[0101] The core objective of pose optimization is to accurately calculate the 6-DOF pose (3 rotations + 3 translations) of each camera by matching feature point data.
[0102] In this formula, T(ξ)·p j The representation transforms the initial coordinate information into panoramic space. π is used to project 3D points in panoramic space onto the 2D image plane. q j This represents the actual observed location of feature points within a panoramic video frame. ||π(·)-q j || Used to calculate the reprojection error of feature points, which can reflect the accuracy of pose estimation. The smaller the value, the closer the projected position is to the actual observed position. The representation sums the reprojection errors of all matching feature points. The goal is to find the optimal pose parameters to minimize the total weight projection error (stitching error).
[0103] The optimization steps may include feature matching, constructing the optimization problem, and iterative optimization solution. Feature matching includes extracting SIFT feature points from video frames captured by the target camera, matching corresponding points between different cameras, and generating point pairs p. j and q j The optimization problem involves obtaining the initial pose parameters of the target camera (i.e., the viewpoint parameters reported by the target camera) and constructing the objective function. The meanings of each character in the objective function can be found in the preceding text. Iterative optimization can be performed using numerical optimization tools (such as those in Python or MATLAB) or the Levenberg-Marquardt (LM) algorithm, ultimately yielding optimized pose parameters. Based on these optimized pose parameters, an optimized pose transformation matrix is determined. This matrix is then used to perform perspective transformation on the video frames captured by the target camera. Finally, the multiple perspective-transformed video frames are fused to generate a seamless panoramic video frame.
[0104] In step S140, the videos captured by multiple target cameras are stitched together based on the optimized pose parameters.
[0105] Based on the optimized pose parameters of each target camera, the transformation matrix corresponding to each target camera is determined. Then, the video captured by each target camera is transformed based on the transformation matrix corresponding to each target camera. The transformed videos are then stitched together to obtain a panoramic video.
[0106] The technical solution adopted in this application embodiment uses a target camera determined from multiple cameras, taking into account the heterogeneity of the multiple cameras; the target camera and its viewing angle parameters are determined considering the real-time state parameters of multiple cameras, which can adapt to the state changes of multiple cameras in real time and make real-time adjustments; the target camera and its viewing angle parameters are determined with the goal of minimizing stitching error and processing latency, so the stitching error is small, the stitching quality is high, and the processing latency is low, which can meet the requirements of real-time video stitching; the pose parameters of multiple target cameras are optimized, which can improve the stitching quality of the video.
[0107] Based on the above technical solution, as an embodiment, the real-time status parameters of the camera include viewing angle parameters; the stitching error can be determined through the following steps: acquiring multiple video frames captured by multiple cameras; acquiring a panoramic video frame stitched together from the multiple video frames; determining the first coordinate information of the scale-invariant feature transform (SIFT) feature points in the multiple video frames; determining the second coordinate information of the target feature points corresponding to the SIFT feature points in the panoramic video frame; determining the pose transformation matrix based on the viewing angle parameters of the multiple cameras; transforming the first coordinate information into third coordinate information in the panoramic space based on the pose transformation matrix; projecting the third coordinate information to obtain fourth coordinate information; and determining the stitching error based on the second and fourth coordinate information.
[0108] In the process of finding the target camera, it is necessary to try different camera combinations, calculate the stitching error corresponding to different camera combinations, and determine the camera in the camera combination that minimizes the stitching error and processing delay as the target camera.
[0109] It can acquire video frames from multiple cameras in a camera array, extract SIFT feature points from each video frame, and determine the first coordinate information of each SIFT feature point in the video frame. It can then stitch together the video frames from multiple cameras into a panoramic video frame.
[0110] For each camera in the camera array, determine the second coordinate information of the target feature point corresponding to the SIFT feature point in the panoramic video frame. When solving for the target camera, a coarse calculation is performed based on the camera's viewpoint parameters; therefore, the pose transformation matrix is directly determined based on the viewpoint parameters of multiple cameras. Based on this pose transformation matrix, the first coordinate information of the SIFT feature point in the video frame is transformed into the third coordinate information in panoramic space. The third coordinate information is projected to obtain the fourth coordinate information. The difference between the second and fourth coordinate information is determined as the reprojection error. The sum of the reprojection errors corresponding to each camera in the camera array is determined as the stitching error corresponding to the camera array. Based on the stitching errors corresponding to multiple camera arrays, and under the constraint of the real-time state parameters of multiple cameras, the target camera array is determined with the goal of minimizing the stitching error and processing latency. The multiple cameras in the target camera array are then identified as the target cameras.
[0111] In one embodiment, the splicing error can be determined using the following formula:
[0112]
[0113] Where E is the splicing error; k is the total number of matched feature points; p j The first coordinate information (3D homogeneous coordinates) of the j-th SIFT feature point in the video frame captured by the camera; q j For p j The second coordinate information (2D or 3D coordinates, depending on the projection) of the target feature point corresponding to the panoramic video frame; T is the pose transformation matrix, belonging to the special Euclidean group SE(3), representing the camera's extrinsic parameters (rotation and translation), which are determined based on the viewpoint parameters; ξ is the Lie algebra. The elements in the text; π is the projection function used to map 3D points onto a 2D image plane (e.g., perspective projection or orthographic projection); ||| 2 It is the Euclidean norm, used to calculate the deviation between the projected position and the desired position of a feature point.
[0114] In this formula, T(ξ)·p j The representation transforms the initial coordinate information into panoramic space. π is used to project 3D points in panoramic space onto the 2D image plane. q j Characterizes the actual observed location of SIFT feature points in a panoramic video frame. ||π(·)-q j || Used to calculate the reprojection error of feature points, which can reflect the accuracy of pose estimation. The smaller the value, the closer the projected position is to the actual observed position.
[0115] Lie groups and Lie algebras offer several advantages, including avoiding singularities, efficient optimization, and global convergence. Specifically, while the rotation matrix R contains nine parameters, the actual rotation is determined by only three degrees of freedom, leading to overparameterization. The Lie algebra ξ, however, can concisely represent the rotation with only three parameters, thus avoiding singularities. By directly differentiating ξ, the processing of rotations and translations can be unified, making the optimization process more efficient. Furthermore, Lie groups and Lie algebras avoid the gimbal lock problem that occurs with traditional Euler angles, providing greater convergence assurance for optimization operations at the global level.
[0116] The technical solution adopted in this application has multiple advantages such as avoiding singularity, efficient optimization and global convergence. Therefore, determining the splicing error based on the Lie group-Lie algebra can make the calculation of the splicing error simpler, more efficient and accurate, and can also avoid the problems of singularity and complex constraints of traditional methods, resulting in better splicing effect and faster calculation.
[0117] Based on the above technical solution, as an embodiment, the processing latency can be determined by the following steps: obtaining the pose recalculation latency; obtaining the frame buffering latency caused by caching multiple video frames; and determining the sum of the pose recalculation latency and the frame buffering latency as the processing latency.
[0118] In multi-view video stitching, the processing latency mainly comes from pose recalculation latency and frame buffer latency. Pose recalculation latency refers to the time interval from triggering pose recalculation to completing the solution of the new pose parameters. Buffer latency refers to the time interval from when a video frame enters the buffer to when it is retrieved from the buffer for stitching processing. The pose recalculation latency and frame buffer latency can be determined separately, and then their sum can be used to determine the processing latency.
[0119] The start time can be recorded when initiating pose recalculation, and then the time it takes to output the optimized pose parameters after calculation can be recorded. The start time can be subtracted from the completion time, and the average of multiple tests (different scenarios, deviations) can be taken to obtain the pose recalculation delay.
[0120] A capture timestamp can be added to each video frame captured by the camera to record the capture time of the video frame; when a video frame is retrieved from the buffer for stitching, the read time is recorded. The single-frame latency is obtained by subtracting the capture time from the read time, and statistical analysis is performed on multiple single-frame latencies to obtain the frame buffer latency.
[0121] The technical solution of this application embodiment determines the sum of pose recalculation latency and frame buffer latency as the processing latency, decomposing the complex processing latency into two core controllable parts, reducing the difficulty of analysis; determining the processing latency makes it easier for relevant personnel to intuitively control the total latency, thereby ensuring smooth output and meeting user needs.
[0122] Based on the above technical solution, as an embodiment, the real-time status parameters of the camera include computing power; satisfying the constraints of the real-time status parameters of multiple cameras may include satisfying computing power constraints. Whether the computing power constraints are satisfied can be determined through the following steps: obtaining the resolution and frame rate of the video streams processed by each of the multiple cameras; obtaining the computing power conversion coefficient; determining the computing load of each of the multiple cameras based on the computing power conversion coefficient, the resolution and frame rate of the video streams processed by each of the multiple cameras; and determining that the computing power constraints of multiple cameras are satisfied when the computing load of each camera does not exceed its own computing power.
[0123] The computational load of the camera can be determined using the following formula:
[0124] w j =μ×res j ×log(fps j );
[0125] Among them, w j The computational load of the j-th camera is μ; the computational power conversion coefficient is res. j The resolution of the video stream processed by the j-th camera; fps j The frame rate of the video stream processed by the j-th camera.
[0126] The computing power reported by the camera refers to the computing power that the camera can use for video stitching tasks. When determining the target camera, in order to meet the computing power constraint, the computing load of the target camera cannot exceed the computing power of the target camera.
[0127] By adopting the technical solution of this application embodiment, and by satisfying the computing power constraint, it is possible to prevent crashes or stuttering caused by the computing load exceeding the computing power of the camera, thereby ensuring the stable operation of multi-view video stitching.
[0128] Based on the above technical solution, as an embodiment, the real-time status parameters of the camera include a bandwidth limit. Meeting the constraints of the real-time status parameters of multiple cameras may include meeting the camera's bandwidth constraints. Specifically, the sum of the bandwidth occupied by the frame rates of one or more video streams output by multiple cameras is obtained; if the sum of the bandwidths corresponding to each of the multiple cameras does not exceed the camera's bandwidth limit, it is determined that the camera's bandwidth constraints are met.
[0129] The bandwidth constraint of a camera can be characterized by the following formula:
[0130]
[0131] Among them, Ф k Γ represents the bandwidth occupied by video stream k; i The bandwidth limit for the i-th camera; fps k The frame rate of the k-th video stream processed by the camera; K i The video stream processed for the i-th camera.
[0132] The bandwidth limit reported by the camera is the maximum bandwidth that the camera can use for video stitching tasks. When determining the target camera, in order to meet the computing power constraints, the sum of the bandwidth occupied by the frame rates of one or more video streams output by the target camera cannot exceed the bandwidth limit of that target camera.
[0133] In one embodiment, the bandwidth constraint corresponding to the switch also needs to be satisfied. The bandwidth constraint of the switch can be characterized by the following formula:
[0134] ∑ i∈N ξ i ×Γ i ≤Γ switch ;
[0135] Among them, Γ switch ξ represents the maximum bandwidth of the switch; N represents the total number of cameras; i The activation state of the camera is represented by ξ, where ξ corresponds to the camera. i =0 indicates that the camera is not the target camera, and the corresponding ξ of the camera is...i =1 indicates that the camera is the target camera.
[0136] By adopting the technical solution of this application embodiment, the bandwidth occupied by the video to be transmitted for the video stitching task undertaken by the camera can be limited to not exceeding the available bandwidth of the camera, thereby ensuring that the bandwidth of multiple cameras can undertake the allocated video stitching task.
[0137] In one embodiment, satisfying the constraints of the real-time status parameters of multiple cameras may include satisfying the resolution requirements of the cameras, ensuring that the total resolution of the multiple target cameras meets the user's requirements.
[0138] Based on the above technical solutions, as one embodiment, minimizing stitching error and processing latency can be achieved by minimizing the sum of stitching error and processing latency. In another embodiment, it can be achieved by minimizing the weighted sum of stitching error and processing latency. In this embodiment, determining the target camera and its viewing angle parameters with the goal of minimizing stitching error and processing latency can include: obtaining error coefficients and latency coefficients; weighting the stitching error based on the error coefficients to obtain a weighted stitching error; weighting the processing latency based on the latency coefficients to obtain a weighted processing latency; and determining the target camera and its viewing angle parameters with the goal of minimizing the sum of the weighted stitching error and the weighted processing latency.
[0139] To minimize the sum of weighted stitching error and weighted processing delay, the camera selection vector and viewpoint parameters are determined, which can be represented by the following formula:
[0140] min X,Θ (αE(X,Θ)+βT(X));
[0141] Where X represents the target camera, Θ represents the viewing angle parameter, E(X,Θ) is the stitching error, T(X) is the processing delay, α is the error coefficient, and β is the delay coefficient.
[0142] In one embodiment, obtaining the error coefficient may include: obtaining the resolution requirement; obtaining the average resolution of multiple cameras; and determining the error coefficient based on the average resolution of the multiple cameras and the resolution requirement.
[0143] The error coefficient α can be determined using the following formula:
[0144]
[0145] Among them, R avg R is the average resolution of multiple cameras. req To meet the resolution requirements requested by the user.
[0146] Rreq It is a two-dimensional vector, in the form of:
[0147] R req =[W req H req ];
[0148] Among them, W req The minimum total width resolution required by the user, in pixels; H req The minimum total height resolution required by the user, in pixels.
[0149] R req This represents the minimum resolution standard required for the final stitching of the panoramic video; for example, live sports broadcasts may require 4K resolution, i.e., W. req =3840, H req =2160.
[0150] For example, if a user requests a panoramic video resolution of 1920×1080 (Full HD), and the average resolution of the device resource pool is 1280×720, then:
[0151]
[0152] In this way, the error coefficient can be determined by the average resolution and resolution requirements of multiple cameras. When the average resolution and resolution requirements change, the error coefficient can be adjusted in real time, thereby ensuring that the optimization target is updated in real time, and thus determining more accurate target camera and viewing angle parameters.
[0153] In one embodiment, obtaining the latency coefficient may include: obtaining the average bandwidth of multiple cameras; and determining the latency coefficient as the reciprocal of the average bandwidth.
[0154] The time delay coefficient β can be determined using the following formula:
[0155]
[0156] Among them, B avg This represents the average bandwidth across multiple cameras.
[0157] The higher the bandwidth, the lower the impact of processing latency. Therefore, the reciprocal of the average bandwidth of multiple cameras can be directly used to determine the latency coefficient.
[0158] The error coefficient can be dynamically adjusted by combining the average resolution of multiple cameras with the user's resolution requirements. When the average resolution of multiple cameras is close to the user's resolution requirements, the error coefficient decreases, and the optimization objective focuses more on reducing latency. When the average resolution of multiple cameras is much lower than the user's resolution requirements, the error coefficient increases, and the optimization objective focuses more on reducing stitching errors.
[0159] The technical solution of this application embodiment can adjust the optimization target in real time based on the error coefficient and the time delay coefficient, thereby guiding the optimization in the desired direction and obtaining more accurate target camera and viewing angle parameters.
[0160] Based on the above technical solution, as an embodiment, the method may further include: obtaining a learning rate; obtaining the partial derivative of the stitching error with respect to the error coefficient; determining an adjustment amount based on the learning rate and the partial derivative; adjusting the error coefficient based on the adjustment amount to obtain the adjusted error coefficient; and using the adjusted error coefficient to redetermine the camera selection vector and viewpoint parameters.
[0161] The partial derivative of the splicing error with respect to the error coefficient reflects the instantaneous rate of change of the splicing error with respect to the error coefficient. The partial derivative of the splicing error with respect to the error coefficient can be positive, negative, or zero.
[0162] The learning rate is positive and can be pre-set. The adjustment amount is determined by multiplying the learning rate and the partial derivative of the stitching error with respect to the error coefficient. The adjustment amount can be positive, negative, or 0. The sum of the adjustment amount and the error coefficient is determined as the adjusted error coefficient. Based on the adjusted error coefficient, the target camera and viewpoint parameters corresponding to the next video frame can be determined.
[0163] The adjusted error coefficient can be determined using the following formula:
[0164]
[0165] Where, α new The adjusted error coefficient; α old The error coefficient before adjustment is η; the learning rate is E; and the splicing error is E. This is the partial derivative of the splicing error with respect to the error coefficient.
[0166] By adjusting the error coefficient, we can achieve the following: when the splicing error is too large, increase the error coefficient so that the next frame prioritizes optimizing the splicing accuracy; when the splicing error is small, decrease the error coefficient so that the next frame focuses on reducing processing latency.
[0167] In this way, the target camera and viewing angle parameters can be determined by the error coefficient, the stitching error can be determined based on the target camera and viewing angle parameters, and the error coefficient can be updated based on the stitching error, forming a closed-loop control. This allows the system to adaptively optimize stitching quality and processing latency.
[0168] Figure 3 This is a dependency diagram provided in one embodiment of this application. As an embodiment, refer to... Figure 3The process can begin with camera selection. Real-time status parameters from multiple cameras are sent to the resource scheduling module. This module can then select a target camera and report its viewing angle parameters. Based on the target camera and its viewing angle parameters, pose parameters can be optimized. These optimized pose parameters are then used to perform geometric transformations on video frames, converting them into panoramic video frames. These panoramic video frames are then stitched together to obtain a panoramic video. After pose optimization, the stitching error corresponding to the optimized pose parameters is simultaneously obtained. This stitching error is used to update the error coefficients, which are then sent to the resource scheduling module. This allows the resource scheduling module to determine the target camera and viewing angle parameters based on the updated error coefficients the next time it needs to be selected.
[0169] The viewpoint parameters are used to initiate the optimization of the pose parameters, providing initial values for the Lie group optimization at the beginning of the optimization phase. The optimized pose parameters are then used to determine the transformation matrix of the video. The stitching error is fed back to the error coefficients, which can drive the re-determination of the target camera in the next frame, realizing dynamic constraints with feedback.
[0170] In this embodiment, Lie group-Lie algebra representation of stitching error is adopted. By minimizing the stitching error, the camera pose can be accurately solved, thus solving the problem of rotation parameter singularity. The stitching error can be dynamically adjusted to optimize the global target, enabling sub-pixel alignment accuracy in more than 95% of multi-view videos, meeting the high requirements of scenarios such as sports live streaming and security monitoring.
[0171] To facilitate better implementation of the video stitching method of this application, this application also provides a video stitching device based on the above-described video stitching method. The meanings of the terms used are the same as in the video stitching method described above, and specific implementation details can be found in the descriptions of the method embodiments.
[0172] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of the video splicing device provided in the embodiments of this application, wherein the video splicing device includes:
[0173] The acquisition module 401 is used to acquire real-time status parameters of multiple cameras;
[0174] The determining module 402 is used to determine the target camera and the viewing angle parameters of the target camera under the constraints of the real-time status parameters of multiple cameras, with the goal of minimizing stitching error and processing delay;
[0175] The optimization module 403 is used to optimize the pose parameters of multiple target cameras based on the target camera and the viewpoint parameters of the target camera;
[0176] The stitching module 404 is used to stitch together videos captured by multiple target cameras according to the optimized pose parameters.
[0177] In one embodiment, the real-time status parameters of the camera include viewing angle parameters; the stitching error is determined through the following steps:
[0178] Acquire multiple video frames captured by the multiple cameras;
[0179] Obtain a panoramic video frame stitched together from multiple video frames;
[0180] Determine the first coordinate information of the scale-invariant feature transform (SIFT) feature points in multiple video frames;
[0181] Determine the second coordinate information of the target feature point corresponding to the SIFT feature point in the panoramic video frame;
[0182] Based on the viewpoint parameters of the multiple cameras, the pose transformation matrix is determined;
[0183] Based on the pose transformation matrix, the first coordinate information is transformed into third coordinate information in the panoramic space;
[0184] The third coordinate information is projected to obtain the fourth coordinate information;
[0185] The splicing error is determined based on the second coordinate information and the fourth coordinate information.
[0186] In one embodiment, the processing delay is determined by the following steps:
[0187] Acquire pose recalculation delay;
[0188] Obtain the frame buffering latency caused by caching multiple video frames;
[0189] The sum of the pose recalculation delay and the frame buffer delay is determined as the processing delay.
[0190] In one embodiment, the determining module 402 includes:
[0191] The coefficient acquisition unit is used to acquire the error coefficient and the time delay coefficient;
[0192] The first weighting unit is used to perform weighted processing on the splicing error based on the error coefficient to obtain the weighted splicing error;
[0193] The second weighting unit is used to perform weighted processing on the processing delay based on the delay coefficient to obtain the weighted processing delay.
[0194] The parameter determination unit is used to determine the target camera and the viewing angle parameters of the target camera with the goal of minimizing the sum of the weighted stitching error and the weighted processing delay.
[0195] In one embodiment, the real-time status parameters of the camera include resolution; the coefficient acquisition unit includes:
[0196] The requirement acquisition subunit is used to acquire resolution requirements;
[0197] A resolution acquisition subunit is used to acquire the average resolution of the multiple cameras;
[0198] The error coefficient determination subunit is used to determine the error coefficient based on the average resolution of the multiple cameras and the resolution requirement.
[0199] In one embodiment, the real-time status parameters of the camera include bandwidth; the coefficient acquisition unit includes:
[0200] A bandwidth acquisition subunit is used to acquire the average bandwidth of the multiple cameras;
[0201] The delay coefficient determination subunit is used to determine the delay coefficient by taking the reciprocal of the average bandwidth.
[0202] In one embodiment, the device further includes:
[0203] The learning rate acquisition module is used to acquire the learning rate;
[0204] A partial derivative acquisition module is used to acquire the partial derivative of the splicing error with respect to the error coefficient;
[0205] The adjustment amount determination module is used to determine the adjustment amount based on the learning rate and the partial derivative;
[0206] The coefficient adjustment module is used to adjust the error coefficient based on the adjustment amount to obtain the adjusted error coefficient; the adjusted error coefficient is used to redetermine the target camera and the viewing angle parameters.
[0207] In one embodiment, the real-time status parameters of the camera include computing power; satisfying the constraints of the real-time status parameters of multiple cameras includes:
[0208] Obtain the resolution and frame rate of the video streams processed by each of the multiple cameras;
[0209] Obtain the computing power conversion coefficient;
[0210] The computing load of each of the multiple cameras is determined based on the computing power conversion coefficient, the resolution and frame rate of the video streams processed by each of the multiple cameras;
[0211] When the computational load of each camera does not exceed its own computational capability, the computational capability constraints of multiple cameras are satisfied.
[0212] In one embodiment, the real-time status parameters of the camera include a bandwidth limit; satisfying the constraints of the real-time status parameters of multiple cameras includes:
[0213] Obtain the sum of the bandwidth occupied by the frame rates of one or more video streams output by the multiple cameras;
[0214] If the sum of the bandwidths corresponding to each of the multiple cameras does not exceed the upper limit of the bandwidth of the camera, then the bandwidth constraint of the camera is determined to be satisfied.
[0215] The technical solution adopted in this application embodiment uses a target camera determined from multiple cameras, taking into account the heterogeneity of the multiple cameras; the target camera and its viewing angle parameters are determined considering the real-time state parameters of multiple cameras, which can adapt to the state changes of multiple cameras in real time and make real-time adjustments; the target camera and its viewing angle parameters are determined with the goal of minimizing stitching error and processing latency, so the stitching error is small, the stitching quality is high, and the processing latency is low, which can meet the requirements of real-time video stitching; the pose parameters of multiple target cameras are optimized, which can improve the stitching quality of the video.
[0216] Specific limitations regarding the video splicing device can be found in the limitations of the video splicing method described above, and will not be repeated here. Each module in the aforementioned video splicing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0217] In addition, this application also provides an electronic device, such as Figure 5 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically:
[0218] The electronic device may include components such as a processor 501 with one or more processing cores and a memory 502 of one or more computer-readable storage media. Those skilled in the art will understand that... Figure 5The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0219] The processor 501 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.
[0220] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0221] In one embodiment, the electronic device further includes a power supply 503 that supplies power to the various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0222] In one embodiment, the electronic device may further include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0223] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 502 according to the following instructions, and the processor 501 runs the applications stored in the memory 502, thereby implementing the steps in any of the video splicing methods provided in the embodiments of this application.
[0224] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0225] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described in any embodiment of this application.
[0226] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of this application.
[0227] In some embodiments, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the methods described in any embodiment of this application.
[0228] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0229] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0230] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the video stitching methods provided in this application.
[0231] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0232] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0233] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the video splicing methods provided in this application, the beneficial effects that any of the video splicing methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0234] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0235] The foregoing has provided a detailed description of a video splicing method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A video stitching method, characterized in that, include: Obtain real-time status parameters from multiple cameras; Under the constraints of satisfying the real-time status parameters of multiple cameras, with the goal of minimizing stitching error and processing delay, the target camera and the viewing angle parameters of the target camera are determined. Based on the target camera and its viewing angle parameters, the pose parameters of multiple target cameras are optimized. Based on the optimized pose parameters, the videos captured by the multiple target cameras are stitched together.
2. The method according to claim 1, characterized in that, The real-time status parameters of the camera include the viewing angle parameter; the stitching error is determined through the following steps: Acquire multiple video frames captured by the multiple cameras; Obtain a panoramic video frame stitched together from multiple video frames; Determine the first coordinate information of the scale-invariant feature transform (SIFT) feature points in multiple video frames; Determine the second coordinate information of the target feature point corresponding to the SIFT feature point in the panoramic video frame; Based on the viewpoint parameters of the multiple cameras, the pose transformation matrix is determined; Based on the pose transformation matrix, the first coordinate information is transformed into third coordinate information in the panoramic space; The third coordinate information is projected to obtain the fourth coordinate information; The splicing error is determined based on the second coordinate information and the fourth coordinate information.
3. The method according to claim 1, characterized in that, The processing delay is determined by the following steps: Acquire pose recalculation delay; Get the frame buffer latency caused by caching multiple video frames; The sum of the pose recalculation delay and the frame buffer delay is determined as the processing delay.
4. The method according to claim 1, characterized in that, The process of determining the target camera and its viewing angle parameters, with the goal of minimizing stitching errors and processing latency, includes: Obtain the error coefficient and delay coefficient; The splicing error is weighted based on the error coefficient to obtain the weighted splicing error; The processing delay is weighted based on the delay coefficient to obtain the weighted processing delay. The target camera and its viewing angle parameters are determined with the goal of minimizing the sum of the weighted stitching error and the weighted processing delay.
5. The method according to claim 4, characterized in that, The real-time status parameters of the camera include resolution; The acquisition of error coefficients includes: Determine the resolution requirements; Obtain the average resolution of the multiple cameras; The error coefficient is determined based on the average resolution of the multiple cameras and the resolution requirement.
6. The method according to claim 4, characterized in that, The real-time status parameters of the camera include bandwidth; obtaining the latency coefficient includes: Obtain the average bandwidth of the multiple cameras; The reciprocal of the average bandwidth is determined as the delay coefficient.
7. The method according to claim 4, characterized in that, The method further includes: Obtain the learning rate; Obtain the partial derivative of the splicing error with respect to the error coefficient; The adjustment amount is determined based on the learning rate and the partial derivative. The error coefficient is adjusted based on the adjustment amount to obtain the adjusted error coefficient; the adjusted error coefficient is used to redetermine the target camera and the viewing angle parameters.
8. The method according to claim 1, characterized in that, The real-time status parameters of the camera include computing power; satisfying the constraints of the real-time status parameters of multiple cameras includes: Obtain the resolution and frame rate of the video streams processed by each of the multiple cameras; Obtain the computing power conversion coefficient; The computing load of each of the multiple cameras is determined based on the computing power conversion coefficient, the resolution and frame rate of the video streams processed by each of the multiple cameras; When the computational load of each camera does not exceed its own computational capability, the computational capability constraints of multiple cameras are satisfied.
9. The method according to claim 1, characterized in that, The real-time status parameters of the camera include a bandwidth limit; constraints satisfying the real-time status parameters of multiple cameras include: Obtain the sum of the bandwidth occupied by the frame rates of one or more video streams output by the multiple cameras; If the sum of the bandwidths corresponding to each of the multiple cameras does not exceed the upper limit of the bandwidth of the camera, then the bandwidth constraint of the camera is determined to be satisfied.
10. A video splicing device, characterized in that, include: The acquisition module is used to acquire real-time status parameters from multiple cameras; The determination module is used to determine the target camera and the viewing angle parameters of the target camera under the constraints of the real-time status parameters of multiple cameras, with the goal of minimizing stitching error and processing delay; An optimization module is used to optimize the pose parameters of multiple target cameras based on the target camera and the viewpoint parameters of the target camera; The stitching module is used to stitch together videos captured by multiple target cameras based on optimized pose parameters.
11. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the video stitching method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the video stitching method as described in any one of claims 1 to 9.