A method for automatic alignment of multi-camera video frames

By combining scoring functions and temporal flow models with deep learning technology, the problem of video frame time synchronization in multi-camera systems is solved, efficient and accurate video frame alignment is achieved, which adapts to complex environments, reduces computational complexity, and expands the scope of application.

CN119484997BActive Publication Date: 2025-09-30UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411584188.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-09-30
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

The temporal synchronization problem of video frames in multi-camera systems leads to poor reconstruction accuracy and effect. Existing methods are costly, inflexible, computationally complex, and lack robustness, making it difficult to achieve real-time processing.

Method used

The similarity of video frames is measured by a scoring function, a temporal flow model is constructed, the time offset of the cameras is estimated using the time offset theorem, and the video frames are adjusted to achieve alignment. Deep learning technology is combined to improve the alignment accuracy and robustness.

Benefits of technology

It improves the accuracy and efficiency of multi-camera systems in 3D reconstruction, adapts to different perspectives, dynamic scenes and lighting conditions, reduces computational complexity, and achieves real-time video frame alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484997B_ABST
    Figure CN119484997B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatic alignment of multi-camera video frames, belonging to the field of computer vision technology. The method comprises: obtaining video frames captured by each camera and, based on a preset scoring function, measuring the similarity between the video frames captured by different cameras to find a standard frame in the video frames as an alignment reference frame; constructing a temporal flow model to model the changes of video frames over time; using the temporal flow model and based on the time offset theorem, estimating the time offset of each camera; and adjusting the video frames of each camera based on the time offset to align them with the standard frame. The present invention significantly improves the accuracy and efficiency of three-dimensional video reconstruction. It can adapt to different viewing angles, dynamic scenes, and changes in lighting conditions, achieving fast and accurate video frame alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a method for automatically aligning video frames of multiple cameras. Background Art

[0002] With the rapid development of information technology, multi-camera systems have been widely used in various fields. By capturing the same scene from different angles, multi-camera systems can provide more comprehensive and richer visual information. They are widely used in monitoring systems, autonomous driving, sports broadcasting, and other fields.

[0003] 4D video reconstruction involves recovering a 3D scene and its time-varying dynamic information from video sequences captured by multiple cameras. Unlike 3D image reconstruction, 4D video reconstruction requires not only recovering the scene's geometric structure but also capturing its dynamic changes. This technology has widespread applications in virtual reality, film special effects, motion analysis, and other fields. However, a major challenge facing multi-camera 4D video reconstruction is the temporal synchronization of video frames. Asynchronous shooting times between different cameras lead to temporal misalignment between video frames, severely impacting the accuracy and effectiveness of the reconstruction.

[0004] In existing multi-camera systems, the time synchronization problem of videos is mainly solved by the following methods: (1) Using dedicated hardware devices, such as synchronization triggers, to ensure that all cameras start shooting at the same time. Although this method can provide high synchronization accuracy, it is expensive and less flexible; (2) Using simple software algorithms to perform time correction on the captured video frames. Common methods include timestamp-based synchronization, image content-based synchronization, etc. However, due to the difference in perspectives of different cameras, it may increase the difficulty of aligning video frames. Especially in the case of large perspective differences, the accuracy and reliability of feature point matching will be affected. In dynamic scenes, the movement and changes of objects will also affect the detection and matching of feature points, increasing the complexity of alignment. For example, in a crowded scene, moving pedestrians and vehicles will cause significant differences between videos. In addition, the different lighting conditions of different cameras will lead to differences in brightness and contrast of video frames, affecting the alignment effect.

[0005] In summary, the existing multi-camera video frame alignment technology has the following deficiencies: (1) Perspective difference: The perspective difference of different cameras may make the alignment between video frames more difficult; (2) Dynamic scene: In a dynamic scene, the movement and change of objects will also affect the detection and matching of feature points, increasing the complexity of alignment; (3) Lighting conditions: Different lighting conditions of different cameras will lead to differences in brightness and contrast of video frames, affecting the alignment effect; (3) Computational complexity: Existing alignment methods usually have high computational complexity and are difficult to achieve real-time processing, which limits their performance in real-time applications; (4) Lack of robustness: Existing alignment methods have poor robustness when facing complex environments and dynamic scenes and are easily affected by external interference. Summary of the Invention

[0006] The present invention provides a method for automatic alignment of multi-camera video frames to solve the technical problems of low alignment accuracy, poor robustness and high computational complexity in the existing technology when dealing with problems such as differences in camera viewing angles, dynamic scenes, and lighting changes.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In one aspect, the present invention provides a method for automatically aligning video frames of multiple cameras, comprising:

[0009] Obtain the video frames shot by each camera, and measure the similarity between the video frames captured by different cameras based on a preset scoring function to find a standard frame in the video frames as an alignment reference frame;

[0010] Construct a temporal flow model to model the changes of video frames over time. The changes of video frames over time include changes in image content and image quality over time. The changes in image content over time include changes in object movement speed and direction. The changes in image quality over time include changes in light intensity and blur level.

[0011] Using the time flow model and based on the time offset theorem, the time offset of each camera is estimated;

[0012] Adjust the video frames of each camera to align with the standard frame based on the time offset.

[0013] Furthermore, the scoring function measures the similarity between video frames captured by different cameras through the feature point matching, color histogram similarity and depth information consistency between the video frames captured by different cameras.

[0014] Furthermore, the expression of the scoring function is:

[0015]

[0016] in, Represents the video frame captured by the jth camera at time t1 and the video frame captured by the kth camera at time t2 similarity between Indicates the feature point matching score; represents the color histogram similarity score; represents the depth information consistency score; w1, w2, and w3 are preset weight parameters, and the values ​​of each weight parameter are adjusted according to the actual application scenario.

[0017] Further, Among them, N matches express and The number of matching point pairs between feature points; N total express and The total number of feature points detected in ;

[0018] The value of The color histogram of The similarity between the color histograms of

[0019] The value of The depth map and The mean absolute error between the depth maps.

[0020] Furthermore, the method of measuring the similarity between video frames captured by different cameras based on a preset scoring function to find a standard frame in the video frames as an alignment reference frame includes:

[0021] Selecting the first n frames of a video frame captured by a camera as candidate standard frames; wherein n is a preset value;

[0022] Calculate the scoring function value between each video frame captured by the current camera and each candidate standard frame;

[0023] The candidate standard frame corresponding to the largest scoring function value is selected as the final standard frame.

[0024] Furthermore, the expression of the time flow model is:

[0025] I(t)=I c (t)+I q (t)

[0026] Where I(t) represents the video frame data at time t; I c(t) represents the part of the image content of I(t) that changes over time; I q (t) represents the part of the image quality of I(t) that changes with time.

[0027] Furthermore, I c (t) is a multi-order polynomial, expressed as:

[0028]

[0029] Where N is the polynomial order; a n are the coefficients of the polynomial;

[0030] I q (t) is a Fourier series, expressed as:

[0031]

[0032] Where L is the number of Fourier series terms; b l is the Fourier coefficient; ω l is the frequency; φ l It's the phase.

[0033] Furthermore, the estimating the time offset of each camera by using the time flow model based on the time offset theorem includes:

[0034] Set the polynomial order N and the number of Fourier series terms L;

[0035] Collect time series data of each camera around the standard frame;

[0036] Using the collected time series data, the least squares method is used to fit I c (t) and I q (t) expression;

[0037] Based on the fitting results, the time offset of each camera is estimated.

[0038] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0039] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the above method.

[0040] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0041] The technical solution provided by this invention significantly improves the accuracy and efficiency of 3D reconstruction in multi-camera systems by introducing efficient image processing and computer vision algorithms, combined with deep learning techniques. It can adapt to varying viewpoints, dynamic scenes, and changing lighting conditions, achieving fast and accurate video frame alignment. It improves the accuracy and reliability of multi-camera video frame alignment, especially in scenes with large viewpoint differences and dynamic scenes. It also enhances the system's robustness in scenes with varying lighting conditions, effectively handling variations in lighting conditions between different cameras. Furthermore, it reduces computational complexity and improves processing efficiency, enabling real-time alignment of multi-camera video frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 Schematic diagram of the principle of the multi-camera video frame automatic alignment method provided by an embodiment of the present invention;

[0044] Figure 2 1 is a schematic diagram of an execution flow of a multi-camera video frame automatic alignment method provided by an embodiment of the present invention;

[0045] Figure 3 This is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0047] First, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present concepts in a concrete manner. In addition, in the embodiments of the present invention, the meaning of "and / or" can be both or either of the two.

[0048] First embodiment

[0049] This embodiment provides a method for automatically aligning video frames of multiple cameras, and its implementation principle is as follows: Figure 1 As shown, the method can be implemented by an electronic device, which can be a terminal or a server.

[0050] Specifically, the execution process of this method is as follows Figure 2 As shown, the following steps are included:

[0051] S1, obtains the video frames captured by each camera, and measures the similarity between the video frames captured by different cameras based on a preset scoring function to find a standard frame in the video frames as an alignment reference frame;

[0052] It should be noted that, since the starting frames taken by each camera are not synchronized, and there are sampling noises or wrong frames, such as Figure 1 As shown, The maximum difference is n frames. The inputs from each camera are fed into the time alignment module, where a scoring function is set to measure the distance between images and find the relationship between them. The scoring function is constructed based on multiple features such as feature point matching, color histograms, and depth information. This scoring function searches through the first n frames to find a standard frame as a reference point, minimizing the number of frames discarded before the standard frame and minimizing the number of frames offset after the standard frame.

[0053] Specifically, the implementation process of the above S1 is as follows:

[0054] S11, Design of Scoring Function

[0055] In order to overcome the problems caused by asynchronous input and sampling noise, this embodiment designs a scoring function This function is used to measure the video frames captured by different cameras j and k at a certain time t. With reference frame The scoring function comprehensively considers the characteristics of feature point matching, color histogram similarity and depth information consistency, as follows:

[0056] Feature point matching: Use the Scale-Invariant Feature Transform (SIFT) or Speeded Up Robust Features (SURF) algorithm to detect key points in the video frames, and use the Random Sample Array Consensus Algorithm (RANSAC) to calculate the homography matrix H between the two video frames to evaluate the quality of feature point matching.

[0057] Color histogram similarity: Calculate the color histograms of two video frames and use Chi-Square Distance or other similarity metrics to evaluate the similarity of color distributions.

[0058] Depth information consistency: If the video frames contain depth information (e.g., data from a depth camera), the depth consistency can be evaluated by comparing the depth maps of two video frames.

[0059] S12, the form of the scoring function

[0060] Scoring function It can be formalized as:

[0061]

[0062] in:

[0063] is the feature point matching score.

[0064] is the color histogram similarity score.

[0065] is the depth information consistency score.

[0066] w1, w2, w3 are weight parameters that can be adjusted according to the actual application scenario. S13, feature point matching score S feat

[0067] Feature point matching score S feat The calculation method is as follows:

[0068] Detection in video frames using SIFT or SURF algorithms and The key points in the video are calculated using the RANSAC algorithm to calculate the homography matrix H between the two video frames. The number of matching point pairs N is calculated. matches .

[0069] Calculate the feature point matching score:

[0070]

[0071] Among them, N total is the total number of keypoints detected in the two video frames.

[0072] S14, color histogram similarity score S color

[0073] Color histogram similarity score S color The calculation method is as follows:

[0074] Video frame and Convert to grayscale image.

[0075] Calculate the color histogram H of two video frames j and H k .

[0076] Compute the similarity of color histograms using chi-square distance:

[0077]

[0078] Where B is the number of bins in the histogram.

[0079] S15, depth information consistency score S depth

[0080] Depth information consistency score S depth The calculation method is as follows:

[0081] Get video frame and The depth map D j and D k .

[0082] Calculate the mean absolute error (MAE) between two depth maps:

[0083]

[0084] Where N is the number of pixels in the depth map.

[0085] S16, Optimization of scoring function

[0086] To find the best standard frame We need to search the first n frames of each camera and select the frame with the highest total score as the standard frame of the standard camera (there is only one standard frame for the standard camera, and all other cameras are aligned to it. If there are multiple frames with the same highest score, the earliest frame is selected. If the time is the same, a random frame is selected). The specific steps are as follows:

[0087] Initialize the variable max_score = 0,

[0088] For each candidate frame

[0089] Calculating the scoring function

[0090] if Update max_score and best_frame.

[0091] Finally, the bestframe is selected as the standard frame.

[0092] Through the above scoring function and optimization algorithm, we can effectively find a standard frame in a standard camera, so that the frames before the standard frame are discarded at the least, and the frames after the standard frame are compensated at the least, thereby achieving accurate alignment of multi-camera video frames.

[0093] Furthermore, after finding the standard frame, considering that the video stream is continuous in time but discontinuous in sampling time, this embodiment proposes to generate curves for k cameras around the standard frame through fitting, thereby estimating the initial point of each camera relative to the standard frame. Curve fitting can more accurately determine the behavior pattern of each camera around the standard frame, improving alignment accuracy.

[0094] S2, builds a temporal flow model to model the changes of video frames over time;

[0095] It should be noted that after finding the standard frame, subsequent frame alignment must ensure that each camera's video frames are temporally consistent with the standard frame. To achieve this, this embodiment assumes that the time stream for each camera is identical and uses a temporal stream model to describe the temporal changes of video frames. These changes include changes in image content and image quality over time. Changes in image content over time include changes in object speed and direction; changes in image quality over time include changes in illumination intensity and blur. This model consists of two parts: a polynomial component and a Fourier series component.

[0096] Polynomial part: describes the changing trend of image content over time, such as the speed and direction of an object's movement.

[0097] Fourier series part: describes the changes in image quality over time, such as light intensity and blur.

[0098] Based on the above, the time flow model can be expressed by the following formula:

[0099] I(t)=I c (t)+I q (t)

[0100] in:

[0101] I(t) represents the video frame data at a certain time t.

[0102] I c (t) represents the part of the image content that changes over time.

[0103] I q (t) represents the part where the image quality changes over time.

[0104] Polynomials Part I c (t) can be expressed as an Nth-order polynomial:

[0105]

[0106] Among them, a n are the coefficients of the polynomial, which can be obtained by least squares fitting.

[0107] Fourier Series Part I q (t) can be expressed as an L-term Fourier series:

[0108]

[0109] Among them, b l is the Fourier coefficient, ω l is the frequency, φ l is the phase, and these parameters can also be obtained by least square fitting.

[0110] S3, using the time flow model and based on the time offset theorem, estimates the time offset of each camera;

[0111] Assume that there is a time offset τ, which allows the video frame data I(t) to be aligned with the time offset τ. Construct an objective function E(τ) to represent the error between the time-shifted video frame data and the reference standard frame. Using the above optimization method, minimize the objective function E(τ) and solve for the time offset τ. The objective function is defined as follows:

[0112]

[0113] Among them, I ref (t) is the data of the standard frame; I(t-τ) is the video frame data after time offset; t is the time variable, which represents the timestamp of the video frame.

[0114] It's important to note that the time shift theorem is a fundamental property of the Fourier transform, describing how a signal's time shift in the time domain affects its frequency domain representation. Suppose there's a signal f(t), whose Fourier transform is F(ω). If the signal is time-shifted (i.e., the new signal is f(t-τ)), then the Fourier transform of this new signal is:

[0115] F(ω)e -iωτ

[0116] Here, i is the imaginary unit, ω is the frequency component, and τ is the time offset.

[0117] When the signal moves forward (τ<0) or backward (τ>0) on the time axis, the corresponding Fourier transform spectrum does not change the amplitude of the signal, but introduces a phase factor e in the frequency domain. -iωτ The phase factor is proportional to the offset τ, and for each frequency component ω, the phase change caused by the offset is -ωτ.

[0118] S4, adjusts the video frame of each camera according to the time offset so that it is aligned with the standard frame.

[0119] Based on the above, the subsequent frame alignment steps are as follows:

[0120] 1. Initialization parameters: set the polynomial order N and the number of Fourier series terms L.

[0121] 2. Data Preparation: Collect time series data for each camera around the standard frame. Ensure the data's completeness and accuracy. This step provides a reliable foundation for subsequent model fitting and time offset estimation.

[0122] 3. Polynomial Fitting: Fitting Polynomials Using Least Squares Method Part I c (t).

[0123] 4. Fourier Series Fitting: Using Least Squares Method to Fit Fourier Series Part I q (t).

[0124] 5. Time offset estimation: Based on the fitting results, estimate the time offset τ of each camera k .

[0125] 6. Frame alignment: according to the time offset τ k , adjust the video frames of each camera to align them with the standard frame.

[0126] Through the above steps, we can achieve subsequent alignment of multi-camera video frames, ensuring that the video frames of each camera are consistent with the standard frame in time, thereby improving the accuracy and robustness of video frame alignment.

[0127] In summary, this embodiment provides a method for automatic alignment of multi-camera video frames. By introducing efficient image processing and computer vision algorithms, combined with deep learning techniques, it significantly improves the accuracy and efficiency of 3D reconstruction in multi-camera systems. It can adapt to varying viewpoints, dynamic scenes, and lighting conditions, achieving fast and accurate video frame alignment. The following are the main benefits of this solution:

[0128] Improve the accuracy and reliability of 3D reconstruction

[0129] 3D reconstruction relies on the accurate alignment of video frames acquired from multiple viewpoints. This invention ensures that video frames captured by different cameras can be accurately aligned by combining multiple features such as feature point matching, color histogram similarity, and depth information consistency. Specifically:

[0130] Feature point matching: Detecting key points in video frames using the Scale-Invariant Feature Transform (SIFT) or Speeded Up Robust Features (SURF) algorithm and calculating the homography matrix using the Random Sample Array Consensus Algorithm (RANSAC) improves the accuracy and reliability of feature point matching. This is crucial for geometric alignment in 3D reconstruction, ensuring that video frames from different perspectives are correctly aligned, thereby improving the accuracy of the 3D model.

[0131] Scoring function: This function comprehensively considers multiple features, including feature point matching, color histogram similarity, and depth information consistency. It uses this scoring function to find the optimal standard frame, ensuring accurate alignment of video frames from different perspectives. This step is particularly important in 3D reconstruction, as any alignment error directly affects the quality of the final model.

[0132] Color Histogram Similarity: By calculating the similarity of color histograms, this method effectively addresses differences in lighting conditions between different cameras and improves alignment. In 3D reconstruction, variations in lighting conditions can lead to significant differences between video frames, affecting reconstruction accuracy. This method ensures accurate alignment of video frames under varying lighting conditions by calculating color histogram similarity.

[0133] Depth Information Consistency: Utilizing the consistency of depth information, we further improve alignment accuracy under varying lighting conditions. Depth information provides important geometric constraints in 3D reconstruction. By checking the consistency of depth information, we ensure that video frames are correctly aligned in 3D space, thereby improving the geometric accuracy of the reconstructed model.

[0134] Improving the robustness and adaptability of 3D reconstruction

[0135] 3D reconstruction requires not only high precision but also good robustness and adaptability to cope with various complex environments and dynamic scenes. This invention significantly improves the robustness and adaptability of the system through the following methods:

[0136] Temporal Flow Model: This model uses a polynomial and Fourier series component to model the temporal changes of video frames, effectively handling the motion and changes of objects in dynamic scenes. The polynomial component describes the temporal changes in image content, such as the speed and direction of an object's movement; the Fourier series component describes the temporal changes in image quality, such as lighting intensity and blur. This temporal flow model enables the system to flexibly adapt to different dynamic scenes, improving the robustness and adaptability of 3D reconstruction.

[0137] Temporal offset estimation: Utilizing the temporal offset theorem, the temporal offset of each camera is accurately estimated, ensuring accurate alignment of video frames in dynamic scenes. Temporal offset estimation is implemented using the least squares method, enabling fast time offset estimation and ensuring the system's real-time processing capabilities. In 3D reconstruction, accurate temporal offset estimation is particularly important for dynamic scenes, significantly improving the stability and reliability of the reconstructed model.

[0138] Multiple features of the scoring function: The scoring function comprehensively considers multiple features, which can reduce the impact of external interference in complex environments and improve the robustness of the system. In 3D reconstruction, external interference (such as lighting changes, occlusions, etc.) can cause alignment errors between video frames, affecting reconstruction accuracy. By using multiple features of the scoring function, the present invention ensures that the most appropriate alignment reference frame can be found under various conditions, ensuring the robustness and stability of the system in complex environments.

[0139] Reduce computational complexity and improve processing efficiency

[0140] 3D reconstruction usually involves a lot of calculations, which places high demands on real-time processing capabilities. The present invention significantly reduces the computational complexity and improves processing efficiency through the following methods.

[0141] Efficiency of polynomials and Fourier series: Polynomial and Fourier series fitting is implemented using the least squares method, which has low computational complexity and can quickly find the optimal solution. In 3D reconstruction, efficient computational methods can significantly shorten reconstruction time and improve the system's real-time processing capabilities.

[0142] Highly efficient scoring functions: The scoring function is implemented using efficient algorithms such as feature point matching, color histogram similarity, and depth information consistency, enabling quick scoring calculations. In 3D reconstruction, fast scoring ensures efficient and real-time processing of large numbers of video frames.

[0143] Fast time offset estimation: Time offset estimation is performed using the least squares method, enabling fast time offset estimation, ensuring the system's real-time processing capabilities. In 3D reconstruction, fast time offset estimation can significantly improve the system's real-time processing capabilities and reconstruction efficiency.

[0144] Expanding the application scope of 3D reconstruction

[0145] This invention significantly expands the application scope of multi-camera systems in 3D reconstruction by improving the accuracy and efficiency of video frame alignment. Specifically:

[0146] Surveillance systems: By accurately aligning video frames, surveillance systems can obtain more comprehensive information from multiple perspectives, improving monitoring effectiveness. In 3D reconstruction, multi-view monitoring can provide richer geometric information, improving the accuracy and integrity of 3D models.

[0147] Autonomous driving: By accurately aligning video frames, autonomous driving systems can better fuse multi-sensor data, improving perception accuracy and decision-making capabilities. In 3D reconstruction, the fusion of multi-sensor data can provide more detailed environmental information, improving the accuracy and reliability of 3D maps.

[0148] Sports broadcasting: By accurately aligning video frames, sports broadcasts can provide richer visual information from multiple perspectives, enhancing the viewer experience. In 3D reconstruction, multi-perspective broadcasts can provide more comprehensive game information, improving the accuracy and visual appeal of the 3D reconstructed model.

[0149] Virtual Reality and Movie Special Effects: Accurate alignment of video frames enables more precise 3D reconstruction in virtual reality and movie special effects, improving visual quality. In 3D reconstruction, alignment of multi-view video frames provides richer geometric and texture information, enhancing the realism and visual quality of 3D models.

[0150] In summary, this invention significantly improves the accuracy and efficiency of 3D reconstruction using multi-camera systems by integrating multiple advanced image processing and computer vision technologies. This improves the system's robustness and adaptability, reduces computational complexity, and expands the application scope of 3D reconstruction. These improvements not only enhance the quality and efficiency of 3D reconstruction but also provide strong support for the application of multi-camera systems in various fields.

[0151] Second embodiment

[0152] This embodiment provides an electronic device, such as Figure 3 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. In addition, the electronic device may also include a transceiver; the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0153] Next, combine Figure 3 A detailed introduction to the various components of the electronic device is given below:

[0154] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0155] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 3 The CPU0 and CPU1 shown in FIG are, of course, only exemplary.

[0156] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0157] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 3 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0158] The transceiver may include a receiver and a transmitter ( Figure 3 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 3 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0159] In addition, it should be noted that Figure 3 The structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown, or may combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment can refer to the technical effects described in the first embodiment above, and therefore will not be repeated here.

[0160] Third embodiment

[0161] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device. The instructions stored therein can be loaded by a processor in a terminal to execute the method described above.

[0162] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention may take the form of a fully or partially hardware embodiment, a fully or partially software embodiment, or an embodiment combining software and hardware aspects. Furthermore, when implemented using software, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.

[0163] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0164] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0165] It should also be noted that, in this document, relational terms such as first and second are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or terminal device comprising the element. In addition, the term "and / or" is merely a description of an associative relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the presence of A alone, the presence of A and B simultaneously, or the presence of B alone, where A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "more" means two or more. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0166] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0167] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0168] In the several embodiments provided herein, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is merely a logical functional division. In actual implementation, other division methods may be used, such as multiple units or components being combined or integrated into another device, or some features being ignored or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interface, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0169] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0170] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention. It should be noted that, although preferred embodiments of the present invention have been described, those skilled in the art, once understanding the basic inventive concepts of the present invention, may make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as covering the preferred embodiments and all variations and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A method for automatic alignment of multi-camera video frames, characterized in that: include: Obtain the video frames shot by each camera, and measure the similarity between the video frames captured by different cameras based on a preset scoring function to find a standard frame in the video frames as an alignment reference frame; Construct a temporal flow model to model the changes of video frames over time. The changes of video frames over time include changes in image content and image quality over time. The changes in image content over time include changes in object movement speed and direction. The changes in image quality over time include changes in light intensity and blur level. Using the time flow model and based on the time offset theorem, the time offset of each camera is estimated; Adjust the video frames of each camera to align with the standard frame according to the time offset; The scoring function measures the similarity between video frames captured by different cameras through feature point matching, color histogram similarity and depth information consistency between video frames captured by different cameras; The expression of the time flow model is: I(t)=I c (t)+I q (t) Where I(t) represents the video frame data at time t; I c (t) represents the part of the image content of I(t) that changes over time; I q (t) represents the part of the image quality of I(t) that changes with time; I c (t) is a multi-order polynomial, expressed as: Where N is the polynomial order; a n are the coefficients of the polynomial; I q (t) is a Fourier series, expressed as: Where L is the number of Fourier series terms; b l is the Fourier coefficient; ω l is the frequency; φ l It is the phase; The method of estimating the time offset of each camera by using the time flow model and based on the time offset theorem includes: Set the polynomial order N and the number of Fourier series terms L; Collect time series data of each camera around the standard frame; Using the collected time series data, the least squares method is used to fit I c (t) and I q (t) expression; Based on the fitting results, the time offset of each camera is estimated.

2. The method for automatically aligning multiple camera video frames according to claim 1, wherein: The expression of the scoring function is: in, Represents the video frame captured by the jth camera at time t1 and the video frame captured by the kth camera at time t2 The similarity between Indicates the matching score of feature points; represents the color histogram similarity score; represents the depth information consistency score; w1, w2, and w3 are preset weight parameters, and the values ​​of each weight parameter are adjusted according to the actual application scenario.

3. The method for automatically aligning multiple camera video frames according to claim 2, wherein: Among them, N matches express and The number of matching point pairs between feature points; N total express and The total number of feature points detected in ; The value of The color histogram of The similarity between the color histograms of The value of The depth map and The mean absolute error between the depth maps.

4. The method for automatically aligning multiple camera video frames according to claim 1, wherein: The method of measuring the similarity between video frames captured by different cameras based on a preset scoring function to find a standard frame in the video frames as an alignment reference frame includes: Selecting the first n frames of a video frame captured by a camera as candidate standard frames; wherein n is a preset value; Calculate the scoring function value between each video frame captured by the current camera and each candidate standard frame; The candidate standard frame corresponding to the largest scoring function value is selected as the final standard frame.