Panoramic video image processing method and apparatus, electronic device, and medium

By acquiring the target prediction stitching parameters and utilizing the changes in stitching parameters in historical images, the stitching parameters for future moments can be predicted, thus solving the problems of image latency and stitching errors in panoramic video live streaming and achieving higher quality image stitching.

CN115841421BActive Publication Date: 2026-06-19RUNBO PANORAMIC CULTURE & TOURISM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RUNBO PANORAMIC CULTURE & TOURISM TECH CO LTD
Filing Date
2022-11-22
Publication Date
2026-06-19

Smart Images

  • Figure CN115841421B_ABST
    Figure CN115841421B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of image processing technology, specifically to a panoramic video image processing method, apparatus, electronic device, and medium. The method includes: acquiring target prediction stitching parameters, wherein the target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images, wherein the historical images are images captured by N cameras positioned in different directions at a first time moment and at a time moment adjacent to the first time moment; acquiring N target image frames, wherein the N target image frames are images captured by the N cameras at the target time moment; and stitching the N target image frames according to the target prediction stitching parameters to obtain a first panoramic video image at the target time moment. Multiple images acquired at future times can be stitched using accurately predicted stitching parameters, improving the quality of the stitched image while ensuring reduced latency in the panoramic video image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, specifically to a panoramic video image processing method, apparatus, electronic device, and medium. Background Technology

[0002] With the continuous development of panoramic camera technology, the application of panoramic cameras is becoming increasingly widespread. For example, deploying panoramic cameras in the tourism and security sectors allows for real-time transmission of panoramic video streams to users, enabling live panoramic video streaming.

[0003] In related technologies, panoramic cameras are equipped with four cameras, each with a corresponding sensor that acquires image frames. The panoramic camera transmits the acquired image frames to a cloud server via a wireless communication module. The cloud server then stitches the acquired image frames together and sends them to the user device via the communication module. Thus, in live streaming scenarios, the panoramic camera can continuously acquire consecutive image frames, stitch them together via the cloud server, and then send them to the user device to achieve real-time output of panoramic video.

[0004] The process of stitching panoramic videos requires image registration and stitching. However, continuous image registration calculations in live streaming systems introduce significant latency. Currently, semi-statically updated stitching parameters can reduce this latency by directly using the stitching parameters from the previous frame to the current frame. However, when image content changes, especially significant changes in image boundary areas, the old stitching parameters may become inapplicable. Using the old parameters will result in serious stitching errors, leading to poor image quality after stitching.

[0005] Therefore, while ensuring reduced latency in panoramic video, improving the quality of the stitched image becomes an urgent problem to be solved. Summary of the Invention

[0006] To address the problems in the related technologies, this disclosure provides a panoramic video image processing method, apparatus, electronic device, and medium.

[0007] In a first aspect, this disclosure provides a panoramic video image processing method.

[0008] Specifically, the method includes:

[0009] Obtain target prediction stitching parameters, which are stitching parameters corresponding to the target time after the first time. The target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images. The historical images are images collected by N cameras set in different directions of the panoramic camera at the first time and the previous time adjacent to the first time.

[0010] N target image frames are acquired, wherein the N target image frames are images captured by the N cameras at the target time respectively;

[0011] Based on the target prediction stitching parameters, the N target image frames are stitched together to obtain the first panoramic video image at the target time.

[0012] In one embodiment of this disclosure, obtaining the target prediction stitching parameters includes:

[0013] N first edge image frames and N second edge image frames are acquired. Any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region. Any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region. The N first edge image frames are images captured by the N cameras at the first moment. Each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames.

[0014] For each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time.

[0015] Based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the splicing parameter change prediction result information is obtained based on the pre-trained splicing parameter change prediction network.

[0016] Based on the prediction results of the splicing parameter changes, the target predicted splicing parameters are obtained.

[0017] In one embodiment of this disclosure, obtaining the target predicted stitching parameters based on the prediction result information of the stitching parameter changes includes:

[0018] If the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

[0019] In one embodiment of this disclosure, obtaining the target predicted stitching parameters based on the prediction result information of the stitching parameter changes includes:

[0020] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

[0021] In one embodiment of this disclosure, when the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, including:

[0022] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network.

[0023] The N predicted image frames are processed using image registration and image stitching algorithms to obtain a predicted stitched image.

[0024] The target predicted stitching parameters are obtained based on the predicted stitched image.

[0025] In one embodiment of this disclosure, before acquiring N first edge image frames and N second edge image frames, the method further includes:

[0026] Obtain N first image frames and N second image frames, wherein each of the N first image frames is a full-size image captured by one of the N cameras at a first moment, and each of the N second image frames is a full-size image captured by one of the N cameras at a second moment, wherein the second moment is the moment preceding the first moment;

[0027] The acquisition of N first edge image frames and N second edge image frames includes:

[0028] Based on the first overlapping image region, the N first image frames are cropped respectively to obtain the N first edge image frames;

[0029] Based on the second overlapping image region, the N second image frames are cropped respectively to obtain the N second edge image frames.

[0030] Secondly, this disclosure provides a panoramic video image processing device.

[0031] Specifically, the device includes:

[0032] The first acquisition module is configured to acquire target prediction stitching parameters, wherein the target prediction stitching parameters are stitching parameters corresponding to the target time after the first time. The target prediction stitching parameters are determined by the change of stitching parameters obtained from historical images. The historical images are images acquired by N cameras set in different directions of the panoramic camera at the first time and the previous time adjacent to the first time.

[0033] The second acquisition module is configured to acquire N target image frames, wherein the N target image frames are images captured by the N cameras at the target time respectively;

[0034] The processing module is configured to perform image stitching on the N target image frames according to the target prediction stitching parameters to obtain the first panoramic video image at the target time.

[0035] In one embodiment of this disclosure, the first acquisition module is specifically configured as follows:

[0036] N first edge image frames and N second edge image frames are acquired. Any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region. Any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region. The N first edge image frames are images captured by the N cameras at the first moment. Each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames.

[0037] For each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time.

[0038] Based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the splicing parameter change prediction result information is obtained based on the pre-trained splicing parameter change prediction network.

[0039] Based on the prediction results of the splicing parameter changes, the target predicted splicing parameters are obtained.

[0040] In one embodiment of this disclosure, the first acquisition module is specifically configured as follows:

[0041] If the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

[0042] In one embodiment of this disclosure, the first acquisition module is specifically configured as follows:

[0043] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

[0044] In one embodiment of this disclosure, the first acquisition module is specifically configured as follows:

[0045] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network.

[0046] The N predicted image frames are processed using image registration and image stitching algorithms to obtain a predicted stitched image.

[0047] The target predicted stitching parameters are obtained based on the predicted stitched image.

[0048] In one embodiment of this disclosure, the apparatus further includes:

[0049] The third acquisition module is configured to acquire N first image frames and N second image frames, wherein each of the N first image frames is a full-size image captured by one of the N cameras at a first moment, and each of the N second image frames is a full-size image captured by one of the N cameras at a second moment, wherein the second moment is the moment preceding the first moment.

[0050] The first acquisition module is specifically configured to crop the N first image frames according to the first overlapping image region to obtain the N first edge image frames; and to crop the N second image frames according to the second overlapping image region to obtain the N second edge image frames.

[0051] Thirdly, embodiments of this disclosure provide an electronic device including a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method as described in any one of the first aspects and possible implementations of the first aspect.

[0052] Fourthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method as described in any one of the first aspect and possible implementations of the first aspect.

[0053] According to the technical solution provided in this disclosure, target prediction stitching parameters are obtained. These target prediction stitching parameters are stitching parameters corresponding to a target time after a first time. The target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images. The historical images are images captured by N cameras positioned in different directions at the first time and at a time immediately preceding the first time. N target image frames are obtained, each captured by the N cameras at the target time. Based on the target prediction stitching parameters, the N target image frames are stitched together to obtain a first panoramic video image at the target time. This solution allows for accurate prediction of future stitching parameters based on changes in historical image stitching parameters, enabling the stitching of multiple images acquired at future times using the accurately predicted parameters. Thus, compared to related technologies, the quality of the stitched image is improved while ensuring reduced latency in the panoramic video image.

[0054] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0055] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:

[0056] Figure 1 A flowchart illustrating a panoramic video image processing method according to an embodiment of the present disclosure is shown;

[0057] Figure 2 This diagram illustrates the process of obtaining an optical flow estimation frame according to an embodiment of the present disclosure;

[0058] Figure 3 A schematic diagram illustrating the process of obtaining splicing parameter change prediction result information according to an embodiment of the present disclosure;

[0059] Figure 4 This diagram illustrates the process of obtaining target prediction stitching parameters according to an embodiment of the present disclosure;

[0060] Figure 5 A structural block diagram of a panoramic video image processing apparatus according to an embodiment of the present disclosure is shown;

[0061] Figure 6 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown;

[0062] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown. Detailed Implementation

[0063] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of exemplary embodiments have been omitted from the drawings.

[0064] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.

[0065] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0066] In this disclosure, any operation involving the acquisition of user information or user data, or the display of user information or user data to others, is an operation authorized or confirmed by the user, or actively selected by the user.

[0067] As mentioned above, with the continuous development of panoramic camera technology, the application of panoramic cameras is becoming increasingly widespread. For example, deploying panoramic cameras in the tourism and security sectors allows for real-time transmission of panoramic video streams to users, enabling live panoramic video streaming.

[0068] In related technologies, panoramic cameras are equipped with four cameras, each with a corresponding sensor that acquires image frames. The panoramic camera transmits the acquired image frames to a cloud server via a wireless communication module. The cloud server then stitches the acquired image frames together and sends them to the user device via the communication module. Thus, in live streaming scenarios, the panoramic camera can continuously acquire consecutive image frames, stitch them together via the cloud server, and then send them to the user device to achieve real-time output of panoramic video.

[0069] The process of stitching panoramic videos requires image registration and stitching. However, continuous image registration calculations in live streaming systems introduce significant latency. Currently, semi-statically updated stitching parameters can reduce this latency by directly using the stitching parameters from the previous frame to the current frame. However, when image content changes, especially significant changes in image boundary areas, the old stitching parameters may become inapplicable. Using the old parameters will result in serious stitching errors, leading to poor image quality after stitching.

[0070] Therefore, while ensuring reduced latency in panoramic video, improving the quality of the stitched image becomes an urgent problem to be solved.

[0071] Based on the aforementioned technical deficiencies, this disclosure provides a panoramic video image processing method. This method can obtain target prediction stitching parameters, which are stitching parameters corresponding to a target time after a first time moment. These parameters are determined by the changes in stitching parameters obtained from historical images. The historical images are images captured by N cameras positioned in different directions at the first time moment and at a time moment immediately preceding the first time moment. The method also acquires N target image frames, which are images captured by the N cameras at the target time moment. Finally, based on the target prediction stitching parameters, the N target image frames are stitched together to obtain a first panoramic video image at the target time moment. This approach allows for accurate prediction of future stitching parameters based on changes in historical image stitching parameters, enabling the stitching of multiple images acquired at future times using the accurately predicted parameters. Thus, compared to related technologies, this method improves the quality of the stitched image while reducing the latency of the panoramic video image.

[0072] Figure 1 A flowchart illustrating a panoramic video image processing method according to an embodiment of the present disclosure is shown. Figure 1 As shown, the panoramic video image processing method includes the following steps S101-S103:

[0073] In S101, obtain the target prediction stitching parameters.

[0074] Wherein, the target prediction stitching parameter is the stitching parameter corresponding to the target time after the first time. The target prediction stitching parameter is determined by the change of stitching parameters obtained from historical images. The historical images are images collected by N cameras set in different directions of the panoramic camera at the first time and the previous time adjacent to the first time.

[0075] In S102, N target image frames are acquired.

[0076] Wherein, the N target image frames are images captured by the N cameras at the target time respectively;

[0077] In S103, the N target image frames are stitched together according to the target prediction stitching parameters to obtain the first panoramic video image at the target time.

[0078] In one embodiment of this disclosure, the panoramic video image processing method can be applied to panoramic cameras that perform panoramic video image processing, or cloud servers corresponding to panoramic cameras, etc.

[0079] In one embodiment of this disclosure, the panoramic camera includes a bottom pillar and a top panoramic camera. The bottom pillar houses a battery module, a control module, and a communication module; the top pillar includes multiple lens groups and corresponding image sensors and image signal processors (ISPs). The control module controls the camera's startup, shooting, and image processing. Of course, the control module can control each lens group and its corresponding image sensor and image signal processor individually or simultaneously.

[0080] In one embodiment of this disclosure, the panoramic camera involved in this embodiment can be understood as a 360-degree panoramic camera.

[0081] In one embodiment of this disclosure, the target time can be the next time adjacent to the first time, or a time that is separated from the first time by a preset time interval.

[0082] Furthermore, since the target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images, the stitching parameters corresponding to the image acquired by the panoramic camera at a certain time in the future can be predicted using historical images.

[0083] It should be noted that, in the disclosed embodiments, by setting the target time and the time difference between the next time adjacent to the first time, it is ensured that after acquiring N target image frames, the corresponding stitching parameters have been acquired in advance.

[0084] Furthermore, the target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images, which can be referred to the specific description in the following embodiments, and will not be repeated here.

[0085] In this embodiment of the disclosure, the acquisition of N target image frames can be implemented in the following two ways:

[0086] (Scenario 1) Images are captured at the target time by N cameras set in different directions using a panoramic camera, and the captured images are obtained by the image sensor corresponding to each of the N cameras, that is, N target image frames are obtained.

[0087] It should be noted that the above situation applies to the panoramic video image processing method provided in the embodiments of this disclosure when applied to a panoramic camera.

[0088] (Scenario 2) Receive N target image frames sent by the panoramic camera.

[0089] It should be noted that the above situation applies to the panoramic video image processing method provided in this embodiment of the invention when applied to the cloud server corresponding to the panoramic camera.

[0090] In one embodiment of this disclosure, the number of target prediction stitching parameters can be one or more. When the number of target prediction stitching parameters is one, the same target prediction stitching parameter is used to stitch the N target image frames together; when the number of target prediction stitching parameters is multiple, multiple stitching parameters are used to stitch corresponding image frames among the N target image frames together.

[0091] In one embodiment of this disclosure, the N target image frames acquired by different image sensors have overlapping regions. Image stitching refers to fusing the overlapping regions of the N target image frames to create a seamless panoramic image, i.e., obtaining the first panoramic video image.

[0092] In one embodiment of this disclosure, after S103, the panoramic video image processing method provided in this embodiment may further include: sending the first panoramic video image to a cloud server of the panoramic camera. Thus, the first panoramic video image can be sent to a user device via the cloud server to achieve panoramic video live streaming.

[0093] This disclosure provides a panoramic video image processing method. The method involves obtaining target prediction stitching parameters, which are stitching parameters corresponding to a target time after a first time point. These parameters are determined by observing changes in stitching parameters obtained from historical images. The historical images are images captured by N cameras positioned in different directions at the first time point and at a time point adjacent to the first time point. N target image frames are acquired, each captured by one of the N cameras at the target time point. Based on the target prediction stitching parameters, the N target image frames are stitched together to obtain a first panoramic video image at the target time point. This method accurately predicts future stitching parameters based on changes in historical image stitching parameters, allowing the stitching of multiple images acquired at future times to be stitched together using the accurately predicted parameters. Therefore, compared to related technologies, this method improves the quality of the stitched image while reducing the latency of the panoramic video image.

[0094] Optionally, in one possible implementation of this disclosure, S101 may specifically include the following in S101a to S101d:

[0095] In S101a, N first edge image frames and N second edge image frames are acquired.

[0096] Wherein, any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region, and any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region, the N first edge image frames are images captured by the N cameras at the first moment, and each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames.

[0097] In S101b, for each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time.

[0098] In S101c, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the stitching parameter change prediction result information is obtained based on the pre-trained stitching parameter change prediction network.

[0099] In S101d, the target predicted splicing parameters are obtained based on the splicing parameter change prediction result information.

[0100] In one embodiment of this disclosure, the optical flow estimation network can be understood as a depth optical flow estimation network. For example, the optical flow estimation network is FlowNet2.0.

[0101] In one embodiment of this disclosure, the stitching parameter change prediction network is used to indicate the changes in stitching parameters obtained based on the acquired historical images and the optical flow estimation frames corresponding to the historical images.

[0102] In one embodiment of this disclosure, the splicing parameter change prediction network can be understood as a U-Net.

[0103] In one embodiment of this disclosure, historical images acquired at a first time and a second time are used as input, and optical flow estimation information corresponding to the historical images is obtained according to an optical flow estimation network, forming the input of a training network, denoted as x. Further, the stitching parameters at the target time and the stitching parameters at the first time are calculated. If the stitching parameters are the same, an output data y = 1 is recorded; if the stitching parameters are different, an output data y = 0 is recorded. This constructs a set of training data (x, y). Thus, a large amount of training data is obtained using existing processing modules to train the stitching parameter change prediction network.

[0104] In one embodiment of this disclosure, in one possible scenario, the splicing parameter change prediction result information indicates that the splicing parameters have not changed; in another possible scenario, the splicing parameter change prediction result information indicates that the splicing parameters have changed.

[0105] It is understandable that the target prediction splicing parameters obtained will be different for the two possible scenarios described above. Therefore, the corresponding target prediction splicing parameters can be accurately obtained by predicting the changes in splicing parameters based on the changes in splicing parameters predicted by the prediction network.

[0106] In one embodiment of this disclosure, when the pixel value of the optical flow estimation frame is 0, which means that the image remains relatively still, the stitching parameter change prediction result information indicates that the stitching parameters have not changed.

[0107] In one embodiment of this disclosure, the splicing parameter change prediction network can be understood as a network used to predict whether the splicing parameters need to be rewritten. That is, the splicing parameter change prediction network acts as a classifier, and different values ​​of the splicing parameter change prediction result information obtained by the splicing parameter change prediction network indicate whether the splicing parameters have changed.

[0108] For example, suppose the obtained prediction result information of the splicing parameter change is y1, and different values ​​of y1 indicate whether the splicing parameters have changed. When y1 = 0, it means that the splicing parameters have changed.

[0109] Furthermore, if a value y is set for each pixel in the image... ab Therefore, when the obtained splicing parameter change prediction result information is y ab At that time, y ab The stitching parameters used to indicate whether the (a, b)th pixel has changed; if a value y is set for each row of pixels in the image. a Therefore, when the obtained splicing parameter change prediction result information is y a At that time, y a This is used to indicate whether the stitching parameters of the pixels in row a have changed.

[0110] In this embodiment, N first edge image frames and N second edge image frames can be acquired. For each first edge image frame, optical flow estimation is performed on the first edge image frame and the corresponding second edge image frame based on an optical flow estimation network to obtain N optical flow estimation frames at the target time. Based on the first edge image frame, the corresponding second edge image frame, and the corresponding optical flow estimation frames, a pre-trained stitching parameter change prediction network is used to obtain stitching parameter change prediction result information. And based on the stitching parameter change prediction result information, the target predicted stitching parameters are obtained. Thus, the corresponding target predicted stitching parameters can be accurately obtained based on the stitching parameter change predicted by the stitching parameter change prediction network.

[0111] Optionally, in one possible implementation of this disclosure, before S101a above, the panoramic video image processing method provided in this embodiment may further include S104; correspondingly, S101a above can be implemented by the following steps (1) and (2):

[0112] In S104, N first image frames and N second image frames are acquired. Each of the N first image frames is a full-size image captured by one of the N cameras at the first moment. Each of the N second image frames is a full-size image captured by one of the N cameras at the second moment. The second moment is the moment preceding the first moment.

[0113] Step (1) Based on the first overlapping image region, crop the N first image frames respectively to obtain the N first edge image frames;

[0114] Step (2) Based on the second overlapping image region, the N second image frames are cropped respectively to obtain the N second edge image frames.

[0115] In this embodiment of the disclosure, the acquisition of N first image frames and N second image frames can be implemented in the following two ways:

[0116] (Scenario 1) The panoramic camera sets up N cameras in different directions to capture images at the first and second moments respectively, and the captured images are obtained through the image sensor corresponding to each of the N cameras, that is, N first image frames and N second image frames are obtained.

[0117] It should be noted that the above situation applies to the panoramic video image processing method provided in the embodiments of this disclosure when applied to a panoramic camera.

[0118] (Case 2) Receive N first image frames and N second image frames sent by the panoramic camera.

[0119] It should be noted that the above situation applies to the panoramic video image processing method provided in this embodiment of the invention when applied to the cloud server corresponding to the panoramic camera.

[0120] It should be noted that since image stitching is based on image registration operations on overlapping areas, calculating the target stitching parameters does not require prediction of the complete image frame. Therefore, it is necessary to crop the images captured by the N cameras respectively.

[0121] In one embodiment of this disclosure, the first overlapping image region is the overlapping image region of any two first image frames captured by cameras at adjacent positions among the N first image frames.

[0122] Furthermore, the first overlapping image region may include one or more overlapping image regions. When the first overlapping image region includes multiple overlapping image regions, the size and content of these multiple overlapping image regions may be the same or different. This is determined specifically based on the actual usage, and the embodiments disclosed herein do not impose limitations on this.

[0123] The description of the second overlapping image region can be referred to the relevant description of the first overlapping image region in the above embodiments, and will not be repeated in this disclosure.

[0124] In this embodiment, N first image frames and N second image frames can be acquired. Based on the first overlapping image region, the N first image frames are cropped to obtain the N first edge image frames. Similarly, based on the second overlapping image region, the N second image frames are cropped to obtain the N second edge image frames. Thus, images acquired by N cameras at different times can be cropped based on the overlapping image region to obtain cropped images that facilitate the prediction of stitching parameters at a future time.

[0125] Example 1, taking N=2 and the target time as TP as an example. For example... Figure 2 The diagram illustrates the process of acquiring optical flow estimation frames. At the first time T1, image frame A1 is acquired through sensor 1 and image frame B1 is acquired through sensor 2, resulting in two first image frames. At the second time T0, image frame AN ​​is acquired through sensor 1 and image frame BN is acquired through sensor 2, resulting in two second image frames. Based on the first overlapping image region of A1 and B1, A1 and B1 are cropped to obtain AC-1 and BC-1, resulting in two first edge image frames. Based on the second overlapping image region of AN and BN, AN and BN are cropped to obtain AC-N and BC-N, resulting in two second edge image frames.

[0126] Furthermore, the optical flow estimation network FlowNet2.0 is used to perform optical flow estimation on AC-1, AC-N, BC-1 and BC-N to obtain the optical flow estimation frames AP and BP at TP time.

[0127] Furthermore, such as Figure 3 The diagram illustrates the process of obtaining the prediction results of stitching parameter changes. Two image frames AC-1, AC-N and optical flow estimation frame AP from sensor 1, and two image frames BC-1, BC-N and optical flow estimation frame BP from sensor 2 are organized into an input image matrix and input to a pre-trained stitching parameter change prediction network U-Net to obtain the prediction results of the stitching parameter changes.

[0128] Thus, the target predicted splicing parameters can be obtained based on the prediction results of the splicing parameter changes.

[0129] Furthermore, in this embodiment of the disclosure, the above-mentioned S101d can be implemented in two possible ways, as follows:

[0130] One possible implementation:

[0131] In S101d1, if the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

[0132] In one embodiment of this disclosure, images can be captured by N cameras of a panoramic camera at a first moment to obtain N image frames; stitching parameters are determined by the changes in stitching parameters obtained from historical images acquired before the first moment; and the N image frames are stitched together according to the stitching parameters to obtain the second panoramic video image.

[0133] It is understandable that determining the stitching parameters corresponding to the second panoramic video image at the first moment as the target prediction stitching parameters means that the image at the target moment and the image at the first moment remain relatively still.

[0134] In this implementation, if the stitching parameter change prediction result indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters. That is, by directly using the stitching parameters of the previous frame image to stitch together multiple images acquired at future moments, the stitching delay of panoramic video can be significantly reduced, bringing a better experience to panoramic video live streaming.

[0135] Another possible implementation:

[0136] In S101d2, when the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

[0137] In one embodiment of this disclosure, the prediction result information of the splicing parameter change indicates that the splicing parameters have changed, which may specifically include the following possible scenarios:

[0138] (1) The stitching parameters of the (a, b)th pixel change.

[0139] (2) The splicing parameters of the pixels in row a have changed.

[0140] In this embodiment, when the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame. This reduces the latency of panoramic video image generation and ensures the quality of image stitching by obtaining the stitching parameters for future time moments in advance.

[0141] Optionally, in this embodiment of the disclosure, S101d2 may specifically include the following steps (a) to (c):

[0142] Step (a) When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network.

[0143] Step (b) uses an image registration algorithm and an image stitching algorithm to process the N predicted image frames to obtain a predicted stitched image.

[0144] Step (c) Obtain the target prediction stitching parameters based on the predicted stitched image.

[0145] Optionally, the image generation network in this embodiment of the present disclosure can be understood as a self-encoder, for example, the image generation network is U-Net.

[0146] In one embodiment of this disclosure, the image registration algorithm refers to identifying feature points in each image, matching feature points using overlapping regions, and completing the registration process. Specific details can be found in related technical descriptions, and this disclosure does not limit the scope of the embodiments.

[0147] Example 2, in conjunction with Example 1 in the above embodiments. For example... Figure 4 The diagram illustrates the process of obtaining the target prediction stitching parameters. After obtaining the prediction result information of the stitching parameter change, when the stitching parameter change prediction result information indicates that the stitching parameters have changed, based on AC-1, AC-N, and the corresponding optical flow estimation frame AP, a predicted image frame APF is obtained based on the pre-acquired image generation network U-Net. Similarly, based on BC-1, BC-N, and the corresponding optical flow estimation frame BP, a predicted image frame BPF is obtained based on the pre-acquired image generation network U-Net. Image registration and image stitching algorithms are used to process the APF and BPF to obtain the predicted stitched image; based on the predicted stitched image, the target prediction stitching parameters SP are obtained.

[0148] In this way, SP can be used to stitch together N target image frames acquired at the target time to obtain the first panoramic video image.

[0149] It should be noted that the target prediction stitching parameters are obtained by performing registration and stitching calculations after extracting feature points from N prediction image frames.

[0150] In this embodiment, when the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network. Then, an image registration algorithm and an image stitching algorithm are used to process the N predicted image frames to obtain a predicted stitched image. Based on the predicted stitched image, the target predicted stitching parameters are obtained. That is, the predicted stitching parameters are obtained through the predicted image for stitching the real image, thereby reducing the latency of panoramic video image generation and ensuring the quality of image stitching.

[0151] It should be noted that the above embodiments are all illustrative examples of stitching together multiple images acquired at future times by predicting stitching parameters. This disclosure also provides another possible panoramic video image processing method: acquiring N target image frames; and stitching the N target image frames together. In this way, multiple acquired images can be stitched together directly.

[0152] Figure 5 A structural block diagram of a panoramic video image processing apparatus according to an embodiment of the present disclosure is shown. This apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both.

[0153] like Figure 5 As shown, the panoramic video image processing device 200 may include a first acquisition module 201, a second acquisition module 202, and a processing module 203. The first acquisition module can be configured to acquire target prediction stitching parameters, which are stitching parameters corresponding to a target time after a first time. These target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images. The historical images are images captured by N cameras positioned in different directions at the first time and at a time immediately preceding the first time. The second acquisition module 202 can be configured to acquire N target image frames, which are images captured by the N cameras at the target time. The processing module 203 can be configured to stitch the N target image frames according to the target prediction stitching parameters to obtain a first panoramic video image at the target time.

[0154] In one embodiment of this disclosure, the first acquisition module 201 may be specifically configured as follows:

[0155] N first edge image frames and N second edge image frames are acquired. Any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region. Any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region. The N first edge image frames are images captured by the N cameras at the first moment. Each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames.

[0156] For each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time.

[0157] Based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the splicing parameter change prediction result information is obtained based on the pre-trained splicing parameter change prediction network.

[0158] Based on the prediction results of the splicing parameter changes, the target predicted splicing parameters are obtained.

[0159] In one embodiment of this disclosure, the first acquisition module 201 may be specifically configured as follows:

[0160] If the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

[0161] In one embodiment of this disclosure, the first acquisition module 201 may be specifically configured as follows:

[0162] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

[0163] In one embodiment of this disclosure, the first acquisition module 201 may be specifically configured as follows:

[0164] When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network.

[0165] The N predicted image frames are processed using image registration and image stitching algorithms to obtain a predicted stitched image.

[0166] The target predicted stitching parameters are obtained based on the predicted stitched image.

[0167] In one embodiment of this disclosure, the device 200 may further include:

[0168] The third acquisition module can be configured to acquire N first image frames and N second image frames, wherein each of the N first image frames is a full-size image captured by one of the N cameras at a first moment, and each of the N second image frames is a full-size image captured by one of the N cameras at a second moment, wherein the second moment is the moment preceding the first moment.

[0169] The first acquisition module can be specifically configured to crop the N first image frames according to the first overlapping image region to obtain the N first edge image frames; and to crop the N second image frames according to the second overlapping image region to obtain the N second edge image frames.

[0170] This disclosure provides a panoramic video image processing apparatus that can accurately predict future stitching parameters based on changes in stitching parameters of historical images. This allows for the stitching of multiple images acquired in the future using the precisely predicted parameters. Thus, while minimizing latency in the panoramic video image processing, the quality of the stitched image is improved.

[0171] This disclosure also discloses an electronic device. Figure 6 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0172] like Figure 6 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to embodiments of the present disclosure.

[0173] Figure 7A schematic diagram of the structure of a computer system suitable for implementing the method according to embodiments of the present disclosure is shown.

[0174] like Figure 7 As shown, the computer system includes a processing unit that can execute various methods described above based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0175] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks, etc.; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processes via a network such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as needed. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.

[0176] In particular, according to embodiments of this disclosure, the methods described above can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code for performing the methods described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.

[0177] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0178] The units or modules described in the embodiments of this disclosure can be implemented in software or programmable hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0179] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system described above; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to perform the methods described in this disclosure.

[0180] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A panoramic video image processing method, characterized by, The method includes: Obtain target prediction stitching parameters, which are stitching parameters corresponding to the target time after the first time. The target prediction stitching parameters are determined by the changes in stitching parameters obtained from historical images. The historical images are images collected by N cameras set in different directions of the panoramic camera at the first time and the previous time adjacent to the first time, where N is an integer greater than 1. The process of obtaining the target prediction splicing parameters includes: N first edge image frames and N second edge image frames are acquired. Any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region. Any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region. The N first edge image frames are images captured by the N cameras at the first moment. Each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames. For each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time; Based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the splicing parameter change prediction result information is obtained based on the pre-trained splicing parameter change prediction network. Based on the prediction results of the splicing parameter changes, the target predicted splicing parameters are obtained; N target image frames are acquired, wherein the N target image frames are images captured by the N cameras at the target time respectively; Based on the target prediction stitching parameters, the N target image frames are stitched together to obtain the first panoramic video image at the target time.

2. The method of claim 1, wherein, The step of obtaining the target predicted splicing parameters based on the splicing parameter change prediction result information includes: If the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

3. The method of claim 1, wherein, The step of obtaining the target predicted splicing parameters based on the splicing parameter change prediction result information includes: When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

4. The method of claim 3, wherein, When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, including: When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network. The N predicted image frames are processed using image registration and image stitching algorithms to obtain a predicted stitched image. The target predicted stitching parameters are obtained based on the predicted stitched image.

5. The method of claim 1, wherein, Before acquiring N first edge image frames and N second edge image frames, the method further includes: Obtain N first image frames and N second image frames, wherein each of the N first image frames is a full-size image captured by one of the N cameras at a first moment, and each of the N second image frames is a full-size image captured by one of the N cameras at a second moment, wherein the second moment is the moment preceding the first moment; The acquisition of N first edge image frames and N second edge image frames includes: Based on the first overlapping image region, the N first image frames are cropped respectively to obtain the N first edge image frames; Based on the second overlapping image region, the N second image frames are cropped respectively to obtain the N second edge image frames.

6. A panoramic video image processing device, characterized in that, The device includes: The first acquisition module is configured to acquire target prediction stitching parameters. The target prediction stitching parameters are the stitching parameters corresponding to the target time after the first time. The target prediction stitching parameters are determined by the change of stitching parameters obtained from historical images. The historical images are images collected by N cameras set in different directions of the panoramic camera at the first time and the previous time adjacent to the first time, respectively, where N is an integer greater than 1. The first acquisition module is specifically configured as follows: N first edge image frames and N second edge image frames are acquired. Any two first edge image frames captured at adjacent shooting positions in the N first edge image frames have a first overlapping image region. Any two second edge image frames captured at adjacent shooting positions in the N second edge image frames have a second overlapping image region. The N first edge image frames are images captured by the N cameras at the first moment. Each second edge image frame in the second edge image frames is the previous frame image of the corresponding first edge image frame in the N first edge image frames. For each first edge image frame, optical flow estimation is performed on the first edge image frame and the second edge image frame corresponding to the first edge image frame based on the optical flow estimation network to obtain N optical flow estimation frames at the target time. Based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, the splicing parameter change prediction result information is obtained based on the pre-trained splicing parameter change prediction network. Based on the prediction results of the splicing parameter changes, the target predicted splicing parameters are obtained; The second acquisition module is configured to acquire N target image frames, wherein the N target image frames are images captured by the N cameras at the target time respectively; The processing module is configured to perform image stitching on the N target image frames according to the target prediction stitching parameters to obtain the first panoramic video image at the target time.

7. The apparatus of claim 6, wherein, The first acquisition module is specifically configured as follows: If the stitching parameter change prediction result information indicates that the stitching parameters have not changed, the stitching parameters corresponding to the second panoramic video image at the first moment will be determined as the target predicted stitching parameters.

8. The apparatus of claim 6, wherein, The first acquisition module is specifically configured as follows: When the stitching parameter change prediction result information indicates that the stitching parameters have changed, the target predicted stitching parameters are obtained based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame.

9. The apparatus of claim 8, wherein, The first acquisition module is specifically configured as follows: When the stitching parameter change prediction result information indicates that the stitching parameters have changed, for each first edge image frame, based on the first edge image frame, the second edge image frame corresponding to the first edge image frame, and the corresponding optical flow estimation frame, N predicted image frames corresponding to the first edge image frame are obtained based on the pre-acquired image generation network. The N predicted image frames are processed using image registration and image stitching algorithms to obtain a predicted stitched image. The target predicted stitching parameters are obtained based on the predicted stitched image.

10. The apparatus of claim 6, wherein, The device further includes: The third acquisition module is configured to acquire N first image frames and N second image frames, wherein each of the N first image frames is a full-size image captured by one of the N cameras at a first moment, and each of the N second image frames is a full-size image captured by one of the N cameras at a second moment, wherein the second moment is the moment preceding the first moment. The first acquisition module is specifically configured to crop the N first image frames according to the first overlapping image region to obtain the N first edge image frames; and to crop the N second image frames according to the second overlapping image region to obtain the N second edge image frames.

11. An electronic device, comprising: It includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method steps of any one of claims 1 to 5.

12. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, they implement the method steps of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for splicing image sequence

    CN102156867A

  • Method and device for splicing image frames

    CN102393953A