Image processing method and system, electronic equipment and storage medium

By acquiring the shooting pose and depth image sequence and combining it with the user's viewing pose, the playback angle is dynamically adjusted, which solves the motion sickness problem caused by a fixed viewing angle when the camera is moving and improves the immersive viewing experience of video playback.

CN121908000APending Publication Date: 2026-04-21GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2025-12-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

When playing videos, the fixed viewing angle caused by the camera's movement can cause motion sickness in users, affecting the immersive viewing experience.

Method used

By acquiring the shooting pose and depth image sequence from the multimedia file and combining it with the user's viewing pose, the playback angle is dynamically adjusted to generate a target image that matches the user's viewing angle, thus achieving dynamic adjustment of the playback angle.

Benefits of technology

It avoids motion sickness caused by a fixed viewing angle, improving the user's immersive viewing experience and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908000A_ABST
    Figure CN121908000A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and system, electronic equipment and a storage medium, and relates to the technical field of image processing.The method comprises the steps that a multimedia file is obtained, the multimedia file is analyzed, and a first video stream representing a space video, a shooting pose parameter and a depth image sequence are obtained, the shooting pose parameter comprises a shooting path comprising a shooting pose, and the first video stream comprises a left-eye video stream and / or a right-eye video stream; obtaining a watching pose of the user, and determining a matched shooting pose matched with the watching pose in the shooting path; determining a first video frame corresponding to the matched shooting pose according to the first video stream, and determining a first depth image corresponding to the first video frame in the depth image sequence; and according to the watching pose, the matched shooting pose, the first video frame and the first depth image, generating a target image of which the playing view angle is a user watching view angle, and displaying the target image. According to the invention, dynamic adjustment of the playing visual angle is realized when the video is played.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to image processing methods, systems, electronic devices and storage media. Background Technology

[0002] When using devices with shooting capabilities (such as binocular cameras, AR devices, VR devices, etc.) for filming, if the camera is in motion (e.g., a user is walking and filming while holding the device), the video recorded by the device contains spatial video with continuously changing motion perspective information. However, when playing this video, it can only be played from a fixed viewing angle. Since the video content is in continuous motion while the user is stationary, it may cause motion sickness in the user, affecting the user's immersive viewing experience and reducing viewing comfort.

[0003] Therefore, how to dynamically adjust the playback perspective while playing videos has become an urgent problem to be solved.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide an image processing method, system, electronic device, and storage medium, which aims to solve the technical problem of how to dynamically adjust the playback perspective when playing video.

[0006] To achieve the above objectives, this application proposes an image processing method, which is applied to a playback device and includes the following steps: Acquire multimedia files, parse multimedia files, and obtain a first video stream representing spatial video, shooting pose parameters, and a depth image sequence. The shooting pose parameters include a shooting path containing the shooting pose, and the first video stream includes a left-eye video stream and / or a right-eye video stream. Obtain the user's viewing pose and determine the matching shooting pose in the shooting path that matches the viewing pose; Based on the first video stream, determine the first video frame corresponding to the matching shooting pose, and determine the first depth image in the depth image sequence corresponding to the first video frame; Based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, a target image is generated with the playback perspective as the user's viewing perspective, and the target image is displayed.

[0007] Optionally, the step of determining the matching shooting pose in the shooting path that matches the viewing pose includes: Coordinate alignment is performed between the first coordinate system corresponding to the viewing pose and the second coordinate system corresponding to the shooting pose in the shooting path. In response to the first and second coordinate systems after coordinate alignment processing, the shooting pose that matches the viewing pose among the shooting poses is determined as the matched shooting pose.

[0008] Optionally, the shooting pose includes the shooting position in the first coordinate system, and the viewing pose includes the viewing position in the second coordinate system. The steps for determining the shooting pose that matches the viewing pose among all shooting poses are as follows: For each shooting position, determine the first coordinate of the shooting position and the second coordinate of the viewing pose in the same coordinate system, and determine the Euclidean distance between the shooting position and the viewing position based on the first and second coordinates. Determine the smallest Euclidean distance among the Euclidean distances corresponding to each shooting pose, and use the shooting pose corresponding to the smallest Euclidean distance as the matching shooting pose.

[0009] Optionally, the step of determining the shooting pose that matches the viewing pose among the various shooting poses as the matched shooting pose further includes: Connect the various shooting poses to obtain a total arc length image representing the total arc length corresponding to the shooting path; Determine the user's initial viewing pose corresponding to the spatial video of the first video stream, and determine the displacement information of the user moving from the initial viewing pose to the viewing position. Determine the mapping scaling factor between the first coordinate system and the second coordinate system, and map the displacement information into a target arc length image based on the mapping scaling factor. The target arc length image represents the target arc length that is in the same coordinate system as the total arc length and has the same starting point and arc length direction. Based on the total arc length image and the target arc length image, determine the total arc length position that matches the endpoint position of the target arc length, and use the shooting pose corresponding to the matching total arc length position as the matching shooting pose.

[0010] Optionally, the step of generating a target image with a playback perspective as the user's viewing perspective based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image includes: Using a preset forward mapping model, the first depth image is converted into a second depth image in the world coordinate system based on the matched shooting pose; Based on the viewing pose, the second depth image is converted into a third depth image in the first coordinate system corresponding to the viewing pose, and the playback view is the user's viewing view. The color parameters of the third depth image are updated based on the color parameters of the RGB image corresponding to the first video frame to obtain the initial composite image, and the target image is determined based on the initial composite image.

[0011] Optionally, the step of determining the target image based on the initial synthesized image includes: Detect whether there are hole pixels in the initial synthesized image, where hole pixels are pixels that need to be filled with color; If there are empty pixels, the direction of the ray emitted from the shooting end through the empty pixels is determined based on the matching shooting pose and the empty pixels, and the position of the ray that the ray reaches after a preset distance along the ray direction is determined. For each candidate keyframe, the ray position is mapped to the candidate keyframe to obtain the matching pixel in the candidate keyframe corresponding to the ray position. The candidate keyframe is a video frame that is related to the first video frame. Determine the fourth depth image in the depth image sequence that corresponds to the candidate keyframe, and determine the reprojection error of the candidate keyframe based on the fourth depth image, matching pixels, and ray positions; Based on the color parameters of the matching pixels in the candidate keyframe corresponding to the minimum reprojection error, the holed pixels are filled with color to obtain the target image.

[0012] Furthermore, to achieve the above objectives, this application also proposes an image processing method, which is applied to the shooting end and includes the following steps: Acquire a first video stream representing spatial video, wherein the first video stream includes a left-eye video stream and / or a right-eye video stream; Determine the shooting pose and depth image corresponding to each video frame in the first video stream, and construct the shooting path based on the shooting pose; The first video stream, shooting pose parameters, and depth image sequence are encapsulated to obtain a multimedia file, wherein the shooting pose parameters include the shooting path, and the depth image sequence includes depth images. The multimodal file is sent to the playback terminal. After parsing the multimedia file, the playback terminal determines the matching shooting pose in the shooting path that matches the user's viewing pose, and determines the first video frame and the first depth image corresponding to the matching shooting pose in the first video stream and the depth image sequence, respectively. The playback viewpoint generated based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image is the target image of the user's viewing perspective.

[0013] Furthermore, to achieve the above objectives, this application also proposes an image processing system, which includes a shooting end and a playback end. The camera is used to acquire a first video stream that represents spatial video, wherein the first video stream includes a left-eye video stream and / or a right-eye video stream; The camera end is used to determine the shooting pose and depth image corresponding to each video frame in the first video stream, and to construct the shooting path based on the shooting pose. The camera end is used to encapsulate the first video stream, camera pose parameters, and depth image sequence to obtain a multimedia file. The camera pose parameters include the camera path, and the depth image sequence includes depth images. The camera is used to send the multi-modal files to the playback device; The playback end is used to acquire multimedia files, parse multimedia files, and obtain the first video stream representing the spatial video, shooting pose parameters, and depth image sequence. On the playback end, it is used to obtain the user's viewing posture and determine the matching shooting posture in the shooting path that matches the viewing posture; The playback end is used to determine the first video frame corresponding to the matching shooting pose based on the first video stream, and to determine the first depth image in the depth image sequence corresponding to the first video frame; On the playback end, a target image is generated based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, so that the playback perspective is the user's viewing perspective, and the target image is displayed.

[0014] In addition, to achieve the above objectives, this application also proposes an electronic device, which includes: a shooting end, a playback end, a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image processing method described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the image processing method described above.

[0016] In this application, a multimedia file containing a first video stream representing spatial video (e.g., a left-eye video stream and / or a right-eye video stream), shooting pose parameters of a shooting path containing shooting poses, and a depth image sequence is acquired by a shooting end. The multimedia file can be parsed at the playback end to obtain the user's viewing pose and the matching shooting pose in the shooting path that matches the viewing pose. Then, a first video frame corresponding to the matching shooting pose is determined in the first video stream, and a first depth image corresponding to it is determined in the depth image sequence. Based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, a target image with the user's viewing perspective is generated and displayed to achieve playback of spatial video. This allows for the real-time generation of a new perspective (i.e., the user's viewing perspective) that conforms to the laws of physical perspective, ensuring that the target image segmentation displayed at the playback end is accurately matched with the user's motion perception. This enables dynamic adjustment of the playback perspective during video playback, avoiding the phenomenon of motion sickness caused by playing video from a fixed playback perspective, which affects the user's immersive viewing experience and reduces viewing comfort. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the first embodiment of the image processing method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the image processing method of this application; Figure 3 This is a flowchart illustrating the third embodiment of the image processing method of this application; Figure 4 A schematic diagram of the system architecture of the image processing system in this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the image processing method in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] Optionally, the image processing method in this application embodiment can be applied to devices with shooting and / or display functions (such as binocular cameras, AR devices, VR devices, etc.). For example, the shooting end is a binocular camera, and the playback end is a mobile phone or computer, etc. Or, the shooting end and the playback end are integrated into the same device, where the camera of the device is the shooting end (e.g., including a left camera and a right camera), and the display instrument of the device is the playback end, such as a display screen, etc.

[0024] Optionally, this application embodiment uses only a binocular camera as an example.

[0025] Based on this, embodiments of this application provide an image processing method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the image processing method of this application.

[0026] In this embodiment, the image processing method is applied to the shooting end, including steps S100~S400.

[0027] Step S100: Acquire the first video stream representing the spatial video; It should be noted that the first video stream includes the left-eye video stream and / or the right-eye video stream; Alternatively, the shooting end can be a device used to capture scene images, such as a professional VR camera array, a smartphone equipped with a depth camera, an RGB-D (Red Green Blue Depth) sensor, a binocular camera, or a binocular camera.

[0028] Optionally, the acquisition operation can be a spatial video capture operation performed by the shooting end. For example, the color video stream and depth information of the scene can be acquired synchronously or asynchronously through hardware devices such as cameras and depth sensors in the shooting end to obtain the first video stream, RGB (Red, Green, Blue) image, depth image, etc.

[0029] Optionally, the first video stream can be a sequence of base video captured by the camera in response to the shooting scene. It can be a single video stream, such as a left-eye video stream for the user's left eye to view and a right-eye video stream for the user's right eye to view. It can also be a left-eye video stream and a right-eye video stream prepared for stereoscopic viewing.

[0030] Optionally, the first video stream may include multiple video frames, each of which may correspond to an RGB image, a depth image, or an IR (Infrared) image.

[0031] Optionally, the left-eye video stream can be a video stream captured by the left-eye camera at the shooting end, and may include multiple video frames. The left-eye video stream may include left-eye video (which may be spatial video), and the left-eye video may include multiple video frames.

[0032] Optionally, the right-eye video stream can be a video stream captured by the right-eye camera at the shooting end, and may include multiple video frames. The right-eye video stream may include right-eye video (which may be spatial video), and the right-eye video may include multiple video frames.

[0033] Optionally, the spatial video can be video captured by the camera in a specific spatial setting, such as video captured at a wedding. The camera is in motion during video capture; for example, it can rotate 60 degrees to capture the video. The shooting angle for each frame in the spatial video can be different or the same.

[0034] Optionally, the spatial video may contain continuously changing motion perspective information, such as video taken when the user is walking with the handheld camera, or video taken when the camera is shaking.

[0035] Optionally, a shooting end processing module can be set in the playback end, and the shooting end processing module may include a spatial video acquisition unit.

[0036] Optionally, the shooting end processing module is used to synchronously acquire and process relevant data during spatial video recording. The spatial video acquisition unit is used to acquire the time-synchronized left-eye video stream and / or right-eye video stream captured by the shooting end (such as a binocular camera) to obtain the first video stream.

[0037] Optionally, the camera can capture spatial video of the scene to be captured. The capture time can be a first time period (e.g., 30 minutes). In this first time period, the left-eye camera can capture the left-eye video stream of the first time period, and the right-eye camera can simultaneously capture the right-eye video stream of the first time period.

[0038] Optionally, if the camera needs to capture both the left-eye and right-eye video streams, the captured left-eye and right-eye video streams are time-synchronized video streams.

[0039] Step S200: Determine the shooting pose and depth image corresponding to each video frame in the first video stream, and construct the shooting path based on the shooting pose; Optionally, a video frame can correspond to a single image frame. For example, when the first video stream consists of a left-eye video stream and a right-eye video stream, a video frame in the left-eye video stream can correspond to a single image frame in the left-eye video stream, such as the image at the 8th second (which can be at least one of an RGB image, a depth image, or an IR image). Similarly, a video frame in the right-eye video stream can correspond to a single image frame in the right-eye video stream, such as the image at the 8th second (which can be at least one of an RGB image, a depth image, or an IR image).

[0040] Optionally, the shooting pose can be the shooting end pose, which may include the spatial position and orientation of the shooting end. For example, when the shooting end is a binocular camera, the shooting pose may include the spatial position of the binocular camera, as well as the orientation and shooting angle of the binocular camera.

[0041] Optionally, the shooting path can be a set of various shooting positions of the shooting end in the shooting space video.

[0042] Optionally, the shooting end processing module in the playback end may include a shooting pose estimation unit. The shooting pose estimation unit may be based on visual inertial odometry (VIO) technology, using camera images and inertial measurement unit (IMU) data to estimate the shooting pose (such as position and orientation) of the camera in three-dimensional space for each video frame in real time, forming a shooting path.

[0043] Optionally, after capturing the first video stream representing the spatial video, the camera can determine the shooting pose corresponding to each video frame in the first video stream through the shooting pose estimation unit, and summarize the various shooting poses to form a shooting path.

[0044] Optionally, the shooting path can be represented in the form of an image, such as constructing a coordinate axis image of the relationship between shooting pose and time, and representing the shooting pose corresponding to each video frame in the first video stream in the coordinate axis image to form an image representing the shooting path.

[0045] Optionally, the shooting path may include the shooting path corresponding to the left eye video stream and / or the shooting path corresponding to the right eye video stream.

[0046] Optionally, the depth image corresponding to each video frame in the first video stream can be determined, and the depth images can be sorted and summarized according to the order of shooting time to obtain a depth image sequence, such as the depth image sequence corresponding to the left eye video stream and the depth image sequence corresponding to the right eye video stream.

[0047] Step S300: The first video stream, shooting pose parameters, and depth image sequence are encapsulated to obtain a multimedia file; It should be noted that the shooting pose parameters include the shooting path, and the depth image sequence includes depth images; Optionally, the shooting pose parameters may include a shooting pose sequence, a shooting path, and a shooting pose, and may also include other parameters related to the first video stream, such as shooting angle, ambient lighting, temperature, etc. The shooting pose sequence may include shooting poses ordered chronologically by shooting time.

[0048] Optionally, the shooting end processing module in the playback end may include a data encapsulation unit. The data encapsulation unit may encapsulate the first video stream, shooting pose parameters and depth image sequence to obtain a multimedia file. For example, the left eye video stream, the right eye video stream, the time-synchronized shooting pose sequence (such as shooting pose and shooting path), and the depth image sequence may be encapsulated together into a multimedia file container to obtain a multimedia file.

[0049] Step S400: Send the multimodal file to the playback device; It should be noted that after parsing the multimedia file, the playback device determines the matching shooting pose in the shooting path that matches the user's viewing pose, and determines the first video frame and the first depth image corresponding to the matching shooting pose in the first video stream and the depth image sequence, respectively. The playback viewpoint generated based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image is the target image of the user's viewing perspective.

[0050] Optionally, after capturing a multimedia file, the camera can send the multimedia file to the playback device for parsing and playback. The playback device can also dynamically encrypt the multimedia file before sending it to the playback device. For example, it can obtain the viewing pose (referred to as historical viewing pose, such as the viewing pose of the previous video viewing) from the user's historical viewing data in the playback device, and search for the index corresponding to the historical viewing pose in a lookup table containing different indexes and viewing pose mappings. It can then obtain the initial filename of the multimedia file and encrypt it using a preset encryption algorithm (such as DSA (Digital Signature Algorithm)) combined with the index corresponding to the historical viewing pose and the initial filename. For instance, it can use a preset encryption algorithm to perform ciphertext conversion on the index corresponding to the historical viewing pose to obtain a key ciphertext, and then concatenate the key ciphertext with the initial filename to obtain the target filename. The target filename is then used as the actual filename of the multimedia file for storage or sent to the playback device.

[0051] Optionally, the camera can also perform permission checks on the playback device. Once it is determined that the playback device has playback permissions, the multimedia file will be sent to the playback device.

[0052] Optionally, the camera can upload multimedia files to the cloud so that the playback device can download the multimedia files from the cloud for parsing and playback.

[0053] In this embodiment, a multimedia file containing a first video stream representing spatial video (e.g., a left-eye video stream and / or a right-eye video stream), shooting pose parameters of a shooting path containing shooting pose, and a depth image sequence is acquired by the shooting end. This multimedia file is then sent to the playback end. The playback end can subsequently parse the multimedia file to obtain the user's viewing pose and the matching shooting pose in the shooting path. Then, a first video frame corresponding to the matching shooting pose is determined in the first video stream, and a first depth image corresponding to it is determined in the depth image sequence. Based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, a target image with the user's viewing perspective is generated and displayed. This enables the playback of spatial video. Furthermore, by generating a new perspective (i.e., the user's viewing perspective) that conforms to the laws of physical perspective in real time, the target image displayed by the playback end is precisely matched with the user's motion perception. This allows for dynamic adjustment of the playback perspective during video playback, avoiding the use of a fixed playback perspective that could cause motion sickness, affect the user's immersive viewing experience, and reduce viewing comfort.

[0054] Based on the above embodiments, this application provides an image processing method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the image processing method of this application.

[0055] In this embodiment, the image processing method is applied to the playback end, including steps S10 to S40.

[0056] Step S10: Acquire multimedia files, parse multimedia files, and obtain the first video stream representing the spatial video, shooting pose parameters, and depth image sequence; It should be noted that the shooting pose parameters include the shooting path containing the shooting pose, and the first video stream includes the left eye video stream and / or the right eye video stream. Optionally, the playback device can be an electronic device used to ultimately display images or videos, such as a virtual reality (VR) headset, augmented reality (AR) glasses, a smartphone, a tablet, a computer, etc. The playback device and the shooting device can be located on the same device or on different devices.

[0057] Optionally, a playback rendering module can be set in the playback terminal to dynamically render the video frame corresponding to the first video stream based on the user's viewing posture.

[0058] Optionally, the playback rendering module may include a data parsing unit for parsing multimedia files.

[0059] Optionally, the playback device can receive multimedia files sent by the shooting device, download multimedia files from the cloud, or obtain multimedia files from other terminals.

[0060] Optionally, the playback device can parse the multimedia file to obtain a time-synchronized first video stream, shooting pose parameters (such as shooting pose and shooting path), and a depth image sequence (such as depth image).

[0061] Optionally, if the multimedia file obtained by the playback terminal is an encrypted multimedia file, the actual filename of the multimedia file can be obtained, and the parts other than the initial filename of the multimedia file can be extracted from the actual filename to obtain the key ciphertext. The user's current actual viewing pose is collected, and the index corresponding to the user's current actual viewing pose is searched in a lookup table containing different indexes and viewing pose mapping relationships. The index corresponding to the user's current actual viewing pose is converted into ciphertext using a preset encryption algorithm to obtain the verification ciphertext. The verification ciphertext and the key ciphertext are checked for consistency. If the verification ciphertext and the key ciphertext are consistent, the playback terminal parses the multimodal file to obtain the first video stream representing the spatial video, the shooting pose parameters, and the depth image sequence. If the verification ciphertext and the key ciphertext are inconsistent, the parsing process of the multimodal file by the playback terminal is paused, and a prompt message is output to remind the user that the viewing pose is incorrect and the multimodal file cannot be parsed.

[0062] Step S20: Obtain the user's viewing pose and determine the matching shooting pose in the shooting path that matches the viewing pose; Optionally, the viewing posture can be the spatial position and face orientation of the user watching the spatial video corresponding to the first video stream through the playback terminal, that is, the user's location and the user's orientation.

[0063] Optionally, the matching shooting pose can be the shooting pose selected from the shooting path that is closest or most relevant to the current user's viewing pose in terms of spatial location and has the same orientation (e.g., both facing north).

[0064] Optionally, the playback rendering module in the playback terminal may also include a viewing pose acquisition unit, used to acquire the user's head viewing pose in physical space in real time. For example, the user's viewing pose can be detected by a position sensor and / or a posture sensor.

[0065] Optionally, the shooting path can be used to represent the shooting poses corresponding to different time periods and different video frames. In this case, the shooting poses and viewing poses can be transformed into the same coordinate system. In the same coordinate system, the viewing poses are compared and matched with each shooting pose to determine the shooting pose with the highest similarity to the viewing pose, and this is used as the matching shooting pose.

[0066] Step S30: Determine the first video frame corresponding to the matching shooting pose based on the first video stream, and determine the first depth image in the depth image sequence corresponding to the first video frame; Optionally, the first video stream may include a left-eye video stream and / or a right-eye video stream. A video frame corresponding to the matching shooting pose may be determined based on the left-eye video stream, and a depth image corresponding to that video frame may be determined from the depth image sequence; alternatively, a video frame corresponding to the matching shooting pose may be determined based on the right-eye video stream, and a depth image corresponding to that video frame may be determined from the depth image sequence.

[0067] Optionally, a time series corresponding to the matching shooting pose can be determined (i.e., determined based on the shooting time, with each shooting time corresponding to a time series), and at least one video frame corresponding to that time series can be found in the first video stream and used as the first video frame. The number of video frames in the first video frame can be one or more.

[0068] Optionally, since the shooting poses along the shooting path change continuously over time, the corresponding first video frame can be found in the first video stream by combining the changes in the viewing pose. This avoids a disconnect between the image corresponding to the selected first video frame and the already displayed image.

[0069] Optionally, when playing the spatial video corresponding to the first video stream on the playback terminal, the video frame corresponding to the currently displayed image can be used as the first video frame.

[0070] Optionally, a depth image corresponding to the time series can be found in the depth image sequence and used as the first depth image.

[0071] Step S40: Generate a target image with the playback viewpoint as the user's viewing viewpoint based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, and display the target image.

[0072] Optionally, the playback viewpoint corresponding to the first depth image (hereinafter referred to as the first playback viewpoint) can be determined based on the matching shooting pose (including the spatial position of the shooting end (i.e., the shooting position) and the orientation of the shooting end (i.e., the orientation of the shooting end)). For example, the first playback viewpoint may be consistent with the orientation of the shooting end. The playback viewpoint corresponding to the viewing pose (including the user's location (i.e., the viewing position) and the user's orientation (hereinafter referred to as the second playback viewpoint) can be determined. For example, the second playback viewpoint may be consistent with the user's orientation, or it may be consistent with the orientation of the user's face.

[0073] Optionally, the playback viewpoint corresponding to the first depth image can be converted from the first playback viewpoint to the second playback viewpoint to obtain a depth image (or grayscale image) of the second playback viewpoint. Then, the color parameters of the depth image of the second playback viewpoint can be updated using the RGB image corresponding to the first video frame. For example, the color parameters of the depth image of the second playback viewpoint can be consistent with the color parameters of the RGB image corresponding to the first video frame. The coordinates of the depth image of the second playback viewpoint after the color parameter update can be used as the target image of the user's viewing viewpoint.

[0074] Optionally, the target image can be displayed on the display screen or display interface of the playback device.

[0075] Optionally, each video frame in the first video stream can be converted using steps S10 to S40 to obtain target images corresponding to multiple video frames, and then played and displayed sequentially according to the playback time order.

[0076] Optionally, in a scenario where the user's viewing posture is constantly changing, a mapping relationship can be established between different shooting postures and different viewing postures. Based on this mapping relationship, the playback angle of the image corresponding to each video frame in the first video stream can be adjusted to obtain a target image that matches the user's viewing angle, which is then played and displayed. In this case, as the user's viewing posture changes, the playback angle corresponding to the target image played by the playback terminal also changes accordingly, ensuring that the changed playback angle corresponds to the changed viewing posture.

[0077] Optionally, in another scenario, if the user's viewing posture remains unchanged, a matching shooting posture that matches the viewing posture in the shooting path can be determined, and the playback angle of the image corresponding to each video frame in the first video stream can be adjusted based on the matching shooting posture to obtain a target image that matches the user's viewing angle and then played and displayed.

[0078] Optionally, in this embodiment, by establishing a mapping relationship between the shooting posture and the viewing posture, the playback angle can be dynamically adjusted in real time, matching visual changes with the user's motion perception, thereby fundamentally eliminating sensory conflict and improving viewing comfort. In this embodiment, a multimedia file containing a first video stream representing spatial video (e.g., a left-eye video stream and / or a right-eye video stream), shooting pose parameters of a shooting path containing shooting poses, and a depth image sequence is acquired by the shooting end. The multimedia file can be parsed at the playback end to obtain the user's viewing pose and the matching shooting pose in the shooting path that matches the viewing pose. Then, a first video frame corresponding to the matching shooting pose is determined in the first video stream, and a first depth image corresponding to it is determined in the depth image sequence. Based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, a target image with the user's viewing perspective is generated and displayed to achieve the playback of spatial video. This allows for the real-time generation of a new perspective (i.e., the user's viewing perspective) that conforms to the laws of physical perspective, ensuring that the target image segmentation displayed at the playback end is accurately matched with the user's motion perception. This enables dynamic adjustment of the playback perspective during video playback, avoiding the phenomenon of motion sickness caused by playing video from a fixed playback perspective, which affects the user's immersive viewing experience and reduces viewing comfort.

[0079] Based on the first or second embodiment of this application, a third embodiment of this application is proposed. In this third embodiment, content that is the same as or similar to the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, refer to... Figure 3 In step S20, the step of determining the matching shooting pose that matches the viewing pose in the shooting path includes steps a10-a20.

[0080] Step a10: Perform coordinate alignment processing on the first coordinate system corresponding to the viewing pose and the second coordinate system corresponding to the shooting pose in the shooting path; Step a20: In response to the first and second coordinate systems after coordinate alignment processing, determine the shooting pose that matches the viewing pose among the shooting poses as the matched shooting pose.

[0081] Optionally, the first coordinate system can be the coordinate system of the physical space where the user or playback device is located (such as a bedroom, a movie theater, etc.) (such as a physical space coordinate system), and can be based on a specific location in the physical space as the origin, such as a corner of a wall. Under specific scenario conditions, the first coordinate system can be the world coordinate system.

[0082] Optionally, the second coordinate system can be a virtual scene coordinate system, a coordinate system of the physical space where the camera is located, or a world coordinate system.

[0083] Optionally, if neither the first coordinate system nor the second coordinate system is a world coordinate system, then both the first coordinate system and the second coordinate system can be transformed to the world coordinate system to achieve alignment between the first coordinate system and the second coordinate system.

[0084] Optionally, the second coordinate system can be transformed to be consistent with the first coordinate system based on the first coordinate system, or the first coordinate system can be transformed to be consistent with the second coordinate system based on the second coordinate system, or the first and second coordinate systems can be transformed simultaneously to be in a common coordinate system, thereby completing the coordinate alignment process between the first and second coordinate systems.

[0085] Optionally, the user's viewing position can be determined in the first coordinate system, and the user's orientation at that viewing position (e.g., the direction the user's face is facing (e.g., facing due north)) can be determined, and the user's viewing position and orientation can be used as the viewing pose. The shooting position of the shooting end can be determined in the second coordinate system, and the shooting end's orientation at that shooting position (e.g., the shooting end facing due south or due north) can be determined, and the shooting position and orientation can be used as the shooting pose.

[0086] Optionally, coordinate alignment is required when mapping the viewing pose to the shooting path.

[0087] Optionally, a coordinate alignment unit can be set in the playback rendering module of the playback terminal to align the user's first coordinate system with the second coordinate system during the initial playback phase, such as scale alignment.

[0088] Optionally, the shooting pose includes the shooting orientation (i.e., the orientation of the shooting end) and the shooting position. The viewing pose includes the viewing orientation (i.e., the user's orientation) and the viewing position. After aligning the coordinates of the shooting position and the viewing position, that is, after aligning the coordinates of the shooting pose and the viewing pose, the shooting pose that matches the viewing pose among the various shooting poses can be determined and used as the matching shooting pose. For example, the matching shooting position that matches the viewing position among the various shooting poses can be determined, and the shooting orientation that is consistent with the viewing orientation among the shooting orientations corresponding to each matching shooting position can be determined and its corresponding shooting pose can be used as the matching shooting pose.

[0089] In this embodiment, by first aligning the coordinates of the first coordinate system and the second coordinate system, and then determining the matching shooting pose that matches the viewing pose among each shooting pose, it is possible to determine the matching shooting pose in the same coordinate system, thus ensuring the accuracy of the determined matching shooting pose.

[0090] Optionally, step a20, which involves determining the shooting pose that matches the viewing pose among the various shooting poses, includes steps a21-a22.

[0091] Step a21: For each shooting position, determine the first coordinate of the shooting position and the second coordinate of the viewing pose in the same coordinate system, and determine the Euclidean distance between the shooting position and the viewing position based on the first coordinate and the second coordinate. Step a22: Determine the smallest Euclidean distance among the Euclidean distances corresponding to each shooting pose, and use the shooting pose corresponding to the smallest Euclidean distance as the matching shooting pose.

[0092] Optionally, the shooting pose includes the shooting position in the first coordinate system, and the viewing pose includes the viewing position in the second coordinate system.

[0093] Optionally, the nearest neighbor mapping method can be used to determine the shooting position with the closest Euclidean distance to the user's viewing position in each shooting position, and the shooting pose coordinates corresponding to the shooting position with the closest Euclidean distance can be matched with the shooting pose.

[0094] Optionally, for each shooting pose, the same method (such as using the Euclidean distance function to calculate the Euclidean distance) can be used to calculate the Euclidean distance between the viewing position and each shooting position.

[0095] Alternatively, the formula for the Euclidean distance function can be as shown in Formula 1 below.

[0096] (Formula 1); Taking the world coordinate system as an example, if the shooting path includes multiple shooting poses, such as P = {P1, P2, ..., Pn}, and in the world coordinate system, the shooting position of shooting pose Pi is Pi = (xi, yi, zi), where i = 1..n, n is an integer greater than 1, xi is the x-coordinate of shooting pose Pi, yi is the y-coordinate of shooting pose Pi, and zi is the z-coordinate of shooting pose Pi. In the world coordinate system, the shooting position corresponding to viewing pose Vhead is Vhead = (xv, yv, zv), where xv is the x-coordinate of viewing pose Vhead, yv is the y-coordinate of viewing pose Vhead, and zv is the z-coordinate of viewing pose Vhead. di is the Euclidean distance between viewing pose Vhead and shooting pose Pi.

[0097] Optionally, after determining the Euclidean distance between the viewing position and each shooting position, that is, after determining the Euclidean distance between the viewing pose and each shooting pose, the various Euclidean distances can be compared with each other to determine the minimum Euclidean distance, and the shooting pose corresponding to the minimum Euclidean distance can be used as the matching shooting pose.

[0098] In this embodiment, by calculating the Euclidean distance between the viewing pose and each shooting pose, and then taking the shooting pose corresponding to the smallest Euclidean distance as the matching shooting pose, the accuracy of the determined matching shooting pose is ensured.

[0099] Optionally, in step a20, the step of determining the shooting pose that matches the viewing pose among the shooting poses as the matching shooting pose further includes steps a23-a26.

[0100] Step a23: Connect each shooting pose to obtain a total arc length image representing the total arc length corresponding to the shooting path; Step a24: Determine the starting viewing pose of the spatial video corresponding to the first video stream when the user starts watching it, and determine the displacement information of the starting viewing pose moving to the viewing position. Step a25: Determine the mapping scaling factor between the first coordinate system and the second coordinate system, and map the displacement information into a target arc length image based on the mapping scaling factor. The target arc length image represents the target arc length that is in the same coordinate system as the total arc length and has the same starting point and arc length direction. Step a26: Based on the total arc length image and the target arc length image, determine the total arc length position that matches the endpoint position of the target arc length, and use the shooting pose corresponding to the matching total arc length position as the matching shooting pose.

[0101] Optionally, the displacement information may include the effective distance the user's face has moved in the first coordinate system.

[0102] Optionally, the shooting positions of each shooting pose in the shooting path can be connected to form the total arc length corresponding to the shooting path. That is, the first coordinates of each shooting pose can be connected to form a total arc length image representing the total arc length of the shooting path. For example, it can be shown in Formula 2.

[0103] (Formula 2); in, For the total arc length, The Euclidean distance between two adjacent shooting poses.

[0104] Optionally, when a user begins watching the spatial video corresponding to the first video stream through the playback terminal, the playback terminal begins recording the user's viewing posture to determine the initial viewing posture, the current viewing posture, and the distance traveled from the initial viewing posture to the current viewing position. This distance is then used as the displacement information from the initial viewing posture to the current viewing position. The displacement information may also include the direction of movement from the initial viewing posture to the current viewing position.

[0105] Alternatively, the movement distance can be represented by a displacement component parallel to the main direction of the shooting path, or by forward, backward, left, and right movement on a horizontal plane.

[0106] Optionally, a mapping scaling factor between the first and second coordinate systems can be pre-set. This scaling factor determines how many meters a user moves on the virtual path (i.e., the movement path in the space corresponding to the second coordinate system, such as the space where the shooting end is located, or the location of the camera) corresponding to a one-meter movement in physical space (i.e., the physical space corresponding to the first coordinate system, or the location of the user or playback device). In other words, it converts the movement distance corresponding to the viewing position to the corresponding displacement distance on the shooting path. For example, if the mapping scaling factor is 1, it means that for every 1 meter the user moves, the corresponding displacement distance on the shooting path is 1 meter. That is, if the movement distance is 1 meter, the target arc length on the shooting path is 1 meter using the mapping scaling factor. If the mapping scaling factor is 0.5, it means that for every 1 meter the user moves, the corresponding displacement distance on the shooting path is 0.5 meters. That is, if the movement distance is 1 meter, the target arc length on the shooting path is 0.5 meters using the mapping scaling factor. For example, as shown in Formula 3 below... Starget=Sstart+ (Formula 3); Where Starget is the target arc length, Sstart is the arc length of the starting point (i.e., the position corresponding to the first shooting pose) in the total arc length corresponding to the shooting path, and Dv is the movement distance. This is the mapping scaling factor.

[0107] Optionally, the target arc length image represents the target arc length that is in the same coordinate system as the total arc length (which can be the world coordinate system, or the first or second coordinate system after coordinate alignment processing), and whose starting point and arc length direction are consistent.

[0108] Optionally, after determining the target arc length image representing the target arc length and the total arc length image representing the total arc length, the first coordinate corresponding to the first shooting pose in the total arc length image can be used as the starting point to extend backward until the target arc length is reached. The shooting position corresponding to the first coordinate corresponding to the first shooting pose in the total arc length, which is the starting point, is taken as the shooting position of the endpoint reached by the target arc length, and the shooting pose corresponding to the endpoint position is taken as the matching shooting pose.

[0109] Optionally, after determining the matching shooting pose, at least one video frame corresponding to the matching shooting pose can be used as a key video frame, such as the first video frame.

[0110] Optionally, the first video frame can be determined using arc length mapping to ensure it is the correct frame for playback. Since the total arc length of the shooting path increases over time (the camera cannot rewind time), as long as the user's movement is continuous, the mapped target arc length also changes continuously. Consequently, the selected first video frame is also temporally continuous or highly proximate. This fundamentally avoids large "jumps" in time, ensuring the smoothness of video content changes. It cleverly transforms the user's spatial movement into a smooth "scrubbing" of the video timeline, thus enabling the playback and browsing of coherent video footage.

[0111] In this embodiment, by converting the shooting path into a total arc length image, the displacement information between the user's initial viewing pose and the current viewing pose is converted into a target arc length in the same coordinate system as the total arc length through a mapping scaling factor, thereby generating a target arc length image. Then, the total arc length image and the target arc length image are compared to determine the matching shooting pose, thus ensuring the accuracy of the determined matching shooting pose.

[0112] Based on any of the above embodiments, a fourth embodiment of this application is proposed. In this fourth embodiment, content that is the same as or similar to the above embodiments can be referred to the above description and will not be repeated hereafter. Based on this, step S40, which generates a target image with a playback perspective of the user's viewing perspective based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, includes steps b10-b30.

[0113] Step b10: Using a preset forward mapping model, the first depth image is converted into a second depth image in the world coordinate system based on the matched shooting pose. Step b20: Based on the viewing pose, the second depth image is converted into a third depth image in the first coordinate system corresponding to the viewing pose, and the playback view is the user's viewing view. Step b30: Update the color parameters of the third depth image according to the color parameters of the RGB image corresponding to the first video frame to obtain the initial composite image, and determine the target image based on the initial composite image.

[0114] Optionally, the function formulas corresponding to the forward mapping model can be as shown in Formulas 4, 5, and 6 below.

[0115] (Formula 4); (Formula 5); (Formula 6); For each pixel coordinate in the first video frame Fm =( , ) T A uv coordinate system is constructed with the top left corner of the first video frame as the origin, the u-axis pointing horizontally to the right, and the v-axis pointing vertically downwards. , () represents the coordinates of a point in the UV coordinate system, and T is the transpose sign. For the pixel position of the first video frame Fm in the world coordinate system, match the shooting pose. = , To match the rotation matrix of the shooting pose, To match the translation vector of the shooting pose (e.g., the translation vector corresponding to the target arc length), observe the pose. , To view the rotation matrix of the pose, K represents the translation vector of the viewing pose (e.g., the vector corresponding to the distance moved), and K is the intrinsic parameter matrix of the shooting end, such as the camera intrinsic parameter matrix. The inverse of the rotation matrix for the shooting pose. For a depth image, such as the first depth image, Let K be the inverse matrix of the intrinsic parameter matrix of the shooting end. These are intermediate parameters (such as the pixel position of a pixel in the first video frame Fm in the first coordinate system). The pixel coordinates of the target image or the initial composite image are used to determine the playback viewpoint, which is the user's viewing perspective. , For pixels The coordinates of the point in the uv coordinate system.

[0116] Optionally, Formula 4 can transform the image corresponding to the first video frame from the uv coordinate system to the first coordinate system.

[0117] Optionally, the first depth image and the matching shooting pose can be input into Formula 4 to convert the first depth image into an image in the first coordinate system and use it as the second depth image. Furthermore, element-wise operation can be performed during the conversion operation, that is, for each pixel in the first video frame, it is converted into the first coordinate system according to the depth pixel in its corresponding depth image.

[0118] Optionally, a first coordinate system corresponding to the viewing pose can be determined, and based on the first and second coordinate systems after coordinate alignment, the second depth image located in the second coordinate system can be transformed to the first coordinate system. At this time, the playback view of the second depth image in the first coordinate system is associated with the shooting pose of the shooting end (e.g., the first playback view). Then, the playback view of the second depth image in the first coordinate system can be transformed from the first playback view to the user's viewing view (e.g., the second playback view) through the viewing pose to obtain the third depth image of the user's viewing view.

[0119] Optionally, the viewing pose and the second depth image can be input into Formula 5, and the result of Formula 5 can be input into Formula 6 to obtain the third depth image with the playback viewpoint as the user's viewing viewpoint.

[0120] Optionally, the RGB image corresponding to the first video frame can also be determined. For any pixel in the first video frame, the color parameters (such as RGB values ​​and color values) corresponding to it in the RGB image can be assigned to the corresponding pixel in the third depth image to perform color update processing on the third depth image and obtain the initial synthesized image.

[0121] Alternatively, the initial synthesized image can be used as the final target image, or the initial synthesized image can be subjected to adaptive processing (such as noise removal) to obtain the target image.

[0122] In this embodiment, the first depth image is converted into a depth image in the first coordinate system by using a forward mapping model and matching the shooting pose and viewing pose. Then, the playback viewpoint corresponding to the viewing pose is used to convert it into a third depth image with the playback viewpoint being the user's viewing viewpoint. The color parameters are then updated using the RGB image corresponding to the first video frame to obtain an initial synthesized image. The target image is then determined based on the initial synthesized image, thereby ensuring the accuracy of the determined target image.

[0123] Optionally, step b50, which involves determining the target image based on the initial synthesized image, includes steps b51-b55.

[0124] Step b51: Detect whether there are hole pixels in the initial synthesized image, where hole pixels are pixels that need to be filled with color; Step b52: If there are empty pixels, then based on the matching shooting pose and the empty pixels, determine the ray direction of the ray emitted from the shooting end that passes through the empty pixels, and determine the ray position that the ray reaches after traveling a preset distance along the ray direction. Step b53: For each candidate keyframe, map the ray position to the candidate keyframe to obtain the matching pixel in the candidate keyframe corresponding to the ray position, wherein the candidate keyframe is a video frame that is associated with the first video frame. Step b54: Determine the fourth depth image in the depth image sequence that corresponds to the candidate keyframe, and determine the reprojection error of the candidate keyframe based on the fourth depth image, matching pixels, and ray positions. Step b55: Based on the color parameters of the matching pixels in the candidate keyframe corresponding to the smallest reprojection error, fill the hole pixels with color to obtain the target image.

[0125] Optionally, the initial synthesized image can be subjected to hole pixel detection processing to identify whether there are pixels or pixel regions in the initial synthesized image that need to be filled with color, thereby determining whether there are hole pixels in the initial synthesized image.

[0126] Optionally, a binary mask of the same size as the initial composite image can be created and initialized to 0 (effective pixels). For each pixel in the initial composite image, it can be judged according to the following formula 7 to determine whether it is a hole pixel.

[0127] (Formula 7); Where M(i,j) can be the effective pixels in the binary mask. These are the pixels in the initial synthesized image.

[0128] Optionally, if the detection finds that there are no hole pixels in the initial synthesized image, the initial synthesized image can be used as the target image. If there are hole pixels, the same processing is performed on each hole pixel in the initial synthesized image. The following example only illustrates the processing of a single hole pixel.

[0129] Optionally, multi-frame geometric filling can be performed on the holed pixels to find the most accurate color parameters for the holed pixels.

[0130] Optionally, a ray equation can be set, and an optimization problem can be constructed based on the ray equation. The ray equation can be used to determine the ray direction of the ray emitted from the camera's viewpoint that passes through the hole pixel based on the inverse matrix of the hole pixel and the intrinsic parameter matrix of the camera. The ray direction in the world coordinate system can be determined based on the inverse matrix of the rotation matrix of the matching camera pose combined with the ray direction in the camera's viewpoint. The ray position reached by the ray along the ray direction and the preset distance can be determined by combining the origin coordinates of the camera corresponding to the matching camera pose, the preset distance, and the ray direction in the world coordinate system.

[0131] Optionally, video frames in the first video stream that are related to the first video frame can be used as candidate keyframes, such as video frames that are 5 positions before and / or 5 positions after the first video frame in time sequence.

[0132] Optionally, for each candidate keyframe, the ray equation can be used to map the ray position to the candidate keyframe to determine the pixel position in the candidate keyframe through which the ray passes, and the pixel at that pixel position is used as the matching pixel.

[0133] Alternatively, the ray equation can be used to determine the reprojection error of the candidate keyframe by combining the depth image (i.e., the fourth depth image) corresponding to the candidate keyframe, the matching pixels, and the ray position. The reprojection error of each candidate keyframe can be determined, and the candidate keyframe with the smallest reprojection error can be selected. Then, the color parameters of the matching pixels in the candidate keyframe with the smallest reprojection error can be used to fill the hole pixels with color. For example, the color parameters of the hole pixels after color filling are consistent with the color parameters of the matching pixels in the candidate keyframe with the smallest reprojection error.

[0134] Optionally, each hole pixel in the initial synthesized image can be filled with color to obtain the target image.

[0135] Optionally, the ray equation can correspond to Equations 8, 9, and 10, and the optimization problem can correspond to Equation 11.

[0136] (Formula 8); (Formula 9); , >0 (Formula 10); (Formula 11); in, The direction of the ray from the camera's perspective. It is the ray direction in the world coordinate system. The purpose is to accurately know the ray direction in space from the shooting end's perspective when there are empty pixels in the initial composite image. It is the inverse of the intrinsic parameter matrix K of the shooting end. These are empty pixels.

[0137] If the viewing pose is set to the target shooting pose of the camera, then the target shooting pose is... , Rotation matrix for target pose The inverse matrix, The translation vector for the target's pose. The origin point of the target shooting end. Represents reprojection error, measuring the depth of a given guess. The degree of inconsistency between the corresponding 3D point (e.g., 3D coordinates) and the geometric information recorded in the source keyframe i (which could be a candidate keyframe) is determined by... The smaller the value, the deeper the guess. The more accurate. The distance from the optical center of the target camera (viewer) along the line of sight to a certain 3D point is defined. These are the core variables that need to be optimized. It is a two-dimensional vector representing the current... The determined 3D points , the projection position on the image plane of the source keyframe i. It is the measured depth of source keyframe i, representing the depth at pixel coordinates from the depth image of source keyframe i (e.g., the first depth image). The depth value retrieved records the actual distance from the source keyframe camera to the surface of the object corresponding to that pixel within its field of view during the shooting process. Represents the rotation matrix of source keyframe i, a 3x3 matrix. It describes the rotation of vectors in the world coordinate system to align with the camera coordinate system (e.g., uv coordinate system) of source keyframe i. This represents the translation vector of source keyframe i, a 3x1 vector. It describes the displacement between the origin of the world coordinate system and the origin of the camera coordinate system (e.g., the uv coordinate system) of source keyframe i. The 3D point on the ray (i.e., a three-dimensional coordinate) represents the specific position reached in the world coordinate system after traveling a distance λ along the line of sight from the optical center of the target camera.

[0138] and 3D points on the ray Project onto source keyframe i to obtain the pixel coordinates (e.g., the pixel coordinates of the matching pixel) under source keyframe i. And according to the following formula (12), the optimal solution and its reprojection error can be found for each candidate keyframe through one-dimensional search or iterative optimization (such as Newton's method).

[0139] Formula (12).

[0140] in, This is the optimal solution.

[0141] Optionally, after obtaining the reprojection errors of multiple candidate keyframes, the color parameters of the matching pixels in the candidate keyframe with the smallest reprojection error can be selected to fill the hole pixels with color, thereby obtaining the target image.

[0142] In this embodiment, when there are empty pixels in the initial synthesized image, the ray direction and position of the ray emitted from the shooting end through the empty pixels are determined based on the matching shooting pose and the empty pixels. Then, the matching pixels corresponding to them in each candidate keyframe are determined. Then, the reprojection error of each candidate keyframe is determined, and the color parameters of the matching pixels in the candidate keyframe corresponding to the smallest reprojection error are selected to fill the empty pixels with color, thereby obtaining the target image. This ensures the accuracy and effectiveness of the determined target image.

[0143] In addition, refer to Figure 4 This application also provides an image processing system, which includes a shooting end A10 and a playback end A20.

[0144] The camera A10 is used to acquire a first video stream characterizing the spatial video, wherein the first video stream includes a left-eye video stream and / or a right-eye video stream; The camera A10 is used to determine the shooting pose and depth image corresponding to each video frame in the first video stream, and to construct the shooting path based on the shooting pose. The camera A10 is used to encapsulate the first video stream, shooting pose parameters, and depth image sequence to obtain a multimedia file. The shooting pose parameters include the shooting path, and the depth image sequence includes depth images. The A10 camera is used to send the multi-modal files to the playback device; The playback device A20 is used to acquire multimedia files, parse multimedia files, and obtain the first video stream representing the spatial video, shooting pose parameters, and depth image sequence. The playback device A20 is used to obtain the user's viewing posture and determine the matching shooting posture in the shooting path that matches the viewing posture; The playback end A20 is used to determine the first video frame corresponding to the matching shooting pose based on the first video stream, and to determine the first depth image in the depth image sequence corresponding to the first video frame; The playback device A20 is used to generate a target image with the user's viewing perspective based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, and then displays the target image.

[0145] The image processing system provided in this application, employing the image processing method described in the above embodiments, can dynamically adjust the playback angle during video playback. Compared with the prior art, the beneficial effects of the image processing system provided in this application are the same as those of the image processing method described in the above embodiments, and other technical features of the image processing system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0146] Furthermore, this application provides an electronic device, which includes: a shooting end, a playback end, at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the image processing method in the above embodiment one.

[0147] The following is for reference. Figure 5 The figure illustrates a structural diagram of an electronic device suitable for implementing embodiments of this application. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The devices shown in the figure are merely examples and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0148] The electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for device operation. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0149] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0150] The electronic device provided in this application, employing the image processing method described in the above embodiments, can solve the technical problem of dynamically adjusting the playback angle when playing video. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the image processing method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0151] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0152] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0153] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the image processing method described in the above embodiments.

[0154] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0155] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0156] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, enable the electronic device to perform the steps of the aforementioned image processing method.

[0157] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0159] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0160] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for performing the above-described image processing method, and can solve the technical problem of how to dynamically adjust the playback perspective when playing video. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the image processing method provided in the above embodiments, and will not be repeated here.

[0161] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the image processing method described above.

[0162] The computer program product provided in this application solves the technical problem of how to dynamically adjust the playback perspective when playing video. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the image processing method provided in the above embodiments, and will not be repeated here.

[0163] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. An image processing method, characterized in that, The image processing method is applied to the playback device and includes the following steps: Acquire a multimedia file, parse the multimedia file, and obtain a first video stream representing spatial video, shooting pose parameters, and a depth image sequence, wherein the shooting pose parameters include a shooting path containing the shooting pose, and the first video stream includes a left-eye video stream and / or a right-eye video stream. Obtain the user's viewing posture and determine the matching shooting posture in the shooting path that matches the viewing posture; Based on the first video stream, determine the first video frame corresponding to the matched shooting pose, and determine the first depth image in the depth image sequence corresponding to the first video frame; Based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, a target image with the playback perspective as the user's viewing perspective is generated, and the target image is displayed.

2. The image processing method as described in claim 1, characterized in that, The step of determining the matching shooting pose in the shooting path that matches the viewing pose includes: Coordinate alignment is performed between the first coordinate system corresponding to the viewing pose and the second coordinate system corresponding to the shooting pose in the shooting path; In response to the first and second coordinate systems after coordinate alignment processing, the shooting pose that matches the viewing pose among the shooting poses is determined as the matched shooting pose.

3. The image processing method as described in claim 2, characterized in that, The shooting pose includes the shooting position in the first coordinate system, and the viewing pose includes the viewing position in the second coordinate system. The step of determining the shooting pose that matches the viewing pose among the shooting poses is the matching shooting pose, including: For each shooting position, a first coordinate of the shooting position and a second coordinate of the viewing pose are determined in the same coordinate system, and the Euclidean distance between the shooting position and the viewing position is determined based on the first coordinate and the second coordinate. Determine the smallest Euclidean distance among the Euclidean distances corresponding to each shooting pose, and use the shooting pose corresponding to the smallest Euclidean distance as the matching shooting pose.

4. The image processing method as described in claim 2, characterized in that, The step of determining the shooting pose that matches the viewing pose among the shooting poses as the matching shooting pose further includes: By connecting the various shooting poses, a total arc length image representing the total arc length corresponding to the shooting path is obtained; Determine the initial viewing pose of the user when they begin watching the spatial video corresponding to the first video stream, and determine the displacement information of the initial viewing pose as it moves to the viewing position. Determine the mapping scaling factor between the first coordinate system and the second coordinate system, and map the displacement information into a target arc length image based on the mapping scaling factor, wherein the target arc length image represents the target arc length that is in the same coordinate system as the total arc length and has the same starting point and arc length direction; Based on the total arc length image and the target arc length image, determine the total arc length position that matches the endpoint position of the target arc length, and use the shooting pose corresponding to the matching total arc length position as the matching shooting pose.

5. The image processing method as described in claim 1, characterized in that, The step of generating a target image with a playback perspective that is the user's viewing perspective based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image includes: Using a preset forward mapping model, the first depth image is converted into a second depth image in the world coordinate system based on the matching shooting pose; Based on the viewing pose, the second depth image is converted into a third depth image in the first coordinate system corresponding to the viewing pose, and the playback view is the user's viewing perspective. The color parameters of the third depth image are updated based on the color parameters of the RGB image corresponding to the first video frame to obtain an initial composite image, and the target image is determined based on the initial composite image.

6. The image processing method as described in claim 5, characterized in that, The step of determining the target image based on the initial synthesized image includes: Detect whether there are hole pixels in the initial synthesized image, wherein the hole pixels are pixels that need to be filled with color; If there are empty pixels, then based on the matched shooting pose and the empty pixels, the ray direction of the ray emitted from the shooting end that passes through the empty pixels is determined, and the ray position reached by the ray along the ray direction after a preset distance is determined. For each candidate keyframe, the ray position is mapped to the candidate keyframe to obtain the matching pixel in the candidate keyframe corresponding to the ray position, wherein the candidate keyframe is a video frame that is associated with the first video frame; Determine the fourth depth image in the depth image sequence that corresponds to the candidate keyframe, and determine the reprojection error of the candidate keyframe based on the fourth depth image, the matching pixel, and the ray position; Based on the color parameters of the matching pixels in the candidate keyframe corresponding to the minimum reprojection error, the holed pixels are filled with color to obtain the target image.

7. An image processing method, characterized in that, The image processing method is applied to the camera and includes the following steps: Acquire a first video stream representing spatial video, wherein the first video stream includes a left-eye video stream and / or a right-eye video stream; Determine the shooting pose and depth image corresponding to each video frame in the first video stream, and construct a shooting path based on the shooting pose; The first video stream, shooting pose parameters, and depth image sequence are encapsulated to obtain a multimedia file, wherein the shooting pose parameters include the shooting path, and the depth image sequence includes the depth image; The multimodal file is sent to the playback terminal. After parsing the multimedia file, the playback terminal determines the matching shooting pose in the shooting path that matches the user's viewing pose, and determines the first video frame and the first depth image corresponding to the matching shooting pose in the first video stream and the depth image sequence, respectively. The playback terminal then displays the target image generated based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, with the playback perspective being the user's viewing perspective.

8. An image processing system, characterized in that, The image processing system includes a shooting end and a playback end. The camera is used to acquire a first video stream representing spatial video, wherein the first video stream includes a left-eye video stream and / or a right-eye video stream; The camera is used to determine the shooting pose and depth image corresponding to each video frame in the first video stream, and to construct a shooting path based on the shooting pose. The camera is used to encapsulate the first video stream, shooting pose parameters, and depth image sequence to obtain a multimedia file, wherein the shooting pose parameters include the shooting path, and the depth image sequence includes the depth image; The shooting end is used to send the multi-modal file to the playback end; The playback terminal is used to acquire multimedia files, parse the multimedia files, and obtain a first video stream representing spatial video, shooting pose parameters, and depth image sequence; The playback device is used to obtain the user's viewing posture and determine the matching shooting posture in the shooting path that matches the viewing posture; The playback terminal is used to determine, based on the first video stream, a first video frame corresponding to the matched shooting pose, and to determine, in the depth image sequence, a first depth image corresponding to the first video frame; The playback terminal is used to generate a target image with the user's viewing perspective based on the viewing pose, the matching shooting pose, the first video frame, and the first depth image, and to display the target image.

9. An electronic device, characterized in that, The electronic device includes: a shooting end, a playback end, a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image processing method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the image processing method as described in any one of claims 1 to 7.