Video Processing Method, Apparatus, Electronic Device, and Storage Medium
By mapping VR videos into two-dimensional images and synthesizing them based on the center point of the significance area, two-dimensional videos of multiple viewing angles are generated, which solves the problem of stuttering playback in devices without VR playback functions, and realizes accurate segmentation and smooth playback of multi-view two-dimensional videos.
Patent Information
- Application Number
- CN202310253010.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-03-09
AI Technical Summary
VR videos have lag problems when playing in electronic devices that do not have VR playback functions, and the prior art has failed to effectively solve the low accuracy of multi-view 2D video segmentation.
By mapping VR video into a two-dimensional image, feature analysis is performed to determine the significance region, and multi-view images are sliced and merged based on the center point of the significance region to generate two-dimensional videos of multiple perspectives.
It ensures smooth playback of two-dimensional videos of multiple perspectives in electronic devices that do not have VR playback function, avoiding the same object or character being divided into two-dimensional videos of different perspectives, and improving the accuracy of segmentation.
Smart Images

Figure CN116320350B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of video processing, and in particular, to a video processing method, apparatus, electronic device, and storage medium. Background Art
[0002] Virtual Reality (VR) videos, also known as panoramic videos, spherical videos, 360-degree videos, such as panoramic videos with a horizontal angle of 360° * a vertical angle of 360°. The playback of VR videos depends on professional hardware devices. For example, VR players, VR glasses, etc. Since the data volume of VR videos is relatively large, when playing with an electronic device that does not have the VR playback function, problems such as stuttering during playback may occur. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a video processing method, apparatus, electronic device, and storage medium to solve the problem of playing VR videos on an electronic device that does not have the VR playback function.
[0004] According to a first aspect, an embodiment of the present disclosure provides a video processing method, including:
[0005] Obtaining a target virtual reality video, and mapping video frames in the target virtual reality video into two-dimensional images;
[0006] Performing feature analysis on the two-dimensional images to determine significant regions in the two-dimensional images;
[0007] Based on the center points of the significant regions, performing segmentation of the two-dimensional images into images of multiple perspectives to obtain images of the multiple perspectives;
[0008] Merging the images of each perspective to obtain two-dimensional videos of each perspective.
[0009] According to a second aspect, an embodiment of the present disclosure further provides a video processing apparatus, including:
[0010] An obtaining module, configured to obtain a target virtual reality video, and map each video frame in the target virtual reality video into a two-dimensional image;
[0011] An analysis module, configured to perform feature analysis on the two-dimensional images to determine significant regions in the two-dimensional images;
[0012] A segmentation module, configured to perform segmentation of the two-dimensional images into images of multiple perspectives based on the center points of the significant regions to obtain images of the multiple perspectives;
[0013] A merging module, configured to merge the images of each perspective to obtain two-dimensional videos of each perspective.
[0014] According to a third aspect, an embodiment of the present disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the video processing method described in the first aspect or any one of the embodiments of the first aspect.
[0015] According to a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to execute the video processing method described in the first aspect or any one of the embodiments of the first aspect.
[0016] The video processing method provided by the embodiment of the present disclosure maps video frames in a target VR video mapping into two-dimensional images, obtains a salient region in the two-dimensional images through feature analysis, and performs segmentation on the two-dimensional images based on the center points of the salient regions, which can avoid splitting the same object or person into two-dimensional videos of different perspectives, ensure the accuracy of the obtained two-dimensional videos of each perspective, and thus enable smooth playback of two-dimensional videos of multiple perspectives on an electronic device without VR playback function. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a flowchart of the video processing method according to an embodiment of the present disclosure;
[0019] Figure 2 is a flowchart of the video processing method according to an embodiment of the present disclosure;
[0020] Figure 3 is a flowchart of the video processing method according to an embodiment of the present disclosure;
[0021] Figure 4 is a block diagram of the structure of the video processing device according to an embodiment of the present disclosure;
[0022] Figure 5 is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0024] When playing a VR video using an electronic device that does not have the VR video playback function, it is usually necessary to convert the VR video into multiple 2D videos as a video collection, so that it can be played on an electronic device that does not have the VR video playback function. However, in the related art, the multi-view 2D video segmentation of the VR video is not performed in combination with the video content information, which easily splits the same object or task into two videos, resulting in a low accuracy of 2D video segmentation.
[0025] Based on this, the video processing method provided by the embodiments of the present disclosure is segmented based on the center points of the salient regions, avoiding splitting the same object or task into different 2D videos. Further, multiple 2D videos can be generated according to the video content as a video collection, so that multiple-view viewing can be experienced on an electronic device that does not have the VR video playback function.
[0026] According to the embodiments of the present disclosure, an embodiment of a video processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0027] In this embodiment, a video processing method is provided, which can be used in electronic devices such as mobile phones, tablet computers, servers, etc. Figure 1 is a flowchart of the video processing method according to the embodiments of the present disclosure, as Figure 1 shown, the process includes the following steps:
[0028] S11, obtain a target virtual reality video, and map the video frames in the target virtual reality video into two-dimensional images.
[0029] The target VR video can be obtained from a VR video server through interaction with the VR video server; it can also be stored locally; or, it can also be obtained from a third-party device after communicating with the third-party device, etc. The acquisition method of the target VR video is not limited here and is specifically set according to actual needs.
[0030] After obtaining the target VR video, map the video frames in the target VR video to obtain two-dimensional images, that is, each video frame is mapped to a two-dimensional image. For VR video mapping, it means representing a spherical panoramic video as a planar video suitable for compression coding. Among them, there are various mapping models for mapping VR video into two-dimensional images, including but not limited to equidistant cylindrical mapping, hexahedron mapping, equiangular cube mapping, octahedron mapping, dodecahedron mapping, and so on.
[0031] Since the image content of equidistant cylindrical mapping is complete without cropping, it is more suitable for analyzing the image content. Based on this, in the embodiments of the present disclosure, for the obtained target VR video, the equidistant cylindrical mapping method is used to map each video frame in the VR video to obtain two-dimensional images corresponding one by one to the video frames. Specifically, first decode each video frame and convert it into an image of equidistant cylindrical mapping, and record the width w of the single-frame image. To reduce the amount of calculation, scale the image proportionally to a unified width, such as 960mm, and record the scaling ratio parameter scale_ratio = w / 960. Among them, since the aspect ratio of equidistant cylindrical mapping is 2:1, after the scaling process, a two-dimensional image of 960 * 480 size is obtained.
[0032] S12. Perform feature analysis on the two-dimensional image to determine the significant region in the two-dimensional image.
[0033] For each of the obtained two-dimensional images, perform feature analysis on it to obtain the significant region in the two-dimensional image. Among them, the methods for determining the significant region include but are not limited to visual attention models based on Gaussian pyramid fusion of image color, brightness, and orientation features, significance detection models based on the frequency domain, significance detection algorithms based on the Euclidean distance between pixel vectors in the Lab color space and the average pixel vector, or color contrast algorithms based on the global color histogram, and so on. No limitation is imposed on the method for determining the significant region here, and it can be specifically set according to actual needs.
[0034] The significant region is the region of interest in the two-dimensional image. For example, a human body or an object, etc. There can be one or more significant regions in the two-dimensional image. If there are multiple significant regions, multi-view 2D video segmentation can be performed for each significant region respectively to provide it to the user for selection; or, after obtaining multiple significant regions, an interactive interface can be provided to provide it to the user for selection before segmentation to determine one significant region, and then multi-view 2D video segmentation is performed based on the significant region selected by the user. Of course, for the case where there are multiple significant regions, other methods can also be used for processing. No limitation is imposed here, as long as it is ensured that subsequent multi-view segmentation of the two-dimensional image is based on the center point of the significant region.
[0035] S13. Based on the center point of the salient region, the two-dimensional image is segmented into images from multiple perspectives to obtain images from multiple perspectives.
[0036] The salient region is the region of interest, generally a human body or an object, etc. For each two-dimensional image, it is necessary to perform multi-perspective segmentation with the center point of the salient region as the origin to obtain images from multiple perspectives. Among them, the multi-perspectives are generally six perspectives, that is, the two-dimensional image is segmented into images from six perspectives.
[0037] S14. Merge the images from each perspective to obtain two-dimensional videos from each perspective.
[0038] As described above, each two-dimensional image obtains images from multiple perspectives. Then, in order to obtain a continuous two-dimensional video, it is necessary to merge the images from each perspective according to the chronological order. For example, for each two-dimensional image that obtains images from six perspectives, it is specifically expressed as follows:
[0039] Two-dimensional image 1 obtains image 1-1, image 1-2, image 1-3, image 1-4, image 1-5, and image 1-6;
[0040] Two-dimensional image 2 obtains image 2-1, image 2-2, image 2-3, image 2-4, image 2-5, and image 2-6;
[0041] ……
[0042] Two-dimensional image N obtains image N-1, image N-2, image N-3, image N-4, image N-5, and image N-6.
[0043] When merging the images from each perspective, the images from the same perspective are merged according to the time sequence. For example, N images from perspective 1 (i.e., image 1-1 to image N-1) are merged to obtain a two-dimensional video from perspective 1. Similar processing is performed for other perspectives to obtain two-dimensional videos from each perspective.
[0044] The video processing method provided in this embodiment maps the video frames in the target VR video mapping into two-dimensional images, obtains the salient region in the two-dimensional image through feature analysis, and performs the segmentation of the two-dimensional image based on the center point of the salient region, which can avoid splitting the same object or person into two-dimensional videos from different perspectives, ensure the accuracy of the two-dimensional videos from each perspective obtained, and thus enable smooth playback of two-dimensional videos from multiple perspectives on electronic devices without VR playback functions.
[0045] In this embodiment, a video processing method is provided, which can be used in electronic devices such as mobile phones, tablet computers, servers, etc. Figure 2is a flowchart of a video processing method according to an embodiment of the present disclosure. As Figure 2 shown, the process includes the following steps:
[0046] S21, obtain a target virtual reality video and map the video frames in the target virtual reality video into two-dimensional images.
[0047] For details, please refer to Figure 1 S11 of the embodiment shown, which will not be elaborated here.
[0048] S22, perform feature analysis on the two-dimensional image to determine the salient regions in the two-dimensional image.
[0049] For details, please refer to Figure 2 S12 of the embodiment shown, which will not be elaborated here.
[0050] S23, based on the center points of the salient regions, segment the two-dimensional image into images of multiple viewpoints to obtain images of multiple viewpoints.
[0051] Specifically, the above S23 includes:
[0052] S231, obtain the mapping center point corresponding to the two-dimensional image and the correspondence between the image coordinates of multiple viewpoints and the virtual reality spherical coordinates.
[0053] The mapping center point corresponding to the two-dimensional image is the center point during the mapping of the video frames in the VR video. If equidistant cylindrical mapping is used, the mapping center point is the equidistant cylindrical mapping center point. The image coordinates of multiple viewpoints are the correspondence with the VR spherical coordinates, that is, the correspondence between the image coordinates (u, v) of viewpoint i (i = 0, 1,..., 5) and the VR spherical coordinates (X, Y, Z) is shown in the following table:
[0054] Condition f u v |X| >= |Y| and |X| >= |Z| and X > 0 0 -Z / |X| -Y / |X| |X| >= |Y| and |X| >= |Z| and X < 0 1 Z / |X| -Y / |X| |Y| >= |X| and |Y| >= |Z| and Y > 0 2 X / |Y| Z / |Y| |Y| >= |X| and |Y| >= |Z| and Y < 0 3 X / |Y| -Z / |Y| |Z| >= |X| and |Z| >= |X| and Z > 0 4 X / |Z| -Y / |Z| |Z| >= |X| and |Z| >= |Y| and Z < 0 5 -X / |Z| -Y / |Z|
[0055] S232, use the mapping center point to convert the center point of the salient region into longitude and latitude coordinates.
[0056] Taking the mapping center point as the origin, convert the center point (cx, cy) of the salient region obtained in S22 above into longitude and latitude coordinates (phi_c, theta_c). The coordinate conversion method is inverse to the calculation method of converting the video frames in the VR video into two-dimensional images in the above text. For example, if the mapping matrix used for video frame mapping is E, then the inverse matrix of E is used here for coordinate conversion to obtain the longitude and latitude coordinates.
[0057] S233, based on the longitude and latitude coordinates and the correspondence, segment the two-dimensional image into images of multiple viewpoints to obtain images of multiple viewpoints.
[0058] The VR spherical coordinates (X, Y, Z) are represented by the following formula:
[0059] X = cos(theta + theta_c) * cos(phi + phi_c);
[0060] Y = sin(theta + theta_c);
[0061] Z = -cos(theta + theta_c) * sin(phi + phi_c).
[0062] In the formula, theta and phi are the spherical pixel longitude and latitude coordinates when mapped to the hexahedron pixel positions.
[0063] Through the above formula and the corresponding relationship described above, images from multiple perspectives can be obtained.
[0064] In some embodiments, the above S233 includes:
[0065] (1) Reproject the two-dimensional image based on the center point of the salient region to obtain the reprojected two-dimensional image.
[0066] (2) Segment the reprojected two-dimensional image into images from multiple perspectives based on the longitude and latitude coordinates and the corresponding relationship to obtain images from multiple perspectives.
[0067] Segmenting the two-dimensional image into images from multiple perspectives is specifically to move the salient region to the origin of the spherical coordinates of the VR video and perform segmentation from multiple perspectives from this origin, thereby obtaining images from multiple perspectives. Specifically, reproject the two-dimensional image based on the center point of the salient region so that the salient region is located at the center position of the front view, thereby obtaining the reprojected two-dimensional image. Then, based on the longitude and latitude coordinates and the corresponding relationship obtained from the above S232, segment the reprojected two-dimensional image into images from multiple perspectives, thereby obtaining images from multiple perspectives.
[0068] Reprojecting the two-dimensional image based on the center point of the salient region so that the salient region is located at the center position of the front view further ensures the image effect of multi-perspective segmentation.
[0069] S24, Merge the images from each perspective to obtain the two-dimensional video of each perspective.
[0070] The images from each perspective include the identifiers of each target. Based on this, the above S24 includes: merging the images from each perspective based on the identifiers of the targets, so as to merge the same target in the two-dimensional video of the same perspective, and obtaining the two-dimensional videos of each perspective. The merging of the same target can also be regarded as tracking the target, so that the same target appears in the same perspective. Merging the images from multiple perspectives using the identifiers of each target ensures that the same target exists in the two-dimensional videos of the same perspective, thereby realizing the tracking of the target and ensuring the continuity of the target movement.
[0071] The video processing method provided in this embodiment divides the two-dimensional image into multi-perspective images based on the longitude and latitude coordinates corresponding to the center point of the significant region, so as to ensure that the significant region exists in the images of the same perspective.
[0072] In some embodiments, the above method further includes:
[0073] (1) For the two-dimensional videos of each perspective, encoding is performed using the same encoding parameters to obtain the encoding results of each perspective.
[0074] (2) Encapsulating the encoding results using the same encapsulation format to obtain the target two-dimensional videos of each perspective.
[0075] After obtaining the two-dimensional videos of multiple perspectives, encoding and packaging the multi-perspective 2D videos so that they can be played on an electronic device. Specifically, for the two-dimensional videos of each perspective, encoding processing is performed using the same encoding parameters, and the encoding results are encapsulated using the same encapsulation format to obtain the target two-dimensional videos of each perspective, and the target two-dimensional videos can be played on an electronic device without VR playback function.
[0076] Processing the two-dimensional videos of each perspective using the same encoding parameters and encapsulation format can ensure smooth switching when the electronic device switches between videos of different perspectives for playback.
[0077] In this embodiment, a video processing method is provided, which can be used in electronic devices such as mobile phones, tablet computers, servers, etc. Figure 3 It is a flowchart of the video processing method according to the embodiment of the present disclosure, as Figure 3 shown, and this process includes the following steps:
[0078] S31, obtaining a target virtual reality video and mapping the video frames in the target virtual reality video into two-dimensional images.
[0079] For details, please refer to Figure 1 S11 of the embodiment shown, which will not be elaborated here.
[0080] S32. Perform feature analysis on the two-dimensional image to determine the significant regions in the two-dimensional image.
[0081] Specifically, the above S32 includes:
[0082] S321. Perform image feature analysis on the two-dimensional image to determine the image feature map of the two-dimensional image.
[0083] The image features used for image feature analysis of the two-dimensional image include, but are not limited to, one or more of color, brightness, texture, etc. The specific types and quantities are set according to actual needs, and no limitation is made thereto herein.
[0084] If the image features used for image analysis are multiple, corresponding image feature vectors are obtained for each image feature, and the image feature vectors corresponding to the multiple image features are concatenated to obtain the image feature map of the two-dimensional image.
[0085] In some embodiments, the above S321 includes:
[0086] (1) Divide the two-dimensional image to obtain a preset number of image regions.
[0087] (2) For each image region, use the pixel values of the pixel points in the image region to obtain at least one first histogram feature of the image region.
[0088] (3) Connect at least one first histogram feature of each image region to obtain at least one second histogram feature of the two-dimensional image, and the image feature map includes at least one second histogram feature.
[0089] The preset number is set according to actual needs. For example, it can be 32 or 64, etc., and no limitation is made thereto herein. For example, if the two-dimensional image is an image frame of 960*480, it is divided into 450 small regions according to 32*32, that is, 450 image regions are obtained.
[0090] For each 32*32 image region, convert the RGB color space to the Y luminance space, and then use the Y luminance space to calculate at least one first histogram feature of the image region. The first histogram features include, but are not limited to, gray histogram features and LBP (Local Binary Pattern) histogram features, etc., which are specifically set according to actual needs.
[0091] Take the grayscale histogram feature and the LBP histogram feature as examples. Calculate the grayscale histogram feature hist1: The range of Y brightness values is 0 to 255. Divide this range into 8 sub-ranges on average, that is, 0 to 31, 32 to 63,..., 224 to 255. Count the number of brightness values in this area within these 8 sub-ranges to form a 1*8 vector. Normalize this vector to obtain the grayscale histogram feature hist1 of a single 32*32 image area;
[0092] Calculate the LBP histogram feature hist2: For each pixel in the 32*32 image area, compare its grayscale value with the grayscale values of the adjacent 8 pixels. If the surrounding pixel value is greater than the central pixel value, the position of this pixel point is marked as 1, otherwise it is 0; After comparing 8 points, an 8-bit binary number can be generated as the LBP value of this pixel point; Count the occurrence frequency of the LBP value of each pixel in this area to form a 1*256 vector. Normalize this vector to obtain the LBP histogram feature hist2 of a single 32*32 image area.
[0093] Connect the grayscale histogram features and LBP histogram features of each obtained 32*32 image area respectively to obtain the grayscale histogram feature vector and LBP feature vector of the entire image. For each pixel point in the image area, use the Euclidean distance to calculate the feature vector within the area and the feature vector of the entire image to obtain the feature value at the current pixel point position. Normalize the feature values of the entire image to obtain the grayscale histogram feature hist_feature and the LBP histogram feature lbp_feature. Among them, the grayscale histogram feature hist_feature and the LBP histogram feature lbp_feature are the second histogram features mentioned above.
[0094] Before performing image feature analysis, first divide the two-dimensional image, and then perform image feature extraction for each image area. Since the size of the image area is relatively small compared to the two-dimensional image, performing image feature extraction on a smaller area can focus on the detailed features and improve the accuracy of the obtained image feature map.
[0095] S322, perform face detection on the two-dimensional image to determine the face detection feature map of the two-dimensional image.
[0096] Face detection methods include but are not limited to recognition algorithms based on facial feature points, recognition algorithms based on the entire facial image, template-based recognition algorithms, and algorithms using neural networks for recognition, etc. When performing face detection on the two-dimensional image, obtain the face detection feature map face_feature. Specifically, first generate a feature map with all values of 960*480 being 0, and set the values within the detected face position area to 1 to obtain the face detection feature map.
[0097] S323. Determine the saliency map of the two-dimensional image based on the fusion result of the image feature map and the face detection feature map.
[0098] The image feature map includes at least one second histogram feature. When fusing at least one second histogram feature with the face detection feature map, a weighted method can be used for fusion. For example, each second histogram feature and the face detection feature map have corresponding weights, and the sum of all weights is 1. Taking the second histogram features as the grayscale histogram feature and the LBP histogram feature as an example, the saliency map of the two-dimensional image saliency_map is expressed by the following formula:
[0099] saliency_map = alpha * hist_feature + belta * lbp_feature + sigma * face_feature;
[0100] In the formula, alpha + belta + sigma = 1.
[0101] S324. Determine the salient region in the two-dimensional image based on the numerical values at various positions in the saliency map.
[0102] After obtaining the saliency map, compare the values at various positions in the saliency map, and determine the region within a preset range where the position with the highest saliency value is located as the salient region in the two-dimensional image.
[0103] In some embodiments, the above S324 includes:
[0104] (1) Divide the saliency map to obtain a preset number of saliency submaps.
[0105] (2) Statistically analyze the numerical values at various positions in the saliency submaps to obtain the saliency of the saliency submaps.
[0106] (3) Compare the saliencies of the saliency submaps to obtain the target saliency submap, so as to determine the salient region in the two-dimensional image.
[0107] Divide the saliency map into saliency submaps of size 32 * 32 to obtain a preset number of saliency submaps. Sum the numerical values at various positions in the saliency submaps to obtain the saliency of each saliency submap. Then compare the saliencies of each saliency submap, and determine the saliency submap with the highest saliency as the target saliency submap. The region corresponding to this target saliency submap in the two-dimensional image is the salient region. Among them, the center coordinates of the salient region are the center coordinates of the two-dimensional image.
[0108] First, divide the saliency map into sub - maps, and the number of divisions is the same as that for dividing the two - dimensional image in the above text. Since the saliency map is obtained based on the feature analysis of the image regions, using the same number of division methods can ensure the consistency of features and improve the reliability of the obtained saliency regions.
[0109] S33. Based on the center points of the saliency regions, split the two - dimensional image into images from multiple perspectives to obtain images from multiple perspectives.
[0110] For details, please refer to Figure 2 S23 of the illustrated embodiment, which will not be elaborated here.
[0111] S34. Merge the images from each perspective to obtain two - dimensional videos from each perspective.
[0112] For details, please refer to Figure 1 S24 of the illustrated embodiment, which will not be elaborated here.
[0113] The video processing method provided in this embodiment, based on image features and combined with the results of face detection, obtains the saliency map of the two - dimensional image, enabling the consideration of face features in the saliency map, thereby further improving the accuracy of the obtained saliency regions.
[0114] As an optional application example of the embodiments of the present disclosure, this method is applied to a VR video server. The VR video production end uploads the produced VR video to the VR video server. When a playback terminal needs to play a VR video, it sends a video acquisition instruction to the VR video server and obtains the VR video from the VR video server. In this process, some playback terminals have the function of playing VR videos, while some playback terminals do not have the function of playing VR videos. In order to be able to view the content of the VR video on playback terminals that do not have the function of playing VR videos, the video processing method described in the embodiments of the present disclosure is applied in the VR video server to process the VR video into target 2D videos of each perspective. Among them, this process can be processed after receiving the video acquisition instruction, or all VR videos can be converted into 2D videos of multiple perspectives. Taking the case of processing after receiving the video acquisition instruction as an example, the video acquisition instruction carries an identifier of the playback terminal, and this identifier is used to indicate whether the playback terminal has the function of playing VR videos. If it does not have the function of playing VR videos, the VR server first extracts the target VR video, and then uses the video processing method in the embodiments of the present disclosure to process it into target 2D videos of multiple perspectives, and sends the target 2D videos of multiple perspectives to the playback terminal. Through the interaction with the playback terminal, the user can select the target 2D video of the corresponding perspective for playback. Of course, since the VR server sends the target 2D videos of multiple perspectives to the playback terminal, the user can also switch the playback perspective to view the target 2D videos of different perspectives.
[0115] As another optional application example of the embodiments of the present disclosure, this method is applied to a playback terminal. When applied in a playback terminal, this method can be encapsulated as a plugin of a video playback application, and by installing this plugin, it is possible to watch 2D videos of different perspectives in the video playback application on a playback device that does not have the function of playing VR videos. Or, this method can also be encapsulated as an independent application, and this application is installed in the playback terminal to convert the VR video into 2D videos of multiple perspectives. Taking the plugin of the video playback application as an example, if a user wants to play a VR video in the video playback application, but the playback terminal does not have the function of playing VR videos, then this plugin is triggered to work, convert the VR video into 2D videos of multiple perspectives, and present the multiple perspectives on the interface for the user to select. Through the interaction with this interface, the user selects a target perspective, and correspondingly, the video playback application plays the 2D video of the target perspective.
[0116] In this embodiment, a video processing device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0117] This embodiment provides a video processing device, as Figure 4 shown, including:
[0118] An acquisition module 41, configured to acquire a target virtual reality video and map video frames in the target virtual reality video into two-dimensional images;
[0119] An analysis module 42, configured to perform feature analysis on the two-dimensional image to determine a salient region in the two-dimensional image;
[0120] A segmentation module 43, configured to segment the two-dimensional image into images of multiple viewpoints based on the center point of the salient region to obtain images of the multiple viewpoints;
[0121] A merging module 44, configured to merge images of each of the viewpoints to obtain two-dimensional videos of each of the viewpoints.
[0122] In some implementation manners, the analysis module 42 includes:
[0123] An analysis unit, configured to perform image feature analysis on the two-dimensional image to determine an image feature map of the two-dimensional image;
[0124] A face detection unit, configured to perform face detection on the two-dimensional image to determine a face detection feature map of the two-dimensional image;
[0125] A fusion unit, configured to determine a saliency map of the two-dimensional image based on a fusion result of the image feature map and the face detection feature map;
[0126] A determination unit, configured to determine a salient region in the two-dimensional image based on the numerical magnitudes of positions in the saliency map.
[0127] In some implementation manners, the analysis unit includes:
[0128] A first partitioning subunit, configured to partition the two-dimensional image to obtain a preset number of image regions;
[0129] A first determination subunit, configured to, for each of the image regions, obtain at least one first histogram feature of the image region by using pixel values of pixel points in the image region;
[0130] A second determination subunit, configured to connect at least one first histogram feature of each of the image regions to obtain at least one second histogram feature of the two-dimensional image, where the image feature map includes the at least one second histogram feature.
[0131] In some embodiments, the determination unit includes:
[0132] A second division subunit, configured to divide the saliency map to obtain a preset number of saliency sub-maps;
[0133] A statistics subunit, configured to statistically analyze the values at each position in the saliency sub-maps to obtain the saliency of the saliency sub-maps;
[0134] A comparison subunit, configured to compare the saliency of the saliency sub-maps to obtain a target saliency sub-map, so as to determine a saliency region in the two-dimensional image.
[0135] In some embodiments, the splitting module 43 includes:
[0136] An obtaining unit, configured to obtain a mapping center point corresponding to the two-dimensional image and a correspondence between the image coordinates of the multiple viewpoints and virtual reality spherical coordinates;
[0137] A conversion unit, configured to convert the center point of the saliency region into longitude and latitude coordinates by using the mapping center point;
[0138] A splitting unit, configured to split the two-dimensional image into images of multiple viewpoints based on the longitude and latitude coordinates and the correspondence, to obtain the images of the multiple viewpoints.
[0139] In some embodiments, the splitting unit includes:
[0140] A reprojection subunit, configured to perform reprojection on the two-dimensional image based on the center point of the saliency region to obtain a reprojected two-dimensional image;
[0141] A splitting subunit, configured to split the reprojected two-dimensional image into images of multiple viewpoints based on the longitude and latitude coordinates and the correspondence, to obtain the images of the multiple viewpoints.
[0142] In some embodiments, each of the images of the viewpoints includes an identifier of each target, and the merging module 44 includes:
[0143] A merging unit, configured to merge the images of the viewpoints based on the identifier of the target, so as to merge the same target in a two-dimensional video of the same viewpoint, to obtain two-dimensional videos of the multiple viewpoints.
[0144] In some embodiments, the device further includes:
[0145] An encoding module, configured to encode the two-dimensional videos of each of the perspectives using the same encoding parameters to obtain the encoding results of each of the perspectives;
[0146] A packaging module, configured to package the encoding results using the same packaging format to obtain the target two-dimensional videos of each of the perspectives.
[0147] The video processing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0148] The further function descriptions of the above-mentioned respective modules are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0149] This embodiment of the present disclosure further provides an electronic device having the above Figure 4 shown video processing device.
[0150] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an alternative embodiment of the present disclosure. As Figure 5 shown, the electronic device may include: at least one processor 51, such as a CPU (Central Processing Unit), at least one communication interface 53, a memory 54, and at least one communication bus 52. Among them, the communication bus 52 is used to implement connection communication between these components. Among them, the communication interface 53 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the communication interface 53 may further include a standard wired interface and a wireless interface. The memory 54 may be a high-speed RAM memory (Random Access Memory, volatile random access memory), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 54 may further be at least one storage device located far from the aforementioned processor 51. Among them, the processor 51 may be combined with Figure 4 the device described, the memory 54 stores an application program, and the processor 51 calls the program code stored in the memory 54 to be used to execute any of the above method steps.
[0151] Among them, the communication bus 52 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 52 can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, Figure 5 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0152] Among them, the memory 54 can include volatile memory, such as random-access memory (RAM); the memory can also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 54 can also include a combination of the above types of memory.
[0153] Among them, the processor 51 can be a Central Processing Unit (CPU), a Network Processor (NP), or a combination of a CPU and an NP.
[0154] Among them, the processor 51 can further include a hardware chip. The above-mentioned hardware chip can be an Application-Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), or a combination thereof. The above-mentioned PLD can be a Complex Programmable Logic Device (CPLD), a Field-Programmable Gate Array (FPGA), a Generic Array Logic (GAL), or any combination thereof.
[0155] Optionally, the memory 54 is also used to store program instructions. The processor 51 can call the program instructions to implement the video processing method as shown in any embodiment of the present application.
[0156] Embodiments of the present disclosure also provide a non-transitory computer storage medium storing computer-executable instructions that can execute the video processing method in any of the above method embodiments. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.
[0157] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for method embodiments, since they are basically similar to device and system embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the device and system embodiments.
[0158] It can be understood that in the specific implementation of the present disclosure, when it comes to obtaining data related to VR videos, etc., when the above embodiments of the present disclosure are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0159] Although the embodiments of the present disclosure are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A video processing method, characterized in that including: obtaining a target virtual reality video and mapping video frames in the target virtual reality video into two-dimensional images; performing feature analysis on the two-dimensional images to determine significant regions in the two-dimensional images; performing segmentation of the two-dimensional images into images of multiple perspectives based on the center points of the significant regions to obtain the images of the multiple perspectives; merging the images of each of the perspectives to obtain two-dimensional videos of each of the perspectives; wherein the performing segmentation of the two-dimensional images into images of multiple perspectives based on the center points of the significant regions to obtain the images of the multiple perspectives includes: obtaining a mapping center point corresponding to the two-dimensional image and the correspondence between the image coordinates of the multiple perspectives and the spherical coordinates of the virtual reality; converting the center point of the significant region into longitude and latitude coordinates by using the mapping center point; performing reprojection on the two-dimensional image based on the center point of the significant region to obtain a reprojected two-dimensional image; performing segmentation of the reprojected two-dimensional image into images of multiple perspectives based on the longitude and latitude coordinates and the correspondence to obtain the images of the multiple perspectives.
2. The method according to claim 1, characterized in that, The performing feature analysis on the two-dimensional images to determine significant regions in the two-dimensional images includes: performing image feature analysis on the two-dimensional images to determine an image feature map of the two-dimensional images; performing face detection on the two-dimensional images to determine a face detection feature map of the two-dimensional images; determining a saliency map of the two-dimensional images based on the fusion result of the image feature map and the face detection feature map; determining the significant regions in the two-dimensional images based on the numerical magnitudes at various positions in the saliency map.
3. The method according to claim 2, wherein The performing image feature analysis on the two-dimensional images to determine an image feature map of the two-dimensional images includes: dividing the two-dimensional images to obtain a preset number of image regions; for each of the image regions, obtaining at least one first histogram feature of the image region by using the pixel values of the pixel points in the image region; connecting at least one first histogram feature of each of the image regions to obtain at least one second histogram feature of the two-dimensional image, and the image feature map includes the at least one second histogram feature.
4. The method according to claim 2, characterized in that The determining the significant regions in the two-dimensional images based on the numerical magnitudes at various positions in the saliency map includes: dividing the saliency map to obtain a preset number of saliency sub-maps; counting the numerical values at various positions in the saliency sub-maps to obtain the saliency of the saliency sub-maps; comparing the magnitudes of the saliency of the saliency sub-maps to obtain a target saliency sub-map to determine the significant regions in the two-dimensional images.
5. The method according to claim 1, characterized in that, Each of the images of the perspectives includes identifiers of each target, and the merging the images of each of the perspectives to obtain two-dimensional videos of each of the perspectives includes: merging the images of each of the perspectives based on the identifiers of the targets to merge the same target in the two-dimensional videos of the same perspective to obtain the two-dimensional videos of each of the perspectives.
6. The method according to claim 1, wherein The method further includes: For the two-dimensional videos of each of the said perspectives, encoding is performed using the same encoding parameters to obtain the encoding results of each of the said perspectives; The encoding results are encapsulated using the same encapsulation format to obtain the target two-dimensional videos of each of the said perspectives.
7. A video processing device, characterized in that, It includes: An acquisition module, configured to acquire a target virtual reality video and map the video frames in the target virtual reality video into two-dimensional images; An analysis module, configured to perform feature analysis on the two-dimensional images to determine the salient regions in the two-dimensional images; A splitting module, configured to split the two-dimensional images into images of multiple perspectives based on the center points of the salient regions to obtain the images of the multiple perspectives; A merging module, configured to merge the images of each of the said perspectives to obtain the two-dimensional videos of each of the said perspectives; Wherein, the splitting module includes: An acquisition unit, configured to acquire the mapping center point corresponding to the two-dimensional image and the correspondence between the image coordinates of the multiple perspectives and the virtual reality spherical coordinates; A conversion unit, configured to convert the center point of the salient region into longitude and latitude coordinates using the mapping center point; A reprojection sub-unit, configured to perform reprojection on the two-dimensional image based on the center point of the salient region to obtain the reprojected two-dimensional image; A splitting sub-unit, configured to split the reprojected two-dimensional image into images of multiple perspectives based on the longitude and latitude coordinates and the correspondence to obtain the images of the multiple perspectives.
8. An electronic device, characterized in that, It includes: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the video processing method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the video processing method according to any one of claims 1-6.
Citation Information
Patent Citations
Video processing method and apparatus
CN108124193A
Panoramic video saliency detection method based on multi-channel features
CN110827193A