Immersive media data processing method, device, equipment and readable storage medium

By generating interactive feedback messages carrying business-critical fields during immersive media consumption, the interactive feedback information types are enriched, the problem of single interactive feedback information is solved, and the accuracy of media content acquisition is improved.

CN116233493BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310228608.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-09-16
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

During immersive media consumption, the interactive feedback information types between the video client and the server are relatively simple, resulting in reduced accuracy in media content acquisition.

Method used

By generating interactive feedback messages carrying business-critical fields, the interactive feedback information types are enriched, including business event information of zoom operations, switching operations, first-position interactive operations, and second-position interactive operations, thereby improving the accuracy of video clients in obtaining media content.

Benefits of technology

By enriching the types of interactive feedback information, the accuracy of media content acquisition by video clients during the interactive feedback process is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116233493B_ABST
    Figure CN116233493B_ABST
Patent Text Reader

Abstract

This application discloses an immersive media data processing method, apparatus, device, and readable storage medium. The method comprises: generating an interactive feedback message corresponding to an interactive operation on first immersive media content in response to the interactive operation; carrying a business key field in the interactive feedback message for describing the business event information indicated by the interactive operation; sending the interactive feedback message to a server, so that the server determines the business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and obtains second immersive media content for responding to the interactive operation based on the business event information indicated by the interactive operation; and receiving the second immersive media content returned by the server. This application can enrich the types of interactive feedback information and improve the accuracy of media content acquisition by video clients during the interactive feedback process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application filed with the China Patent Office on September 29, 2021, with application number 202111149860.8 and application name “Data processing method, device, equipment and readable storage medium for immersive media”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, and readable storage medium for immersive media. Background Art

[0003] Immersive media (also known as immersive media) refers to media content that can bring an immersive experience to business objects (e.g., users). Immersive media can be divided into 3DoF media, 3DoF+ media, and 6DoF media according to the degree of freedom (DoF) of business objects (e.g., users) when consuming media content.

[0004] During immersive media consumption, a video client and a server can communicate by sending interaction feedback messages. For example, a video client can send back an interaction feedback message describing user location information (e.g., user location) to the server so that the video client can receive media content returned by the server based on the user location information. However, the inventors have discovered in practice that, in existing immersive media consumption, only the interaction feedback message of user location information exists. As a result, the type of feedback information is relatively simple during a conversation between a video client and a server, thereby reducing the accuracy with which the video client can obtain media content during the interaction feedback process. Summary of the Invention

[0005] The embodiments of the present application provide an immersive media data processing method, apparatus, device, and readable storage medium, which can enrich the information types of interactive feedback and improve the accuracy of video clients in acquiring media content during the interactive feedback process.

[0006] An embodiment of the present application provides, on one hand, a method for processing immersive media data, including:

[0007] In response to an interactive operation on the first immersive media content, generating an interactive feedback message corresponding to the interactive operation; the interactive feedback message carries a business key field for describing business event information indicated by the interactive operation;

[0008] Sending the interaction feedback message to the server, so that the server determines the business event information indicated by the interaction operation based on the business key field in the interaction feedback message, and obtains the second immersive media content for responding to the interaction operation based on the business event information indicated by the interaction operation;

[0009] The second immersive media content returned by the server is received.

[0010] An embodiment of the present application provides, on one hand, a method for processing immersive media data, including:

[0011] Receiving an interaction feedback message sent by a video client; the interaction feedback message is a message corresponding to an interaction operation generated by the video client in response to an interaction operation on the first immersive media content; the interaction feedback message carries a business key field for describing business event information indicated by the interaction operation;

[0012] Determining business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and acquiring second immersive media content for responding to the interactive operation based on the business event information indicated by the interactive operation;

[0013] The second immersive media content is returned to the video client.

[0014] In one aspect, an embodiment of the present application provides an immersive media data processing device, including:

[0015] A message generation module, configured to respond to an interactive operation on the first immersive media content and generate an interactive feedback message corresponding to the interactive operation; the interactive feedback message carries a business key field for describing business event information indicated by the interactive operation;

[0016] A message sending module, configured to send the interaction feedback message to the server, so that the server determines the business event information indicated by the interaction operation based on the business key field in the interaction feedback message, and obtains the second immersive media content used to respond to the interaction operation based on the business event information indicated by the interaction operation;

[0017] The content receiving module is configured to receive the second immersive media content returned by the server.

[0018] The device further comprises:

[0019] The video request module is used to respond to a video playback operation for an immersive video in a video client, generate a playback request corresponding to the video playback operation, send the playback request to a server, so that the server obtains the first immersive media content of the immersive video based on the playback request; receive the first immersive media content returned by the server, and play the first immersive media content on the video playback interface of the video client.

[0020] Among them, the business key field includes a first key field, a second key field, a third key field and a fourth key field; the first key field is used to represent the zoom ratio when the zoom event indicated by the zoom operation is executed when the interactive operation includes a zoom operation; the second key field is used to represent the event label and event status corresponding to the switching event indicated by the switching operation when the interactive operation includes a switching operation; the third key field is used to represent the first object position information of the business object of the first immersive media content belonging to the panoramic video when the interactive operation includes a first position interactive operation; the fourth key field is used to represent the second object position information of the business object of the first immersive media content belonging to the volumetric video when the interactive operation includes a second position interactive operation.

[0021] The message generation module includes:

[0022] A first determining unit is configured to, in response to a triggering operation on the first immersive media content, determine a first information type field of the service event information indicated by the triggering operation, and record an operation timestamp of the triggering operation;

[0023] a first adding unit, configured to add the first information type field and the operation timestamp to an interaction signaling table associated with the first immersive media content, and use the first information type field added to the interaction signaling table as a service key field for describing service event information indicated by the interaction operation;

[0024] The first generating unit is configured to generate an interactive feedback message corresponding to a triggering operation based on a service key field and an operation timestamp in an interactive signaling table.

[0025] In which, when the triggering operation includes a zoom operation, the business event information indicated by the zoom operation is a zoom event, and when the field value of the first information type field corresponding to the zoom operation is the first field value, the field mapped by the first information type field with the first field value is used to represent the zoom ratio when executing the zoom event.

[0026] In which, when the triggering operation includes a switching operation, the business event information indicated by the switching operation is a switching event, and when the field value of the first information type field corresponding to the switching operation is the second field value, the field mapped by the first information type field with the second field value is used to represent the event label and event status of the switching event.

[0027] Among them, when the state value of the event state is the first state value, the event state with the first state value is used to represent that the switching event is in the event trigger state; when the state value of the event state is the second state value, the event state with the second state value is used to represent that the switching event is in the event end state.

[0028] The message generation module includes:

[0029] a second determining unit configured to, upon detecting object position information of a business object viewing the first immersive media content, use a position interaction operation directed to the object position information as an interaction operation in response to the first immersive media content; determine a second information type field of business event information indicated by the interaction operation, and record an operation timestamp of the interaction operation;

[0030] a second adding unit, configured to add the second information type field and the operation timestamp to an interaction signaling table associated with the first immersive media content, and use the second information type field added to the interaction signaling table as a service key field for describing service event information indicated by the interaction operation;

[0031] The second generating unit is configured to generate an interaction feedback message corresponding to the interaction operation based on the service key field and the operation timestamp in the interaction signaling table.

[0032] Among them, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a panoramic video, the field value of the second information type field corresponding to the object position information is a third field value, and the second information type field with the third field value includes a first type of position field, and the first type of position field is used to describe the position change information of the business object watching the first immersive media content belonging to the panoramic video.

[0033] In which, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a volumetric video, the field value of the second information type field corresponding to the object position information is a fourth field value, and the second information type field with the fourth field value includes a second-type position field, and the second-type position field is used to describe position change information of a business object viewing the first immersive media content belonging to the volumetric video.

[0034] Among them, the interactive feedback message also includes an extended description field newly added at the system layer of the video client; the extended description field includes a signaling table quantity field, a signaling table identification field, a signaling table version field, and a signaling table length field; the signaling table quantity field is used to represent the total number of interactive signaling tables included in the interactive feedback message; the signaling table identification field is used to represent the identifier of each interactive signaling table included in the interactive feedback message; the signaling table version field is used to represent the version number of each interactive signaling table; and the signaling table length field is used to represent the length of each interactive signaling table.

[0035] The interactive feedback message also includes a resource group attribute field and a resource group identification field; the resource group attribute field is used to represent the subordinate relationship between the first immersive media content and the immersive media content set included in the target resource group; the resource group identification field is used to represent the identifier of the target resource group.

[0036] Among them, when the field value of the resource group attribute field is the first attribute field value, the resource group attribute field with the first attribute field value is used to represent that the first immersive media content belongs to the immersive media content set; when the field value of the resource group attribute field is the second attribute field value, the resource group attribute field with the second attribute field value is used to represent that the first immersive media content does not belong to the immersive media content set.

[0037] In one aspect, an embodiment of the present application provides an immersive media data processing device, including:

[0038] A message receiving module, configured to receive an interaction feedback message sent by a video client; the interaction feedback message is a message corresponding to an interaction operation generated by the video client in response to an interaction operation on the first immersive media content; the interaction feedback message carries a business key field for describing business event information indicated by the interaction operation;

[0039] A content acquisition module is configured to determine the business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and acquire second immersive media content for responding to the interactive operation based on the business event information indicated by the interactive operation;

[0040] The content returning module is configured to return the second immersive media content to the video client.

[0041] In one aspect, an embodiment of the present application provides a computer device, including: a processor and a memory;

[0042] The processor is connected to a memory, wherein the memory is used to store a computer program. When the computer program is executed by the processor, the computer device executes the method provided in the embodiment of the present application.

[0043] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.

[0044] In one aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the embodiments of the present application.

[0045] In an embodiment of the present application, a video client (i.e., a decoding end) can respond to an interactive operation on a first immersive media content by generating an interactive feedback message corresponding to the interactive operation, wherein the interactive feedback message carries a business key field for describing the business event information indicated by the interactive operation. Furthermore, the video client can send the interactive feedback message to a server (i.e., an encoding end), so that the server can determine the business event information indicated by the interactive operation based on the business key field in the interactive feedback message and obtain the second immersive media content used to respond to the interactive operation based on the business event information indicated by the interactive operation. Ultimately, the video client can receive the second immersive media content returned by the server. It can be seen that in the process of interaction between the video client and the server, the video client can feedback business event information indicated by different types of interactive operations to the server. It should be understood that the interactive operations here can not only include operations related to the user's location (for example, changes in the user's location), but also include other operations for the immersive media content currently played by the video client (for example, zoom operations). Therefore, through the business key fields carried in the interactive feedback message, the video client can feedback multiple types of business event information to the server. In this way, the server can determine the immersive media content that responds to the interactive operation based on these different types of business event information, rather than relying solely on user location information, thereby enriching the information types of the interactive feedback and improving the accuracy of the video client in obtaining media content during the interactive feedback process. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 This is an architectural diagram of a panoramic video system provided by an embodiment of the present application;

[0048] Figure 2 is a schematic diagram of 3DoF provided in an embodiment of the present application;

[0049] Figure 3 This is an architectural diagram of a volumetric video system provided by an embodiment of the present application;

[0050] Figure 4 is a schematic diagram of 6DoF provided in an embodiment of the present application;

[0051] Figure 5 is a schematic diagram of 3DoF+ provided in an embodiment of the present application;

[0052] Figure 6 This is a flow chart of a data processing method for immersive media provided in an embodiment of the present application;

[0053] Figure 7 This is a flow chart of a data processing method for immersive media provided in an embodiment of the present application;

[0054] Figure 8 This is an interactive schematic diagram of an immersive media data processing method provided by an embodiment of the present application;

[0055] Figure 9 This is a schematic structural diagram of an immersive media data processing device provided in an embodiment of the present application;

[0056] Figure 10 This is a schematic structural diagram of an immersive media data processing device provided in an embodiment of the present application;

[0057] Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application;

[0058] Figure 12 It is a structural diagram of a data processing system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0060] Embodiments of the present application relate to data processing technology for immersive media. The so-called immersive media (also referred to as immersive media) refers to a media file that can provide immersive media content, so that the business objects immersed in the media content can obtain visual, auditory and other sensory experiences in the real world. Immersive media can be divided into 3DoF media, 3DoF+ media and 6DoF media according to the degree of freedom of the business objects when consuming media content. Among them, common 6DoF media include multi-perspective video and point cloud media. Immersive media content includes video content represented in various forms in three-dimensional (3-Dimension, 3D) space, such as three-dimensional video content represented in spherical form. Specifically, immersive media content can be VR (Virtual Reality) video content, panoramic video content, spherical video content, 360-degree video content or volumetric video content. In addition, the immersive media content also includes audio content synchronized with the video content represented in the three-dimensional space.

[0061] Panoramic video / imagery uses multiple cameras to capture, stitch, and map a scene. This allows for the provision of partial media content based on the viewing orientation or viewport of the user, providing up to 360 degrees of spherical video or imagery. Panoramic video / imagery is a typical form of immersive media that provides a three-degrees-of-freedom (3DoF) experience.

[0062] V3C volumetric media (visual volumetric video-based coding media) refers to immersive media that captures visual content from three-dimensional space and provides a 3DoF+ or 6DoF viewing experience. It is encoded using traditional video and contains volumetric video type tracks in the file encapsulation. Specifically, it can include multi-view video and video coding point clouds.

[0063] Multi-view video, also known as multi-viewpoint video, uses multiple camera arrays to capture a scene from multiple angles, capturing the scene's texture information (such as color) and depth information (such as spatial distance). Multi-view / multi-viewpoint video, also known as free-viewpoint / free-viewpoint video, is an immersive media that provides a six-degree-of-freedom (6DoF) experience.

[0064] Among them, a point cloud is a set of discrete points that are randomly distributed in space and express the spatial structure and surface properties of a three-dimensional object or scene. Each point in a point cloud has at least three-dimensional position information, and may also have color, material or other information depending on the application scenario. Usually, each point in a point cloud has the same number of additional attributes. Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes, and therefore have a wide range of applications, including virtual reality games, computer-aided design (CAD), geographic information systems (GIS), autonomous navigation systems (ANS), digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive telepresence, and three-dimensional reconstruction of biological tissues and organs. Among them, point clouds are mainly obtained through the following methods: computer generation, 3D laser scanning, 3D photogrammetry, etc.

[0065] See Figure 1 , Figure 1 This is an architecture diagram of a panoramic video system provided by an embodiment of the present application. Figure 1 As shown, the panoramic video system may include an encoding device (e.g., encoding device 100A) and a decoding device (e.g., decoding device 100B). The encoding device may refer to a computer device used by a provider of the panoramic video, which may be a terminal (e.g., PC (Personal Computer), smart mobile device (e.g., smartphone), etc.) or a server. The decoding device may refer to a computer device used by a user of the panoramic video, which may be a terminal (e.g., PC (Personal Computer), smart mobile device (e.g., smartphone), VR device (e.g., VR helmet, VR glasses, etc.)). The data processing process of the panoramic video includes a data processing process on the encoding device side and a data processing process on the decoding device side.

[0066] The data processing process on the encoding device side mainly includes: (1) the process of acquiring and producing the media content of the panoramic video; (2) the process of encoding and encapsulating the panoramic video file. The data processing process on the decoding device side mainly includes: (1) the process of decapsulating and decoding the panoramic video file; (2) the process of rendering the panoramic video. In addition, the transmission process of the panoramic video between the encoding device and the decoding device can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, smart media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.

[0067] The following will be combined Figure 1 , each process involved in the data processing of panoramic video is introduced in detail.

[0068] 1. Data processing on the encoding device side:

[0069] (1) The process of acquiring and producing the media content of panoramic videos.

[0070] 1) The process of acquiring the media content of panoramic video.

[0071] The media content of panoramic videos is obtained by capturing the real-world audio and visual scenes using a capture device. In one implementation, the capture device can refer to a hardware component within the encoding device, such as a microphone, camera, or sensor on a terminal. In another implementation, the capture device can also be a hardware device connected to the encoding device, such as a camera connected to a server, which provides the encoding device with a service for acquiring the media content of the panoramic video. The capture device may include, but is not limited to, audio equipment, video equipment, and sensor equipment. Audio equipment may include audio sensors, microphones, and the like. Video equipment may include conventional cameras, stereo cameras, light field cameras, and the like. Sensor equipment may include laser equipment, radar equipment, and the like. There may be multiple capture devices, deployed at specific locations in real space to simultaneously capture audio and video content from different angles within that space, with the captured audio and video content synchronized in both time and space. In embodiments of the present application, the media content in a three-dimensional space captured by capture devices deployed at specific locations to provide a three-degree-of-freedom viewing experience may be referred to as a panoramic video.

[0072] For example, Figure 1 As shown, the real-world sound-visual scene 10A can be captured by multiple audio sensors and a camera array in the encoding device 100A, or by a camera device with multiple cameras and sensors connected to the encoding device 100A. The acquisition result can be a set of digital image / video signals 10B i (ie video content) and digital audio signal 10B a (i.e. audio content). The cameras / cameras here usually cover all directions around the center point of the camera array or camera device, so panoramic video can also be called 360-degree video.

[0073] 2) The production process of the media content of panoramic video.

[0074] It should be understood that the production process of the media content of the panoramic video involved in the embodiments of the present application can be understood as the process of content production of the panoramic video. The captured audio content itself is content suitable for performing audio encoding of the panoramic video. The captured video content can only become content suitable for performing video encoding of the panoramic video after a series of production processes. The production process may include:

[0075] ① Stitching. Since the captured video content is shot from different angles, stitching means combining these videos from various angles into a complete video that reflects a 360-degree visual panorama of the real space. In other words, the stitched video is a spherical video represented in three-dimensional space. Alternatively, multiple captured images are stitched together to create a spherical image represented in three-dimensional space.

[0076] ② Rotation. This is an optional processing operation in the production process. Each video frame in the spherical video obtained by the above stitching is a spherical image based on the unit sphere of the global coordinate axis. Rotation refers to rotating the unit sphere on the global coordinate axis. The angle of rotation is used to represent the rotation angle required to convert the local coordinate axis to the global coordinate axis. Among them, the local coordinate axis of the unit sphere is the axis of the coordinate system after rotation. It should be understood that if the local coordinate axis and the global coordinate axis are the same, no rotation is required.

[0077] ③ Projection. Projection refers to the process of mapping a spliced ​​3D video (or a rotated 3D video) onto a 2D image. The resulting 2D image is called a projected image. Projection methods include, but are not limited to, latitude and longitude projection and regular hexahedron projection.

[0078] ④ Region encapsulation. The projected image can be encoded directly, or the projected image can be region encapsulated before encoding. In practice, it has been found that in the data processing process of immersive media, encoding the two-dimensional projected image after region encapsulation can greatly improve the video encoding efficiency of the immersive media. Therefore, the region encapsulation technology is widely used in the video processing process of immersive media. The so-called region encapsulation refers to the process of performing conversion processing on the projected image by region. The region encapsulation process converts the projected image into an encapsulated image. The region encapsulation process specifically includes: dividing the projected image into multiple mapping regions, and then performing conversion processing on the multiple mapping regions to obtain multiple encapsulated regions, and mapping the multiple encapsulated regions into a 2D image to obtain an encapsulated image. Among them, the mapping region refers to the region obtained by dividing the projected image before performing region encapsulation; the encapsulated region refers to the region located in the encapsulated image after performing region encapsulation. The conversion processing may include but is not limited to: mirroring, rotation, rearrangement, upsampling, downsampling, changing the resolution of the region and moving it.

[0079] For example, Figure 1 As shown, the encoding device 100A can encode the digital image / video signal 10B i The images belonging to the same time instance in are stitched, (possibly) rotated, projected and mapped onto the package image 10D.

[0080] It should be noted that after the panoramic video obtained through the above acquisition and production process is processed by the encoding device and transmitted to the decoding device for corresponding data processing, the business objects on the decoding device side can only view the 360-degree video information by performing certain actions (such as head rotation). In other words, panoramic video is an immersive media that provides three degrees of freedom. Please also refer to Figure 2 , Figure 2 3DoF is a schematic diagram of the embodiment of the present application. Figure 2 As shown, 3DoF means that the business object is fixed at the center point of a three-dimensional space, and the business object's head rotates along the X-axis, Y-axis, and Z-axis to view the image provided by the media content. In the embodiments of this application, users who consume immersive media (such as panoramic video and volumetric video) can be collectively referred to as business objects.

[0081] (2) The process of encoding and file packaging of panoramic videos.

[0082] The captured audio content can be directly audio-encoded to form an audio stream of the panoramic video. After the above-mentioned production process ①-④ (may not include ②), the encapsulated image is video-encoded to obtain a video stream of the panoramic video. The audio stream and the video stream are encapsulated in a file container according to the file format of the panoramic video (such as ISOBMFF (ISO Based Media FileFormat, a media file format based on the ISO standard)) to form a media file resource of the panoramic video. The media file resource can be a media file of the panoramic video formed by a media file or a media fragment; and according to the file format requirements of the panoramic video, the media presentation description information (MPD) is used to record the metadata of the media file resource of the panoramic video. The metadata here is a general term for information related to the presentation of the panoramic video. The metadata may include description information of the media content, description information of the window, and signaling information related to the presentation of the media content, etc. As Figure 1 As shown, the encoding device stores the media presentation description information and media file resources formed after the data processing process.

[0083] For example, Figure 1 As shown, the encoding device 100A can encode the captured digital audio signal 10B. a Perform audio encoding to obtain audio code stream 10E a At the same time, the packaged image 10D can be video-encoded to obtain a video stream 10E v Alternatively, the packaged image 10D may be image-encoded to obtain the encoded image 10E. i Then, the encoding device 100A can convert the encoded image 10E obtained after encoding into a specific media file format (such as ISOBMFF). i 、Video stream 10E v and / or audio stream 10E a Combined into a media file 10F for file playback or combined into a segment sequence 10F containing an initialization segment and multiple media segments for streaming s Among them, the media file 10F and the segment sequence 10F s All of them belong to the media file resources of the panoramic video. In addition, the file encapsulator in the encoding device 100A can also add metadata to the media file 10F or the segment sequence 10F. s For example, the metadata here may include projection information and region encapsulation information, which will help the subsequent decoding device to render the encapsulated image obtained after decoding. Subsequently, the encoding device 100A may use a specific transmission mechanism (such as DASH, SMTP) to transmit the segment sequence 10F to the encoding device 100A. sThe media file 10F is transmitted to the decoding device 100B, and the media file 10F is also transmitted to the decoding device 100B. Optionally, the decoding device 100B here can be an OMAF (Omnidirectional Media Application Format) player.

[0084] 2. Data processing on the decoding device side:

[0085] (3) The process of decapsulating and decoding panoramic video files.

[0086] The decoding device can dynamically and adaptively obtain the panoramic video's media file resources and corresponding media presentation description information from the encoding device based on recommendations from the encoding device or according to the needs of the business object on the decoding device. For example, the decoding device can determine the orientation and position of the business object based on the business object's head / eye tracking information, and then dynamically request the corresponding media file resources from the encoding device based on the determined orientation and position. The media file resources and media presentation description information are transmitted from the encoding device to the decoding device via a transport mechanism (such as DASH or SMT (Smart Media Transport)). The file decapsulation process on the decoding device is the opposite of the file encapsulation process on the encoding device. The decoding device decapsulates the media file resources according to the panoramic video file format (e.g., ISOBMFF) requirements to obtain audio and video streams. The decoding process on the decoding device is the opposite of the encoding process on the encoding device. The decoding device performs audio decoding on the audio stream to restore the audio content, and the decoding device performs video decoding on the video stream to restore the video content.

[0087] For example, Figure 1 As shown, the media file 10F output by the file encapsulator in the encoding device 100A is the same as the media file 10F' input to the file decapsulator in the decoding device 100B. The file decapsulator processes the media file 10F' or the received segment sequence 10F'. s Decapsulate the file and extract the encoded code stream, including the audio code stream 10E' a 、Video stream 10E' v , coded image 10E' i , while parsing the corresponding metadata. Among them, the window-related video may be carried in multiple tracks. Before decoding, these tracks can be merged into a single video code stream in stream rewriting 10E' v Then, the decoding device 100B can process the audio code stream 10E' a Perform audio decoding to obtain audio signal 10B' a(ie, the restored audio content); for the video stream 10E' v Perform video decoding, or decode the coded image 10E' i Image decoding is performed to obtain an image / video signal 10D' (ie, restored video content).

[0088] (4) Rendering process of panoramic video.

[0089] The decoding device renders the audio content obtained from audio decoding and the video content obtained from video decoding based on the rendering-related metadata in the media presentation description information. Once the rendering is completed, the image is played and output. In particular, because panoramic video uses 3DoF production technology, the decoding device mainly renders the image based on the current viewpoint, disparity, depth information, etc. The viewpoint refers to the viewing position of the service object, and the disparity refers to the difference in line of sight between the two eyes of the service object or the difference in line of sight caused by movement.

[0090] The panoramic video system supports data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. A panoramic video can include multiple data boxes, such as a Sphere Region Zooming Box, which contains metadata describing sphere region zooming information; a 2D Region Zooming Box, which contains metadata describing 2D region zooming information; a Region Wise Packing Box, which contains metadata describing the corresponding information during the region packing process; and so on.

[0091] For example, Figure 1 As shown, the decoding device 100B can be based on the current viewing direction or window (ie, viewing area) and projection, spherical coverage, rotation, and from the media file 10F' or the segment sequence 10F' s The obtained regional encapsulation metadata is parsed and the decoded encapsulated image 10D' (ie, the image / video signal 10D') is projected onto the screen of a head-mounted display or any other display device. Similarly, the audio signal 10B' is decoded according to the current viewing direction. a Rendering is performed (e.g., via headphones or speakers). The current viewing direction is determined by head tracking and possibly eye tracking. In addition to being used by the renderer to render appropriate portions of the decoded video and audio signals, the current viewing direction can also be used by the video decoder and audio decoder for decoding optimization. In window-related transmission, the current viewing direction is also passed to the policy module in the decoding device 100B, which can determine the video track to receive based on the current viewing direction.

[0092] Further, please see Figure 3 , Figure 3 This is an architectural diagram of a volumetric video system provided by an embodiment of the present application. Figure 3 As shown, the volumetric video system includes an encoding device (e.g., encoding device 200A) and a decoding device (e.g., decoding device 200B). The encoding device can be a computer device used by a volumetric video provider, which can be a terminal (e.g., a PC (Personal Computer), a smart mobile device (e.g., a smartphone), or a server. The decoding device can be a computer device used by a volumetric video user, which can be a terminal (e.g., a PC (Personal Computer), a smart mobile device (e.g., a smartphone), or a VR device (e.g., a VR helmet, VR glasses), etc.). The volumetric video data processing process includes data processing on the encoding device side and data processing on the decoding device side.

[0093] The data processing process on the encoding device side mainly includes: (1) the process of acquiring and producing the media content of the volumetric video; (2) the process of encoding and encapsulating the volumetric video file. The data processing process on the decoding device side mainly includes: (1) the process of decapsulating and decoding the volumetric video file; (2) the process of rendering the volumetric video. In addition, the transmission process involving volumetric video between the encoding device and the decoding device can be carried out based on various transmission protocols. The transmission protocols here may include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic bit rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, smart media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.

[0094] The following will be combined Figure 3 , each process involved in the data processing of volumetric video is introduced in detail.

[0095] 1. Data processing on the encoding device side:

[0096] (1) The process of acquiring and producing the media content of volumetric video.

[0097] 1) The process of acquiring the media content of volumetric video.

[0098] The media content of volumetric video is obtained by capturing the real-world audio and visual scene using a capture device. In one implementation, the capture device can refer to a hardware component within the encoding device, such as a microphone, camera, or sensor on a terminal. In another implementation, the capture device can also be a hardware device connected to the encoding device, such as a camera connected to a server, which provides the encoding device with a service for acquiring the media content of the volumetric video. The capture device may include, but is not limited to, audio equipment, imaging equipment, and sensor equipment. Audio equipment may include audio sensors and microphones. Camera equipment may include standard cameras, stereo cameras, light field cameras, and other devices. Sensor equipment may include laser equipment, radar equipment, and other devices. There may be multiple capture devices, deployed at specific locations in real space to simultaneously capture audio and video content from different angles within the space, with the captured audio and video content synchronized in both time and space. In embodiments of the present application, media content in a three-dimensional space captured by capture devices deployed at specific locations to provide a multi-degree-of-freedom (e.g., 3DoF+, 6DoF) viewing experience may be referred to as volumetric video.

[0099] For example, take the example of obtaining the video content of the volumetric video. Figure 3 As shown, a visual scene 20A (including a real-world visual scene or a synthetic visual scene) can be captured by a camera array connected to an encoding device 200A, or by a camera device having multiple cameras and sensors connected to the encoding device 200A, or by multiple virtual cameras. The captured result can be source volume data 20B (i.e., video content of a volumetric video).

[0100] 2) The production process of volumetric video media content.

[0101] It should be understood that the process of producing the media content of the volumetric video involved in the embodiments of the present application can be understood as the process of producing the content of the volumetric video, and the content production of the volumetric video here is mainly produced by multi-viewpoint video, point cloud data, light field and other forms of content captured by cameras or camera arrays deployed in multiple locations. For example, the encoding device can convert the volumetric video from a three-dimensional representation to a two-dimensional representation. The volumetric video here can contain geometric information, attribute information, placeholder information and atlas data, etc. The volumetric video generally requires specific processing before encoding. For example, point cloud data requires cutting, mapping and other processes before encoding. For example, before encoding, the different viewpoints of the multi-view video generally need to be grouped to distinguish between the main viewpoint and the auxiliary viewpoint within each group.

[0102] Specifically, ① the 3D representation data of the captured input volumetric video (i.e., the aforementioned point cloud data) is projected onto a 2D plane, typically using orthogonal projection, perspective projection, or ERP (equirectangular projection). The volumetric video projected onto the 2D plane is represented by data from a geometric component, a placeholder component, and an attribute component. The geometric component data provides the position information of each point in the volumetric video in 3D space, the attribute component data provides additional attributes of each point in the volumetric video (such as texture or material information), and the placeholder component data indicates whether the data in other components is associated with the volumetric video.

[0103] ② Process the component data of the 2D representation of the volumetric video to generate tiles. Based on the position of the volumetric video represented in the geometric component data, the 2D plane area where the 2D representation of the volumetric video is located is divided into multiple rectangular areas of different sizes. Each rectangular area is a tile, and the tile contains the necessary information to back-project the rectangular area into 3D space.

[0104] ③ Pack the tiles to generate atlases, placing the tiles in a two-dimensional grid and ensuring that the valid parts of each tile do not overlap. The tiles generated by a volumetric video can be packaged into one or more atlases;

[0105] ④ Generate corresponding geometric data, attribute data, and placeholder data based on the atlas data, and combine the atlas data, geometric data, attribute data, and placeholder data to form the final representation of the volumetric video in a two-dimensional plane.

[0106] It should be noted that, during the content production process of volumetric video, the geometry component is mandatory, the placeholder component is conditionally mandatory, and the attribute component is optional.

[0107] In addition, it should be noted that since panoramic videos can be captured by capture devices, such videos are processed by encoding devices and transmitted to decoding devices for corresponding data processing. Business objects on the decoding device side need to perform some specific actions (such as head rotation) to view 360-degree video information, while performing non-specific actions (such as moving the head) cannot obtain corresponding video changes, and the VR experience is not good. Therefore, it is necessary to provide additional depth information that matches the panoramic video to enable business objects to obtain better immersion and a better VR experience, which involves 6DoF production technology. When business objects can move more freely in a simulated scene, it is called 6DoF. When using 6DoF production technology to produce volumetric video content, the capture device generally uses light field cameras, laser equipment, radar equipment, etc. to capture point cloud data or light field data in space. Please also refer to Figure 4 , Figure 4 6DoF is a schematic diagram of the embodiment of the present application. Figure 4 As shown, 6DoF is divided into window 6DoF, omnidirectional 6DoF and 6DoF, among which window 6DoF means that the rotational movement of the business object in the X-axis and Y-axis is limited, and the translation in the Z-axis is limited. For example, the business object cannot see the scene outside the window frame, and the business object cannot pass through the window. Omnidirectional 6DoF means that the rotational movement of the business object in the X-axis, Y-axis and Z-axis is limited. For example, the business object cannot freely pass through the three-dimensional 360-degree VR content in the restricted movement area. 6DoF means that the business object can freely translate along the X-axis, Y-axis and Z-axis on the basis of 3DoF. For example, the business object can move freely in the three-dimensional 360-degree VR content. Similar to 6DoF, there are 3DoF and 3DoF+ production technologies. Figure 5 This is a schematic diagram of 3DoF+ provided by the embodiment of this application. Figure 5 As shown, 3DoF+ means that when the virtual scene provided by the immersive media has a certain depth information, the business object head can move in a limited space based on 3DoF to view the picture provided by the media content. Figure 2 , I will not go into details here.

[0108] (2) The process of encoding and file encapsulation of volumetric video.

[0109] The captured audio content can be directly audio-encoded to form an audio stream of the volumetric video. The captured video content can be video-encoded to obtain a video stream of the volumetric video. It should be noted here that if 6DoF production technology is used, a specific encoding method (such as a point cloud compression method based on traditional video encoding) needs to be used for encoding during the video encoding process. The audio stream and the video stream are encapsulated in a file container according to the file format of the volumetric video (such as ISOBMFF) to form a media file resource of the volumetric video. The media file resource can be a media file of the volumetric video formed by a media file or a media fragment; and the media presentation description information (i.e. MPD) is used to record the metadata of the media file resource of the volumetric video according to the file format requirements of the volumetric video. The metadata here is a general term for information related to the presentation of the volumetric video. The metadata may include description information of the media content, timing metadata information that describes the mapping relationship between each constructed viewpoint group and the spatial position information of the viewed media content, description information of the view window, and signaling information related to the presentation of the media content, etc. As Figure 1 As shown, the encoding device stores the media presentation description information and media file resources formed after the data processing process.

[0110] Specifically, the captured audio is encoded into a corresponding audio bitstream. The volumetric video's geometry, attributes, and placeholder information can be encoded using traditional video encoding methods, while the volumetric video atlas data can be entropy coded. The encoded media is then encapsulated in a file container according to a specific format (such as ISOBMFF or HNSS) and combined with metadata describing the media content attributes and viewport metadata to form a media file or an initialization segment and media segments, based on a specific media file format.

[0111] For example, Figure 3 As shown, the encoding device 200A performs volumetric video encoding on one or more volumetric video frames in the source volumetric video data 20B to obtain an encoded VC3 code stream 20E. v (i.e., video code stream), including an atlas code stream (i.e., a code stream obtained by encoding the atlas data), at most one occupancy code stream (i.e., a code stream obtained by encoding the placeholder map information), a geometry code stream (i.e., a code stream obtained by encoding the geometry information), and zero or more attribute code streams (i.e., a code stream obtained by encoding the attribute information). Subsequently, the encoding device 200A can encapsulate one or more encoded code streams into a media file 20F for local playback or into a segment sequence 20F for streaming, which includes an initialization segment and multiple media segments, according to a specific media file format (e.g., ISOBMFF). s In addition, the file encapsulator in the encoding device 200A may also add metadata to the media file 20F or the segment sequence 20F. s Furthermore, the encoding device 200A may use a certain transmission mechanism (such as DASH, SMTP) to transmit the segment sequence 20F to the encoding device 200A. s The media file 20F is transmitted to the decoding device 200B, and the media file 20F is also transmitted to the decoding device 200B. Optionally, the decoding device 200B here can be a player.

[0112] 2. Data processing on the decoding device side:

[0113] (3) The process of decapsulating and decoding volumetric video files.

[0114] The decoding device can dynamically and adaptively obtain the volumetric video's media file resources and corresponding media presentation description information from the encoding device based on recommendations from the encoding device or according to the needs of the service object on the decoding device. For example, the decoding device can determine the orientation and position of the service object based on its head / eye tracking information, and then dynamically request the corresponding media file resources from the encoding device based on the determined orientation and position. The media file resources and media presentation description information are transmitted from the encoding device to the decoding device via a transport mechanism (such as DASH and SMT). The file decapsulation process on the decoding device is the inverse of the file encapsulation process on the encoding device. The decoding device decapsulates the media file resources according to the volumetric video file format (e.g., ISOBMFF) to obtain audio and video streams. The decoding process on the decoding device is the inverse of the encoding process on the encoding device. The decoding device performs audio decoding on the audio stream to restore the audio content, and the decoding device performs video decoding on the video stream to restore the video content.

[0115] For example, Figure 3 As shown, the media file 20F output by the file encapsulator in the encoding device 200A is the same as the media file 20F' input to the file decapsulator in the decoding device 200B. The file decapsulator processes the media file 20F' or the received segment sequence 20F'. s Decapsulate the file and extract the encoded VC3 stream 20E' v , and parse the corresponding metadata at the same time, and then VC3 stream 20E' v Volumetric video decoding is performed to obtain a decoded video signal 20D′ (ie, restored video content).

[0116] (4) Volumetric video rendering process.

[0117] The decoding device renders the audio content obtained by audio decoding and the video content obtained by video decoding according to the rendering-related metadata in the media presentation description information corresponding to the media file resource. Once the rendering is completed, the playback output of the image is realized.

[0118] Volumetric video systems support data boxes, which are data blocks or objects containing metadata. Specifically, a data box contains metadata about the corresponding media content. Volumetric video can include multiple data boxes, such as the ISO Base Media File Format (ISOBMFF) box, which contains metadata describing the file encapsulation.

[0119] For example, Figure 3As shown, decoding device 200B can reconstruct the decoded video signal 20D' based on the current viewing direction or viewport to obtain reconstructed volumetric video data 20B'. This reconstructed volumetric video data 20B' can then be rendered and displayed on a head-mounted display (HMD) or any other display device. The current viewing direction is determined by head tracking and, possibly, eye tracking. In viewport-related transmission, the current viewing direction is also communicated to a policy module in decoding device 200B, which can determine which tracks to receive based on the current viewing direction.

[0120] Through the above Figure 1 The process described in the corresponding embodiment or the above Figure 3 In the process described in the corresponding embodiment, the decoding device can dynamically obtain the media file resources corresponding to the immersive media from the encoding device side. Since the media file resources are obtained by the encoding device encoding and encapsulating the captured audio and video content, after the decoding device receives the media file resources returned by the encoding device, it needs to first decapsulate the media file resources to obtain the corresponding audio and video code stream, and then decode the audio and video code stream, and finally present the decoded audio and video content to the business object. The immersive media here includes but is not limited to panoramic video and volumetric video, among which volumetric video can specifically include multi-view video, VPCC (Video-based Point Cloud Compression, point cloud compression based on traditional video coding) point cloud media, and GPCC (Geometry-based Point Cloud Compression, point cloud compression based on geometric models) point cloud media.

[0121] It should be understood that when a business object consumes immersive media, interactive feedback can be continuously performed between the decoding device and the encoding device. For example, the decoding device can feed back the business object state (for example, object location information) to the encoding device, so that the encoding device can provide the business object with corresponding media file resources based on the content of the interactive feedback. In an embodiment of the present application, the playable media content (including audio content and video content) obtained after decapsulating and decoding the media file resources of the immersive media can be collectively referred to as immersive media content. For the decoding device, the decoding device can play the immersive media content restored from the acquired media file resources on the video playback interface. In other words, one media file resource can correspond to one immersive media content. Therefore, in an embodiment of the present application, the immersive media content corresponding to the first media file resource can be referred to as the first immersive media content, and the immersive media content corresponding to the second media file resource can be referred to as the second immersive media content. Similar naming can also be used for other media file resources and corresponding immersive media content.

[0122] In order to support richer interactive feedback scenarios, an embodiment of the present application provides a method for indicating an immersive media interactive feedback message. Specifically, a video client can be run on a decoding device (such as a user terminal), and the first immersive media content can be played on the video playback interface of the video client. It should be understood that the first immersive media content here is obtained by the decoding device after decapsulating and decoding the first media file resource, and the first media file resource is obtained by the encoding device (such as a server) after encoding and encapsulating the relevant audio and video content in advance. During the playback of the first immersive media content, the decoding device can respond to the interactive operation on the first immersive media content and generate an interactive feedback message corresponding to the interactive operation, wherein the interactive feedback message here carries a business key field for describing the business event information indicated by the interactive operation. Further, the decoding device can send the interactive feedback message to the encoding device so that the encoding device can determine the business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and can obtain a second media file resource for responding to the interactive operation based on the business event information indicated by the interactive operation. Wherein, the second media file resource here is obtained by the encoding device after encoding and encapsulating the relevant audio and video content in advance. Finally, the decoding device can receive the second media file resource returned by the encoding device, and decapsulate and decode the second media file resource to obtain a playable second immersive media content, and then play the second immersive media content on its video playback interface. The specific process of decapsulating and decoding the second media file resource can be found in the above Figure 1 The relevant process described in the corresponding embodiment or Figure 3 The relevant processes described in the corresponding embodiments.

[0123] In an embodiment of the present application, the above-mentioned interactive operations may include not only operations related to the user's position (for example, a change in the user's position), but also other operations on the immersive media content currently played by the video client (for example, a zoom operation). Therefore, through the business key fields carried in the interactive feedback message, the video client on the decoding device can feedback various types of business event information to the encoding device. In this way, the encoding device can determine the immersive media content in response to the interactive operation based on these different types of business event information, rather than relying solely on the user's position information, thereby enriching the information type of the interactive feedback and improving the accuracy of the video client in obtaining media content during the interactive feedback process.

[0124] The method provided in the embodiment of the present application can be applied to the server side (i.e., encoding device side), player side (i.e., decoding device side), and intermediate nodes (e.g., SMT receiving entity, SMT sending entity) of the immersive media system. The specific process of interactive feedback between the decoding device and the encoding device can be found in the following Figure 6-Figure 8 Description of the corresponding embodiment.

[0125] Further, see Figure 6 , Figure 6 This is a flow chart of a data processing method for immersive media provided by an embodiment of the present application. The method can be executed by a decoding device in an immersive media system (for example, a panoramic video system or a volumetric video system), which can be the above-mentioned Figure 1 The decoding device 100B in the corresponding embodiment may also be the above Figure 3 The decoding device 200B in the corresponding embodiment may be a user terminal integrated with a video client, and the method may include at least the following steps S101 to S103:

[0126] Step S101: In response to an interactive operation on a first immersive media content, an interactive feedback message corresponding to the interactive operation is generated; the interactive feedback message carries a business key field for describing business event information indicated by the interactive operation;

[0127] Specifically, after obtaining the first media file resource returned by the server, the video client on the user terminal can decapsulate and decode the first media file resource to obtain first immersive media content, and then play the first immersive media content on the video playback interface of the video client. The first immersive media content here refers to the immersive media content currently being viewed by the service object, which can be the user consuming the first immersive media content. Optionally, the first immersive media content can belong to an immersive video, which can be a video collection containing one or more immersive media content. This embodiment of the application does not limit the number of contents contained in the immersive video. For example, assume that an immersive video provided by the server includes N immersive media content, where N is an integer greater than 1, namely: immersive media content A1 associated with scene P1, immersive media content A2 associated with scene P2, ..., immersive media content AN associated with scene P1. The video client can obtain any one or more immersive media content from the N immersive media content, such as immersive media content A1, based on the server's recommendation or the service object's needs. In this case, immersive media content A1 can serve as the current first immersive media content.

[0128] It should be noted that the immersive video may be a panoramic video; optionally, the immersive video may be a volumetric video. The embodiment of the present application does not limit the specific video type of the immersive video.

[0129] Furthermore, during the process of playing the first immersive media content on the video playback interface of the video client, the video client can respond to the interactive operation for the first immersive media content currently being played and generate an interactive feedback message corresponding to the interactive operation. It should be understood that the interactive feedback message can also be called interactive feedback signaling, which can provide interactive feedback between the video client and the server during immersive media consumption. For example, when the first immersive media content consumed is a panoramic video, the SMT receiving entity can periodically send entity feedback virtual camera direction information to the SMT to notify the current VR virtual camera direction. In addition, the corresponding direction information is also sent when the FOV (Field of view) changes. For another example, when the first immersive media content consumed is a volumetric video, the SMT receiving entity can periodically send entity feedback virtual camera position or business object position and viewing direction information to the SMT so that the video client can obtain the corresponding media content. Among them, the SMT receiving entity and the SMT sending entity are intermediate nodes between the video client and the server.

[0130] In an embodiment of the present application, an interactive operation refers to an operation performed by a business object on the first immersive media content currently being consumed, including but not limited to a zoom operation, a switching operation, and a position interaction operation. Among them, a zoom operation refers to an operation of reducing or enlarging the screen size of the first immersive media content. For example, the screen size of the immersive media content A1 can be enlarged by double-clicking the immersive media content A1; for another example, the screen size of the immersive media content A1 can be reduced or enlarged by sliding and stretching the immersive media content A1 in different directions with two fingers at the same time. The switching operation here may include a playback rate switching operation, a picture quality switching operation (i.e., a clarity switching operation), a flip operation, a content switching operation, and other event-based triggering operations predefined at the application layer, such as a click operation on a target position in the screen, a triggering operation when the business object faces a target direction, and the like. The position interaction operation here refers to an operation on the object position information (i.e., user position information) generated by the business object when viewing the first immersive media content, such as a change in real-time position, a change in viewing direction, a change in viewing angle direction, and the like. To facilitate subsequent understanding and distinction, in this embodiment of the application, when the first immersive media content is a panoramic video, the corresponding position interaction operation is referred to as a first position interaction operation; when the first immersive media content is a volumetric video, the corresponding position interaction operation is referred to as a second position interaction operation. It should be noted that this embodiment of the application does not limit the specific triggering methods for zooming, switching, and position interaction operations.

[0131] It should be understood that the interaction feedback message may carry a business key field for describing business event information indicated by the interaction operation.

[0132] In an optional embodiment, the interactive feedback message may directly include business key fields, where the business key fields may include a first key field, a second key field, a third key field, and a fourth key field. Among them, the first key field is used to represent the zoom ratio when the zoom event indicated by the zoom operation is executed when the interactive operation includes a zoom operation; the second key field is used to represent the event label and event status corresponding to the switching event indicated by the switching operation when the interactive operation includes a switching operation; the third key field is used to represent the first object position information (for example, the real-time position of the business object, the viewing direction, etc.) of the business object of the first immersive media content belonging to the panoramic video when the interactive operation includes a first-position interactive operation; the fourth key field is used to represent the second object position information (for example, the real-time viewing direction of the business object) of the business object of the first immersive media content belonging to the volumetric video when the interactive operation includes a second-position interactive operation. It can be seen that the business event information indicated by the interactive operation can be a zoom ratio, an event label and an event status, the first object position information, or the second object position information.

[0133] It should be understood that an interaction feedback message may carry a business key field for describing business event information indicated by one or more interaction operations. The embodiment of the present application does not limit the number and type of interaction operations corresponding to the interaction feedback message.

[0134] It can be understood that since the video types corresponding to the first immersive media content are different, the first position interaction operation and the second position interaction operation cannot exist at the same time, that is, in the same interaction feedback message, the valid third key field and the valid fourth key field cannot exist at the same time.

[0135] It should be understood that, optionally, the business key fields carried by an interactive feedback message may include any one of the first key field, the second key field, the third key field, and the fourth key field. For example, each time an interactive operation occurs, the video client generates a corresponding interactive feedback message. Optionally, in a scenario where the first immersive media content belongs to a panoramic video, the business key fields carried by an interactive feedback message may include any one or more of the first key field, the second key field, and the third key field. Similarly, optionally, in a scenario where the first immersive media content belongs to a volumetric video, the business key fields carried by an interactive feedback message may include any one or more of the first key field, the second key field, and the fourth key field. For example, within a period of time, the business object performed a zoom operation on the above-mentioned immersive media content A2, and in the process of watching the immersive media content A2, the object position information of the business object changed. For example, the business object watched while walking. If the immersive media content A2 belongs to a volumetric video, the interactive feedback message generated at this time will simultaneously include a first key field reflecting the zoom ratio and a fourth key field reflecting the second object position information. Therefore, the second immersive media content finally obtained is determined based on the zoom ratio and the second object position information. In other words, the immersive media content responding to the interactive operation can be determined based on different types of business event information, thereby improving the accuracy of the video client in obtaining media content during the interactive feedback process.

[0136] Optionally, the above-mentioned interactive feedback message may also include an information identification field, which is used to characterize the information type of the business event information indicated by each interactive operation. For example, the field value of the information identification field can be the information name corresponding to each type of business event information. In this way, when an interactive feedback message carries multiple types of business event information at the same time, the information type can be distinguished by the information identification field.

[0137] It should be understood that the embodiments of this application do not limit the timing of interactive feedback, and can be agreed upon at the application layer based on actual needs. For example, upon detecting a certain interactive operation, the video client can immediately generate a corresponding interactive feedback message and send it to the server. Alternatively, the video client can periodically send interactive feedback messages to the server, for example, once every 30 seconds.

[0138] In another optional implementation, the interaction feedback message may include an interaction signaling table associated with the interaction operation, and the interaction signaling table may include a service-critical field for describing the service event information indicated by the interaction operation. In other words, the interaction feedback message may redefine and organize various types of service event information in the form of an interaction signaling table.

[0139] Optionally, in scenarios where the interactive operation is a triggering operation, the video client responds to the triggering operation on the first immersive media content by determining the first information type field of the service event information indicated by the triggering operation and recording the operation timestamp of the triggering operation. The triggering operation here can refer to a contact operation or certain specific non-contact operations on the first immersive media content. For example, the triggering operation can include a zooming operation, a switching operation, etc. Furthermore, the video client can add the first information type field and the operation timestamp to an interactive signaling table associated with the first immersive media content and use the first information type field added to the interactive signaling table as a service key field for describing the service event information indicated by the interactive operation. Subsequently, the video client can generate an interactive feedback message corresponding to the triggering operation based on the service key field and the operation timestamp in the interactive signaling table. The first information type field can be used to indicate the information type of the service event information indicated by the triggering operation. It is understood that each triggering operation can correspond to an interactive signaling table. Therefore, the same interactive feedback message can include one or more interactive signaling tables. This embodiment of the present application does not limit the number of interactive signaling tables included in the interactive feedback message.

[0140] It should be understood that, optionally, when the triggering operation includes a zoom operation, the business event information indicated by the zoom operation is a zoom event, and when the field value of the first information type field corresponding to the zoom operation is the first field value, the field mapped by the first information type field having the first field value is used to represent the zoom ratio when executing the zoom event.

[0141] It should be understood that, optionally, when the triggering operation includes a switching operation, the business event information indicated by the switching operation is a switching event, and when the field value of the first information type field corresponding to the switching operation is a second field value, the field mapped to the first information type field having the second field value is used to represent the event label and event status of the switching event. Specifically, when the state value of the event state is a first state value, the event state having the first state value is used to represent that the switching event is in an event trigger state; and when the state value of the event state is a second state value, the event state having the second state value is used to represent that the switching event is in an event end state.

[0142] For ease of understanding, the following is further explained using the SMT signaling message format as an example. Please refer to Table 1, which is used to indicate the syntax of an interactive signaling table provided in an embodiment of the present application:

[0143] Table 1

[0144]

[0145]

[0146] The semantics of the syntax shown in Table 1 above are as follows: table_id is the signaling table identification field, which is used to represent the identifier of the interactive signaling table. version is the signaling table version field, which is used to represent the version number of the interactive signaling table. length is the signaling table length field, which is used to represent the length of the interactive signaling table. table_type is the first information type field, which is used to represent the information type carried by the interactive signaling table (for example, a zoom event or a switching event). timestamp is the operation timestamp, which is used to indicate the timestamp generated by the current trigger operation. UTC time (Universal Time Coordinated) can be used here. As shown in Table 1, when the field value of the first information type field (i.e., table_type) is the first field value (for example, 2), the field mapped by the first information type field is zoom_ratio. zoom_ratio indicates the ratio of the business object zoom behavior, that is, the zoom ratio when executing the zoom event (also called screen zoom information). Optionally, zoom_ratio can be 2 -3 For example, if user 1 (i.e., the business object) zooms in on the immersive media content F1 (i.e., the first immersive media content), the corresponding interactive feedback message will carry an interactive signaling table with table_type==2, and zoom_ratio=16, which means that the current zoom ratio is 16*2. -3=2 times. Optionally, zoom_ratio can also be used as the first key field described in the aforementioned optional implementation manner. As shown in Table 1, when the field value of the first information type field is the second field value (for example, 3), the fields mapped by the first information type field are event_label and event_trigger_flag, event_label indicates the event label triggered by the business object interaction, and event_trigger_flag indicates the event state triggered by the business object interaction. Optionally, when event_trigger_flag takes the value of 1 (i.e., the first state value), it indicates that the event is triggered (i.e., the switching event is in the event triggered state), and when event_trigger_flag takes the value of 0 (i.e., the second state value), it indicates that the event ends (i.e., the switching event is in the event end state). For example, assuming that user 2 (i.e., the business object) clicks on the content switch control in the video playback interface while watching immersive media content F2 (i.e., the first immersive media content), the corresponding interactive feedback message will carry an interactive signaling table with table_type==3, and event_label="content switch", event_trigger_flag=1, which means that user 2 has triggered the content switching operation and hopes to switch the currently playing immersive media content F2 to other immersive media content. Optionally, event_label and event_trigger_flag can also be used as the second key field described in the aforementioned optional implementation. In addition, reserved indicates the reserved byte position.

[0147] Among them, the embodiment of the present application does not limit the specific numerical values ​​of the above-mentioned first field value and the second field value, and does not limit the specific numerical values ​​of the first state value and the second state value. It should be understood that the embodiment of the present application can support R&D personnel to pre-define the required switching events at the application layer. The specific content of the event label can be determined according to the immersive media content. The embodiment of the present application does not limit this. It should be noted that the relevant immersive media content needs to support customized switching events before it is possible to realize event triggering in the subsequent interaction process. For example, when the above-mentioned immersive media content F2 supports content switching, the corresponding content switching control will be displayed in the video playback interface of the immersive media content F2.

[0148] Optionally, for scenarios where the interactive operation is a location-based interactive operation, embodiments of the present application may also support carrying object location information in the interactive feedback message. In this case, this information may also be defined in the form of an interactive signaling table. Specifically, upon detecting the object location information of a service object for viewing first immersive media content, the video client may treat the location-based interactive operation targeting the object location information as an interactive operation in response to the first immersive media content. The client may then determine the second information type field of the service event information indicated by the interactive operation and record the operation timestamp of the interactive operation. Furthermore, the second information type field and the operation timestamp may be added to an interactive signaling table associated with the first immersive media content, and the second information type field added to the interactive signaling table may be used as a service key field to describe the service event information indicated by the interactive operation. Subsequently, the video client may generate an interactive feedback message corresponding to the interactive operation based on the service key field and operation timestamp in the interactive signaling table. It will be understood that each location-based interactive operation may correspond to an interactive signaling table. Therefore, the same interactive feedback message may include one or more interactive signaling tables. However, the same interactive feedback message cannot contain both an interactive signaling table carrying the first object location information and an interactive signaling table carrying the second object location information. It is understandable that if the video client periodically feeds back the object location information of the business object to the server, then during the period of time that the business object consumes the first immersive media content, its object location information may or may not change. When the object location information does not change, the server can still obtain the corresponding immersive media content based on the object location information. In this case, the immersive media content obtained may be the same as the first immersive media content. Similarly, in addition to the object location information, if the video client also feeds back other information to the server during this period, such as the zoom ratio when performing a zoom operation on the first immersive media content, the server can obtain the corresponding immersive media content based on the object location information and the zoom ratio. In this case, the immersive media content obtained is different from the first immersive media content.

[0149] It should be understood that, optionally, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a panoramic video, the field value of the second information type field corresponding to the object position information is a third field value, and the second information type field with the third field value includes a first type of position field, which is used to describe the position change information of the business object viewing the first immersive media content belonging to the panoramic video.

[0150] It should be understood that, optionally, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a volumetric video, the field value of the second information type field corresponding to the object position information is a fourth field value, and the second information type field with the fourth field value includes a second type of position field, which is used to describe the position change information of the business object viewing the first immersive media content belonging to the volumetric video.

[0151] For ease of understanding, further reference is made to Table 2, which is used to indicate the syntax of an interactive signaling table provided in an embodiment of the present application:

[0152] Table 2

[0153]

[0154]

[0155] The semantics of the syntax shown in Table 2 above are as follows: table_id is the signaling table identification field, which is used to characterize the identifier of the interactive signaling table. version is the signaling table version field, which is used to characterize the version number of the interactive signaling table. length is the signaling table length field, which is used to characterize the length of the interactive signaling table. table_type is the second information type field, which is used to characterize the type of information carried by the interactive signaling table (such as the first object position information or the second object position information). timestamp is the operation timestamp, which is used to indicate the timestamp generated by the current position interaction operation. UTC time can be used here. As shown in Table 2, when the field value of table_type is 0 (that is, the third field value), the first type of position fields it contains are: 3DoF+_flag indicates 3DoF+ video content; interaction_target is the interaction target field, which indicates the target of the current interaction of the video client, including the current status of the helmet device (HMD_status), the business object focus target (Object of interests), the current status of the business object (User_status), etc. interaction_type is the interaction type field, which is set to 0 in the embodiment of the present application. The value of the interaction target field interaction_target can be found in Table 3, which is used to indicate a value table of the interaction target field provided in an embodiment of the present application:

[0156] Table 3

[0157] type Value describe Null 0 The interaction target is empty, that is, there is no specific interaction target HMD_status 1 The interaction target is the current status of the helmet device Object of interests 2 The interaction target is the current state of the business object's focus area User_status 3 The interaction target is the current state of the business object

[0158] In conjunction with Table 3, please continue to refer to Table 2. When the interaction target field value is 1, it indicates that the interaction target is the current state of the helmet device. Correspondingly, ClientRegion is the window information, indicating the size and screen resolution of the video client window. For its specific syntax, please refer to Table 4, which is used to indicate the syntax of a window information provided in an embodiment of the present application:

[0159] Table 4

[0160]

[0161] The semantics of Table 4 above are as follows: Region_width_angle indicates the horizontal angle of the video client window, with an accuracy of 2 -16 Degrees, the value range is (-90*2 16 , 90*2 16 ). Region_height_angle indicates the vertical angle of the video client window, with an accuracy of 2 -16 Degrees, the value range is (-90*2 16 , 90*2 16 Region_width_resolution indicates the horizontal resolution of the video client window, and its value range is (0, 2 16 -1). Region_height_resolution indicates the vertical resolution of the video client window, and its value range is (0, 2 16 -1).

[0162] Please continue to refer to Table 2. When the interaction target field is set to 2, it indicates that the interaction target is the current state of the business object's focus area. Correspondingly, ClientRotation is the viewing direction, indicating the real-time change of the business object's viewing angle relative to the initial viewing angle. For its specific syntax, please refer to Table 5, which is used to indicate a syntax of a viewing angle provided in an embodiment of the present application:

[0163] Table 5

[0164]

[0165] The semantics of the above Table 5 are as follows: 3D_rotation_type indicates the representation type of the rotation information. A value of 0 indicates that the rotation information is given in the form of Euler angles; a value of 1 indicates that the rotation information is given in the form of quaternions; other values ​​are reserved. rotation_yaw indicates the yaw angle of the real-time perspective of the business object relative to the initial perspective along the x-axis, and the value range is (-180*2 16 , 180*2 16-1). rotation_pitch indicates the pitch angle of the business object's real-time viewing angle relative to the initial viewing angle along the y-axis, and the value range is (-90*2 16 , 90*2 16 rotation_roll indicates the roll angle of the business object's real-time viewing angle relative to the initial viewing angle along the z-axis, and the value range is (-180*2 16 , 180*2 16 -1). rotation_x, rotation_y, rotation_z, and rotation_w indicate the values ​​of the quaternion x, y, z, and w components, respectively, representing the rotation information of the business object's real-time perspective relative to the initial perspective.

[0166] Please continue to refer to Table 2. When the interaction target field value is 3 and the 3DoF+_flag value is 1, it means that the interaction target is the current state of the business object. Correspondingly, ClientPosition is the real-time position of the business object, indicating the displacement of the business object relative to the starting position in the virtual scene. In 3DoF (i.e., 3DoF+_flag value is 0), the field values ​​of all fields in the structure are 0. In 3DoF+ (i.e., 3DoF+_flag value is 1), the field values ​​of all fields in the structure are non-zero values, and the value range should be within the constraint range. behavior_coefficient defines an amplification behavior coefficient. Among them, the specific syntax of ClientPosition is shown in Table 6, which is used to indicate the syntax of the real-time position of a business object provided in an embodiment of the present application:

[0167] Table 6

[0168]

[0169] The semantics of the above table 6 are as follows: position_x indicates the displacement of the real-time position of the business object relative to the starting position along the x-axis, and the value range is (-2 15 , 2 15 -1) mm. position_y indicates the displacement of the real-time position of the business object relative to the starting position along the y-axis, and the value range is (-2 15 , 2 15 -1) mm. position_z indicates the displacement of the real-time position of the business object relative to the starting position along the z axis, and the value range is (-2 15 , 2 15 -1) mm.

[0170] Optionally, the first type of position field included when the field value of table_type is 0 can be used as the third key field in the aforementioned optional implementation manner.

[0171] Please continue to refer to Table 2. As shown in Table 2, when the field value of table_type is 1 (i.e., the fourth field value), the second type of position fields it contains are: ClientPosition indicates the current position of the business object in the global coordinate system. Its specific syntax can be found in Table 6 above. V3C_orientation indicates the viewing direction of the business object in the Cartesian coordinate system established with the current position. last_processed_media_timestamp indicates the timestamp of the last media unit added to the decoder buffer. The SMT sending entity uses this field to determine the next transmitted media unit from the new asset (i.e., the new immersive media content) of the volumetric video player. The next media unit is a media unit with a timestamp or sequence number that immediately follows this timestamp. Starting from the subsequent media timestamp, the SMT sending entity switches from transmitting the previous asset (determined according to the previous view window) to transmitting the new asset (determined according to the new view window) to reduce the delay in receiving the media content corresponding to the new view window. Among them, for the specific syntax of V3C_orientation, please refer to Table 7, which is used to indicate the syntax of the real-time viewing direction of a business object provided in an embodiment of the present application:

[0172] Table 7

[0173]

[0174] The semantics of Table 7 are as follows: dirx indicates that a Cartesian coordinate system is established with the business object's location as the origin, with the coordinates of the business object's viewing direction on the x-axis. diry indicates that a Cartesian coordinate system is established with the business object's location as the origin, with the coordinates of the business object's viewing direction on the y-axis. dirz indicates that a Cartesian coordinate system is established with the business object's location as the origin, with the coordinates of the business object's viewing direction on the z-axis.

[0175] Optionally, the second type of position field included when the field value of table_type is 1 can be used as the fourth key field in the aforementioned optional implementation manner.

[0176] It can be understood that the embodiment of the present application can also combine the above Table 1 and Table 2 to obtain an interactive signaling table that can represent at least four types of information, so that the interactive feedback message corresponding to the interactive operation can be generated based on the interactive signaling table. The specific process can be seen in the following Figure 7 This corresponds to step S203 in the embodiment.

[0177] Step S102: Send the interaction feedback message to the server, so that the server determines the business event information indicated by the interaction operation based on the business key field in the interaction feedback message, and obtains the second immersive media content for responding to the interaction operation based on the business event information indicated by the interaction operation;

[0178] Specifically, the video client can send an interactive feedback message to the server. After the server receives the interactive feedback message, it can determine the business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and then obtain the second media file resource corresponding to the second immersive media content used to respond to the interactive operation based on the business event information indicated by the interactive operation. The second media file resource is obtained by the server after encoding and packaging the relevant audio and video content in advance. It corresponds to the second immersive media content. The specific process of encoding and packaging the audio and video content can be found in the above Figure 1 or Figure 3 For example, if the interactive operation is a definition switching operation, the server can obtain a media file resource that matches the resolution indicated by the definition switching operation as a second media file resource in response to the definition switching operation.

[0179] Step S103: receiving the second immersive media content returned by the server.

[0180] Specifically, the video client can receive the second immersive media content returned by the server and can play the second immersive media content on the video playback interface. In conjunction with the above step S102, it should be understood that the server first obtains the media file resource corresponding to the second immersive media content based on the business event information, that is, the second media file resource, and can return the second media file resource to the video client. Therefore, after the video client receives the second media file resource, it can play the second immersive media content through the above step S102. Figure 1 or Figure 3 In the relevant description of the corresponding embodiment, the second media file resource is decapsulated and decoded to obtain the second immersive media content that can be played on the video playback interface of the video client. The specific process of decapsulation and decoding is not repeated here.

[0181] Further, see Figure 7 , Figure 7 This is a flow chart of a data processing method for immersive media provided by an embodiment of the present application. The method can be executed by a decoding device in an immersive media system (for example, a panoramic video system or a volumetric video system), which can be the above-mentioned Figure 1 The decoding device 100B in the corresponding embodiment may also be the above Figure 3The decoding device 200B in the corresponding embodiment may be a user terminal integrated with a video client, and the method may include at least the following steps:

[0182] Step S201: In response to a video playback operation for an immersive video in a video client, generating a playback request corresponding to the video playback operation, and sending the playback request to a server, so that the server obtains first immersive media content of the immersive video based on the playback request;

[0183] Specifically, when a business user wishes to experience an immersive video, they can request the corresponding immersive media content through the video client on the user terminal. For example, the video client can respond to a video playback operation for an immersive video on the video client by generating a play request corresponding to the video playback operation. The play request can then be sent to the server, so that the server can obtain the first media file resource corresponding to the first immersive media content in the immersive video based on the play request. The first media file resource here refers to the data obtained by the server after encoding and packaging the relevant audio and video content.

[0184] Step S202: receiving the first immersive media content returned by the server, and playing the first immersive media content on the video playback interface of the video client;

[0185] Specifically, after the server obtains the first media file resource corresponding to the first immersive media content based on the playback request, the first media file resource can be returned to the video client, so that the video client can receive the first media file resource returned by the server and decapsulate and decode the first media file resource, thereby obtaining the first immersive media content that can be played on the video playback interface of the video client.

[0186] Step S203: When the first immersive media content is played on the video playback interface of the video client, in response to an interactive operation on the first immersive media content, an interactive feedback message corresponding to the interactive operation is generated; the interactive feedback message carries an interactive signaling table associated with the interactive operation;

[0187] Specifically, when playing the first immersive media content on the video playback interface of a video client, the video client can respond to an interactive operation on the first immersive media content and generate an interactive feedback message corresponding to the interactive operation. For example, in response to the interactive operation on the first immersive media content, the video client determines the information type field of the service event information indicated by the interactive operation and records the operation timestamp of the interactive operation. Furthermore, the information type field and operation timestamp can be added to the interactive signaling table associated with the first immersive media content, and the information type field added to the interactive signaling table is used as a service key field for describing the service event information indicated by the interactive operation. Subsequently, an interactive feedback message corresponding to the interactive operation can be generated based on the service key field and operation timestamp in the interactive signaling table.

[0188] It should be understood that the interaction operation here may include one or more of a zoom operation, a switch operation, and a position interaction operation, wherein the position interaction operation may be a first position interaction operation or a second position interaction operation.

[0189] In this embodiment of the present application, an interaction feedback message may carry an interaction signaling table associated with the interaction operation, and the information type field contained in the interaction signaling table may serve as a business-critical field for describing the business event information indicated by the interaction operation. The information type field may include a first information type field related to the triggering operation and a second information type field related to the location interaction operation. In this embodiment of the present application, the first information type field and the second information type field are collectively referred to as the information type field.

[0190] It should be understood that in embodiments of the present application, Tables 1 and 2 can be combined to form an interactive signaling table that can represent at least four information types. In this way, interactive feedback messages of different information types can be integrated through the interactive signaling table without causing confusion due to the diversity of information types. Please refer to Table 8, which is used to indicate the syntax of an interactive signaling table provided in an embodiment of the present application:

[0191] Table 8

[0192]

[0193]

[0194] The table_type shown in Table 8 above is the information type field, which can be used to characterize the type of information carried by the interactive signaling table. The specific semantics of other fields can be found in the above Figure 3 Table 1 and Table 2 in the corresponding embodiment are not described here in detail. Optionally, the value of table_type can refer to Table 9, which is used to indicate a value table of an information type field provided in the embodiment of the present application:

[0195] Table 9

[0196] Value describe 0 Panoramic video user location change information 1 Volumetric video user location change information 2 Screen zoom information 3 Interaction event trigger information 4…255 Undefined

[0197] As shown in Table 9, the field value of the information type field can be the first field value (for example, 2), the second field value (for example, 3), the third field value (for example, 0), the fourth field value (for example, 1), etc. The panoramic video user position change information in Table 9 is the position change information described by the first type of position field, the volumetric video user position change information is the position change information described by the second type of position field, the screen zoom information is the zoom ratio when executing the zoom event, and the interactive event trigger information includes the event label and event status of the switching event. Other values ​​of information may be added in the future.

[0198] It should be understood that based on Table 1, Table 2 or Table 8 above, the interactive feedback message generated by the embodiment of the present application can support a richer interactive feedback scenario. Please also refer to Table 10, which is used to indicate the syntax of an interactive feedback message provided by the embodiment of the present application:

[0199] Table 10

[0200]

[0201]

[0202] The semantics of the syntax shown in Table 10 above are as follows: message_id indicates the identifier of the interactive feedback message. version indicates the version of the interactive feedback message. The information carried by the new version will overwrite any previous old version. length indicates the length of the interactive feedback message in bytes, that is, the length from the next field to the last byte of the interactive feedback message. The value "0" is invalid in this field. number_of_tables is the signaling table number field, which indicates the number of interactive signaling tables included in the interactive feedback message, represented by N1 here. The specific value of N1 is not limited in this embodiment of the application. table_id is the signaling table identification field, which indicates the identifier of each interactive signaling table included in the interactive feedback message. This is a copy of the table_id field in the interactive signaling table included in the payload of the interactive feedback message. table_version is the signaling table version field, which indicates the version number of each interactive signaling table included in the interactive feedback message. This is a copy of the interactive signaling table version field included in the payload of the interactive feedback message. table_length is the signaling table length field, which indicates the length of each interactive signaling table included in the interactive feedback message. This is a copy of the interactive signaling table length field included in the payload of the interactive feedback message. message_source indicates the message source. 0 indicates that the interactive feedback message is sent from the video client to the server, and 1 indicates that the interactive feedback message is sent from the server to the video client. This value is 0. asset_group_flag is a resource group attribute field, which is used to characterize the subordinate relationship between the first immersive media content and the immersive media content set included in the target resource group. For example, when the field value of the resource group attribute field is a first attribute field value (for example, 1), the resource group attribute field with the first attribute field value is used to characterize that the first immersive media content belongs to the immersive media content set; when the field value of the resource group attribute field is a second attribute field value (for example, 0), the resource group attribute field with the second attribute field value is used to characterize that the first immersive media content does not belong to the immersive media content set. In other words, an asset_group_flag value of 1 indicates that the content currently consumed by the video client (i.e., the first immersive media content) belongs to a resource group (such as the target resource group), and a value of 0 indicates that the content currently consumed by the video client does not belong to any resource group.A resource group refers to a collection of multiple immersive media contents. In embodiments of the present application, an immersive video may include multiple immersive media contents (e.g., first immersive media content). These multiple immersive media contents can be further subdivided into resource groups as needed. For example, the immersive video itself can serve as a resource group, meaning that all immersive media contents in the immersive video belong to one resource group. Alternatively, the immersive video can be divided into multiple resource groups, each of which can include multiple immersive media contents in the immersive video. The asset_group_id field is a resource group identifier, indicating the resource group identifier of the content currently being consumed by the video client, i.e., the identifier of the resource group (e.g., the target resource group) corresponding to the immersive media content set to which the first immersive media content belongs. The asset_id field indicates the identifier of the content currently being consumed by the video client. It should be understood that each immersive media content has a unique corresponding asset_id. When a first immersive media content belongs to a resource group, the video client may be consuming more than one first immersive media content. In this case, feeding back the asset_id of a single first immersive media content is clearly inappropriate. Therefore, the identifiers of the resource groups to which multiple first immersive media contents belong can be fed back. table() is an interactive signaling table entity. The interactive signaling table in the payload appears in the same order as table_id in the extended domain. An interactive signaling table can be used as an instance of table(). Among them, the order of the interactive signaling tables can be sorted according to the corresponding operation timestamps, or according to the table_id corresponding to the interactive signaling table, or other sorting methods can be used. The embodiments of the present application do not limit this. It can be seen that a loop statement is used in the interactive feedback message shown in Table 10, so the service event information carried by one or more interactive signaling tables contained in the interactive feedback message can be fed back in an orderly manner. That is to say, when the interactive feedback message contains multiple interactive signaling tables, the server will read each interactive signaling table in turn according to the order of the interactive signaling tables presented in the loop statement.

[0203] Among them, the above-mentioned signaling table quantity field, signaling table identification field, signaling table version field, signaling table length field, resource group attribute field and resource group identification field are all extended description fields newly added in the system layer of the video client.

[0204] As can be seen from the above, the embodiments of the present application, based on the existing technology, redefine and organize the interactive feedback messages, and add two types of feedback information, namely zoom and event triggering, to the types of interactive feedback to support richer interactive feedback scenarios and improve the accuracy of the video client in obtaining media content during the interactive feedback process.

[0205] Step S204: Send the interaction feedback message to the server, so that the server extracts the interaction signaling table, determines the service event information indicated by the interaction operation according to the information type field in the interaction signaling table, and obtains the second immersive media content for responding to the interaction operation based on the service event information indicated by the interaction operation;

[0206] Specifically, the video client can send an interaction feedback message to the server. After receiving the interaction feedback message, the server can sequentially extract the interaction signaling table from the interaction feedback message, read the information type field from the extracted interaction signaling table, and then determine the service event information indicated by the interaction operation based on the information type field. Finally, based on the service event information indicated by the interaction operation, second immersive media content for responding to the interaction operation can be obtained from the immersive video and returned to the video client. For example, when the field value of the information type field is a first field value, the zoom ratio corresponding to the zoom event can be obtained as service event information; when the field value of the information type field is a second field value, the event label and event status of the switching event can be obtained as service event information; when the field value of the information type field is a third field value, position change information of the service object viewing the first immersive media content belonging to the panoramic video can be obtained as service event information; when the field value of the information type field is a fourth field value, position change information of the service object viewing the first immersive media content belonging to the volumetric video can be obtained as service event information.

[0207] Step S205: receiving the second immersive media content returned by the server, and playing the second immersive media content on the video playback interface.

[0208] Specifically, since what the server returns is actually the second media file resource corresponding to the second immersive media content in the immersive video, the video client can receive the second media file resource returned by the server, and decapsulate and decode the second media file resource, thereby obtaining the playable second immersive media content, and can play it on the video playback interface of the video client.

[0209] As can be seen from the above, during the interaction between the video client and the server, the video client can feedback business event information indicated by different types of interactive operations to the server. It should be understood that the interactive operations here can not only include operations related to the user's location (for example, changes in the user's location), but also include other operations for the immersive media content currently played by the video client (for example, zoom operations). Therefore, through the business key fields carried in the interactive feedback message, the video client can feedback multiple types of business event information to the server. In this way, the server can determine the immersive media content in response to the interactive operation based on these different types of business event information, rather than relying solely on user location information, thereby enriching the information types of the interactive feedback and improving the accuracy of the video client in obtaining media content during the interactive feedback process.

[0210] For further information, see Figure 8 , Figure 8 This is an interactive diagram of an immersive media data processing method provided by an embodiment of the present application. The method can be performed by a decoding device and an encoding device in an immersive media system (for example, a panoramic video system or a volumetric video system). The decoding device can be the above-mentioned Figure 1 The decoding device 100B in the corresponding embodiment may also be the above Figure 3 The decoding device 200B in the corresponding embodiment. The encoding device can be the above Figure 1 The decoding device 100A in the corresponding embodiment may also be the above Figure 3 The decoding device 200A in the corresponding embodiment. The decoding device may be a user terminal integrated with a video client, and the encoding device may be a server. The method may include at least the following steps:

[0211] Step S301: The video client initiates a play request to the server;

[0212] The specific implementation of this step can be found in the above Figure 7 Step S201 in the corresponding embodiment will not be described in detail here.

[0213] Step S302: The server obtains first immersive media content of the immersive video based on the playback request;

[0214] Specifically, based on the target content identifier (i.e., target asset_id) carried in the play request, the server may retrieve, from the immersive video, immersive media content that matches the target content identifier as the first immersive media content. Alternatively, based on the current object position information of the service object carried in the play request, the server may retrieve, from the immersive video, immersive media content that matches the object position information as the first immersive media content.

[0215] Step S303: The server returns the first immersive media content to the video client;

[0216] Step S304: the video client plays the first immersive media content on the video playback interface;

[0217] Step S305: The video client responds to the interactive operation on the first immersive media content and generates an interactive feedback message corresponding to the interactive operation;

[0218] The specific implementation of this step can be found in the above Figure 6 The corresponding step S101 in the embodiment, or refer to the above Figure 7 Step S203 in the corresponding embodiment will not be described in detail here.

[0219] Step S306: The video client sends the interaction feedback message to the server;

[0220] Step S307: The server receives the interactive feedback message sent by the video client;

[0221] Step S308: The server determines the service event information indicated by the interactive operation based on the service key field in the interactive feedback message, and acquires the second immersive media content for responding to the interactive operation based on the service event information indicated by the interactive operation;

[0222] Specifically, after receiving the interaction feedback message, the server can determine the service event information indicated by the interaction operation based on the service key field in the interaction feedback message, and then, based on the service event information indicated by the interaction operation, obtain the second immersive media content used to respond to the interaction operation from the immersive video. It will be understood that when the interaction feedback message is represented in the form of an interaction signaling table, the service key field in the interaction feedback message is the information type field added to the interaction signaling table; when the interaction feedback message is not represented in the form of an interaction signaling table, the service key field is directly added to the interaction feedback message.

[0223] It should be understood that if the first immersive media content belongs to the immersive media content set included in the target resource group, the second immersive media content finally obtained may belong to the same immersive media content set, or the second immersive media content may belong to the immersive media content set included in other resource groups, or the second immersive media content may not belong to the immersive media content set included in any resource group. The embodiments of the present application do not limit this.

[0224] Step S309: the server returns the second immersive media content to the video client;

[0225] Step S310: The video client receives the second immersive media content returned by the server, and plays the second immersive media content on the video playback interface.

[0226] For ease of understanding, the above steps are briefly explained using immersive video T as an example. Assume that a video client requests immersive video T from a server. After the server receives the request (e.g., a play request), it can send the immersive media content T1 (i.e., the first immersive media content) in the immersive video T to the video client based on the request. After the video client receives the immersive media content T1, it can play the immersive media content T1 on the corresponding video playback interface. The business object (e.g., user 1) starts to consume the immersive media content T1 and can generate interactive behavior during the consumption process (i.e., perform interactive operations on the immersive media content T1). As a result, the video client can generate an interactive feedback message corresponding to the interactive behavior and send it to the server. Furthermore, the server receives the interactive feedback message sent by the video client and, based on the message content of the interactive feedback message (e.g., a business key field), can select other immersive media content (i.e., the second immersive media content, e.g., immersive media content T2) from the immersive video T and send it to the video client, so that the business object can experience the new immersive media content. For example, assuming user 1 performs a zoom operation on immersive media content T1, such as enlarging the content of immersive media content T1 with a corresponding zoom ratio of 3, the server can select immersive media content with higher color accuracy (for example, immersive media content T2) from immersive video T based on the zoom ratio indicated by the zoom operation and send it to user 1. For another example, assuming user 1 performs a content switching operation on immersive media content T1, the server can select immersive media content (for example, immersive media content T3) corresponding to an alternative version of the content from immersive video T based on the content switching operation and send it to user 1.

[0227] As can be seen from the above, the embodiments of the present application, based on the existing technology, reorganize and define the interactive feedback messages, and add two types of feedback information, namely zoom and switch (or event triggering), to the types of interactive feedback, so as to support richer interactive feedback scenarios and improve the accuracy of the video client in obtaining media content during the interactive feedback process.

[0228] See Figure 9 , is a schematic diagram of the structure of an immersive media data processing device provided in an embodiment of the present application. The immersive media data processing device can be a computer program (including program code) running on a decoding device. For example, the immersive media data processing device can be an application software in the decoding device; the immersive media data processing device can be used to execute the corresponding steps of the immersive media data processing method provided in an embodiment of the present application. Further, as Figure 9As shown, the immersive media data processing device 1 may include: a message generating module 11, a message sending module 12, and a content receiving module 13;

[0229] The message generation module 11 is configured to generate an interaction feedback message corresponding to the interaction operation in response to the interaction operation on the first immersive media content; the interaction feedback message carries a business key field for describing business event information indicated by the interaction operation;

[0230] A message sending module 12 is configured to send the interaction feedback message to the server, so that the server determines the business event information indicated by the interaction operation based on the business key field in the interaction feedback message, and obtains the second immersive media content used to respond to the interaction operation based on the business event information indicated by the interaction operation;

[0231] The content receiving module 13 is configured to receive the second immersive media content returned by the server.

[0232] The specific implementation of the message generation module 11, the message sending module 12, and the content receiving module 13 can be found in the above Figure 6 Steps S101 to S103 in the corresponding embodiment, or, refer to the above Figure 7 Steps S203 to S205 in the corresponding embodiment will not be described in detail here. In addition, the description of the beneficial effects of the same method will not be repeated here either.

[0233] Optional, such as Figure 9 As shown, the immersive media data processing device 1 may further include: a video request module 14;

[0234] The video request module 14 is configured to respond to a video playback operation for an immersive video in a video client, generate a playback request corresponding to the video playback operation, and send the playback request to a server so that the server obtains the first immersive media content of the immersive video based on the playback request; receive the first immersive media content returned by the server, and play the first immersive media content on a video playback interface of the video client.

[0235] The specific implementation of the video request module 14 can be found in the above Figure 7 Steps S201 and S202 in the corresponding embodiment will not be further described here.

[0236] Among them, the business key field includes a first key field, a second key field, a third key field and a fourth key field; the first key field is used to represent the zoom ratio when the zoom event indicated by the zoom operation is executed when the interactive operation includes a zoom operation; the second key field is used to represent the event label and event status corresponding to the switching event indicated by the switching operation when the interactive operation includes a switching operation; the third key field is used to represent the first object position information of the business object of the first immersive media content belonging to the panoramic video when the interactive operation includes a first position interactive operation; the fourth key field is used to represent the second object position information of the business object of the first immersive media content belonging to the volumetric video when the interactive operation includes a second position interactive operation.

[0237] Further, such as Figure 9 As shown, the message generating module 11 may include: a first determining unit 111, a first adding unit 112, a first generating unit 113, a second determining unit 114, a second adding unit 115, and a second generating unit 116;

[0238] A first determining unit 111 is configured to respond to a triggering operation on the first immersive media content, determine a first information type field of the service event information indicated by the triggering operation, and record an operation timestamp of the triggering operation;

[0239] A first adding unit 112 is configured to add the first information type field and the operation timestamp to an interaction signaling table associated with the first immersive media content, and use the first information type field added to the interaction signaling table as a service key field for describing service event information indicated by the interaction operation;

[0240] The first generating unit 113 is configured to generate an interaction feedback message corresponding to the triggering operation based on the service key field and the operation timestamp in the interaction signaling table.

[0241] In which, when the triggering operation includes a zoom operation, the business event information indicated by the zoom operation is a zoom event, and when the field value of the first information type field corresponding to the zoom operation is the first field value, the field mapped by the first information type field with the first field value is used to represent the zoom ratio when executing the zoom event.

[0242] In which, when the triggering operation includes a switching operation, the business event information indicated by the switching operation is a switching event, and when the field value of the first information type field corresponding to the switching operation is the second field value, the field mapped by the first information type field with the second field value is used to represent the event label and event status of the switching event.

[0243] Among them, when the state value of the event state is the first state value, the event state with the first state value is used to represent that the switching event is in the event trigger state; when the state value of the event state is the second state value, the event state with the second state value is used to represent that the switching event is in the event end state.

[0244] A second determining unit 114 is configured to, upon detecting object position information of a business object viewing the first immersive media content, use a position interaction operation directed to the object position information as an interaction operation in response to the first immersive media content; determine a second information type field of business event information indicated by the interaction operation, and record an operation timestamp of the interaction operation;

[0245] A second adding unit 115 is configured to add the second information type field and the operation timestamp to the interaction signaling table associated with the first immersive media content, and use the second information type field added to the interaction signaling table as a service key field for describing service event information indicated by the interaction operation;

[0246] The second generating unit 116 is configured to generate an interaction feedback message corresponding to the interaction operation based on the service key field and the operation timestamp in the interaction signaling table.

[0247] Among them, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a panoramic video, the field value of the second information type field corresponding to the object position information is a third field value, and the second information type field with the third field value includes a first type of position field, and the first type of position field is used to describe the position change information of the business object watching the first immersive media content belonging to the panoramic video.

[0248] In which, when the first immersive media content is immersive media content in an immersive video, and the immersive video is a volumetric video, the field value of the second information type field corresponding to the object position information is a fourth field value, and the second information type field with the fourth field value includes a second-type position field, and the second-type position field is used to describe position change information of a business object viewing the first immersive media content belonging to the volumetric video.

[0249] Among them, the interactive feedback message also includes an extended description field newly added at the system layer of the video client; the extended description field includes a signaling table quantity field, a signaling table identification field, a signaling table version field, and a signaling table length field; the signaling table quantity field is used to represent the total number of interactive signaling tables included in the interactive feedback message; the signaling table identification field is used to represent the identifier of each interactive signaling table included in the interactive feedback message; the signaling table version field is used to represent the version number of each interactive signaling table; and the signaling table length field is used to represent the length of each interactive signaling table.

[0250] The interactive feedback message also includes a resource group attribute field and a resource group identification field; the resource group attribute field is used to represent the subordinate relationship between the first immersive media content and the immersive media content set included in the target resource group; the resource group identification field is used to represent the identifier of the target resource group.

[0251] Among them, when the field value of the resource group attribute field is the first attribute field value, the resource group attribute field with the first attribute field value is used to represent that the first immersive media content belongs to the immersive media content set; when the field value of the resource group attribute field is the second attribute field value, the resource group attribute field with the second attribute field value is used to represent that the first immersive media content does not belong to the immersive media content set.

[0252] The specific implementation of the video request module 14 can be found in the above Figure 6 Step S101 in the corresponding embodiment will not be further described here.

[0253] See Figure 10 , is a schematic diagram of the structure of an immersive media data processing device provided in an embodiment of the present application. The immersive media data processing device can be a computer program (including program code) running on an encoding device. For example, the immersive media data processing device can be an application software in the encoding device; the immersive media data processing device can be used to execute the corresponding steps of the immersive media data processing method provided in an embodiment of the present application. Further, as Figure 10 As shown, the immersive media data processing device 2 may include: a message receiving module 21, a content acquiring module 22, and a content returning module 23;

[0254] A message receiving module 21 is configured to receive an interaction feedback message sent by a video client; the interaction feedback message is a message corresponding to an interaction operation generated by the video client in response to an interaction operation on the first immersive media content; the interaction feedback message carries a business key field for describing business event information indicated by the interaction operation;

[0255] The content acquisition module 22 is configured to determine the business event information indicated by the interactive operation based on the business key field in the interactive feedback message, and acquire the second immersive media content for responding to the interactive operation based on the business event information indicated by the interactive operation;

[0256] The content returning module 23 is configured to return the second immersive media content to the video client.

[0257] The specific implementation of the message receiving module 21, the content obtaining module 22, and the content returning module 23 can be found in the above Figure 8Steps S307 to S309 in the corresponding embodiment will not be described in detail here. In addition, the description of the beneficial effects of adopting the same method will not be described in detail either.

[0258] See Figure 11 , is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 11 As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 11 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0259] In such Figure 11 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to execute the above Figure 6 、 Figure 7 、 Figure 8 The description of the data processing method of the immersive media in any corresponding embodiment can also be performed as described above. Figure 9 The description of the data processing device 1 for the immersive media in the corresponding embodiment can also execute the aforementioned Figure 10 The description of the immersive media data processing device 2 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here.

[0260] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the immersive media data processing device 1 or the immersive media data processing device 2 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above-mentioned Figure 6 、 Figure 7 、 Figure 8 The description of the data processing method for immersive media in any corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiments involved in this application, please refer to the description of the method embodiments of this application.

[0261] The computer-readable storage medium may be the data processing device of the immersive media provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0262] In addition, it should be noted that the present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned Figure 6 、 Figure 7 、 Figure 8 The method provided by any corresponding embodiment. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer program product or computer program embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0263] For further information, see Figure 12 , Figure 12Schematic diagram of a data processing system provided by an embodiment of the present application. The data processing system 3 may include a data processing device 1a and a data processing device 2a. The data processing device 1a may be the above-mentioned Figure 9 The data processing device 1 for immersive media in the corresponding embodiment can be understood as follows: the data processing device 1a can be integrated into the above-mentioned Figure 1 The decoding device 100B in the corresponding embodiment or the above Figure 3 The decoding device 200B in the corresponding embodiment is not described in detail here. Figure 10 The data processing device 2 for immersive media in the corresponding embodiment can be understood as follows: the data processing device 2a can be integrated into the above-mentioned Figure 1 The encoding device 100A in the corresponding embodiment or the above Figure 3 In the encoding device 200A in the corresponding embodiment, therefore, no further description will be given here. In addition, the description of the beneficial effects of adopting the same method will not be given in detail. For technical details not disclosed in the data processing system embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0264] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0265] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0266] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A data processing method for immersive media, characterized in that: The method is executed by a video client and includes: In response to an interactive operation on first immersive media content, generating an interactive feedback message corresponding to the interactive operation; the interactive feedback message carries an interactive signaling table associated with the interactive operation, and the interactive signaling table includes an information type field for describing service event information indicated by the interactive operation; the information type field is used to represent position change information of a service object viewing the first immersive media content when the interactive operation includes a position interactive operation; the position interactive operation is an operation on object position information of the service object; sending the interaction feedback message to a server, so that the server extracts the interaction signaling table from the interaction feedback message, determines the service event information indicated by the interaction operation according to the information type field in the interaction signaling table, and acquires second immersive media content for responding to the interaction operation based on the service event information indicated by the interaction operation; Receive the second immersive media content returned by the server.

2. The method according to claim 1, characterized in that The step of generating, in response to an interactive operation on the first immersive media content, an interactive feedback message corresponding to the interactive operation includes: In response to an interactive operation on the first immersive media content, determining an information type field of service event information indicated by the interactive operation, and recording an operation timestamp of the interactive operation; Adding the information type field and the operation timestamp to an interaction signaling table associated with the first immersive media content, and using the information type field added to the interaction signaling table as a service key field for describing service event information indicated by the interaction operation; Based on the service key field and the operation timestamp in the interaction signaling table, an interaction feedback message corresponding to the interaction operation is generated.

3. The method according to claim 1, characterized in that The interactive operation includes one or more of the position interactive operation, zoom operation and switching operation; the position interactive operation is a first position interactive operation or a second position interactive operation; the zoom operation refers to an operation of reducing or enlarging the screen size of the first immersive media content; the switching operation includes one or more of a playback rate switching operation, a picture quality switching operation, a flip operation, and a content switching operation for the first immersive media content; the interactive feedback message includes one or more interactive signaling tables; and one interactive signaling table corresponds to one interactive operation.

4. The method according to claim 3, characterized in that The information type field is used to represent the type of information carried by the interactive signaling table; when the field value of the information type field is the first field value, the service event information includes a zoom ratio corresponding to the zoom event indicated by the zoom operation; When the field value of the information type field is the second field value, the service event information includes an event label and an event status corresponding to the switching event indicated by the switching operation.

5. The method according to claim 1, wherein The information type field is used to represent the type of information carried by the interactive signaling table; when the field value of the information type field is the third field value, the service event information includes position change information of the service object viewing the first immersive media content belonging to the panoramic video; When the field value of the information type field is the fourth field value, the service event information includes position change information of a service object for viewing the first immersive media content belonging to the volumetric video.

6. The method according to claim 5, characterized in that When the field value of the information type field is the third field value, the information type field with the third field value includes a first type of location field; the first type of location field includes a field for indicating 3DoF+ video content, an interaction target field, and an interaction type field; the interaction target field is used to indicate the interaction target of the video client; the interaction target includes any one of the current state of the helmet device, the current state of the business object focus area, and the current state of the business object.

7. The method according to claim 5, characterized in that When the field value of the information type field is the fourth field value, the information type field with the fourth field value includes a second type of position field; the second type of position field is used to indicate the position of the business object in the global coordinate system and the viewing direction of the business object in the Cartesian coordinate system.

8. The method according to claim 1, characterized in that The interaction feedback message carries a business key field for describing business event information indicated by the interaction operation; when the first immersive media content is a panoramic video, the business key field carried in the interaction feedback message includes one or more of a first key field, a second key field, and a third key field; when the first immersive media content is a volumetric video, the business key field carried in the interaction feedback message includes one or more of the first key field, the second key field, and a fourth key field; the first key field is used to represent, when the interaction operation includes a zoom operation, a zoom ratio when executing a zoom event indicated by the zoom operation; The second key field is used to represent the event label and event status corresponding to the switching event indicated by the switching operation when the interactive operation includes a switching operation; the third key field is used to represent the first object position information of the business object of the first immersive media content belonging to the panoramic video when the interactive operation includes a first position interactive operation; the fourth key field is used to represent the second object position information of the business object of the first immersive media content belonging to the volumetric video when the interactive operation includes a second position interactive operation.

9. The method according to claim 8, characterized in that The interaction feedback message further includes an information identification field, and the information identification field is used to indicate the information type of the service event information indicated by the interaction operation.

10. The method according to claim 1, characterized in that Also includes: In response to a video playback operation for an immersive video in a video client, generating a playback request corresponding to the video playback operation, and sending the playback request to a server, so that the server obtains a first media file resource corresponding to a first immersive media content in the immersive video based on the playback request; The first media file resource returned by the server is received, the first media file resource is decapsulated and decoded to obtain the first immersive media content, and the first immersive media content is played on a video playback interface of the video client.

11. A data processing method for immersive media, characterized in that: The method is executed by a server and includes: Receiving an interaction feedback message sent by a video client; the interaction feedback message is a message corresponding to an interaction operation generated by the video client in response to an interaction operation on first immersive media content; the interaction feedback message carries an interaction signaling table associated with the interaction operation, and the interaction signaling table includes an information type field for describing service event information indicated by the interaction operation; the information type field is used to represent position change information of a service object viewing the first immersive media content when the interaction operation includes a position interaction operation; the position interaction operation is an operation on object position information of the service object; extracting the interaction signaling table from the interaction feedback message, determining service event information indicated by the interaction operation according to the information type field in the interaction signaling table, and acquiring second immersive media content for responding to the interaction operation based on the service event information indicated by the interaction operation; The second immersive media content is returned to the video client.

12. The method according to claim 11, characterized in that The extracting the interaction signaling table from the interaction feedback message, determining the service event information indicated by the interaction operation according to the information type field in the interaction signaling table, and acquiring the second immersive media content for responding to the interaction operation based on the service event information indicated by the interaction operation includes: extracting the interaction signaling table from the interaction feedback message in the order of the interaction signaling table, reading the information type field from the extracted interaction signaling table, and determining the service event information indicated by the interaction operation according to the information type field; the order of the interaction signaling table is determined by the operation timestamp of the interaction operation or the identifier of the interaction signaling table; Acquire, based on the service event information indicated by the interactive operation, a second media file resource corresponding to the second immersive media content for responding to the interactive operation; Then, returning the second immersive media content to the video client includes: The second media file resource carrying the second immersive media content is returned to the video client, so that the video client decapsulates and decodes the second media file resource to obtain the second immersive media content, and plays the second immersive media content on a video playback interface of the video client.

13. A computer device, characterized in that: include: processor and memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program to enable the computer device to execute the method according to any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 12.

15. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. The computer instructions are suitable for being read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method for adjusting video and terminal

    CN106992004A

  • Method, device and system for processing media information

    CN108616751A