Video rendering processing method, system, computer-readable storage medium, electronic device, and computer program product

By coordinating perception between the XR terminal and the radio access network (RAN), and utilizing multi-base station collaboration and sensor data fusion, the problem of wide-area XR rendering is solved, achieving high-quality, continuous video rendering effects suitable for various scenarios.

WO2026130063A1PCT designated stage Publication Date: 2026-06-25ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ZTE CORP
Filing Date
2025-11-26
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve wide-area rendering of extended reality (XR), especially in large spaces and complex environments, resulting in discontinuous and delayed video rendering effects for XR devices.

Method used

By coordinating perception among the XR terminal, the radio access network (RAN), and the core network, and utilizing multi-base station collaboration, multiple-input multiple-output (MIMO) technology, and sensor data fusion, the location and attitude information of the XR terminal can be obtained in real time, and video rendering can be performed in combination with the perception data.

Benefits of technology

It achieves high-quality, continuous rendering of XR videos in a wide-area environment, reduces network latency, and is suitable for outdoor activities, emergency rescue, and other scenarios, providing an immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025137923_25062026_PF_FP_ABST
    Figure CN2025137923_25062026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a video rendering processing method, a system, a computer-readable storage medium, an electronic device, and a computer program product. The method comprises: receiving a sensing request initiated by an XR terminal; sending the sensing request to a sensing function (SF) of a core network, wherein the sensing request is used for requesting the SF to perform sensing and communication sensing on an XR service of the XR terminal; receiving sensing data of the XR terminal from the SF; and rendering an XR video of the XR terminal on the basis of the sensing data. The method can solve the problem in the related art of how to perform wide-area rendering on an XR video, and can realize wide-area XR video rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Video rendering processing methods, systems, computer-readable storage media, electronic devices and computer program products

[0001] Cross-reference to related applications

[0002] This disclosure is based on and claims priority to Chinese Patent Application No. 2024119170313, filed on December 20, 2024, entitled “Video Rendering Processing Method, System, Computer-Readable Storage Medium, Electronic Device and Computer Program Product”, and incorporates the entire contents of that patent application by reference. Technical Field

[0003] This disclosure relates to the field of communications, and more specifically, to a video rendering processing method, system, computer-readable storage medium, electronic device, and computer program product. Background Technology

[0004] With the continuous enrichment and evolution of Extended Reality (XR) applications, the future holds scenarios such as large-scale Virtual Reality (VR) and wide-area Augmented Reality (AR). Current video rendering, typically combined with head-mounted displays' gyroscopes and gravimeters, is suitable for localized XR posture perception and motion. However, wide-area XR may construct a digital world parallel to the physical world, or a large-scale virtual space. How to solve the problem of wide-area XR rendering, allowing XR to transcend localized limitations, is a problem that needs to be addressed. Summary of the Invention

[0005] This disclosure provides a video rendering processing method, system, computer-readable storage medium, electronic device, and computer program product to at least solve the problem of how to perform wide-area rendering of XR videos based on synesthesia in related technologies.

[0006] According to one embodiment of this disclosure, a video rendering processing method is provided, applied to an extended reality (XR) platform, comprising:

[0007] Receive perception requests initiated by XR terminals;

[0008] The sensing request is sent to the sensing function SF of the core network, wherein the sensing request is used to request the SF to perform sensory sensing of the XR service of the XR terminal.

[0009] Receive sensing data from the XR terminal from the SF;

[0010] The XR video of the XR terminal is rendered based on the perceived data.

[0011] According to another embodiment of this disclosure, a video rendering processing method is provided, applied to a radio access network (RAN), comprising:

[0012] The core network receives perception requests, wherein the perception requests are sent by the XR terminal to the extended reality XR platform and then sent to the core network by the XR platform.

[0013] The XR services of the XR terminal are sensed by the sensor, and the sensed data of the XR terminal is sent to the XR platform through the core network to render the XR video of the XR terminal.

[0014] According to another embodiment of this disclosure, a video rendering processing system is provided, including: an XR platform, a radio access network (RAN), and a core network.

[0015] The XR platform is used to receive perception requests initiated by XR terminals and send the perception requests to the perception function SF of the core network.

[0016] The core network is used to receive the sensing request through the SF, send the sensing request to the radio access network RAN, receive the sensing data of the XR terminal through the radio access network RAN, and send the sensing data to the XR platform.

[0017] The XR platform is also used to render the XR video of the XR terminal based on the perceived data.

[0018] According to yet another embodiment of this disclosure, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.

[0019] According to yet another embodiment of this disclosure, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0020] According to yet another embodiment of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments. Attached Figure Description

[0021] Figure 1 is an architecture diagram of synesthesia-based XR rendering according to an embodiment of the present disclosure;

[0022] Figure 2 is a flowchart of video rendering processing according to an embodiment of the present disclosure;

[0023] Figure 3 is a flowchart of video rendering processing according to an embodiment of the present disclosure;

[0024] Figure 4 is a schematic diagram of the structure of a video rendering processing system according to an embodiment of the present disclosure;

[0025] Figure 5 is a flowchart of the perception data reporting of an XR terminal according to an embodiment of the present disclosure;

[0026] Figure 6 is a flowchart of synesthetic XR video co-rendering according to an embodiment of the present disclosure;

[0027] Figure 7 is a flowchart of remote collaborative processing of XR video rendering based on synesthesia according to an embodiment of the present disclosure. Detailed Implementation

[0028] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings and examples.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0030] This disclosure embodiment can operate on the network architecture shown in Figure 1. Figure 1 is an architecture diagram of XR rendering based on sensing according to this disclosure embodiment. As shown in Figure 1, the XR rendering architecture based on sensing includes three parts: XR terminal, 5G / 6G network, and XR application. The XR terminal type can be AR / MR glasses, VR headset, naked-eye 3D tablet, or even AI mobile phone, smart cockpit, humanoid robot, etc. The terminal supports sensing and can determine the position of the XR terminal or the posture of the person in real time. The 5G / 6G network provides sensing function, supports sensing the XR terminal, obtains the position information of the XR terminal, and supports high-speed transmission of XR video. The XR application specifically includes a media platform (also called an XR platform), spatial computing, and video rendering. The platform supports rendering XR based on the position and posture obtained by sensing, combined with the position information reported by the XR terminal and the collected video.

[0031] This embodiment provides a video rendering processing method running on the aforementioned mobile terminal or network architecture. Figure 2 is a flowchart of video rendering processing according to an embodiment of this disclosure. As shown in Figure 2, it is applied to an extended reality (XR) platform, and the process includes the following steps:

[0032] Step S202: Receive a perception request initiated by the XR terminal;

[0033] When an XR terminal needs to render video, it sends a perception request to the XR platform.

[0034] Step S204: Send a perception request to the perception function SF of the core network, wherein the perception request is used to request the SF to perform sensory perception of the XR services of the XR terminal.

[0035] Step S206: Receive sensing data from the XR terminal from the SF;

[0036] The sensing data in step S206 above can be obtained by the XR terminal through sensing, or by the RAN through sensing, or by the XR terminal and the RAN working together to obtain sensing data.

[0037] Step S208: Render the XR video of the XR terminal based on the perception data.

[0038] Through the above steps S202 to S208, the XR platform can acquire the perception data of the XR terminal in real time, and render the XR video of the XR terminal using the perception data. This can solve the problem of how to perform wide-area rendering of XR video in related technologies, and realize wide-area XR video rendering.

[0039] The XR terminals in this disclosure include, but are not limited to, augmented reality (AR) glasses, mixed reality (MR) glasses, virtual reality (VR) headsets, glasses-free 3D tablets, artificial intelligence (AI) mobile phones, smart cockpits, or humanoid robots. This broad terminal compatibility enables the technology to be applied to various XR devices, meeting the needs of different user groups, such as applications in industrial maintenance, telemedicine, and virtual conferencing scenarios.

[0040] XR terminals include various types, including but not limited to AR glasses, MR glasses, VR headsets, glasses-free 3D tablets, AI phones, smart cockpits, or humanoid robots. This broad terminal support enables the system to cover a wide range of fields, from personal entertainment to industrial applications, such as in gaming, design, and manufacturing.

[0041] In one embodiment of this disclosure, after sending a sensing request to the sensing function SF of the core network, the sensing request can be sent to the radio access network (RAN) through the access and mobility management function (AMF) or the network control unit (NCU); or the sensing request can be sent directly to the radio access network RAN. After receiving the sensing request, the RAN can cooperate with the XR terminal to sense the XR terminal in order to obtain sensing data.

[0042] The collaborative sensing between the RAN and the XR terminal can specifically include: after receiving a sensing request, the base station in the RAN begins signal interaction with the XR terminal, including but not limited to transmitting specific detection signals, receiving feedback signals from the XR terminal, and analyzing the characteristics of these signals, such as signal strength, time of arrival, and phase changes. The XR terminal also participates in signal measurement and feedback, and may need to adjust its antenna direction or power to cooperate with the base station's detection. In wide-area coverage scenarios, multiple base stations may participate in sensing the XR terminal. Multiple base stations use MIMO (Multiple-Input Multiple-Output) technology to collaboratively transmit and receive signals to improve the accuracy and coverage of sensing. Multiple base stations share sensing data and use data fusion algorithms to determine the precise location of the XR terminal. The base station uses its multiple sensors (such as radar and cameras) to collect information about the XR terminal. For example, radar can provide a rough location of the XR terminal, while cameras can use image recognition technology to help determine the precise attitude and orientation of the XR terminal. Based on signal interaction data and information collected by multiple sensors, the base station uses positioning algorithms (such as triangulation and multi-base station positioning) to calculate the absolute position (i.e., location information), speed, altitude, and direction of motion (X-axis, Y-axis, Z-axis) of the XR terminal, thus obtaining the sensing data.

[0043] The collaborative perception mechanism can ensure the real-time nature and accuracy of perception data, and maintain high-quality rendering of XR videos even in fast-moving or complex environments. It is particularly suitable for scenarios such as outdoor activities, live sports events, and emergency rescue.

[0044] The XR terminal and RAN report the sensed data they obtain to the sensing function module (SF) of the core network, which then reports it to the XR platform. The SF is responsible for aggregating and processing the sensed data collected from multiple base stations, ensuring data consistency and integrity. In one embodiment of this disclosure, step S206 may specifically include: if the RAN receives a sensed request, identifies the XR terminal based on the terminal identifier in the sensed request, and collaborates with the XR terminal to perform sensory perception, obtains sensed data, and then reports it to the SF, then the XR platform receives the sensed data reported by the RAN from the SF; if the XR terminal receives a sensed request, collaborates with the RAN to perform sensory perception, obtains sensed data, and then reports the sensed data to the SF, then the XR platform receives the sensed data reported by the XR terminal from the SF. This bidirectional data transmission mechanism ensures the comprehensiveness and reliability of the sensed data, providing a data foundation for personalized rendering of XR videos.

[0045] In one embodiment of this disclosure, step S208 may specifically include: receiving rotation information reported by the XR terminal; if the XR terminal has sensing service permissions, performing XR video rendering based on the rotation information reported by the XR terminal and the sensing data; furthermore, firstly determining the target position in the digital space of the XR video based on the position information in the sensing data, specifically, converting the latitude and longitude coordinates in the position information into target position coordinates in the digital space. For example, if the XR service involves a virtual stadium, then the latitude and longitude coordinates need to be matched with the coordinate system of the virtual stadium to determine the specific position of the XR terminal in the virtual scene (i.e., the target position); then obtaining the viewpoint to be rendered based on the position information in the sensing data and the rotation information reported by the XR terminal. If the sensing data also includes translation information, direction of motion, speed, and height, at least one of the translation information, direction of motion, speed, and height can be combined to determine the viewpoint to be rendered. Specifically, combining the rotation information reported by the XR terminal with the position information and rotation information in the sensing data, the current viewpoint of the XR terminal is calculated, thus obtaining the viewpoint to be rendered. This calculation process determines the current direction and angle facing the XR terminal user, as well as the viewpoint that should be displayed in the virtual world relative to their position in the physical world. Then, based on the target position and the viewpoint to be rendered, information about the virtual objects to be rendered in the digital space is obtained. Specifically, in the digital space, based on the target position and the viewpoint, it is determined which virtual objects should appear in the XR video. This may involve querying and analyzing the 3D model of the digital space to find virtual objects that match the current position and viewpoint. Detailed information about the selected virtual objects is extracted from the digital space database, resulting in the information about the virtual objects to be rendered, which may include their appearance, position, and dynamic characteristics. This information will be used for subsequent XR video rendering. Finally, the virtual object information is superimposed onto the XR video to obtain the rendered video. Specifically, the extracted virtual object information is superimposed on the real-world XR video to achieve XR video fusion rendering. After the virtual object superposition is completed, the rendered video is generated. This step may require further optimization processing, such as video compression and format conversion, to adapt to network transmission and the display requirements of XR terminals. Rendering strategies based on target location and viewpoint can provide a more natural and smooth XR video experience, suitable for scenarios that require high-precision positioning and viewpoint matching, such as virtual reality navigation and remote virtual operation.

[0046] The sensing data in this embodiment includes at least one of the following: translation information, position information, direction of motion, speed, and altitude; the rotation information reported by the XR terminal includes at least one of the following: rotation information about the X-axis, rotation information about the Y-axis, and rotation information about the Z-axis. The rotation information may specifically include pitch about the X-axis, yaw about the Y-axis, and roll about the Z-axis. This information is usually collected by sensors (such as gyroscopes and accelerometers) inside the XR terminal.

[0047] This embodiment renders video from a specific perspective based on the user's viewpoint and location, ensuring that the user sees a virtual scene of their current location. The XR platform dynamically renders the video based on real-time perception data, including the position and angle of virtual objects and their degree of integration with the real world, ensuring that the image is synchronized with the user's movement. The XR platform encodes the rendered video, and the encoded video is transmitted to the RAN (Radio Access Network) via UPF (User Plane Function), and then pushed to the XR terminal by the RAN. After receiving the rendered video, the XR terminal displays it to the user through its built-in display device (such as a head-up display, glasses, etc.), achieving an immersive experience.

[0048] Through the above process, XR video rendering can adapt to the movement of the terminal in a wide-area environment, providing a continuous and lag-free immersive experience. Whether it is VR virtual world exploration or AR real world augmentation, high-quality video presentation can be obtained.

[0049] The perception request in this embodiment carries at least one of the following: service type, service requirements, perceived target information (e.g., terminal identifier), and perception node information. The service type can be XR wide-area perception, indicating that the perception request is used to request XR wide-area perception. The perception node information is an XR terminal and / or RAN, meaning that XR perception is performed through the XR terminal and / or RAN to obtain perception data. The service requirements include at least one of the following: perception resolution, perception accuracy, frame rate, duration, target area information, and latency. The detailed information carried in the perception request enables fine-grained control of the perception process, ensuring the quality and efficiency of XR video rendering. This is suitable for scenarios requiring highly customized services, such as virtual reality live streaming and customized AR advertising.

[0050] The aforementioned perception resolution refers to the perception system's ability to represent details of the physical world in digital space. In XR applications, perception resolution determines the size or level of detail of objects that can be identified and tracked. Perception accuracy refers to the accuracy with which the perception system measures position, orientation, or object state. Frame rate refers to the speed at which the perception system updates data, typically measured in frames per second (FPS). Duration defines the length of time the perception service needs to run. Target area information specifies the geographical area or specific space that the perception service needs to cover. Latency is the time it takes for perception data to be collected, processed, and returned to the XR terminal. Perceived target information refers to the terminal identifier of the object or user (i.e., the XR terminal) that needs to be perceived, such as IMSI (International Mobile Subscriber Identity), temporary identifiers, etc.

[0051] Another aspect of this disclosure provides a video rendering processing method. Figure 3 is a flowchart of a video rendering processing method according to an embodiment of this disclosure. As shown in Figure 3, the method is applied to a Radio Access Network (RAN) and includes the following steps:

[0052] Step S302: Receive a perception request through the core network, wherein the perception request is sent by the XR platform to the core network after being sent by the XR terminal to the extended reality XR platform;

[0053] Step S304: Perform sensory perception on the XR services of the XR terminal and send the perception data of the XR terminal to the XR platform through the core network to render the XR video of the XR terminal.

[0054] RAN's sensing capabilities not only improve the efficiency of acquiring sensing data, but also adapt to different network environments and service requirements, providing solid technical support for high-quality rendering of XR videos.

[0055] One aspect of this disclosure provides a video rendering processing system. FIG4 is a schematic diagram of the structure of the video rendering processing system according to an embodiment of this disclosure. As shown in FIG4, it includes: an XR platform, a radio access network (RAN), and a core network.

[0056] The XR platform is used to receive perception requests initiated by XR terminals. The perception request carries at least one of the following: service type, service requirements, information of the perceived target, and information of the perception node. The perception request is sent to the perception function SF of the core network.

[0057] The core network is used to receive sensing requests through the SF, send the sensing requests to the radio access network RAN, receive sensing data from XR terminals through the radio access network RAN, and send the sensing data to the XR platform.

[0058] The XR platform is also used to render XR videos from XR terminals based on perception data.

[0059] This system architecture not only optimizes data transmission paths and reduces network latency, but also enables efficient processing and distribution of sensing data through the sensing function SF of the core network, providing a powerful technical platform for the widespread application and innovation of XR services.

[0060] In one embodiment of this disclosure, the RAN is used to receive a sensing request sent by the Access and Mobility Management Function (AMF) or the Network Control Unit (NCU) through the core network; or directly to receive a sensing request sent by the SF through the core network; and to cooperate with the XR terminal to sense the XR terminal according to the sensing request.

[0061] RAN's sensing mechanism can quickly respond to the dynamic changes of XR terminals, and maintain the accuracy and real-time nature of sensing data even in high-density user environments, making it possible for large-scale XR applications, such as providing high-quality XR video services in large exhibitions, concerts and other occasions.

[0062] A system for XR rendering and delivery based on synesthetic fusion can perform wide-area XR rendering with the assistance of synesthesia, specifically including:

[0063] The XR terminal sends a perception request to the XR platform to request that the XR terminal be perceived. The perception request carries the service type, service requirements, and information about the specified perception node (XR terminal or RAN).

[0064] If the XR platform determines whether the current XR terminal has the permission to trigger the perception service and needs to trigger wide-area XR services, then the XR platform will trigger the perception of the XR terminal to the core network perception function (SF).

[0065] The SF sends the XR terminal's perception request to the RAN either through control network elements (such as AMF, NCU) or directly. The RAN or the XR terminal reports the XR terminal's perception data according to the perception request and transmits it to the XR platform.

[0066] The XR platform combines rotation information reported by the XR terminal with sensory data from synesthesia to render XR videos.

[0067] Furthermore, the XR platform controls triggering permissions by setting whether the XR terminal supports the triggering capability of the sensing service. If the XR terminal does not have sensing service permissions, the XR platform will only render video based on the position and rotation information reported by the XR terminal.

[0068] In one embodiment, the perception of the XR terminal includes, but is not limited to, the base station perceiving all XR terminals in the area and then identifying the XR terminal that needs to be based on sensor rendering through the perceived target information (such as IMSI, temporary identifier, etc.) in the perception request. Alternatively, the XR terminal may cooperate with the base station in the RAN to perform perception according to the RAN's perception request and directly report the perception data.

[0069] In one embodiment, the XR platform performs video rendering, including but not limited to combining the rotation information reported by the XR terminal with the translation information, position information, and direction of motion perceived by synesthesia, to perform wide-area, continuous video rendering in outdoor or high-speed scenes.

[0070] For XR scenarios, XR terminals can capture real-world video and send it to the XR platform. The XR platform combines the sensory data obtained from perception with the virtual object information to form a video that blends the virtual and real worlds.

[0071] The video rendering provided in this disclosure not only achieves high precision and efficiency in XR video rendering, but also significantly enhances the flexibility and adaptability of XR services by introducing sensing and collaborative mechanisms. This technical solution can meet the needs of different industries and scenarios, promoting the widespread application and in-depth development of XR technology, and providing users with a richer, more realistic, and immersive XR experience. Furthermore, by using the sensing function (SF) of the core network as an intermediate bridge, data transmission between the XR platform and the RAN is effectively coordinated, reducing network latency and improving the transmission efficiency of sensing data. This provides a solid technical foundation for the widespread application of XR services and also offers new ideas and directions for the design of future 5G and 6G network architectures.

[0072] This embodiment of extended reality (XR) content rendering and delivery based on a communication network requests the activation of a sensing service from the communication network to collect and fuse wide-area location information, translation information, etc., of the XR terminal; renders XR content based on the fused information to generate a video stream; and pushes the video stream to the XR terminal through the user plane function and radio access network (RAN) of the communication network. By accurately sensing the location information and translation information of the XR terminal, precise positioning and real-time rendering of XR content in outdoor environments can be achieved, providing users with a more immersive XR experience, such as in tourism, education, and entertainment scenarios.

[0073] Figure 5 is a flowchart of the perception data reporting of an XR terminal according to an embodiment of the present disclosure. As shown in Figure 5, it includes:

[0074] S501, the XR terminal / UE initiates a perception request to the XR platform, carrying the service type (e.g., XR wide-area perception), service requirements (e.g., perception resolution, perception accuracy, frame rate, duration, target area information, latency, perceived target information (XR terminal information)), and specified perception node information (e.g., XR terminal and RAN).

[0075] In S502, if the XR platform determines that the XR terminal is a registered user and needs to trigger wide-area XR services, it sends a sensing request to the core network sensing function (SF) to sense the XR terminal. If the XR platform is a third-party application, the request is initiated through a capability open platform, such as NEF; if the XR platform is a trusted platform of the operator, the sensing request is directly initiated to the SF. The request carries the service type (e.g., XR wide-area sensing), service requirements (e.g., sensing resolution, sensing accuracy, frame rate, duration, target area information, latency, information of the sensed target (e.g., XR terminal)), and information of the specified sensing node (e.g., XR terminal and / or RAN), etc.

[0076] S503, SF sends a perception request to AMF / NCU, specifying the XR terminal to be perceived, service type, service requirements, etc.

[0077] S504, the AMF / NCU sends a perception request to the RAN, specifying the XR terminals, service types, service requirements, etc. that need to be perceived.

[0078] S505, SF can directly send a perception request to RAN, specifying the XR terminal to be perceived, service type, service requirements, etc.

[0079] In S506, the XR terminal and the RAN work together to sense the XR terminal and acquire its sensing data. This sensing data can specifically include location information, translation information such as speed, altitude, and direction of movement. The XR terminal's sensing methods include, but are not limited to:

[0080] The RAN senses all XR terminals in the region, obtains point cloud information of the RAN coverage area, and obtains the current XR terminal's location information, translation information, etc. by analyzing the point cloud. Then, it filters and identifies the XR terminals that need to be based on synesthetic rendering by using the terminal identifier (such as IMSI, temporary identifier, etc.) in the sensing request.

[0081] The XR terminal, in accordance with the RAN's sensing request, collaborates with the RAN to perform sensing. The XR terminal and the RAN exchange information, and based on the time delay, direction, angle, etc., transmitted with the RAN, the XR terminal calculates its relative position information to the RAN, and then combines this with the position of the base station in the RAN to convert it into absolute position information in the physical world, including latitude, longitude, speed, etc.

[0082] S507, RAN reports the sensing data to SF.

[0083] S508, the XR terminal reports the perceived data to SF.

[0084] S509, SF reports the perceived data to the XR platform.

[0085] Through the above steps S501 to S507, after triggering the XR service through the perception request, the XR terminal and RAN will report the perception data obtained based on the perception request to the SF of the core network, and then report it to the XR platform through the SF, providing a data foundation for the XR platform to perform XR video rendering.

[0086] Figure 6 is a flowchart of sensor-based XR video joint rendering according to an embodiment of the present disclosure, as shown in Figure 6, including:

[0087] S601, the XR terminal or RAN will report the sensing data to SF.

[0088] S602, SF reports the perceived data to the XR platform.

[0089] S603, the XR terminal reports rotation information to the XR platform through RAN and UPF respectively. The rotation information includes, but is not limited to, rotation around the X-axis (pitch), rotation around the Y-axis (yaw), and rotation around the Z-axis (roll).

[0090] The S604 XR platform combines sensory data acquired through synesthesia with rotational information reported by the XR terminal to render VR videos.

[0091] The methods used by XR platforms to render video include, but are not limited to:

[0092] If the speed data from the perception data indicates that the current XR scene is low-speed (such as walking), the XR platform combines the rotation information from the XR terminal with the perception data for rendering. First, the XR platform determines the target location in the digital world of the video based on the position information in the perception data, such as a shopping mall, sports center, or museum, and obtains information about virtual objects in the wide-area digital space to be rendered. Combining the rotation information reported by the XR terminal, the position information, translation information, and direction of motion from the perception data, the platform calculates the current XR terminal's viewpoint, i.e., the viewpoint to be rendered. In the digital space, based on the target location and the viewpoint to be rendered, it determines which virtual objects should appear in the XR video and overlays the virtual object information onto the XR video to obtain the rendered video.

[0093] By combining latitude, longitude, speed and other data from the sensing data to determine whether the current XR terminal is on a high-speed train or in a smart cockpit of a car, the system can quickly predict the forward position based on speed, location information and direction of movement, and perform continuous video rendering in advance to avoid dizziness caused by video stuttering or delay.

[0094] S605, the XR platform sends the rendered video to the UPF.

[0095] S606, UPF sends the rendered video data to RAN.

[0096] S607, the RAN sends the rendered data to the XR terminal.

[0097] Through the above steps S601 to S607, the joint rendering of information reported by the sensor and the XR terminal can realize XR large space services, enabling users to move in a large-scale digital space.

[0098] If the XR terminal is AR / MR glasses, in the industrial internet, multiple designers in different locations can access the same digital space through AR / MR glasses. They can collaboratively design and operate virtual objects in the digital space, such as airplanes and cars, achieving a "face-to-face, shoulder-to-shoulder, hand-in-hand" effect. Figure 7 is a flowchart of remote collaborative processing of XR video rendering based on synesthesia according to an embodiment of this disclosure. As shown in Figure 7, it includes:

[0099] S701, XR terminal or RAN reports the sensing data to SF.

[0100] S702, SF reports the perceived data to the XR platform.

[0101] The S703 XR terminal uses a camera to capture real-world video of its surrounding environment and sends it to the XR platform via the User Plane Function (UPF).

[0102] The S704 XR platform acquires virtual object information based on perception data, overlays it onto real video, and performs video rendering to obtain the rendered video. The specific rendering process is similar to the rendering process described above, and will not be repeated here.

[0103] S705, the XR platform sends the rendered video to the UPF.

[0104] S706, UPF forwards the rendered video to RAN.

[0105] S707, RAN forwards the rendered video to the XR terminal.

[0106] Through the above S701 to S707, synesthesia can better determine the position of AR / MR glasses in the real world, and also better determine the position of AR / MR glasses in the virtual world, thereby realizing virtual-real interaction and collaborative design.

[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.

[0108] Embodiments of this disclosure also provide a computer-readable storage medium storing a computer program configured to perform the steps in any of the above method embodiments when executed.

[0109] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0110] Embodiments of this disclosure also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.

[0111] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0112] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0113] It is obvious to those skilled in the art that the modules or steps of this disclosure described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this disclosure is not limited to any particular combination of hardware and software.

[0114] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A video rendering processing method applied to an extended reality (XR) platform, comprising: Receive perception requests initiated by XR terminals; The sensing request is sent to the sensing function SF of the core network, wherein the sensing request is used to request the SF to perform sensory sensing of the XR service of the XR terminal. Receive sensing data from the XR terminal from the SF; The XR video of the XR terminal is rendered based on the perceived data.

2. The method according to claim 1, wherein, After sending the sensing request to the sensing function (SF) of the core network, the method further includes: The sensing request is sent to the Radio Access Network (RAN) via the Access and Mobility Management Function (AMF) or the Network Control Unit (NCU); or The sensing request is sent directly to the Radio Access Network (RAN); The RAN is used to work with the XR terminal to sense the XR terminal.

3. The method according to claim 1, wherein, The sensing data received from the XR terminal from the SF includes: The sensing data is received from the SF via the Radio Access Network (RAN), wherein the sensing data is reported by the RAN after it has identified the XR terminal according to the sensing request and has cooperated with the XR terminal to perform sensing; or The XR terminal reports the sensing data received from the SF, wherein the sensing data is reported by the XR terminal after it performs sensory perception in collaboration with the RAN according to the RAN's sensing request.

4. The method according to claim 1, wherein, Rendering the XR video of the XR terminal based on the perceived data includes: Receive rotation information reported by the XR terminal; If the XR terminal has sensing service permissions, XR video rendering is performed based on the rotation information reported by the XR terminal and the sensing data.

5. The method according to claim 4, wherein, XR video rendering based on the rotation information reported by the XR terminal and the perceived data includes: The target location in the digital space of the XR video is determined based on the location information in the perceived data. The viewpoint to be rendered is obtained based on the position information and rotation information in the perceived data; Information about the virtual object to be rendered in the digital space is obtained based on the target location and the viewpoint to be rendered. The virtual object information is superimposed onto the XR video to obtain the rendered video.

6. The method according to claim 5, wherein, The rotation information includes at least one of the following: rotation information about the X-axis, rotation information about the Y-axis, and rotation information about the Z-axis.

7. The method according to any one of claims 1 to 6, wherein, The perception request carries at least one of the following: service type, service requirements, information of the perceived target, and information of the perception node.

8. A video rendering processing method applied to a Radio Access Network (RAN), wherein, include: The core network receives perception requests, wherein the perception requests are sent by the XR terminal to the extended reality XR platform and then sent to the core network by the XR platform. The XR services of the XR terminal are sensed by the sensor, and the sensed data of the XR terminal is sent to the XR platform through the core network to render the XR video of the XR terminal.

9. A video rendering processing system, comprising: XR platform, Radio Access Network (RAN), and core network The XR platform is used to receive perception requests initiated by XR terminals and send the perception requests to the perception function SF of the core network. The core network is used to receive the sensing request through the SF, send the sensing request to the radio access network RAN, receive the sensing data of the XR terminal through the radio access network RAN, and send the sensing data to the XR platform. The XR platform is also used to render the XR video of the XR terminal based on the perceived data.

10. The system according to claim 9, wherein, The RAN is used to receive the sensing request sent by the Access and Mobility Management Function (AMF) or Network Control Unit (NCU) of the core network; Alternatively, the sensing request can be sent directly to the SF via the core network; and the XR terminal can be sensed in coordination with the XR terminal according to the sensing request.

11. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 8.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method according to any one of claims 1 to 8.

13. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1 to 8.