XR space calculation multipath transmission method and system

By decomposing video frames in XR space computing and selecting the optimal transmission path using the multi-arm gambling algorithm, the edge server performs feature extraction and the cloud server performs pose matching, solving the problems of insufficient computing resources and high network latency in the prior art, and achieving efficient real-time and smoothness of XR applications.

CN120390103APending Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409999.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing XR space computing solutions have problems such as insufficient computing resources, high network latency, low task execution efficiency and inconsistent latency in mobile terminals, edge computing, cloud computing and distributed computing, which are difficult to meet real-time and user experience needs.

Method used

The video frame is decomposed into data packets, and the optimal transmission path is selected using the improved multi-arm gambling algorithm, and the data packet is sent to the edge server for feature extraction. The edge server is reorganized and the feature points and descriptors are transmitted to the cloud server for matching. The cloud server estimates the position information and transmits it to the mobile device through the wide area network for virtual information rendering.

Benefits of technology

It significantly reduces network bandwidth consumption and transmission delay, improves system performance and user experience, reduces the computing burden of mobile devices, and ensures real-time and smoothness of XR applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390103A_ABST
    Figure CN120390103A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and discloses an XR space calculation multipath transmission method and system, and the method comprises the steps: selecting an optimal transmission path through an improved multi-arm bandit algorithm, and transmitting a data package to a corresponding edge server through the optimal transmission path; the edge server receives and recombines the data packet into a video frame image, extracts feature points and descriptors of the video frame image by using a scale invariant feature transformation algorithm, and transmits the feature points and the descriptors to a cloud server; matching is carried out in combination with a pre-constructed three-dimensional point cloud map, pose information of the user is estimated, and the pose information is transmitted to a mobile device through a wide area network; and superposing the virtual information to a video frame by using a rendering engine, and displaying the video frame to a user for interaction. According to the invention, the collected video frames are unloaded to a plurality of edge server nodes, and only the feature extraction task is executed on the nodes, so that the network bandwidth consumption and the transmission delay are obviously reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an XR spatial computing multi-path transmission method and system. Background Art

[0002] Extended Reality (XR) spatial computing is a comprehensive computing platform that combines technologies such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). It processes and perceives the three-dimensional space through computers and sensors to create immersive and interactive three-dimensional environments, enabling users to interact with elements in the virtual and real worlds. XR spatial computing typically involves key processes such as three-dimensional modeling, spatial perception, and rendering interaction. Among them, three-dimensional modeling generates a three-dimensional space model through sensors and data fusion technologies. Common methods include generating a three-dimensional point cloud map using devices such as RGB cameras and LiDAR scans, providing a basis for the construction of the virtual world; spatial perception relies on technologies such as computer vision and deep learning to identify objects or track the position and pose of users in the three-dimensional space; rendering interaction involves seamlessly integrating virtual elements with the real world to enhance the user experience. For example, in a mobile AR navigation application, the user captures the real-world view through the camera of a mobile device such as a smartphone or tablet. The XR spatial computing system identifies environmental elements such as roads and buildings through spatial perception technologies, and at the same time generates real-time virtual navigation information (such as guiding lines, road conditions, etc.) and superimposes it on the user's field of view to help the user perform accurate navigation and path planning. The XR spatial computing tasks covered by the entire process include: pre-construction of the point cloud map, video frame acquisition, feature extraction, feature matching and pose estimation, and virtual information rendering.

[0003] However, existing XR spatial computing solutions can be divided into the following four types according to the computing task allocation mode, but each solution has significant defects and faces certain technical challenges and limitations:

[0004] 1. XR spatial computing based on mobile terminals: In this solution, all tasks (including pre-construction of the point cloud map, video frame acquisition, feature extraction, feature matching and pose estimation, and virtual information rendering, etc.) are locally executed on the user's mobile device. However, the computing resources of mobile terminals are limited and the processing capabilities are weak, making it difficult to handle complex XR spatial computing tasks with high real-time requirements. The GPU or CPU of modern smartphones can usually only process computing tasks of 1 - 5 GFLOPS, while feature extraction and feature matching tasks usually require processing capabilities of dozens to hundreds of GFLOPS. Executing all tasks on mobile terminals may increase the computing latency, fail to meet the real-time requirements of XR applications, and at the same time significantly increase battery consumption and cause the device to overheat.

[0005] 2. Edge Computing-based XR Spatial Computing: This solution offloads some computing tasks (such as feature extraction, pose estimation, etc.) to edge devices close to users, thus reducing the computing burden on local devices. However, edge computing still needs to communicate with user devices through the network, and network latency and bandwidth limitations can have a significant impact on task execution efficiency. Especially in the case of multi-user concurrency, the computing resources of the edge server may be insufficient to meet the requirements of complex XR computing tasks.

[0006] 3. Cloud Computing-based XR Spatial Computing: Cloud servers have powerful computing capabilities and are suitable for processing large-scale XR computing tasks. However, since cloud servers are usually far from users and data transmission needs to go through wide area networks, they will face high network latency problems. Although cloud computing provides high-performance computing, network transmission latency may lead to a poor user experience. Especially when the latency exceeds 20 milliseconds, it may cause discomfort to users.

[0007] 4. Distributed XR Spatial Computing: This solution combines edge computing and cloud computing, distributes computing tasks to multiple edge devices and cloud servers for execution, and finally aggregates the computing results to user devices. However, the complexity of task partitioning, data synchronization, and result aggregation may lead to latency and data inconsistency problems. In XR spatial computing tasks, the dependencies between multiple computing steps are relatively strong, and the order and execution results of tasks will affect the accuracy of subsequent steps.

[0008] Therefore, how to provide an XR spatial computing multi-path transmission method and system is an urgent problem to be solved at present. Summary of the Invention

[0009] Embodiments of the present invention provide an XR spatial computing multi-path transmission method and system to solve the above technical problems existing in the prior art.

[0010] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0011] According to the first aspect of the embodiments of the present invention, an XR spatial computing multi-path transmission method is provided.

[0012] In one embodiment, an XR spatial computing multi-path transmission method includes:

[0013] Decompose video frames into several data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server through the optimal transmission path;

[0014] The edge server receives and reorganizes the data packets into video frame images, extracts the feature points and descriptors of the video frame images by using the Scale-Invariant Feature Transform (SIFT) algorithm, and transmits the feature points and descriptors to the cloud server;

[0015] The cloud server receives the feature points and descriptors, performs matching in combination with the pre-constructed three-dimensional point cloud map, estimates the pose information of the user, and transmits the pose information to the mobile device through the wide area network;

[0016] The mobile device, according to the received pose information, uses the rendering engine to superimpose the virtual information on the video frame and display it to the user for interaction.

[0017] In one embodiment, the step of decomposing the video frame into several data packets, selecting the optimal transmission path by using the improved multi-armed bandit algorithm, and sending the data packets to the corresponding edge server through the optimal transmission path includes:

[0018] Capture scene images from different perspectives by using a camera, generate a three-dimensional point cloud map by using three-dimensional reconstruction technology, and store the three-dimensional point cloud map on the cloud server;

[0019] Start the extended reality space computing application through the mobile device and authorize the application to access the multi-modal network terminal;

[0020] Capture the video frame of the user's current environment by using the mobile device and decompose the video frame into several data packets, where the data packets include image data and identification information;

[0021] Use the improved multi-armed bandit algorithm to select the optimal transmission path for each data packet, and in combination with the preset transmission protocol, send the data packets to the corresponding edge server through the optimal transmission path.

[0022] In one embodiment, the step of using the improved multi-armed bandit algorithm to select the optimal transmission path for each data packet, and in combination with the preset transmission protocol, send the data packets to the corresponding edge server through the optimal transmission path includes:

[0023] According to the decomposed data packets, collect the network status and edge device performance characteristics of each transmission path to form a state vector, and construct an action set of the multi-armed bandit, where each action in the action set corresponds to the selection of a transmission path;

[0024] Use the improved multi-armed bandit algorithm to process the action set, calculate and output the optimal action of each data packet within the current scheduling interval to determine the optimal transmission path;

[0025] Based on the determined optimal transmission path, in combination with the preset transmission protocol, send the data packets to the corresponding edge server.

[0026] In one embodiment, processing the action set by using the improved multi-armed bandit algorithm, calculating and outputting the optimal action of each data packet within the current scheduling interval to determine the optimal transmission path includes:

[0027] Based on each action in the action set, initialize the covariance matrix and vector, and use the multi-armed bandit algorithm to calculate the corresponding parameter vector and expected reward;

[0028] According to the identification information of each data packet, determine whether the currently scheduled data packet is the first data packet of the video frame;

[0029] If the identification information is zero, indicating that the current video frame is the first data packet, then calculate the action with the maximum expected reward, generate a random number within a preset interval, and compare it with a preset parameter. Based on the comparison result, construct a constrained action space to determine the optimal action for the current scheduling interval;

[0030] If the identification information is not zero, indicating that the current video frame is not the first data packet, then calculate the action with the maximum expected reward as the optimal action for the current scheduling interval;

[0031] According to the proportional formula, calculate the time taken by the data packet, and update the covariance matrix and vector corresponding to the selected action to ensure that the optimal action of each data packet is continuously output in subsequent scheduling intervals to determine the optimal transmission path.

[0032] In one embodiment, the calculation formula for using the multi-armed bandit algorithm to calculate the corresponding parameter vector and expected reward is:

[0033]

[0034] In the formula, θ represents the parameter vector; a represents the action in the action set; A represents the covariance matrix; represents the inverse matrix of the covariance matrix; b represents the vector; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; represents the transpose matrix of the state vector; α represents the hyperparameter; t represents the time.

[0035] In one embodiment, the comparison results include:

[0036] If the random number is less than the preset parameter, then the action is the optimal action for the current scheduling interval;

[0037] If the random number is greater than the preset parameter, then randomly select an action from the action set as the optimal action for the current scheduling interval.

[0038] In one embodiment, the formula for calculating the action with the maximum expected reward is as follows:

[0039]

[0040] In the formula, a t represents the action with the maximum expected reward; a represents an action in the action set; Y represents the action space; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; t represents the time.

[0041] In one embodiment, the ratio formula is as follows:

[0042]

[0043] In the formula, CT(d t |a t ) represents the computing time for performing feature extraction tasks on the video frame F t on the edge device to which the corresponding path of the action a t leads; CT represents the computing time; d represents the data packet; t represents the time; a represents an action in the action set; F represents the video frame; B represents the size of the data in bytes; B(d t ) represents the data size of the data packet; B(F t ) represents the data size of the video frame.

[0044] In one embodiment, the edge server receives and reorganizes the data packets into video frame images, extracts the feature points and descriptors of the video frame images by using the scale-invariant feature transform algorithm, and transmits the feature points and descriptors to the cloud server, including:

[0045] The edge server receives the decomposed data packets through a socket program and stores the data packets in a buffer;

[0046] Parse the identification information of each data packet in the buffer and sort the data packets according to the identification information;

[0047] Based on the sorting result and all the data packets of the same video frame, extract and splice the data packet payloads and reorganize them into video frame images;

[0048] Use the scale-invariant feature transform algorithm to detect the feature points in the video frame image and calculate the descriptors of each feature point to capture the local image information of the feature point area;

[0049] Transmit the detected feature points and the calculated descriptors to the edge server.

[0050] According to the second aspect of the embodiments of the present invention, an XR space computing multipath transmission system is provided.

[0051] In one embodiment, the XR spatial computing multipath transmission system includes:

[0052] A data packet processing and transmission module, which is used to decompose a video frame into several data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server through the optimal transmission path;

[0053] An edge server recombination and feature extraction module, which is used for the edge server to receive and recombine the data packets into a video frame image, extract feature points and descriptors of the video frame image using the scale-invariant feature transform algorithm, and transmit the feature points and descriptors to the cloud server;

[0054] A cloud server matching and pose estimation module, which is used for the cloud server to receive the feature points and descriptors, perform matching in combination with a pre-constructed three-dimensional point cloud map, estimate the pose information of the user, and transmit the pose information to the mobile device through the wide area network;

[0055] A mobile device rendering and interaction module, which is used for the mobile device to overlay virtual information on the video frame according to the received pose information and display it to the user for interaction.

[0056] According to the third aspect of the embodiments of the present invention, a computer device is provided.

[0057] In some embodiments, the computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0058] According to the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided.

[0059] In one embodiment, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0060] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0061] 1. By offloading the collected video frames to multiple edge server nodes, only performing feature extraction tasks on these nodes, and transmitting the extracted feature points or descriptors to the remote cloud server nodes, the present invention not only greatly reduces the computing and storage burdens of the edge servers, but also avoids transmitting complete video frames over the wide area network, thus significantly reducing network bandwidth consumption and transmission latency.

[0062] 2. By storing the pre-constructed offline point cloud map on the cloud server, the present invention effectively reduces the storage pressure on the edge server; meanwhile, by parallelly processing the feature extraction tasks on the edge server, the computing and storage resources are fully utilized, thereby improving the task execution efficiency.

[0063] 3. By transmitting feature points instead of complete video frames in the network, the present invention significantly reduces the bandwidth consumption and transmission delay, and improves the system performance and user experience; in addition, by reasonably allocating the computing tasks, the computing burden and battery consumption of the mobile device are effectively reduced, ensuring the real-time performance and smoothness of the XR application.

[0064] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0066] Figure 1 is a flowchart of an XR spatial computing multi-path transmission method shown according to an exemplary embodiment;

[0067] Figure 2 is a schematic block diagram of an XR spatial computing multi-path transmission system shown according to an exemplary embodiment;

[0068] Figure 3 is a schematic structural diagram of a computer device shown according to an exemplary embodiment;

[0069] Figure 4 is a schematic diagram of a spatial computing system in an XR spatial computing multi-path transmission method shown according to an exemplary embodiment;

[0070] Figure 5 is a flowchart of an improved multi-armed bandit algorithm in an XR spatial computing multi-path transmission method shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The following description and the accompanying drawings fully disclose specific embodiments herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. In this document, the terms "first", "second", etc. are only used to distinguish one element from another, without requiring or implying any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a structure, device or equipment comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the structure, device or equipment comprising the said element. The various embodiments herein are described in a progressive manner, with each embodiment highlighting the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0072] In this document, the orientation or positional relationship indicated by terms such as "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing this document and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention. In the description herein, unless otherwise specified and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0073] In this document, unless otherwise stated, the term "plurality" means two or more.

[0074] In this document, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.

[0075] In this document, the term "and / or" is a description of the associated relationship of an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.

[0076] It should be understood that although the various steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or sub-steps or stages of other steps.

[0077] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0078] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0079] Figure 1 An embodiment of a multi-path transmission method for XR spatial computing according to the present invention is shown.

[0080] In this alternative embodiment, the multi-path transmission method for XR spatial computing includes:

[0081] Step S101: Decompose a video frame into several data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server through the optimal transmission path;

[0082] Step S102: The edge server receives and recombines the data packets into a video frame image, extracts feature points and descriptors of the video frame image using the scale-invariant feature transform algorithm, and transmits the feature points and descriptors to the cloud server;

[0083] Step S103: The cloud server receives the feature points and descriptors, performs matching in combination with a pre-constructed three-dimensional point cloud map, estimates the pose information of the user, and transmits the pose information to the mobile device through the wide area network;

[0084] Step S104: The mobile device superimposes virtual information on the video frame according to the received pose information and displays it to the user for interaction.

[0085] In this alternative embodiment, the steps of decomposing a video frame into a number of data packets, selecting an optimal transmission path using an improved multi-armed bandit algorithm, and sending the data packets to the corresponding edge server through the optimal transmission path include:

[0086] Use a camera to capture scene images from different perspectives, generate a three-dimensional point cloud map using three-dimensional reconstruction technology, and store the three-dimensional point cloud map on a cloud server;

[0087] Start an extended reality space computing application on a mobile device and authorize the application to access the multimodal network terminal;

[0088] Use the mobile device to capture a video frame of the user's current environment and decompose the video frame into a number of data packets, where the data packets include image data and identification information;

[0089] Use an improved multi-armed bandit algorithm to select an optimal transmission path for each data packet, and combine it with a preset transmission protocol to send the data packets to the corresponding edge server through the optimal transmission path.

[0090] In this alternative embodiment, the steps of using an improved multi-armed bandit algorithm to select an optimal transmission path for each data packet, and combining it with a preset transmission protocol to send the data packets to the corresponding edge server through the optimal transmission path include:

[0091] According to the decomposed data packets, collect the network status and edge device performance characteristics of each transmission path, form a state vector, and construct an action set of the multi-armed bandit, where each action in the action set corresponds to the selection of a transmission path;

[0092] Use an improved multi-armed bandit algorithm to process the action set, calculate and output the optimal action for each data packet within the current scheduling interval to determine the optimal transmission path;

[0093] Based on the determined optimal transmission path, combine it with a preset transmission protocol to send the data packets to the corresponding edge server.

[0094] In this alternative embodiment, the steps of using an improved multi-armed bandit algorithm to process the action set, calculate and output the optimal action for each data packet within the current scheduling interval to determine the optimal transmission path include:

[0095] Based on each action in the action set, initialize the covariance matrix and vector, and use the multi-armed bandit algorithm to calculate the corresponding parameter vector and expected reward;

[0096] According to the identification information of each data packet, determine whether the currently scheduled data packet is the first data packet of the video frame;

[0097] If the identification information is zero, indicating that the current video frame is the first data packet, then calculate the action with the maximum expected reward, generate a random number within a preset range, compare it with a preset parameter, and construct a constrained action space based on the comparison result to determine the optimal action for the current scheduling interval;

[0098] If the identification information is not zero, indicating that the current video frame is not the first data packet, then calculate the action with the maximum expected reward as the optimal action for the current scheduling interval;

[0099] According to the proportional formula, calculate the time taken for the data packet, and update the covariance matrix and vector corresponding to the selected action to ensure that the optimal action for each data packet is continuously output in subsequent scheduling intervals and determine the optimal transmission path.

[0100] In this alternative embodiment, the calculation formula for the corresponding parameter vector and expected reward using the multi-armed bandit algorithm is:

[0101]

[0102] In the formula, θ represents the parameter vector; a represents the action in the action set; A represents the covariance matrix; represents the inverse matrix of the covariance matrix; b represents the vector; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; represents the transpose matrix of the state vector; α represents the hyperparameter; t represents the time.

[0103] In this alternative embodiment, the comparison results include:

[0104] If the random number is less than the preset parameter, then the action is the optimal action for the current scheduling interval;

[0105] If the random number is greater than the preset parameter, then randomly select an action from the action set as the optimal action for the current scheduling interval.

[0106] In this alternative embodiment, the formula for calculating the action with the maximum expected reward is:

[0107]

[0108] In the formula, a t represents the action with the maximum expected reward; a represents the action in the action set; Y represents the action space; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; t represents the time.

[0109] In this alternative embodiment, the proportional formula is:

[0110]

[0111] In the formula, CT(d t |a t ) represents the computing time for performing feature extraction tasks on the video frame F on the edge device to which the corresponding path leads during the action a t ; CT represents the computing time; d represents the data packet; t represents the moment; a represents the action in the action set; F represents the video frame; B represents the byte size occupied by the data; B(d t ) represents the data size of the data packet; B(F t ) represents the data size of the video frame. t )

[0112] In this alternative embodiment, the edge server receives and reorganizes data packets into video frame images, extracts feature points and descriptors of the video frame images using the scale-invariant feature transform algorithm, and transmits the feature points and descriptors to the cloud server, including:

[0113] The edge server receives the decomposed data packets through a socket program and stores the data packets in a buffer;

[0114] Parse the identification information of each data packet in the buffer and sort the data packets according to the identification information;

[0115] Based on the sorting result and all the data packets of the same video frame, extract and splice the data packet payloads and reorganize them into video frame images;

[0116] Use the scale-invariant feature transform algorithm to detect feature points in the video frame image and calculate the descriptors of each feature point to capture the local image information of the feature point region;

[0117] Transmit the detected feature points and the calculated descriptors to the edge server.

[0118] It should be noted that, as Figure 4 shown, the specific embodiments of an XR spatial computing multipath transmission method of the present invention are as follows:

[0119] Step 1: The mapping personnel use a camera to capture scene images from different perspectives, generate a three-dimensional point cloud map using three-dimensional reconstruction technologies such as the existing structure from motion (SFM) algorithm, and store it on a remote cloud server; this step is pre-completed in a non-real-time offline environment; if the three-dimensional point cloud map already exists, this step can be omitted.

[0120] Step 2: The user starts the XR spatial computing application through a mobile device (such as a smartphone or wearing AR glasses) and authorizes the application to access the WiFi network, 5G cellular network, and camera.

[0121] Step 3: When the application starts, the camera of the mobile device captures the user's current view (video frame); the high-resolution video frame is decomposed into several data packets because it exceeds the maximum transmission unit (MTU) of the network data packet. The improved multi-armed bandit algorithm provided by the present invention is used to determine the best offloading path for each data packet, and each data packet is transmitted from the determined best offloading path to the corresponding edge server device based on the network socket program through the QUIC protocol. Specifically as follows:

[0122] 1. Video frame unpacking

[0123] The captured video frame is unpacked and decomposed according to the constraint that the size of a single data packet does not exceed the maximum transmission unit (MTU) of the network, where each data packet contains partial image data and identification information.

[0124] 2. Edge computing resource recording and transmission

[0125] The resource manager on each edge device records the real-time CPU and GPU resource utilization rates at a certain frequency and transmits them to the user's mobile device through the QUIC protocol. The recording frequency is equal to the reciprocal of the scheduling time interval of the data packet. The QUIC protocol is a network transmission protocol proposed by Google and standardized by the IETF. Compared with the traditional TCP protocol, QUIC is based on UDP (User Datagram Protocol) and has characteristics such as low-latency connection, multiplexing, built-in encryption, and packet loss resistance. In this process, the edge device first encodes the CPU and GPU resource utilization rate data into a format suitable for transmission (such as JSON) and sends it to the user's mobile device through the stream mechanism of the QUIC protocol. The specific implementation steps include establishing a QUIC protocol encryption connection, performing a TLS handshake, transmitting resource data through the QUIC stream, the mobile device receiving and parsing the resource data sent by the edge device, and handling and retransmitting potential errors. The multiplexing and congestion control mechanisms of the QUIC protocol ensure the efficient transmission of data and network adaptability.

[0126] 3. Network path condition monitoring

[0127] The network condition monitoring module on the user's mobile device also records the real-time network conditions of each offloading path at the above frequency, including the congestion window size (CWND), the number of data packets being transmitted (InP), the send window size (SWND), and the round-trip delay (RTT).

[0128] 4. Path optimization selection

[0129] The improved multi-armed bandit algorithm provided by the present invention is used to select the best offloading path for each data packet. Taking three edge devices and two heterogeneous network types (WiFi and 5G) as examples, a state vector x of the scheduling interval t is constructed t=(c t,1 , f t,1 , w t,1 , τ t,1 , cg t,1 , cu t,1 , gg t,1 , gu t,1 , …, c t,6 , f t,6 , w t,6 , τ t,6 , cg t,6 , cu t,6 , gg t,6 , gu t,6 ), where c t,i is the congestion window size of the i-th path, f t,i is the number of data packets being transmitted on the i-th path, w t,i is the sending window size of the i-th path, τ t,i is the round-trip delay of the i-th path, cg t,i is the inherent CPU computing performance GOPS (billions of operations per second) of the edge device that the i-th path goes to, cu t,i is the real-time CPU utilization rate of the edge device that the i-th path goes to, gg t,i is the inherent GPU computing performance GOPS (billions of operations per second) of the edge device that the i-th path goes to, gu t,i is the real-time GPU utilization rate of the edge device that the i-th path goes to. It can be seen that x t is a 48*1 column vector. Construct the action set Y of the multi-armed bandit. The number of elements in the set Y is equal to the number of offloading paths, which is 6 in this embodiment. Each element ɑ i in the set Y is a 6*1 action vector. When the i-th path is selected, the i-th value in the action vector ɑ i is set to 1, and the rest of the values are set to 0. Input the state matrix x t , the action set Y, together with the hyperparameters α and the random factor ε into the improved multi-armed bandit algorithm to obtain the optimal action ɑ t of the scheduling interval t. In this embodiment, α = 0.8 and ε = 0.2. As Figure 5 shown, specifically as follows:

[0130] (1) For each element ɑ in the action set Y, initialize the covariance matrix A ɑ as the l×l identity matrix I, and initialize the vector b ɑ as the l×1 zero vector, which is used to update the selection parameters of the path later. In this embodiment, l is equal to 48.

[0131] (2) For each element ɑ in the action set Y, calculate the parameter vector θɑ , the calculation formula is as follows:

[0132]

[0133] In the formula, θ represents the parameter vector; a represents the action in the action set; A represents the covariance matrix; A a -1 represents the inverse matrix of the covariance matrix; b represents the vector.

[0134] (3) For each element ɑ in the action set Y, calculate the expected reward E[R t,ɑ |x t , and the calculation formula is as follows:

[0135]

[0136] In the formula, E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; represents the transpose matrix of the state vector; α represents the hyperparameter; t represents the time; θ represents the parameter vector; a represents the action in the action set; represents the inverse matrix of the covariance matrix.

[0137] (4) According to the identification field (intra-frame sequence number) set in the header information of each video frame data packet previously, determine whether the data packet to be scheduled currently is the first data packet of the video frame. If the intra-frame sequence number is 0, it indicates that it is the first data packet of the current video frame, then go to (5), otherwise go to (6).

[0138] (5) Calculate the action a t with the maximum expected reward, and the formula is as follows:

[0139]

[0140] In the formula, a t represents the action with the maximum expected reward; a represents the action in the action set; Y represents the action space; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; t represents the time.

[0141] Next, generate a random number between 0 and 1. If the random number is less than the parameter ε, then a t is the optimal action for the scheduling interval t;

[0142] If the random number is greater than the parameter ε, randomly select an action from the action space Y as the optimal action a t for the scheduling interval t; meanwhile, construct the constrained action space Y', and the elements in Y' are the action a tThe corresponding network path leads to the network path between the edge device and the user's mobile device. In this embodiment, the number of elements in Y' is 2 (corresponding to two networks, WiFi and 5G).

[0143] (6) Calculate the action a with the maximum expected reward t As the optimal action for the scheduling interval t, the formula is expressed as follows:

[0144]

[0145] In the formula, a t represents the action with the maximum expected reward; a represents the action in the action set; Y represents the action space; E represents the expected reward; R t,a Represents the reward obtained after selecting an action; x represents the state vector; t represents the time.

[0146] (7) Update action a t The corresponding covariance matrix A ɑt and vector b ɑt , the update formulas are as follows:

[0147]

[0148] In the formula, A at represents the covariance matrix; a t represents the action with the maximum expected reward; x represents the state vector; t represents the time; represents the transposed matrix of the state vector; b at Represents a vector; R t,a Indicates selection of action a t The reward obtained after is the current scheduled data packet d t From action a t The time TT(d) from the sending of the corresponding path to the receiving of the ACK mark t |ɑ t ) and data packet d t In action t The computation time CT(d t |ɑ t ). Since the minimum input unit for feature extraction is the video frame, for the data packet d t The computation time CT(d t |ɑ t ) To obtain it in proportion, the calculation formula is as follows:

[0149]

[0150] Where, CT(d t |a t)Indicates on the edge device to which the corresponding path leads, for the video frame F t The computing time for performing the feature extraction task; CT represents the computing time; d represents the data packet; t represents the moment; a represents the action in the action set; F represents the video frame; B represents the size in bytes of the data; B(d t )Represents the data size of the data packet; B(F t )Represents the data size of the video frame. t )Represents the data size of the video frame.

[0151] 5. Data packet transmission

[0152] The network communication module on the mobile device, based on the network socket program, sends the data packet corresponding to the scheduling interval t to ɑ t The corresponding network path, and transmits it to the corresponding edge server device through the QUIC protocol.

[0153] Step 4: After the edge device receives all the data packets that make up the same video frame, it recombines them into a complete video frame, and uses the Scale-Invariant Feature Transform (SIFT) algorithm to extract the key feature points and descriptors of the image.

[0154] Specifically as follows:

[0155] 1. Receive video data packets

[0156] The edge device first uses the socket program to receive the video frame data split into multiple data packets. Each time a new data packet is received, the received data packet is stored in the buffer.

[0157] 2. Parse the location identifier

[0158] Parse out that each data packet contains its location identifier in the video frame (including the inter-frame sequence number and the intra-frame sequence number). The edge device sorts the data packets in the buffer according to these identifiers.

[0159] 3. Recombine data packets

[0160] After all the data packets to form the same video frame arrive in the buffer, extract the payload parts of all the data packets and directly splice them together to combine the individual data packets into a complete image for subsequent processing.

[0161] 4. Feature point detection

[0162] Use the existing SIFT algorithm to detect the key feature points in the image. The SIFT algorithm determines the feature points by finding the extreme points (including corner points, edges, etc.) in the local area of the image, that is, the areas with significant visual changes. The advantage of the SIFT algorithm is that it is invariant to scale, rotation, and illumination changes.

[0163] 5. Feature descriptor calculation

[0164] For each detected key feature point, the SIFT algorithm calculates a descriptor. A descriptor is a vector that captures the local image information in the area around the feature point and is described by the local gradient direction and intensity. These descriptors can distinguish different feature points, and the feature points remain invariant even under different scales and rotation angles.

[0165] Step Five: The edge device transmits the extracted feature points and descriptors to the remote cloud server via a high-speed wired link; the cloud server matches these features with a pre-constructed three-dimensional point cloud map. Specifically, through the SIFT descriptor matching algorithm (such as the brute-force matching algorithm, for each two-dimensional feature point, calculate the distance between its descriptor and all descriptors in the three-dimensional point cloud, and select the point corresponding to the minimum distance as the matching point) to find the 3D coordinates corresponding to the extracted feature points in the three-dimensional point cloud map; use the matching results to estimate the precise position and pose of the user, and transmit the pose result back to the user's mobile device via a wide area network. The specific implementation steps of pose estimation include that the cloud server uses the existing PnP (Perspective-n-Point) algorithm to dock the two-dimensional image coordinates with the three-dimensional world coordinates based on the known matching point pairs, and solve the rotation matrix and displacement vector of the camera, so as to obtain the user's pose (rotation matrix) and position (translation vector).

[0166] Step Six: The rendering engine on the mobile device renders and superimposes the virtual information onto the original video frame according to the pose result returned by the cloud server, presenting virtual content related to specific downstream tasks in the user's field of view to achieve the interaction function. The specific implementation steps are as follows: First, use the user's pose information to convert the virtual guidance information (such as arrows, markers, paths, etc.) from the world coordinate system to the camera coordinate system, and project the virtual guidance information onto the 2D image plane through the camera internal parameters. This step uses existing graphics rendering libraries (such as OpenGL, Vulkan, or use the mobile device's graphics API, such as Metal or OpenGL ES) to fuse the rendered guidance elements with the pixel data in the original video frame. According to actual needs, transparent overlay or mask overlay can be selected; among them, transparent overlay uses the alpha blending technique to blend the virtual guidance elements with the video frame, so that the guidance information is displayed on the upper layer of the video and ensures that the background video content is visible; mask overlay is to set the background color or background transparency for the virtual guidance in some scenarios to avoid the graphic elements affecting other contents of the video frame.

[0167] Step Seven: Repeat Steps Three to Six. The mobile device continues to select the offloading path for each data packet in the subsequent video frame through the improved multi-armed bandit algorithm and obtains the pose result returned by the cloud server to achieve the rendering interaction at the next moment.

[0168] Figure 2 An embodiment of an XR spatial computing multi-path transmission system of the present invention is shown.

[0169] In this alternative embodiment, the XR spatial computing multi-path transmission system includes:

[0170] A data packet processing and transmission module 201, configured to decompose a video frame into a plurality of data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server through the optimal transmission path;

[0171] An edge server reorganization and feature extraction module 202, configured to receive and reorganize the data packets into a video frame image by the edge server, extract feature points and descriptors of the video frame image using a scale-invariant feature transform algorithm, and transmit the feature points and descriptors to the cloud server;

[0172] A cloud server matching and pose estimation module 203, configured to receive the feature points and descriptors by the cloud server, perform matching in combination with a pre-constructed three-dimensional point cloud map, estimate the pose information of the user, and transmit the pose information to the mobile device through a wide area network;

[0173] A mobile device rendering and interaction module 204, configured to overlay virtual information on the video frame according to the received pose information by the mobile device and display it to the user for interaction.

[0174] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 3 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the steps in the above method embodiment.

[0175] Those skilled in the art can understand that Figure 3 the structure shown in

[0176] In addition, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0177] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0178] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0179] The present invention is not limited to the structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. An XR space calculation multipath transmission method, characterized in that, The method includes: Decompose a video frame into a number of data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server via the optimal transmission path; The edge server receives and recombines the data packets into a video frame image, extracts feature points and descriptors of the video frame image using the Scale-Invariant Feature Transform (SIFT) algorithm, and transmits the feature points and descriptors to the cloud server; The cloud server receives the feature points and descriptors, performs matching in combination with a pre-constructed 3D point cloud map, estimates the pose information of the user, and transmits the pose information to the mobile device via a wide area network; The mobile device, based on the received pose information, uses a rendering engine to overlay virtual information on the video frame and display it to the user for interaction.

2. The multi-path transmission method for XR spatial computing according to claim 1, wherein The step of decomposing a video frame into a number of data packets, selecting an optimal transmission path using an improved multi-armed bandit algorithm, and sending the data packets to the corresponding edge server via the optimal transmission path includes: Capture scene images from different perspectives using a camera, generate a 3D point cloud map using 3D reconstruction technology, and store the 3D point cloud map on the cloud server; Start an extended reality space computing application on the mobile device and authorize the application to access the multi-modal network terminal; Capture a video frame of the user's current environment using the mobile device and decompose the video frame into a number of data packets, where the data packets include image data and identification information; Use an improved multi-armed bandit algorithm to select an optimal transmission path for each data packet, and in combination with a preset transmission protocol, send the data packets to the corresponding edge server via the optimal transmission path.

3. The XR spatial computing multipath transmission method according to claim 2, characterized in that, The step of using an improved multi-armed bandit algorithm to select an optimal transmission path for each data packet, and in combination with a preset transmission protocol, send the data packets to the corresponding edge server via the optimal transmission path includes: According to the decomposed data packets, collect the network status and edge device performance characteristics of each transmission path to form a state vector, and construct an action set for the multi-armed bandit, where each action in the action set corresponds to the selection of a transmission path; Use an improved multi-armed bandit algorithm to process the action set, calculate and output the optimal action for each data packet within the current scheduling interval to determine the optimal transmission path; Based on the determined optimal transmission path, in combination with a preset transmission protocol, send the data packets to the corresponding edge server.

4. The XR spatial computing multipath transmission method according to claim 3, wherein The step of using an improved multi-armed bandit algorithm to process the action set, calculate and output the optimal action for each data packet within the current scheduling interval to determine the optimal transmission path includes: Based on each action in the action set, initialize the covariance matrix and vector, and use the multi-armed bandit algorithm to calculate the corresponding parameter vector and expected reward; According to the identification information of each data packet, determine whether the currently scheduled data packet is the first data packet of the video frame; If the identification information is zero, indicating that the current video frame is the first data packet, then calculate the action with the maximum expected reward, generate a random number within a preset interval, and compare it with a preset parameter. Based on the comparison result, construct a constrained action space to determine the optimal action for the current scheduling interval; If the identification information is not zero, indicating that the current video frame is not the first data packet, then calculate the action with the maximum expected reward as the optimal action for the current scheduling interval; According to the proportional formula, calculate the time taken for the data packet, and update the covariance matrix and vector corresponding to the selected action to ensure that the optimal action for each data packet is continuously output in subsequent scheduling intervals, and determine the optimal transmission path.

5. The XR spatial computing multipath transmission method according to claim 4, wherein The calculation formula for the corresponding parameter vector and expected reward using the multi-armed bandit algorithm is: Where, θ represents the parameter vector; a represents the action in the action set; A represents the covariance matrix; represents the inverse matrix of the covariance matrix; b represents the vector; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; represents the transpose matrix of the state vector; α represents the hyperparameter; t represents the time instant.

6. The XR spatial computing multipath transmission method according to claim 5, wherein, The comparison results include: If the random number is less than the preset parameter, the action is the optimal action for the current scheduling interval; If the random number is greater than the preset parameter, randomly select an action from the action set as the optimal action for the current scheduling interval.

7. A multi-path transmission method for XR spatial computing according to claim 6, characterized in that The formula for calculating the action with the maximum expected reward is: where a t represents the action with the maximum expected reward; a represents an action in the action set; Y represents the action space; E represents the expected reward; R t,a represents the reward obtained after selecting the action; x represents the state vector; t represents the time instant.

8. A multi-path transmission method for XR spatial computing according to claim 7, characterized in that The proportional formula is: where CT(d t |a t ) represents the computing time for performing the feature extraction task on the edge device to which the corresponding path of action a t leads for video frame F t ; CT represents the computing time; d represents the data packet; t represents the moment; a represents the action in the action set; F represents the video frame; B represents the size of the data in bytes; B(d t ) represents the data size of the data packet; B(F t ) represents the data size of the video frame.

9. The XR space calculation multipath transmission method according to claim 1, characterized in that The edge server receives and reorganizes the data packets into video frame images, extracts the feature points and descriptors of the video frame images using the scale-invariant feature transform algorithm, and transmits the feature points and descriptors to the cloud server, including: The edge server receives the decomposed data packets through a socket program and stores the data packets in a buffer; Parse the identification information of each data packet in the buffer and sort the data packets according to the identification information; Based on the sorting result and all the data packets of the same video frame, extract and splice the data packet payloads and reorganize them into video frame images; Use the scale-invariant feature transform algorithm to detect the feature points in the video frame images and calculate the descriptors of each feature point to capture the local image information of the feature point region; Transmit the detected feature points and the calculated descriptors to the edge server.

10. An XR spatial computing multipath transmission system, characterized in that, The system includes: A data packet processing and transmission module, which is used to decompose a video frame into several data packets, select an optimal transmission path using an improved multi-armed bandit algorithm, and send the data packets to the corresponding edge server through the optimal transmission path; An edge server reorganization and feature extraction module, which is used for the edge server to receive and reorganize the data packets into video frame images, extract the feature points and descriptors of the video frame images using the scale-invariant feature transform algorithm, and transmit the feature points and descriptors to the cloud server; A cloud server matching and pose estimation module, which is used for the cloud server to receive the feature points and descriptors, perform matching in combination with a pre-constructed three-dimensional point cloud map, estimate the pose information of the user, and transmit the pose information to the mobile device through a wide area network; A mobile device rendering and interaction module, which is used for the mobile device to overlay virtual information on the video frame using a rendering engine according to the received pose information and display it to the user for interaction.