Application program management system and method driven by virtual engine
The application management system driven by the virtual engine adopts a four-dimensional isolation environment and intelligent resource allocation technology to solve the resource scheduling, rendering efficiency, audio and video processing and interactive response problems in the virtual engine, and improve the intelligence level of application management and user experience.
Patent Information
- Application Number
- CN202510807975.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing virtual engines have shortcomings in resource scheduling, rendering efficiency, audio and video processing, and interactive response, resulting in uneven CPU/memory utilization, imbalance between rendering quality and performance, insufficient multimodal signal processing accuracy, and delayed interactive response, affecting user experience.
It adopts a spatiotemporal virtualization engine, a semantic rendering decider, a multimodal signal processor, a GPU rendering unit, an intelligent interaction predictor and a metadata hub. Through technologies such as four-dimensional isolation environment, semantic rendering, multimodal signal synchronization, and user behavior prediction, it achieves dynamic optimization and intelligent allocation of resources, thereby improving rendering quality and interactive response.
It achieves refined resource scheduling, improves the balance between rendering quality and performance, enhances the immersiveness of audio and video processing and the timeliness of interactive responses, and significantly improves the intelligence level of application management and user experience.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual engines, and in particular to an application management system and method driven by a virtual engine. Background Art
[0002] With the development of mobile internet and cloud computing technologies, virtualization engines are increasingly being used in gaming, industrial simulation, the metaverse, and other fields. However, existing virtualization engine-driven application management systems generally suffer from the following technical bottlenecks:
[0003] Extensive resource scheduling: Traditional two-dimensional sandbox environments (such as Android's native VirtualApp) only support spatial isolation and cannot achieve refined resource allocation in the time dimension. This leads to uneven CPU / memory utilization and prone to lag in high-load scenarios (for example, frame rate fluctuations exceeding 20% when switching between multiple tasks).
[0004] Imbalance between rendering quality and performance: Low-end devices struggle to support high-resolution rendering, while high-end devices waste resources in simple scenes. Traditional LOD technology switches levels of detail based solely on geometric distance, without incorporating semantic information (such as object importance). This results in no difference in rendering priority between key characters and background objects, impacting the user experience.
[0005] Inadequate multimodal signal processing accuracy: Audio and video synchronization relies on a fixed Presentation Time Stamp (PTS) mechanism, resulting in synchronization errors typically exceeding 25ms. This can easily cause audio and video misalignment in fast-moving scenes. Traditional audio noise reduction algorithms (such as Wiener filtering) are ineffective against non-stationary noise (such as sudden human voices), improving the signal-to-noise ratio by less than 10dB.
[0006] Interaction response lag: User action hotspot prediction relies on heuristic rules (such as a fixed area at the bottom of the screen) without integrating historical behavior data. Interaction delays generally exceed 100ms, making it difficult to meet real-time interaction requirements (such as aiming in FPS games). Summary of the Invention
[0007] The present invention provides a virtual engine-driven application management system and method. Through multi-dimensional technological innovation, it systematically solves the core problems of existing virtual engines in resource scheduling, rendering efficiency, audio and video processing, and interactive response, significantly improving the intelligence level of application management and user experience, and has significant industrial application value.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A virtual engine-driven application management system, comprising a spatiotemporal virtualization engine, a semantic rendering decider, a multimodal signal processor, a GPU rendering unit, an intelligent interaction predictor, and a metadata hub;
[0010] The spatiotemporal virtualization engine is used to create a four-dimensional isolated environment consisting of a three-dimensional spatial grid and time slices, generate spatiotemporal units, allocate initial resource allocation tables, and output them to the metadata hub for storage. It also receives the optimized resource allocation tables from the metadata hub and dynamically adjusts spatiotemporal unit resources.
[0011] The semantic rendering decision maker is connected to the spatiotemporal virtualization engine, receives the spatiotemporal environment handle from the spatiotemporal virtualization engine and the initial resource allocation table from the metadata hub, performs semantic segmentation and instance detection on the scene, builds an octree spatial index, and combines the frustum culling algorithm to generate a visible object list. The visible object list contains semantic labels and spatial coordinates and is output to the metadata hub and GPU rendering unit. Based on the semantic labels of the visible objects and their distance from the viewpoint, the LOD level configuration table is generated and input into the GPU rendering unit.
[0012] The multimodal signal processor, connected to the GPU rendering unit and the spatiotemporal virtualization engine, collects the rendered low-resolution video stream and the original audio stream in the four-dimensional isolation environment, calculates the time offset and synchronizes the audio and video through a cross-correlation algorithm; extracts video semantic features, audio Mel-spectrograms, and motion vectors, and fuses them into a multimodal feature vector for input into the metadata hub; receives dynamic bitrate control parameters generated by the metadata hub based on the initial resource allocation table and scene complexity, and adjusts the H.265 encoding bitrate;
[0013] The intelligent interaction predictor is connected to the spatiotemporal virtualization engine and the metadata hub, collects user interaction data, generates interaction feature vectors, and inputs them into the metadata hub. It receives the predicted interaction instructions generated by the metadata hub based on the initial resource allocation table and user profiles, and triggers the spatiotemporal virtualization engine to increase resource quotas for spatiotemporal units in high-frequency interaction areas.
[0014] The metadata hub is bidirectionally connected to the spatiotemporal virtualization engine, semantic rendering decision maker, multimodal signal processor and intelligent interaction predictor, and stores the initial resource allocation table, model parameters and user portrait data; receives the visible object list of the semantic rendering decision maker, the multimodal feature vector of the multimodal signal processor and the interaction feature vector of the intelligent interaction predictor, and trains and updates the scene complexity model and behavior prediction model in combination with the initial resource allocation table; outputs the optimized resource allocation table based on user behavior to the spatiotemporal virtualization engine, provides dynamic bit rate control parameters to the multimodal signal processor, and outputs predicted interaction instructions to the intelligent interaction predictor.
[0015] In this specification, the metadata hub calculates the importance weights of objects in each spatiotemporal unit based on the initial resource allocation table and the visible object list of the semantic rendering decider, marks the spatiotemporal units containing high-importance objects as high priority, and increases their CPU cycle quota by 20%-40% in the optimized resource allocation table.
[0016] In this specification, when the multimodal signal processor receives the dynamic bit rate control parameters from the metadata hub, if the memory bandwidth quota of the corresponding space-time unit in the initial resource allocation table is lower than the threshold, the encoding resolution is automatically reduced to 720p to match the resource limit.
[0017] In this specification, the intelligent interaction predictor queries the initial resource allocation table through the metadata center to obtain the resource usage history data of the space-time unit corresponding to the current interaction hot zone. If the utilization rate of three consecutive time slices exceeds 85%, the space-time virtualization engine is triggered to reclaim resources from low-priority units and reallocate them.
[0018] In this specification, the spatiotemporal virtualization engine requests an initial resource allocation table template from the metadata hub during the system initialization phase, generates an instantiated initial resource allocation table based on device hardware parameters and application type, and stores it in the non-volatile storage area of the metadata hub.
[0019] A virtual engine-driven application management method, comprising:
[0020] S1: Spatiotemporal environment modeling and resource pre-allocation, creating a four-dimensional isolation environment consisting of a three-dimensional spatial grid and time slices, generating a spatiotemporal environment handle and an initial resource allocation table; the spatiotemporal environment handle is used to identify the four-dimensional isolation environment, and the initial resource allocation table contains the CPU cycle and memory bandwidth allocation parameters for each spatiotemporal unit;
[0021] S2: Scene semantic parsing and dynamic occlusion culling. The Mask R-CNN model is used to perform semantic segmentation and instance detection on the scene within the four-dimensional isolation environment generated by S1. The spatial coordinate reference is obtained based on the spatiotemporal environment handle of S1. An octree index is constructed and invisible objects outside the view frustum are eliminated to generate a list of visible objects. The list of visible objects contains semantic labels and spatial coordinates.
[0022] S3: Content-aware super-resolution and dynamic LOD rendering. This method dynamically selects the LOD level based on the semantic labels of objects in the visible object list of S2 and their distance from the viewpoint. It then uses the ESRGAN model to generate a super-resolution image from the low-resolution image processed by S2 to obtain the rendering result.
[0023] S4: Spatiotemporal synchronization and feature extraction of audio and video signals. This process synchronizes the video stream from the rendering result of S3 with the original audio stream in the four-dimensional isolation environment using a cross-correlation algorithm, calculates the time offset, and generates synchronized audio and video streams. It also extracts video semantic features, audio Mel-spectrograms, and motion vectors, fuses them into a multimodal feature vector, and inputs it into the metadata hub.
[0024] S5: Dynamic bitrate control and intelligent noise reduction. This uses the ResNet-50 model to calculate the scene complexity score for S4's multimodal feature vectors and dynamically adjust the H.265 encoding bitrate. It also uses the WaveNet model to perform noise reduction on S4's synchronized audio stream, combining it with the semantic labels in S2's visible object list to generate the encoded audio and video streams.
[0025] S6: User behavior prediction and interaction optimization. This system uses the LSTM-Attention model to analyze S4's multimodal feature vectors and historical user interaction data stored in the metadata hub, predict future user actions, and generate predicted interaction instructions. It adjusts the interface layout based on the prediction results and sends resource preloading requests to S1's four-dimensional isolation environment.
[0026] S7: Dynamic adjustment and feedback loop of spatiotemporal resources. Based on the bitrate fluctuation data of the encoded audio and video streams from S5 and the predicted interaction instructions from S6, the resource utilization deviation of each spatiotemporal unit is calculated to generate an optimized resource allocation table. The initial resource allocation table of S1 is updated through the exponential moving average mechanism.
[0027] In this specification, the division of the three-dimensional space grid in S1 is performed by calculating the number of grids by the physical space range and the grid precision, and the time slice uses a circular buffer to store resource allocation records.
[0028] In this specification, an initial resource allocation table is generated based on viewpoint prediction and historical data, where the historical data comes from resource usage patterns of the first 100 user sessions stored in a metadata hub.
[0029] In this specification, the scene complexity score range in S5 is [0,100]. When the score is >80, the bit rate is increased to 20Mbps, and when the score is <30, the bit rate is reduced to 1Mbps. The WaveNet model improves the signal-to-noise ratio of the denoised audio by 18dB by separating the ambient noise from the noisy audio. The model training data contains 10,000 noisy environment audio samples from historical data.
[0030] In this specification, S6 generates a set of space-time units corresponding to the hot zone based on the hot zone coordinates in the predicted interaction instruction and the three-dimensional space grid mapping of S1. The set of space-time units corresponding to the hot zone includes the hot zone center unit and its adjacent 3×3×3 grid units. S7 calculates the resource utilization deviation of each space-time unit based on the bit rate fluctuation data of the encoded audio and video stream in S5 and the set of space-time units corresponding to the hot zone, and generates an optimized resource allocation table; executes the resource recovery strategy: for space-time units with a utilization rate <30% and a distance from the viewpoint >10 meters, 20% of their CPU cycles and memory bandwidth are recovered and reallocated to the set of space-time units corresponding to the hot zone.
[0031] In summary, this invention solves the core problems of traditional virtual engines through technologies such as four-dimensional spatiotemporal isolation and semantic-aware rendering, improving the intelligent level of application management and user experience. The specific technical effects are as follows:
[0032] Resource scheduling is refined, and utilization is significantly improved:
[0033] Four-dimensional space-time isolation mechanism: Breaking through the limitations of traditional two-dimensional sandboxes, achieving refined allocation of resources in the time dimension, significantly reducing lag in high-load scenarios, and lowering frame rate fluctuations when switching between multiple tasks.
[0034] Dynamic priority marking: Increases the CPU cycle quota for high-priority time and space units, improving overall resource utilization by approximately.
[0035] Intelligent resource recovery and allocation: Resource quotas in high-frequency interaction areas are increased, resources in low-priority units are recovered and redistributed, and the dynamic allocation capability of system resources is enhanced.
[0036] Balance between rendering quality and performance, and optimize the experience:
[0037] Semantic Fusion LOD Technology: Combines semantic tags and distance to generate LOD levels, giving higher rendering priority to key objects and improving the rendering frame rate of complex scenes on the same hardware.
[0038] Content-Aware Super-Resolution Rendering: The ESRGAN model improves image clarity and significantly improves the richness of details in high-resolution scenes.
[0039] Precise audio and video processing for enhanced immersion:
[0040] Time-space synchronization technology: effectively controls audio and video synchronization errors, reduces sound and image misalignment in fast-moving scenes, and enhances immersion.
[0041] Intelligent noise reduction and dynamic bitrate control: The WaveNet model improves the audio signal-to-noise ratio and dynamically adjusts the bitrate to save bandwidth, ensuring better audio and video transmission in different network environments.
[0042] Interactive response is timely and operation is smooth:
[0043] User behavior prediction technology: reduces interaction delays, meets real-time interaction needs, and responds to operations promptly.
[0044] Resource preloading and layout adjustment: Prepare resources for high-frequency interactive areas in advance to reduce loading delays and ensure smoother operations.
[0045] The system is intelligently closed-loop optimized and highly adaptable:
[0046] The metadata hub forms a feedback loop, dynamically adjusts resource allocation and processing strategies, improves the level of intelligence, and enhances adaptability to diverse scenarios and management efficiency, making it valuable for industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 This is a schematic diagram of the framework of the application management system driven by the virtual engine involved in the present invention.
[0049] Figure 2 It is a flowchart of the virtual engine driven application management method involved in the present invention.
[0050] Figure 3 This is a schematic diagram of the process of generating a visible object list involved in the present invention.
[0051] Figure 4 Schematic diagram of the process of obtaining a multimodal feature vector involved in the present invention. DETAILED DESCRIPTION
[0052] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the embodiments of the present invention. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0053] The disclosure below provides many different embodiments or examples for implementing different structures of the embodiments of the present invention. In order to simplify the disclosure of the embodiments of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the embodiments of the present invention. In addition, the embodiments of the present invention may repeat reference numerals and / or reference letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0054] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0055] like Figure 1 As shown, this embodiment provides an application management system driven by a virtual engine, including a spatiotemporal virtualization engine, a semantic rendering decider, a multimodal signal processor, a GPU rendering unit, an intelligent interaction predictor, and a metadata hub;
[0056] The spatiotemporal virtualization engine is used to create a four-dimensional isolated environment consisting of a three-dimensional spatial grid and time slices, generate spatiotemporal units, allocate initial resource allocation tables, and output them to the metadata hub for storage. It also receives the optimized resource allocation tables from the metadata hub and dynamically adjusts spatiotemporal unit resources.
[0057] The semantic rendering decision maker is connected to the spatiotemporal virtualization engine, receives the spatiotemporal environment handle from the spatiotemporal virtualization engine and the initial resource allocation table from the metadata hub, performs semantic segmentation and instance detection on the scene, builds an octree spatial index, and combines the frustum culling algorithm to generate a visible object list. The visible object list contains semantic labels and spatial coordinates and is output to the metadata hub and GPU rendering unit. Based on the semantic labels of the visible objects and their distance from the viewpoint, the LOD level configuration table is generated and input into the GPU rendering unit.
[0058] The multimodal signal processor, connected to the GPU rendering unit and the spatiotemporal virtualization engine, collects the rendered low-resolution video stream and the original audio stream in the four-dimensional isolation environment, calculates the time offset and synchronizes the audio and video through a cross-correlation algorithm; extracts video semantic features, audio Mel-spectrograms, and motion vectors, and fuses them into a multimodal feature vector for input into the metadata hub; receives dynamic bitrate control parameters generated by the metadata hub based on the initial resource allocation table and scene complexity, and adjusts the H.265 encoding bitrate;
[0059] The intelligent interaction predictor is connected to the spatiotemporal virtualization engine and the metadata hub, collects user interaction data, generates interaction feature vectors, and inputs them into the metadata hub. It receives the predicted interaction instructions generated by the metadata hub based on the initial resource allocation table and user profiles, and triggers the spatiotemporal virtualization engine to increase resource quotas for spatiotemporal units in high-frequency interaction areas.
[0060] The metadata hub is bidirectionally connected to the spatiotemporal virtualization engine, semantic rendering decision maker, multimodal signal processor, and intelligent interaction predictor. It stores the initial resource allocation table, model parameters, and user profile data. It receives the visible object list from the semantic rendering decision maker, the multimodal feature vector from the multimodal signal processor, and the interaction feature vector from the intelligent interaction predictor. It uses this information to train and update the scene complexity model and behavior prediction model based on the initial resource allocation table. It outputs an optimized resource allocation table based on user behavior to the spatiotemporal virtualization engine, provides dynamic bitrate control parameters to the multimodal signal processor, and outputs predicted interaction instructions to the intelligent interaction predictor. By training the scene complexity model and behavior prediction model, it dynamically adjusts resource allocation strategies (for example, increasing the CPU quota of high-priority units by 20%-40%).
[0061] In some embodiments, the metadata hub calculates the importance weights of objects in each spatiotemporal unit based on the initial resource allocation table and the visible object list of the semantic rendering decider, marks the spatiotemporal units containing high-importance objects as high priority, and increases their CPU cycle quota by 20%-40% in the optimized resource allocation table.
[0062] In some embodiments, when the multimodal signal processor receives the dynamic bit rate control parameters from the metadata hub, if the memory bandwidth quota of the corresponding space-time unit in the initial resource allocation table is lower than the threshold, the encoding resolution is automatically reduced to 720p to match the resource limitation.
[0063] In some embodiments, the intelligent interaction predictor queries the initial resource allocation table through the metadata hub to obtain the resource usage history data of the space-time unit corresponding to the current interaction hot zone. If the utilization rate of three consecutive time slices exceeds 85%, the space-time virtualization engine is triggered to reclaim resources from low-priority units and reallocate them.
[0064] In some embodiments, the spatiotemporal virtualization engine requests an initial resource allocation table template from the metadata hub during the system initialization phase, generates an instantiated initial resource allocation table based on device hardware parameters and application type, and stores it in the non-volatile storage area of the metadata hub.
[0065] In some embodiments, the spatiotemporal virtualization engine:
[0066] Four-dimensional isolation environment construction:
[0067] Three-dimensional spatial grid division: The number of grid cells is calculated as 100×100×100 based on the physical space range (such as 100m×100m×100m) and grid accuracy (such as 1m / grid). Each grid cell corresponds to a spatial unit. Time slicing uses a circular buffer, with each slice lasting 50ms, which stores resource allocation records for the most recent 10 slices.
[0068] Initial resource allocation table generation:
[0069] During the system initialization phase, templates are instantiated based on device hardware parameters (such as the number of CPU cores and memory capacity) and application types (such as gaming or industrial simulation).
[0070] For example, in a game scenario, the 3×3×3 grid cell at the viewpoint center is initially allocated a CPU cycle of 500 ms / slice and a memory bandwidth of 200 MB / s. In an industrial simulation, the cell where the core device is located is initially allocated a CPU cycle of 800 ms / slice and a memory bandwidth of 300 MB / s.
[0071] Dynamic resource adjustment: Receives the optimization table from the metadata hub, updates resource allocation every 100ms, supports dynamic CPU cycle adjustment with a step size of 100ms / slice, and memory bandwidth adjustment with a step size of 50MB / s.
[0072] In some embodiments, the semantic rendering decider:
[0073] Semantic Segmentation and Instance Detection:
[0074] The Mask R-CNN model is used, with an input image size of 640×480, and outputs 80 common object categories (such as people and mechanical parts). It is pre-trained using the COCO dataset and has an inference speed of 25FPS.
[0075] Construct an octree spatial index with a maximum depth of 6. The leaf nodes store the object list. The frustum culling uses the OpenGL Frustum algorithm to calculate the intersection of the bounding box and the 6 planes of the frustum.
[0076] Dynamic LOD generation:
[0077] Semantic tag priority mapping: The LOD level of the game protagonist (labeled "Player") is two levels higher than that of background objects; the industrial simulation core device (labeled "CoreDevice") is forced to maintain the highest LOD.
[0078] Distance threshold setting: Near view (<5m), medium view (5-15m), and far view (>15m) correspond to LOD levels 3, 2, and 1, and the number of triangles is reduced by 50% at each level.
[0079] In some embodiments, the multimodal signal processor:
[0080] Audio and video synchronization:
[0081] The cross-correlation algorithm uses normalized cross correlation (NCC) with a window size of 500ms and a step size of 10ms to calculate the time offset between the video frame and the audio frame, and the synchronization error is controlled within 8ms.
[0082] Intelligent noise reduction and bit rate control:
[0083] WaveNet model: 10-layer network, 128 filters per layer, 80-dimensional Mel-spectrogram as input audio, and 10,000 training data samples from noisy environments (such as factory noise and crowd sounds). After denoising, the signal-to-noise ratio is improved by 18 dB.
[0084] Dynamic bitrate strategy:
[0085] Scene complexity score > 80 (complex scene): H.265 bitrate 20Mbps, encoding resolution 1080p;
[0086] Score < 30 (simple scenario): The bitrate is 1 Mbps, and the resolution is automatically reduced to 720p (if the memory bandwidth is less than the threshold of 100 MB / s).
[0087] In some embodiments, the intelligent interaction predictor:
[0088] User behavior prediction:
[0089] LSTM-Attention model: Input multimodal feature vector (128-dimensional video semantic features + 80-dimensional audio features + 32-dimensional motion vectors = 240 dimensions), LSTM layer with 256 units, attention mechanism weight matrix dimension 256 × 240, predicting actions 200 ms in the future, with latency < 50 ms.
[0090] Hotspot mapping: The predicted hotspot coordinates are mapped to a three-dimensional grid, generating a set of central cells and adjacent 3×3×3 cells (a total of 3×3×3×10 slices = 270 spatiotemporal cells).
[0091] Resource preloading: Send a preloading request to the spatiotemporal virtualization engine 100ms in advance to load hot zone related model resources (such as button icons and interactive interface data). The amount of preloaded data is limited to 10% of the memory capacity.
[0092] In some embodiments, the metadata hub:
[0093] Model training and updating:
[0094] Scenario complexity model: Based on ResNet-50, it inputs a multimodal feature vector, and the fully connected layer outputs a score. It uses the mean squared error (MSE) loss function and updates the model parameters every second using the exponential moving average (EMA, smoothing factor 0.9) with new data.
[0095] Behavior prediction model: LSTM-Attention weights are fine-tuned every 5 seconds based on the latest 100 interaction data. The Adam optimizer is used with a learning rate of 0.001.
[0096] Resource allocation optimization:
[0097] The CPU cycle of high-priority units is increased by 30% (for example, from 500ms to 650ms / slice), and the memory bandwidth is increased by 20% (for example, from 200MB to 240MB / s).
[0098] Resource recovery strategy: For units with utilization rates less than 30% and distances greater than 10m from the viewpoint, 20% of resources are recovered (e.g., CPU cycle from 400ms to 320ms / slice) and redistributed to hot zone units.
[0099] like Figure 2 As shown, a virtual engine driven application management method includes:
[0100] S1: Spatiotemporal environment modeling and resource pre-allocation, creating a four-dimensional isolation environment consisting of a three-dimensional spatial grid and time slices, generating a spatiotemporal environment handle and an initial resource allocation table; the spatiotemporal environment handle is used to identify the four-dimensional isolation environment, and the initial resource allocation table contains the CPU cycle and memory bandwidth allocation parameters for each spatiotemporal unit;
[0101] S2: Scene semantic parsing and dynamic occlusion culling. The Mask R-CNN model is used to perform semantic segmentation and instance detection on the scene in the four-dimensional isolation environment generated by S1. The spatial coordinate reference is obtained based on the spatiotemporal environment handle of S1. An octree index is constructed and invisible objects outside the view frustum are eliminated to generate a list of visible objects. The list of visible objects contains semantic labels and spatial coordinates.
[0102] S3: Content-aware super-resolution and dynamic LOD rendering. This method dynamically selects the LOD level based on the semantic labels of objects in the visible object list of S2 and their distance from the viewpoint. It then uses the ESRGAN model to generate a super-resolution image from the low-resolution image processed by S2 to obtain the rendering result.
[0103] S4: Spatiotemporal synchronization and feature extraction of audio and video signals. This process synchronizes the video stream from the rendering result of S3 with the original audio stream in the four-dimensional isolation environment using a cross-correlation algorithm, calculates the time offset, and generates synchronized audio and video streams. It also extracts video semantic features, audio Mel-spectrograms, and motion vectors, fuses them into a multimodal feature vector, and inputs it into the metadata hub.
[0104] S5: Dynamic bitrate control and intelligent noise reduction. This uses the ResNet-50 model to calculate the scene complexity score for S4's multimodal feature vectors and dynamically adjust the H.265 encoding bitrate. It also uses the WaveNet model to perform noise reduction on S4's synchronized audio stream, combining it with the semantic labels in S2's visible object list to generate the encoded audio and video streams.
[0105] S6: User behavior prediction and interaction optimization. This system uses the LSTM-Attention model to analyze S4's multimodal feature vectors and historical user interaction data stored in the metadata hub, predict future user actions, and generate predicted interaction instructions. It adjusts the interface layout based on the prediction results and sends resource preloading requests to S1's four-dimensional isolation environment.
[0106] S7: Dynamic adjustment and feedback loop of spatiotemporal resources. Based on the bitrate fluctuation data of the encoded audio and video streams from S5 and the predicted interaction instructions from S6, the resource utilization deviation of each spatiotemporal unit is calculated to generate an optimized resource allocation table. The initial resource allocation table of S1 is updated through the exponential moving average mechanism.
[0107] In some embodiments, the division of the three-dimensional space grid in S1 is performed by calculating the number of grids by using the physical space range and the grid precision, and the time slice uses a circular buffer to store resource allocation records.
[0108] In some embodiments, an initial resource allocation table is generated based on viewpoint prediction and historical data from resource usage patterns of the previous 100 user sessions stored in a metadata hub.
[0109] In some embodiments, the scene complexity score range in S5 is [0,100]. When the score is >80, the bit rate is increased to 20Mbps, and when the score is <30, the bit rate is reduced to 1Mbps. The WaveNet model separates the environmental noise from the noisy audio, thereby improving the audio signal-to-noise ratio by 18dB after noise reduction. The model training data includes 10,000 noisy environment audio samples in historical data.
[0110] In some embodiments, S6 generates a set of space-time units corresponding to the hot zone based on the hot zone coordinates in the predicted interaction instruction and the three-dimensional space grid mapping of S1. The set of space-time units corresponding to the hot zone includes the hot zone center unit and its adjacent 3×3×3 grid units. S7 calculates the resource utilization deviation of each space-time unit based on the bit rate fluctuation data of the encoded audio and video stream in S5 and the set of space-time units corresponding to the hot zone, and generates an optimized resource allocation table; executes a resource recovery strategy: for space-time units with a utilization rate <30% and a distance from the viewpoint >10 meters, 20% of their CPU cycles and memory bandwidth are recovered and reallocated to the set of space-time units corresponding to the hot zone.
[0111] In some embodiments, spatiotemporal environment modeling and resource pre-allocation
[0112] Create a four-dimensional isolated environment consisting of a three-dimensional spatial grid and time slices. First, divide the three-dimensional spatial grid according to the physical space range (such as length, width, and height) and grid accuracy, and calculate the total number of grids. The time dimension is divided into slices of fixed length, and a circular buffer is used to store resource allocation records for the most recent slices. Based on device hardware parameters (such as the number of CPU cores and memory capacity), application type (such as gaming / industrial simulation), and historical data (hot zone distribution of the first 100 user sessions), an initial resource allocation table is generated, containing the CPU cycle and memory bandwidth quotas for each spatial and temporal unit, and stored in the metadata hub.
[0113] Calculation of the number of grids in three-dimensional space:
[0114] ;
[0115] W, H, D: length, width, and height of the physical space (unit: m); 、 、 : Grid accuracy of each axis (unit: m / grid), usually = = ; : Total number of grids.
[0116] Time-sliced circular buffer:
[0117] Slice duration: =50ms; buffer capacity: stores the latest M slices (e.g. M=20, corresponding to 1 second of data); data structure: Buffer[M][time-space unit ID, CPU cycle quota, memory bandwidth quota, utilization].
[0118] Initial resource allocation formula:
[0119] ;
[0120] C: number of CPU cores of the device, : Single-core benchmark period (e.g. 1000ms); : Priority coefficient (hot zone unit =1.5, ordinary unit =1); : Initial CPU cycle quota for a single space-time unit (unit: ms / slice).
[0121] In some embodiments, scene semantic parsing and dynamic occlusion culling:
[0122] The Mask R-CNN model is used to perform semantic segmentation and instance detection on scene images in a four-dimensional isolated environment, and output the semantic labels (such as "Player", "CoreDevice"), spatial coordinates, and confidence levels of the objects. An octree spatial index is constructed, and the scene space is recursively divided until the maximum depth (such as 6 layers) is reached. Each leaf node stores a list of objects. Combined with the frustum culling algorithm, the positional relationship between the object bounding box and the 6 planes of the frustum is calculated, completely invisible objects are eliminated, and a list of visible objects containing semantic labels and spatial coordinates is generated. Figure 3 shown.
[0123] Mask R-CNN reasoning process:
[0124] ;
[0125] N: total number of objects detected, : The three-dimensional center coordinates of the i-th object
[0126] Frustum culling judgment:
[0127] The equation of the viewing cone plane is: ; Corresponding to the left, right, bottom, top, near and far planes of the viewing cone, 、 、 : The normal vector component of the j-th plane of the frustum (unit vector), : The offset of the jth plane of the viewing frustum, used to determine the distance between the plane and the origin.
[0128] The distance from the center of the object's bounding box to the plane:
[0129] ;
[0130] : The three-dimensional center coordinates of the object's bounding box, the bounding box radius: r = diagonal length / 2; elimination condition: if dist>r, the object is completely invisible.
[0131] In some embodiments, content-aware super-resolution and dynamic LOD rendering:
[0132] The LOD level is dynamically selected based on the semantic labels of visible objects and their distance from the viewpoint: key objects (such as the protagonist and core equipment) are prioritized +1, high LOD levels are used for near-field objects, and low LOD levels are used for distant objects. The ESRGAN model is used to generate super-resolution images from low-resolution images, enhancing detail clarity.
[0133] Calculation of the distance between the object and the viewpoint:
[0134] ;
[0135] : viewpoint coordinates;
[0136] LOD level mapping rules:
[0137] ;
[0138] The number of triangles decreases as the LOD level decreases (e.g. LOD3→2 decreases by 50%).
[0139] ESRGAN model architecture:
[0140] Input: low-resolution image (dimensions H times W);
[0141] Output: Super-resolution image (Size sH times sW, s=4 is the scaling factor);
[0142] Loss function: ;
[0143] : Adversarial loss, output by the discriminator; : pixel-level L1 loss; : Feature loss based on VGG network; Weight coefficient in the total loss =0.1, Weight coefficient in the total loss =0.01. > This indicates that pixel-level accuracy takes precedence over semantic feature consistency (because super-resolution tasks focus more on detail clarity); smaller Avoid over-constraining the generator due to perceptual loss (high-level features are more difficult to optimize, and excessive weights can easily lead to unstable training).
[0144] In some embodiments, spatiotemporal synchronization and feature extraction of audio and video signals:
[0145] The time offset between the video and audio streams is calculated using a cross-correlation algorithm, and the synchronization error is controlled within 10ms. Video semantic features (such as the ROI pooling output of Mask R-CNN), audio Mel spectrum (80 dimensions), and motion vector (32 dimensions) are extracted and fused into a 240-dimensional multimodal feature vector for input into the metadata hub. Figure 4 shown.
[0146] Normalized Cross Correlation (NCC) calculation:
[0147] ;
[0148] : A sequence of video frame brightness values (normalized to [0,1]), : audio frame energy sequence (extracted by short-time Fourier transform), T: window length (such as 500ms), : Time offset (unit: ms), synchronization offset: ;
[0149] Multimodal feature vector:
[0150] : Video semantic features (128 dimensions), : Audio Mel spectrum (80 dimensions), : Motion vector (32 dimensions, such as mouse / touch coordinate changes).
[0151] In some embodiments, dynamic bit rate control and intelligent noise reduction:
[0152] A ResNet-50 model calculates a scene complexity score (range [0,100]) for multimodal feature vectors and dynamically adjusts the H.265 encoding bitrate based on the score: 20 Mbps for complex scenes (score > 80) and 1 Mbps for simple scenes (score < 30). Furthermore, a WaveNet model reduces noise in the audio stream, combining semantic labeling to preserve key sounds (such as human voices), improving the signal-to-noise ratio by 18 dB.
[0153] Scene complexity score: ;
[0154] Dynamic rate control rules:
[0155] ;
[0156] If the memory bandwidth quota is less than 100MB / s, the resolution is forcibly reduced to 720p;
[0157] WaveNet denoising model: ;
[0158] : Noisy frequency-time domain signal, : audio signal after noise reduction;
[0159] Training data: 10,000 noisy frequency samples (SNR 5-10dB);
[0160] Signal-to-noise ratio improvement: The signal-to-noise ratio of the output audio after noise reduction by the WaveNet model = the original signal-to-noise ratio of the input audio + 18dB.
[0161] In some embodiments, user behavior prediction and interaction optimization:
[0162] The LSTM-Attention model analyzes multimodal feature vectors and historical user interaction data to predict user actions (such as click locations) 200ms in the future. The predicted hotspot coordinates are mapped to a three-dimensional grid, generating a hotspot set consisting of a central unit and adjacent 3×3×3 units. This model then preloads resources and adjusts the interface layout.
[0163] LSTM-Attention prediction process: ;
[0164] M: input sequence length (e.g. 20, corresponding to 1 second of data), : Predict the probability distribution of actions (such as click coordinate probability);
[0165] Hot zone coordinate mapping: ;
[0166] : three-dimensional coordinates of the hot zone, : Predict the three-dimensional grid cell coordinates corresponding to the center of the hot zone;
[0167] Hot zone unit collection:
[0168] ;
[0169] t: current time slice index (e.g., the t-th slice in a circular buffer), : The three-dimensional grid unit coordinates after the hot zone is expanded, the range is the center unit and its adjacent front, back, left, right, top, and bottom units (a total of 3 x 3 x 3 = 27 spatial units); : The time slice index involved in the hot zone, including the current slice t and the next future slice (t+1), a total of 2 time dimension slices.
[0170] In some embodiments, spatiotemporal resources are dynamically adjusted and feedback looped:
[0171] The resource utilization deviation (the difference between actual utilization and the target of 70%) is calculated for each spatial and temporal unit. Adjustments are triggered when the deviation exceeds ±15%. Resource quotas are updated using an exponential moving average (EMA). 20% of resources are reclaimed from low-priority units (utilization <30% and distance >10m) and reallocated to hotspot units.
[0172] Resource utilization calculation: ;
[0173] : actual CPU cycles or memory bandwidth used, : allocated CPU cycles or memory bandwidth;
[0174] Quota target utilization: =70%, with an allowable deviation of ±15%;
[0175] Exponential Moving Average Update: ;
[0176] : Updated resource quota (new quota); : Resource quota before update (old quota), smoothing factor =0.9, : Current demand quota (calculated based on hot zone or complexity);
[0177] Resource recovery strategy: ;condition: ;
[0178] : The amount of resources recovered.
[0179] In industrial simulation scenarios:
[0180] Resource scheduling: The time and space units where core equipment (such as machine tools) are located are marked as high priority, and the CPU cycle is increased by 40% (from 800ms to 1120ms / slice), ensuring that real-time dynamic calculations are not stalled.
[0181] Rendering optimization: Machine tool components are identified through semantic tags, maintaining the highest LOD level, and ESRGAN super-resolution rendering of details such as screw lines improves operator observation accuracy.
[0182] Interaction optimization: We predict the location of the "parameter adjustment" button that engineers often click, preload related interface resources in advance, and reduce the interaction delay from 120ms to 45ms.
[0183] In a multiplayer online game scenario:
[0184] Audio and video processing: In fast-moving gunfight scenes, the cross-correlation algorithm achieved a synchronization error of 8ms, reducing the audio and video misalignment rate from 15% with traditional solutions to 3%. WaveNet eliminated ambient noise other than gunfire, improving voice chat clarity.
[0185] Resource recycling: If the utilization rate of grass and trees in units far away from players (>10m) is less than 25%, 20% of the resources will be recycled and allocated to player gathering areas. The frame rate in complex team battle scenes will be increased from 45FPS to 60FPS.
[0186] The above embodiments are intended to illustrate the present invention, not to limit the present invention. Therefore, changes in illustrative values or substitutions of equivalent components should still fall within the scope of the present invention.
[0187] From the above detailed description, it will be clear to those skilled in the art that the present invention can indeed achieve the aforementioned objectives and is in compliance with the provisions of the Patent Law.
[0188] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as covering the preferred embodiments and all changes and modifications that fall within the scope of the invention. The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
[0189] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.
[0190] The basic concepts have been described above. It will be apparent to those skilled in the art after reading this application that the above disclosures are merely illustrative and do not constitute limitations on this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and amendments to this application. Such modifications, improvements, and amendments are suggested in this application and remain within the spirit and scope of the exemplary embodiments of this application.
[0191] At the same time, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or more in different places in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.
[0192] Furthermore, those skilled in the art will appreciate that various aspects of the present application may be illustrated and described in terms of a number of patentable categories or situations, including any new and useful process, machine, product, or combination of substances, or any new and useful improvement thereof. Thus, various aspects of the present application may be implemented entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software. Each of the above hardware and software may be referred to as a "unit," "module," or "system." Furthermore, various aspects of the present application may take the form of a computer program product embodied in one or more computer-readable media, with computer-readable program code embodied therein.
[0193] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C programming language, Visual Basic, Fortran2103, Perl, COBOL2102, PHP, ABAP, dynamic programming languages such as Python, Ruby and Groovy, or other programming languages. The program code can be run entirely on the user's computer, or as a standalone software package on the user's computer, or partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0194] In addition, unless expressly stated in the claims, the order of the processing elements and sequences described in this application, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and methods of this application. Although the above disclosure discusses some embodiments of the invention that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this application. For example, although the implementation of the various components described above can be embodied in a hardware device, it can also be implemented as a pure software solution, for example, installation on an existing server or mobile device.
[0195] Similarly, it should be noted that in order to simplify the presentation of this disclosure and thereby facilitate understanding of one or more of the invention's embodiments, the foregoing descriptions of the embodiments of this disclosure sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this approach should not be interpreted as reflecting an intention that the claimed subject matter requires more features than expressly recited in each claim. Rather, the subject matter of the invention may possess fewer features than the single embodiment described above.
Claims
1. A virtual engine driven application management system, characterized in that: It includes a spatiotemporal virtualization engine, a semantic rendering decision maker, a multimodal signal processor, a GPU rendering unit, an intelligent interaction predictor, and a metadata hub; The spatiotemporal virtualization engine is used to create a four-dimensional isolated environment consisting of a three-dimensional spatial grid and time slices, generate spatiotemporal units, allocate initial resource allocation tables, and output them to the metadata hub for storage. It also receives the optimized resource allocation tables from the metadata hub and dynamically adjusts spatiotemporal unit resources. The semantic rendering decision maker is connected to the spatiotemporal virtualization engine, receives the spatiotemporal environment handle from the spatiotemporal virtualization engine and the initial resource allocation table from the metadata hub, performs semantic segmentation and instance detection on the scene, builds an octree spatial index, and combines the frustum culling algorithm to generate a visible object list. The visible object list contains semantic labels and spatial coordinates and is output to the metadata hub and GPU rendering unit. Based on the semantic labels of the visible objects and their distance from the viewpoint, the LOD level configuration table is generated and input into the GPU rendering unit. The multimodal signal processor is connected to the GPU rendering unit and the spatiotemporal virtualization engine to collect the low-resolution video stream output by the rendering and the original audio stream in the four-dimensional isolation environment, calculate the time offset through the cross-correlation algorithm, and synchronize the audio and video; Extract video semantic features, audio Mel-spectrograms, and motion vectors, fuse them into multimodal feature vectors, and input them into the metadata hub. The metadata hub receives dynamic rate control parameters generated based on the initial resource allocation table and scene complexity, and adjusts the H.265 encoding bitrate. The intelligent interaction predictor is connected to the spatiotemporal virtualization engine and the metadata hub, collects user interaction data, generates interaction feature vectors, and inputs them into the metadata hub. It receives the predicted interaction instructions generated by the metadata hub based on the initial resource allocation table and user profiles, and triggers the spatiotemporal virtualization engine to increase resource quotas for spatiotemporal units in high-frequency interaction areas. The metadata hub is bidirectionally connected to the spatiotemporal virtualization engine, semantic rendering decision maker, multimodal signal processor and intelligent interaction predictor, and stores the initial resource allocation table, model parameters and user portrait data; receives the visible object list of the semantic rendering decision maker, the multimodal feature vector of the multimodal signal processor and the interaction feature vector of the intelligent interaction predictor, and trains and updates the scene complexity model and behavior prediction model in combination with the initial resource allocation table; outputs the optimized resource allocation table based on user behavior to the spatiotemporal virtualization engine, provides dynamic bit rate control parameters to the multimodal signal processor, and outputs predicted interaction instructions to the intelligent interaction predictor.
2. The virtual engine driven application management system according to claim 1, characterized in that: The metadata hub calculates the importance weights of objects in each spatiotemporal unit based on the initial resource allocation table and the visible object list of the semantic rendering decider, marks spatiotemporal units containing high-importance objects as high priority, and increases their CPU cycle quota by 20%-40% in the optimized resource allocation table.
3. The virtual engine driven application management system according to claim 1, characterized in that: When the multimodal signal processor receives the dynamic bit rate control parameters from the metadata hub, if the memory bandwidth quota of the corresponding spatiotemporal unit in the initial resource allocation table is lower than the threshold, the encoding resolution is automatically reduced to 720p to match the resource limitation.
4. The virtual engine driven application management system according to claim 1, characterized in that: The intelligent interaction predictor queries the initial resource allocation table through the metadata hub to obtain the resource usage history data of the space-time unit corresponding to the current interaction hot zone. If the utilization rate of three consecutive time slices exceeds 85%, the space-time virtualization engine is triggered to reclaim resources from low-priority units and reallocate them.
5. The virtual engine driven application management system according to claim 1, characterized in that: The spatiotemporal virtualization engine requests an initial resource allocation table template from the metadata hub during the system initialization phase, generates an instantiated initial resource allocation table based on device hardware parameters and application types, and stores the instantiated initial resource allocation table in a non-volatile storage area of the metadata hub.
6. A virtual engine driven application management method, characterized in that: include: S1: Spatiotemporal environment modeling and resource pre-allocation, creating a four-dimensional isolation environment consisting of a three-dimensional spatial grid and time slices, generating a spatiotemporal environment handle and an initial resource allocation table; the spatiotemporal environment handle is used to identify the four-dimensional isolation environment, and the initial resource allocation table contains the CPU cycle and memory bandwidth allocation parameters for each spatiotemporal unit; S2: Scene semantic parsing and dynamic occlusion culling. The Mask R-CNN model is used to perform semantic segmentation and instance detection on the scene within the four-dimensional isolation environment generated by S1. The spatial coordinate reference is obtained based on the spatiotemporal environment handle of S1. An octree index is constructed and invisible objects outside the view frustum are eliminated to generate a list of visible objects. The list of visible objects contains semantic labels and spatial coordinates. S3: Content-aware super-resolution and dynamic LOD rendering. This method dynamically selects the LOD level based on the semantic labels of objects in the visible object list of S2 and their distance from the viewpoint. It then uses the ESRGAN model to generate a super-resolution image from the low-resolution image processed by S2 to obtain the rendering result. S4: Spatiotemporal synchronization and feature extraction of audio and video signals. This process synchronizes the video stream from the rendering result of S3 with the original audio stream in the four-dimensional isolation environment using a cross-correlation algorithm, calculates the time offset, and generates synchronized audio and video streams. It also extracts video semantic features, audio Mel-spectrograms, and motion vectors, fuses them into a multimodal feature vector, and inputs it into the metadata hub. S5: Dynamic bitrate control and intelligent noise reduction. This uses the ResNet-50 model to calculate the scene complexity score for S4's multimodal feature vectors and dynamically adjust the H.265 encoding bitrate. It also uses the WaveNet model to perform noise reduction on S4's synchronized audio stream, combining it with the semantic labels in S2's visible object list to generate the encoded audio and video streams. S6: User behavior prediction and interaction optimization. This system uses the LSTM-Attention model to analyze S4's multimodal feature vectors and historical user interaction data stored in the metadata hub, predict future user actions, and generate predicted interaction instructions. It adjusts the interface layout based on the prediction results and sends resource preloading requests to S1's four-dimensional isolation environment. S7: Dynamic adjustment and feedback loop of spatiotemporal resources. Based on the bitrate fluctuation data of the encoded audio and video streams from S5 and the predicted interaction instructions from S6, the resource utilization deviation of each spatiotemporal unit is calculated to generate an optimized resource allocation table. The initial resource allocation table of S1 is updated through the exponential moving average mechanism.
7. The virtual engine driven application management method according to claim 6, characterized in that: The three-dimensional space grid division in S1 is performed by calculating the number of grids according to the physical space range and grid accuracy, and the time slice uses a circular buffer to store resource allocation records.
8. The virtual engine driven application management method according to claim 6, characterized in that: An initial resource allocation table is generated based on viewpoint predictions and historical data from resource usage patterns of the previous 100 user sessions stored in a metadata hub.
9. The virtual engine driven application management method according to claim 6, characterized in that: The scene complexity score in S5 ranges from [0,100]. When the score is greater than 80, the bitrate is increased to 20 Mbps, and when the score is less than 30, the bitrate is reduced to 1 Mbps. The WaveNet model improves the signal-to-noise ratio of the denoised audio by 18 dB by separating the ambient noise from the noisy audio. The model training data contains 10,000 noisy environment audio samples from historical data.
10. The virtual engine driven application management method according to claim 6, characterized in that: In S6, based on the hot zone coordinates in the predicted interaction instruction and the three-dimensional space grid mapping of S1, a set of space-time units corresponding to the hot zone is generated. The set of space-time units corresponding to the hot zone includes the hot zone center unit and its adjacent 3×3×3 grid units. In S7, based on the bit rate fluctuation data of the encoded audio and video stream in S5 and the set of space-time units corresponding to the hot zone, the resource utilization deviation of each space-time unit is calculated to generate an optimized resource allocation table; and a resource recovery strategy is executed: for space-time units with a utilization rate <30% and a distance from the viewpoint >10 meters, 20% of their CPU cycles and memory bandwidth are recovered and reallocated to the set of space-time units corresponding to the hot zone.
Citation Information
Patent Citations
Virtual reality interactive training system and method based on multi-modal feedback
CN119937798A
Intelligent optimization system and method integrating video resource scheduling and voice emotion recognition
CN120151548A
Methods and Apparatus for Autonomous Robotic Control
US20170024877A1
Server structure for supporting multiple sessions of virtualization
US20190158892A1