A method and system for optimizing and scheduling streaming media in remote desktop collaboration
By combining three-level rendering assets and a vital sign state machine, the rendering level is dynamically adjusted, which solves the problem of visual feedback lag in remote collaboration, ensures image quality and interaction stability during network fluctuations and changes in vital signs, and reduces the risk of medical accidents.
Patent Information
- Application Number
- CN202610665897.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-06-26
AI Technical Summary
In high-risk remote collaboration scenarios, existing technologies can cause audio and video streams to become out of sync with real-time changes in the patient's vital signs, resulting in delayed visual feedback due to network fluctuations. This can affect the accuracy of doctors' decisions and pose a risk of medical accidents.
By employing a three-tiered rendering asset (high-fidelity layer in the cloud, medium-fidelity layer in edge computing, and minimalist skeleton proxy layer on the client) combined with a state machine and network degradation prediction, the rendering level is dynamically adjusted to ensure image quality and interactive stability.
It enables smooth screen switching during network fluctuations and changes in vital signs, avoiding visual lag, ensuring the security and real-time performance of remote collaborative control, and reducing the risk of medical accidents.
Smart Images

Figure CN122293658A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of streaming media communication and remote collaboration, and in particular to a method and system for optimizing and scheduling streaming media video in remote desktop collaboration. Background Technology
[0002] With the deep integration of digital healthcare construction and remote collaboration technology, highly real-time interactive scenarios such as remote rehabilitation training and assisted therapy have been widely applied. To assist remote doctors in making accurate clinical decisions, the system typically needs to display high-fidelity 3D medical models in real time, including complex physiological processes such as muscle deformation and blood flow dynamics. Given the limited computing power of terminals, the industry generally relies on cloud-based graphics processing unit (GPU) clusters for real-time rendering, and then encodes the rendered images into video streams for multi-terminal collaborative presentation, such as real-time web communication media streams built using efficient video coding standards.
[0003] However, existing technologies often fail to keep pace with complex remote collaborative scenarios, resulting in a disconnect between the underlying audio and video transmission control and actual medical operations. For example, patent application CN113891080A discloses a streaming media adaptive transmission and bitrate adjustment method. This method relies solely on network transmission perspectives, evaluating performance based on network layer metrics such as latency and packet loss, and passively reducing the video encoding bitrate or frame rate when encountering network fluctuations.
[0004] This passive adaptive approach, relying solely on network metrics, has serious flaws when applied to high-risk remote dynamic rehabilitation scenarios. Under degraded network conditions, the underlying buffering mechanisms of audio and video inevitably cause data packet accumulation, resulting in long delays of hundreds of milliseconds or even seconds. Simultaneously, rehabilitation patients may experience sudden muscle spasms or rigidity, among other emergency physiological conditions. Because existing underlying transmission controls cannot perceive the patient's real-time vital signs, the system cannot proactively intervene in the transmission pipeline and conserve network bandwidth at critical moments of sudden changes in the patient's condition. This results in the doctor's visual perception of movements and postures lagging significantly behind the patient's actual physical state. This severe lag in visual feedback can easily lead to doctors making incorrect intervention decisions or failing to issue timely stop commands, potentially causing medical accidents such as muscle strains in patients due to the robotic arm or exoskeleton.
[0005] In summary, the core technical problem that the existing technology urgently needs to solve is: how to achieve coordinated linkage between the underlying transmission scheduling of audio and video streams and the real-time changes in the patient's vital signs, so as to completely eliminate the visual feedback lag caused by video buffering under network fluctuations and ensure the security of remote collaborative control. Summary of the Invention
[0006] To address the challenge of coordinating the underlying transmission scheduling of audio and video streams with real-time changes in the patient's vital signs, thereby completely eliminating the problem of visual feedback lag caused by video buffering due to network fluctuations, this invention provides a method and system for optimizing the scheduling of streaming media images in remote desktop collaboration.
[0007] In a first aspect, the present invention provides a method for optimizing and scheduling streaming media images in remote desktop collaboration, employing the following technical solution: A method for optimizing and scheduling streaming media in remote desktop collaboration includes the following steps: S1: Pre-built three-level rendering assets, including a high-fidelity layer deployed in the cloud, a medium-fidelity layer deployed on multiple access edge computing nodes, and a minimalist skeleton proxy layer pre-loaded to the client; S2: The client drives the vital signs state machine to migrate between five vital signs states based on the vital signs sampling data. The five vital signs states are idle state, early warning candidate state, early warning state, candidate alarm state and emergency state in sequence. The early warning state serves as the preheating trigger condition for the skeleton channel. The client uploads the current vital signs state to the audio and video forwarding server. S3: The audio and video forwarding server periodically collects network indicators and weights them to synthesize a continuous network degradation degree, and predicts the network degradation degree at future moments based on historical slopes; S4: The audio and video forwarding server uses the physical condition and network degradation as input to determine the target circuit breaker level according to the mapping rules, and executes the level change action; when the target circuit breaker level is zero, the main screen is rendered with a high-fidelity layer; when the target circuit breaker level is one, the high-fidelity main screen is maintained and the skeleton channel is warmed up; when the target circuit breaker level is two, it switches to a medium-fidelity layer and overlays the skeleton screen; when the target circuit breaker level is three, the video stream is suspended and the local skeleton is rendered by the minimalist skeleton proxy layer.
[0008] This invention provides a multi-layered foundation for system visual presentation by pre-setting three levels of rendering assets: a high-fidelity layer in the cloud, a mid-fidelity layer in edge computing, and a simplified skeleton proxy layer on the client side. It utilizes five vital signs from the client's uplink and the network degradation level, periodically synthesized and predicted by the audio / video forwarding server, as decision inputs. A target circuit breaker level is determined according to mapping rules, and changes are executed at target circuit breaker levels zero to three, respectively: maintaining high-fidelity rendering, warming up the skeleton channel, overlaying the skeleton with mid-fidelity rendering, and rendering the skeleton locally. This design allows streaming media screen scheduling to no longer solely rely on network bandwidth performance, but incorporates the user's real-time physiological state into the scheduling control process. When faced with network deterioration or abnormal vital signs, the system can orderly mobilize rendering assets at different levels, avoiding instantaneous screen freezes or black screens that result in loss of interactive context. Simultaneously, the pre-warming of the skeleton channel triggered by the warning state ensures smooth screen switching and continuous data transmission when entering higher-level circuit breaker levels.
[0009] Preferably, the client is connected to a grip force sensor as a vital sign sensor, and the client drives the vital sign state machine transition based on the vital sign sampling data collected by the grip force sensor, specifically including: The grip force sensor outputs a raw force value at a preset sampling rate. The raw force value is then low-pass filtered at a preset cutoff frequency and then averaged over a preset time span to obtain a stable grip force value. The transition of the vital sign state machine is driven by the relative relationship between the stable grip strength value and the lower warning threshold and the upper alarm threshold in the prescription threshold tensor; the lower warning threshold is taken as a preset ratio of the upper alarm threshold.
[0010] This invention utilizes a combination of pre-processing low-pass filtering and moving average algorithms to effectively eliminate hardware electrical noise and interference from momentary muscle tremors during normal user exertion, ensuring that the force data used for state determination has a high signal-to-noise ratio and is truly representative. The introduction of a prescription threshold tensor allows for setting differentiated force value boundaries for users at different rehabilitation stages, expanding the applicability of the scheduling. The preset ratio division method establishes a clear early warning buffer zone between normal exertion and severe overload, providing ample judgment intervals for the server to initiate pre-screen warm-up scheduling in advance.
[0011] Preferably, the transition conditions of the vital sign state machine specifically include: When the stable grip force value falls into the warning band formed by the lower warning threshold and the upper alarm threshold, it enters the warning candidate state. If it remains in the warning band for a continuous preset number of sampling periods, it transitions to the warning state and sends a warning signal to the audio and video forwarding server, but does not trigger the circuit breaker. When the stable grip force value exceeds the alarm upper limit threshold, it enters the candidate alarm state. If it continues to exceed the limit for a preset duration, it will transition to the emergency state and set the emergency flag bit. Emergency reset requires the stable grip force value to fall below the hysteresis threshold of the alarm upper limit threshold minus the preset hysteresis amount and to remain there for a preset duration.
[0012] Preferably, the network metrics include round-trip time, transmission queue delay, packet loss rate, and bandwidth gap, and the network degradation degree is normalized and obtained by weighting and summing the round-trip time, transmission queue delay, packet loss rate, and bandwidth gap according to their respective preset weight coefficients; The round-trip delay is taken from the current round-trip delay field in the network connection statistics interface; the sending queue delay is obtained by dividing the increment of the total data packet sending delay in the outbound real-time transmission stream statistics by the number of packets sent in the same period; the packet loss rate is taken from the packet loss rate field of the remote inbound real-time transmission stream statistics; and the bandwidth gap is obtained by the difference between the available bandwidth estimated by the congestion control algorithm and the target coding bitrate of the video stream. The audio and video forwarding server collects the aforementioned network metrics at a preset collection period.
[0013] Preferably, the prediction of network degradation at future moments based on historical slope specifically involves: Based on the slope of the change in the network degradation sampling sequence within a preset historical time window, the predicted value of network degradation after a preset prediction time is predicted by first-order extrapolation. When the predicted network degradation value exceeds the preset degradation threshold, even if the current measured network degradation is still in the safe zone, it will be upgraded to the first level in advance and the preheating action of the skeleton channel will be triggered.
[0014] This invention enables transmission scheduling to have advanced defense capabilities based on the extrapolation prediction of historical slopes. In the early stage when the trend of weak network is just emerging but has not yet caused serious communication blockage, it can establish data connection channels in advance and load end-side skeleton assets. This effectively ensures that the local compensatory rendering environment is ready in the instant when the carrying capacity of the transmission channel drops sharply, eliminating the risk of long-term image freeze or operation failure caused by sudden network outage.
[0015] Preferably, the range of network degradation is divided into three levels: low, medium, and high by a preset first threshold and a preset second threshold, and the preset first threshold is less than the preset second threshold. When the vital signs are in an idle state, the target circuit breaker level remains at level zero regardless of the network degradation level. When the vital signs are in a warning state, the network degradation level remains at level zero when it is at a low level, rises to level one when it is at a medium level, and rises to level two when it is at a high level. When the vital signs state is in an emergency, the network degradation level rises from low to level one, from medium to level two, and from high to level three.
[0016] This invention utilizes preset first and second thresholds to classify network degradation into three levels: low, medium, and high, and constructs a two-dimensional judgment matrix that combines the network degradation status with the network degradation level. This two-dimensional scheduling matrix establishes a control logic that prioritizes physiological characteristics: in an emergency, the main video quality of the streaming media is forcibly compressed or the video stream is directly suspended, maximizing the allocation of underlying network transmission resources and edge computing resources to critical command interactions and status monitoring. In an idle state without abnormal conditions, the system fully trusts the current network capacity, maintaining a continuous and stable output of high-fidelity video from the cloud, thus achieving a balance between bandwidth allocation efficiency and visual fidelity.
[0017] Preferably, after determining the target circuit breaker level according to the mapping rule, the level change action is performed according to the asymmetric hysteresis condition of fast upgrade and slow rollback, specifically including: When the level is upgraded, the change will be executed once the triggering conditions corresponding to the target circuit breaker level are continuously met for the preset upgrade duration. When rolling back a level, the triggering conditions corresponding to the adjacent lower level must be continuously satisfied for the preset rollback retention time corresponding to the adjacent lower level before the change is executed, and any of the preset rollback retention times is greater than the preset upgrade retention time. The audio and video forwarding server performs the judgment and modification actions of the mapping rules at a preset decision cycle.
[0018] This invention establishes an asymmetric hysteresis control process that features fast upgrades and slow rollbacks when executing changes to the target circuit breaker level. The shorter upgrade hold time ensures that the system can switch to a low-bandwidth overhead level with extremely fast response speed when encountering sudden channel congestion or when users are in dangerous physiological states, rapidly ensuring the real-time issuance of control commands. The longer rollback hold time avoids the system repeatedly switching between video streams and skeleton streams during network oscillation recovery periods, ensuring a smooth transition in remote visual presentation and long-term interface stability.
[0019] Preferably, the execution actions of the fourth-level circuit breaker specifically include: The high-fidelity layer at level zero is used to transmit a high-efficiency video encoded stream via the main media stream channel; The first level maintains the high-fidelity layer as the main screen, and sends joint quaternions through the data channel at a preset first frequency. The client pre-renders the skeleton screen in the off-screen canvas but does not display it for the time being. When the main screen of the second level is switched to the mid-fidelity layer, the quaternion transmission frequency of the data channel is increased to a preset second frequency, and the skeleton is superimposed on the video with a preset transparency. The preset second frequency is higher than the preset first frequency. The third level suspends the media video stream and sends joint quaternions and angular velocities only through the data channel at the preset second frequency. The client switches to the minimalist skeleton proxy layer to render the webpage graphics library skeleton.
[0020] This invention clearly defines the underlying data flow and specific rendering tasks corresponding to four target circuit breaker levels. This refined division of labor enables precise, step-by-step control of streaming media transmission overhead. In the first level, low-frequency transmission and implicit rendering are used to complete data connection and memory readiness without significantly increasing channel load. In the second and third levels, the transmission frequency of kinematic data is actively increased to ensure that skeleton movements still possess extremely high smoothness and low latency even when the main image quality is reduced or completely lost. In the third level, only a tiny joint data stream is transmitted, enabling continuous synchronization of basic coordinated movements under extremely poor network channel conditions.
[0021] Preferably, the method further includes a doctor-side reverse emergency stop closed-loop step: The doctor generates an emergency stop command by clicking the emergency stop button or by pre-training a voice command. The emergency stop command is transmitted to the audio and video forwarding server through a dedicated data channel with a high priority attribute. The audio and video forwarding server bypasses the conventional processing pipeline and sends the command to the edge gateway through a dedicated user data packet protocol channel. The edge gateway then sends the command to the exoskeleton or electrical stimulation device on the patient's end via the local area network to perform unlocking, decompression, or stop training. The emergency stop command is repeatedly sent a preset number of times at a preset redundancy interval, and an independent emergency stop signaling queue, isolated from the regular uplink channel, is maintained on the audio / video forwarding server side.
[0022] Secondly, this invention provides a streaming media screen optimization and scheduling system for remote desktop collaboration, employing the following technical solution: A streaming media screen optimization and scheduling system for remote desktop collaboration includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the aforementioned streaming media screen optimization and scheduling method for remote desktop collaboration.
[0023] By adopting the above technical solution, a computer program is generated from the above-mentioned method for optimizing and scheduling streaming media in remote desktop collaboration, and stored in the memory so that it can be loaded and executed by the processor. This allows for the creation of a terminal device based on the memory and the processor, making it convenient to use.
[0024] The present invention has the following technical effects: This invention no longer passively waits for severe network buffer congestion before making image quality adjustments. Instead, it achieves adaptive hierarchical response through pre-emptive judgment using a state machine and extrapolation prediction of network degradation based on historical slopes. This dual-dimensional predictive logic can allocate heterogeneous rendering assets in advance before a crisis occurs, ensuring the security and determinism of remote interaction processes.
[0025] Furthermore, by utilizing a three-tiered architecture of cloud, edge nodes, and client, coupled with a progressive quaternion data stream recovery action, the screen layers are smoothly switched sequentially as the transmission environment deteriorates. Even in extreme cases where the video stream is completely suspended, the edge device, using a pre-loaded skeleton proxy layer and lightweight angular velocity parameters, can still provide accurate and coherent posture and motion effect restoration, effectively eliminating the visual lag caused by long latency accumulation in the screen.
[0026] Furthermore, asymmetric hysteresis level control ensures rapid response and smooth recovery during hazardous conditions. Based on this, a completely independent dedicated processing stream and underlying queue are established to carry emergency stop signaling. This control method, combining service isolation and hardware linkage, guarantees a high delivery rate of commands in congested channels and provides underlying system-level security protection for remote physical interactions. Attached Figure Description
[0027] Figure 1 This is a flowchart of a method in a streaming media screen optimization and scheduling method provided in an embodiment of the present invention for remote desktop collaboration; Figure 2 This is a diagram illustrating the grip strength data provided in an embodiment of the present invention. Figure 3 The diagram illustrates the effect of multi-state cooperative transition of the physical state machine provided in this embodiment of the invention. Detailed Implementation
[0028] The system architecture supporting this method is distributed in a three-tiered physical topology of cloud, edge, and endpoint. Its hardware configuration and interconnection relationships are as follows: The first layer is the acquisition and execution layer. This layer comprises two branches: patient-side hardware and doctor-side hardware. The patient-side hardware uses an industrial-grade resistive strain gauge grip force sensor as the vital signs sensor. This sensor incorporates a Wheatstone bridge structure, with a preferred measurement range of 0 to 50 kg and a preferred resolution of 0.01 kg. The output of the sensor is amplified by a pre-amplifier circuit, then converted to digital by a 24-bit high-precision analog-to-digital converter, and finally connected to the patient-side edge control box via a Universal Serial Bus interface or Bluetooth Low Energy protocol. The patient-side edge control box uses a quad-core high-performance microcontroller as the main control chip. This chip runs an embedded real-time operating system and connects to the patient-side exoskeleton or electrical stimulation device via a controller area network bus or industrial Ethernet bus. The exoskeleton device has servo motors and inertial measurement units embedded in its joints. The inertial measurement units output joint angular velocity and quaternion posture data at a 1 kHz sampling rate. The electrical stimulation device is equipped with a relay switch array to support hardware-level power failure in emergency situations. Finally, the patient-side edge control box further connects to the network via Gigabit Ethernet or a 5G mobile communication module. Corresponding to the patient-side hardware, the doctor-side hardware is a workstation terminal equipped with a multi-core central processing unit, a dedicated graphics processor, a display screen with an image resolution of at least 1920 x 1080 pixels, a microphone array, and a touchscreen. This workstation terminal runs a browser or dedicated client program that supports a real-time web communication protocol stack and connects to the hospital's backbone network via Gigabit Ethernet.
[0029] The second layer is the edge processing layer. This layer consists of multiple multi-access edge computing nodes distributed across the geographical area where the patient is located. Each multi-access edge computing node is equipped with a graphics processor (GPU) with tensor acceleration capabilities. The GPU runs a containerized rendering service process to handle the geometric model calculations and rasterization tasks of the mid-fidelity layer. Considering link latency constraints, the physical distance between the multi-access edge computing nodes and the client is preferably no more than 50 kilometers, and the corresponding network round-trip latency is preferably controlled within 20 milliseconds. Furthermore, load migration and failover are achieved between the multi-access edge computing nodes through a software-defined network controller.
[0030] Thirdly, there is the cloud core layer. This cloud core layer includes a cloud rendering server, the audio / video forwarding server, a signaling server, cloud audit storage, and a local hospital information system access gateway. The cloud rendering server is deployed with a multi-GPU cluster to handle offline modeling and real-time rendering tasks for the high-fidelity layer. The audio / video forwarding server uses a selective forwarding unit architecture as a media stream relay node and integrates a circuit breaker decision engine module, a network quality monitoring module, an emergency stop signaling distribution module, and an audit log module. The signaling server is responsible for session establishment and prescription threshold tensor distribution based on the Socket communication protocol. The cloud audit storage uses a distributed object storage cluster and enables version-immutable write characteristics. The local hospital information system interfaces with the audio / video forwarding server via a dedicated line and a health information exchange standard interface.
[0031] Communication between different layers occurs via Real-Time Transport Protocol (RTP), Data Channel Protocol (DSP), User Datagram Protocol (UDP), and Transmission Control Protocol (TCP). Specifically, the main media stream channel carries video streams based on RTP, the data channel transmits joint quaternions, control signaling, and emergency stop commands, and the dedicated UDP channel is used for low-latency direct transmission of reverse emergency stops. This constructs a remote medical monitoring physical link characterized by high availability and low latency.
[0032] Example 1: This invention discloses a method for optimizing and scheduling streaming media feeds in remote desktop collaboration, referring to... Figure 1 This includes steps S1-S4: S1: Pre-built three-level rendering assets, including a high-fidelity layer deployed in the cloud, a medium-fidelity layer deployed on multiple access edge computing nodes, and a minimalist skeleton proxy layer pre-loaded to the client.
[0033] During the system initialization phase, rendering assets need to be distributed at different physical layers to build a visual buffer. Therefore, this step breaks down the rendering assets into three layers and deploys them separately.
[0034] In the first phase, the high-fidelity layer is deployed on the graphics processing unit cluster of the cloud rendering server. The high-fidelity layer carries a high-precision soft tissue medical model, which includes muscle tissue deformation, blood flow dynamics, and advanced texture material elements. It is composed of subdivided surface features, and the polygon count of the high-precision soft tissue medical model is preferably between 500,000 and 1,000,000 faces, accompanied by a physically based rendering texture map with a resolution of 4,000 x 4,000. The overall size of the high-precision soft tissue medical model is preferably between 200 and 500 megabytes. During runtime, the high-precision soft tissue medical model is rendered in real-time by the cloud rendering server using the graphics rendering interface, and transmitted via the main media stream channel at a rate of 60 frames per second after compression using a high-efficiency video encoding protocol, such as the second-generation high-efficiency video encoding protocol, dedicated to high-definition rendering under normal conditions.
[0035] In the second stage, the mid-fidelity layer is deployed on the aforementioned multi-access edge computing nodes. The mid-fidelity layer carries a mid-fidelity skeleton model. Compared to the high-precision soft tissue medical model, the overall volume of the mid-fidelity skeleton model is preferably approximately 20 megabytes, the number of polygons is preferably 50,000 to 100,000, and the texture mapping resolution is preferably 1024 x 1024. Only the human skeleton and basic skin are retained, while muscle details and advanced textures are stripped away. During runtime, the rendering of the mid-fidelity skeleton model is handled by the lightweight graphics processor of the multi-access edge computing nodes. The rendering pipeline runs on a web graphics library or Vulkan graphics interface environment, with an output frame rate preferably of 30 frames per second. The model is compressed using a video encoding protocol, such as Part 10 Advanced Video Coding Protocol, and then sent back to the client as a visual transition between the high-fidelity and minimalist models.
[0036] In the third stage, the simplified skeleton proxy layer is preloaded into the client's browser memory. This simplified skeleton proxy layer carries a simplified topological skeleton model, preferably with a file size of less than 1 megabyte, containing only seventeen core joints and their rigid body connections. These seventeen core joints specifically include the head, neck, left and right shoulders, left and right elbows, left and right wrists, torso, left and right hips, left and right knees, and left and right ankles. In terms of data organization, the simplified topological skeleton model uses JSON data format to describe the topological skeleton, and is accompanied by a simplified GLB binary file describing the mesh. During the session establishment phase, the client pre-downloads the simplified topological skeleton model via Hypertext Transfer Protocol and stores it in the browser's index database. During rendering, the client calls the web graphics library to directly draw the simplified topological skeleton model in real time, with the end-to-end rendering latency preferably controlled within 30 milliseconds.
[0037] The high-precision soft tissue medical model, the medium-fidelity skeleton model, and the minimalist topological skeleton model are respectively located on the cloud rendering server, the multi-access edge computing node, and the client, serving as the rendering source for subsequent steps.
[0038] S2: The client drives the vital sign state machine to migrate between five vital sign states based on the vital sign sampling data. The five vital sign states are, in order, idle state, early warning candidate state, early warning state, candidate alarm state and emergency state. The early warning state serves as the preheating trigger condition for the skeleton channel. The client then uploads the current vital sign state to the audio and video forwarding server.
[0039] It should be noted that before the vital signs state machine formally starts the evaluation, both the client and the audio / video forwarding server need to hold prescription threshold tensors in advance. These tensors serve as the vital signs-side benchmark for the vital signs state machine's transition determination and the network-side benchmark for threshold division in the subsequent mapping rules. Therefore, prescription threshold tensors need to be sent to the client and the audio / video forwarding server first. The initial distribution of the prescription threshold tensors is uniformly scheduled by the signaling server, and its distribution data stream specifically includes the following five stages.
[0040] Phase 1: The rehabilitation physician inputs the patient's personalized prescription parameters into the prescription configuration console on the workstation terminal of the doctor's end hardware. The workstation terminal converts the input parameters into a three-dimensional prescription threshold tensor. This three-dimensional prescription threshold tensor is described using nested objects in JSON data format, covering three dimensions: first, the vital signs dimension, including an upper alarm threshold and a lower warning threshold; second, the network dimension, including a hard latency threshold and a soft latency threshold; and third, the vital signs state machine dimension, including a minimum duration and a reset hold time. The lower warning threshold is set by default to a preset percentage of the upper alarm threshold; in this embodiment, the preset percentage is preferably 80%.
[0041] Phase Two: The doctor's hardware uploads the prescription threshold tensor to the signaling server via the hospital's backbone network. Upon receiving the tensor, the signaling server uses a preset asymmetric signature algorithm to sign it. In this embodiment, the preset asymmetric signature algorithm preferably uses RSA-SHA256 with a key length of 2048 bits. Then, the signaling server encapsulates the signed prescription threshold tensor and a monotonically increasing prescription version number into a single data packet for distribution.
[0042] Phase 3: The signaling server simultaneously distributes the data packets to both the audio / video forwarding server and the client in broadcast mode. The link to the audio / video forwarding server passes through the low-latency backplane network within the cloud core layer; the link to the client passes through the hospital backbone network and the public network or a 5G mobile communication network.
[0043] Phase Four: After receiving the transmitted data packet, the audio / video forwarding server and the client respectively use the pre-configured public key of the signaling server to verify the signature of the transmitted data packet. Only after successful verification will the process proceed to the next phase; if verification fails, the transmitted data packet is discarded and a verification failure receipt is sent back to the signaling server, which then triggers a retransmission mechanism.
[0044] Phase 5: After successful verification, the client writes the prescription threshold tensor into the browser index database for local persistent storage; the audio / video forwarding server writes the prescription threshold tensor into an in-memory database, preferably using Redis with master-slave replication for high availability.
[0045] In this way, both ends complete the solidification of the prescription threshold tensor, and can be restored from the persistence layer of their respective browser index database and memory database in case of disconnection and reconnection, and the mismatch between the old and new versions is avoided by comparing the prescription version number.
[0046] After the above five stages, the client copy of the prescription threshold tensor will be directly read by the judgment logic of the vital sign state machine, and driven by the vital sign state machine inside the client to transition between the five vital sign states based on continuous vital sign sampling data, thereby finely characterizing the patient's spasticity risk. The specific process is as follows: The client is connected to a grip force sensor, which acts as a vital sign sensor. A strain gauge within the sensor generates a microvolt-level voltage difference under grip force. The built-in Wheatstone bridge structure of the sensor outputs a microvolt-level voltage difference under grip force. This microvolt-level voltage difference is amplified differentially by the pre-amplifier circuit and then input to the 24-bit high-precision analog-to-digital converter for digitization. The 24-bit high-precision analog-to-digital converter continuously outputs the original force value at a preset sampling rate; in this embodiment, the preset sampling rate is preferably 200 Hz. The original force value is transmitted to the client via the Universal Serial Bus interface or the Bluetooth Low Energy protocol, received by the JavaScript acquisition thread in the client through a WebUSB interface or a WebBLE interface, and finally stored in a circular buffer queue.
[0047] To filter out high-frequency muscle tremor noise, the client inputs the original force value into a low-pass filter with a preset cutoff frequency. In this embodiment, the low-pass filter is preferably a second-order Butterworth low-pass filter with a cutoff frequency of 50 Hz. Considering the performance of JavaScript execution, the difference equation of the second-order Butterworth low-pass filter is implemented by the WebAssembly module to ensure real-time calculation. Then, a moving average calculation over a preset time span is performed on the filtered result. In this embodiment, the preset time span is preferably 40 milliseconds. Finally, a smooth and stable grip force value is obtained. The stable grip force value and timestamp are written to the main thread message queue.
[0048] Next, the vital signs state machine performs a five-sign state transition based on the relative relationship between the stable grip strength value and the lower warning threshold and the upper alarm threshold in the prescription threshold tensor. The vital signs state machine is implemented using a finite state machine pattern, internally maintaining the following state variables: current state, previous state, time of entering the current state, consecutive hit count, protective lock flag, and emergency flag. Whenever a new stable grip strength value is reached, the vital signs state machine performs a transition evaluation according to the following logic: First, when the stable grip strength value is in the safe zone, that is, below the warning lower limit threshold, the vital signs state machine is anchored to the idle state.
[0049] Secondly, once the stable grip strength value suddenly increases and exceeds the lower warning threshold, falling into the warning band formed by the lower warning threshold and the upper alarm threshold, the vital signs state machine immediately switches to the warning candidate state, sets the warning candidate state flag internally, and starts the continuous hit counter; if the stable grip strength value does not fall out of the warning band within a subsequent preset number of sampling periods, the vital signs state machine transitions to the warning state, sets the warning state flag internally, and sends a warning signal to the audio / video forwarding server but does not directly trigger the circuit breaker. At the same time, the warning state serves as the preheating trigger condition for the skeleton channel. In this embodiment, the preset number of sampling periods is preferably two sampling periods.
[0050] Furthermore, when the stable grip strength value further climbs above the alarm upper limit threshold, the vital signs state machine enters the candidate alarm state, internally sets the candidate alarm state flag, and starts an internal timer. If the stable grip strength value remains in the over-limit state for a subsequent preset duration, the vital signs state machine eventually transitions to the emergency state, internally sets the emergency state flag, and then sends a high-priority alarm signal. In this embodiment, the preset duration is preferably 30 milliseconds. If the stable grip strength value fails to remain in the over-limit state within the preset duration, i.e., falls back to or below the alarm upper limit threshold at any sampling moment within the preset duration, the vital signs state machine returns to the warning state and clears the internal timer.
[0051] Finally, the reset of the emergency state must meet the stringent hysteresis fall-off condition, that is, the stable grip force value must fall back to the hysteresis threshold below the alarm upper limit threshold minus the preset hysteresis amount, and the preset holding time must not exceed the limit before it can return to the degraded state; in this embodiment, the preset hysteresis amount is preferably 0.2 kg, and the preset holding time is preferably 500 milliseconds.
[0052] Finally, the client packages the vital sign sampling data, carrying the status identifier, timestamp, and the current stable grip strength value, into JSON data format and uploads it to the audio / video forwarding server in real time through the data channel. The reliability parameters of the data channel are configured for ordered transmission and a maximum retransmission count of 0, i.e., unreliable transmission mode, to ensure minimal end-to-end latency.
[0053] Furthermore, to prevent physical device failure, the client additionally maintains a sensor health watchdog task. When the vital signs sensor experiences a continuous communication loss exceeding a preset silence time, or when the raw force value output by the 24-bit high-precision analog-to-digital converter is continuously in the full-scale saturation region, the sensor health watchdog task triggers the system to automatically intercept the current erroneous data stream, forcibly push the vital signs state machine into a protective lock state, and alert the doctor and patient with a prominent color through the user interface; until the connection is restored and a certain number of valid heartbeat packets are continuously received, in this embodiment, the number of valid heartbeat packets is preferably 10, and the preset silence time is preferably 200 milliseconds.
[0054] S3: The audio and video forwarding server periodically collects network indicators and weights them to synthesize a continuous network degradation degree, and predicts the network degradation degree at future moments based on historical slopes.
[0055] While the vital signs assessment is performed in parallel, the system also needs to assess the transmission channel quality. The logical carrier of this step is the network quality monitoring module in the audio / video forwarding server, which resides as an independent thread within the audio / video forwarding server process.
[0056] Specifically, the network quality monitoring module calls the network connection statistics interface to collect multi-dimensional raw indicators at a preset collection period. In this embodiment, the preset collection period is preferably 200 milliseconds. The network connection statistics interface corresponds to the peer-to-peer connection acquisition statistics application interface in the web real-time communication protocol. The multi-dimensional raw indicators include round-trip time, sending queue delay, packet loss rate, and bandwidth gap. The round-trip time is taken from the current round-trip time field of the candidate pair statistical objects returned by the network connection statistics interface; the sending queue delay is obtained by dividing the current period increment of the total data packet sending delay field in the outbound real-time transmission statistical object by the current period increment of the field of sent data packets in the same period; the packet loss rate is taken from the packet loss score field of the remote inbound real-time transmission statistical object; and the bandwidth gap is obtained from the difference between the available bandwidth estimated by the Google congestion control algorithm or the bandwidth estimation-based BBR algorithm and the target encoding bitrate of the video stream. The network quality monitoring module writes the above four raw values along with the collection timestamp into a fixed-length circular buffer queue for subsequent trend analysis.
[0057] Subsequently, the network quality monitoring module normalizes the above four raw data items sequentially using reference values: the round-trip time is normalized using the reference maximum round-trip time as the denominator, which is preferably 300 milliseconds in this embodiment; the transmission queue delay is normalized using the reference maximum queue delay as the denominator, which is preferably 100 milliseconds in this embodiment; the packet loss rate is already normalized to the range of 0 to 1; and the bandwidth gap is normalized to the range of 0 to 1 using the target coding bitrate of the video stream as the denominator. After processing, the round-trip time, transmission queue delay, packet loss rate, and bandwidth gap are each assigned their own preset weight coefficients and linearly weighted and summed to obtain the continuous network degradation degree of the normalized values. In this embodiment, the weight coefficients for the round-trip time are preferably 30%, the transmission queue delay is 35%, the packet loss rate is 20%, and the bandwidth gap is 15%.
[0058] Considering the sensing delay between network degradation occurrence and actual detection, and the processing delay between actual detection and circuit breaker decision response, if the circuit breaker decision engine module only makes decisions based on the currently measured continuous network degradation degree, its triggering timing will inevitably lag behind the actual degradation event, thus missing the optimal time window for skeleton channel warm-up. To eliminate this response lag, this step introduces a first-order trend predictor for prediction, thereby shifting the judgment benchmark from the current measured value to the future predicted value.
[0059] Specifically, the first-order trend predictor runs as a subroutine of the network quality monitoring module, and its workflow is as follows: First, the first-order trend predictor extracts all continuous network degradation samples within a preset historical time window from the circular buffer queue to construct a time series. In this embodiment, the preset historical time window is preferably the past 1 second.
[0060] The slope is then fitted using the least squares method. The formula for calculating the slope in this embodiment is:
[0061] In the formula, This refers to the i-th time value of the time series within the historical time window. The degradation degree of the i-th continuous network within the historical time window is the time series. and These are the mean value of the time series and the mean value of its degradation within the preset historical time window, respectively. This represents the total number of samples included within the historical time window. Indicates the slope.
[0062] The first-order trend predictor uses the slope as an extrapolation coefficient to predict the network degradation value after a preset prediction time. The calculation formula is as follows:
[0063] In the formula, The continuous network degradation degree is obtained from the latest data collection, and Δt is the preset prediction duration, which is preferably 500 milliseconds in this embodiment. The predicted network degradation value after a preset prediction time.
[0064] Once the predicted network degradation value exceeds the preset degradation threshold, even if the currently measured continuous network degradation does not exceed the preset degradation threshold, the network quality monitoring module sends a preheating trigger signal to the client, independent of the main link for level determination. Upon receiving the preheating trigger signal, even if the client's main screen rendering source is still provided by the high-fidelity layer, it silently pre-renders the skeleton image in the off-screen canvas area outside the screen. Simultaneously, the audio / video forwarding server begins sending joint quaternions via the data channel at the preset first frequency as input for the skeleton image pre-rendering. In this embodiment, the preset degradation threshold is preferably 0.7.
[0065] This demonstrates that this step transforms the traditional passive and delayed response into a proactive and predictive deployment, thus securing an extremely valuable time window for the advance scheduling of downstream resources.
[0066] S4: The audio and video forwarding server uses the physical condition and network degradation as input to determine the target circuit breaker level according to the mapping rules, and executes the level change action; when the target circuit breaker level is zero, the main screen is rendered with a high-fidelity layer; when the target circuit breaker level is one, the high-fidelity main screen is maintained and the skeleton channel is warmed up; when the target circuit breaker level is two, it switches to a medium-fidelity layer and overlays the skeleton screen; when the target circuit breaker level is three, the video stream is suspended and the local skeleton is rendered by the minimalist skeleton proxy layer.
[0067] It should be noted that after acquiring the vital signs and the continuous network degradation, the circuit breaker decision engine module in the audio / video forwarding server is activated. This step maps the above two-dimensional input to a progressive circuit breaker ladder from level zero to level three, thereby avoiding visual cliffs caused by sudden drops in video quality.
[0068] Specifically, the circuit breaker decision engine module runs the mapping rules in an independent polling thread with a preset decision period. In this embodiment, the preset decision period is preferably 100 milliseconds. The value range of the continuous network degradation degree is divided into three levels: low, medium, and high by a preset first threshold and a preset second threshold. In this embodiment, the preset first threshold is preferably 0.3, and the preset second threshold is preferably 0.7.
[0069] The mapping rules are pre-stored in memory in the form of a two-dimensional lookup table, and the rules are as follows: When the status indicator is in the idle state, the target circuit breaker level remains at level zero regardless of the level of continuous network degradation. When the status indicator is in the warning state, the continuous network degradation remains at level zero if it is at a low level, rises to level one if it is at a medium level, and rises to level two if it is at a high level. When the status indicator is in the emergency state, the continuous network degradation rises to level one if it is at a low level, rises to level two if it is at a medium level, and enters level three if it is at a high level.
[0070] After determining the target circuit breaker level according to the mapping rules, the level change action is executed according to the asymmetric hysteresis condition of fast upgrade and slow rollback. Specifically, the circuit breaker decision engine module maintains an upgrade timer and a rollback timer for each circuit breaker level. When the target circuit breaker level determined in real time is higher than the current actual execution level, the upgrade logic is triggered: the system checks whether the triggering conditions corresponding to the target circuit breaker level continuously meet the preset upgrade hold duration. In this embodiment, the preset upgrade hold duration is preferably 100 milliseconds. Once it is met, the change is executed. When the target circuit breaker level determined in real time is lower than the current actual execution level, the rollback logic is triggered: the system requires that the triggering conditions corresponding to the adjacent lower level must continuously meet the preset rollback hold duration corresponding to the adjacent lower level before the change is executed. Through this asymmetric hysteresis mechanism, the frequent upgrade and downgrade oscillations of the screen caused by instantaneous network jitter are effectively filtered out. After the above hysteresis condition determination confirms that a level change needs to be executed, the system triggers the corresponding execution action according to different circuit breaker levels. In this embodiment, the preset rollback hold duration is preferably 2000 milliseconds.
[0071] For level zero, the cloud rendering server performs real-time rendering of the high-precision soft tissue medical model in the high-fidelity layer. The rendered video frames are compressed into a high-efficiency video encoded bitstream at a rate of 60 frames per second using an efficient video encoding protocol. This bitstream is then sent to the client via the main media stream channel based on a real-time transmission protocol. Finally, the client decodes the bitstream and displays it as the main screen on the doctor's hardware display. At the bitrate control level, the audio / video forwarding server calls the sending parameter configuration interface of the web real-time communication protocol to set the maximum encoding bitrate field in the real-time transmission encoding parameters of the video sending track to 4 megabits per second. This serves as the hardware operating upper limit constraint threshold for the target encoding bitrate, thereby strictly locking the maximum encoding bitrate under normal conditions.
[0072] Upon entering the first level, the main media stream channel remains unchanged, but the audio / video forwarding server actively calls the aforementioned interface to compress the value of the maximum encoding bitrate field to 3.5 megabits per second, thereby simultaneously lowering the allowable dynamic fluctuation limit of the target encoding bitrate, thus freeing up approximately 500 kilobits per second of spare bandwidth. Simultaneously, the audio / video forwarding server activates a covert channel: first, it sends a data channel control frame to notify the client to enter preheating mode; after the client responds, the audio / video forwarding server then sends the patient's real-time joint quaternions via the data channel at a preset first frequency. In this embodiment, the preset first frequency is preferably 30 Hz. The joint quaternions are transmitted in a binary array buffer format, with each record containing quaternions for 17 joints and a timestamp; after receiving the data, the client calls the web graphics library to silently pre-render the skeleton image in an off-screen canvas area but does not display it immediately. Through this processing, a hot-standby minimalist rendering pipeline is immediately ready, allowing the main screen to switch within a single frame cycle when an emergency circuit breaker is actually activated.
[0073] When the situation deteriorates to level two, the audio / video forwarding server switches the main screen rendering source from the high-fidelity layer in the cloud rendering server to the mid-fidelity layer in the multi-access edge computing node through session description protocol renegotiation or dynamic switching technology based on the cascade stream. Simultaneously, the joint quaternion sending frequency of the data channel is increased from the preset first frequency to the preset second frequency; in this embodiment, the preset second frequency is preferably 60 Hz. After receiving the video source output from the mid-fidelity layer and the joint quaternion dual sources, the client's rendering and compositing thread uses canvas overlay technology to place the mid-fidelity skeleton model rendering video corresponding to the mid-fidelity layer at the bottom layer and the minimalist topological skeleton model at the top layer. The minimalist topological skeleton model is then overlaid on the mid-fidelity skeleton model rendering video with a preset transparency, thereby forming visual redundancy through dual-source fusion; in this embodiment, the preset transparency is preferably 30%.
[0074] Given that the time difference between the midfidelity layer video and the joint quaternion is different, the alignment submodule of the client further extracts the absolute timestamps attached to the midfidelity layer video and the joint quaternion respectively, and performs millisecond-level alignment through a cache queue; for cases where the skeleton frame is lagging, a spherical linear interpolation algorithm is used to perform kinematic frame interpolation compensation on the joint quaternion.
[0075] Upon reaching the third level, the audio / video forwarding server invokes the sending track replacement interface of the web real-time communication protocol and sets the input parameter track object of this interface to a null value, thereby disabling the video sending track and suspending the media stream video stream. Simultaneously, the rendering pipeline of the high-fidelity layer in the cloud rendering server enters a paused state to conserve graphics processor resources. The audio / video forwarding server continues to send the joint quaternions and angular velocities through the data channel at the preset second frequency, drastically reducing the overall bandwidth usage to approximately 10 kilobits per second. At this point, the client instantly switches the main screen display source to the pre-warmed minimalist skeleton proxy layer, and the web graphics library performs local rendering of the minimalist topological skeleton model, with end-to-end latency controlled within 30 milliseconds.
[0076] During multi-level streaming media scheduling, if the doctor observes an irreversible abnormality in the patient through video or a local skeleton, hardware-level intervention can be executed through the system. Specifically, this includes: the doctor's workstation generating a high-priority emergency stop command via clicking the emergency stop button on the UI or triggering a pre-trained voice command. This command is directly marked with the highest priority DSCP and transmitted directly to the audio / video forwarding server via a dedicated data channel. The audio and video forwarding server maintains an independent emergency stop signaling queue that is physically isolated from the regular uplink channel. When an emergency stop command is detected, the scheduling engine bypasses the processing pipeline of the regular media stream and data stream and sends it to the edge gateway where the patient is located via a dedicated UDP direct connection channel with extremely low network encapsulation overhead. The edge gateway rapidly distributes the command via the local area network to the servo motor driver or electrical stimulation device relay of the patient's exoskeleton, executing hardware-level unlocking, force release, or power-off to stop training. To ensure extremely high reliability, after the emergency stop command is triggered, the system automatically retransmits it a preset number of times at a preset redundancy interval, ensuring absolute delivery even under severe packet loss. In this embodiment, the preset redundancy interval is preferably 10 milliseconds, and the preset number of redundancy attempts is 5.
[0077] To demonstrate the effectiveness of the solution, relevant experiments were conducted. Below are the images obtained from the experiments: Figure 2 This is a graph illustrating the grip strength data. Thin dotted lines represent the original force values; smooth solid lines represent stable grip strength values; short dashed lines represent the upper warning threshold; and dotted-dash lines represent the lower warning threshold. Double dotted-dash lines represent hysteresis thresholds. The background fill layer represents the risk buffer zone between normal force application and extreme overload, i.e., the warning zone.
[0078] The image clearly shows that the solid line perfectly isolates the disordered muscle tremors and electrical noise within hundreds of milliseconds, closely matching the actual physical force envelope. This demonstrates that the filtering algorithm plays a decisive role in improving the signal-to-noise ratio of the core input data involved in state determination.
[0079] The inclusion of a warning zone filler layer in the image provides a multi-level grayscale prediction space for the physical force application process. The system does not need to wait for the force curve to violently impact the top dotted line boundary before taking defensive action; instead, it can trigger the collaborative contingency plan by sending an uplink signal when the stable force value enters the zone, demonstrating an extremely high level of safety redundancy.
[0080] Figure 3 The diagram shows the effect of multi-state collaborative transfer of the vital signs state machine. The continuous broken lines in the diagram, which are distributed in a stepped pattern and marked with diamonds, represent the evolution and transfer process of the client's internal vital signs state machine between five discrete physiological characteristic states: idle state, early warning candidate state, early warning state, candidate alarm state, and emergency state, driven by continuous and stable grip force values.
[0081] As can be seen from the images, at the initial stage of entering the warning zone and at the alarm threshold, the state machine did not immediately and blindly jump to the next state, but instead underwent a time-span verification of the candidate states. This strongly demonstrates that the time threshold mechanisms, such as the continuous preset sampling period and preset duration, effectively played their role in preventing spoofing actions and ensuring the high authenticity of the uplink status signaling.
[0082] Observing the peak region of the image, once an emergency state is confirmed, even if the applied force value instantly drops to the upper limit, the curve remains in the highest position until the stringent condition of dropping below the hysteresis limit and maintaining it for a preset duration is fully met, at which point it smoothly switches back to the lower level state. This trajectory fully demonstrates the strong protection logic designed by the system for patient spasm scenarios, completely avoiding the potential for secondary control and physiological damage caused by prematurely resuming high-definition, high-computing-power media streams.
[0083] This invention also discloses a streaming media screen optimization and scheduling system for remote desktop collaboration, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a streaming media screen optimization and scheduling method for remote desktop collaboration according to the present invention is implemented.
[0084] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
Claims
1. A method for optimizing and scheduling streaming media in remote desktop collaboration, characterized in that, include: S1: Pre-built three-level rendering assets, including a high-fidelity layer deployed in the cloud, a medium-fidelity layer deployed on multiple access edge computing nodes, and a minimalist skeleton proxy layer pre-loaded to the client; S2: The client drives the vital signs state machine to migrate between five vital signs states based on the vital signs sampling data. The five vital signs states are idle state, early warning candidate state, early warning state, candidate alarm state and emergency state in sequence. The early warning state serves as the preheating trigger condition for the skeleton channel. The client uploads the current vital signs state to the audio and video forwarding server. S3: The audio and video forwarding server periodically collects network indicators and weights them to synthesize a continuous network degradation degree, and predicts the network degradation degree at future moments based on historical slopes; S4: The audio and video forwarding server uses the physical condition and network degradation as input to determine the target circuit breaker level according to the mapping rules, and executes the level change action; when the target circuit breaker level is zero, the main screen is rendered with a high-fidelity layer; when the target circuit breaker level is one, the high-fidelity main screen is maintained and the skeleton channel is warmed up; when the target circuit breaker level is two, it switches to a medium-fidelity layer and overlays the skeleton screen; when the target circuit breaker level is three, the video stream is suspended and the local skeleton is rendered by the minimalist skeleton proxy layer.
2. The method of claim 1, wherein, The client is connected to a grip force sensor, which acts as a vital sign sensor. The client drives the vital sign state machine transition based on the vital sign sampling data collected by the grip force sensor, specifically including: The grip force sensor outputs a raw force value at a preset sampling rate. The raw force value is then low-pass filtered at a preset cutoff frequency and then averaged over a preset time span to obtain a stable grip force value. The transition of the vital sign state machine is driven by the relative relationship between the stable grip strength value and the lower warning threshold and the upper alarm threshold in the prescription threshold tensor; the lower warning threshold is taken as a preset ratio of the upper alarm threshold.
3. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The transition conditions of the vital signs state machine specifically include: When the stable grip force value falls into the warning band formed by the lower warning threshold and the upper alarm threshold, it enters the warning candidate state. If it remains in the warning band for a continuous preset number of sampling periods, it transitions to the warning state and sends a warning signal to the audio and video forwarding server, but does not trigger the circuit breaker. When the stable grip force value exceeds the alarm upper limit threshold, it enters the candidate alarm state. If it continues to exceed the limit for a preset duration, it will transition to the emergency state and set the emergency flag bit. Emergency reset requires the stable grip force value to fall below the hysteresis threshold of the alarm upper limit threshold minus the preset hysteresis amount and to remain there for a preset duration.
4. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The network metrics include round-trip time, transmission queue delay, packet loss rate, and bandwidth gap. The network degradation degree is normalized and obtained by weighting and summing the round-trip time, transmission queue delay, packet loss rate, and bandwidth gap according to their respective preset weight coefficients. The round-trip delay is taken from the current round-trip delay field in the network connection statistics interface; the sending queue delay is obtained by dividing the increment of the total data packet sending delay in the outbound real-time transmission stream statistics by the number of packets sent in the same period; the packet loss rate is taken from the packet loss rate field of the remote inbound real-time transmission stream statistics; and the bandwidth gap is obtained by the difference between the available bandwidth estimated by the congestion control algorithm and the target coding bitrate of the video stream. The audio and video forwarding server collects the aforementioned network metrics at a preset collection period.
5. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The prediction of network degradation at future moments based on historical slopes is specifically as follows: Based on the slope of the change in the network degradation sampling sequence within a preset historical time window, the predicted value of network degradation after a preset prediction time is predicted by first-order extrapolation. When the predicted network degradation value exceeds the preset degradation threshold, even if the current measured network degradation is still in the safe zone, it will be upgraded to the first level in advance and the preheating action of the skeleton channel will be triggered.
6. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The range of network degradation is divided into three levels: low, medium, and high by a preset first threshold and a preset second threshold, and the preset first threshold is less than the preset second threshold. When the vital signs are in an idle state, the target circuit breaker level remains at level zero regardless of the network degradation level. When the vital signs are in a warning state, the network degradation level remains at level zero when it is at a low level, rises to level one when it is at a medium level, and rises to level two when it is at a high level. When the vital signs state is in an emergency, the network degradation level rises from low to level one, from medium to level two, and from high to level three.
7. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, After determining the target circuit breaker level according to the mapping rules, the level change action is executed according to the asymmetric hysteresis condition of fast upgrade and slow rollback, specifically including: When the level is upgraded, the change will be executed once the triggering conditions corresponding to the target circuit breaker level are continuously met for the preset upgrade duration. When rolling back a level, the triggering conditions corresponding to the adjacent lower level must be continuously satisfied for the preset rollback retention time corresponding to the adjacent lower level before the change is executed, and any of the preset rollback retention times is greater than the preset upgrade retention time. The audio and video forwarding server performs the judgment and modification actions of the mapping rules at a preset decision cycle.
8. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The specific actions to be executed for the fourth level of circuit breaker include: The high-fidelity layer at level zero is used to transmit a high-efficiency video encoded stream via the main media stream channel; The first level maintains the high-fidelity layer as the main screen, and sends joint quaternions through the data channel at a preset first frequency. The client pre-renders the skeleton screen in the off-screen canvas but does not display it for the time being. When the main screen of the second level is switched to the mid-fidelity layer, the quaternion transmission frequency of the data channel is increased to a preset second frequency, and the skeleton is superimposed on the video with a preset transparency. The preset second frequency is higher than the preset first frequency. The third level suspends the media video stream and sends joint quaternions and angular velocities only through the data channel at the preset second frequency. The client switches to the minimalist skeleton proxy layer to render the webpage graphics library skeleton.
9. The method for optimizing and scheduling streaming media in remote desktop collaboration according to claim 1, characterized in that, The method also includes a doctor-side reverse emergency stop closed-loop step: The doctor generates an emergency stop command by clicking the emergency stop button or by pre-training a voice command. The emergency stop command is transmitted to the audio and video forwarding server through a dedicated data channel with a high priority attribute. The audio and video forwarding server bypasses the conventional processing pipeline and sends the command to the edge gateway through a dedicated user data packet protocol channel. The edge gateway then sends the command to the exoskeleton or electrical stimulation device on the patient's end via the local area network to perform unlocking, decompression, or stop training. The emergency stop command is repeatedly sent a preset number of times at a preset redundancy interval, and an independent emergency stop signaling queue, isolated from the regular uplink channel, is maintained on the audio / video forwarding server side.
10. A streaming media video optimization and scheduling system for remote desktop collaboration, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement a streaming media screen optimization scheduling method in remote desktop collaboration according to any one of claims 1-9.
Citation Information
Patent Citations
Optimization method for image processing in SPICE cloud desktop
CN113891080A