Cloud desktop audio and video synchronization method and system, electronic equipment and storage medium
By evaluating network status parameters in real time and adopting dynamic adjustment strategies, the synchronization problem of cloud desktop audio and video synchronization in a weak network environment is solved, efficient synchronous transmission is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510514271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
In weak network environments, cloud desktop audio and video synchronization problems often occur, affecting the user experience, and traditional methods are difficult to dynamically adapt to complex network changes.
By evaluating network status parameters in real time, weighted calculations and hierarchical divisions are performed, and different audio and video synchronization strategies are adopted, such as timestamp alignment, dynamic buffer compensation, redundancy rate adjustment, audio priority processing and bit rate adaptive processing, ensuring efficient synchronous transmission of audio and video when network conditions are poor.
In the case of poor network conditions, efficient synchronous transmission of cloud desktop audio and video is achieved, improving the user experience.
Smart Images

Figure CN120390112A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of data synchronous transmission, and in particular, to a method, a system, an electronic device, and a storage medium for synchronizing audio and video in a cloud desktop. Background Art
[0002] With the rapid development of cloud computing and remote desktop technologies, cloud desktop applications have been widely used in various industries. Especially in the fields of remote work, education, etc., the synchronization of audio and video has become a key factor in user experience. However, in a weak network environment (such as large network latency, insufficient bandwidth, packet loss, jitter, etc.), audio and video synchronization problems often occur, affecting the user experience. Traditional methods rely on fixed threshold adjustment or a single parameter (such as latency) for synchronization compensation, and it is difficult to dynamically adapt to complex network changes.
[0003] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] Embodiments of the present disclosure aim to provide a method, a system, an electronic device, and a storage medium for synchronizing audio and video in a cloud desktop, so as to at least overcome one or more problems caused by the limitations and defects of the related technologies to a certain extent.
[0006] According to a first aspect of the embodiments of the present disclosure, there is provided a method for synchronizing audio and video in a cloud desktop, the method including the following steps:
[0007] Real-time evaluate the network quality according to network state parameters to obtain a parameter evaluation result, where the network state parameters include: network latency, jitter, packet loss rate, and bandwidth;
[0008] Perform weighted calculation on the parameter evaluation result to obtain a network quality score, and classify the network quality according to the network quality score, where the network quality levels include: good network, mild weak network, and severe weak network;
[0009] When the network quality level is a good network, align the audio and video timestamps and do not need to perform buffer compensation on the audio and video;
[0010] When the network quality level is a mild weak network, perform dynamic buffer compensation, redundancy rate adjustment, and audio priority processing on the audio and video;
[0011] When the network quality level is severely weak network, perform prediction compensation, selective frame dropping, and bitrate adaptation processing on the audio and video timestamps.
[0012] In an exemplary embodiment of the present disclosure, the parameter evaluation result is represented by a fractional value; when performing weighted calculation on the parameter evaluation result, the sum of the weight coefficients of the network state parameters is 1.
[0013] In an exemplary embodiment of the present disclosure, the step of aligning the audio and video timestamps includes:
[0014] Synchronize the audio and video timestamps based on the NTP protocol.
[0015] In an exemplary embodiment of the present disclosure, the step of performing dynamic buffer compensation on the audio and video includes:
[0016] Calculate the audio and video synchronization error according to the buffer capacity and playback rate of the audio and video, and perform dynamic buffer compensation on the audio and video according to the synchronization error. The compensation method includes: increasing or decreasing the buffer capacity of the audio and video.
[0017] In an exemplary embodiment of the present disclosure, the step of adjusting the redundancy rate includes:
[0018] Add redundant packets to increase the redundancy rate so that the number of redundant packets is not less than the expected number of lost packets.
[0019] In an exemplary embodiment of the present disclosure, the step of performing prediction compensation on the audio and video timestamps includes:
[0020] Predict the delay trend based on Kalman filtering and dynamically correct the audio and video playback timestamps.
[0021] According to a second aspect of the embodiments of the present disclosure, there is provided a cloud desktop audio and video synchronization system, the system includes:
[0022] A network quality evaluation module, configured to perform real-time evaluation on the network quality according to network state parameters to obtain a parameter evaluation result, where the network state parameters include: network delay, jitter, packet loss rate, and bandwidth;
[0023] A network quality scoring module, configured to perform weighted calculation on the parameter evaluation result to obtain a network quality score, and perform grade division on the network quality according to the network quality score, where the network quality levels include: good network, mildly weak network, and severely weak network;
[0024] An audio - video synchronization control module is used to align the audio - video timestamps when the network quality level is a good network, and there is no need to perform buffer compensation on the audio - video; when the network quality level is a slightly weak network, perform dynamic buffer compensation, redundancy rate adjustment, and audio priority processing on the audio - video; when the network quality level is a severely weak network, perform prediction compensation on the audio - video timestamps, selective frame dropping, and bitrate adaptation processing.
[0025] In an exemplary embodiment of the present disclosure, the audio - video synchronization control module includes:
[0026] An audio - video compression unit is used to perform bitrate adaptation processing on the video;
[0027] A quality adjustment unit is used to perform redundancy rate adjustment, selective frame dropping processing, and audio priority processing on the video;
[0028] A buffer management unit is used to perform dynamic buffer compensation on the audio - video;
[0029] An audio - video synchronization unit is used to align the audio - video timestamps and perform prediction compensation on the audio - video timestamps.
[0030] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including:
[0031] A processor; and
[0032] A memory for storing executable instructions of the processor;
[0033] Wherein, the processor is configured to execute the steps of the cloud desktop audio - video synchronization method described in any one of the above - mentioned embodiments by executing the executable instructions.
[0034] According to a fourth aspect of the embodiments of the present disclosure, a computer - readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the cloud desktop audio - video synchronization method described in any one of the above embodiments are implemented.
[0035] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0036] In the embodiments of the present disclosure, according to multiple network state parameters such as network latency, jitter, packet loss rate, and bandwidth, the network quality is evaluated in real - time, and then the quality and transmission strategy of audio - video data are automatically adjusted according to the level of network quality, ensuring that even under poor network conditions, the cloud desktop audio - video can still be transmitted efficiently and synchronously, improving the user experience.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0039] Figure 1 A flowchart showing the method for synchronizing audio and video of a cloud desktop in an exemplary embodiment of the present disclosure;
[0040] Figure 2 A module structure diagram showing the cloud desktop audio and video synchronization system in an exemplary embodiment of the present disclosure;
[0041] Figure 3 A schematic structural diagram showing an electronic device in an exemplary embodiment of the present disclosure;
[0042] Figure 4 A schematic structural diagram showing a program product for implementing the method for synchronizing audio and video of a cloud desktop in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0044] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0045] In this example embodiment, a method for synchronizing audio and video of a cloud desktop is first provided. Referring to Figure 1 as shown, the method may include the following steps:
[0046] Step S101, perform real-time evaluation of network quality according to network status parameters to obtain a parameter evaluation result, where the network status parameters include: network latency, jitter, packet loss rate, and bandwidth;
[0047] Step S102: Perform weighted calculation on the parameter evaluation result to obtain a network quality score, and classify the network quality according to the network quality score. The network quality levels include: good network, slightly weak network, and severely weak network.
[0048] Step S103: When the network quality level is a good network, align the audio and video timestamps and do not need to perform buffer compensation on the audio and video.
[0049] Step S104: When the network quality level is a slightly weak network, perform dynamic buffer compensation, redundancy rate adjustment, and audio priority processing on the audio and video.
[0050] Step S105: When the network quality level is a severely weak network, perform predictive compensation on the audio and video timestamps, selective frame dropping, and bitrate adaptation processing.
[0051] In this embodiment, according to multiple network state parameters such as network latency, jitter, packet loss rate, and bandwidth, the network quality is evaluated in real time, and then the quality and transmission strategy of the audio and video data are automatically adjusted according to the network quality level, ensuring that even under poor network conditions, the cloud desktop audio and video can still be transmitted efficiently and synchronously, improving the user experience.
[0052] Next, each step of the above method in this exemplary embodiment will be described in more detail.
[0053] In step S101, according to the actual situation of the network, network state parameters such as latency (RTT), jitter (Jitter), packet loss rate (Loss), and bandwidth (Bandwidth) are selected to evaluate the network quality in real time, and provide a basis for subsequent audio and video synchronization adjustment strategies. The detection content of each network state parameter is as follows:
[0054] Bandwidth detection: Regularly detect the upload and download bandwidth of the network.
[0055] Latency detection: Calculate the network latency by measuring the round-trip time of data packets.
[0056] Packet loss rate detection: Calculate the proportion of lost data packets.
[0057] Jitter: Evaluate the network volatility by analyzing the real-time changes of latency and bandwidth.
[0058] The parameter evaluation results obtained for the above network status parameters are represented by fractional values. For example, the fractional values of bandwidth, latency, packet loss rate, and jitter are all in the range of 0 to 100 points, and their fractional values are respectively represented by Bandwidth_Score, RTT_Score, Loss_Score, and Jitter_Score.
[0059] Next, in step S102, a weighted calculation is performed on the parameter evaluation results obtained for the above network status parameters. The weighting coefficients are represented by α, β, γ, and δ respectively. The calculated network quality score C_score is as follows:
[0060] C_score = α·RTT_Score + β·Jitter_Score + γ·Loss_Score + δ·Bandwidth_Score
[0061] Among them, the weighting coefficients α, β, γ, and δ are dynamically optimized and adjusted according to engineering data, and are used to represent the influence degree of the corresponding network status parameters. The sum of the weighting coefficients α, β, γ, and δ is 1.
[0062] The network quality is classified according to the network quality score. Among them, the network quality levels include: good network, slightly weak network, and severely weak network, which are specifically as follows:
[0063] Good network (C_score ≥ 80): high bandwidth, low latency, and low packet loss;
[0064] Slightly weak network (60 ≤ C_score < 80): medium bandwidth, moderate latency, and tolerable jitter;
[0065] Severely weak network (C_score < 60): low bandwidth, high packet loss, and high latency.
[0066] According to the network level classification, steps S103 to S105 are taken to dynamically adjust the network status to enable synchronous audio and video transmission, which is specifically as follows:
[0067] (1) Good network
[0068] Standard timestamp alignment: Synchronize the audio and video timestamps based on the NTP protocol; there is no need to buffer the audio and video, and the real-time experience is pursued. The network status at this time hardly affects the user experience.
[0069] Synchronizing audio and video timestamps based on the NTP protocol achieves synchronized playback of audio and video data by aligning the time base of the audio and video streams with a precise Network Time Protocol (NTP) time source. The NTP protocol provides microsecond-level time synchronization using layered time sources (such as atomic clocks and GPS), effectively addressing local device clock drift and preventing audio and video tearing caused by clock desynchronization.
[0070] In scenarios with unstable network jitter or latency, NTP timestamps provide a unified reference standard for playback, enabling dynamic adjustment of buffering strategies (such as adaptive frame loss or interpolation compensation) to reduce lag. In distributed systems or multi-terminal playback scenarios (such as live streaming and video conferencing), NTP synchronization ensures that different devices process data based on the same timeline, enabling synchronized playback on multiple screens.
[0071] Furthermore, combined with RTCP's feedback mechanism, NTP timestamps can be used as a fault-tolerance mechanism. For example, when network packet loss causes timestamp loss, an interpolation algorithm can be used to estimate the time position of the lost frame and maintain synchronization continuity.
[0072] (2) Mildly weak network
[0073] ① Dynamic buffer compensation: Dynamically adjust the buffer size (i.e., buffer capacity) based on the obtained jitter value.
[0074] Define the buffer model: Assume the audio buffer size is Ba, the video buffer size is Bv, the audio playback rate is Ra, and the video playback rate is Rv, then the synchronization error ΔT is:
[0075] ΔT=Ba / Ra–Bv / Rv
[0076] The buffer size is adjusted, that is, the goal is to achieve ΔT=0, which requires adjusting the audio buffer size Ba and the video buffer size Bv.
[0077] Specifically, when ΔT>0, the video buffer size Bv is increased or the audio buffer size Ba is decreased. If ΔT<0, the audio buffer size Ba is increased or the video buffer size Bv is decreased.
[0078] ②Forward Error Correction (FEC): Adds redundant packets to resist packet loss.
[0079] Let the original data packet be k, the redundant data packet be nk, and the total data packet be n.
[0080] FEC coding model: uses Reed-Solomon coding, and the decoding conditions are:
[0081] e≤nk
[0082] Where e is the number of lost packets.
[0083] The redundancy rate D is: D = (n - k) / k
[0084] According to the network packet loss rate p, select D and satisfy: D ≥ p / (1 - p).
[0085] Methods such as forward error correction coding or redundant frame insertion can be used to adjust the redundancy rate, thereby adjusting the video compression efficiency and fault tolerance. After increasing the redundancy rate, the packet loss resistance of the video during transmission can be improved. At the same time, some bitrate resources will be sacrificed, which may lead to a decrease in picture quality.
[0086] ③ Audio priority processing
[0087] During the transmission of audio and video, video is more likely to freeze. At this time, methods such as temporarily reducing the video clarity are needed to allow the video stream to temporarily reduce the quality and ensure the continuity of the audio stream first.
[0088] In addition, according to the comparison between the received audio and video timestamps and the local NTP time, the buffer is dynamically adjusted: if the audio timestamp is ahead, the audio playback speed is reduced or silence is inserted; if the video timestamp is behind, frame dropping or accelerated rendering can be adopted, etc.
[0089] (3) Severe weak network
[0090] ① Timestamp prediction compensation: Based on Kalman filtering to predict the delay trend and dynamically correct the playback timestamp. The Kalman filter can be used to estimate the state of the audio and video synchronization error and correct these errors according to real-time feedback. The Kalman filter can estimate the current delay trend based on the previous state, then update the prediction error covariance matrix, calculate the Kalman gain according to the actually measured network delay, and finally compensate the timestamp.
[0091] Timestamp compensation strategy:
[0092] Positive delay: If the predicted delay increases, decode future frames in advance and extend the buffer time;
[0093] Negative delay: If the predicted delay decreases, accelerate playback or discard redundant frames;
[0094] Dynamic interpolation: Smooth the timestamp jump through audio resampling or video frame interpolation.
[0095] ② Selective frame dropping: Discard non-key video frames (such as B frames or P frames) to ensure the smoothness of the audio first.
[0096] When it is detected that the packet loss rate of the network status is high and the bandwidth is low, resulting in the inability to synchronously transmit audio and video, the method of discarding non-critical video frames is started to give priority to ensuring smooth audio. It is possible to trigger the discard of all B frames in the current buffer once, or it is also possible to continuously discard all P frames and B frames between two key frames, only retaining the first and last I frames. After discarding the frames, the PTS (Presentation Time Stamp) of the subsequent frames is advanced to avoid audio-visual asynchronization caused by frame loss.
[0097] By adopting selective frame dropping processing, the bandwidth occupancy can be reduced. At the same time, methods such as repeated frames or motion compensation interpolation can be used to alleviate the short-term freezing or blurring of the picture; by dynamically shortening the key frame interval, continuous audio playback can be ensured; dynamic bitrate adjustment can also be combined to prevent the visual jumping feeling caused by frequent frame dropping, so as to achieve the purpose of reducing the end-to-end delay.
[0098] ③ Bitrate adaptation: Reduce the video resolution and frame rate (such as 1080p → 720p, etc.).
[0099] Dynamically adjust the encoding parameters according to the network conditions to reduce the audio-visual synchronization error.
[0100] Frame rate adjustment: Let the target frame rate be f target , the network bandwidth be r(t), then: f target = r(t) / R per_frame, where, R per_frame is the average bitrate per frame.
[0101] Resolution adjustment: Let the target resolution be S target , then: S target = r(t) / R per_pixel , where R per_pixel is the average bitrate per pixel.
[0102] In the implementation manner of the present application, secondly, a cloud desktop audio-visual synchronization system is provided. Please refer to Figure 2 , the system includes:
[0103] A network quality evaluation module, which is used to evaluate the network quality in real time according to network status parameters to obtain a parameter evaluation result. The network status parameters include: network delay, jitter, packet loss rate, and bandwidth;
[0104] A network quality scoring module, which is used to perform weighted calculation on the parameter evaluation result to obtain a network quality score, and classify the network quality according to the network quality score. Among them, the network quality levels include: good network, mild weak network, and severe weak network;
[0105] The audio - video synchronization control module is used to align the audio - video timestamps when the network quality level is a good network, without the need for buffering compensation for audio - video; when the network quality level is a slightly weak network, perform dynamic buffering compensation, redundancy rate adjustment, and audio priority processing on the audio - video; when the network quality level is a severely weak network, perform predictive compensation on the audio - video timestamps, selective frame dropping, and bitrate adaptation processing.
[0106] In this embodiment, based on multiple network state parameters such as network latency, jitter, packet loss rate, and bandwidth, the network quality is evaluated in real - time, and then the quality and transmission strategy of audio - video data are automatically adjusted according to the network quality level, ensuring that even under poor network conditions, the cloud desktop audio - video can still be efficiently and synchronously transmitted, improving the user experience.
[0107] Among them, the audio - video synchronization control module includes:
[0108] The audio - video compression unit is used to perform bitrate adaptation processing on the video;
[0109] The quality adjustment unit is used to perform redundancy rate adjustment, selective frame dropping processing, and audio priority processing on the video;
[0110] The buffer management unit is used to perform dynamic buffering compensation on the audio - video;
[0111] The audio - video synchronization unit is used to align the audio - video timestamps and perform predictive compensation on the audio - video timestamps.
[0112] By using the above - mentioned units to simultaneously perform multiple network adjustment processing measures, the purpose of audio - video synchronization is achieved.
[0113] The following describes the coordination and cooperation method among the above - mentioned multiple processing strategies.
[0114] Taking the severely weak network as an example: When the severely weak network is activated, perform timestamp prediction compensation (always exists) processing for a duration of δ, and the duration can be set as needed. If it is detected that the network state has not improved, then according to the detected network situation, perform selective frame dropping or bitrate adaptation processing. Among them, bitrate adaptation processing is given priority, and selective frame dropping processing is performed when the network condition has not improved.
[0115] If the network condition of the severely weak network improves, for example, it changes to a slightly weak network, then perform at least one of dynamic buffering compensation, forward error correction, and audio priority processing to adjust the network condition so as to improve the data transmission effect.
[0116] Regarding the system in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0117] It should be noted that although several modules of the system for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described modules can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.
[0118] See Figure 3 , an embodiment of the present invention further provides an electronic device 300, which includes at least one memory 310, at least one processor 320, and a bus 330 connecting different platform systems.
[0119] The memory 310 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 311 and / or a cache memory 312, and may further include a read-only memory (ROM) 313.
[0120] Among them, the memory 310 also stores a computer program, which can be executed by the processor 320, so that the processor 320 executes the steps of the cloud desktop audio-video synchronization method in any embodiment of the present invention. The specific implementation manner is consistent with the implementation manner and the achieved technical effects recorded in the embodiments of the above cloud desktop audio-video synchronization method, and some content will not be elaborated.
[0121] The memory 310 may further include a utility 314 having at least one program module 315. Such program modules 315 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0122] Correspondingly, the processor 320 can execute the above computer program and can also execute the utility 314.
[0123] The bus 330 may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures.
[0124] The electronic device 300 can also communicate with one or more external devices 340 such as a keyboard, a pointing device, a Bluetooth device, etc., and can also communicate with one or more devices capable of interacting with the electronic device 300, and / or communicate with any device (such as a router, a modem, etc.) that enables the electronic device 300 to communicate with one or more other computing devices. Such communication can be carried out through the input / output interface 350. Moreover, the electronic device 300 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 360. The network adapter 360 can communicate with other modules of the electronic device 300 through the bus 330. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 300, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0125] An embodiment of the present invention also provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed, the steps of the cloud desktop audio-video synchronization method in the embodiment of the present invention are realized. Its specific implementation manner is consistent with the implementation manner and the achieved technical effects recorded in the embodiment of the above cloud desktop audio-video synchronization method, and some contents will not be elaborated herein.
[0126] Figure 4 A program product 400 for implementing the above cloud desktop audio-video synchronization method provided in this embodiment is shown. It can adopt a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product 400 of the present invention is not limited thereto. In the present invention, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program product 400 can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0127] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0128] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these fall within the protection scope of the present invention.
Claims
1. A method for synchronizing audio and video of a cloud desktop, characterized in that, The method includes the following steps: Perform real-time evaluation of network quality based on network status parameters to obtain parameter evaluation results. The network status parameters include network latency, jitter, packet loss rate, and bandwidth; Perform weighted calculation on the parameter evaluation results to obtain a network quality score, and classify the network quality according to the network quality score. Among them, the network quality levels include: good network, mild weak network, and severe weak network; When the network quality level is a good network, align the audio and video timestamps, and there is no need to perform buffer compensation on the audio and video; When the network quality level is a mild weak network, perform dynamic buffer compensation, redundancy rate adjustment, and audio priority processing on the audio and video; When the network quality level is a severe weak network, perform prediction compensation on the audio and video timestamps, selective frame dropping, and bitrate adaptation processing; 2. The cloud desktop audio and video synchronization method according to claim 1, wherein The parameter evaluation results are represented by score values; when performing weighted calculation on the parameter evaluation results, the sum of the weight coefficients of the network status parameters is 1.
3. The cloud desktop audio-video synchronization method according to claim 1, wherein The step of aligning the audio and video timestamps includes: Synchronize the audio and video timestamps based on the NTP protocol.
4. The cloud desktop audio-video synchronization method according to claim 1, wherein The step of performing dynamic buffer compensation on the audio and video includes: Calculate the audio and video synchronization error according to the buffer capacity and playback rate of the audio and video, and perform dynamic buffer compensation on the audio and video according to the synchronization error. The compensation method includes increasing or decreasing the buffer capacity of the audio and video.
5. The cloud desktop audio and video synchronization method according to claim 1, wherein The step of redundancy rate adjustment includes: Add redundant packets to increase the redundancy rate so that the number of redundant packets is not less than the expected number of lost packets.
6. The cloud desktop audio and video synchronization method according to claim 1, wherein The step of performing prediction compensation on the audio and video timestamps includes: Predict the delay trend based on Kalman filtering and dynamically correct the audio and video playback timestamps.
7. Cloud desktop audio-video synchronization system, characterized in that, The system includes: A network quality evaluation module for performing real-time evaluation of network quality based on network status parameters to obtain parameter evaluation results. The network status parameters include network latency, jitter, packet loss rate, and bandwidth; A network quality scoring module for performing weighted calculation on the parameter evaluation results to obtain a network quality score, and classifying the network quality according to the network quality score. Among them, the network quality levels include: good network, mild weak network, and severe weak network; An audio and video synchronization control module for aligning the audio and video timestamps when the network quality level is a good network, and there is no need to perform buffer compensation on the audio and video; when the network quality level is a mild weak network, perform dynamic buffer compensation, redundancy rate adjustment, and audio priority processing on the audio and video; when the network quality level is a severe weak network, perform prediction compensation on the audio and video timestamps, selective frame dropping, and bitrate adaptation processing.
8. The cloud desktop audio-video synchronization system according to claim 7, wherein, The audio and video synchronization control module includes: An audio and video compression unit for performing bitrate adaptation processing on the video; A quality adjustment unit for performing redundancy rate adjustment, selective frame dropping processing, and audio priority processing on the video; A buffer management unit for performing dynamic buffer compensation on the audio and video; An audio and video synchronization unit for aligning the audio and video timestamps and performing prediction compensation on the audio and video timestamps.
9. An electronic device, characterized in that, Includes: A processor; And A memory for storing executable instructions of the processor; Among them, the processor is configured to execute the steps of the cloud desktop audio-video synchronization method according to any one of claims 1 to 6 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the cloud desktop audio-video synchronization method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Remote video missing processing method and related device
CN121567923A
Video stable transmission method and system in high-speed moving scene and electronic equipment
CN121604025A