Audio transmission quality determination method, electronic device, storage medium, and product

By monitoring parameters such as delay, restore degree and lag rate of cloud desktop audio transmission quality, the problem of difficulty in comprehensively monitoring audio transmission quality in the existing technology is solved, and the user experience of cloud desktop audio transmission is improved.

WO2025177083A1PCT designated stage Publication Date: 2025-08-28CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050756
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-01-24
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The existing technology is difficult to fully monitor the quality of cloud desktop audio transmission, making it difficult for operation and maintenance personnel to efficiently adjust their strategies to improve user experience.

Method used

Provides a method for determining audio transmission quality, by obtaining transmission information and playback information between the source and destination, monitoring parameters such as audio transmission delay, audio restoration degree and audio lag rate, and drawing time-related line charts to help operation and maintenance personnel discover and solve problems that affect audio transmission quality in a timely manner.

Benefits of technology

It realizes comprehensive monitoring of the audio transmission quality of cloud desktops, helping operation and maintenance personnel to quickly discover and formulate adjustment strategies and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050756_28082025_PF_FP_ABST
    Figure IB2025050756_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of Internet, and provides an audio transmission quality determination method, an electronic device, a computer readable storage medium, and a computer program product. The audio transmission quality determination method comprises: acquiring transmission information of a target audio between a source end and a destination end, and playback information of the target audio on the destination end, wherein one of the source end and the destination end is a cloud desktop server, and the other is a cloud desktop client; and determining a transmission quality parameter of the target audio on the basis of the transmission information and / or the playback information, wherein the transmission quality parameter includes at least one of an audio transmission delay parameter, audio fidelity, and an audio stuttering rate; the audio transmission delay parameter is used for characterizing the delay between the transmission of the target audio from the source end and the playback thereof on the destination end; the audio fidelity is used for characterizing the difference between the target audio played on the destination end and the target audio transmitted from the source end; and the audio stuttering rate is used for characterizing the playback stuttering of the target audio on the destination end.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure claims priority to a Chinese patent application entitled "Method for Determining Audio Transmission Quality, Electronic Device, Storage Medium, and Product," filed with the China Patent Office on February 20, 2024, under application number 202410194265. The entire contents of this application are incorporated herein by reference. Technical Field: This disclosure relates to the field of Internet technology, and more particularly to a method for determining audio transmission quality, an electronic device, a computer-readable storage medium, and a computer program product. Background: Cloud desktops are a new type of service where computing power is deployed on cloud servers and interactions are performed locally. They offer advantages such as high availability, high security, easy specification expansion, and multi-terminal connectivity. Cloud desktop audio streams are encoded in the cloud and then transmitted to the local server via a transmission protocol for decoding and playback. Because different users access cloud desktops through different terminals and in diverse usage environments, the performance of audio uplinks and downlinks often varies. Reflecting the quality of a user's cloud desktop audio uplinks and downlinks through data is crucial for improving the operation, maintenance, and customization capabilities of cloud desktop service providers. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a method for determining audio transmission quality, an electronic device, a computer-readable storage medium, and a computer program product to alleviate or resolve one or more technical problems existing in the prior art. In a first aspect, embodiments of the present disclosure provide a method for determining audio transmission quality, comprising: obtaining transmission information of a target audio signal between a source and a destination, and obtaining playback information of the target audio signal on the destination, wherein one of the source and destination is a cloud desktop server and the other is a cloud desktop client; determining transmission quality parameters of the target audio signal based on the transmission information and / or playback information, wherein the transmission quality parameters include at least one of an audio transmission delay parameter, audio fidelity, and audio stuttering rate; the audio transmission delay parameter represents the delay between the target audio signal being transmitted from the source and being played by the destination; the audio fidelity represents the difference between the target audio signal played by the destination and the target audio signal sent by the source; and the audio stuttering rate represents the stuttering of the target audio signal played by the destination. In a second aspect, embodiments of the present disclosure provide an electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements any of the methods of the embodiments of the present disclosure. In a third aspect, embodiments of the present disclosure provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements any of the methods described in the embodiments of the present disclosure. In a fourth aspect, embodiments of the present disclosure provide a computer program product including the computer program. When the computer program is executed by a processor, the computer program implements any of the methods described in the embodiments of the present disclosure.Based on the method for determining audio transmission quality according to the embodiments of the present disclosure, by monitoring and integrating multiple relevant data points of uplink and downlink audio data in a cloud desktop scenario as audio transmission quality parameters, various factors affecting audio transmission quality can be more comprehensively monitored. This allows operations and maintenance personnel to promptly identify issues such as low audio transmission quality caused by network jitter, high network latency, and improper client audio buffer settings, and develop adjustment strategies to mitigate or eliminate these issues, thereby providing cloud desktop users with a better experience. The above description is only an overview of the technical solution of the present disclosure. To better understand the technical means of the present disclosure, implementation should be carried out in accordance with the contents of the specification. To further enhance the understanding of the above and other objectives, features, and advantages of the present disclosure, specific embodiments of the present disclosure are described below. BRIEF DESCRIPTION OF THE DRAWINGS In the accompanying drawings, unless otherwise specified, identical reference numerals throughout the multiple drawings indicate identical or similar components or elements. The drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments of the present disclosure and should not be construed as limiting the scope of the present disclosure. Figure 1 is a schematic diagram of an application scenario of an embodiment of the present disclosure; Figure 2 is a schematic diagram of a cloud desktop monitoring log system for determining audio transmission quality according to an embodiment of the present disclosure; Figure 3 is a flow chart of a method for determining audio transmission quality according to an embodiment of the present disclosure; Figure 4 is a schematic diagram of an apparatus for determining audio transmission quality according to an embodiment of the present disclosure; and Figure 5 is a block diagram of an electronic device used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS The following briefly describes certain exemplary embodiments. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the spirit or scope of the present disclosure. Therefore, the drawings and description are to be regarded as illustrative rather than restrictive in nature. To facilitate understanding of the technical solutions of the embodiments of the present disclosure, the following describes related technologies. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any manner and fall within the scope of protection of the embodiments of the present disclosure. The application scenario of cloud desktops is based on virtualization technology that virtualizes computer hardware resources into multiple virtual computers, on which unmodified desktop operating systems can be directly run. Cloud desktops can transform the traditional software usage mode of "local installation and local computing" into a "ready-to-use" service. Clients do not need to install the corresponding software and can connect to and control remote service clusters through the Internet or local area network to complete business logic or computing tasks, such as communication applications, online shopping applications, and system service applications on the cloud desktop.For audio applications running on a cloud desktop (e.g., radio or music playback applications, and audio and video call applications), the cloud desktop server can store audio data in the cloud and transmit it to the client upon request. In audio and video call scenarios (e.g., remote conferencing applications or voice messages in communication applications), the client collects the user's audio data and transmits it to the cloud desktop server, which then stores the audio data (e.g., voice conference minutes) and forwards it to other clients as needed. Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure, illustrating the audio transmission link between the cloud desktop server and the cloud desktop client. As shown in Figure 1, for the audio uplink, the cloud desktop client encodes local audio data, such as microphone data collected or audio data transmitted from other local devices to the client's device, and then uploads the processed audio data to the cloud desktop service through the transmission protocol. For the audio downlink, the cloud desktop server encodes the audio data stored in the cloud and sends the processed audio data to the client through the transmission protocol so that the user can play it locally. After receiving the audio data, the client or server needs to use an audio buffer to eliminate audio instability caused by network jitter or unstable audio processing threads. It will also perform various processing such as resampling and noise reduction on the audio as needed. Understandably, for cloud desktop service providers, ensuring audio transmission quality in cloud desktop scenarios is crucial to enhancing the competitiveness of cloud desktop products. Factors affecting audio transmission quality in cloud desktop scenarios include network conditions, client-side decoding performance, and client-side audio playback jitter. Traditional cloud desktop audio transmission quality monitoring often only focuses on factors such as network latency and audio data volume, which cannot fully reflect the client experience. This makes it difficult for operations and maintenance personnel to efficiently identify transmission issues and adjust relevant configurations and policies based on existing monitoring data. Therefore, embodiments of the present disclosure provide a method for determining audio transmission quality that can monitor and integrate multiple relevant data points from uplink and downlink audio data in cloud desktop scenarios as audio transmission quality parameters. This allows operations and maintenance personnel to quickly and accurately identify problems arising from cloud desktop audio transmission based on these audio transmission quality parameters and develop optimization strategies to mitigate or resolve them. Figure 2 shows a schematic diagram of a cloud desktop monitoring log system for determining audio transmission quality according to an embodiment of the present disclosure. As shown in Figure 2, the audio transmission quality parameters monitored by the cloud desktop monitoring log include audio transmission delay, audio fidelity, and audio stuttering rate.Audio transmission delay can include audio transmission link delay. Audio transmission link delay is specifically the time required for an audio frame to be sent from the source end to the destination end. Its main influencing factor is network delay. Therefore, network delay can also be used to represent audio transmission link delay. Audio transmission delay can also include total audio transmission delay. Total audio transmission delay is specifically the time required for an audio frame to be sent from the source end to be played by the destination end. It specifically includes audio transmission link delay, audio processing delay, and audio buffer delay. The audio processing delay refers to the time required for the destination end to perform resampling, noise reduction, and other processing on the received audio / audio frame. The audio buffer delay refers to the time required for the audio buffer set by the destination end to buffer the audio / audio frame. For the audio uplink, the source end is the cloud desktop client and the destination end is the cloud desktop server. For the audio downlink, the source end is the cloud desktop server and the destination end is the cloud desktop client. The audio stutter rate indicates the presence of stuttering in the audio / audio frames received by the destination end. Audio stuttering refers to interruptions or discontinuities during audio playback. The audio stutter rate can be determined by counting the number of stutters experienced per unit time. Audio fidelity refers to the difference between the audio received and played by the destination end and the original audio sent by the source end. A smaller difference indicates higher audio fidelity, meaning less loss of audio quality during transmission and processing. Audio fidelity can be determined through audio evaluation, which includes both subjective and objective audio evaluation. Subjective audio evaluation involves inviting a sufficient number of people to listen to the original audio and the processed audio played by the destination end and rate the audio quality. Objective audio evaluation can be conducted with or without a reference. With a reference, the processed audio played by the destination end is evaluated based on an undamaged reference signal (the original audio), while without a reference, the quality is assessed by directly analyzing the processed audio played by the destination end. For the three audio transmission quality parameters mentioned above, the cloud desktop monitoring log system can obtain these transmission quality parameters at a fixed period during the use of the cloud desktop and display the changes in the audio transmission quality parameters to operation and maintenance personnel by drawing a time-related line chart. This allows operation and maintenance personnel to intuitively understand the status of the cloud desktop audio uplink and downlink, and promptly identify and resolve issues that affect the cloud desktop user experience.Based on this, the cloud desktop monitoring log system, employing the audio transmission quality determination method according to the embodiments of the present disclosure, can more comprehensively monitor various factors affecting audio transmission quality. This allows operations and maintenance personnel to promptly identify issues such as low audio transmission quality caused by network jitter, high network latency, and improper client audio buffer settings, and develop adjustment strategies to mitigate or eliminate these issues, providing cloud desktop users with a better experience. It should be noted that the aforementioned application scenarios or examples of the audio transmission quality determination method provided in the embodiments of the present disclosure are provided for ease of understanding, and the embodiments of the present disclosure do not specifically limit the application of the audio transmission quality determination method. The following detailed description of the technical solution of the present disclosure and how it solves the aforementioned technical problems will be provided using specific embodiments. The several specific embodiments listed above may be combined with each other, and identical or similar concepts or processes may not be described in detail in certain embodiments. Figure 3 shows a flow chart of the audio transmission quality determination method according to an embodiment of the present disclosure. This audio transmission quality determination method can be applied to an audio transmission quality determination device. As shown in FIG3 , the method for determining audio transmission quality may include the following steps: Step S301: Acquiring transmission information of target audio between a source and a destination, and playback information of the target audio on the destination. One of the source and destination is a cloud desktop server, and the other is a cloud desktop client. In the disclosed embodiment, a cloud desktop client refers to an electronic device used by a user with communication functions, such as a mobile phone, tablet computer, personal computer, wearable device, or Internet of Things (IoT) device. A cloud desktop server refers to a computer device that can provide cloud-side application services and generally has the ability to undertake and guarantee services. The server can respond to service requests from the client and provide users with services related to cloud-side applications. The server can be a single server device, a cloud-based server array, or a virtual machine (VM) running in the cloud-based server array. Furthermore, the server can also refer to other computing devices with corresponding service capabilities, such as terminal devices such as computers (running service programs). Among them, for the audio uplink, the source end is the cloud desktop client and the destination end is the cloud desktop server. For example, the client uploads local audio to the server for cloud storage; and for the audio downlink, the source end is the cloud desktop server and the destination end is the cloud desktop client. For example, the cloud desktop server responds to the request in the client and transmits the audio data stored in the cloud to the client for playback.Step S302: Determine transmission quality parameters of the target audio based on the transmission information and / or playback information, where the transmission quality parameters include at least one of an audio transmission delay parameter, audio fidelity, and an audio stutter rate. The audio transmission delay parameter represents the delay between the target audio being sent from the source and being played by the destination. The audio fidelity represents the difference between the target audio played by the destination and the target audio sent by the source. The audio stutter rate represents the amount of playback stuttering of the target audio by the destination. Exemplarily, in step S301, the transmission information may include the time when the source end sends the target audio and the time when the destination end receives the target audio, and the playback information may include the time when the destination end plays the target audio. Therefore, in step S302, the transmission quality parameters of the target audio are determined based on the transmission information and / or the playback information, including: determining the audio transmission delay parameters of the target audio based on the transmission information and / or the playback information; wherein the audio transmission delay parameters include the total audio transmission delay and / or the audio transmission link delay, the total audio transmission delay is the time difference between the source end sending the audio frame in the target audio and the destination end playing the audio frame, and the audio transmission link delay is the time difference between the source end sending the audio frame and the destination end receiving the audio frame. It's understandable that audio transmission link latency, or the time it takes for an audio frame in the target audio to be sent from the source to the destination, is primarily influenced by network latency. After receiving audio data, both the client and server need to use an audio buffer to mitigate audio instability caused by network jitter or unstable audio processing threads. They also perform various other processing operations, such as resampling and noise reduction, on the audio as needed. Audio transmission link latency only reflects network transmission conditions and doesn't reflect how the destination buffers and processes the target audio. Total audio transmission latency can be calculated by calculating the time difference between the source sending the target audio frame and the destination playing it. It essentially represents the sum of audio transmission link latency, audio processing latency, and audio buffer latency. It more intuitively demonstrates the time it takes for a user to request audio data from the cloud desktop server through a client and for the local client to be able to play the requested audio data. Therefore, monitoring audio transmission link latency can reflect network conditions during audio data transmission in cloud desktop scenarios, while monitoring total audio transmission latency allows users to more closely perceive the efficiency of cloud desktop responses to user requests. In one embodiment, in step S302, determining the transmission quality parameter of the target audio according to the transmission information and / or the playback information includes: determining the audio freeze rate of the target audio according to the playback information; wherein the playback information includes the number of freezes in the playback of the target audio per unit time.To determine stuttering, a time interval between the rendering of two audio frames can be defined as greater than a fixed time (e.g., 100ms). The audio stutter rate can then be calculated by counting the number of stutters per unit time. Preferably, audio stuttering can also be defined as occurring when the remaining audio frames in the audio buffer (hereinafter referred to as the remaining audio buffer time) are less than the expected audio frame arrival interval. The expected audio frame arrival interval refers to the time difference between the arrival time of the next audio frame in the audio buffer and the arrival time of the previous audio frame. When the remaining audio buffer time is less than the expected audio frame arrival interval, it indicates that the audio frames in the audio buffer will be exhausted before the arrival of the next audio frame, resulting in stuttering. The stuttering time is the difference between the expected audio frame arrival interval and the remaining audio buffer time. The audio stutter rate can be determined by calculating the ratio of the stuttering time per unit time. Based on this, the audio stutter rate of the target audio played by the destination end is monitored. While ensuring that the user experiences the requested audio content in a timely manner, the impact of the audio data transmission process on the playback quality of the audio content is monitored. In addition to the efficiency of the cloud desktop's response to user requests, the cloud desktop's response quality to user requests can also be monitored. It is understandable that the user experience of cloud desktop audio applications in the cloud desktop not only includes audio response speed and audio playback stutter rate, but some users often also have requirements for audio quality. In one embodiment, the playback information may also include the audio signal of the target audio played by the destination end. In step S302, the audio restoration degree of the target audio may be determined by performing an audio evaluation on the audio signal. Audio evaluation includes subjective audio evaluation and objective audio evaluation, and objective audio evaluation may include both reference evaluation and non-reference evaluation. Since subjective audio evaluation requires inviting a sufficient number of people to listen to both the original audio and the audio played after processing on the destination end, and reference evaluation requires the original audio without any lossy processing, in cloud desktop applications, all audio on the cloud desktop must undergo lossy processing such as audio encoding. Non-reference evaluation, on the other hand, can be performed by directly analyzing the audio played after processing on the destination end to derive a quality evaluation. Therefore, in the disclosed embodiments, non-reference evaluation is preferred for determining audio fidelity. Based on this, evaluating the audio fidelity of the target audio played on the destination end can reflect the degree of loss of the target audio during audio transmission and processing, allowing cloud desktop operations and maintenance personnel to intuitively understand the listening experience of cloud desktop users based on the audio fidelity.For example, during the use of a cloud desktop, these audio transmission quality parameters can be collected at regular intervals to generate a monitoring log. Time-correlated line graphs can be plotted in the monitoring log to display changes in the audio transmission quality parameters to operations and maintenance personnel. This allows operations and maintenance personnel to intuitively understand the status of the cloud desktop's audio uplink and downlink links, allowing them to promptly identify and resolve issues that affect the cloud desktop user experience. It should be understood that while the disclosed embodiments propose methods for determining three audio transmission quality parameters: audio transmission delay, audio fidelity, and audio stuttering rate, in actual use, it may not be possible to fully monitor each audio transmission quality parameter at every cycle. Therefore, the disclosed embodiments also propose methods for identifying and resolving audio transmission issues based on each audio transmission quality parameter alone or in combination with other audio transmission quality parameters. In Case 1 (Monitoring Only the Audio Transmission Delay Parameter), since the audio transmission delay parameter can include both the audio transmission link delay and the total audio transmission delay, delay time thresholds can be set for each to detect audio transmission quality issues. For example, if the audio transmission link delay exceeds the first preset duration, this indicates that network latency is severe and has impacted the efficiency of responding to user-requested audio data. In this case, possible adjustment strategies include reducing the network latency between the source and destination endpoints. Specific methods may include checking and resolving network connection issues, changing network connection lines, or notifying the network operator for repairs, among other approaches, though these are not specifically defined here. If the audio transmission link delay does not exceed the first preset duration, but the total audio transmission delay exceeds the second preset duration, this indicates that while the cloud desktop's response speed to the user's requested audio data does not meet requirements, network issues can be ruled out. Therefore, possible issues include playback delays caused by an oversized audio buffer on the destination endpoint or a lag in the audio processing thread on the destination endpoint. Possible adjustment strategies include reducing the destination endpoint's audio buffer length and / or adjusting the destination endpoint's audio processing thread. Specifically, all audio played within the cloud desktop is mixed into one audio channel, encoded by the streaming server within the cloud desktop, and then transmitted to the client over the network. Therefore, regardless of the audio played within the cloud desktop, the client can calculate the corresponding playback duration of the audio within the audio buffer based on the audio buffer length, audio sampling rate, number of audio channels, and audio data bit depth. Therefore, after determining that the cause of the high total audio transmission delay does not include network latency, the audio buffer latency can be further calculated to determine whether the cause of the high total audio transmission delay is an unreasonable audio buffer setting or an audio processing thread problem.Case 2 (Only monitoring the audio stutter rate): Since the audio stutter rate reflects the match between the destination's audio frame reception rate and its audio frame consumption rate, a high audio stutter rate could be caused by excessive network jitter, an undersized audio buffer that fails to address the jitter, or unstable audio processing (including audio processing and playback) rates in the audio processing thread. Therefore, if the audio stutter rate exceeds the preset ratio, possible adjustments include increasing the destination's audio buffer length and / or adjusting the destination's audio processing thread. Furthermore, excessive network jitter can also be caused by local device issues, such as router or network cable failures or local plug-in settings, in addition to carrier issues. Therefore, if the audio stutter rate exceeds the preset ratio, operations personnel can also implement adjustments by sending self-diagnosis prompts to cloud desktop users, reminding them to follow the instructions to gradually troubleshoot local device issues and alleviate network jitter, thereby achieving a better cloud desktop experience. Case 3 (Monitoring Audio Transmission Delay Parameters, Audio Fidelity, and Audio Stuttering Rate) Audio fidelity reflects the overall listening experience and, at a macro level, demonstrates the degree of audio degradation, such as audio speed variation and audio stuttering. Cloud desktop cloud audio links generally have some adjustment mechanisms. For example, they adjust the speed when the audio reception and consumption rates do not match or when audio latency is too high, thereby reducing the probability of audio interruptions. Therefore, low audio fidelity is generally judged based on audio stuttering rate and latency data.

[0002] 1) When the audio transmission delay parameter only includes the total audio transmission delay, if the total audio transmission delay does not exceed the second preset duration, the audio stuttering rate does not exceed the preset ratio, but the audio fidelity does not meet the preset fidelity threshold, then since audio quality issues caused by audio stuttering and audio delay can be ruled out, if the audio fidelity still does not meet the requirements, the target audio source can be determined to be abnormal. Operations and maintenance personnel can report the target audio error, so that the backend can update the target audio source to ensure the subsequent user experience of the cloud desktop. If the total audio transmission delay does not exceed the second preset duration, the audio stuttering rate exceeds the preset ratio, and the audio fidelity does not meet the preset fidelity threshold, possible issues include an undersized audio buffer on the destination end, causing stuttering and affecting audio fidelity, or a stall in the audio processing thread on the destination end, affecting audio fidelity. Possible adjustments include increasing the audio buffer length on the destination end and / or adjusting the audio processing thread on the destination end.

[0003] 2) Audio transmission delay parameters include total audio transmission delay and audio transmission link delay. If the audio transmission link delay does not exceed a first preset duration, but the total audio transmission delay exceeds a second preset duration, the audio stuttering rate does not exceed a preset ratio, and the audio fidelity does not meet the preset fidelity threshold, given normal network latency, the high total delay may be due to a stall in the destination audio processing thread or an oversized destination audio buffer, resulting in low audio fidelity. Possible adjustments include reducing the destination audio buffer length and / or adjusting the destination audio processing thread. Alternatively, in this case, the destination audio buffer length can be first obtained to determine whether the buffer is oversized. If so, the destination audio buffer length can be reduced, and the total audio transmission delay and audio fidelity can be further tested. If these still do not meet the requirements, the destination audio processing thread can be repaired and adjusted for more precise control. If the audio transmission link delay does not exceed the first preset duration, but the total audio transmission delay exceeds the second preset duration, and the audio stuttering rate exceeds the preset ratio, but the audio fidelity meets the preset fidelity threshold, then since both network latency and audio fidelity are normal, the high total latency and high stuttering rate are not caused by a network issue. Possible adjustments include adjusting the destination audio processing thread. Alternatively, the audio stuttering rate exceeding the preset ratio in this case may be due to the audio buffer being slightly smaller than normal, resulting in stuttering but no significant impact on audio fidelity. However, increasing the audio buffer length will increase the total audio transmission delay. The audio buffer length can be adjusted appropriately based on the actual stuttering and latency. For example, if the user prioritizes audio stuttering, the audio buffer length can be increased at the expense of response time to avoid stuttering. If the user prioritizes efficiency, the audio buffer length can be maintained or appropriately reduced to reduce the total audio transmission delay. If the audio transmission link delay does not exceed the first preset duration, but the total audio transmission delay exceeds the second preset duration, the audio freeze rate exceeds a preset ratio, and the audio fidelity does not meet the preset fidelity threshold, possible issues include a freeze in the destination audio processing thread or a significantly shorter audio buffer length. Possible adjustments include increasing the destination audio buffer length and / or adjusting the destination audio processing thread. The above adjustment strategies based on audio transmission quality parameters are intended only to better illustrate and support the intuitive benefits that the method for determining audio transmission quality parameters in the disclosed embodiments can provide to cloud desktop operators. Due to the complex use cases of cloud desktops, in actual applications, operators may also make corresponding configuration and policy changes based on the conditions indicated by the parameters. This disclosure is not limited to such adjustments.Based on the technical solutions of the embodiments of the present disclosure, by monitoring and integrating multiple relevant data points from the uplink and downlink of audio data in cloud desktop scenarios as audio transmission quality parameters, various factors affecting audio transmission quality can be more comprehensively monitored. This allows operations and maintenance personnel to promptly identify issues such as low audio transmission quality caused by network jitter, high network latency, and improper client audio buffer settings, and formulate adjustment strategies to mitigate or eliminate these issues, providing cloud desktop users with a better experience. Another beneficial effect of the embodiments of the present disclosure is that even if the monitored audio transmission quality parameters are partially missing during actual use, the required adjustment strategy can be determined based on some of these parameters, providing fault tolerance for the cloud desktop monitoring system. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of these data must comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation portals are provided for users to select, edit, authorize, or reject. Corresponding to the application scenarios and methods for determining audio transmission quality provided in the embodiments of the present disclosure, the embodiments of the present disclosure further provide an apparatus 400 for determining audio transmission quality. The apparatus may include: an information acquisition module 401, configured to acquire transmission information of target audio between a source end and a destination end and playback information of the target audio on the destination end, wherein one of the source end and the destination end is a cloud desktop server end, and the other is a cloud desktop client end; a determination module 402, configured to determine transmission quality parameters of the target audio based on the transmission information and / or playback information, wherein the transmission quality parameters include at least one of an audio transmission delay parameter, an audio restoration degree, and an audio stuttering rate; the audio transmission delay parameter is used to characterize the delay between the target audio being emitted from the source end and being played by the destination end; the audio restoration degree is used to characterize the difference between the target audio played by the destination end and the target audio sent by the source end; and the audio stuttering rate is used to characterize the stuttering of the target audio played by the destination end. In one embodiment, the information acquisition module is configured to determine an audio transmission delay parameter of the target audio based on the transmission information and / or the playback information. The audio transmission delay parameter includes a total audio transmission delay and / or an audio transmission link delay. The total audio transmission delay is the time difference between the source end sending an audio frame in the target audio and the destination end playing the audio frame. The audio transmission link delay is the time difference between the source end sending an audio frame and the destination end receiving the audio frame.In one embodiment, the device is further configured to reduce the network delay between the source and destination ends when the audio transmission link delay exceeds a first preset duration; or, when the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds a second preset duration, reduce the destination end's audio buffer length and / or adjust the destination end's audio processing thread. In one embodiment, the information acquisition module is configured to determine the audio stutter rate of the target audio based on playback information; the playback information includes the number of stutters in the target audio playback per unit time. In one embodiment, the device is further configured to increase the destination end's audio buffer length and / or adjust the destination end's audio processing thread when the audio stutter rate exceeds a preset ratio. In one embodiment, when the audio transmission delay parameter includes the total audio transmission delay, and the transmission quality parameter of the target audio also includes the audio restoration degree and the audio stuttering rate, the device is further used to: determine that the sound source of the target audio is abnormal when the total audio transmission delay does not exceed a second preset duration, the audio stuttering rate does not exceed a preset ratio, but the audio restoration degree does not reach a preset restoration degree threshold; and increase the audio buffer length of the destination end and / or adjust the audio processing thread of the destination end when the total audio transmission delay does not exceed the second preset duration, the audio stuttering rate exceeds a preset ratio, and the audio restoration degree does not reach the preset restoration degree threshold. In one embodiment, when the audio transmission delay parameter also includes an audio transmission link delay, the device is further configured to: reduce the destination end's audio buffer length and / or adjust the destination end's audio processing thread if the audio transmission link delay does not exceed a first preset duration but the total audio transmission delay exceeds a second preset duration, the audio stuttering rate does not exceed a preset ratio, and the audio fidelity does not reach a preset fidelity threshold; adjust the destination end's audio processing thread if the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds a second preset duration, the audio stuttering rate exceeds a preset ratio, and the audio fidelity does not reach a preset fidelity threshold; and increase the destination end's audio buffer length and / or adjust the destination end's audio processing thread if the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds a second preset duration, the audio stuttering rate exceeds a preset ratio, and the audio fidelity does not reach a preset fidelity threshold. The functions of the various modules in the various devices of the present disclosure embodiments can be found in the corresponding descriptions of the aforementioned methods, and they possess corresponding beneficial effects, and are not further described here. Figure 5 is a block diagram of an electronic device for implementing the present disclosure embodiments. As shown in FIG5 , the electronic device includes a memory 501 and a processor 502 , wherein the memory 501 stores a computer program that can be run on the processor 502 .When the processor 502 executes the computer program, the method in the above embodiment is implemented. The number of memory 501 and processor 502 can be one or more. The electronic device also includes a communication interface 503 for communicating with external devices and exchanging data. If the memory 501, processor 502, and communication interface 503 are implemented independently, the memory 501, processor 502, and communication interface 503 can be interconnected via a bus to enable communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG5 shows only one thick line, but this does not mean that there is only one bus or only one type of bus. Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, the memory 501, processor 502, and communication interface 503 can communicate with each other via an internal interface. The embodiments of the present disclosure provide a computer-readable storage medium storing a computer program. When executed by a processor, the program implements the methods provided in the embodiments of the present disclosure. The embodiments of the present disclosure also provide a chip including a processor configured to retrieve and execute instructions stored in a memory, so that a communication device equipped with the chip performs the methods provided in the embodiments of the present disclosure. The embodiments of the present disclosure also provide a chip including an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected via an internal connection path. The processor is configured to execute code in the memory. When the code is executed, the processor performs the methods provided in the embodiments of the present disclosure. It should be understood that the above-mentioned processor may be a CPU, or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.The general-purpose processor may be a microprocessor or any conventional processor. It is worth noting that the processor may be a processor supporting the Advanced Reduced Instruction Set Machine (ARM) architecture. Furthermore, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example and not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM RAM (DR RAM). In the above embodiments, all or part of them can be implemented through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in or transferred from one computer-readable storage medium to another.In the description of this disclosure, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate different embodiments or examples described in this disclosure, as well as features from different embodiments or examples, unless otherwise specified. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed to indicate or imply relative importance or implicitly specify the number of technical features being referred to. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of those features. In the description of this disclosure, "plurality" means two or more, unless otherwise specifically defined. Any process or method described in a flowchart or otherwise herein can be understood to represent a module, segment, or portion of code comprising one or more executable instructions for implementing a specific logical function or process step. Furthermore, the scope of the preferred embodiments of the present disclosure includes alternative implementations in which functions may be performed out of the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved. The logic and / or steps described in a flowchart or otherwise herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such an instruction execution system, apparatus, or device. It should be understood that various aspects of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above-described embodiment can be completed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiment. In addition, the functional units in the various embodiments of the present disclosure can be integrated into a single processing module, each unit can exist physically separately, or two or more units can be integrated into a single module.The above-mentioned integrated modules can be implemented in either hardware or software functional modules. If implemented as software functional modules and sold or used as independent products, the integrated modules can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a magnetic disk, or an optical disk. The above description is merely an exemplary embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Anyone skilled in the art can easily conceive of various variations and substitutions within the technical scope of the present disclosure, and such variations and substitutions are intended to be encompassed by the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

Claims 1. A method for determining audio transmission quality, comprising: Obtain transmission information of target audio between a source end and a destination end and playback information of the target audio on the destination end, wherein one of the source end and the destination end is a cloud desktop server end and the other is a cloud desktop client end; determine transmission quality parameters of the target audio based on the transmission information and / or the playback information, wherein the transmission quality parameters include at least one of an audio transmission delay parameter, an audio restoration degree, and an audio stuttering rate; the audio transmission delay parameter is used to characterize the delay between the target audio being emitted from the source end and being played by the destination end, the audio restoration degree is used to characterize the difference between the target audio played by the destination end and the target audio sent by the source end; the audio stuttering rate is used to characterize the stuttering of the playback of the target audio by the destination end.

2. The method according to claim 1, wherein: Determining the transmission quality parameters of the target audio according to the transmission information and / or the playback information includes: determining the audio transmission delay parameters of the target audio according to the transmission information and / or the playback information; wherein the audio transmission delay parameters include a total audio transmission delay and / or an audio transmission link delay, the total audio transmission delay being a time difference between the source end sending an audio frame in the target audio and the destination end playing the audio frame, and the audio transmission link delay being a time difference between the source end sending the audio frame and the destination end receiving the audio frame.

3. The method according to claim 2, further comprising: When the audio transmission link delay exceeds a first preset duration, reducing the network delay between the source end and the destination end; Alternatively, when the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds a second preset duration, the audio buffer length of the destination end is reduced and / or the audio processing thread of the destination end is adjusted.

4. The method according to claim 1, wherein: Determining a transmission quality parameter of the target audio according to the transmission information and / or the playback information includes: determining an audio jam rate of the target audio according to the playback information; wherein the playback information includes a number of jams in playing the target audio per unit time.

5. The method according to claim 4, further comprising: When the audio jam rate exceeds a preset ratio, the audio buffer length of the destination terminal is increased and / or the audio processing thread of the destination terminal is adjusted.

6. The method according to claim 2, wherein: When the audio transmission delay parameter includes the total audio transmission delay, and the transmission quality parameter of the target audio also includes the audio restoration degree and the audio stuttering rate, it also includes: when the total audio transmission delay does not exceed the second preset duration, the audio stuttering rate does not exceed the preset ratio, but the audio restoration degree does not reach the preset restoration degree threshold, determining that the sound source of the target audio is abnormal; when the total audio transmission delay does not exceed the second preset duration, the audio stuttering rate exceeds the preset ratio, and the audio restoration degree does not reach the preset restoration degree threshold, increasing the audio buffer length of the destination end and / or adjusting the audio processing thread of the destination end.

7. The method according to claim 6, wherein: In the case where the audio transmission delay parameter also includes the audio transmission link delay, it also includes: when the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds the second preset duration, the audio stuttering rate does not exceed the preset ratio and the audio restoration degree does not reach the preset restoration degree threshold, reducing the audio buffer length of the destination end and / or adjusting the audio processing thread of the destination end; when the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds the second preset duration, the audio stuttering rate exceeds the preset ratio but the audio restoration degree reaches the preset restoration degree When the audio transmission link delay does not exceed the first preset duration but the total audio transmission delay exceeds the second preset duration, the audio stuttering rate exceeds the preset ratio and the audio restoration degree does not reach the preset restoration degree threshold, increase the audio buffer length of the destination end and / or adjust the audio processing thread of the destination end.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, wherein when executed by a processor, the computer program implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice quality evaluation method and device, electronic device and storage medium

    CN110401622A

  • Audio parameter adjustment method, equipment, device and storage medium

    CN116527613A

  • Automated testing of audio and multimedia over remote desktop protocol

    US20080244081A1

  • Voice quality assessment system

    US20220311867A1