Audio sharing method and device
By obtaining the decoded data of shared audio on the cloud desktop client, the problems of sound quality damage and resource overhead caused by audio separation in cloud desktop conferences are solved, high-quality and smooth audio sharing is achieved, and user experience and compatibility are improved.
Patent Information
- Application Number
- CN202410291085.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-16
AI Technical Summary
When sharing sound in a cloud desktop conference, the existing audio separation technology causes echo cancellation to damage the sound quality, and has high resource consumption, poor audio compatibility, and a poor user experience.
Obtain the decoded data of shared audio through the cloud desktop client, avoiding echo cancellation and directly obtaining uplink audio data based on the decoded data. This is applicable to various operating systems and improves audio quality and compatibility.
It improves the sound quality and smoothness of audio sharing, reduces resource overhead, avoids audio interruptions, and enhances user experience and compatibility.
Smart Images

Figure CN120658732A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to an audio sharing method and device. Background Art
[0002] With the development of cloud desktop (Workspace) technology, more and more users can log in to cloud desktops using a variety of terminal devices and use conferencing software to conduct meetings through cloud desktops. Meetings often require audio sharing: the presenter plays video or audio on a cloud desktop shared with attendees and enables the audio sharing function in the conferencing software so that other attendees can hear the audio of the content being played.
[0003] In related technologies, these conferences typically utilize audio separation technology: a conference plug-in within a terminal device captures the conference audio played from the terminal's speakers, performs audio enhancement, and then sends it to the conference server, ensuring a better audio experience for attendees. In scenarios where audio is shared, audio separation technology can cause the captured audio of the shared content to be mixed with the voices of other attendees, creating an echo. To address this, echo cancellation can be performed on the audio captured from the speakers to remove the voices of other attendees and maintain the quality of the shared audio.
[0004] However, echo cancellation can damage the sound quality to a certain extent. Therefore, how to improve the audio quality of meetings conducted via cloud desktops in shared sound scenarios is an urgent problem to be solved. Summary of the Invention
[0005] To address the above technical issues, the present application provides an audio sharing method and apparatus. In this audio sharing method, a conference plug-in subscribes to shared audio from a cloud desktop client used to access the conference client, obtaining shared audio that is not mixed with the voices of other participants, thereby improving the audio quality of sound sharing in conferences conducted via the cloud desktop.
[0006] In a first aspect, the present application provides an audio sharing method, which includes: upon receiving an audio sharing request for shared content from a conference client accessed through a cloud desktop client, sending a subscription request for shared audio to the cloud desktop client; receiving decoded data of the shared audio sent by the cloud desktop client; obtaining uplink audio data based on the decoded data of the shared audio, and sending the uplink audio data to a conference server corresponding to the conference client.
[0007] In an embodiment of the present application, in a scenario where a conference client is accessed through a cloud desktop client, the shared content containing audio shared in the conference is presented on the cloud desktop client, and the audio in the shared content, i.e., the decoded data of the shared audio, is obtained from the cloud desktop client. This decoded data can be used to obtain the shared audio without the voices of the conference participants being superimposed. In this way, the uplink audio data that can be sent to the conference server can be obtained directly based on the decoded data of the shared audio without the need for additional echo cancellation, thereby avoiding damage to the sound quality of the shared audio, thereby improving the sound quality of the shared audio sent by the conference server to the participants. In addition, this can avoid the large resource overhead caused by echo cancellation and further improve the performance of audio sharing.
[0008] Furthermore, compared to capturing shared audio from speakers, obtaining decoded data for shared audio through the cloud desktop client eliminates the need to interact with the speaker driver. This avoids audio interruptions caused by users switching speakers, ensuring smoother shared audio for attendees and further improving the user audio experience. Furthermore, this eliminates the need to develop adapter interfaces for audio device drivers on different operating systems, making it compatible with a wide range of operating systems, thereby improving the compatibility of audio sharing methods.
[0009] According to the first aspect, the decoded data of the shared audio is stored in the cache; after receiving the decoded data of the shared audio sent by the cloud desktop client, the method also includes: receiving audio attribute information sent by the cloud desktop client; obtaining uplink audio data based on the decoded data of the shared audio, including: when the decoded data of the shared audio does not match the audio attribute information, collecting the decoded data of the shared audio from the cache according to the audio attribute information to obtain resampled decoded data; obtaining uplink audio data based on the resampled decoded data.
[0010] In an embodiment of the present application, when the decoded data of the cached shared audio does not match the audio attribute information provided by the cloud desktop client, the decoded data of the cached shared audio is sampled according to the audio attribute information, thereby ensuring that the resampled decoded data can be played normally in the cloud desktop client, further improving the quality of the shared audio provided to the participants.
[0011] According to the first aspect, or any implementation method of the first aspect above, after sending a subscription request for shared audio to the cloud desktop client, the method also includes: when receiving a request from the conference client to end sharing of shared content, stopping the execution of decoding data based on the shared audio to obtain uplink audio data; and sending a cancellation request for shared audio to the cloud desktop client.
[0012] In an embodiment of the present application, when audio sharing is stopped, the acquisition of uplink audio data is stopped, and a subscription cancellation request is sent to the cloud desktop client to instruct the cloud desktop client to stop sending decoded data of the shared audio, thereby avoiding occupation of cloud desktop client resources and further improving the user experience of the meeting.
[0013] According to the first aspect, or any implementation method of the first aspect above, uplink audio data is obtained based on the decoded data of the shared audio, including: when the audio data of the microphone is collected, the audio data of the microphone and the decoded data of the shared audio are mixed to obtain uplink data; wherein, the microphone is connected to the running device of the cloud desktop client.
[0014] In an embodiment of the present application, when the audio data of the speaker's microphone is collected, the decoded data of the shared audio is mixed with the audio data of the microphone to obtain uplink audio data, thereby meeting the improvement of the shared audio quality in the scenario where the speaker shares the sound while speaking.
[0015] In a second aspect, an embodiment of the present application provides an audio sharing method, the method comprising: receiving a subscription request from a conference plug-in for shared audio played through a cloud desktop client; wherein the shared audio is used to be shared in a conference corresponding to the conference plug-in, and the conference is conducted by a conference client accessed through a cloud desktop client; obtaining decoded data of the shared audio, and sending the decoded data of the shared audio to the conference plug-in.
[0016] According to the second aspect, after receiving a subscription request from the conference plug-in for shared audio played through the cloud desktop client, the method further includes: sending audio attribute information to the conference plug-in; wherein the audio attribute information is used for playing audio data on the cloud desktop client.
[0017] According to the second aspect, or any implementation of the second aspect above, after receiving a subscription request from the conference plug-in for the shared audio played through the cloud desktop client, the method also includes: upon receiving a cancellation request for the shared audio sent by the conference plug-in, stopping sending the decoded data of the shared audio to the conference plug-in; wherein the cancellation request is sent by the conference plug-in when receiving a request to end sharing.
[0018] According to the second aspect, or any implementation method of the above second aspect, obtaining the decoded data of the shared audio includes: when receiving a playback indication of the shared audio, sending a request for obtaining the shared audio to the cloud desktop server corresponding to the cloud desktop client; receiving the shared audio sent by the cloud desktop server; decoding the shared audio to obtain the decoded data of the shared audio.
[0019] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0020] In a third aspect, an embodiment of the present application provides an audio sharing method, which includes: when the conference plug-in receives an audio sharing request for shared content from a conference client accessed through a cloud desktop client, the conference plug-in sends a subscription request for shared audio to the cloud desktop client; when the cloud desktop client receives the subscription request for shared audio, the cloud desktop client obtains the decoded data of the shared audio and sends the decoded data of the shared audio to the conference plug-in; when the conference plug-in receives the decoded data of the shared audio, the conference plug-in obtains uplink audio data based on the decoded data of the shared audio and sends the uplink audio data to the conference server corresponding to the conference client.
[0021] The third aspect and any implementation of the third aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the third aspect and any implementation of the third aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0022] In a fourth aspect, an embodiment of the present application provides an audio sharing device, which includes: a subscription module for sending a subscription request for shared audio to a cloud desktop client when receiving an audio sharing request for shared content from a conference client accessed through a cloud desktop client; a receiving module for receiving decoded data of shared audio sent by the cloud desktop client; an audio module for obtaining uplink audio data based on the decoded data of shared audio; and an upload module for sending uplink audio data to a conference server corresponding to the conference client.
[0023] The fourth aspect and any implementation of the fourth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0024] In the fifth aspect, an embodiment of the present application provides an audio sharing device, which includes: a receiving module for receiving a subscription request from a conference plug-in for shared audio played through a cloud desktop client; wherein the shared audio is used to be shared in a conference corresponding to the conference plug-in, and the conference is conducted by a conference client accessed through a cloud desktop client; an acquisition module for obtaining decoded data of the shared audio; and a sending module for sending decoded data of the shared audio to the conference plug-in.
[0025] The fifth aspect and any implementation of the fifth aspect correspond to the second aspect and any implementation of the second aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can be referred to the technical effects corresponding to the above-mentioned second aspect and any implementation of the second aspect, and will not be repeated here.
[0026] In a sixth aspect, an embodiment of the present application provides an audio sharing device, which includes: a conference plug-in and a cloud desktop client; the conference plug-in is used to send a subscription request for shared audio to the cloud desktop client when receiving an audio sharing request for shared content from a conference client accessed through the cloud desktop client; the cloud desktop client is used to obtain the decoded data of the shared audio when receiving the subscription request for shared audio, and send the decoded data of the shared audio to the conference plug-in; the conference plug-in is used to obtain uplink audio data based on the decoded data of the shared audio when receiving the decoded data of the shared audio, and send the uplink audio data to the conference server corresponding to the conference client.
[0027] The sixth aspect and any implementation of the sixth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the sixth aspect and any implementation of the sixth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0028] In the seventh aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a transceiver; a memory for storing one or more programs; when the one or more programs are executed by one or more processors, the one or more processors implement a method such as the first aspect, the second aspect, the third aspect, any one of the implementations of the first aspect, any one of the implementations of the second aspect, and any one of the implementations of the third aspect.
[0029] In an eighth aspect, embodiments of the present application provide a computer-readable storage medium comprising a computer program, the computer program comprising instructions for executing the method of the first aspect or any possible implementation of the first aspect. When the computer program is executed in an electronic device, the electronic device executes the method of any one of the first aspect, the second aspect, the third aspect, any implementation of the first aspect, any implementation of the second aspect, and any implementation of the third aspect.
[0030] In the ninth aspect, an embodiment of the present application provides a computer program, which includes instructions for executing the method in the first aspect or any possible implementation of the first aspect. When the computer program is run in an electronic device, the electronic device executes the method in the first aspect or any possible implementation of the first aspect.
[0031] In the tenth aspect, an embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors; the interface circuit is used to receive signals from a memory of an electronic device and send signals to the processor, the signals comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the instructions of the method in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0033] Figure 1 This is a schematic diagram of the audio sharing scenario in a meeting;
[0034] Figure 2 This is an example diagram of the processing flow of audio sharing in a meeting;
[0035] Figure 3 This is one of the example diagrams of the audio sharing scenario in a meeting provided by the embodiment of the present application;
[0036] Figure 4 Schematic diagram of the flow of the audio sharing method provided in the embodiment of the present application;
[0037] Figure 5 This is one of the example diagrams of the audio sharing scenario in a meeting provided by the embodiment of the present application;
[0038] Figure 6 This is a schematic diagram of the processing flow of audio sharing in a meeting provided by an embodiment of the present application;
[0039] Figure 7 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application;
[0040] Figure 8 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application;
[0041] Figure 9 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application;
[0042] Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0044] To facilitate understanding of this embodiment, some technical terms and background technologies involved in this embodiment are first introduced:
[0045] Audio separation: Conference audio capture, playback, and audio enhancement are implemented on devices running the cloud desktop client, such as TC (Thin Client) devices. The conference client (such as the conference app) runs on the computing node (such as a server, virtual machine, or container) where the cloud desktop server resides. The conference client sends control signaling to the conference plug-in in the cloud desktop client on the TC device through the Huawei Desktop Protocol (HDP) virtual channel. The plug-in controls the audio engine on the TC device (integrated in the cloud desktop client as an SDK) to operate audio devices such as speakers and microphones.
[0046] Shared Audio: The conference presenter shares the audio output of the local TC device with other participants, allowing other participants to hear the audio played by the presenter's local TC device.
[0047] Driver: A computer software term referring to the software program that drives computer hardware. A driver, short for device driver, is a special program added to the operating system that contains information about hardware devices. This information enables the computer to communicate with the corresponding device. A driver is a configuration file written by the hardware manufacturer for the operating system. Without a driver, the hardware in the computer will not function.
[0048] A plug-in (also known as addin, add-in, addon, or add-on) is a program written using a standardized application programming interface (API). A plug-in requires access to libraries or data provided by the native system. Therefore, it must run on the platform specified by the program (which may support multiple platforms simultaneously) and cannot run independently of the specified platform.
[0049] Software Development Kit (SDK): In a broad sense, an SDK is a collection of development tools used to build applications for a specific software package, software framework, hardware platform, operating system, etc. In a narrower sense, an SDK is a new, independent set of tools developed and packaged based on the system SDK that can perform specific functions and return relevant data. An SDK can be thought of as a collection of plug-ins.
[0050] For example, Figure 1 This is a diagram of the audio sharing scenario in a conference. Figure 1 As shown in the figure, in the audio separation scenario, the presenter uses the cloud desktop to remotely access the conference client on the virtual machine to hold a remote conference with at least one ordinary participant. At this time, if the presenter shares the desktop with the ordinary participant and plays a video on the shared desktop, the process of sharing the sound is involved:
[0051] S11, the speaker (sharing sound) shares the desktop and plays a video on the desktop;
[0052] S12, the voices of ordinary participants will also be played on the speaker's computer;
[0053] S13, the audio played by the speaker's computer is shared with ordinary participants (shared audio receiving end);
[0054] S14, the speaker's computer excludes the voices of ordinary participants (echo cancellation).
[0055] Specifically, the sound played by the presenter's computer (such as a TC device) is collected through the speakers of that computer, that is, the sound played by the actual audio device is collected. The sound played by the speakers includes the sound of the shared video and the voices of ordinary participants. In this case, if the voices of ordinary participants are not removed, ordinary participants will hear their own voices. Therefore, it is necessary to remove the voices of ordinary participants from the sound played by the speakers before sending it to ordinary participants. This way, the sound effect is acceptable to ordinary participants.
[0056] For example, Figure 2 This is an example diagram of the audio sharing process in a meeting. Figure 2As shown, the speaker side can introduce an echo cancellation module to eliminate the voices of other participants. For example, the speaker side refers to the voices of other participants, performs echo cancellation on the shared audio obtained by collecting the speaker's voice, and performs audio enhancement, also known as voice quality enhancement (VQE). In addition, the speaker side can use a microphone to collect the speaker's voice when speaking and perform audio enhancement, and then mix it with the shared audio after VQE. The mixed result is packaged and encoded uplink and sent to other participants via the network, thereby ensuring that other participants can hear the shared sound and that the shared sound does not contain the voices of ordinary participants themselves.
[0057] However, the above-mentioned echo cancellation may damage the sound quality, especially in music scenes, where the texture is significantly affected. In addition, echo cancellation will also occupy a certain amount of resource overhead on the speaker side. In one case, the sound played by the speaker is likely to involve the switching of audio devices such as speakers and headphones, for example, the speaker plugs in and unplugs headphones, speakers, etc. The switching of audio devices will cause a brief interruption in the played sound, and accordingly, the sound collected from the speakers will also be interrupted, reducing the user's audio experience. In addition, the operating systems of the speaker side, such as TC devices, are often diverse, such as Windows, Ubuntu, UOS and other operating systems, resulting in diverse drivers for audio devices. It is also necessary to develop adapter interfaces for drivers of audio devices of different operating systems, resulting in reduced audio sharing compatibility and inconvenience.
[0058] An embodiment of the present application provides an audio sharing method to solve the above-mentioned problem. In this method, in the scenario of a conference client accessed through a cloud desktop client, the shared content containing audio shared in the conference is presented on the cloud desktop client, and the audio in the shared content, that is, the decoded data of the shared audio, is obtained from the cloud desktop client. The decoded data can be used to obtain the shared audio without superimposing the voices of the conference participants. In this way, the uplink audio data that can be sent to the conference server can be obtained directly based on the decoded data of the shared audio without the need for additional echo cancellation, thereby avoiding damage to the sound quality of the shared audio, thereby improving the sound quality of the shared audio sent by the conference server to the participants. In addition, this can avoid the large resource overhead caused by echo cancellation and further improve the performance of audio sharing.
[0059] Furthermore, compared to capturing shared audio from speakers, obtaining decoded data for shared audio through the cloud desktop client eliminates the need to interact with the speaker driver. This avoids audio interruptions caused by users switching speakers, ensuring smoother shared audio for attendees and further improving the user audio experience. Furthermore, this eliminates the need to develop adapter interfaces for audio device drivers on different operating systems, making it compatible with a wide range of operating systems, thereby improving the compatibility of audio sharing methods.
[0060] The following first introduces the scenario of the audio sharing method provided in the embodiment of the present application.
[0061] For example, Figure 3 This is an example diagram of an audio sharing scenario in a conference provided by the embodiment of the present application. Figure 3 As shown, in the scenario where a user uses a TC device to log in to a cloud desktop client for a remote meeting, participants can include the speaker side that is currently sharing the sound and other participants other than the speaker, that is, the participant side. The speaker side and the participant side conduct conference-related communications through the conference server. The speaker side includes a cloud desktop server side (for example, a virtual machine) and a cloud desktop client based on HDP remote communication, a conference APP running on the cloud desktop server side, that is, the conference client, and a conference plug-in running on the cloud desktop client. The conference APP can send signaling to the conference plug-in through the conference server to control the audio device of the device where the cloud desktop client is located.
[0062] The following combination Figures 4 to 10 The audio sharing method provided in the embodiment of the present application is described in detail.
[0063] Figure 4 Schematic diagram of the flow of the audio sharing method provided in the embodiment of the present application. Figure 4 As shown, the audio sharing method may include:
[0064] S401, the conference plug-in receives an audio sharing request for shared content from a conference client;
[0065] For example, Figure 5 This is one of the example diagrams of audio sharing scenarios in a conference provided by the embodiment of this application. Figure 5 As shown, on the speaker side, also known as the conference sharing side, the speaker accesses the conference app through the cloud desktop client and enables audio sharing via the HDP virtual channel between the conference app and the conference plug-in. At this point, the conference plug-in receives the audio sharing request from the conference client for shared content. This audio sharing request can be, for example, a control signaling signal sent via the HDP virtual channel.
[0066] In addition, shared content is content shared by the presenter to the participants, and specifically may be content including audio, video, etc., and is played by the presenter on the cloud desktop client.
[0067] S402, the conference plug-in sends a subscription request for shared audio to the cloud desktop client;
[0068] For example, see Figure 3 The process of playing audio on the cloud desktop client may include S31 to S34:
[0069] S31, the cloud desktop server collects the virtual speaker sound;
[0070] Virtual speaker sound refers to the audio stream from the virtual speakers of an operating system, such as Windows. The cloud desktop server can call the virtual speaker interface of the operating system to collect the audio stream.
[0071] S32, encoding and sending;
[0072] The cloud desktop server encodes the audio stream collected from the virtual speaker and transmits it to the cloud desktop client according to HDP.
[0073] S33, receiving;
[0074] The cloud desktop client receives the encoded audio stream sent by the cloud desktop server according to HDP.
[0075] S34, decoding;
[0076] The cloud desktop client decodes the received audio stream and sends it to the actual audio device of the TC for playback.
[0077] Based on the aforementioned audio playback process on the cloud desktop client, when the presenter enables the audio sharing feature, the conference plug-in subscribes to the decoded audio stream data from the cloud desktop client, thereby subscribing to the shared audio. The subscription request can instruct the cloud desktop client to continuously send the decoded shared audio data to the conference plug-in during the shared audio playback period.
[0078] S403, the cloud desktop client obtains decoded data of the shared audio;
[0079] In an optional implementation, the cloud desktop client obtains the decoded data of the shared audio, which may specifically include:
[0080] When the cloud desktop client receives the instruction to play the shared audio, it sends a request to obtain the shared audio to the cloud desktop server corresponding to the cloud desktop client;
[0081] The cloud desktop client receives the shared audio sent by the cloud desktop server;
[0082] The cloud desktop client decodes the shared audio to obtain decoded data of the shared audio.
[0083] For example, see Figure 5 When the speaker plays the content containing audio on the cloud desktop client, the cloud desktop client requests the content to be played from the cloud desktop server. Figure 3 The same acquisition and encoding process is used, communicating with the cloud desktop client based on the HDP protocol to transmit the encoded audio data. The cloud desktop client decodes the received audio data to generate an audio stream that can be played on the actual audio device. This audio stream contains the audio data of the shared audio itself and does not contain other sounds played on the actual audio device.
[0084] S404, the cloud desktop client sends decoded data of the shared audio to the conference plug-in;
[0085] For example, Figure 6 This is a schematic diagram of the processing flow of audio sharing in a conference provided by the embodiment of this application. Figure 6 As shown, after the cloud desktop client obtains the decoded data of the shared audio, it can send the decoded data of the shared audio to the conference plug-in. The conference plug-in can directly obtain the decoded audio stream sent by the cloud desktop client to the actual audio device before playback, that is, the decoded data used for playback, and thus obtain a purer video sound, thereby ensuring that the shared audio is not mixed with the voices of ordinary participants speaking received by the conference sharing end.
[0086] In an optional implementation, the conference plug-in stores the received decoded data of the shared audio in a cache;
[0087] After the cloud desktop client receives the subscription request from the conference plug-in for the shared audio played through the cloud desktop client, the audio sharing method provided in the embodiment of the present application may further include:
[0088] The cloud desktop client sends audio attribute information to the conference plug-in; wherein the audio attribute information is used for playing audio data on the cloud desktop client;
[0089] The conference plug-in receives audio attribute information sent by the cloud desktop client;
[0090] When the decoded data of the shared audio does not match the audio attribute information, the conference plug-in collects the decoded data of the shared audio from the cache according to the audio attribute information to obtain resampled decoded data;
[0091] The conference plug-in obtains uplink audio data based on the resampled decoded data.
[0092] Exemplarily, the audio attribute information may be, for example, the sampling rate, number of channels, bit width, decoded data format, and other audio configuration parameters of the audio data that can ensure normal playback of the audio data on the cloud desktop client. The cloud desktop client caches the decoded audio data as the decoded data of the shared audio and sends it to the conference plug-in, and the conference plug-in stores the received decoded data of the shared audio in the cache. After receiving the subscription request sent by the conference plug-in, the cloud desktop client forwards the sampling rate, number of channels, decoded data format and other audio attribute information of the decoded data to the conference plug-in. The conference plug-in can match the attribute information such as the sampling rate, number of channels, bit width, etc. of the cached decoded data of the shared audio with the audio attribute information sent by the cloud desktop client according to a review cycle, for example, 10ms. If there is any inconsistency, a resampling operation is performed: the decoded data of the shared audio is collected from the cache according to the audio attribute information to obtain resampled decoded data.
[0093] In an embodiment of the present application, when the decoded data of the cached shared audio does not match the audio attribute information provided by the cloud desktop client, the decoded data of the cached shared audio is sampled according to the audio attribute information, thereby ensuring that the resampled decoded data can be played normally in the cloud desktop client, further improving the quality of the shared audio provided to the participants.
[0094] S405, the conference plug-in obtains uplink audio data based on the decoded data of the shared audio;
[0095] In one example, if the speaker is not speaking, the conference sharing end will not collect audio data from the microphone connected to the device running the cloud desktop client. At this time, the conference plug-in can perform VQE processing on the decoded data of the shared audio to obtain enhanced decoded data, and encode and package the enhanced decoded data to obtain uplink audio data.
[0096] In an optional implementation, the conference plug-in obtains uplink audio data based on the decoded data of the shared audio, which may specifically include:
[0097] When audio data from the microphone is collected, the audio data from the microphone is mixed with decoded data of the shared audio to obtain uplink data; wherein the microphone is connected to a running device of the cloud desktop client.
[0098] For example, see Figure 6 After the conference plug-in obtains the pure audio stream (i.e., the video sound) through subscription, it can mix it with the audio data collected by the microphone. The mixed data is then packaged and encoded for uplink transmission to the conference server via the network. It is understood that the microphone connected to the device running the cloud desktop client can be the microphone on the device, or it can be an external microphone connected to the device.
[0099] In an embodiment of the present application, when the audio data of the speaker's microphone is collected, the decoded data of the shared audio is mixed with the audio data of the microphone to obtain uplink audio data, thereby meeting the improvement of the shared audio quality in the scenario where the speaker shares the sound while speaking.
[0100] S406: The conference plug-in sends the uplink audio data to the conference server corresponding to the conference client.
[0101] For example, see Figure 3 The conference plug-in can send uplink audio data to the conference server corresponding to the conference client, such as the conference server, through the communication network. The conference server sends the uplink audio data to the participant end. The participant end receives and plays the audio and / or video sent by the other end, and can hear the voice of the speaker sharing the content.
[0102] In an optional implementation, after the conference plug-in sends a subscription request for audio sharing to the cloud desktop client, the audio sharing method provided in the embodiment of the present application may further include:
[0103] When the conference plug-in receives a request from the conference client to end sharing of the shared content, it stops decoding the shared audio data and obtains the uplink audio data.
[0104] The conference plug-in sends a request to unsubscribe from shared audio to the cloud desktop client.
[0105] Accordingly, after the cloud desktop client receives the subscription request from the conference plug-in for the shared audio played through the cloud desktop client, it can stop sending the decoded data of the shared audio to the conference plug-in when it receives the unsubscribe request for the shared audio sent by the conference plug-in; wherein, the unsubscribe request is sent by the conference plug-in when it receives the end sharing request.
[0106] For example, if the presenter stops sharing video, the cloud conference app sends a request to the conference plugin to stop sharing audio, using the corresponding control signaling. The conference plugin then stops mixing audio and sends a cancel subscription request to the cloud desktop client. Upon receiving the cancel subscription request, the cloud desktop client stops forwarding the audio stream to the conference plugin.
[0107] In an embodiment of the present application, when audio sharing is stopped, the acquisition of uplink audio data is stopped, and a subscription cancellation request is sent to the cloud desktop client to instruct the cloud desktop client to stop sending decoded data of the shared audio, thereby avoiding occupation of cloud desktop client resources and further improving the user experience of the meeting.
[0108] An embodiment of the present application also provides an audio sharing device.
[0109] For example, Figure 7 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application. Figure 7 As shown, the audio sharing device includes:
[0110] The subscription module 701 is configured to send a subscription request for shared audio to the cloud desktop client upon receiving an audio sharing request for shared content from the conference client accessed through the cloud desktop client;
[0111] The receiving module 702 is configured to receive decoded data of the shared audio sent by the cloud desktop client;
[0112] The audio module 703 is configured to obtain uplink audio data based on the decoded data of the shared audio;
[0113] The uploading module 704 is configured to send uplink audio data to the conference server corresponding to the conference client.
[0114] For example, Figure 8 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application. Figure 8 As shown, the audio sharing device includes:
[0115] Receiving module 801 is used to receive a subscription request from a conference plug-in for shared audio played through a cloud desktop client; wherein the shared audio is used to share in a conference corresponding to the conference plug-in, and the conference is conducted by a conference client accessed through the cloud desktop client;
[0116] An acquisition module 802 is configured to acquire decoded data of the shared audio;
[0117] The sending module 803 is configured to send the decoded data of the shared audio to the conference plug-in.
[0118] For example, Figure 9 This is one of the structural block diagrams of the audio sharing device provided in the embodiment of the present application. Figure 9 As shown, the audio sharing device includes: a conference plug-in 901 and a cloud desktop client 902;
[0119] The conference plug-in 901 is used to send a subscription request for shared audio to the cloud desktop client when receiving an audio sharing request for shared content from the conference client accessed through the cloud desktop client;
[0120] The cloud desktop client 902 is configured to obtain decoded data of the shared audio upon receiving a subscription request for the shared audio, and send the decoded data of the shared audio to the conference plug-in;
[0121] The conference plug-in 901 is used to obtain uplink audio data based on the decoded data of the shared audio when receiving the decoded data of the shared audio, and send the uplink audio data to the conference server corresponding to the conference client.
[0122] Among them, all relevant contents of each step involved in the above-mentioned method embodiment of the present application can be referred to the functional description of the corresponding functional modules in the above-mentioned device embodiments of the present application, and will not be repeated here.
[0123] Figures 7 to 9 The audio sharing devices shown can all be implemented through software or hardware.
[0124] As an example of a software functional unit, a module may include code running on a computing instance, wherein the computing instance may be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices.
[0125] As an example of a hardware functional unit, a module can include at least one computing device, such as a TC. Alternatively, the YY device can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be implemented using a CPLD, FPGA, GAL, or any combination thereof.
[0126] For example, Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device 1000 may be a device or a chip or functional module in the device. Figure 10 As shown, the electronic device 1000 includes a processor 1001 , a transceiver 1002 and a communication circuit 1003 .
[0127] Among them, the processor 1001 is used to execute any step in the audio sharing method embodiment provided by the above-mentioned application, and when executing processes such as receiving encoded image bit streams, it can optionally call the transceiver 1002 and the communication line 1003 to complete the corresponding operations.
[0128] Furthermore, the electronic device 1000 may further include a memory 1004 . The processor 1001 , the memory 1004 and the transceiver 1002 may be connected via a communication line 1003 .
[0129] Transceiver 1002 is used to communicate with other devices or other communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Transceiver 1002 can be a module, circuit, transceiver, or any device capable of implementing communication.
[0130] The transceiver 1002 is mainly used for transmitting and receiving bit streams, etc., and may include a transmitter and a receiver for transmitting and receiving bit streams, etc. respectively; operations other than transmitting and receiving bit streams, etc. are implemented by the processor, such as enabling at least one target channel in the channel group of the chip, etc.
[0131] The communication line 1003 is used to transmit information between the components included in the electronic device 1000.
[0132] In one design, the processor can be considered as the logic circuit and the transceiver as the interface circuit.
[0133] The memory 1004 is used to store instructions, where the instructions may be computer programs.
[0134] It should be noted that the memory 1004 can exist independently of the processor 1001 or can be integrated with the processor 1001. The memory 1004 can be used to store instructions, program code, or some data. The memory 1004 can be located inside the electronic device 1000 or outside the electronic device 1000, without limitation. The processor 1001 is configured to execute the instructions stored in the memory 1004 to implement the methods provided in the above embodiments of the present application.
[0135] In one example, the processor 1001 may include one or more processors, such as Figure 10 Processor 0 and processor 1 in.
[0136] As an optional implementation, the electronic device 1000 includes multiple processors, for example, Figure 10 In addition to the processor 1001, the processor 1007 may also be included.
[0137] As an optional implementation, the electronic device 1000 further includes an output device 1005 and an input device 1006. For example, the input device 1006 includes at least a microphone, and may also be a keyboard, mouse, or joystick, and the output device 1005 is a display screen and a speaker.
[0138] It is understandable that when the electronic device 1000 does not include the output device 1005 and the input device 1006 , the speaker and microphone used for cloud conferencing and the display screen of the cloud desktop can be externally connected to the electronic device 1000 .
[0139] It should be noted that the electronic device 1000 can be a chip system or a Figure 10 Devices with similar structures in the chip system. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only examples, and other names can also be used in specific implementations without limitation. In addition, Figure 10 The composition structure shown in the figure does not constitute a limitation on the electronic device 1000. Figure 10 In addition to the components shown, the electronic device 1000 may include Figure 10 More or fewer components may be shown, or certain components may be combined, or the components may be arranged differently.
[0140] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), n-type metal oxide semiconductor (NMOS), p-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0141] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the audio sharing method provided in any of the above embodiments of the present application.
[0142] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the audio sharing method provided in any of the embodiments of the present application.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
[0144] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0145] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0146] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0147] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
Claims
1. An audio sharing method, characterized in that: The method comprises: Upon receiving an audio sharing request for shared content from a conference client accessed through a cloud desktop client, sending a subscription request for shared audio to the cloud desktop client; Receiving decoded data of the shared audio sent by the cloud desktop client; Based on the decoded data of the shared audio, uplink audio data is obtained, and the uplink audio data is sent to the conference server corresponding to the conference client.
2. The method according to claim 1, characterized in that The decoded data of the shared audio is stored in a cache; After receiving the decoded data of the shared audio sent by the cloud desktop client, the method further includes: Receiving audio attribute information sent by the cloud desktop client; The obtaining uplink audio data based on the decoded data of the shared audio includes: In a case where the decoded data of the shared audio does not match the audio attribute information, collecting the decoded data of the shared audio from the cache according to the audio attribute information to obtain resampled decoded data; Uplink audio data is obtained based on the resampled decoded data.
3. The method according to claim 1 or 2, characterized in that After sending the subscription request for shared audio to the cloud desktop client, the method further includes: When receiving a request from the conference client to end sharing of the shared content, stopping the decoding of the shared audio data to obtain uplink audio data; Send a request to cancel subscription of shared audio to the cloud desktop client.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining uplink audio data based on the decoded data of the shared audio includes: When audio data from a microphone is collected, the audio data from the microphone is mixed with the decoded data of the shared audio to obtain the uplink data; wherein the microphone is connected to a running device of the cloud desktop client.
5. An audio sharing method, characterized in that: The method comprises: Receiving a subscription request from a conference plug-in for shared audio played through a cloud desktop client; wherein the shared audio is used to be shared in a conference corresponding to the conference plug-in, and the conference is conducted by a conference client accessed through the cloud desktop client; Obtain decoded data of the shared audio, and send the decoded data of the shared audio to the conference plug-in.
6. The method according to claim 5, characterized in that After receiving the subscription request from the conference plug-in for the shared audio played by the cloud desktop client, the method further includes: Sending audio attribute information to the conference plug-in; wherein the audio attribute information is used for playing audio data on the cloud desktop client.
7. The method according to claim 5 or 6, characterized in that After receiving a subscription request from the conference plug-in for the shared audio played through the cloud desktop client, the method further includes: Upon receiving a shared audio unsubscription request sent by the conference plug-in, stop sending the decoded data of the shared audio to the conference plug-in; wherein the unsubscription request is sent by the conference plug-in upon receiving an end sharing request.
8. The method according to any one of claims 5 to 7, characterized in that The obtaining the decoded data of the shared audio includes: Upon receiving the play instruction of the shared audio, sending a request for obtaining the shared audio to the cloud desktop server corresponding to the cloud desktop client; Receiving the shared audio sent by the cloud desktop server; The shared audio is decoded to obtain decoded data of the shared audio.
9. An audio sharing method, characterized in that: The method comprises: When the conference plug-in receives an audio sharing request for shared content from a conference client accessed through a cloud desktop client, it sends a subscription request for shared audio to the cloud desktop client; When receiving the subscription request for the shared audio, the cloud desktop client obtains the decoded data of the shared audio and sends the decoded data of the shared audio to the conference plug-in; When receiving the decoded data of the shared audio, the conference plug-in obtains uplink audio data based on the decoded data of the shared audio, and sends the uplink audio data to the conference server corresponding to the conference client.
10. An audio sharing device, characterized in that: The device comprises: A subscription module, configured to, upon receiving an audio sharing request for shared content from a conference client accessed through a cloud desktop client, send a subscription request for shared audio to the cloud desktop client; A receiving module, configured to receive decoded data of the shared audio sent by the cloud desktop client; an audio module, configured to obtain uplink audio data based on the decoded data of the shared audio; The uploading module is used to send the uplink audio data to the conference server corresponding to the conference client.
11. An audio sharing device, characterized in that: The device comprises: A receiving module, configured to receive a subscription request from a conference plug-in for shared audio played through a cloud desktop client; wherein the shared audio is used to be shared in a conference corresponding to the conference plug-in, and the conference is conducted by a conference client accessed through the cloud desktop client; An acquisition module, configured to acquire decoded data of the shared audio; A sending module is used to send the decoded data of the shared audio to the conference plug-in.
12. An audio sharing device, characterized in that: The device includes: a conference plug-in and a cloud desktop client; The conference plug-in is used to send a subscription request for shared audio to the cloud desktop client upon receiving an audio sharing request for shared content from the conference client accessed through the cloud desktop client; The cloud desktop client is configured to obtain decoded data of the shared audio upon receiving a subscription request for the shared audio, and send the decoded data of the shared audio to the conference plug-in; The conference plug-in is used to obtain uplink audio data based on the decoded data of the shared audio when receiving the decoded data of the shared audio, and send the uplink audio data to the conference server corresponding to the conference client.
13. An electronic device, characterized in that: include: processors and transceivers; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
14. A computer-readable storage medium, characterized in that The method comprises a computer program, wherein when the computer program is run on a camera, the camera is caused to execute the method according to any one of claims 1 to 9.
15. A computer program product, characterized in that The method comprises a computer program, which, when executed by an electronic device, causes the electronic device to perform the method according to any one of claims 1 to 9.