Audio processing methods, cloud desktop servers, clients and systems

By establishing multiple recording channels between the cloud desktop server and multiple clients and performing audio mixing, the problem of not being able to acquire audio from multiple clients simultaneously in existing technologies is solved, improving user experience and audio processing efficiency.

CN116055580BActive Publication Date: 2026-04-03ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In cloud desktop servers, existing technologies cannot simultaneously acquire audio from multiple clients, resulting in a poor user experience.

Method used

By establishing multiple recording channels between the cloud desktop server and multiple clients, the system receives and processes the audio from multiple clients, performs mixing, and generates mixed audio for each client and conference terminal.

Benefits of technology

This improves the user experience, enabling cloud desktop servers to process audio from multiple clients simultaneously, thus enhancing the efficiency and quality of audio processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116055580B_ABST
    Figure CN116055580B_ABST
Patent Text Reader

Abstract

This application provides an audio processing method, a cloud desktop server, a client, and a system. Multiple clients can access the cloud desktop server, and each client has a recording channel with the cloud desktop server. The method includes: receiving at least one first audio signal from multiple clients through the recording channel; receiving a second audio signal from a conference server, the second audio signal being obtained by the conference server processing audio from at least one conference terminal; determining a mixed audio signal corresponding to each client and conference terminal based on the at least one first audio signal and the second audio signal, and sending the corresponding mixed audio signal to each client and the conference server. This improves the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio processing method, a cloud desktop server, a client, and a system. Background Technology

[0002] In some scenarios, users can log in to the cloud desktop server on the client and conduct audio and video conferences through the cloud desktop server.

[0003] In related technologies, multiple clients can log in to a cloud desktop server. However, because there is usually a recording channel between the cloud desktop server and the clients, only one client's audio can be captured at a time through the recording channel, making it impossible to capture the audio of multiple clients simultaneously, resulting in a poor user experience. Summary of the Invention

[0004] This application provides an audio processing method, a cloud desktop server, a client, and a system to improve the user experience.

[0005] In a first aspect, embodiments of this application provide an audio processing method applied to a cloud desktop server, wherein multiple clients access the cloud desktop server, and each client has a recording channel with the cloud desktop server; the method includes:

[0006] Receive at least one first audio signal sent by the plurality of clients through the recording channel between the plurality of clients;

[0007] Receive a second audio message sent by the conference server, the second audio message being obtained by the conference server processing audio from at least one conference terminal;

[0008] Based on the at least one first audio and the second audio, determine the mixed audio corresponding to each client and the conference terminal, and send the corresponding mixed audio to each client and the conference server.

[0009] Secondly, embodiments of this application provide an audio processing method applied to a cloud desktop server, the cloud desktop server being used to connect to multiple clients, the method comprising:

[0010] Receive at least one first audio signal sent by the plurality of clients through the recording channel between the plurality of clients;

[0011] Receive the second audio sent by the conference server;

[0012] Based on the at least one first audio and the second audio, determine the mixed audio corresponding to each client and the conference terminal, and send the corresponding mixed audio to each client and the conference server.

[0013] In one possible implementation, based on the at least one first audio and the second audio, a mixed audio corresponding to each client and the conference terminal is determined, and the corresponding mixed audio is sent to each client and the conference server, including:

[0014] The at least one first audio and the second audio are mixed to obtain a first mixed audio corresponding to each client, and the corresponding first mixed audio is sent to each client;

[0015] The at least one first audio is mixed to obtain a second mixed audio corresponding to the conference server, and the second mixed audio is sent to the conference server.

[0016] In one possible implementation, for any client; the at least one first audio and the second audio are mixed to obtain a first mixed audio corresponding to the client, including:

[0017] Determine the first client identifier of the client, and the client identifier corresponding to each first audio;

[0018] Based on the first client identifier and the client identifier corresponding to each first audio, a target audio is determined from the at least one first audio, wherein the client identifier corresponding to the target audio is different from the first client identifier;

[0019] The target audio and the second audio are mixed to obtain the first mixed audio corresponding to the client.

[0020] In one possible implementation, receiving at least one first audio signal sent by the plurality of clients via a recording channel between the plurality of clients includes:

[0021] Obtain the channel status of multiple recording channels between the cloud desktop server and the multiple clients, wherein the channel status is either muted or unmute;

[0022] Based on the channel status of the plurality of recording channels, at least one target recording channel is determined among the plurality of recording channels, wherein the target recording channel is in a non-mute state;

[0023] The at least one first audio signal is received through the at least one target recording channel.

[0024] In one possible implementation, the method further includes:

[0025] Receive a mute request from the client, and set the recording channel status between the client and the cloud desktop server to mute according to the mute request; or,

[0026] The system receives a non-mute request from the client and, based on the non-mute request, sets the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

[0027] In one possible implementation, before receiving at least one first audio signal sent by the plurality of clients via a recording channel between the plurality of clients, the method further includes:

[0028] Receive a login request sent by a client, the login request including verification information;

[0029] The verification information is verified, and after the verification information is verified, a recording channel is created between the cloud desktop server and the client. The recording channel is a socket channel.

[0030] Thirdly, embodiments of this application provide an audio processing method applied to a client, wherein the client is equipped with a microphone, and the method includes:

[0031] The first audio signal is obtained by capturing audio through the microphone.

[0032] Determine the channel status of the recording channel between the client and the cloud desktop server, wherein the channel status is either muted or non-mute;

[0033] When the channel is in a non-mute state, the first audio is sent to the cloud desktop server through the recording channel.

[0034] In one possible implementation, the method further includes:

[0035] Display a first page, which includes the status control for the recording channel;

[0036] In response to a mute operation on the status control, a mute request is sent to the cloud desktop server; or, in response to a demute operation on the status control, a demute request is sent to the cloud desktop server.

[0037] In one possible implementation, the method further includes:

[0038] When the channel is in a mute state, the first audio is stored in a preset storage space or the first audio is discarded.

[0039] In one possible implementation, the method further includes:

[0040] The system receives a mixing result sent by the cloud desktop server, which is obtained by the cloud desktop server mixing audio from other clients and the conference server.

[0041] Fourthly, embodiments of this application provide an audio processing device applied to a cloud desktop server. The audio processing device includes: a first receiving module, a second receiving module, and a determining module, wherein...

[0042] The first receiving module is configured to receive at least one first audio signal sent by the plurality of clients through a recording channel between the plurality of clients;

[0043] The second receiving module is used to receive a second audio signal sent by the conference server, wherein the second audio signal is obtained by the conference server processing audio signals from at least one conference terminal;

[0044] The determining module is used to determine the mixed audio corresponding to each client and the conference terminal based on the at least one first audio and the second audio, and send the corresponding mixed audio to each client and the conference server.

[0045] In one possible implementation, the determining module is specifically used for:

[0046] The at least one first audio and the second audio are mixed to obtain a first mixed audio corresponding to each client, and the corresponding first mixed audio is sent to each client;

[0047] The at least one first audio is mixed to obtain a second mixed audio corresponding to the conference server, and the second mixed audio is sent to the conference server.

[0048] In one possible implementation, the determining module is specifically used for:

[0049] Determine the first client identifier of the client, and the client identifier corresponding to each first audio;

[0050] Based on the first client identifier and the client identifier corresponding to each first audio, a target audio is determined from the at least one first audio, wherein the client identifier corresponding to the target audio is different from the first client identifier;

[0051] The target audio and the second audio are mixed to obtain the first mixed audio corresponding to the client.

[0052] In one possible implementation, the determining module is specifically used for:

[0053] Obtain the channel status of multiple recording channels between the cloud desktop server and the multiple clients, wherein the channel status is either muted or unmute;

[0054] Based on the channel status of the plurality of recording channels, at least one target recording channel is determined among the plurality of recording channels, wherein the target recording channel is in a non-mute state;

[0055] The at least one first audio signal is received through the at least one target recording channel.

[0056] In one possible implementation, the first receiving module is further configured to:

[0057] Receive a mute request from the client, and set the recording channel status between the client and the cloud desktop server to mute according to the mute request; or,

[0058] The system receives a non-mute request from the client and, based on the non-mute request, sets the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

[0059] In one possible implementation, the audio processing device further includes a third receiving module and a creation module.

[0060] The third receiving module is used to receive a login request sent by the client, the login request including verification information;

[0061] The creation module is used to verify the verification information and, after the verification information is verified, create a recording channel between the cloud desktop server and the client. The recording channel is a socket channel.

[0062] Fifthly, embodiments of this application provide an audio processing device applied to a client, the audio processing device comprising: an acquisition module, a determination module, and a transmission module, wherein...

[0063] The acquisition module is used to acquire audio through the microphone to obtain a first audio signal;

[0064] The determining module is used to determine the channel status of the recording channel between the client and the cloud desktop server, wherein the channel status is either muted or non-mute.

[0065] The sending module is used to send the first audio to the cloud desktop server through the recording channel when the channel state is non-mute.

[0066] In one possible implementation, the audio processing device further includes a display module:

[0067] The display module is used to display a first page, the first page including the status control of the recording channel;

[0068] The sending module is used to send a mute request to the cloud desktop server in response to a mute operation on the status control; or, in response to a non-mute operation on the status control, send a non-mute request to the cloud desktop server.

[0069] In one possible implementation, the audio processing device further includes a storage module:

[0070] The storage module is used to store the first audio to a preset storage space or discard the first audio when the channel is in a mute state.

[0071] In one possible implementation, the audio processing device further includes a receiving module:

[0072] The receiving module is used to receive the mixing result sent by the cloud desktop server, the mixing result being obtained by the cloud desktop server mixing audio from other clients and the conference server.

[0073] Sixthly, embodiments of this application provide a cloud desktop server, including: a memory and a processor;

[0074] The memory stores computer-executed instructions;

[0075] The processor executes computer execution instructions stored in the memory, causing the processor to perform the audio processing method as described in either the first or second aspect.

[0076] In a seventh aspect, embodiments of this application provide a client, including: a memory, a processor, and a microphone;

[0077] The microphone is used to capture the first audio signal;

[0078] The memory stores computer-executed instructions;

[0079] The processor executes computer execution instructions stored in the memory, causing the processor to perform the audio processing method as described in any of the third aspects.

[0080] Eighthly, embodiments of this application provide an audio processing system, including the cloud desktop server described in the sixth aspect and at least one client described in the seventh aspect, the client being used to access the cloud desktop server.

[0081] Ninthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the audio processing method described in either the first or second aspect.

[0082] In a tenth aspect, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the audio processing method described in any of the third aspects.

[0083] Eleventhly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the audio processing method described in either the first or second aspect.

[0084] In a twelfth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the audio processing method described in any of the third aspects.

[0085] This application provides an audio processing method, a cloud desktop server, a client, and a system. The client can capture audio through a microphone to obtain a first audio value and determine the channel status of the recording channel between the client and the cloud desktop server. When the channel status is non-mute, the client can send the first audio value to the cloud desktop server through the recording channel, enabling the cloud desktop server to receive the first audio value. The cloud desktop server can receive a second audio value sent by a conference server. Based on at least one first audio value and a second audio value, the cloud desktop server can determine the mixed audio value corresponding to each client and conference terminal, and send the corresponding mixed audio value to each client and conference server. Since the cloud desktop server can receive at least one audio value sent by multiple clients through multiple recording channels, compared to receiving audio sent by one client through a single recording channel, the user experience is improved. Attached Figure Description

[0086] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0087] Figure 1 A schematic diagram illustrating an application scenario provided for an exemplary embodiment of this application;

[0088] Figure 2 A flowchart illustrating an audio processing method provided for an exemplary embodiment of this application;

[0089] Figure 3A flowchart illustrating another audio processing method provided for an exemplary embodiment of this application;

[0090] Figure 4 A schematic diagram of an audio processing system provided for an exemplary embodiment of this application. Figure 1 ;

[0091] Figure 5A A schematic diagram of the first page provided for an exemplary embodiment of this application. Figure 1 ;

[0092] Figure 5B A schematic diagram of the first page provided for an exemplary embodiment of this application. Figure 2 ;

[0093] Figure 6 A schematic diagram of an audio processing system provided for an exemplary embodiment of this application. Figure 2 ;

[0094] Figure 7 A schematic diagram of the structure of an audio processing apparatus provided for an exemplary embodiment of this application;

[0095] Figure 8 A schematic diagram of another audio processing apparatus provided as an exemplary embodiment of this application;

[0096] Figure 9 A schematic diagram of the structure of a cloud desktop server provided as an exemplary embodiment of this application;

[0097] Figure 10 A schematic diagram of the structure of a client provided for an exemplary embodiment of this application. Detailed Implementation

[0098] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0099] Figure 1 This is a schematic diagram illustrating an application scenario provided for an exemplary embodiment of this application. Please refer to [link / reference]. Figure 1 This includes multiple clients, a cloud desktop server, and a conferencing server. For example, multiple clients can include client 1, client 2, ..., client n. The cloud desktop server and any one of the clients can communicate with each other. The cloud desktop server and the conferencing server can also communicate with each other.

[0100] For any given client, there can be a recording channel between the client and the cloud desktop server. The client can send the first audio file to the cloud desktop server through the corresponding recording channel. For example, client 1 can send the first audio file 1 to the cloud desktop server through recording channel 1, client 2 can send the first audio file 2 to the cloud desktop server through recording channel 2, and so on, client n can send the first audio file n to the cloud desktop server through recording channel n.

[0101] The conference server can send a second audio message to the cloud desktop server. This second audio message can be obtained by the conference server processing audio from at least one conference terminal.

[0102] After receiving at least one first audio and one second audio, the cloud desktop server can determine the mixed audio corresponding to each client and conference terminal based on the first and second audio, and send the corresponding mixed audio to each client and conference server. For example, the cloud desktop server can determine the mixed audio 1 corresponding to client 1, the mixed audio 2 corresponding to client 2, ..., the mixed audio n corresponding to client n, and the mixed audio a corresponding to the conference server based on the first and second audio, and send the corresponding mixed audio to each client and conference server.

[0103] In related technologies, multiple clients can log in to a cloud desktop server. However, because there is usually a recording channel between the cloud desktop server and the clients, only one client's audio can be captured at a time through the recording channel, making it impossible to capture the audio of multiple clients simultaneously, resulting in a poor user experience.

[0104] In this embodiment, each client can have a recording channel with the cloud desktop server. For any given client, the cloud desktop server can receive the first audio file sent by the client through the client's corresponding recording channel. Since the cloud desktop server can receive at least one first audio file sent by multiple clients through multiple recording channels, compared to receiving audio sent by only one client through a single recording channel, the user experience is improved.

[0105] The technical solutions shown in this application will now be described in detail through specific embodiments. It should be noted that the following embodiments may exist independently or in combination with each other; for identical or similar content, the description will not be repeated in different embodiments.

[0106] Figure 2 This is a flowchart illustrating an audio processing method provided for an exemplary embodiment of this application. Please refer to [link / reference]. Figure 2 The method may include:

[0107] S201. The client acquires audio through the microphone and obtains the first audio.

[0108] The client can be a lightweight electronic device with a display screen and microphone. For example, the client can be a mobile phone or a television.

[0109] Users can input voice in the client, and the client can then capture the audio through the microphone to obtain the first audio.

[0110] For example, a user can input the voice message "The weather is nice today" in the client. The client can then capture the audio through the microphone to obtain the first audio, which may include the voice data of "The weather is nice today".

[0111] S202. The client determines the channel status of the recording channel between the client and the cloud desktop server.

[0112] A recording channel is a channel for transmitting audio between the client and the cloud desktop server. Each client can have a recording channel with the cloud desktop server.

[0113] The channel status of the recording channel can be either mute or non-mute.

[0114] Optionally, the client can determine the channel status of the recording channel between the client and the cloud desktop server. When the recording channel status is muted, the client can either store the first audio file in a preset storage space or discard the first audio file.

[0115] For example, if there are two clients, Client 1 and Client 2, and the recording channel between Client 1 and the cloud desktop server is Recording Channel 1, while the recording channel between Client 2 and the cloud desktop server is Recording Channel 2, then Client 1 can determine that the channel status of Recording Channel 1 is muted, and Client 2 can determine that the channel status of Recording Channel 2 is not muted. If Client 1 captures the first audio 1, since Recording Channel 1 is muted, Client 1 can either store the first audio 1 in its preset storage space or discard it.

[0116] S203. When the channel is in a non-mute state, the client sends the first audio to the cloud desktop server through the recording channel.

[0117] Optionally, multiple clients can send at least one first audio file to the cloud desktop server through corresponding recording channels, so that the cloud desktop server can receive at least one first audio file.

[0118] For example, if there are two clients, Client 1 and Client 2, and the recording channel between Client 1 and the cloud desktop server is Recording Channel 1, and the recording channel between Client 2 and the cloud desktop server is Recording Channel 2, if Client 1 captures the first audio 1 and Client 2 captures the first audio 2, and both Recording Channel 1 and Recording Channel 2 are in a non-mute state, then Client 1 can send the first audio 1 to the cloud desktop server through Recording Channel 1, and Client 2 can send the first audio 2 to the cloud desktop server through Recording Channel 2.

[0119] S204, The cloud desktop server receives the second audio sent by the conference server.

[0120] A conference server can be used to process audio from conference terminals.

[0121] The conferencing terminal can be an electronic device. For example, a conferencing terminal can be a computer.

[0122] The second audio can be the audio obtained by the conference server processing audio from at least one conference terminal.

[0123] In one optional embodiment, the conference server can receive audio sent by at least one conference terminal and process the audio to obtain a second audio. The conference server can then send the second audio to a cloud desktop server, enabling the cloud desktop server to receive the second audio from the conference server.

[0124] For example, User 1 can input the voice message "The weather is nice today" in Conference Terminal 1, causing Conference Terminal 1 to capture Audio 1; User 2 can input the voice message "Yes, the sun is shining brightly" in Conference Terminal 2, causing Conference Terminal 2 to capture Audio 2. Conference Terminal 1 can send Audio 1 to the conference server, and Conference Terminal 2 can send Audio 2 to the conference server. After receiving Audio 1 and Audio 2, the conference server can process them to obtain a second audio message. The second audio message may include "The weather is nice today" and "Yes, the sun is shining brightly." The conference server can then send this second audio message to the cloud desktop server, allowing the cloud desktop server to receive it.

[0125] S205. The cloud desktop server determines the mixed audio corresponding to each client and conference terminal based on at least one first audio and one second audio, and sends the corresponding mixed audio to each client and conference server.

[0126] Mixed audio refers to audio obtained by processing at least one first audio and one second audio. Mixed audio may include at least one first audio and one second audio. For example, if there are first audio 1 and first audio 2, and the second audio includes audio 1 and audio 2, then mixed audio 1 may include first audio 2, audio 1, and audio 2.

[0127] In an optional embodiment, the mixed audio corresponding to each client and conference terminal can be determined and sent to each client and conference server in the following manner: at least one first audio and a second audio are mixed to obtain a first mixed audio corresponding to each client, and the corresponding first mixed audio is sent to each client; at least one first audio is mixed to obtain a second mixed audio corresponding to the conference server, and the second mixed audio is sent to the conference server.

[0128] For example, if there are two clients, Client 1 and Client 2, and if the cloud desktop server can receive a first audio 1 sent by Client 1 and a first audio 2 sent by Client 2, and the cloud desktop server can receive a second audio 2 sent by the conference server, which includes audio 1 and audio 2, then the cloud desktop server can perform audio mixing processing on these two audio 1 and audio 2. It can determine that the first mixed audio 1 corresponding to Client 1 includes first audio 2, audio 1, and audio 2, and send this first mixed audio 1 to Client 1; it can also determine that the first mixed audio 2 corresponding to Client 2 includes first audio 1, audio 1, and audio 2, and send this first mixed audio 2 to Client 2. The cloud desktop server can also perform audio mixing processing on these two audio 1 to determine that the second mixed audio corresponding to the conference server includes first audio 1 and first audio 2, and send this second mixed audio to the conference server.

[0129] In this embodiment, the client can capture audio through a microphone to obtain a first audio signal and determine the channel status of the recording channel between the client and the cloud desktop server. When the channel status is non-mute, the client can send the first audio signal to the cloud desktop server through the recording channel, enabling the cloud desktop server to receive the first audio signal. The cloud desktop server can receive a second audio signal sent by the conference server. Based on at least one first audio signal and one second audio signal, the cloud desktop server can determine the mixed audio signal corresponding to each client and conference terminal, and send the corresponding mixed audio signal to each client and conference server. Since the cloud desktop server can receive at least one audio signal sent by multiple clients through multiple recording channels, compared to receiving audio signal sent by one client through a single recording channel, the user experience is improved.

[0130] Below, in Figure 2 Based on the illustrated embodiments, combined with Figure 3The above audio processing method will be further explained in detail.

[0131] Figure 3 A flowchart illustrating another audio processing method provided for an exemplary embodiment of this application. Please refer to... Figure 3 The method may include:

[0132] S301, The client sends a login request to the cloud desktop server.

[0133] For any client, a login request can be sent to the cloud desktop server. The login request may include verification information, which may include the user's username and password.

[0134] For example, if there are two clients, client 1 and client 2, client 1 can send login request 1 to the cloud desktop server. Login request 1 can include verification information 1, namely user account 1 and password 1. Client 2 can send login request 2 to the cloud desktop server. Login request 2 can include verification information 2, namely user account 2 and password 2.

[0135] S302. The cloud desktop server verifies the verification information and, after the verification information is verified, creates a recording channel between the cloud desktop server and the client.

[0136] Alternatively, the recording channel can be a socket channel.

[0137] In one optional embodiment, the cloud desktop server can determine the user account and its corresponding preset password. If the verification information includes the user account and password, the cloud desktop server can perform verification processing on the user account and password, that is, the cloud desktop server can determine whether the preset password corresponding to the user account matches the password in the verification information. If the preset password corresponding to the user account matches the password in the verification information, the verification passes; if the preset password corresponding to the user account does not match the password in the verification information, the verification fails.

[0138] For example, if client 1 sends a login request 1 to the cloud desktop server that includes verification information 1, which includes user account 1 and password 1, and if the cloud desktop server determines that the preset password corresponding to user account 1 is password 1, then the verification will pass because the preset password corresponding to user account 1 is the same as the password 1 in verification information 1; if the cloud desktop server determines that the preset password corresponding to user account 1 is password 2, then the verification will fail because the preset password corresponding to user account 1 is different from the password 1 in verification information 1.

[0139] Optionally, after the cloud desktop server verifies the authentication information, it can create a recording channel between the cloud desktop server and the client. For example, if client 1 sends login request 1 to the cloud desktop server, which includes authentication information 1, and the cloud desktop server verifies authentication information 1, then the cloud desktop server can create a recording channel 1 between itself and the client. This recording channel 1 can be a socket channel.

[0140] S303: The client acquires audio through the microphone to obtain the first audio.

[0141] After successfully logging into the cloud desktop server, the user can input voice in the client, and the client can then capture the audio through the microphone to obtain the first audio.

[0142] For example, user 1 can input the voice "The weather is nice today" in client 1. Then client 1 can collect the audio through the microphone to obtain the first audio 1, which may include the voice data of "The weather is nice today".

[0143] S304. The client determines the channel status of the recording channel between the client and the cloud desktop server.

[0144] It should be noted that the execution process of step S304 can be referred to the execution process of step S202, and will not be repeated here.

[0145] S305. When the channel is in a non-mute state, the client sends the first audio to the cloud desktop server through the recording channel.

[0146] It should be noted that the execution process of step S305 can be referred to the execution process of step S203, and will not be repeated here.

[0147] S306, The conference server sends a second audio message to the cloud desktop server.

[0148] Optionally, the conference server can receive audio sent by at least one conference terminal, process the audio, and obtain a second audio. The conference server can then send this second audio to the cloud desktop server, enabling the cloud desktop server to receive the second audio from the conference server.

[0149] For example, conference terminal 1 can send audio 1 to the conference server, and conference terminal 2 can send audio 2 to the conference server. After receiving audio 1 and audio 2, the conference server can process audio 1 and audio 2 to obtain a second audio, which may include audio 1 and audio 2. The conference server can then send this second audio to a cloud desktop server so that the cloud desktop server can receive the second audio.

[0150] S307, the cloud desktop server performs mixing processing on at least one first audio and one second audio to obtain the first mixed audio corresponding to each client.

[0151] For any client, at least one first audio and a second audio can be mixed to obtain the first mixed audio corresponding to the client in the following way: determine the first client identifier of the client and the client identifier corresponding to each first audio; determine the target audio in at least one first audio based on the first client identifier and the client identifier corresponding to each first audio; mix the target audio and the second audio to obtain the first mixed audio corresponding to the client.

[0152] The cloud desktop server can identify the first client among multiple clients. For example, if there are three clients, namely client 1, client 2, and client 3, and the first client is client 1, then the cloud desktop server can identify the first client as client 1.

[0153] For any given first audio file, it can have a corresponding client identifier. The cloud desktop server can determine the client identifier corresponding to the first audio file. For example, if client 1 sends first audio file 1 to the cloud desktop server through recording channel 1, the cloud desktop server can determine that the client identifier corresponding to first audio file 1 is client 1.

[0154] Optionally, the cloud desktop server can receive at least one first audio message sent by at least one client, and determine the target audio message from the at least one first audio message based on the first client identifier and the client identifier corresponding to each first audio message. The client identifier corresponding to the target audio message may be different from the first client identifier.

[0155] For example, if the first client identifier is Client 1, and the cloud desktop server can receive the first audio 1 sent by Client 1, the first audio 2 sent by Client 2, and the first audio 3 sent by Client 3, then the cloud desktop server can determine that the client identifier corresponding to the first audio 1 is Client 1, the client identifier corresponding to the first audio 2 is Client 2, and the client identifier corresponding to the first audio 3 is Client 3. Since the first client identifier is Client 1, the cloud desktop server can determine that the target audio 1 includes both the first audio 2 and the first audio 3 among the first audio 1, first audio 2, and first audio 3. The client identifiers corresponding to the target audio 1 include Client 2 and Client 3, which are different from the first client identifier, i.e., Client 1.

[0156] Optionally, after the cloud desktop server determines the target audio, it can perform a mixing process on the target audio and the second audio to obtain the first mixed audio corresponding to the client.

[0157] For example, if the cloud desktop server receives a first audio 1 sent by client 1, a first audio 2 sent by client 2, and a first audio 3 sent by client 3, and if the cloud desktop server receives a second audio sent by the conference server, the second audio including audio 1 and audio 2, and if the first client is identified as client 1, the cloud desktop server determines that the target audio 1 includes first audio 2 and first audio 3. Therefore, the cloud desktop server can perform a mixing process on the target audio 1 and the second audio to obtain a first mixed audio 1. That is, the cloud desktop server can determine that the first mixed audio 1 includes first audio 2, first audio 3, audio 1, and audio 2. Similarly, if the first client is identified as client 2, the cloud desktop server determines that the target audio 2 includes first audio 1 and first audio 3. Therefore, the cloud desktop server can perform a mixing process on the target audio 2 and the second audio to obtain a first mixed audio 2. That is, the cloud desktop server can determine that the first mixed audio 2 includes first audio 1, first audio 3, audio 1, and audio 2.

[0158] S308, the cloud desktop server sends the corresponding first mixed audio to each client.

[0159] For example, if the cloud desktop server determines that the first mixed audio 1 includes the first audio 2, the first audio 3, the audio 1 and the audio 2, then the cloud desktop server can send the first mixed audio 1 to client 1; if the cloud desktop server determines that the first mixed audio 2 includes the first audio 1, the first audio 3, the audio 1 and the audio 2, then the cloud desktop server can send the first mixed audio 2 to client 2.

[0160] S309. The cloud desktop server performs mixing processing on at least one first audio to obtain a second mixed audio corresponding to the conference server.

[0161] For example, if the cloud desktop server receives a first audio 1 sent by client 1, a first audio 2 sent by client 2, and a first audio 3 sent by client 3, the cloud desktop server can mix these three first audios to obtain a second mixed audio. That is, the cloud desktop server can determine that the second mixed audio includes the first audio 1, the first audio 2, and the first audio 3.

[0162] The S310 cloud desktop server sends the second mixed audio to the conference server.

[0163] For example, if the cloud desktop server determines that the second mixed audio includes the first audio 1, the first audio 2, and the first audio 3, then the cloud desktop server sends the second mixed audio to the conference server.

[0164] Optionally, after receiving the second mixed audio, the conference server can perform mixing processing on the second mixed audio and the second audio to determine the mixed audio corresponding to each conference terminal.

[0165] For example, if there are two conference terminals, conference terminal 1 and conference terminal 2, and the conference server can receive audio 1 sent by conference terminal 1 and audio 2 sent by conference terminal 2, then the second audio includes audio 1 and audio 2. If the second mixed audio includes first audio 1, first audio 2, and first audio 3, then the conference server can perform mixing processing on the second mixed audio and the second audio to determine the mixed audio corresponding to each conference terminal. That is, the conference server can determine that the mixed audio corresponding to conference terminal 1 includes first audio 1, first audio 2, first audio 3, and audio 2; and it can determine that the mixed audio corresponding to conference terminal 2 includes first audio 1, first audio 2, first audio 3, and audio 1.

[0166] In this embodiment, the client can send a login request to the cloud desktop server. The cloud desktop server can verify the authentication information and, after successful verification, create a recording channel between the cloud desktop server and the client. The client can capture audio through a microphone to obtain a first audio value and determine the channel status of the recording channel between the client and the cloud desktop server. When the channel status is non-mute, the client can send the first audio value to the cloud desktop server through the recording channel. The conference server can send a second audio value to the cloud desktop server. The cloud desktop server can mix at least one first audio value and the second audio value to obtain a first mixed audio value corresponding to each client and send the corresponding first mixed audio value to each client; the cloud desktop server can also mix at least one first audio value to obtain a second mixed audio value corresponding to the conference server and send the second mixed audio value to the conference server. Since the cloud desktop server can receive at least one audio value sent by multiple clients through multiple recording channels, compared to receiving audio sent by one client through a single recording channel, the user experience is improved.

[0167] Below, based on any of the above embodiments, combined with Figure 4 This provides an audio processing system.

[0168] Figure 4 A schematic diagram of an audio processing system provided for an exemplary embodiment of this application. Figure 1 Please see. Figure 4 It includes a cloud desktop server and at least one client. For example, the at least one client can be client 1, client 2, ..., client n.

[0169] For any client, a client can send a login request to the cloud desktop server, which may include verification information. Upon receiving the login request, the cloud desktop server can verify the verification information. If the verification is successful, the cloud desktop server can create a recording channel with the client; this recording channel can be a socket channel. For example, client 1 can send login request 1 to the cloud desktop server, which may include verification information 1. Upon receiving login request 1, the cloud desktop server can verify the verification information 1. If the verification is successful, the cloud desktop server can create a recording channel 1 with client 1.

[0170] Once a client successfully logs into the cloud desktop server, it can participate in the meeting through the cloud desktop server.

[0171] Optionally, during the meeting, the client can display a first page, which includes status controls for the recording channel.

[0172] Below, in conjunction with Figure 5A and Figure 5B This section provides an explanation of the first page.

[0173] Figure 5A A schematic diagram of the first page provided for an exemplary embodiment of this application. Figure 1 Please see Figure 5A The client can display a first page, which may include status controls for the recording channel. Users can mute these status controls on the client, such as... Figure 5A As shown in the diagram, the client can send a mute request to the cloud desktop server in response to a mute operation on the status control. The cloud desktop server can receive the mute request from the client and, based on the mute request, set the channel status of the recording channel between the client and the cloud desktop server to the mute state.

[0174] Figure 5B A schematic diagram of the first page provided for an exemplary embodiment of this application. Figure 2 Please see Figure 5B The client can display a first page, which may include status controls for the recording channel. Users can perform de-mute operations on these status controls on the client, such as... Figure 5B As shown in the diagram, the client can send a non-mute request to the cloud desktop server in response to a non-mute operation on the status control. The cloud desktop server can receive the client's non-mute request and, based on the request, set the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

[0175] For example, client 1 can respond to user 1's non-mute operation on the status control by sending a non-mute request 1 to the cloud desktop server. Upon receiving non-mute request 1, the cloud desktop server can set the channel status of recording channel 1 between client 1 and the cloud desktop server to a non-mute state. Client 2 can respond to user 2's mute operation on the status control by sending a mute request 2 to the cloud desktop server. Upon receiving mute request 2, the cloud desktop server can set the channel status of recording channel 2 between client 2 and the cloud desktop server to a mute state; ...; client n can respond to user n's non-mute operation on the status control by sending a non-mute request n to the cloud desktop server. Upon receiving non-mute request n, the cloud desktop server can set the channel status of recording channel n between client n and the cloud desktop server to a non-mute state.

[0176] Each client can be equipped with a microphone. During the meeting, the client can capture audio through the microphone to obtain the initial audio.

[0177] It should be noted that during audio capture, the client's microphone can remain on. The cloud desktop server can configure the recording channel's state, allowing it to switch between mute and non-mute modes to meet user needs. Because the client's microphone is always on, there's no need to turn it on when switching from mute to non-mute mode, reducing recording latency on the client side.

[0178] Since the cloud desktop server can set the recording channel status between a client and the cloud desktop server to a mute state based on a mute request sent by any client, or set the recording channel status between a client and the cloud desktop server to a non-mute state based on a non-mute request sent by any client, the recording channel status between multiple clients and the cloud desktop server can include a mute state and a non-mute state.

[0179] Optionally, the cloud desktop server can obtain the channel status of multiple recording channels with multiple clients, where the channel status is either muted or non-mute; based on the channel status of the multiple recording channels, it can determine at least one target recording channel among the multiple recording channels, where the target recording channel is non-mute; and receive at least one first audio signal through the at least one target recording channel. When the channel status of a recording channel is muted, the client can store the first audio signal in a preset storage space or discard the first audio signal.

[0180] For example, if the recording channel 1 between client 1 and the cloud desktop server is in a non-mute state, the recording channel 2 between client 2 and the cloud desktop server is in a mute state, and so on, and the recording channel n between client n and the cloud desktop server is in a non-mute state, then since the recording channels 1, ..., and recording channel n are in a non-mute state, the cloud desktop server can determine that the target recording channels include recording channels 1, ..., and recording channel n. The cloud desktop server can receive the first audio 1 sent by client 1 through recording channel 1, ..., and can receive the first audio n sent by client n through recording channel n. Since the recording channel 2 is in a mute state, client 2 can store the first audio 2 in a preset storage space or discard the first audio 2 without sending the first audio 2 to the cloud desktop server.

[0181] Because the cloud desktop server can determine the channel status of each recording channel before obtaining the first audio, it can obtain at least one first audio through the non-mute recording channel, thus reducing power consumption.

[0182] Optionally, the cloud desktop server can also receive a second audio message sent by the conference server. This second audio message can be obtained by the conference server processing audio from at least one conference terminal.

[0183] After receiving at least one first audio and one second audio, the cloud desktop server can perform mixing processing on the at least one first audio and one second audio to obtain the first mixed audio corresponding to each client, and send the corresponding first mixed audio to each client.

[0184] Specifically, the cloud desktop server can determine the first client identifier of the client and the client identifier corresponding to each first audio, and determine the target audio in at least one first audio based on the first client identifier and the client identifier corresponding to each first audio. The cloud desktop server can perform audio mixing processing on the target audio and the second audio to obtain the first mixed audio corresponding to the client, and send the corresponding first mixed audio to the client so that the client can receive the mixing result sent by the cloud desktop server.

[0185] For example, if the cloud desktop server can receive first audio 1 sent by client 1, ..., first audio n sent by client n, and if the cloud desktop server receives second audio sent by the conference server, the second audio including audio 1 and audio 2, then the cloud desktop server can determine that the client identifier corresponding to the first audio 1 is client 1, ..., and the client identifier corresponding to the first audio n is client n. If the first client identifier is client 1, the cloud desktop server can determine the target audio 1 from the first audio 1, ..., first audio n, and perform audio mixing on the target audio 1 and the second audio to obtain the first mixed audio 1, and send the first mixed audio 1 to client 1; if the first client identifier is client 2, the cloud desktop server determines the target audio 2, and performs audio mixing on the target audio 2 and the second audio to obtain the first mixed audio 2, and sends the first mixed audio 2 to client 2; ...; if the first client identifier is client n, the cloud desktop server determines the target audio n, and performs audio mixing on the target audio n and the second audio to obtain the first mixed audio n, and sends the first mixed audio n to client n.

[0186] Optionally, the cloud desktop server can also perform mixing processing on at least one first audio to obtain a second mixed audio corresponding to the conference server, and send the second mixed audio to the conference server.

[0187] For example, if the cloud desktop server receives the first audio 1 sent by client 1, ..., the first audio n sent by client n, the cloud desktop server can perform mixing processing on the first audio 1, ..., the first audio n to obtain the second mixed audio, and send the second mixed audio to the conference server.

[0188] In this embodiment, the client can send a login request to the cloud desktop server. The cloud desktop server can verify the authentication information and, after successful verification, create a recording channel between the cloud desktop server and the client. The client can capture audio through a microphone to obtain a first audio signal and determine the channel status of the recording channel between the client and the cloud desktop server. When the channel status is non-mute, the cloud desktop server can receive the first audio signal sent by the client through the recording channel. The conference server can send a second audio signal to the cloud desktop server. The cloud desktop server can mix at least one first audio signal and the second audio signal to obtain a first mixed audio signal corresponding to each client and send the corresponding first mixed audio signal to each client; the cloud desktop server can also mix at least one first audio signal to obtain a second mixed audio signal corresponding to the conference server and send the second mixed audio signal to the conference server. Since the cloud desktop server can receive at least one audio signal sent by multiple clients through multiple recording channels, compared to receiving audio sent by only one client through a single recording channel, the user experience is improved.

[0189] Below, based on any of the above embodiments, combined with Figure 6 The audio processing system will be described in further detail.

[0190] Figure 6 A schematic diagram of an audio processing system provided for an exemplary embodiment of this application. Figure 2 Please see. Figure 6 It includes at least one client, a cloud desktop server, a conference server, and at least one conference terminal. For example, at least one client can be client 1, client 2, ..., client n; at least one conference terminal can be conference terminal 1, conference terminal 2, ..., conference terminal m.

[0191] For any client, a login request can be sent to the cloud desktop server, and the login request may include verification information. After receiving the login request, the cloud desktop server can verify the verification information in the login request. If the verification is successful, the cloud desktop server can create a recording channel with the client, which can be a socket channel. For example, the cloud desktop server can establish recording channel 1 for client 1, recording channel 2 for client 2, ..., recording channel n for client n.

[0192] Once a client successfully logs into the cloud desktop server, it can participate in the meeting through the cloud desktop server.

[0193] Optionally, during the meeting, the client can display a first page, which includes a status control for the recording channel. The client can respond to the user's mute action on the status control by sending a mute request to the cloud desktop server; or, in response to the user's unmute action on the status control, send a non-mute request to the cloud desktop server.

[0194] The cloud desktop server can receive a mute request from the client and, based on the mute request, set the channel status of the recording channel between the client and the cloud desktop server to mute; or, the cloud desktop server can receive a non-mute request from the client and, based on the non-mute request, set the channel status of the recording channel between the client and the cloud desktop server to non-mute.

[0195] For example, a cloud desktop server can set recording channel 1 to non-mute, recording channel 2 to mute, ..., and recording channel n to non-mute.

[0196] Each client can be equipped with a microphone. During the meeting, the client can capture audio through the microphone to obtain the initial audio.

[0197] The cloud desktop server can obtain the channel status of multiple recording channels with multiple clients, and determine at least one target recording channel among the multiple recording channels based on the channel status of the multiple recording channels. Then, it can receive at least one first audio signal sent by the client through at least one target recording channel.

[0198] For example, if the cloud desktop server can determine that recording channel 1 is in a non-mute state, recording channel 2 is in a mute state, and recording channel n is in a non-mute state, then the cloud desktop server can receive the first audio 1 sent by client 1 through recording channel 1, and so on, and can receive the first audio n sent by client n through recording channel n. Since the recording channel 2 is in a mute state, client 2 can store the first audio 2 in a preset storage space or discard the first audio without sending the first audio 2 to the cloud desktop server.

[0199] Optionally, the conference server can receive at least one audio message sent by at least one conference terminal, and process the at least one audio message to obtain a second audio message. The conference server can send the second audio message to the cloud desktop server so that the cloud desktop server can receive the second audio message.

[0200] For example, the conference server can receive audio 1 sent by conference terminal 1, audio 2 sent by conference terminal 2, ..., audio m sent by conference terminal m. The conference server can process these m audio files to obtain a second audio file, which may include audio 1, audio 2, ..., audio m. The conference server can then send this second audio file to a cloud desktop server so that the cloud desktop server can receive it.

[0201] After receiving at least one first audio and one second audio, the cloud desktop server can perform mixing processing on the at least one first audio and one second audio to obtain the first mixed audio corresponding to each client, and send the corresponding first mixed audio to each client.

[0202] For example, a cloud desktop server can perform mixing processing on at least one first audio and one second audio to obtain first mixed audio 1 corresponding to client 1, first mixed audio 2 corresponding to client 2, ..., first mixed audio n corresponding to client n, and send the corresponding first mixed audio to each client.

[0203] Optionally, the cloud desktop server can perform mixing processing on at least one first audio to obtain a second mixed audio corresponding to the conference server, and send the second mixed audio to the conference server.

[0204] Optionally, after receiving the second mixed audio, the conference server can perform mixing processing on the second mixed audio and the second audio, determine the mixed audio corresponding to each conference terminal, and send the corresponding mixed audio to each conference terminal.

[0205] For example, after receiving the second mixed audio, the conference server can perform mixing processing on the second mixed audio and the second audio to obtain mixed audio 1 corresponding to conference terminal 1, mixed audio 2 corresponding to conference terminal 2, ..., mixed audio m corresponding to conference terminal m, and send the corresponding mixed audio to each conference terminal.

[0206] In this embodiment, the client can send a login request to the cloud desktop server. The cloud desktop server can verify the authentication information and, after successful verification, create a recording channel between the cloud desktop server and the client. The client can capture audio through a microphone to obtain a first audio signal and determine the channel status of the recording channel between the client and the cloud desktop server. When the channel status is non-mute, the cloud desktop server can receive the first audio signal sent by the client through the recording channel. The conference server can send a second audio signal to the cloud desktop server. The cloud desktop server can mix at least one first audio signal and the second audio signal to obtain a first mixed audio signal corresponding to each client and send the corresponding first mixed audio signal to each client; the cloud desktop server can also mix at least one first audio signal to obtain a second mixed audio signal corresponding to the conference server and send the second mixed audio signal to the conference server. Since the cloud desktop server can receive at least one audio signal sent by multiple clients through multiple recording channels, compared to receiving audio sent by only one client through a single recording channel, the user experience is improved.

[0207] Figure 7 A schematic diagram of an audio processing apparatus provided as an exemplary embodiment of this application can be applied to a cloud desktop server. Please refer to [link / reference]. Figure 7 The audio processing device 10 includes: a first receiving module 11, a second receiving module 12, and a determining module 13, wherein,

[0208] The first receiving module 11 is configured to receive at least one first audio signal sent by the plurality of clients through a recording channel between the plurality of clients;

[0209] The second receiving module 12 is used to receive a second audio signal sent by the conference server, wherein the second audio signal is obtained by the conference server processing audio signals from at least one conference terminal.

[0210] The determining module 13 is used to determine the mixed audio corresponding to each client and the conference terminal based on the at least one first audio and the second audio, and send the corresponding mixed audio to each client and the conference server.

[0211] The audio processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0212] In one possible implementation, the determining module 13 is specifically used for:

[0213] The at least one first audio and the second audio are mixed to obtain a fifth mixed audio corresponding to each client, and the corresponding first mixed audio is sent to each client;

[0214] The at least one first audio is mixed to obtain a second mixed audio corresponding to the conference server, and the second mixed audio is sent to the conference server.

[0215] In one possible implementation, the determining module 13 is specifically used for:

[0216] Determine the first client identifier of the client and the client identifier corresponding to each first audio; 0 Based on the first client identifier and the client identifier corresponding to each first audio, in the at least one first

[0217] The target audio is identified from the audio, and the client identifier corresponding to the target audio is different from the first client identifier;

[0218] The target audio and the second audio are mixed to obtain the first mixed audio corresponding to the client.

[0219] In one possible implementation, the determining module 13 is specifically used for:

[0220] Obtain the channel status of multiple recording channels between the cloud desktop server and the multiple clients, wherein the channel status is either silent or non-silent.

[0221] Based on the channel status of the plurality of recording channels, at least one target recording channel is determined among the plurality of recording channels, wherein the target recording channel is in a non-mute state;

[0222] The at least one first audio signal is received through the at least one target recording channel.

[0223] In one possible implementation, the first receiving module 11 is further configured to: receive a mute request sent by the client, and, based on the mute request, connect the client to the cloud desktop service.

[0224] The recording channel between servers is set to mute; or,

[0225] The system receives a non-mute request from the client and, based on the non-mute request, sets the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

[0226] In one possible implementation, the audio processing device further includes a third receiving module 14 and a creation module 15, wherein the third receiving module 14 is used to receive a login request sent by a client, the login request including verification information;

[0227] The creation module 15 is used to verify the verification information and, after the verification information is verified, create a recording channel between the cloud desktop server and the client, wherein the recording channel is a socket channel.

[0228] The audio processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0229] 0 Figure 8 A schematic diagram of another audio processing apparatus provided as an exemplary embodiment of this application can be applied to customers.

[0230] For the end, please see Figure 8 The audio processing device 20 includes: a collection module 21, a determination module 22, and a transmission module 23, wherein...

[0231] The acquisition module 21 is used to acquire audio through the microphone to obtain a first audio signal;

[0232] The determining module 22 is used to determine the channel status of the recording channel between the client and the cloud desktop server, wherein the channel status is either a silent state or a non-silent state.

[0233] The sending module 23 is used to send the first audio to the cloud desktop server through the recording channel when the channel state is non-mute.

[0234] The audio processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0235] In one possible implementation, the audio processing device further includes a display module 24.

[0236] The display module 24 is used to display a first page, the first page including the status control of the recording channel;

[0237] The sending module 23 is used to send a mute request to the cloud desktop server in response to a mute operation on the status control; or, in response to a non-mute operation on the status control, send a non-mute request to the cloud desktop server.

[0238] In one possible implementation, the audio processing device further includes a storage module 25.

[0239] The storage module 25 is used to store the first audio to a preset storage space or discard the first audio when the channel state is muted.

[0240] In one possible implementation, the audio processing device further includes a receiving module 26.

[0241] The receiving module 26 is used to receive the mixing result sent by the cloud desktop server, wherein the mixing result is obtained by the cloud desktop server mixing audio from other clients and the conference server.

[0242] The audio processing apparatus provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0243] Figure 9 A schematic diagram of the structure of a cloud desktop server provided as an exemplary embodiment of this application can be found in [link to schematic diagram]. Figure 9 The cloud desktop server 30 may include a processor 31 and a memory 32. For example, the processor 31 and the memory 32 are interconnected via a bus 33.

[0244] The memory 32 stores computer-executed instructions;

[0245] The processor 31 executes the computer execution instructions stored in the memory 32, causing the processor 31 to perform the audio processing method as shown in the above method embodiment.

[0246] Figure 10 For a schematic diagram of a client structure provided as an exemplary embodiment of this application, please refer to [link / reference]. Figure 10 The client 40 may include a processor 41, a memory 42, and a microphone 43. Exemplarily, the processor 41, the memory 42, and the microphone 43 are interconnected via a bus 44.

[0247] The microphone 43 is used to collect the first audio signal;

[0248] The memory 42 stores computer-executed instructions;

[0249] The processor 41 executes the computer execution instructions stored in the memory 42, causing the processor 41 to perform the audio processing method as shown in the above method embodiment.

[0250] Accordingly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the audio processing method described in the above method embodiments.

[0251] Accordingly, embodiments of this application may also provide a computer program product, including a computer program, which, when executed by a processor, can implement the audio processing method shown in the above method embodiments.

[0252] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0253] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0254] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0255] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0256] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0257] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0258] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0259] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0260] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. An audio processing method, characterized in that, Applied to a cloud desktop server, multiple clients access the cloud desktop server, and each client has a recording channel with the cloud desktop server; the method includes: Obtain the channel status of multiple recording channels between the cloud desktop server and the multiple clients, wherein the channel status is either muted or unmute; Based on the channel status of the plurality of recording channels, at least one target recording channel is determined among the plurality of recording channels, wherein the target recording channel is in a non-mute state; Through the at least one target recording channel, at least one first audio signal sent by the plurality of clients is received, wherein when the channel state of the recording channel is muted, the client stores the first audio signal in a preset storage space or discards the first audio signal; Receive a second audio message sent by the conference server, the second audio message being obtained by the conference server processing audio from at least one conference terminal; Based on the at least one first audio and the second audio, determine the mixed audio corresponding to each client and the conference terminal, and send the corresponding mixed audio to each client and the conference server; The method further includes: Receive a mute request from the client, and set the recording channel status between the client and the cloud desktop server to mute according to the mute request; or, The system receives a non-mute request from the client and, based on the non-mute request, sets the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

2. An audio processing method, characterized in that, Applied to a cloud desktop server for connecting to multiple clients, the method includes: Obtain the channel status of multiple recording channels between the cloud desktop server and the multiple clients, wherein the channel status is either muted or unmute; Based on the channel status of the plurality of recording channels, at least one target recording channel is determined among the plurality of recording channels, wherein the target recording channel is in a non-mute state; Through the at least one target recording channel, at least one first audio signal sent by the plurality of clients is received, wherein when the channel state of the recording channel is muted, the client stores the first audio signal in a preset storage space or discards the first audio signal; Receive the second audio sent by the conference server; Based on the at least one first audio and the second audio, determine the mixed audio corresponding to each client and conference terminal, and send the corresponding mixed audio to each client and the conference server; The method further includes: Receive a mute request from the client, and set the recording channel status between the client and the cloud desktop server to mute according to the mute request; or, The system receives a non-mute request from the client and, based on the non-mute request, sets the channel status of the recording channel between the client and the cloud desktop server to a non-mute state.

3. The method according to claim 1 or 2, characterized in that, Based on the at least one first audio and the second audio, determine the mixed audio corresponding to each client and the conference terminal, and send the corresponding mixed audio to each client and the conference server, including: The at least one first audio and the second audio are mixed to obtain a first mixed audio corresponding to each client, and the corresponding first mixed audio is sent to each client; The at least one first audio is mixed to obtain a second mixed audio corresponding to the conference server, and the second mixed audio is sent to the conference server.

4. The method according to claim 3, characterized in that, For any client; Mixing the at least one first audio and the second audio to obtain the first mixed audio corresponding to the client includes: Determine the first client identifier of the client, and the client identifier corresponding to each first audio; Based on the first client identifier and the client identifier corresponding to each first audio, a target audio is determined from the at least one first audio, wherein the client identifier corresponding to the target audio is different from the first client identifier; The target audio and the second audio are mixed to obtain the first mixed audio corresponding to the client.

5. The method according to any one of claims 1-4, characterized in that, Before receiving at least one first audio file sent by the plurality of clients via the recording channel between the plurality of clients, the method further includes: Receive a login request sent by a client, the login request including verification information; The verification information is verified, and after the verification information is verified, a recording channel is created between the cloud desktop server and the client. The recording channel is a socket channel.

6. An audio processing method, characterized in that, The method is applied to a client-side application, wherein the client is equipped with a microphone and is used to access a cloud desktop server, and each client has an independent recording channel with the cloud desktop server. The method includes: The first audio signal is obtained by capturing audio through the microphone. Determine the channel status of the recording channel between the client and the cloud desktop server, wherein the channel status is either muted or non-mute; When the channel is in a non-mute state, the first audio is sent to the cloud desktop server through the recording channel, so that the cloud desktop server can receive at least one first audio from multiple clients through multiple recording channels with multiple clients. When the channel is in a mute state, the first audio is stored in a preset storage space or discarded. The cloud desktop server receives mixed audio sent by the cloud desktop server. The mixed audio is determined by the cloud desktop server based on first audio from multiple clients and second audio sent by the conference server. The second audio is obtained by the conference server processing audio from at least one conference terminal.

7. The method according to claim 6, characterized in that, The method further includes: Display a first page, which includes the status control for the recording channel; In response to a mute operation on the status control, a mute request is sent to the cloud desktop server; or, in response to a demute operation on the status control, a demute request is sent to the cloud desktop server.

8. A cloud desktop server, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the audio processing method as described in any one of claims 1-5.

9. A client, characterized in that, include: Memory, processor, and microphone; The microphone is used to capture the first audio signal; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the audio processing method as described in any one of claims 6-7.

10. An audio processing system, characterized in that, It includes the cloud desktop server as described in claim 8, and at least one client as described in claim 9, wherein the client is used to access the cloud desktop server.

Citation Information

Patent Citations

  • Dynamic locale based aggregation of full duplex media streams

    US20160014373A1

  • User Interaction with Shared Content During a Virtual Meeting

    US20200293261A1