Real-time Application Method, System, Device and Medium for Cloud-based Audio Processing
By establishing a network connection between the audio relay device and the cloud server, and using the collaborative work of the audio relay device and the cloud server, the audio processing problem of the mobile terminal application large model in the instant messaging scenario is solved, and a low-cost and high-compatibility cloud audio real-time processing solution is realized, providing users with a good experience.
Patent Information
- Application Number
- CN202510511070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to implement a large model of mobile terminal application for audio processing in real-time communication scenarios, resulting in poor user experience and high cost.
By establishing a network connection between the audio relay device and the cloud server, the audio relay device receives function selection information, obtains the connection mode, and communicates with the cloud server to obtain model parameters and audio processing configuration information. Then, the audio transfer device collects external voice and forwards it to the cloud server. The cloud server performs audio processing and feedbacks the results to the audio transfer device, and finally outputs the processed audio.
It realizes a low-cost and high-compatibility cloud audio real-time processing solution, solves the problem that large-model audio processing is difficult to apply to mobile terminals in portability, and realizes the seamless combination of cloud large-model voice processing capabilities and instant communication scenarios, providing users with a good user experience.
Smart Images

Figure CN120050265B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, and in particular, to a real-time application method, system, device, and medium based on cloud audio processing. Background Art
[0002] With the rapid development of large model technologies, significant progress has been made in the capabilities of audio processing models such as speech recognition, voice cloning, voice conversion, voice beautification, and simultaneous interpretation. However, due to the high computational resource requirements of these high-quality audio processing models, they are currently mainly deployed on cloud servers or high-performance computing devices. This has led to the current large model audio processing capabilities being limited to the audio processing level and being difficult to seamlessly integrate into instant messaging application scenarios of intelligent terminals, such as daily usage scenarios like phone calls, video conferences, voice messages, live broadcasts, and media playback. In the prior art, there are mainly two ways to apply large model audio processing capabilities to terminal devices. One way is to configure a computer virtual sound card. By installing virtual sound card software on a computer, the audio is sent to the cloud for processing and then transmitted back, and then output through the virtual sound card. This method is complex to operate, only applicable to computer devices, and cannot be implemented on mobile terminals. Another way is to connect an external sound card. Through an external sound card device, the audio of one device is transmitted to another device. This method has a cumbersome connection, usually only exists as an audio input device, has a high cost, and a poor user experience. Therefore, the prior art methods cannot meet the actual needs of applying large models for audio processing on mobile terminals in instant messaging scenarios. Summary of the Invention
[0003] Embodiments of the present invention provide a real-time application method, system, device, and medium based on cloud audio processing, aiming to solve the problem in the prior art methods that large models cannot be applied for audio processing on mobile terminals in instant messaging scenarios.
[0004] In a first aspect, embodiments of the present invention provide a real-time application method based on cloud audio processing. The method is applied in a real-time application system, and the real-time application system includes an audio relay device and a cloud server. The audio relay device establishes network connections with an intelligent terminal and the cloud server respectively to achieve data information transmission. The method includes:
[0005] If the audio relay device receives the input function selection information, obtain the connection mode with the intelligent terminal;
[0006] The audio relay device sends the function selection information to the cloud server;
[0007] The cloud server obtains the model parameters of the cloud model matching in the preset model library according to the function selection information and sends them to the audio relay device;
[0008] The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table;
[0009] The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information;
[0010] The cloud server processes the external voices through the cloud model to obtain audio processing information and feeds it back to the audio relay device;
[0011] The audio relay device outputs the audio processing information according to the audio processing configuration information.
[0012] In a second aspect, an embodiment of the present invention further provides a real-time application system based on cloud audio processing. The real-time application system includes a connection mode acquisition unit, a sending unit, an audio processing configuration information acquisition unit, a forwarding unit, and an output unit configured in the audio relay device. The real-time application system further includes a model parameter sending unit and a processing unit configured in the cloud server. The audio relay device establishes network connections with the intelligent terminal and the cloud server respectively to implement data information transmission. The real-time application system is used to execute the real-time application method based on cloud audio processing described in the first aspect above;
[0013] The connection mode acquisition unit is configured to receive function selection information from the intelligent terminal and acquire the connection mode with the intelligent terminal;
[0014] The sending unit is configured to send the function selection information to the cloud server;
[0015] The model parameter sending unit is configured to acquire model parameters of a cloud model matching in a preset model library according to the function selection information and send them to the audio relay device;
[0016] The audio processing configuration information acquisition unit is configured to configure audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table;
[0017] The forwarding unit is configured to collect external voices according to the function selection information and forward the external voices to the cloud server according to the audio processing configuration information;
[0018] The processing unit is configured to process the external voices through the cloud model to obtain audio processing information and feed it back to the audio relay device;
[0019] The output unit is configured to output the audio processing information according to the audio processing configuration information.
[0020] In a third aspect, an embodiment of the present invention further provides a computer device, where the device includes a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory communicate with each other through the communication bus;
[0021] The memory is used to store a computer program;
[0022] When the processor executes the program stored in the memory, it implements the steps of the real-time application method for cloud-based audio processing described in the first aspect above.
[0023] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the real-time application method for cloud-based audio processing described in the first aspect above.
[0024] An embodiment of the present invention provides a real-time application method, system, device, and medium for cloud-based audio processing. The method includes: if function selection information input by a smart terminal is received, obtaining a connection mode, obtaining audio processing configuration information according to the function selection information, model parameters, and the connection mode, sending the function selection information to a cloud server and collecting external voice and sending it to the cloud server, the cloud server processes the external voice according to a cloud model to obtain audio processing information and feeds it back to an audio relay device, and the audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, solving the problem that it is difficult to port the audio processing of large models to mobile terminals; through audio relay and streaming processing technologies, it realizes the seamless combination of the voice processing ability of cloud large models and instant messaging scenarios, providing users with a good user experience. Description of the Drawings
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention;
[0027] Figure 2Schematic diagram of the application scenario of the real-time application method based on cloud audio processing provided by the embodiments of the present invention;
[0028] Figure 3 Application scenario diagram of the real-time application method based on cloud audio processing provided by the embodiments of the present invention;
[0029] Figure 4 Schematic diagram of the implementation of the real-time application method based on cloud audio processing provided by the embodiments of the present invention;
[0030] Figure 5 Another implementation illustration of the real-time application method based on cloud audio processing provided by the embodiments of the present invention;
[0031] Figure 6 Another implementation illustration of the real-time application method based on cloud audio processing provided by the embodiments of the present invention;
[0032] Figure 7 Schematic block diagram of the real-time application system based on cloud audio processing provided by the embodiments of the present invention;
[0033] Figure 8 Schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed implementation manners
[0034] In this application, an application (APP) is also an application program installed on a terminal device (such as the intelligent terminal 10) and running relying on the system program in the terminal device. The application program can run on the terminal device and implement corresponding usage functions, such as live broadcast, conference, voice conversation or multimedia playback.
[0035] In this application, the intelligent terminal 10 can be any device with computing and processing capabilities. Of course, the terminal device can also have audio and video playback and interface display functions. For example, the intelligent terminal 10 can be a mobile phone, a tablet computer, a vehicle-mounted device, a wearable device, an industrial device, etc.
[0036] The audio transfer device 20 can be any device with audio acquisition and audio processing. For example, the audio transfer device 20 can be a USB microphone, a wireless collar microphone, an earphone box, a Dangel earphone, an expansion dock with an audio transmission port, a power bank, MIFI, or a terminal with a screen. The audio transfer device 20 establishes a wired (such as USB, 3.5mm audio interface) or wireless (such as Bluetooth) communication connection with the smart terminal 10. At this time, the audio transfer device 20 is identified as a standard audio input / output device (such as earphones) on the smart terminal 10. The audio transfer device 20 can realize functions such as audio pre-processing, cloud large model interaction, audio post-processing, and audio link redirection, so that the audio processed by the cloud large model configured on the cloud server can be transmitted to the application scenario in real time.
[0037] USB Audio Class (UAC) is a device category specification in the USB protocol standard, formulated by the USB Implementers Forum (USB-IF) to standardize the way audio devices (such as the above-mentioned audio transfer device 20) communicate with smart terminals 10 (such as mobile terminals such as mobile phones and tablets) through USB interfaces. Its core goal is to standardize audio data transmission and control, so that devices that meet this specification can be recognized and used by the operating system without installing a dedicated driver, realizing "plug and play". Its communication protocol versions include UAC1.0, UAC2.0, UAC3.0 or higher versions.
[0038] Bluetooth HFP (Hands-Free Profile) is one of the core protocols in Bluetooth technology for voice call control and audio transmission. It is designed to support hands-free call scenarios (such as vehicle-mounted systems and Bluetooth headsets). HFP defines how to establish a voice call connection, control the call process (such as answering / hanging up the call), and transmit two-way voice signals (for example, communication between a mobile phone and a vehicle-mounted system or headset) between Bluetooth devices (such as the above-mentioned audio transfer device 20 and the smart terminal 10). Its communication protocol versions include FHP1.0, FHP1.6+, FHP1.7+ or higher versions.
[0039] Bluetooth A2DP (Advanced Audio Distribution Profile) is the core protocol for high-quality unidirectional audio transmission in Bluetooth technology, designed specifically for wireless transmission of stereo audio (such as music, podcasts). A2DP defines how high-fidelity audio streams are transmitted between Bluetooth devices (such as between the above-mentioned audio relay device 20 and the smart terminal 10), and supports unidirectional transmission of stereo signals from an audio source device (such as a mobile phone, computer) to an audio receiving device (such as headphones, speakers). Its core goal is to enable wireless music playback, rather than two-way calls (which are the responsibility of the HFP protocol), and its communication protocol versions include A2DP1.0, A2DP1.2, A2DP1.3 or higher versions.
[0040] The 3.5mm audio transmission interface (also known as the headphone jack or TRS interface) is a standardized physical interface widely used for analog audio signal transmission. It is named after its diameter of approximately 3.5 millimeters and can be used to transmit analog audio signals.
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0043] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0044] It should be further understood that the term " / and" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0045] Please refer to Figure 1, as shown in the figure, an embodiment of the present invention provides a real-time application method based on cloud audio processing. This method is applied to the terminal device of the real-time application system and is executed by the application software installed in the terminal device. As Figure 2 and Figure 3 shown, the real-time application system includes an audio relay device 20 and a cloud server 30; the audio relay device 20 establishes network connections with the smart terminal 10 and the cloud server 30 respectively to achieve data information transmission; the local user is the user of the smart terminal 10. The smart terminal 10 can be a terminal device using operating systems such as Android, iOS, Windows, Mac, Linux, HarmonyOS, etc.; the audio relay device 20 is integrated with a communication module, an audio system, a processing control unit and an external interface; the audio relay device 20 can communicate with an external audio device through Bluetooth / USB. For example, one of the earphones in a pair of earphones is used as the audio relay device 20 and the other earphone is used as the external audio device; or the earphone case is used as the audio relay device 20 and both earphones are used as external audio devices. The cloud server 30 integrates a management server and multiple cloud large models, and the management server manages the input, operation and output of the multiple cloud large models. The audio relay device 20 establishes a WiFi communication connection with the cloud server 30 through the WiFi communication module in the communication module, and the audio relay device 20 can also establish a 4G / 5G communication connection with the cloud server 30 through the 4G / 5G communication module in the communication module. As Figure 1 shown, the method includes steps S110 to S170.
[0046] S110. If the audio relay device receives the input function selection information, obtain the connection mode with the smart terminal.
[0047] If the audio relay device receives the input function selection information, obtain the connection mode with the smart terminal. The user can input function selection information on the smart terminal and transmit it to the audio relay device 20; in the specific application process, the user can also directly input function selection information on the audio relay device 20, or the user can use a remote control to input function selection information and transmit it to the audio relay device 20 that supports remote control, so as to transparently transmit the corresponding function selection information to the application program in the audio relay device 20. When the audio relay device 20 receives the function selection information, it obtains the connection mode with the smart terminal. The function selection information includes the selection information obtained by selecting functions such as voice conversion, voice beautification, simultaneous interpretation, etc., and the function parameter information for setting the target language and the local language.
[0048] In a specific embodiment, step S110 includes sub-steps: obtaining the type corresponding to the connection port between the intelligent terminal as the connection type; obtaining the version of the driver protocol corresponding to the connection port; combining the connection type and the driver protocol version as the connection mode.
[0049] Specifically, the audio relay device can obtain the type corresponding to the connection port between the intelligent terminal as the connection type. For example, if the audio relay device is connected to the intelligent terminal through a USB connection port, the corresponding connection type is UAC; if the audio relay device is wirelessly connected to the intelligent terminal through Bluetooth, the corresponding connection type is Bluetooth; if the audio relay device is connected to the intelligent terminal through a 3.5mm audio transmission interface, the corresponding connection type is analog signal.
[0050] Further obtain the version of the driver protocol corresponding to the connection port. Different connection types support different protocol versions. Obtaining the version of the script program for communicating with the driver connection port can obtain the driver protocol version; if the connection port is a 3.5mm audio transmission interface, it is determined that the driver protocol version is analog signal.
[0051] Combining the connection type and the driver protocol version can obtain the connection mode between the audio relay device and the intelligent terminal. The versions and characteristics of the script programs corresponding to each connection type are shown in Tables 1 to 3:
[0052] Table 1: UAC (USB Audio Class) Driver Protocol Version
[0053]
[0054] Among them, USB connection (UAC protocol) supports USB 2.0 and above standards; through the UAC protocol, the device is recognized as a USB audio device; supports audio quality up to 48kHz / 16bit; adapts to Windows, Mac, and Linux systems to achieve plug-and-play:
[0055] Table 2: Bluetooth HFP (Hands-Free Profile) Driver Protocol Version
[0056]
[0057] Table 3: Bluetooth A2DP (Advanced Audio Distribution Profile) Driver Protocol Version
[0058]
[0059] Among them, the Bluetooth connection (HFP / HSP protocol) supports Bluetooth version 4.2 and above; the device is recognized as a Bluetooth headset through the HFP / HSP protocol; two-way transmission of Bluetooth audio is achieved; the connection strategy is optimized for Android and iOS systems to adapt to the Bluetooth stacks of different systems.
[0060] The specification information transmitted through the 3.5mm audio interface is shown in Table 4:
[0061] Table 4
[0062]
[0063] Among them, the 3.5mm audio interface connection supports the CTIA-standard TRRS interface; audio input and output are achieved through analog signal conversion.
[0064] S120. The audio relay device sends the function selection information to the cloud server.
[0065] The audio relay device sends the function selection information to the cloud server.
[0066] S130. The cloud server obtains the model parameters of the cloud model that matches in the pre-set model library according to the function selection information and sends them to the audio relay device.
[0067] The cloud server obtains the cloud model that matches in the pre-set model library according to the function selection information and the function selection information. If there are multiple large language models configured in the model library of the cloud server, the model that matches in the model library can be obtained as the cloud model according to the function selection information and the function selection information, and the model parameters of the cloud model are obtained and sent to the audio relay device.
[0068] In a specific embodiment, step S130 includes sub-steps: obtaining the model that matches the number of audio channels in the function selection information as an alternative model from the model library; matching the function selection information with the application functions of each alternative model to obtain the alternative model whose application function matches the function selection information as the cloud model.
[0069] Specifically, if the number of audio channels in the function selection information is two-channel, the model that supports two channels needs to be obtained from the model library as an alternative model; if the number of audio channels in the function selection information is single-channel, the model that supports mono is obtained from the model library as an alternative model.
[0070] Furthermore, each alternative model has its corresponding application function. Therefore, it is necessary to match the application function of the alternative model through function selection information, so as to obtain an alternative model whose application function matches the function selection information as the cloud model finally used. Taking the foreign language recognition in the simultaneous interpretation scenario as an example, when the user selects Italian, only Model A has the corresponding streaming recognition ability at this time, and this model can be selected as the cloud model. Similarly, if Model A is selected for use, the audio pre-processing parameters and audio post-processing parameters can be correspondingly configured according to the interface target format of this model, and the sending configuration parameters (related to the size of the audio sending packet) and sending interval time (related to the sending interval of the audio sending packet) in the audio link parameters can also be correspondingly configured according to the audio transmission format of this model.
[0071] The real-time application method in the technology of this application can support multiple cloud processing models. The common cloud processing models and related features are shown in Table 5:
[0072] Table 5
[0073]
[0074] S140. The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table.
[0075] The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table. An audio configuration table is configured in the audio relay device, and corresponding audio processing configuration information can be configured according to the audio configuration table. The audio processing configuration information is used to configure the processing flow of the audio.
[0076] In a specific embodiment, step S140 includes sub-steps: obtaining audio format information matching the connection mode in the audio configuration table; correspondingly configuring audio link parameters according to the audio format information and the function selection information; correspondingly configuring audio pre-processing parameters and audio post-processing parameters according to the audio format information and the model parameters; combining the audio link parameters, the audio pre-processing parameters and the audio post-processing parameters into the audio processing configuration information.
[0077] Specifically, audio format information matching the connection mode can be obtained from the audio configuration table, and the audio configuration table includes the content involved in Tables 1 to 4. Then, according to the connection mode, information such as the audio format, sampling rate, bit depth, and number of channels supported by this connection mode can be correspondingly matched and obtained as the corresponding audio format information.
[0078] According to the audio format information and function selection information, the audio link parameters can be correspondingly configured. The audio link parameters can be used to indicate in what format and by what method the collected external voice is transmitted. If the function selection information is the AI voice conversion / voice beautification function, only one-way link configuration is required, and the corresponding audio processing flow is as follows Figure 4 shown. Its audio link parameters can be described in detail as follows: collect the user's voice according to the audio format information → preprocessing → transmit to the cloud → cloud AI voice beautification / voice conversion processing → return the processed audio → postprocessing → use it as the microphone input and transmit it to the application on the intelligent terminal side. Here, the large model is also the cloud model (the same below), and here this device is also the audio relay device (the same below). The corresponding specific embodiments include: (1) The user connects the audio relay device to the intelligent terminal; (2) Open the live application or game application on the intelligent terminal side; (3) Select the "voice beautification" or "voice conversion" function on the audio relay device; (4) Automatically configure the audio processing configuration information; (5) The user's voice is processed in real time through the cloud large model, and other users hear the voice after AI beautification or voice conversion. The end-to-end delay of the whole processing process is controlled within 200 ms, which does not affect normal communication. Specifically, its audio processing link is divided into three segments. The first segment refers to the delay time when the audio processed locally by the audio relay device is sent to the cloud server. The third segment refers to the delay time when the cloud server sends the processed audio to the audio relay device. The delay times of the first segment and the third segment can both be controlled within 200 ms. The second segment refers to the time for audio processing through the cloud model in the cloud server, and this processing time varies in the actual application of each model. The test results show that in the 4G / 5G network environment, the end-to-end delay of this solution is stable at 150 - 180 ms, and the audio quality is significantly improved.
[0079] If the function selection information is the real-time translation function / sensitive word filtering function, only one-way link configuration is also required, and the corresponding audio processing flow is as follows Figure 5As shown in the figure, its audio link parameters can be described in detail as follows: Collect in-app audio according to audio format information → Pre-processing → Transmit to the cloud → Cloud translation / sensitive content detection → Obtain translation information / replace or block content → Speech synthesis → Return the processed audio → Post-processing → Use it as the microphone input and transmit it to the application on the smart terminal side. The corresponding specific embodiments include: (1) The user connects the audio relay device to the smart terminal; (2) Open the live application or conference application on the smart terminal side; (3) Select "Real-time Translation" on the audio relay device and set the target language or the "Content Filtering" function; (4) Automatically configure the audio processing configuration information; (5) The in-app audio is translated into the corresponding response audio in the target language through the cloud / the voice content containing sensitive words will be automatically replaced with a prompt tone or a neutral expression. The test results show that when performing speech translation with this solution, the accuracy rate is high and the end-to-end delay is stably within 200 ms. Here, the end-to-end delay corresponds to the first and third segments of the audio processing link in the above description, excluding the processing time of the cloud model in the second segment.
[0080] If the function selection information is two-way call simultaneous interpretation, two-way link configuration is required, and the corresponding audio processing process is as Figure 6 shown in the figure, its audio link parameters can be described in detail as follows: Collect local voice according to audio format information → Pre-processing → Translate to the target language in the cloud → Return the translated audio → Post-processing → Output to the application on the smart terminal side; Remote voice → Pre-processing → Translate to the local language in the cloud → Return the translated audio → Post-processing → Local playback. The corresponding specific embodiments include: (1) The user connects the audio relay device to the smart terminal; (2) Open the phone application or video conference application on the smart terminal side; (3) Select the "Simultaneous Interpretation" function on the audio relay device and set the source language and target language; (4) Automatically configure the audio processing configuration information; (5) Both parties can communicate in their respective mother tongues, and the system automatically completes real-time translation. The test results show that this solution supports real-time translation between 21 languages such as Chinese, English, Japanese, Korean, French, and German. The translation accuracy rate reaches over 85% in mainstream scenarios, and the end-to-end delay is controlled within 300 ms. Here, the end-to-end delay corresponds to the first and third segments of the audio processing link in the above description, excluding the processing time of the cloud model in the second segment. The number of languages only depends on the capabilities of the accessed models and can be expanded at any time.
[0081] Further, audio pre - processing parameters and audio post - processing parameters can be configured correspondingly according to the audio format information. The source of the original audio may be a microphone, Bluetooth, or USB, and the format may be PCM, AAC, MP3, etc. The audio pre - processing parameters are the parameter information for pre - processing the audio output to the cloud server to make it meet the transmission requirements. The audio post - processing parameters are the parameter information for post - processing the processed audio from the cloud server to make it meet the local playback / output requirements.
[0082] For example, in the audio configuration table, the audio format information corresponding to the connection mode is: audio format SBC, sampling rate 48 kHz, bit depth 24 bits, and stereo channels; in the model parameters, the interface target format is: audio format PCM, sampling rate 16 kHz, bit depth 16 bits, and single - channel. Then, the audio pre - processing parameters need to be configured as audio format information → interface target format, and the configuration process of the audio post - processing parameters is similar.
[0083] Combine the audio link parameters, the audio pre - processing parameters, and the audio post - processing parameters into the audio processing configuration information. S150: The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information.
[0084] The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information. The audio relay device collects external voices corresponding to the function selection information. For example, if the function selection information is set to the AI voice - changing / voice - beautifying function, the user's audio is collected through the microphone as the external voice. If the function selection information is set to the real - time translation function / sensitive word filtering function, the audio relay device obtains the in - app audio from the terminal device through the communication connection as the external voice. The audio relay device forwards the obtained external voice to the cloud server according to the audio processing configuration information.
[0085] In a specific embodiment, step S150 includes sub - steps: transcoding the external voice according to the audio pre - processing parameters in the audio processing configuration information to obtain the corresponding transcoded audio; splitting the transcoded audio according to the sending configuration parameters of the audio link parameters in the audio processing configuration information to obtain the corresponding audio sending packets; and sending the audio sending packets to the cloud server according to the sending link of the audio link parameters in the audio processing configuration information.
[0086] Specifically, transcoding processing can be performed on the external voice according to the audio preprocessing parameters to obtain transcoded audio. Transcoding processing is to convert the audio format, sampling rate, bit depth, and channels of the external voice. For example, if the external voice is stereo and a mono audio needs to be transcoded and obtained as the transcoded audio, the two channels of audio in the stereo can be mixed to obtain the corresponding mono audio.
[0087] The transcoded audio is segmented according to the sending configuration parameters in the audio link parameters to obtain redundant audio sending packets. For the convenience of sending audio information, audio sending packets of a preset size can be segmented according to the transmission requirements. For example, if the preset size is set to 1 kB, then audio sending packets of 1 kB in size can be segmented and sent accordingly.
[0088] The audio sending packets are sent to the cloud server according to the sending link of the audio link parameters. Further, the audio sending packets can also be sent according to the sending interval time configured in the audio link parameters, so that the interval between the sending times of adjacent audio sending packets is not less than the sending interval time. For example, if the sending interval time is 20 ms, it can be ensured that the interval time between the sending of two audio sending packets is not less than 20 ms to avoid interference between audio sending packets during the sending and processing.
[0089] S160. The cloud server processes the external voice through the cloud model to obtain audio processing information and feeds it back to the audio relay device.
[0090] The cloud server processes the external voice through the cloud model to obtain audio processing information and feeds it back to the audio relay device. The cloud server processes the received external voice by matching the obtained cloud model to obtain audio processing information. The cloud server feeds the obtained audio processing information back to the audio relay device, and the audio processing information includes the output audio. During the process of the cloud model processing the external voice, it includes audio feature extraction, audio recognition, recognition information processing, and speech synthesis. First, the audio features of the external voice are extracted; and the audio features are recognized to obtain corresponding recognition information. At this time, the recognition information is the initial text content recognized based on the external voice. The recognition information is processed, such as translation / sensitive word filtering, etc., to obtain processing information, and then speech synthesis is performed based on the processing information (the synthesis parameters of the cloud model are involved in the speech synthesis process), and the corresponding audio processing information can be obtained.
[0091] In a specific embodiment, before the cloud model processes the external voice to obtain audio processing information and feeds it back to the audio relay device, the following steps are further included: mapping the initial synthesis parameters according to a preset parameter mapping table to obtain mapping parameters corresponding to the cloud model; configuring the parameters of the cloud model according to the mapping parameters.
[0092] For the voice synthesis part of each model, there are differences in the parameter ranges of speech rate, volume, and pitch. To facilitate user use, the synthesis parameters of the model can be set to be unified. This parameter configuration process is achieved through a parameter mapping table, and parameter transmission with the cloud model can be completed through the parameter mapping table. The initial synthesis parameters include: speech rate - "0" (the settable range is -10 to 10), volume - "10" (the settable range is 0 to 10), pitch - "0" (the settable range is -10 to 10).
[0093] The parameter mapping table includes the mapping relationship between the initial synthesis parameters and the synthesis parameters in each model. For example, if the cloud model is model A, according to the initial synthesis parameters, the mapping parameters of this cloud model can be obtained as: speech rate - "0" (the settable range is -500 to 500), volume - "100" (the settable range is 0 to 100), pitch - "0" (the settable range is -500 to 500).
[0094] Configure the parameters of the cloud model according to the obtained mapping parameters, so that the parameters of the cloud model for voice synthesis match the initial synthesis parameters.
[0095] S170. The audio relay device outputs the audio processing information according to the audio processing configuration information.
[0096] The audio relay device outputs the audio processing information according to the audio processing configuration information. After receiving the audio processing information from the cloud server, the audio relay device can output the obtained audio processing information according to the audio processing configuration information.
[0097] In a specific embodiment, step S170 includes sub-steps: transcoding the output audio in the audio processing information according to the audio post-processing parameters in the audio processing configuration information to obtain a corresponding target audio; sending the target audio to the intelligent terminal and / or playing and outputting it according to the output link of the audio link parameters in the audio processing configuration information.
[0098] Specifically, the output audio in the audio processing information can be transcoded according to the audio post-processing parameters to obtain the target audio. The transcoding process here is similar to the process of transcoding external speech in the above steps, except that there are differences in the audio formats during transcoding. Further, the target audio is output according to the output link in the audio link parameters. According to the output link, the target audio can be sent to the applications of the smart terminal. For example, in a real-time meeting application scenario, the voice of the local user on the terminal device side is converted into the target audio corresponding to the language of the remote user in the application, and the target audio is sent to the remote user through the application. It can also be that the target audio is directly played in the audio relay device according to the output link. For example, after the voice of the remote user is processed to obtain the target audio corresponding to the language of the local user on the terminal device side, the target audio can be directly output and played through the audio relay device, and the user listens to the target audio through the audio relay device. The target audio can also be sent to the smart terminal and directly output and played in the audio relay device according to the output link to form a multi-directional stereo. Under normal use conditions, the user can freely switch to output the original sound or the target sound in the audio relay device, or mix the two audio signals and output and play the mixed audio.
[0099] Further, the audio processing information further includes text information. Then, the audio relay device can send the text information in the audio processing information to the smart terminal and / or perform display output according to the output link. The text information in the audio processing information is also the processing information obtained before the cloud server performs speech synthesis. This processing information can be used as the text content to assist the local user in viewing and understanding. The text information can be sent to the smart terminal for display, or directly displayed and output in the audio relay device, or can also be sent to the smart terminal for display and displayed and output in the audio relay device at the same time.
[0100] In the real-time application method based on cloud audio processing disclosed in the above embodiments, the method includes: if function selection information input by the smart terminal is received, obtain the connection mode, obtain the audio processing configuration information according to the function selection information, model parameters, and connection mode, send the function selection information to the cloud server, and collect external speech and send it to the cloud server. The cloud server processes the external speech according to the cloud model to obtain audio processing information and feedback it to the audio relay device. The audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, which solves the problem that it is difficult to apply large model audio processing to mobile terminals portably; through audio relay and streaming processing technologies, it realizes the seamless combination of the voice processing ability of the cloud large model and the instant messaging scenario, providing a good user experience for users.
[0101] An embodiment of the present invention further provides a real-time application system based on cloud audio processing, as Figure 7 shown. The real-time application system 100 of the real-time application system based on cloud audio processing includes a connection mode acquisition unit 210, a sending unit 220, an audio processing configuration information acquisition unit 230, a forwarding unit 240, and an output unit 250 configured in the audio relay device. The real-time application system further includes a model parameter sending unit 310 and a processing unit 320 configured in the cloud server 30. The audio relay device establishes network connections with the smart terminal and the cloud server respectively to realize the transmission of data information. This real-time application system based on cloud audio processing is used to execute any embodiment of the foregoing real-time application method based on cloud audio processing. Specifically, please refer to Figure 7 , Figure 7 which is a schematic block diagram of the real-time application system based on cloud audio processing provided by the embodiment of the present invention.
[0102] The connection mode acquisition unit 210 is configured to receive function selection information from the smart terminal and acquire the connection mode with the smart terminal.
[0103] The sending unit 220 is configured to send the function selection information to the cloud server.
[0104] The model parameter sending unit 310 is configured to acquire the model parameters of the cloud model matching in the preset model library according to the function selection information and send them to the audio relay device.
[0105] The audio processing configuration information acquisition unit 230 is configured to configure audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table.
[0106] The forwarding unit 240 is configured to collect external voices according to the function selection information and forward the external voices to the cloud server according to the audio processing configuration information.
[0107] The processing unit 320 is configured to process the external voices through the cloud model to obtain audio processing information and feedback it to the audio relay device.
[0108] The output unit 250 is configured to output the audio processing information according to the audio processing configuration information.
[0109] In the real-time application system based on cloud audio processing provided by the embodiments of the present invention, when the above real-time application method based on cloud audio processing is applied, if function selection information input by a smart terminal is received, the connection mode is obtained, and audio processing configuration information is obtained according to the function selection information, model parameters, and connection mode. The function selection information is sent to the cloud server, and external voice is collected and sent to the cloud server. The cloud server processes the external voice according to the cloud model to obtain audio processing information and feeds it back to the audio relay device. The audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, solving the problem that it is difficult to apply large model audio processing to mobile terminals in a portable manner; through audio relay and streaming processing technologies, seamless integration of the cloud large model voice processing ability and the instant messaging scenario is achieved, providing a good user experience for users.
[0110] In the above real-time application system based on cloud audio processing, each unit module can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 8 the following.
[0111] Please refer to Figure 8 , Figure 8 which is a schematic block diagram of the computer device provided by the embodiments of the present invention. This computer device can be an audio relay device and a cloud server used to execute the real-time application method based on cloud audio processing to realize real-time application of audio processing based on a cloud large model.
[0112] Referring to Figure 8 , the computer device 500 includes a processor 502, a memory, and a communication interface 505 connected through a communication bus 501. Among them, the memory can include a storage medium 503 and an internal memory 504.
[0113] The storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can be made to execute the real-time application method based on cloud audio processing. Among them, the storage medium 503 can be a volatile storage medium or a non-volatile storage medium.
[0114] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0115] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute the real-time application method based on cloud audio processing.
[0116] The communication interface 505 is used for network communication, such as providing the transmission of data information, etc. Those skilled in the art can understand that Figure 8 The structure shown in Figure 8 is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. Specifically, the computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0117] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the corresponding functions in the above-mentioned real-time application method based on cloud audio processing.
[0118] Those skilled in the art can understand that Figure 8 The embodiments of the computer device shown in Figure 8 do not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are the same as those of Figure 8 the embodiments shown, and will not be described in detail here.
[0119] It should be understood that in the embodiments of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0120] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps included in the above-mentioned real-time application method based on cloud audio processing are implemented.
[0121] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0122] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, or units with the same function can be aggregated into one unit. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.
[0123] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0124] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes.
[0126] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A real-time application method based on cloud audio processing, characterized in that: The method is applied in a real-time application system, which includes an audio transfer device and a cloud server; The audio transfer device establishes network connections with the smart terminal and the cloud server respectively to realize the transmission of data information, and the method includes: If the audio transfer device receives the input function selection information, it obtains the connection mode between the audio transfer device and the smart terminal; The audio transfer device sends the function selection information to the cloud server; The cloud server obtains the model parameters of the matching cloud model in the preset model library according to the function selection information and sends them to the audio transfer device; The audio transfer device configures the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table; The audio transfer device collects external voice according to the function selection information and forwards the external voice to the cloud server according to the audio processing configuration information; The cloud server processes the external voice through the cloud model to obtain audio processing information which is fed back to the audio transfer device; The audio transfer device outputs the audio processing information according to the audio processing configuration information.
2. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The acquiring of the connection mode with the smart terminal includes: Acquire a type corresponding to a connection port between the intelligent terminals as a connection type; Obtain the driver protocol version corresponding to the connection port; The connection type and the driver protocol version are combined into the connection mode.
3. The real-time application method based on cloud audio processing according to claim 1 or 2, characterized in that: The configuring the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameter according to the preset audio configuration table includes: Acquire audio format information matching the connection mode in the audio configuration table; Obtaining audio link parameters according to the corresponding configuration of the audio format information and the function selection information; Obtaining audio pre-processing parameters and audio post-processing parameters according to the audio format information and the corresponding configuration of the model parameters; The audio link parameter, the audio pre-processing parameter, and the audio post-processing parameter are combined into the audio processing configuration information.
4. The real-time application method based on cloud audio processing according to claim 1 or 2, characterized in that: The forwarding the external voice to the cloud server according to the audio processing configuration information includes: Performing transcoding processing on the external voice according to the audio pre-processing parameters in the audio processing configuration information to obtain corresponding transcoded audio; Segment the transcoded audio according to the sending configuration parameters of the audio link parameters in the audio processing configuration information to obtain corresponding audio sending packets; The audio sending package is sent to the cloud server according to the sending link of the audio link parameter in the audio processing configuration information.
5. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The step of acquiring model parameters of a matching cloud model in a preset model library according to the function selection information and sending the model parameters to the audio transfer device includes: According to the audio path of the function selection information, a model matching the audio path in the model library is obtained as a candidate model; The function selection information is matched with the application function of each candidate model to obtain a candidate model whose application function matches the function selection information as a cloud model.
6. The real-time application method based on cloud audio processing according to claim 1 or 5, characterized in that: Before the external voice is processed by the cloud model to obtain audio processing information and then fed back to the audio transfer device, the method further includes: Mapping the initial synthesis parameters according to a preset parameter mapping table to obtain mapping parameters corresponding to the cloud model; Parameters of the cloud model are configured according to the mapping parameters.
7. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The outputting the audio processing information according to the audio processing configuration information includes: Performing transcoding processing on the output audio in the audio processing information according to the audio post-processing parameters in the audio processing configuration information to obtain corresponding target audio; The target audio is sent to the smart terminal and / or played and output according to the output link of the audio link parameter in the audio processing configuration information.
8. A real-time application system based on cloud audio processing, characterized in that: The real-time application system includes a connection mode acquisition unit, a sending unit, an audio processing configuration information acquisition unit, a forwarding unit and an output unit configured in the audio transfer device, and the real-time application system also includes a model parameter sending unit and a processing unit configured in the cloud server. The audio transfer device establishes a network connection with the smart terminal and the cloud server respectively to realize the transmission of data information. The real-time application system is used to execute the real-time application method based on cloud audio processing according to any one of claims 1 to 7; The connection mode acquisition unit is used to receive function selection information from the smart terminal and acquire the connection mode between the smart terminal and the smart terminal; The sending unit is used to send the function selection information to the cloud server; The model parameter sending unit is used to obtain the model parameters of the matching cloud model in the preset model library according to the function selection information and send them to the audio transfer device; The audio processing configuration information acquisition unit is used to configure the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table; The forwarding unit is used to collect external voice according to the function selection information and forward the external voice to the cloud server according to the audio processing configuration information; The processing unit is used to process the external voice through the cloud model to obtain audio processing information and feed it back to the audio transfer device; The output unit is used to output the audio processing information according to the audio processing configuration information.
9. A computer device, characterized in that: The device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of the real-time application method based on cloud audio processing described in any one of claims 1 to 7 when executing the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the real-time application method based on cloud audio processing are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
IMS audio and video transcoding and transmission control system and implementation method
CN113301025A
Splicing persistent connections
US20020120743A1