Real-time application method, system and equipment based on cloud audio processing and medium

By establishing a network connection between the audio transit device and the cloud server and using the cloud model to process external voice, the problem of difficult to achieve large-scale audio processing of mobile terminals is solved, and a low-cost and high-compatibility cloud audio real-time processing solution is realized, providing users with a good user experience.

CN120050265AActive Publication Date: 2025-05-27SHENZHEN EASTAI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510511070.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing technology is difficult to implement a large model of mobile terminal application for audio processing in real-time communication scenarios, and the existing methods are complex in operation, high in cost and poor user experience.

Method used

By establishing a network connection between the audio transfer device and the cloud server, the transmission of function selection information and the configuration of audio processing configuration are realized, external voice is processed using the cloud model, and the processing results are fed back to the audio transfer device for output.

Benefits of technology

It provides a low-cost and high-compatibility cloud audio real-time processing solution, solving the problem that large-model audio processing is difficult to apply to mobile terminals in portability, and realizes the seamless combination of cloud large-model voice processing capabilities and instant communication scenarios, providing users with a good user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050265A_ABST
    Figure CN120050265A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time application method, system and device based on cloud audio processing and a medium, and the method comprises the steps: obtaining a connection mode if function selection information inputted by an intelligent terminal is received, and obtaining audio processing configuration information according to the function selection information, model parameters and the connection mode, the function selection information is sent to the cloud server, external voice is collected and sent to the cloud server, the cloud server processes the external voice according to a cloud model to obtain audio processing information and feeds the audio processing information back to the audio transfer device, and the audio transfer device outputs the audio processing information according to the audio processing configuration information. According to the method, a cloud audio real-time processing scheme which is low in cost and high in compatibility is provided, and the problem that large-model audio processing is difficult to be conveniently applied to the mobile terminal is solved; through the audio transfer and streaming processing technology, seamless combination of the cloud large model voice processing capability and the instant messaging scene is realized, and good use experience is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and in particular, to a real-time application method, system, device, and medium based on cloud audio processing. Background Art

[0002] With the rapid development of large model technologies, significant progress has been made in the capabilities of audio processing models such as speech recognition, voice cloning, voice conversion, voice beautification, and simultaneous interpretation. However, due to the high computational resource requirements of these high-quality audio processing models, they are currently mainly deployed on cloud servers or high-performance computing devices. This results in the current large model audio processing capabilities being only at the audio processing level and being difficult to seamlessly integrate into instant messaging application scenarios of intelligent terminals, such as daily usage scenarios like phone calls, video conferences, voice messages, live broadcasts, and media playback. In the prior art, there are mainly two ways to apply large model audio processing capabilities to terminal devices. One way is to configure a computer virtual sound card. By installing virtual sound card software on a computer, the audio is sent to the cloud for processing and then returned, and then output through the virtual sound card. This method is complex to operate, only applicable to computer devices, and cannot be implemented on mobile terminals. Another way is to connect an external sound card. Through an external sound card device, the audio of one device is transmitted to another device. This method has a cumbersome connection, is usually only used as an audio input device, has a high cost, and a poor user experience. Therefore, the prior art methods cannot meet the actual needs of applying large models for audio processing on mobile terminals in instant messaging scenarios. Summary of the Invention

[0003] Embodiments of the present invention provide a real-time application method, system, device, and medium based on cloud audio processing, aiming to solve the problem in the prior art methods that large models cannot be applied for audio processing on mobile terminals in instant messaging scenarios.

[0004] In a first aspect, embodiments of the present invention provide a real-time application method based on cloud audio processing. The method is applied to a real-time application system, and the real-time application system includes an audio relay device and a cloud server. The audio relay device establishes network connections with an intelligent terminal and the cloud server respectively to achieve data information transmission. The method includes: If the audio relay device receives the input function selection information, obtain the connection mode with the intelligent terminal; The audio relay device sends the function selection information to the cloud server; The cloud server obtains the model parameters of the cloud model matching the function selection information in the preset model library and sends them to the audio relay device; The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table; The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information; The cloud server processes the external voices through the cloud model and feeds back audio processing information to the audio relay device; The audio relay device outputs the audio processing information according to the audio processing configuration information.

[0005] In a second aspect, an embodiment of the present invention further provides a real-time application system based on cloud audio processing. The real-time application system includes a connection mode acquisition unit, a sending unit, an audio processing configuration information acquisition unit, a forwarding unit, and an output unit configured in the audio relay device. The real-time application system further includes a model parameter sending unit and a processing unit configured in the cloud server. The audio relay device establishes network connections with the intelligent terminal and the cloud server respectively to implement data information transmission. The real-time application system is used to execute the real-time application method based on cloud audio processing described in the first aspect above; The connection mode acquisition unit is configured to receive function selection information from the intelligent terminal and acquire the connection mode with the intelligent terminal; The sending unit is configured to send the function selection information to the cloud server; The model parameter sending unit is configured to acquire model parameters of a cloud model matching in a preset model library according to the function selection information and send them to the audio relay device; The audio processing configuration information acquisition unit is configured to configure audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table; The forwarding unit is configured to collect external voices according to the function selection information and forward the external voices to the cloud server according to the audio processing configuration information; The processing unit is configured to process the external voices through the cloud model and feed back audio processing information to the audio relay device; The output unit is configured to output the audio processing information according to the audio processing configuration information.

[0006] In a third aspect, an embodiment of the present invention further provides a computer device. The device includes a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus; A memory for storing computer programs; A processor, when executing the programs stored in the memory, implements the steps of the real-time application method for cloud-based audio processing described in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the steps of the real-time application method for cloud-based audio processing described in the first aspect above.

[0008] An embodiment of the present invention provides a real-time application method, system, device, and medium for cloud-based audio processing. The method includes: if function selection information input by an intelligent terminal is received, then obtain a connection mode, obtain audio processing configuration information according to the function selection information, model parameters, and connection mode, send the function selection information to a cloud server and collect external voices and send them to the cloud server. The cloud server processes the external voices according to a cloud model to obtain audio processing information and feedback it to an audio relay device. The audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, solving the problem that it is difficult to portably apply large model audio processing to mobile terminals; through audio relay and streaming processing technologies, it realizes the seamless combination of the speech processing ability of the cloud large model and the instant messaging scenario, providing users with a good user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0010] Figure 1 It is a flowchart of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of an application scenario of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention; Figure 3 It is an application scenario diagram of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention; Figure 4 It is an implementation schematic diagram of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention; Figure 5 It is another implementation schematic of the real-time application method for cloud-based audio processing provided by an embodiment of the present invention; Figure 6 Another implementation schematic of the real-time application method based on cloud audio processing provided by an embodiment of the present invention; Figure 7 Schematic block diagram of the real-time application system based on cloud audio processing provided by an embodiment of the present invention; Figure 8 Schematic block diagram of the computer device provided by an embodiment of the present invention. Detailed implementation manners

[0011] In this application, an application (APP) is also an application program installed on a terminal device (such as the intelligent terminal 10) and running relying on the system program in the terminal device. The application program can run on the terminal device and implement corresponding usage functions, such as live broadcast, meeting, voice conversation, or multimedia playback.

[0012] In this application, the intelligent terminal 10 can be any device with computing and processing capabilities. Of course, the terminal device can also have audio and video playback and interface display functions. For example, the intelligent terminal 10 can be a mobile phone, a tablet computer, a vehicle-mounted device, a wearable device, an industrial device, etc.

[0013] The audio relay device 20 can be any device with audio collection and audio processing capabilities. For example, the audio relay device 20 can be a USB microphone, a wireless lapel microphone, a headphone case, a Dangel headphone, a docking station with an audio transmission port, a power bank, a MIFI, a terminal with a screen. The audio relay device 20 establishes a wired (such as USB, 3.5mm audio interface) or wireless (such as Bluetooth) communication connection with the intelligent terminal 10. At this time, the audio relay device 20 is recognized as a standard audio input / output device (such as a headphone) at the intelligent terminal 10 end. The audio relay device 20 can implement functions such as audio pre-processing, cloud large model interaction, audio post-processing, and audio link redirection inside, so that the audio processed by the cloud large model configured in the cloud server can be transmitted to the application scenario in real time.

[0014] USB Audio Class (UAC, USB audio class) is a device category specification in the USB protocol standard, formulated by the USB Implementers Forum (USB-IF), and is used to standardize the way audio devices (such as the above-mentioned audio relay device 20) communicate with intelligent terminals 10 (such as mobile terminals like mobile phones and tablet computers) through the USB interface. Its core goal is to standardize audio data transmission and control, so that devices that conform to this specification can be recognized and used by the operating system without installing special drivers, achieving "plug and play". Its communication protocol versions include UAC1.0, UAC2.0, UAC3.0, or higher versions.

[0015] Bluetooth HFP (Hands-Free Profile) is one of the core protocols in Bluetooth technology for implementing voice call control and audio transmission. It is designed specifically to support hands-free call scenarios (such as in-vehicle systems, Bluetooth headsets). HFP defines how to establish a voice call connection between Bluetooth devices (such as between the above-mentioned audio relay device 20 and the smart terminal 10), control the call process (such as answering / hanging up a call), and transmit two-way voice signals (for example: communication between a mobile phone and an in-vehicle system, headphones). Its communication protocol versions include FHP1.0, FHP1.6+, FHP1.7+ or higher versions.

[0016] Bluetooth A2DP (Advanced Audio Distribution Profile) is the core protocol in Bluetooth technology for high-quality unidirectional audio transmission. It is designed specifically for wireless transmission of stereo audio (such as music, podcasts). A2DP defines how to transmit a high-fidelity audio stream between Bluetooth devices (such as between the above-mentioned audio relay device 20 and the smart terminal 10), and supports unidirectional transmission of stereo signals from an audio source device (such as a mobile phone, computer) to an audio receiving device (such as headphones, speakers). Its core goal is to enable wireless music playback, rather than two-way calls (which are the responsibility of the HFP protocol). Its communication protocol versions include A2DP1.0, A2DP1.2, A2DP1.3 or higher versions.

[0017] The 3.5mm audio transmission interface (also known as the headphone jack or TRS interface) is a standardized physical interface widely used for analog audio signal transmission. It is named after its diameter of approximately 3.5 millimeters and can be used to transmit analog audio signals.

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0020] It should also be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0021] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0022] Please refer to Figure 1 , as shown in the figure, an embodiment of the present invention application provides a real-time application method based on cloud audio processing, which is applied to a real-time application system terminal device, and the method is executed by an application software installed in the terminal device. As Figure 2 and Figure 3 shown, the real-time application system includes an audio relay device 20 and a cloud server 30; the audio relay device 20 establishes network connections with the smart terminal 10 and the cloud server 30 respectively to realize the transmission of data information; the local user is the user of the smart terminal 10. The smart terminal 10 can be a terminal device using operating systems such as Android, iOS, Windows, Mac, Linux, HarmonyOS, etc.; the audio relay device 20 is integrated with a communication module, an audio system, a processing control unit and an external interface; the audio relay device 20 can communicate with an external audio device through Bluetooth / USB, such as one of a pair of headphones as the audio relay device 20 and the other headphone as the external audio device; or the headphone case as the audio relay device 20 and both headphones as the external audio devices. The cloud server 30 integrates a management server and multiple cloud large models, and the management server manages the input, operation and output of the multiple cloud large models. The audio relay device 20 establishes a WiFi communication connection with the cloud server 30 through the WiFi communication module in the communication module, and the audio relay device 20 can also establish a 4G / 5G communication connection with the cloud server 30 through the 4G / 5G communication module in the communication module. As Figure 1 shown, the method includes steps S110 to S170.

[0023] S110. If the audio relay device receives the input function selection information, obtain the connection mode with the smart terminal.

[0024] If the audio relay device receives the input function selection information, it obtains the connection mode with the smart terminal. The user can input function selection information on the smart terminal and transmit it to the audio relay device 20; in the specific application process, the user can also directly input function selection information on the audio relay device 20, or the user can use a remote control to input function selection information and transmit it to the audio relay device 20 that supports remote control, so as to transparently transmit the corresponding function selection information to the application program in the audio relay device 20. When the audio relay device 20 receives the function selection information, it obtains the connection mode with the smart terminal. The function selection information includes the selection information obtained by selecting functions such as voice conversion, voice beautification, and simultaneous interpretation, as well as the function parameter information for setting the target language and the native language of the interpreter.

[0025] In a specific embodiment, step S110 includes sub-steps: obtaining the type corresponding to the connection port with the smart terminal as the connection type; obtaining the driver protocol version corresponding to the connection port; and combining the connection type and the driver protocol version as the connection mode.

[0026] Specifically, the audio relay device can obtain the type corresponding to the connection port with the smart terminal as the connection type. For example, if the audio relay device is connected to the smart terminal through a USB connection port, the corresponding connection type is UAC; if the audio relay device is wirelessly connected to the smart terminal through Bluetooth, the corresponding connection type is Bluetooth; if the audio relay device is connected to the smart terminal through a 3.5mm audio transmission interface, the corresponding connection type is analog signal.

[0027] Further obtain the driver protocol version corresponding to the connection port. Different connection types support different protocol versions. The driver protocol version can be obtained by obtaining the version of the script program for communicating with the driver connection port; if the connection port is a 3.5mm audio transmission interface, it is determined that the driver protocol version is analog signal.

[0028] Combining the connection type and the driver protocol version can obtain the connection mode between the audio relay device and the smart terminal. The versions and characteristics of the script programs corresponding to each connection type are shown in Tables 1 to 3: Table 1: UAC (USB Audio Class) Driver Protocol Version

[0029] Among them, USB connection (UAC protocol) supports USB 2.0 and above standards; through the UAC protocol, the device is recognized as a USB audio device; it supports audio quality up to 48kHz / 16bit; it is adapted to Windows, Mac, and Linux systems and realizes plug-and-play: Table 2: Bluetooth HFP (Hands-Free Profile) Driver Protocol Version Table 3: Bluetooth A2DP (Advanced Audio Distribution Profile) Driver Protocol Version

[0030] Among them, Bluetooth connection (HFP / HSP protocol) supports Bluetooth version 4.2 and above; the device is recognized as a Bluetooth headset through the HFP / HSP protocol; two-way transmission of Bluetooth audio is achieved; the connection strategy is optimized for Android and iOS systems to adapt to the Bluetooth stacks of different systems.

[0031] The specification information transmitted through the 3.5mm audio interface is shown in Table 4: Table 4

[0032] Among them, the 3.5mm audio interface connection supports the CTIA-standard TRRS interface; audio input and output are achieved through analog signal conversion.

[0033] S120. The audio relay device sends the function selection information to the cloud server.

[0034] The audio relay device sends the function selection information to the cloud server.

[0035] S130. The cloud server obtains the model parameters of the cloud model matching the function selection information in the preset model library and sends them to the audio relay device.

[0036] The cloud server obtains the cloud model matching the function selection information in the preset model library according to the function selection information. If there are multiple large language models configured in the model library of the cloud server, the cloud model matching the function selection information can be obtained from the model library according to the function selection information, and the model parameters of the cloud model are obtained and sent to the audio relay device.

[0037] In a specific embodiment, step S130 includes sub-steps: obtaining the model matching the number of audio channels in the model library according to the number of audio channels of the function selection information as an alternative model; matching the function selection information with the application functions of each alternative model to obtain the alternative model with the application function matching the function selection information as the cloud model.

[0038] Specifically, if the number of audio channels in the function selection information is two-channel, then it is necessary to obtain a model that supports two channels from the model library as an alternative model; if the number of audio channels in the function selection information is single-channel, then a model that supports mono-channel is obtained from the model library as an alternative model.

[0039] Furthermore, each alternative model has its corresponding application function. Therefore, it is necessary to match the application function of the alternative model through the function selection information to obtain an alternative model whose application function matches the function selection information as the cloud model finally used. Taking the foreign language recognition in the simultaneous interpretation scenario as an example, when the user selects Italian, only model A has the corresponding streaming recognition ability at this time, so this model can be selected as the cloud model. Similarly, if model A is selected for use, the audio pre-processing parameters and audio post-processing parameters can be configured according to the interface target format of this model, and the sending configuration parameters (related to the size of the audio sending packet) and sending interval time (related to the sending interval of the audio sending packet) in the audio link parameters can also be configured according to the audio transmission format of this model.

[0040] The real-time application method in the technology of this application can support multiple cloud processing models. The common cloud processing models and related features are shown in Table 5: Table 5

[0041] S140. The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a pre-set audio configuration table.

[0042] The audio relay device configures audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a pre-set audio configuration table. An audio configuration table is configured in the audio relay device, and the corresponding audio processing configuration information can be configured according to the audio configuration table. The audio processing configuration information is used to configure the processing flow of the audio.

[0043] In a specific embodiment, step S140 includes sub-steps: obtaining audio format information matching the connection mode in the audio configuration table; configuring audio link parameters according to the audio format information and the function selection information; configuring audio pre-processing parameters and audio post-processing parameters according to the audio format information and the model parameters; and combining the audio link parameters, the audio pre-processing parameters, and the audio post-processing parameters into the audio processing configuration information.

[0044] Specifically, the audio format information matching the connection mode can be obtained from the audio configuration table, and the audio configuration table includes the content involved in Tables 1 to 4. Then, according to the connection mode, the audio format, sampling rate, bit depth, channels, etc. supported by the connection mode can be correspondingly matched and obtained as the corresponding audio format information.

[0045] According to the audio format information and the function selection information, the audio link parameters can be correspondingly configured, and the audio link parameters can be used to indicate in what format and manner the collected external voice is transmitted. For example, if the function selection information is the AI voice conversion / voice beautification function, only a one-way link configuration is required, and the corresponding audio processing flow is as Figure 4 shown. Its audio link parameters can be described in detail as follows: collect the user's voice according to the audio format information → pre-processing → transmit to the cloud → cloud AI voice beautification / voice conversion processing → return the processed audio → post-processing → use it as the microphone input and transmit it to the application on the smart terminal side. Here, the large model is also the cloud model (the same below), and here this device is also the audio relay device (the same below). The corresponding specific embodiments include: (1) The user connects the audio relay device to the smart terminal; (2) Open the live application or game application on the smart terminal side; (3) Select the "voice beautification" or "voice conversion" function on the audio relay device; (4) Automatically configure the audio processing configuration information; (5) The user's voice is processed in real time through the cloud large model, and other users hear the voice after AI beautification or voice conversion. The end-to-end delay of the entire processing process is controlled within 200 ms, which does not affect normal communication. Specifically, its audio processing link is divided into three segments. The first segment refers to the delay time when the audio processed locally by the audio relay device is sent to the cloud server. The third segment refers to the delay time when the cloud server sends the audio processing to the audio relay device. The delay times of the first segment and the third segment can both be controlled within 200 ms. The second segment refers to the time for audio processing through the cloud model in the cloud server, and this processing time varies in the actual application of each model. The test results show that in the 4G / 5G network environment, the end-to-end delay of this solution is stable at 150 - 180 ms, and the audio quality is significantly improved.

[0046] If the function selection information is the real-time translation function / sensitive word filtering function, only a one-way link configuration is also required, and the corresponding audio processing flow is as Figure 5As shown, its audio link parameters can be described in detail as follows: Collect in-app audio according to audio format information → Pre-processing → Transmit to the cloud → Cloud translation / sensitive content detection → Obtain translation information / replace or block content → Speech synthesis → Return the processed audio → Post-processing → Use it as the microphone input and transmit it to the application on the smart terminal side. The corresponding specific embodiments include: (1) The user connects the audio relay device to the smart terminal; (2) Open the live application or conference application on the smart terminal side; (3) Select "Real-time Translation" on the audio relay device and set the target language or the "Content Filtering" function; (4) Automatically configure the audio processing configuration information; (5) The in-app audio is translated into the corresponding audio in the target language through the cloud / the voice content containing sensitive words will be automatically replaced with a prompt tone or a neutral expression. The test results show that when this solution is used for speech translation, the accuracy rate is high and the end-to-end delay is stably within 200 ms. Here, the end-to-end delay corresponds to the first and third segments of the audio processing link in the above description, excluding the processing time of the cloud model in the second segment.

[0047] If the function selection information is two-way call simultaneous interpretation, two-way link configuration is required, and the corresponding audio processing process is as Figure 6 shown. Its audio link parameters can be described in detail as follows: Collect local voice according to audio format information → Pre-processing → Translate to the target language in the cloud → Return the translated audio → Post-processing → Output to the application on the smart terminal side; Remote voice → Pre-processing → Translate to the local language in the cloud → Return the translated audio → Post-processing → Local playback. The corresponding specific embodiments include: (1) The user connects the audio relay device to the smart terminal; (2) Open the phone application or video conference application on the smart terminal side; (3) Select the "Simultaneous Interpretation" function on the audio relay device and set the source language and target language; (4) Automatically configure the audio processing configuration information; (5) Both parties can communicate in their respective native languages, and the system automatically completes real-time translation. The test results show that this solution supports real-time translation between 21 languages such as Chinese, English, Japanese, Korean, French, and German. The translation accuracy rate reaches over 85% in mainstream scenarios, and the end-to-end delay is controlled within 300 ms. Here, the end-to-end delay corresponds to the first and third segments of the audio processing link in the above description, excluding the processing time of the cloud model in the second segment. The number of languages only depends on the capabilities of the accessed models and can be expanded at any time.

[0048] Further, audio pre - processing parameters and audio post - processing parameters can be configured according to the corresponding audio format information. The source of the original audio may be a microphone, Bluetooth, or USB, and the format may be PCM, AAC, MP3, etc. The audio pre - processing parameters are the parameter information for pre - processing the audio output to the cloud server to make it meet the transmission requirements. The audio post - processing parameters are the parameter information for post - processing the processed audio from the cloud server to make it meet the local playback / output requirements.

[0049] For example, in the audio configuration table, the audio format information corresponding to the connection mode is: audio format SBC, sampling rate 48 kHz, bit depth 24 bits, and stereo channels; in the model parameters, the interface target format is: audio format PCM, sampling rate 16 kHz, bit depth 16 bits, and mono channel. Then, the audio pre - processing parameters need to be configured as audio format information → interface target format, and the configuration process of the audio post - processing parameters is similar.

[0050] Combine the audio link parameters, the audio pre - processing parameters, and the audio post - processing parameters into the audio processing configuration information. S150. The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information.

[0051] The audio relay device collects external voices according to the function selection information and forwards the external voices to the cloud server according to the audio processing configuration information. The audio relay device collects external voices corresponding to the function selection information. For example, if the function selection information is set to the AI voice - changing / voice - beautifying function, the user's audio is collected through the microphone as the external voice. If the function selection information is set to the real - time translation function / sensitive word filtering function, the audio relay device obtains the in - app audio from the terminal device through the communication connection as the external voice. The audio relay device forwards the obtained external voice to the cloud server according to the audio processing configuration information.

[0052] In a specific embodiment, step S150 includes sub - steps: transcoding the external voice according to the audio pre - processing parameters in the audio processing configuration information to obtain the corresponding transcoded audio; splitting the transcoded audio according to the sending configuration parameters of the audio link parameters in the audio processing configuration information to obtain the corresponding audio sending packets; and sending the audio sending packets to the cloud server according to the sending link of the audio link parameters in the audio processing configuration information.

[0053] Specifically, transcoding processing can be performed on the external voice according to the audio preprocessing parameters to obtain transcoded audio. Transcoding processing is to convert the audio format, sampling rate, bit depth, and channels of the external voice. For example, if the external voice is stereo and a mono audio needs to be transcoded and obtained as the transcoded audio, the two channels of audio in the stereo can be mixed to obtain the corresponding mono audio.

[0054] The transcoded audio is segmented according to the sending configuration parameters in the audio link parameters to obtain redundant audio sending packets. For the convenience of sending audio information, audio sending packets of a preset size can be segmented according to the transmission requirements. For example, if the preset size is set to 1 kB, audio sending packets of 1 kB size can be segmented and sent accordingly.

[0055] The audio sending packets are sent to the cloud server according to the sending link of the audio link parameters. Further, the audio sending packets can also be sent according to the sending interval time configured in the audio link parameters, so that the interval between the sending times of adjacent audio sending packets is not less than the sending interval time. For example, if the sending interval time is 20 ms, it can be ensured that the interval time between the sending of two audio sending packets is not less than 20 ms to avoid mutual interference during the sending and processing of audio sending packets.

[0056] S160. The cloud server processes the external voice through the cloud model to obtain audio processing information and feeds it back to the audio relay device.

[0057] The cloud server processes the external voice through the cloud model to obtain audio processing information and feeds it back to the audio relay device. The cloud server processes the received external voice by matching the obtained cloud model to obtain audio processing information. The cloud server feeds the obtained audio processing information back to the audio relay device, and the audio processing information includes the output audio. During the process of the cloud model processing the external voice, it includes audio feature extraction, audio recognition, recognition information processing, and speech synthesis. First, the audio features of the external voice are extracted; and the audio features are recognized to obtain corresponding recognition information. At this time, the recognition information is the initial text content recognized based on the external voice. The recognition information is processed, such as translation / sensitive word filtering, etc. to obtain processed information, and based on the processed information, speech synthesis (the synthesis parameters of the cloud model are involved in the speech synthesis process) is performed to obtain the corresponding audio processing information.

[0058] In a specific embodiment, before the cloud model processes the external voice to obtain audio processing information and feeds it back to the audio relay device, the following steps are further included: mapping the initial synthesis parameters according to a preset parameter mapping table to obtain mapping parameters corresponding to the cloud model; configuring the parameters of the cloud model according to the mapping parameters.

[0059] For the voice synthesis part of each model, the parameter ranges for speech rate, volume, and pitch are different. For the convenience of users, the synthesis parameters of the model can be set to be unified. This parameter configuration process is achieved through a parameter mapping table, and parameter transmission with the cloud model can be completed through the parameter mapping table. The initial synthesis parameters include: speech rate - "0" (the settable range is -10 to 10), volume - "10" (the settable range is 0 to 10), pitch - "0" (the settable range is -10 to 10).

[0060] The parameter mapping table includes the mapping relationship between the initial synthesis parameters and the synthesis parameters in each model. For example, if the cloud model is Model A, according to the initial synthesis parameters, the mapping parameters of this cloud model can be obtained as: speech rate - "0" (the settable range is -500 to 500), volume - "100" (the settable range is 0 to 100), pitch - "0" (the settable range is -500 to 500).

[0061] Configure the parameters of the cloud model according to the obtained mapping parameters, so that the parameters of the cloud model for voice synthesis match the initial synthesis parameters.

[0062] S170. The audio relay device outputs the audio processing information according to the audio processing configuration information.

[0063] The audio relay device outputs the audio processing information according to the audio processing configuration information. After receiving the audio processing information from the cloud server, the audio relay device can output the obtained audio processing information according to the audio processing configuration information.

[0064] In a specific embodiment, step S170 includes sub-steps: transcoding the output audio in the audio processing information according to the audio post-processing parameters in the audio processing configuration information to obtain a corresponding target audio; sending the target audio to the intelligent terminal and / or playing and outputting it according to the output link of the audio link parameters in the audio processing configuration information.

[0065] Specifically, the output audio in the audio processing information can be transcoded according to the audio post-processing parameters to obtain the target audio. The transcoding process here is similar to the process of transcoding external speech in the above steps, except that there are differences in the audio formats during transcoding. Further, the target audio is output according to the output link in the audio link parameters. According to the output link, the target audio can be sent to the applications of the intelligent terminal. For example, in a real-time conference application scenario, the speech of the local user on the terminal device side is converted into the target audio corresponding to the language of the remote user in the application, and the target audio is sent to the remote user through the application. It can also be that the target audio is directly played in the audio relay device according to the output link. For example, after the speech of the remote user is processed to obtain the target audio corresponding to the language of the local user on the terminal device side, the target audio can be directly output and played through the audio relay device, and the user listens to the target audio through the audio relay device. The target audio can also be sent to the intelligent terminal and directly output and played in the audio relay device according to the output link, so as to form a multi-directional stereo. Under normal use conditions, the user can freely switch to output the original sound or the target sound in the audio relay device, or mix the two audio signals and output and play the mixed audio signal.

[0066] Further, the audio processing information further includes text information. Then, the audio relay device can send the text information in the audio processing information to the intelligent terminal and / or perform display output according to the output link. The text information in the audio processing information is also the processing information obtained before the above cloud server performs speech synthesis. This processing information can be used as the text content to assist the local user to view and understand. The text information can be sent to the intelligent terminal for display, or directly displayed and output in the audio relay device, or can also be sent to the intelligent terminal for display and displayed and output in the audio relay device at the same time.

[0067] In the real-time application method based on cloud audio processing disclosed in the above embodiments, the method includes: if function selection information input by the intelligent terminal is received, the connection mode is obtained, the audio processing configuration information is obtained according to the function selection information, the model parameters and the connection mode, the function selection information is sent to the cloud server, and the external speech is collected and sent to the cloud server. The cloud server processes the external speech according to the cloud model to obtain the audio processing information and feeds it back to the audio relay device. The audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, which solves the problem that it is difficult to apply large model audio processing to mobile terminals in a portable manner; through audio relay and streaming processing technologies, it realizes the seamless combination of the speech processing ability of the cloud large model and the instant messaging scenario, providing a good user experience for users.

[0068] An embodiment of the present invention further provides a real-time application system based on cloud audio processing, as Figure 7 shown. The real-time application system 100 of the real-time application system based on cloud audio processing includes a connection mode acquisition unit 210, a sending unit 220, an audio processing configuration information acquisition unit 230, a forwarding unit 240, and an output unit 250 configured in an audio relay device. The real-time application system further includes a model parameter sending unit 310 and a processing unit 320 configured in a cloud server 30. The audio relay device establishes network connections with a smart terminal and the cloud server respectively to implement data information transmission. The real-time application system based on cloud audio processing is used to execute any one of the foregoing embodiments of the real-time application method based on cloud audio processing. Specifically, please refer to Figure 7 , Figure 7 which is a schematic block diagram of the real-time application system based on cloud audio processing provided by an embodiment of the present invention.

[0069] The connection mode acquisition unit 210 is configured to receive function selection information from the smart terminal and acquire the connection mode with the smart terminal.

[0070] The sending unit 220 is configured to send the function selection information to the cloud server.

[0071] The model parameter sending unit 310 is configured to acquire model parameters of a cloud model matching in a preset model library according to the function selection information and send them to the audio relay device.

[0072] The audio processing configuration information acquisition unit 230 is configured to configure audio processing configuration information corresponding to the function selection information, the connection mode, and the model parameters according to a preset audio configuration table.

[0073] The forwarding unit 240 is configured to collect external voices according to the function selection information and forward the external voices to the cloud server according to the audio processing configuration information.

[0074] The processing unit 320 is configured to process the external voices through the cloud model to obtain audio processing information and feedback it to the audio relay device.

[0075] The output unit 250 is configured to output the audio processing information according to the audio processing configuration information.

[0076] In the real-time application system based on cloud audio processing provided by the embodiments of the present invention, the above real-time application method based on cloud audio processing is applied. If function selection information input by a smart terminal is received, the connection mode is obtained. According to the function selection information, model parameters, and connection mode, audio processing configuration information is obtained. The function selection information is sent to the cloud server, and external voice is collected and sent to the cloud server. The cloud server processes the external voice according to the cloud model to obtain audio processing information and feedbacks it to the audio relay device. The audio relay device outputs the audio processing information according to the audio processing configuration information. The above method provides a low-cost and highly compatible cloud audio real-time processing solution, solving the problem that it is difficult to apply large model audio processing to mobile terminals in a portable manner; through audio relay and streaming processing technologies, seamless integration of the voice processing ability of the cloud large model and the instant messaging scenario is achieved, providing a good user experience for users.

[0077] Each unit module in the above real-time application system based on cloud audio processing can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 8 shown.

[0078] Please refer to Figure 8 , Figure 8 which is a schematic block diagram of the computer device provided by the embodiments of the present invention. This computer device can be an audio relay device and a cloud server for executing the real-time application method based on cloud audio processing to implement real-time application of audio processing based on a cloud large model.

[0079] Refer to Figure 8 , the computer device 500 includes a processor 502, a memory, and a communication interface 505 connected through a communication bus 501. Among them, the memory can include a storage medium 503 and an internal memory 504.

[0080] The storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can be made to execute the real-time application method based on cloud audio processing. Among them, the storage medium 503 can be a volatile storage medium or a non-volatile storage medium.

[0081] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0082] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute the real-time application method based on cloud audio processing.

[0083] The communication interface 505 is used for network communication, such as providing data information transmission, etc. Those skilled in the art can understand that Figure 8 The structure shown in Figure 8 is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0084] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the corresponding functions in the above-mentioned real-time application method based on cloud audio processing.

[0085] Those skilled in the art can understand that Figure 8 The embodiment of the computer device shown in Figure 8 does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are the same as those in Figure 8 the embodiment shown, and will not be elaborated here.

[0086] It should be understood that in the embodiments of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0087] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps included in the above-mentioned real-time application method based on cloud audio processing are implemented.

[0088] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0089] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, or units with the same function can be aggregated into one unit. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.

[0090] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the objectives of the embodiments of the present invention.

[0091] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0092] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes.

[0093] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A real-time application method based on cloud audio processing, characterized in that: The method is applied in a real-time application system, which includes an audio transfer device and a cloud server; The audio transfer device establishes network connections with the smart terminal and the cloud server respectively to realize the transmission of data information, and the method includes: If the audio transfer device receives the input function selection information, it obtains the connection mode between the audio transfer device and the smart terminal; The audio transfer device sends the function selection information to the cloud server; The cloud server obtains the model parameters of the matching cloud model in the preset model library according to the function selection information and sends them to the audio transfer device; The audio transfer device configures the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table; The audio transfer device collects external voice according to the function selection information and forwards the external voice to the cloud server according to the audio processing configuration information; The cloud server processes the external voice through the cloud model to obtain audio processing information which is fed back to the audio transfer device; The audio transfer device outputs the audio processing information according to the audio processing configuration information.

2. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The acquiring of the connection mode with the smart terminal includes: Acquire a type corresponding to a connection port between the intelligent terminals as a connection type; Obtain the driver protocol version corresponding to the connection port; The connection type and the driver protocol version are combined into the connection mode.

3. The real-time application method based on cloud audio processing according to claim 1 or 2, characterized in that: The configuring the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameter according to the preset audio configuration table includes: Acquire audio format information matching the connection mode in the audio configuration table; Obtaining audio link parameters according to the corresponding configuration of the audio format information and the function selection information; Obtaining audio pre-processing parameters and audio post-processing parameters according to the audio format information and the corresponding configuration of the model parameters; The audio link parameter, the audio pre-processing parameter, and the audio post-processing parameter are combined into the audio processing configuration information.

4. The real-time application method based on cloud audio processing according to claim 1 or 2, characterized in that: The forwarding the external voice to the cloud server according to the audio processing configuration information includes: Performing transcoding processing on the external voice according to the audio pre-processing parameters in the audio processing configuration information to obtain corresponding transcoded audio; Segment the transcoded audio according to the sending configuration parameters of the audio link parameters in the audio processing configuration information to obtain corresponding audio sending packets; The audio sending package is sent to the cloud server according to the sending link of the audio link parameter in the audio processing configuration information.

5. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The step of acquiring model parameters of a matching cloud model in a preset model library according to the function selection information and sending the model parameters to the audio transfer device includes: According to the audio path of the function selection information, a model matching the audio path in the model library is obtained as a candidate model; The function selection information is matched with the application function of each candidate model to obtain a candidate model whose application function matches the function selection information as a cloud model.

6. The real-time application method based on cloud audio processing according to claim 1 or 5, characterized in that: Before the external voice is processed by the cloud model to obtain audio processing information and then fed back to the audio transfer device, the method further includes: Mapping the initial synthesis parameters according to a preset parameter mapping table to obtain mapping parameters corresponding to the cloud model; Parameters of the cloud model are configured according to the mapping parameters.

7. The real-time application method based on cloud audio processing according to claim 1, characterized in that: The outputting the audio processing information according to the audio processing configuration information includes: Performing transcoding processing on the output audio in the audio processing information according to the audio post-processing parameters in the audio processing configuration information to obtain corresponding target audio; The target audio is sent to the smart terminal and / or played and output according to the output link of the audio link parameter in the audio processing configuration information.

8. A real-time application system based on cloud audio processing, characterized in that: The real-time application system includes a connection mode acquisition unit, a sending unit, an audio processing configuration information acquisition unit, a forwarding unit and an output unit configured in the audio transfer device, and the real-time application system also includes a model parameter sending unit and a processing unit configured in the cloud server. The audio transfer device establishes a network connection with the smart terminal and the cloud server respectively to realize the transmission of data information. The real-time application system is used to execute the real-time application method based on cloud audio processing according to any one of claims 1 to 7; The connection mode acquisition unit is used to receive function selection information from the smart terminal and acquire the connection mode between the smart terminal and the smart terminal; The sending unit is used to send the function selection information to the cloud server; The model parameter sending unit is used to obtain the model parameters of the matching cloud model in the preset model library according to the function selection information and send them to the audio transfer device; The audio processing configuration information acquisition unit is used to configure the audio processing configuration information corresponding to the function selection information, the connection mode and the model parameters according to a preset audio configuration table; The forwarding unit is used to collect external voice according to the function selection information and forward the external voice to the cloud server according to the audio processing configuration information; The processing unit is used to process the external voice through the cloud model to obtain audio processing information and feed it back to the audio transfer device; The output unit is used to output the audio processing information according to the audio processing configuration information.

9. A computer device, characterized in that: The device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of the real-time application method based on cloud audio processing described in any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the real-time application method based on cloud audio processing are implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • IMS audio and video transcoding and transmission control system and implementation method

    CN113301025A

  • Splicing persistent connections

    US20020120743A1