Audio data processing method, device and system and electronic equipment
By dynamically adjusting compression parameters based on the audio data type, the problem of large data volume in Bluetooth audio transmission is solved, enabling fast transmission of audio data between different devices.
Patent Information
- Application Number
- CN202410925812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2026-01-13
AI Technical Summary
In the existing Bluetooth audio transmission specifications, the compression parameters are fixed, resulting in a large amount of compressed data for both voice and non-voice signals, which consumes a lot of bandwidth and affects the fast transmission of audio data between different devices.
The compression parameters are dynamically adjusted based on the type of audio data, and different compression ratios are used to compress speech and non-speech signals to reduce the total amount of compressed audio data.
While ensuring the quality of the voice signal, the amount of audio data transmitted was reduced, and the transmission speed of audio data between different devices was increased.
Smart Images

Figure CN121334633A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer and network communication technology, and in particular to an audio data processing method, apparatus, system and electronic device. Background Technology
[0002] Bluetooth technology enables short-range wireless connections between different devices. Bluetooth provides various audio transmission specifications, such as the Advanced Audio Distribution Profile (A2DP), which defines high-quality (stereo or mono) wireless transmission methods; and the Audio / Video Remote Control Profile (AVRCP), which provides features for remotely controlling audio or video devices, such as play, pause, and volume control. However, these specifications have certain limitations, such as limitations in sound quality, encoding, and bandwidth. Furthermore, the compression algorithms used in these standards result in relatively large amounts of compressed data. Summary of the Invention
[0003] This disclosure provides an audio data processing method, apparatus, system, and electronic device.
[0004] In a first aspect, embodiments of this disclosure provide an audio data processing method applied to a first terminal. The method includes: acquiring first audio data; determining target compression parameters based on target type information corresponding to the first audio data; compressing the first audio data based on the target compression parameters to obtain first compressed data; wherein the target type information includes a voice type or a non-voice type; and sending the first compressed data to a second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0005] Secondly, embodiments of this disclosure provide an audio data processing method applied to a second terminal. The method includes: acquiring first transmission data from a first terminal, wherein the second terminal and the first terminal are wirelessly connected via Bluetooth; decompressing first compressed data in the first transmission data based on compression parameter information indicating target compression parameters in the first transmission data to obtain second audio data, wherein the first compressed data is obtained by compressing the first audio data using the target compression parameters.
[0006] Thirdly, this disclosure provides an audio data processing device, disposed on a first terminal, the device comprising: a first acquisition unit, configured to acquire first audio data; a compression unit, configured to determine target compression parameters based on target type information corresponding to the first audio data, and compress the first audio data based on the target compression parameters to obtain first compressed data; wherein the target type information includes a voice type or a non-voice type; and a sending unit, configured to send the first compressed data to a second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0007] Fourthly, this disclosure provides an audio data processing device disposed on a second terminal. The device includes: a second acquisition unit, configured to acquire first transmission data from a first terminal, wherein the second terminal and the first terminal are wirelessly connected via Bluetooth; and a decompression unit, configured to decompress the first compressed data in the first transmission data based on compression parameter information indicating target compression parameters in the first transmission data to obtain second audio data.
[0008] Fifthly, embodiments of this disclosure provide an audio data processing system, including a first terminal and a second terminal; the first terminal and the second terminal are wirelessly connected via Bluetooth; the first terminal includes a data acquisition device, a data processing device, and a first Bluetooth device, wherein the data acquisition device acquires first audio data; the data processing device determines target compression parameters based on target type information corresponding to the first audio data, and compresses the first audio data based on the target compression parameters to obtain first compressed data; wherein the target type information includes a voice type or a non-voice type; the first Bluetooth device sends the first compressed data to the second terminal; the second terminal includes a second Bluetooth device and a second data processing device, the second Bluetooth device acquires first transmission data from the first terminal; the second data processing device decompresses the first compressed data in the first transmission data based on compression parameter information indicating target compression parameters in the first transmission data to obtain second audio data.
[0009] Sixthly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0010] The memory stores computer-executed instructions;
[0011] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the methods described in the first aspect, the second aspect, and various possible designs of the first and second aspects.
[0012] In a seventh aspect, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the methods described in the first aspect, the second aspect, and various possible designs of the first and second aspects.
[0013] Eighthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in the first aspect, the second aspect, and various possible designs of the first and second aspects.
[0014] The audio data processing method, apparatus, system, and electronic device provided in this embodiment determine target type information for the acquired first audio data through a first terminal, determine target compression parameters based on the voice or non-voice type indicated by the target type information, compress the first audio data based on the compression parameters to obtain first compressed data, and send the first compressed data to a second terminal wirelessly connected to the first terminal via Bluetooth. Therefore, for audio data streams, the compression parameters used for compression and transmission vary depending on the type of audio data when compressing and transmitting audio data sampled at different sampling periods. Using different target compression parameters for voice and non-voice audio data helps reduce the total amount of compressed audio data while ensuring minimal loss of the voice signal, thus facilitating rapid transmission of audio data between different devices. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 Flowchart of the audio data processing method provided in the embodiments of this disclosure Figure 1 ;
[0017] Figure 2 Flowchart of the audio data processing method provided in the embodiments of this disclosure Figure 2 ;
[0018] Figure 3 Flowchart of the audio data processing method provided in the embodiments of this disclosure Figure 3 ;
[0019] Figure 4 A flowchart illustrating the audio data processing method provided in this embodiment of the disclosure;
[0020] Figure 5 This is a schematic diagram of the structure of the audio data processing apparatus provided in the embodiments of this disclosure;
[0021] Figure 6 This is a schematic diagram of the structure of the audio data processing apparatus provided in the embodiments of this disclosure;
[0022] Figure 7 This is a schematic diagram of the structure of an audio data processing system provided in an embodiment of the present disclosure;
[0023] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] Devices supporting Bluetooth wireless communication technology can transmit data. For example, a mobile terminal and a Bluetooth headset can transmit audio data. The Bluetooth headset can collect audio signals and transmit them to the mobile terminal. Before transmitting audio data from the Bluetooth headset to the mobile terminal, the collected audio data is usually compressed. In one embodiment, the standard specifications provided by Bluetooth technology can be used to compress the audio signal before transmission. However, in the compression schemes provided by the aforementioned standard specifications, the compression parameters are fixed. Both voice and non-voice signals need to be compressed using the set compression parameters. Typically, to ensure the quality of the restored voice signal, compression parameters with good voice signal restoration quality are used. Therefore, the compressed data obtained from audio data compression is large, and transmitting compressed data occupies a large data transmission bandwidth, which is not conducive to the rapid transmission of audio data between different devices.
[0026] To address the aforementioned issues, this disclosure can determine target compression parameters based on the target type information of the acquired first audio data, and then compress the audio data according to the target compression parameters. This allows for the application of different target compression parameters to audio signals of speech type and non-speech type, which helps to reduce the total amount of audio compressed data while ensuring minimal loss of the speech signal, and facilitates rapid transmission of audio data between different devices.
[0027] refer to Figure 1, Figure 1 Flowchart of the audio data processing method provided in the embodiments of this disclosure Figure 1 The method of this embodiment can be applied to a first terminal, and the audio data processing method includes:
[0028] S101: Obtain the first audio data.
[0029] The first terminal can be any terminal device that supports Bluetooth technology, and a Bluetooth device can be installed in the first terminal device. The first terminal can wirelessly connect with other terminal devices that support Bluetooth technology via Bluetooth.
[0030] Schematic illustration: The first terminal mentioned above can be a wearable device, such as headphones, smart glasses, smart bracelets, or smartwatches. Furthermore, the wearable device may also include a processor for processing audio data, in which audio data processing algorithms can be run. These algorithms may include, for example, audio classification algorithms and audio data compression algorithms.
[0031] In some embodiments, the first terminal described above may be equipped with an audio signal acquisition device, such as a microphone. The audio signal acquisition device may periodically acquire audio signals, convert the acquired audio signals into analog electrical signals, and then sample and quantize the analog electrical signals to obtain first audio data.
[0032] S102: Determine the target compression parameters based on the target type information corresponding to the first audio data, and compress the first audio data based on the target compression parameters to obtain the first compressed data; wherein, the target type information includes speech type or non-speech type.
[0033] The aforementioned first terminal can parse the type of the first audio data to obtain target type information. For each piece of first audio data, the target type information indicates whether the first audio data is speech or non-speech.
[0034] In some implementations, the aforementioned voice types include human voices, such as voices spoken by a user using the first terminal, and non-voice types include ambient sounds.
[0035] Ambient sounds can be any sounds other than human voices, such as the sounds of musical instruments, background music, and environmental noise.
[0036] In some implementations, the audio data processing method further includes the following steps:
[0037] First, the first audio data is classified based on a preset audio classification algorithm.
[0038] Secondly, the target type information of the first audio data is determined based on the classification results.
[0039] In these embodiments, a preset audio classification algorithm can be set in the first terminal to classify the first audio data.
[0040] To illustrate, the aforementioned preset audio classification algorithm can determine the spectrogram of the first audio data, and then determine the category of the first audio data based on the spectrogram and the pre-acquired spectral information corresponding to different types of speech. For example, human speech audio is usually distributed within a certain frequency range (e.g., 500Hz to 4kHz), while non-speech signals (e.g., ambient sound) may cover a wider frequency range. If the spectrogram of the first audio data falls within the frequency band corresponding to speech audio, then the classification result of the first audio data is speech type. If the spectrogram of the first audio data covers a wider frequency range, then the classification result of the first audio data can be considered as non-speech type.
[0041] For example, the aforementioned preset speech classification algorithm can also count the zero-crossing rate of the first audio data to determine whether the first audio data is a speech signal or a non-speech signal. The zero-crossing rate is the number of times a signal crosses from positive to negative or from negative to positive. Generally, non-speech signals have a higher zero-crossing rate, while speech signals have a lower zero-crossing rate. If the zero-crossing rate of the first audio data is greater than a preset zero-crossing rate threshold, the classification result corresponding to the first audio data is non-speech; otherwise, the classification result corresponding to the first audio data is speech.
[0042] For example, the aforementioned preset speech classification algorithm may also include a pre-trained machine learning algorithm. The machine learning algorithm extracts features from the first audio data, and then classifies the first audio data based on these features. The features of the first audio data include one or more of the following: spectral features, temporal features, and harmonic features. Spectral features include, but are not limited to: Mel-frequency cepstral coefficients, linear predictive coding, spectral center point, and spectral bandwidth. Temporal features include, but are not limited to: zero-crossing rate, energy, and entropy. Harmonic features: Speech-type audio data typically has a clear harmonic structure, while non-speech-type audio data may not have a clear harmonic structure.
[0043] In these embodiments, by setting a preset speech classification algorithm in the first terminal, the first audio data can be quickly classified to obtain the target type information corresponding to the first audio data.
[0044] The target compression parameters here can include the compression ratio. The compression ratio refers to the ratio of the compressed data size to the original data size, usually expressed as a percentage. If an audio file is originally 100MB and the compressed file is 50MB, then the compression ratio is 50%. A higher compression ratio results in a smaller compressed audio file size. However, if the compression ratio is too high, the loss of audio data is also greater, and the distortion of the audio data reconstructed from the compressed data is also greater.
[0045] In some embodiments, the compression ratio of audio data of the speech type is lower than that of audio data of the non-speech type.
[0046] Assume that audio data of speech type corresponds to the first compression ratio, and audio data of non-speech type corresponds to the second compression ratio.
[0047] If the first audio data is speech-type audio data, then the first audio data is compressed using a first compression ratio. If the first audio data is non-speech-type audio data, then the second compression ratio is used to compress the first audio data. That is, for non-speech-type audio data and speech-type audio data of the same size, due to the different compression ratios, the size of the first compressed data after compression of the non-speech-type audio data is smaller, and the size of the first compressed data after compression of the speech-type audio data is larger. In other words, in the audio data stream, speech data is compressed using the first compression ratio, and non-speech data is compressed using the second compression ratio. The second compression ratio is greater than the first compression ratio. Therefore, the amount of compressed data transmitted from the first terminal to the second terminal can be reduced, reducing data transmission bandwidth and facilitating fast audio data transmission from the first terminal to the second terminal.
[0048] S103: Send the first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0049] In this embodiment, the second terminal can be any terminal used by the user that supports Bluetooth technology. In some implementations, the second terminal can be a mobile terminal, such as a mobile phone, laptop, or tablet.
[0050] The first terminal and the second terminal can exchange data via Bluetooth wireless connection. The first terminal can send first compressed data to the second terminal via Bluetooth wireless connection. That is, the first terminal compresses the acquired first audio data and transmits it to the second terminal. The second terminal can decompress the aforementioned first compressed data to obtain the decompressed second audio data.
[0051] In this embodiment, the first terminal determines target type information for the acquired first audio data, determines target compression parameters based on the voice or non-voice type indicated by the target type information, compresses the first audio data based on the compression parameters to obtain first compressed data, and sends the first compressed data to a second terminal wirelessly connected to the first terminal via Bluetooth. Therefore, for audio data streams, the compression parameters used for compression transmission vary depending on the type of audio data when compressing audio data sampled at different sampling periods. Using different target compression parameters for voice and non-voice audio data helps reduce the total amount of compressed audio data while ensuring minimal loss of the voice signal, thus facilitating rapid transmission of audio data between different devices.
[0052] Please refer to Figure 2 , Figure 2 Flowchart of the audio data processing method provided in this disclosure Figure 2 This audio data processing method is applied to the first terminal, such as... Figure 2 As shown, the method includes the following steps:
[0053] S201: Obtain the first audio data.
[0054] For details on the implementation of step S201, please refer to [link / reference]. Figure 1 The description of step S101 in the illustrated embodiment will not be repeated here.
[0055] Audio data can be collected periodically, and the audio data collected in each period can be regarded as the first audio data corresponding to that period.
[0056] S202: In response to the target type information of the first audio data being speech type, determine the first proportion of the speech data.
[0057] S203: Determine the first target compression parameter corresponding to the first proportion.
[0058] In this embodiment, the first terminal can first determine the target type information corresponding to the first audio data. A preset audio classification algorithm set in the first terminal can be used to determine the target type information corresponding to the first audio data. The target type information indicates whether the first audio data belongs to a speech type or a non-speech type.
[0059] It is understandable that the first audio data of the speech type can consist entirely of speech data, or it can consist of speech data and non-speech data. For example, the first audio data of 20ms can include 0-10ms of non-speech data and 10ms-20ms of speech data.
[0060] For any sampling period, the first terminal can determine the first proportion of speech data in the first audio data.
[0061] The aforementioned first percentage can be greater than a first preset threshold and less than or equal to 1. For example, the first preset threshold may include 30%.
[0062] In one implementation, the first audio data may include two or more frames of audio data. For each frame of audio data, the corresponding audio data type can be determined. Then, based on the audio data types of each frame in the first audio data, the first proportion of speech data in the first audio data is determined. For example, if the first audio data includes six frames of audio data, and four of these frames are of speech data type, then the first proportion of speech data in the first audio data is approximately 67%.
[0063] As another implementation, the first audio data is sent to an audio classification algorithm. This algorithm determines a first score for the first audio data belonging to the speech type and a second score for it belonging to a non-speech type. The first and second scores can then be used to determine the first proportion of the speech data.
[0064] For example, the first score of the first audio data obtained by the audio classification algorithm is 70% for human voices and 30% for non-human voices. The first score of 70% can be taken as the first proportion of human voices.
[0065] A first mapping relationship between the proportion of voice data and compression parameters can be preset. After determining the first proportion of voice data in the first audio data, a first target compression parameter can be determined based on the aforementioned first mapping relationship and the first proportion. The first target compression parameter includes a first target compression ratio. The larger the first proportion, the smaller the first target compression ratio.
[0066] One implementation approach is to pre-configure compression parameters corresponding to multiple speech data percentage intervals, and then associate and store these intervals along with their corresponding compression parameters. After determining the first percentage of speech data in the first audio data, the target speech data percentage interval corresponding to that first percentage can be determined, and the compression parameters associated with and stored for that target speech data percentage interval can be used as the first target compression parameter.
[0067] Indicatively, the speech data here refers to human voice data. A first proportion of human voice data can be determined in the first audio data, and then a first target compression parameter for the human voice data can be determined based on the first proportion of human voice data.
[0068] S204: Compress the first audio data based on the first target compression parameters to obtain the first compressed data.
[0069] After obtaining the first target compression parameters, the first audio data can be compressed to obtain the first compressed data.
[0070] S205: Send first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0071] This embodiment describes how, when the first audio data belongs to the speech type, a first proportion corresponding to the speech data is determined, a first target compression parameter is determined based on the first proportion, and the first audio data is compressed according to the first target compression parameter. Therefore, the first target compression parameter can be different for the first audio data of the speech type acquired at different times. This achieves dynamic compression of the speech type audio data according to a dynamic compression ratio. Compressing the speech type first audio data according to dynamic compression parameters has two advantages: firstly, the dynamic compression ratio can effectively control the loudness fluctuation of the speech data, making the sound more balanced; secondly, since the first target compression parameter is determined based on the first proportion, it is beneficial to increase the compression ratio of the speech type audio data, which can further reduce the size of the compressed audio data.
[0072] exist Figure 1 and Figure 2 In some embodiments of the audio data processing method shown, the method further includes the following steps:
[0073] First, in response to the target type information being non-voice type, a second proportion of non-voice data is determined.
[0074] Secondly, determine the second target compression parameter corresponding to the second proportion.
[0075] Finally, the first audio data is compressed based on the second target compression parameters to obtain the first compressed data.
[0076] In these embodiments, the first terminal can first determine the target type information corresponding to the first audio data. Indicatively, a preset audio classification algorithm can be used to classify the first audio data, obtaining a first score indicating whether the first audio data belongs to the speech type and a second score indicating whether it belongs to a non-speech type. If the second score is greater than the first score, then the first audio data belongs to the non-speech type.
[0077] If the first audio data is a non-speech type, determine the second proportion of non-speech data in the first audio data.
[0078] The aforementioned second percentage can be greater than the second preset threshold and less than or equal to 1. For example, the first preset threshold here can be 50%.
[0079] As one implementation method, a second mapping relationship between the proportion of non-speech data and the compression parameters of non-speech data can be preset. After determining the second proportion, the second target compression parameter corresponding to the second proportion can be determined according to the above second mapping relationship.
[0080] The second target compression parameter includes the compression ratio. The compression ratio can be determined based on the second mapping relationship described above.
[0081] As one implementation method, multiple compression parameters corresponding to different non-speech data percentage intervals can be pre-configured, and these intervals, along with their corresponding compression parameters, can be stored together. After determining the second percentage of non-speech data in the first audio data, a target non-speech data percentage interval corresponding to the second percentage can be determined. Then, the compression parameters stored in association with the target non-speech data percentage interval are determined as the second target compression parameters corresponding to the second percentage.
[0082] Schematic illustration: Here, non-speech refers to ambient sound. A second proportion of ambient sound data can be determined from the first audio data, and then a second target compression parameter for the ambient sound data can be determined based on this second proportion.
[0083] This embodiment describes how, when the first audio data is non-speech type, a second proportion corresponding to the non-speech data is determined, a second target compression parameter is determined based on the second proportion, and the first audio data is compressed according to the second target compression parameter. Thus, the compression parameter corresponding to non-speech type audio data can be different at different times. This allows for dynamic compression of non-speech type audio data, helping to further reduce the size of the compressed data.
[0084] Please continue to refer to this. Figure 3 , Figure 3 Schematic flowchart of the audio data processing method provided in this disclosure Figure 3 This method is applied to the first terminal, such as... Figure 3 As shown, the method includes the following steps:
[0085] S301: Obtain the first audio data.
[0086] For details on the implementation of step S301, please refer to Figure 1 The description of step S101 in the illustrated embodiment will not be repeated here.
[0087] S302: In response to the first audio data being classified as a speech type based on the target type information corresponding to the first audio data, the first audio data is parsed to determine the target semantics of the first audio data.
[0088] S303: Determine the third target compression parameter corresponding to the target semantics.
[0089] S304: Compress the first speech data using the third target compression parameters to obtain the first compressed data.
[0090] In this embodiment, the first terminal can first determine the target type information corresponding to the first audio data. A preset audio classification algorithm set in the first terminal can be used to determine the target type information corresponding to the first audio data. The target type information indicates whether the first audio data belongs to the speech type or is not speech.
[0091] In some implementations, the first audio data can be sent to an audio classification algorithm, which can then determine whether the first audio data belongs to a speech type or a non-speech type.
[0092] If the first audio data belongs to the speech type, the semantics of the first audio data can be parsed.
[0093] In some embodiments, step S302 includes:
[0094] Identify preset keywords in the first audio data, and determine the target semantics based on the preset keywords; or...
[0095] Semantic understanding is performed on the first audio data, and the target semantics are determined based on the semantic understanding results.
[0096] In some application scenarios, preset keywords in the first audio data can be identified, and the target semantics can be determined based on the preset keywords.
[0097] The preset keywords here can be pre-set words that represent certain semantics, such as words used to wake up the target application or words used to instruct the closing of the target application. These preset keywords include custom wake words for the target application (e.g., a voice assistant) and end keywords for ending the interaction.
[0098] In these application scenarios, the first audio data is first converted into text, and one or more preset keywords are matched against this text. The target semantics of the first audio data are then determined based on the successfully matched target preset keywords.
[0099] In these application scenarios, the semantics of the first audio data can be determined by using preset keywords included in the first audio data, which can quickly determine the target semantics of the first audio data.
[0100] In some application scenarios, semantic understanding can be performed on the first audio data, and the target semantics can be determined based on the semantic understanding results.
[0101] In these application scenarios, the first audio data can be converted into text, and then semantic understanding can be performed based on the text. For example, the text can be segmented, labeled with parts of speech, named entity recognition, parsed, and semantically analyzed to obtain the target semantics of the first audio data.
[0102] As an example, the text corresponding to the first audio data can be input into a pre-trained semantic recognition model, and the semantic recognition model can output the target semantics corresponding to the first audio data.
[0103] Different target semantics can correspond to different third target compression parameters. The third target compression parameter can include the compression ratio. For example, if the target semantic of the first audio data is to wake up the target application, it needs to be transmitted to the second terminal quickly to obtain a response. Therefore, a larger compression ratio can be used to compress the first audio data, resulting in a smaller amount of compressed data, which allows for faster transmission to the second terminal.
[0104] For example, the target semantics of the first audio data instruct the target application to respond to a given question. In order for the target application to respond accurately to the question, the first audio data needs to be transmitted to the target application with minimal loss; therefore, a lower compression ratio can be used to compress the first audio data.
[0105] S305: Send first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0106] This example describes the target semantics of first audio data to determine the speech type, the third target compression parameters to determine based on the target semantics, and the compression of the first audio data based on the third target compression parameters. This helps the target application to quickly respond to or accurately reply to the first audio data based on the semantics.
[0107] exist Figure 1 , Figure 2 and Figure 3 In some embodiments of the audio data processing method provided in the illustrated examples, the method further includes the following steps:
[0108] First transmission data, including first compressed data, is constructed according to a custom data transmission protocol and sent to a second terminal. The first transmission data further includes at least one of compression parameter information indicating target compression parameters, voice audio start time information, and voice audio end time information.
[0109] In some application scenarios, the aforementioned compression parameter information includes target type information for the first audio data. This target type information indicates whether the first audio data is speech or non-speech. In this embodiment, the second terminal may pre-store compression parameters corresponding to speech or non-speech types.
[0110] In one example, the compression parameter information mentioned above includes the target compression parameters corresponding to the first audio data.
[0111] In another example, the compression parameter information mentioned above may also include target type information and target compression parameters.
[0112] By setting compression parameters in the first transmitted data, the second terminal can decompress the first transmitted data according to the compression parameters.
[0113] The aforementioned audio start time information can be time information relative to the starting time point, where the starting time point can be the start time of the first audio data.
[0114] Similarly, the end time information of the voice / audio can also be time information relative to the aforementioned start time point.
[0115] In these embodiments, the voice audio start time information and voice audio end time information in the first transmitted data help the second terminal extract and parse the voice data from the first compressed data.
[0116] In some embodiments, the first transmitted data does not contain encapsulated data.
[0117] After obtaining the target compression parameters, the first audio data can be compressed using the OPUS compression format to obtain the first compressed data. Using the OPUS compression format to compress the first audio data yields the first compressed data without encapsulation, thereby further reducing the size of the first transmitted data.
[0118] exist Figure 1 , Figure 2 and Figure 3 In some embodiments of the audio data processing method provided in the illustrated examples, the method further includes the following steps:
[0119] Transmit one or more of the following to the second terminal based on a custom data transmission protocol:
[0120] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0121] In this embodiment, one or more of the aforementioned data can be sent to the second terminal via a custom data transmission protocol before acquiring the first audio data. Alternatively, one or more of the aforementioned data can be transmitted along with the first data transmission.
[0122] By transmitting the aforementioned data to the second terminal, the second terminal can decompress the first transmitted data based on the aforementioned data.
[0123] Please refer to Figure 4 , Figure 4 A flowchart illustrating the audio data processing method provided in this disclosure is shown. This audio data processing method is applied to a second terminal and includes the following steps:
[0124] S401: Obtain first transmission data from the first terminal, wherein the second terminal is wirelessly connected to the first terminal via Bluetooth.
[0125] S402: Based on the compression parameter information indicating the target compression parameters in the first transmission data, decompress the first compressed data in the first transmission data to obtain the second audio data, wherein the first compressed data is obtained by compressing the first audio data using the target compression parameters.
[0126] The second terminal mentioned above can be a mobile terminal or a fixed terminal that supports Bluetooth communication. The mobile terminal can be, for example, a mobile phone, a laptop, or a tablet. The fixed terminal can be, for example, a desktop computer.
[0127] The aforementioned first terminal can be a wearable device, such as headphones, smart bracelets, smart glasses, smartwatches, etc.
[0128] The first terminal can send first transmission data to the second terminal. The first transmission data includes first compressed data. The first transmission data also includes compression parameter information. Based on the compression parameter information, the target compression parameters corresponding to the first compressed data can be determined.
[0129] After determining the target compression parameters, the second terminal can decompress the first compressed data.
[0130] The steps executed by the first terminal can be referred to Figures 1-3 The descriptions in the illustrated embodiments are not repeated here.
[0131] In some embodiments, the first audio data is collected by a first terminal; and / or, the target type information includes a voice type or a non-voice type.
[0132] In these embodiments, an audio data acquisition device, such as a microphone, may be provided in the first terminal. This audio data acquisition device can periodically acquire audio data. The first terminal acquires first audio data, compresses the acquired first audio data, and transmits it to the second terminal in real time, enabling the second terminal to respond quickly to the first audio data.
[0133] The target type information corresponding to the first audio data includes either speech type or non-speech type. The speech type includes human voice, and the non-speech type includes ambient sound.
[0134] The target compression parameters for the first audio data of speech type and the first audio data of non-speech type can be different. This is beneficial to reduce the amount of audio compressed data while ensuring that the speech signal has minimal loss, and helps to achieve fast transmission of audio data between the first terminal and the second terminal.
[0135] In some embodiments, the target compression parameters are related to the target type information corresponding to the first audio data, and the first transmission data includes the first compressed data after the first audio data is compressed by the target compression parameters.
[0136] In these embodiments, the target compression parameters are related to the target type information of the first audio data. The target type information indicates whether the first audio data belongs to a speech type or a non-speech type. That is, different types of first audio data correspond to different target compression parameters.
[0137] The target compression parameters include the compression ratio. The compression ratio for the first audio data of the speech type is less than the compression ratio for the first audio data of the non-speech type. This reduces the total size of the compressed audio data in the audio stream.
[0138] In some embodiments, the first transmitted data further includes at least one of voice audio start time information and voice audio end time information.
[0139] The aforementioned audio start time information can be time information relative to the starting time point, where the starting time point can be the start time of the first audio data.
[0140] Similarly, the end time information of the voice / audio can also be time information relative to the aforementioned start time point.
[0141] In these embodiments, the voice audio start time information and voice audio end time information in the first transmitted data help the second terminal extract and parse the voice data from the first compressed data.
[0142] In some embodiments, second transmission data is obtained from a first terminal based on a custom data transmission protocol, and the second transmission data further includes one or more of the following:
[0143] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0144] In this embodiment, one or more of the aforementioned data can be sent to the second terminal via a custom data transmission protocol before acquiring the first audio data. Alternatively, one or more of the aforementioned data can be transmitted along with the first data transmission.
[0145] The second terminal decompresses the first transmitted data based on one or more of the aforementioned data.
[0146] In some embodiments, the unencapsulated data in the first transmitted data can be obtained by compressing the first audio data using the OPUS compression format. Compressing the first audio data using the OPUS compression format yields compressed first data without encapsulation, thereby further reducing the size of the first transmitted data.
[0147] In some embodiments, the compression parameter information in step S402 includes target type information of the first audio data, wherein the target type information indicates whether the first audio data is a speech type or a non-speech type. In this embodiment, the second terminal may pre-store compression parameters corresponding to speech types or non-speech types respectively. After receiving the target type information, if the target type information indicates a speech type, the target compression parameter corresponding to the speech type is obtained from the locally stored compression parameters. If the target type information indicates a non-speech type, the target compression parameter corresponding to the non-speech type is obtained from the locally stored compression parameters.
[0148] In some implementations, the compression parameter information mentioned above includes the target compression parameters corresponding to the first audio data.
[0149] In some embodiments, the method further includes: transmitting the second audio data to a target application on a second terminal, wherein the target application generates a response of the second audio data.
[0150] In these embodiments, a target application may run on the second terminal. The target application is used to interact with a user using the first terminal device.
[0151] The target application can generate a response based on the second audio data. This response may include audio and / or text responses.
[0152] In these embodiments, the response of generating second audio data through the target application enables human-computer interaction between the user and the target application, making it easier for the user to obtain corresponding services from the target application.
[0153] In some implementations, the response is generated by the target application performing semantic understanding on the second audio data and based on the semantic understanding results; wherein the response includes response audio and / or response text.
[0154] In these implementations, after receiving the second audio data, the target application can perform semantic understanding on the second audio data. For example, the target application can invoke a natural language semantic understanding model to perform semantic understanding on the second audio data and obtain a semantic understanding result. Based on the semantic understanding result, a response audio and / or response text can be generated.
[0155] In some examples, the target application connects to a web server and transmits second audio data to the web server for semantic understanding.
[0156] In these examples, the aforementioned web server can be configured with various natural language processing (NLP) models. These NLP models can convert the input second audio data into text, perform semantic understanding on the text, and then output the semantic understanding results. The web server can then send the semantic understanding results output by the NLP models to the target application, which can generate response audio and / or response text based on the semantic understanding results.
[0157] In some application scenarios, the target application can generate response audio based on semantic understanding results and transmit the response audio to the first terminal for playback via a second terminal. This allows the user to hear the response audio through the first terminal.
[0158] In some application scenarios, the target application can also generate response text based on the semantic understanding results and display the response text on the target application's interface. This allows users to view the response text on a second terminal.
[0159] In these implementations, the second audio data is sent to the target application via the second terminal, and the target application generates a response containing the second audio data. This enables interaction between the user and the target application via the first terminal, providing convenience for the user to obtain information through the target application.
[0160] In this embodiment, the second terminal obtains first transmitted data from the first terminal and decompresses the first compressed data in the first transmitted data according to the compression parameter information in the first transmitted data to obtain second audio data. The first compressed data is obtained by compressing the first audio data according to target compression parameters. The amount of the first transmitted data is relatively small, therefore it can be quickly obtained from the first terminal to enable a rapid response to the first audio data.
[0161] Corresponding to the above text Figure 1 The audio data processing method of the embodiment, Figure 5 This is a schematic structural block diagram of an audio data processing apparatus provided according to an embodiment of the present disclosure. For ease of explanation, only the parts relevant to the embodiments of the present disclosure are shown. (Refer to...) Figure 5 The device 50 includes: a first acquisition unit 501, a compression unit 502, and a transmission unit 503. Among them,
[0162] The first acquisition unit 501 is used to acquire the first audio data;
[0163] Compression unit 502 is used to determine target compression parameters based on target type information corresponding to the first audio data, and compress the first audio data based on the target compression parameters to obtain first compressed data; wherein, the target type information includes speech type or non-speech type;
[0164] The transmitting unit 503 is used to send first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0165] In some embodiments, the compression unit 502 is further configured to:
[0166] In response to the target type information being voice type, determine the first proportion of voice data;
[0167] Determine the first target compression parameter corresponding to the first proportion;
[0168] The first audio data is compressed based on the first target compression parameters.
[0169] In some embodiments, the compression unit 502 is further configured to:
[0170] In response to the target type information being non-voice type, a second proportion of non-voice data is determined;
[0171] Determine the second target compression parameter corresponding to the second proportion;
[0172] The first audio data is compressed based on the second target compression parameters.
[0173] In some embodiments, the compression unit 502 is further configured to:
[0174] In response to the target type information being speech type, the first audio data is parsed to determine the target semantics of the first audio data;
[0175] Determine the third target compression parameters corresponding to the target semantics;
[0176] The first speech data is compressed using the third target compression parameters.
[0177] In some embodiments, the compression unit 502 is further configured to:
[0178] Identify preset keywords in the first audio data, and determine the target semantics based on the preset keywords; or...
[0179] Semantic understanding is performed on the first audio data, and the target semantics are determined based on the semantic understanding results.
[0180] In some embodiments, the device 50 further includes a sorting unit (not shown), the sorting unit being used for:
[0181] The first audio data is classified based on a preset audio classification algorithm;
[0182] The target type information of the first audio data is determined based on the classification results.
[0183] In some embodiments, the speech type includes human voice, and the non-speech type includes ambient sound.
[0184] In some embodiments, the target compression parameter includes a compression ratio, wherein...
[0185] The compression ratio of speech-type audio data is lower than that of non-speech-type audio data.
[0186] In some embodiments, the sending unit 503 is further configured to:
[0187] First transmission data, including first compressed data, is constructed according to a custom data transmission protocol and sent to a second terminal. The first transmission data further includes at least one of compression parameter information indicating target compression parameters, audio start time information of voice type, and audio end time information of voice type.
[0188] In some embodiments, the sending unit 503 is further configured to: transmit one or more of the following to the second terminal based on a custom data transmission protocol:
[0189] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0190] In some embodiments, the first transmitted data does not contain encapsulated data.
[0191] Corresponding to the above text Figure 4 The audio data processing method of the embodiment, Figure 6 This is a schematic structural block diagram of an audio data processing apparatus provided according to an embodiment of the present disclosure. For ease of explanation, only the parts relevant to the embodiments of the present disclosure are shown. (Refer to...) Figure 6 The device 60 includes a second acquisition unit 601 and a decompression unit 602. Wherein,
[0192] The second acquisition unit 601 is used to acquire first transmission data from the first terminal, wherein the second terminal and the first terminal are wirelessly connected via Bluetooth.
[0193] The decompression unit 602 is used to decompress the first compressed data in the first transmitted data based on the compression parameter information indicating the target compression parameters in the first transmitted data to obtain the second audio data, wherein the first compressed data is obtained by compressing the first audio data using the target compression parameters.
[0194] In some embodiments, the device 60 further includes a response unit (not shown), which is used to: transmit the second audio data to a target application on a second terminal, and have the target application generate a response of the second audio data.
[0195] In some embodiments, the response is generated by the target application performing semantic understanding on the second audio data and based on the semantic understanding result; wherein the response includes response audio and / or response text.
[0196] In some embodiments, the target application is connected to a web server, and the target application transmits second audio data to the web server for semantic understanding.
[0197] In some embodiments, the target compression parameters are related to the target type information corresponding to the first audio data, and the first transmission data includes the first compressed data after the first audio data is compressed by the target compression parameters.
[0198] In some embodiments, the first audio data is collected by a first terminal; and / or, the target type information includes a voice type or a non-voice type.
[0199] In some embodiments, the first transmitted data further includes at least one of voice audio start time information and voice audio end time information.
[0200] In some embodiments, second transmission data is obtained from a first terminal based on a custom data transmission protocol, and the second transmission data further includes one or more of the following:
[0201] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0202] In some embodiments, the first transmitted data does not contain encapsulated data.
[0203] Please refer to Figure 7 , Figure 7 This is a schematic diagram of the audio data processing system provided in this disclosure. Figure 7 As shown, the audio data processing system includes a first terminal and a second terminal. Among them,
[0204] The first terminal includes a data acquisition device, a data processing device, and a first Bluetooth device. The data acquisition device acquires first audio data. The data processing device determines target compression parameters based on target type information corresponding to the first audio data, and compresses the first audio data based on the target compression parameters to obtain first compressed data. The target type information includes voice type or non-voice type. The first Bluetooth device sends the first compressed data to the second terminal.
[0205] The second terminal includes a second Bluetooth device and a second data processing device. The second Bluetooth device obtains first transmission data from the first terminal. The second data processing device decompresses the first compressed data in the first transmission data based on the compression parameter information indicating the target compression parameter in the first transmission data to obtain second audio data.
[0206] The data acquisition device in the first terminal can acquire the first audio data. For details on how the data acquisition device acquires the first audio data, please refer to... Figure 1 The following is a description of step S101 in the illustrated embodiment.
[0207] The data processing device of the first terminal determines the target compression parameters based on the target type information corresponding to the first audio data, and compresses the first audio data according to the target compression parameters to obtain the first compressed data. For a description of this process, please refer to [link to relevant documentation]. Figures 1-3 The relevant parts of the illustrated embodiment will not be described in detail here.
[0208] The first terminal can be a wearable device, such as headphones. The second terminal can be a mobile terminal, such as a mobile phone, laptop, or tablet. To implement the above embodiments, this disclosure also provides an electronic device.
[0209] refer to Figure 8The diagram illustrates a structural schematic of an electronic device 800 suitable for implementing embodiments of the present disclosure. The electronic device 800 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0210] like Figure 8 As shown, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0211] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0212] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0213] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0214] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0215] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0216] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0217] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0218] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0219] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0220] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0221] In a first aspect, according to one or more embodiments of the present disclosure, an audio data processing method is provided, comprising: acquiring first audio data;
[0222] The target compression parameters are determined based on the target type information corresponding to the first audio data, and the first audio data is compressed based on the target compression parameters to obtain the first compressed data; wherein, the target type information includes speech type or non-speech type;
[0223] First compressed data is sent to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0224] According to one or more embodiments of this disclosure, determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes:
[0225] In response to the target type information being voice type, determine the first proportion of voice data;
[0226] Determine the first target compression parameter corresponding to the first proportion;
[0227] The first audio data is compressed based on the first target compression parameters.
[0228] According to one or more embodiments of this disclosure, determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes:
[0229] In response to the target type information being non-voice type, a second proportion of non-voice data is determined;
[0230] Determine the second target compression parameter corresponding to the second proportion;
[0231] The first audio data is compressed based on the second target compression parameters.
[0232] According to one or more embodiments of this disclosure, determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes:
[0233] In response to the target type information being speech type, the first audio data is parsed to determine the target semantics of the first audio data;
[0234] Determine the third target compression parameters corresponding to the target semantics;
[0235] The first speech data is compressed using the third target compression parameters.
[0236] According to one or more embodiments of this disclosure, parsing first audio data to determine the target semantics of the first audio data includes:
[0237] Identify preset keywords in the first audio data, and determine the target semantics based on the preset keywords; or...
[0238] Semantic understanding is performed on the first audio data, and the target semantics are determined based on the semantic understanding results.
[0239] According to one or more embodiments of this disclosure, the method further includes:
[0240] The first audio data is classified based on a preset audio classification algorithm;
[0241] The target type information of the first audio data is determined based on the classification results.
[0242] According to one or more embodiments of this disclosure, the speech type includes human voice, and the non-speech type includes ambient sound.
[0243] According to one or more embodiments of this disclosure, the target compression parameters include a compression ratio, wherein,
[0244] The compression ratio of speech-type audio data is lower than that of non-speech-type audio data.
[0245] According to one or more embodiments of this disclosure, sending first compressed data to a second terminal includes:
[0246] First transmission data, including first compressed data, is constructed according to a custom data transmission protocol and sent to a second terminal. The first transmission data further includes at least one of compression parameter information indicating target compression parameters, voice audio start time information, and voice audio end time information.
[0247] According to one or more embodiments of this disclosure, the method further includes: transmitting one or more of the following to a second terminal based on a custom data transmission protocol:
[0248] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0249] According to one or more embodiments of this disclosure, the first transmitted data does not contain encapsulated data.
[0250] Secondly, according to one or more embodiments of this disclosure, an audio data processing method is provided, comprising:
[0251] First transmission data is obtained from the first terminal, wherein the second terminal is wirelessly connected to the first terminal via Bluetooth;
[0252] Based on the compression parameter information indicating the target compression parameters in the first transmitted data, the first compressed data in the first transmitted data is decompressed to obtain the second audio data, wherein the first compressed data is obtained by compressing the first audio data using the target compression parameters.
[0253] According to one or more embodiments of this disclosure, the method further includes:
[0254] The second audio data is transmitted to the target application on the second terminal, and the target application generates a response containing the second audio data.
[0255] According to one or more embodiments of this disclosure, the response is generated by the target application performing semantic understanding on the second audio data and based on the semantic understanding result; wherein, the response includes response audio and / or response text.
[0256] According to one or more embodiments of this disclosure, a target application is connected to a web server, and the target application transmits second audio data to the web server for semantic understanding.
[0257] According to one or more embodiments of this disclosure, the target compression parameters are related to target type information corresponding to the first audio data, and the first transmission data includes first compressed data after the first audio data is compressed by the target compression parameters.
[0258] According to one or more embodiments of this disclosure, the first audio data is collected by a first terminal; and / or, the target type information includes a voice type or a non-voice type.
[0259] According to one or more embodiments of this disclosure, the first transmission data further includes at least one of voice audio start time information and voice audio end time information.
[0260] According to one or more embodiments of this disclosure, second transmission data is obtained from a first terminal based on a custom data transmission protocol, and the second transmission data further includes one or more of the following:
[0261] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0262] According to one or more embodiments of this disclosure, the first transmitted data does not contain encapsulated data.
[0263] Thirdly, according to one or more embodiments of this disclosure, an audio data processing apparatus is provided, disposed in a first terminal, the apparatus comprising:
[0264] The first acquisition unit is used to acquire the first audio data;
[0265] The compression unit is used to determine target compression parameters based on the target type information corresponding to the first audio data, and to compress the first audio data based on the target compression parameters to obtain first compressed data; wherein, the target type information includes speech type or non-speech type;
[0266] The transmitting unit is used to send first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0267] According to one or more embodiments of this disclosure, the compression unit is further configured to:
[0268] In response to the target type information being voice type, determine the first proportion of voice data;
[0269] Determine the first target compression parameter corresponding to the first proportion;
[0270] The first audio data is compressed based on the first target compression parameters.
[0271] According to one or more embodiments of this disclosure, the compression unit is further configured to:
[0272] In response to the target type information being non-voice type, a second proportion of non-voice data is determined;
[0273] Determine the second target compression parameter corresponding to the second proportion;
[0274] The first audio data is compressed based on the second target compression parameters.
[0275] According to one or more embodiments of this disclosure, the compression unit is further configured to:
[0276] In response to the target type information being speech type, the first audio data is parsed to determine the target semantics of the first audio data;
[0277] Determine the third target compression parameters corresponding to the target semantics;
[0278] The first speech data is compressed using the third target compression parameters.
[0279] According to one or more embodiments of this disclosure, the compression unit is further configured to:
[0280] Identify preset keywords in the first audio data, and determine the target semantics based on the preset keywords; or...
[0281] Semantic understanding is performed on the first audio data, and the target semantics are determined based on the semantic understanding results.
[0282] According to one or more embodiments of this disclosure, the device further includes a sorting unit, the sorting unit being used for:
[0283] The first audio data is classified based on a preset audio classification algorithm;
[0284] The target type information of the first audio data is determined based on the classification results.
[0285] According to one or more embodiments of this disclosure, the speech type includes human voice, and the non-speech type includes ambient sound.
[0286] According to one or more embodiments of this disclosure, the target compression parameters include a compression ratio, wherein,
[0287] The compression ratio of speech-type audio data is lower than that of non-speech-type audio data.
[0288] According to one or more embodiments of this disclosure, the transmitting unit is further configured to:
[0289] First transmission data, including first compressed data, is constructed according to a custom data transmission protocol and sent to a second terminal. The first transmission data further includes at least one of compression parameter information indicating target compression parameters, audio start time information of voice type, and audio end time information of voice type.
[0290] According to one or more embodiments of this disclosure, the sending unit is further configured to: transmit one or more of the following to the second terminal based on a custom data transmission protocol:
[0291] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0292] According to one or more embodiments of this disclosure, the first transmitted data does not contain encapsulated data.
[0293] Fourthly, according to one or more embodiments of this disclosure, an audio data processing apparatus is provided, disposed in a second terminal, the apparatus comprising:
[0294] The second acquisition unit is used to acquire first transmission data from the first terminal, wherein the second terminal and the first terminal are wirelessly connected via Bluetooth.
[0295] The decompression unit is used to decompress the first compressed data in the first transmitted data based on the compression parameter information indicating the target compression parameters in the first transmitted data, to obtain the second audio data.
[0296] According to one or more embodiments of this disclosure, the apparatus further includes a response unit, which is configured to: transmit the second audio data to a target application on a second terminal, and have the target application generate a response of the second audio data.
[0297] According to one or more embodiments of this disclosure, the response is generated by the target application performing semantic understanding on the second audio data and based on the semantic understanding result; wherein, the response includes response audio and / or response text.
[0298] According to one or more embodiments of this disclosure, a target application is connected to a web server, and the target application transmits second audio data to the web server for semantic understanding.
[0299] According to one or more embodiments of this disclosure, the target compression parameters are related to target type information corresponding to the first audio data, and the first transmission data includes first compressed data after the first audio data is compressed by the target compression parameters.
[0300] According to one or more embodiments of this disclosure, the first audio data is collected by a first terminal; and / or, the target type information includes a voice type or a non-voice type.
[0301] According to one or more embodiments of this disclosure, the first transmission data further includes at least one of voice audio start time information and voice audio end time information.
[0302] According to one or more embodiments of this disclosure, second transmission data is obtained from a first terminal based on a custom data transmission protocol, and the second transmission data further includes one or more of the following:
[0303] Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
[0304] According to one or more embodiments of this disclosure, the first transmitted data does not contain encapsulated data.
[0305] Fifthly, according to one or more embodiments of this disclosure, an audio data processing system is provided, including a first terminal and a second terminal; the first terminal and the second terminal are wirelessly connected via Bluetooth.
[0306] The first terminal includes a data acquisition device, a data processing device, and a first Bluetooth device. The data acquisition device acquires first audio data. The data processing device determines target compression parameters based on target type information corresponding to the first audio data, and compresses the first audio data based on the target compression parameters to obtain first compressed data. The target type information includes voice type or non-voice type. The first Bluetooth device sends the first compressed data to the second terminal.
[0307] The second terminal includes a second Bluetooth device and a second data processing device. The second Bluetooth device obtains first transmission data from the first terminal. The second data processing device decompresses the first compressed data in the first transmission data based on the compression parameter information indicating the target compression parameter in the first transmission data to obtain second audio data.
[0308] According to one or more embodiments of this disclosure, the first terminal includes a wearable device;
[0309] The second terminal includes mobile terminals.
[0310] In a sixth aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0311] The memory stores the instructions that the computer executes;
[0312] At least one processor executes computer execution instructions stored in memory, causing at least one processor to perform the methods described above as the first aspect, the second aspect, and various possible designs of the first and second aspects.
[0313] In a seventh aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the information display method described above as the first aspect, the second aspect, and various possible designs of the first and second aspects.
[0314] Eighthly, according to one or more embodiments of this disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, it implements the first aspect, the second aspect, and various possible designs of the first and second aspects as described above.
[0315] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0316] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0317] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An audio data processing method, applied to a first terminal, characterized in that, include: Obtain the first audio data; Target compression parameters are determined based on the target type information corresponding to the first audio data, and the first audio data is compressed based on the target compression parameters to obtain first compressed data; wherein, the target type information includes speech type or non-speech type; The first compressed data is sent to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
2. The method according to claim 1, characterized in that, The step of determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes: In response to the target type information being voice type, a first proportion of voice data is determined; Determine the first target compression parameter corresponding to the first proportion; The first audio data is compressed based on the first target compression parameters.
3. The method according to claim 1, characterized in that, The step of determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes: In response to the target type information being non-voice type, a second proportion of non-voice data is determined; Determine the second target compression parameter corresponding to the second proportion; The first audio data is compressed based on the second target compression parameters.
4. The method according to claim 1, characterized in that, The step of determining target compression parameters based on target type information corresponding to the first audio data, and compressing the first audio data based on the target compression parameters, includes: In response to the target type information being a speech type, the first audio data is parsed to determine the target semantics of the first audio data; Determine the third target compression parameter corresponding to the target semantics; The first speech data is compressed using the third target compression parameters.
5. The method according to claim 4, characterized in that, The step of parsing the first audio data to determine the target semantics of the first audio data includes: Identify preset keywords in the first audio data, and determine the target semantics based on the preset keywords; or... The first audio data is semantically understood, and the target semantics are determined based on the semantic understanding results.
6. The method according to claim 1, characterized in that, The method further includes: The first audio data is classified based on a preset audio classification algorithm; The target type information of the first audio data is determined based on the classification results.
7. The method according to claim 1, characterized in that, The speech type includes human voice, and the non-speech type includes ambient sound.
8. The method according to claim 1, characterized in that, The target compression parameters include the compression ratio, where, The compression ratio of speech-type audio data is lower than that of non-speech-type audio data.
9. The method according to any one of claims 1-8, characterized in that, Sending the first compressed data to the second terminal includes: First transmission data, including the first compressed data, is constructed according to a custom data transmission protocol and sent to the second terminal. The first transmission data further includes at least one of compression parameter information indicating the target compression parameters, voice / audio start time information, and voice / audio end time information.
10. The method according to any one of claims 1-8, characterized in that, The method further includes: transmitting one or more of the following to the second terminal based on a custom data transmission protocol: Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
11. The method according to any one of claims 1-8, characterized in that, The first transmitted data does not contain any encapsulated data.
12. An audio data processing method, applied to a second terminal, characterized in that, include: First transmission data is obtained from the first terminal, wherein the second terminal is wirelessly connected to the first terminal via Bluetooth; Based on the compression parameter information indicating the target compression parameters in the first transmitted data, the first compressed data in the first transmitted data is decompressed to obtain the second audio data, wherein the first compressed data is obtained by compressing the first audio data using the target compression parameters.
13. The method according to claim 12, characterized in that, The method further includes: The second audio data is transmitted to a target application on the second terminal, and the target application generates a response to the second audio data.
14. The method according to claim 13, characterized in that, The response is generated by the target application performing semantic understanding on the second audio data and based on the semantic understanding results; wherein, the response includes response audio and / or response text.
15. The method according to claim 14, characterized in that, The target application is connected to a network server, and the target application transmits the second audio data to the network server for semantic understanding.
16. The method according to claim 12, characterized in that, The target compression parameter is related to the target type information corresponding to the first audio data, and the first transmission data includes the first compressed data after the first audio data is compressed by the target compression parameter.
17. The method according to claim 16, characterized in that, The first audio data is collected by the first terminal; and / or, the target type information includes voice type or non-voice type.
18. The method according to claim 16, characterized in that, The first transmitted data also includes at least one of voice audio start time information and voice audio end time information.
19. The method according to claim 12, characterized in that, Second transmission data is obtained from the first terminal based on a custom data transmission protocol, wherein the second transmission data further includes one or more of the following: Sampling rate, frame length, number of channels, packet interval, supported packet length, supported compression ratio.
20. The method according to claim 12, characterized in that, The first transmitted data does not contain any encapsulated data.
21. An audio data processing apparatus, characterized in that, include: The first acquisition unit is used to acquire the first audio data; A compression unit is configured to determine target compression parameters based on target type information corresponding to the first audio data, and compress the first audio data based on the target compression parameters to obtain first compressed data; wherein, the target type information includes speech type or non-speech type; The transmitting unit is used to transmit the first compressed data to the second terminal, wherein the first terminal and the second terminal are wirelessly connected via Bluetooth.
22. An audio data processing device, disposed in a second terminal, characterized in that, include: The second acquisition unit is used to acquire first transmission data from the first terminal, wherein the second terminal and the first terminal are wirelessly connected via Bluetooth; The decompression unit is used to decompress the first compressed data in the first transmitted data based on the compression parameter information indicating the target compression parameters in the first transmitted data, to obtain the second audio data.
23. An audio data processing system, characterized in that, include: A first terminal and a second terminal; the first terminal and the second terminal are wirelessly connected via Bluetooth; The first terminal includes a data acquisition device, a data processing device, and a first Bluetooth device. The data acquisition device acquires first audio data; the data processing device determines target compression parameters based on target type information corresponding to the first audio data, and compresses the first audio data based on the target compression parameters to obtain first compressed data; wherein the target type information includes voice type or non-voice type; the first Bluetooth device sends the first compressed data to the second terminal. The second terminal includes a second Bluetooth device and a second data processing device. The second Bluetooth device obtains first transmission data from the first terminal. The second data processing device decompresses the first compressed data in the first transmission data based on the compression parameter information indicating the target compression parameter in the first transmission data to obtain second audio data.
24. The audio data processing system according to claim 23, characterized in that, The first terminal includes a wearable device; The second terminal includes a mobile terminal.
25. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 11, or the method as described in any one of claims 12 to 20.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 11, or the method as described in any one of claims 12 to 20.
27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 11, or the method as described in any one of claims 12 to 20.