Audio data processing method and device and electronic equipment

By mixing the received first and second audio data in the first electronic device and determining the adjustment parameters according to their respective audio types, the problem of playback effect differences caused by different preset algorithms is solved, and the effect of audio data processing is improved.

CN121306154APending Publication Date: 2026-01-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410917828.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In electronic devices, different preset algorithms are used, resulting in different playback effects when playing target audio data processed by different preset algorithms through headphones, and the audio data processing effect is poor.

Method used

The first electronic device wirelessly connects with the second electronic device, receives first audio data and second audio data from the second electronic device, determines the corresponding preset adjustment parameters according to their respective audio types, performs mixing processing on the audio data, and generates target audio data.

Benefits of technology

By mixing the two types of audio data using the first electronic device, the playback effect difference caused by the second electronic device using different preset algorithms is avoided, thus improving the audio data processing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306154A_ABST
    Figure CN121306154A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio data processing method and device and electronic equipment. The method is applied to first electronic equipment and comprises the following steps: wirelessly receiving first audio data and second audio data from second electronic equipment; wherein the first electronic equipment is wirelessly connected with the second electronic equipment through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type; performing sound mixing processing on the first audio data and the second audio data to obtain target audio data; and playing the target audio data. And the audio data processing effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer, and particularly, to an audio data processing method and device and electronic equipment. BACKGROUND

[0002] The electronic equipment can establish a Bluetooth connection with the earphone and play audio data sent by the electronic equipment to the earphone through the earphone. In some scenarios, when multiple audio of different sources are played,

[0003] In actual application, the electronic equipment usually mixes multiple audio of different sources through a preset algorithm to obtain target audio, and sends target audio data to the earphone. The earphone receives and plays the target audio data. In the above process, different preset algorithms are used in different types of electronic equipment, and the target audio data obtained by different preset algorithms is different. When the target audio data obtained by different preset algorithms is played through the earphone, the playing effect will be different. This leads to poor effect of audio data processing. SUMMARY

[0004] Embodiments of the present disclosure provide an audio data processing method and device and electronic equipment to solve the problem of poor effect of audio data processing.

[0005] In a first aspect, embodiments of the present disclosure provide an audio data processing method applied to a first electronic equipment, and the method comprises:

[0006] wirelessly receiving first audio data and second audio data from a second electronic equipment, wherein the first electronic equipment is wirelessly connected to the second electronic equipment through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type;

[0007] mixing the first audio data and the second audio data to obtain target audio data;

[0008] playing the target audio data.

[0009] In a second aspect, embodiments of the present disclosure provide an audio data processing device, and the device comprises:

[0010] a receiving module configured to wirelessly receive first audio data and second audio data from a second electronic equipment, wherein the second electronic equipment is wirelessly connected to the audio data processing device through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type;

[0011] a processing module configured to mix the first audio data and the second audio data to obtain target audio data;

[0012] The playing module is configured to play the target audio data.

[0013] In a third aspect, the present disclosure provides a chip, wherein the chip stores a computer program, and the computer program is executed by the chip to implement the method according to any one of the first aspect.

[0014] In a fourth aspect, the present disclosure provides a chip module, wherein the chip module stores a computer program, and the computer program is executed by the chip module to implement the method according to any one of the first aspect.

[0015] In a fifth aspect, the present disclosure provides an electronic device, comprising:

[0016] at least one processor; and

[0017] a memory connected with the at least one processor; wherein

[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of the first aspect.

[0019] In a sixth aspect, the present disclosure provides a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions are used to make the computer execute the method according to any one of the first aspect.

[0020] In a seventh aspect, the present disclosure provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the method according to any one of the first aspect.

[0021] The audio data processing method, device and electronic device provided by the present disclosure can receive first audio data and second audio data from a second electronic device wirelessly by a first electronic device. The first electronic device and the second electronic device are connected wirelessly through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type. The first audio data and the second audio data are mixed to obtain target audio data. The target audio data is played. In the above process, the first electronic device can mix the two types of audio data to obtain the target audio data. By processing and playing through the first electronic device, the effect of playing through the first electronic device can be different when the preset algorithm used by the second electronic device is different and the target audio data obtained by processing through different preset algorithms is different. The effect of audio data processing is improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0023] Figure 1 The schematic diagram of the application scenario provided by the embodiment of the present disclosure is shown in the following figure.

[0024] Figure 2 The flowchart of the audio data processing method provided by the embodiment of the present disclosure is shown in the following figure.

[0025] Figure 3 The flowchart of another audio data processing method provided by the embodiment of the present disclosure is shown in the following figure.

[0026] Figure 4 The process diagram of the audio data decompression processing provided by the embodiment of the present disclosure is shown in the following figure.

[0027] Figure 5 The process diagram of the audio data decoding processing provided by the embodiment of the present disclosure is shown in the following figure.

[0028] Figure 6 The process diagram of the audio mixing processing provided by the embodiment of the present disclosure is shown in the following figure.

[0029] Figure 7 The schematic diagram of the audio data processing process provided by the embodiment of the present disclosure is shown in the following figure.

[0030] Figure 8 The structural diagram of the audio data processing device provided by the embodiment of the present disclosure is shown in the following figure.

[0031] Figure 9 The structural diagram of the electronic device provided by the embodiment of the present disclosure is shown in the following figure. DETAILED DESCRIPTION

[0032] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0033] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards, and provide corresponding operation portal for user to choose authorization or refusal.

[0034] For ease of understanding, below, combined with Figure 1 The application scenarios to which the embodiments of the present disclosure are applied are described.

[0035] Figure 1 A schematic diagram of the application scenario provided by the embodiments of the present disclosure is shown in FIG. 1. Please refer to Figure 1 , including an electronic device 101 and an earphone 102. The electronic device 101 can be a mobile terminal. For example, the mobile terminal can be a mobile phone, a tablet computer, etc. The earphone 102 can be a Bluetooth earphone, a True Wireless Stereo (TWS) earphone, an open earphone, etc. After the electronic device 101 and the earphone 102 establish a connection through Bluetooth, the electronic device 101 sends audio data corresponding to a song 1 to the earphone 102 in response to a user's audio playing operation. After receiving the audio data sent by the electronic device 101, the earphone 102 can play the song 1 according to the audio data. When playing the song 1, if the electronic device 101 also runs a voice assistant application in the electronic device 101, the earphone 102 needs to play audio data corresponding to the song 1 and audio data corresponding to the voice assistant at the same time. At this time, the two types of audio data need to be mixed.

[0036] In actual application process, the multimedia audio data and the audio data of the voice assistant can be mixed in the electronic device through a preset algorithm to obtain target audio data. The target audio data is sent to the earphone. The earphone receives and plays the target audio data. In the above process, the preset algorithms used by different types of electronic devices are different, and the target audio data obtained by processing through different preset algorithms is different. When playing the target audio data obtained by processing through different preset algorithms through the earphone, the playing effect will be different. This leads to poor effect of audio data processing.

[0037] In the embodiment of the present disclosure, the first electronic device wirelessly receives first audio data and second audio data from the second electronic device. The first electronic device and the second electronic device are connected wirelessly through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type. The first audio data and the second audio data are mixed to obtain target audio data. The target audio data is played. In the above process, the first electronic device can mix the two types of audio data to obtain the target audio data. By processing and playing through the first electronic device, the difference in the effect of playing through the first electronic device can be avoided when the preset algorithm used by the second electronic device is different and the target audio data obtained by processing through different preset algorithms is different. The effect of audio data processing is improved.

[0038] In the following, the method shown in the present disclosure is described through specific embodiments. It should be noted that the following embodiments can exist independently or be combined with each other. For the same or similar content, it will not be repeated in different embodiments.

[0039] Figure 2 A flowchart of an audio data processing method provided by the embodiment of the present disclosure is shown. Please refer to Figure 2 The method can include:

[0040] S201, wirelessly receiving first audio data and second audio data from the second electronic device.

[0041] The execution subject of the embodiment of the present disclosure can be an electronic device, or a chip, a chip module or an audio data processing device provided in the electronic device. The audio data processing device can be realized by software, or by the combination of software and hardware.

[0042] The first electronic device and the second electronic device are connected wirelessly through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type.

[0043] The first electronic device includes a wearable device, and the second electronic device includes a mobile terminal. The wearable device includes at least one of an earphone, smart glasses, a smart watch, and a smart bracelet; and the second electronic device includes at least one of a mobile phone, a notebook computer, and a tablet computer.

[0044] The first audio data includes media audio data played by an audio and video player or call data for voice communication of the second electronic device; and the second audio data includes assistant audio data, which includes response audio data of the second electronic device to user voice data transmitted by the first electronic device. The assistant audio data includes human voice.

[0045] The first electronic device is provided with a voice assistant application. A user can wake up the voice assistant application in the first electronic device through a preset voice. The voice assistant is woken up and generates prompt information. After the user obtains the prompt information, the user inputs voice data to the voice assistant. The prompt information can be a ringtone or voice information.

[0046] After the first electronic device receives the voice data of the user, the voice data is sent to the second electronic device. The second electronic device determines the response audio data corresponding to the voice data according to the voice data, and sends the response audio data to the first electronic device. The first electronic device receives and plays the response audio data.

[0047] For example, the first electronic device is earphone A, and the second electronic device is mobile phone A. After the user establishes a Bluetooth connection between mobile phone A and earphone A, the user wakes up the voice assistant of earphone A through a preset voice A. The voice assistant of earphone A generates prompt information and plays the preset voice A prompt information through earphone A. After the user obtains the prompt information, the user inputs voice data "today's weather" to the voice assistant of earphone A. The voice assistant of earphone A obtains the voice data and sends the voice data to mobile phone A. Mobile phone A obtains the response audio data corresponding to the voice data from a cloud server, and sends the response audio data corresponding to the voice data to earphone A. After earphone A receives the response audio data, earphone A plays the response audio data "today's weather is XX".

[0048] The first audio data is transmitted to the first electronic device through at least one of a Hands-free Profile (HFP) protocol, a Headset Profile (HSP) protocol, and an Advanced Audio Distribution Profile (A2DP) protocol; and the second audio data is encoded in an Opus format at the second electronic device.

[0049] If the first audio data is call data of voice communication, the first audio data is transmitted between the first electronic device and the second electronic device through the HFP protocol or the HSP protocol. If the first audio data is media audio data played by a video player, the first audio data is transmitted between the first electronic device and the second electronic device through the A2DP protocol.

[0050] The user voice collected by the first electronic device is pulse code modulation (PCM) audio data. The first electronic device can perform encoding and compression processing on the collected user voice by Opus to obtain user voice data. The format of the user voice data sent by the first electronic device and received by the second electronic device is OGG format. The second electronic device performs decompression processing on the user voice data in OGG format to obtain PCM audio data. The second electronic device obtains the response voice corresponding to the user data according to the PCM audio data. The second electronic device performs encoding and compression processing on the response voice by Opus to obtain second audio data. The second electronic device sends the second audio data to the first electronic device.

[0051] For example, assuming that the voice to be compressed is voice A. The Opus encoding and compression is performed according to 20 ms 960 bytes per frame, the compression rate is 5 times, the length of the compressed data is 192, and every 3 frames form a data packet. Therefore, a compressed frame can be determined as 24000*2*20 / 1000 / 5=192 bytes. Every 3 frames form a data packet, so the compressed packet obtained by processing can be determined as 576 bytes. The compressed packet obtained by Opus encoding and compression processing can be specifically as shown in Table 1:

[0052] Table 1

[0053]

[0054] It should be noted that the encoding and compression manner of the second electronic device for the first audio data can be Opus or other manners, and the present disclosure does not make any limitation.

[0055] S202, performing mixing processing on the first audio data and the second audio data to obtain target audio data.

[0056] The first audio data and the second audio data can be mixed and processed to obtain target audio data in the following manner: the first audio data and the second audio data are processed by using adjusting parameters of different sizes respectively, and the adjusting parameters include preset volume gain and / or preset frequency gain.

[0057] The first audio data and the second audio data can be processed by using adjusting parameters of different sizes in the following manner: a first preset adjusting parameter corresponding to the first audio data is determined according to a first audio type corresponding to the first audio data; a second preset adjusting parameter corresponding to the second audio data is determined according to a second audio type corresponding to the second audio data, and the first preset adjusting parameter and the second preset adjusting parameter are different in size.

[0058] The audio types can be human voice type, multimedia type, call type, etc. The user can determine the adjustment parameters corresponding to each audio type according to the priority of the audio type. For example, the user expects to hear the human voice type audio more clearly, and when setting the adjustment parameters, the human voice type adjustment parameters (preset volume gain and / or preset frequency gain) can be set larger. The adjustment parameters (preset volume gain and / or preset frequency gain) of other audio types are set smaller. In this way, after the first audio data and the second audio data are processed by using the adjustment parameters of different sizes respectively, the volume of the human voice type audio in the target audio data is larger, and the volume of the non-human voice type audio is smaller.

[0059] The plurality of audio types and the preset adjustment parameters corresponding to each audio type can be set in advance and stored in the preset storage space of the first electronic device.

[0060] When mixing the first audio data and the second audio data, the volume and the frequency of the first audio data and the volume and the frequency of the second audio data can be adjusted according to the preset adjustment parameters set by the user. In this way, the volume of the first audio data and the second audio data in the target audio data obtained by mixing is different, and the sound effect is different. In this way, the user can hear the audio with higher priority more clearly.

[0061] For example, according to the example shown in the above, it is determined that the first electronic device is earphone A, and the second audio data is "today's weather is XX". Assuming that the first audio data is the audio data corresponding to music A. After the earphone A receives the first audio data and the second audio data, the first audio data and the second audio data are processed by using adjustment parameters of different sizes respectively. The volume of the first audio data is reduced according to the first preset adjustment parameter, and the volume of the second audio data is increased according to the second preset adjustment parameter, thereby obtaining target audio data A.

[0062] S203, playing the target audio data.

[0063] The volume of the first audio type and the volume of the second audio type are different when the target audio data is played.

[0064] After obtaining the target audio data, the first electronic device can directly play the target audio data through the loudspeaker of the earphone.

[0065] For example, according to the example described above, the target audio data A is determined. The target audio data A includes the audio of "today's weather is XX" and the audio of music A. The earphone A plays the target audio data A through the speaker, in which the audio volume of "today's weather is XX" is larger, and the audio volume of music A is smaller. In this way, the user can clearly obtain the response audio data corresponding to the user voice data.

[0066] The audio data processing method provided by the embodiments of the present disclosure receives the first audio data and the second audio data wirelessly from the second electronic device. The first audio data and the second audio data are mixed to obtain target audio data. The target audio data is played. In the above process, the first electronic device can be used to mix the two kinds of audio data to obtain the target audio data. By processing and playing through the first electronic device, the difference in the effect of playing through the first electronic device can be avoided when the preset algorithm used by the second electronic device is different, and the target audio data obtained by processing through different preset algorithms is different. The effect of audio data processing is improved.

[0067] Based on any one of the above embodiments, the following will be described in combination with Figure 3 The detailed process of processing the audio data will be described.

[0068] Figure 3 Another flowchart of an audio data processing method provided by the embodiments of the present disclosure is shown in FIG. 6. Please refer to Figure 3 The method includes the following steps.

[0069] S301, receiving the first audio data and the second audio data wirelessly from the second electronic device.

[0070] It should be noted that the execution steps of S301 can refer to S201, which will not be described here.

[0071] S302, determining the first preset adjustment parameter corresponding to the first audio data according to the first audio type corresponding to the first audio data.

[0072] For example, it is assumed that the first electronic device is earphone B, the first audio data received by the earphone B is the audio data of video B, and the second audio data is the response audio data of the voice assistant of the earphone B. The earphone B determines that the first audio type of the first audio data is a multimedia type. The first preset adjustment parameter corresponding to the first audio data in the preset storage space of the earphone B includes a preset volume gain of 0.5 and a preset frequency gain of 0.8.

[0073] S303, determining the second preset adjustment parameter corresponding to the second audio data according to the second audio type corresponding to the second audio data.

[0074] The first preset adjustment parameter and the second preset adjustment parameter are different in size.

[0075] For example, according to the example described above, it is determined that the second audio data is the response audio data of the voice assistant of the earphone B. The earphone B determines that the second audio type of the second audio data is a human voice type. And the second preset adjustment parameter corresponding to the second audio data in the preset storage space of the earphone B includes a preset volume gain 1.2.

[0076] S304, the first audio data and the second audio data are respectively decompressed and decoded to obtain the first intermediate audio data corresponding to the first audio data and the second intermediate audio data corresponding to the second audio data.

[0077] Next, combined with Figure 4 The process of decompressing the audio data is described. Figure 4 The process of decompressing the audio data provided by the embodiments of the present disclosure is shown in the figure. Please refer to Figure 4 , including audio data 401 and audio data 402. The audio data 401 can be the first audio data sent by the second electronic device to the first electronic device through the A2DP protocol. Or, the second electronic device sends the second audio data to the first electronic device through the serial port data transparent transmission (Serial Port Profile, SPP) protocol implemented by Bluetooth. The audio data 401 includes a plurality of data packets, each data packet is compressed by Opus. After the first electronic device receives the audio data 401, the audio data 401 is decompressed to obtain the audio data 402. Each sequence in the decompressed audio data 402 includes 960 bytes of audio data.

[0078] The SPP protocol is a private Bluetooth protocol. Users can customize the specific content of the SPP protocol according to the use scenario. In the embodiments of the present disclosure, the audio data corresponding to the voice assistant application in the first electronic device can be transmitted through the SPP protocol.

[0079] For example, the voice assistant of the first electronic device is provided with a wake-up algorithm, and user voice data is obtained through the wake-up algorithm. When it is determined that the voice of the user is a preset voice, the user voice data and at least one operation instruction can be sent to the second electronic device through the artificial intelligence timing control in the SPP protocol. After the second electronic device receives the user voice data and at least one operation instruction, the user voice data is sent to the cloud server. The cloud server obtains corresponding response audio data according to the user voice data. After the cloud server determines the response audio data, a preparation instruction is sent to the second electronic device. After the second electronic device receives the preparation instruction, the preparation instruction is sent to the first electronic device. In this way, the second electronic device and the first electronic device perform preparation operations (for example, player startup, memory allocation, etc.) to receive the response audio data. The cloud server sends the response audio data to the second electronic device through the SSP protocol, and the second electronic device sends the response audio data to the first electronic device through the SSP protocol.

[0080] Next, the process of decoding the audio data will be described in combination with Figure 5 The process of decoding the audio data will be described in combination with Figure 5 The process of decoding the audio data provided by the embodiments of the present disclosure is shown in the following figure. Please refer to Figure 5 , including decoding queue 501 and linked list 502. The decompressed audio data shown in the above Figure 4 is cached in the linked list 402. In response to the data preparation instruction of the decoding thread in the first electronic device, the decoding thread obtains the header field from the list once to perform opus decoding, and obtains the decoded audio data. The decoded audio data is cached in the decoding queue 501. The queue 501 can cache the audio data corresponding to 3 data packets after decompression processing. When the decompressed audio data is cached in the linked list 502, if 25 data packets are cached, the decoding speed cannot be accurately controlled. If the decoding speed is too fast, it will cause packet overflow and frame loss. If the decoding speed is too slow, it will cause lag and asynchronization. Therefore, the number of cached data packets can be set to 10.

[0081] S305, the first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, to obtain the first selected audio data corresponding to the first intermediate audio data and the second selected audio data corresponding to the second intermediate audio data, respectively.

[0082] For any one of the first intermediate audio data and the second intermediate audio data, if the preset adjustment parameter corresponding to the intermediate audio data includes a preset volume gain; the intermediate audio data can be processed by using the preset adjustment parameter in the following manner: determining the volume size corresponding to the intermediate audio data; processing the volume of the intermediate audio data with the preset volume gain to obtain the to-be-selected audio data corresponding to the intermediate audio data.

[0083] For any one of the first intermediate audio data and the second intermediate audio data, if the preset adjustment parameter corresponding to the intermediate audio data includes a preset frequency gain; the intermediate audio data can be processed by using the preset adjustment parameter in the following manner: determining the frequency size corresponding to the intermediate audio data; processing the frequency of the intermediate audio data with the preset frequency gain to obtain the to-be-selected audio data corresponding to the intermediate audio data.

[0084] For any one of the first intermediate audio data and the second intermediate audio data, if the preset adjustment parameter corresponding to the intermediate audio data includes a preset volume gain and a preset frequency gain; the intermediate audio data can be processed by using the preset adjustment parameter in the following manner: determining the volume size and the frequency size corresponding to the intermediate audio data; processing the volume of the intermediate audio data with the preset volume gain, and processing the frequency of the intermediate audio data with the preset frequency gain, to obtain the to-be-selected audio data corresponding to the intermediate audio data.

[0085] When the volume of the intermediate audio data is processed with the preset volume gain, the volume of the intermediate audio data and the preset volume gain can be added or multiplied, which is not limited by the present disclosure.

[0086] When the frequency of the intermediate audio data is processed with the preset frequency gain, the frequency of the intermediate audio data and the preset frequency gain can be added or multiplied, which is not limited by the present disclosure.

[0087] For example, according to the example shown in the above, it is determined that the first preset adjustment parameter includes a preset volume gain 0.5 and a preset frequency gain 0.8, and the second preset adjustment parameter includes a preset volume gain 1.2. The volume size and the frequency size of each intermediate audio data can be determined as shown in Table 2:

[0088] Table 2

[0089] Intermediate audio data Volume Frequency First intermediate audio data a1 f1 Second intermediate audio data a2 f2

[0090] For the first intermediate audio data, according to Table 2, the volume a1 of the first intermediate audio data is multiplied by the preset volume gain 0.5, and the frequency f1 of the first intermediate audio data is multiplied by the preset frequency gain 0.8, to obtain the first selected audio data A1. The volume of the first selected audio data A1 is 0.5*a1, and the frequency is 0.8*f1. For the second intermediate audio data, according to Table 2, the volume a2 of the second intermediate audio data is multiplied by the preset volume gain 1.2, to obtain the second selected audio data A2. The volume of the second selected audio data A2 is 1.2*a2.

[0091] In S306, the first selected audio data and the second selected audio data are mixed to obtain the target audio data.

[0092] The first selected audio data and the second selected audio data can be mixed to obtain the target audio data in the following manner: a preset sampling rate is obtained; the first selected audio data and the second selected audio data are respectively sampled according to the preset sampling rate, to obtain M pieces of sampling audio data corresponding to the first selected audio data and M pieces of sampling audio data corresponding to the second selected audio data, M being an integer greater than or equal to 1; and the M pieces of sampling audio data corresponding to the first selected audio data and the M pieces of sampling audio data corresponding to the second selected audio data are mixed to obtain the target audio data.

[0093] The preset sampling rate can be set in advance and stored in a preset storage space of the first electronic device. The preset sampling rate is greater than a first sampling rate at which the first audio data is generated and a second sampling rate at which the second audio data is generated. The preset sampling rate can be 384K. After processing by the preset sampling rate, the sampling rates of the sampling audio data corresponding to each selected audio data are the same. In this way, the timing of mixing the M pieces of sampling audio data is the same.

[0094] The M pieces of sampling audio data corresponding to the first selected audio data and the M pieces of sampling audio data corresponding to the second selected audio data can be mixed to obtain the target audio data in the following manner: M groups of audio data are determined, the i-th group of audio data including the i-th sampling audio data in the first selected audio data and the i-th sampling audio data in the second selected audio data; each sampling audio data in each group of audio data is mixed to obtain M mixed audio data; and it is determined that the target audio data includes the M mixed audio data.

[0095] For example, according to the preset sampling rate, the first candidate audio data is sampled to obtain 64 sample audio data corresponding to the first candidate audio data, including sample data A1 to sample data A64. According to the preset sampling rate, the second candidate audio data is sampled to obtain 64 sample audio data corresponding to the second candidate audio data, including sample data B1 to sample data B64. The 64 groups of audio data can be determined as shown in Table 3:

[0096] Table 3

[0097] Audio data First candidate audio Second candidate audio First group of audio data First sample audio data First sample audio data Second group of audio data Second sample audio data Second sample audio data …… …… …… Sixty-fourth group of audio data Sixty-fourth sample audio data Sixty-fourth sample audio data

[0098] Each sample audio data in each group of audio data shown in Table 3 is mixed to obtain 64 mixed audio data. The target audio data includes 64 mixed audio data.

[0099] The mixing process can be implemented in the chip in the first electronic device. The process of mixing is described below. Figure 6 The process of mixing is described below. Figure 6 The process of mixing provided by the embodiments of the present disclosure is shown in the figure. Please refer to Figure 6 , including chip 601. The chip 601 is arranged in the first electronic device, and the chip 601 includes a central processing unit (CPU) and a codec. The multimedia / call data stream of the CPU is used to process the first audio data, and the voice assistant data stream of the CPU is used to process the second audio data. The codec of the chip 601 obtains the first intermediate audio data and the second intermediate audio data through direct memory access (DMA). The codec of the chip 601 processes the first intermediate audio data through the digital-to-analog converter 1 and the first preset adjustment parameter to obtain the first candidate audio data. And through the digital-to-analog converter 2 and the second preset adjustment parameter, the second intermediate audio data is processed to obtain the second candidate audio data. The codec of the chip 601 mixes the first candidate audio data and the second candidate audio data to obtain the target audio data. And the target audio data is sent to the player of the first electronic device, so that the player plays the target audio data.

[0100] S307, playing the target audio data.

[0101] If the first electronic device is a TWS earphone, when playing the target audio data and determining that the master earphone and the slave earphone of the earphone need to play at the same time, the two earphones need to be synchronized. Avoiding the target audio data played by the two earphones out of sync, resulting in poor user experience.

[0102] The two earphones can be synchronously processed in an event synchronization manner. The earphones can be synchronously processed in the following manner: after the target audio data is generated, a target moment is determined according to a Bluetooth clock; a target data frame corresponding to the target moment is determined in the target audio data; and when the current moment is the target moment, the master earphone and the slave earphone of the earphone simultaneously play the target data frame.

[0103] After the earphones simultaneously play the target data frame, each data frame in the target audio data is continuously played in turn.

[0104] For example, when the earphones determine the target audio data, it is determined that the target moment is moment 2 and the current moment is moment 1. The earphones determine that the target data frame corresponding to moment 2 is data frame 1 in the target audio data. Then, when the current moment is moment 2, the master earphone and the slave earphone in the earphones synchronously play data frame 1.

[0105] The audio data processing method provided in the embodiments of the present disclosure receives first audio data and second audio data from a second electronic device wirelessly. According to a first audio type corresponding to the first audio data, a first preset adjustment parameter corresponding to the first audio data is determined. According to a second audio type corresponding to the second audio data, a second preset adjustment parameter corresponding to the second audio data is determined. The first audio data and the second audio data are respectively processed by decompression and decoding to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data. The first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, to respectively obtain first selected audio data corresponding to the first intermediate audio data and second selected audio data corresponding to the second intermediate audio data. The first selected audio data and the second selected audio data are mixed to obtain target audio data. The target audio data is played. In the above process, the two kinds of audio data can be mixed by the first electronic device to obtain the target audio data. By processing and playing through the first electronic device, when the preset algorithms used by the second electronic device are different, the target audio data obtained by processing through different preset algorithms is different, and the effect of playing through the first electronic device is different. The effect of audio data processing is improved.

[0106] On the basis of any one of the above embodiments, the following will be described in combination with Figure 7 The process of processing audio data is exemplified.

[0107] Figure 7 The schematic diagram of the process of processing audio data provided in the embodiments of the present disclosure is shown in FIG. 1. Please refer to FIG. 1. Figure 7, including a first electronic device 701 and a second electronic device 702. The first electronic device 701 can be a wearable device. For example, the first electronic device 701 can be a headset, smart glasses, a smart watch, a smart bracelet, etc. The first electronic device 701 is provided with a chip and an interface for data transmission. The first electronic device 701 is also provided with a voice assistant. The second electronic device 702 can be a mobile terminal. For example, the second electronic device 702 can be a mobile phone, a notebook computer, a tablet computer, etc.

[0108] After the Bluetooth connection is established between the first electronic device 701 and the second electronic device 702, the second electronic device 702 plays song 2 through the first electronic device 701 in response to a user input play instruction. At this time, the voice assistant of the first electronic device 701 sends the user voice data to the second electronic device 702 through the SPP protocol after obtaining the user voice data. The second electronic device 702 obtains the response audio data corresponding to the user voice data from the cloud server. The second electronic device 702 sends the response audio data to the first electronic device 701 through the SSP protocol, and sends the audio data corresponding to song 2 to the first electronic device 701 through the A2DP protocol.

[0109] After the first electronic device 701 receives the audio data corresponding to song 2 and the response audio data, it determines the audio type and the preset adjustment parameter corresponding to each kind of audio data, which can be as shown in Table 4:

[0110] Table 4

[0111] Audio data Audio type Pre-set adjustment parameter Audio data corresponding to song 2 Multimedia type Volume gain 1 Response audio data Vocal type Volume gain 2

[0112] For the audio data corresponding to song 2, the chip of the first electronic device 701 processes the audio data corresponding to song 2 using the volume gain 1 shown in Table 4 to obtain the candidate audio data 1 corresponding to the audio data corresponding to song 2. For the response audio data, the chip of the first electronic device 701 processes the response audio data using the volume gain 2 shown in Table 4 to obtain the candidate audio data 2 corresponding to the response audio data. The chip of the first electronic device 701 obtains a preset sampling rate. And according to the preset sampling rate, the candidate audio data 1 and the candidate audio data 2 are respectively sampled to obtain 128 sampling audio data corresponding to the candidate audio data 1 and 128 sampling audio data corresponding to the candidate audio data 1. The chip of the first electronic device 701 mixes the 128 sampling audio data corresponding to the candidate audio data 1 and the 128 sampling audio data corresponding to the candidate audio data 1 to obtain target audio data. The chip of the first electronic device 701 sends the target audio data to the loudspeaker of the first electronic device 701, and plays the target audio data through the loudspeaker.

[0113] The audio data processing process provided by the embodiments of the present disclosure receives first audio data and second audio data from a second electronic device wirelessly. According to a first audio type corresponding to the first audio data, a first preset adjustment parameter corresponding to the first audio data is determined. According to a second audio type corresponding to the second audio data, a second preset adjustment parameter corresponding to the second audio data is determined. The first audio data and the second audio data are respectively processed by decompression and decoding to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data. The first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, to obtain first selected audio data corresponding to the first intermediate audio data and second selected audio data corresponding to the second intermediate audio data. The first selected audio data and the second selected audio data are mixed to obtain target audio data. The target audio data is played. In the above process, the first electronic device can mix the two types of audio data to obtain the target audio data. By processing and playing through the first electronic device, the difference in the effect of playing through the first electronic device can be avoided when the preset algorithms used by the second electronic device are different and the target audio data obtained by processing through different preset algorithms is different. The effect of audio data processing is improved.

[0114] Figure 8 A structural schematic diagram of an audio data processing apparatus provided by the embodiments of the present disclosure is provided. Please refer to Figure 8 The audio data processing apparatus 800 includes:

[0115] The receiving module 801 is configured to receive first audio data and second audio data from a second electronic device wirelessly, the second electronic device is wirelessly connected to the audio data processing apparatus through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type.

[0116] The processing module 802 is configured to mix the first audio data and the second audio data to obtain target audio data.

[0117] The playing module 803 is configured to play the target audio data.

[0118] According to one or more embodiments of the present disclosure, the first electronic device includes a wearable device, and the second electronic device includes a mobile terminal.

[0119] According to one or more embodiments of the present disclosure, the wearable device includes at least one of a headset, smart glasses, a smart watch, and a smart bracelet, and the second electronic device includes at least one of a mobile phone, a notebook computer, and a tablet computer.

[0120] According to one or more embodiments of the present disclosure, the first audio data comprises media audio data played by an audio-video player or call data of voice communication performed by the second electronic device; and the second audio data comprises assistant audio data, which comprises response audio data of the second electronic device to user voice data transmitted by the first electronic device.

[0121] According to one or more embodiments of the present disclosure, the first audio data is transmitted to the first electronic device through at least one of a hands-free call / hands-free operation telephone (HFP) protocol, a telephone call / operation telephone (HSP) protocol, or a Bluetooth audio transmission model agreement (A2DP) protocol; and the second audio data is encoded by the second electronic device in an Opus format.

[0122] According to one or more embodiments of the present disclosure, the assistant audio data comprises human voice.

[0123] According to one or more embodiments of the present disclosure, when the target audio data is played, the volume of the first audio type is different from the volume of the second audio type.

[0124] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0125] The first audio data and the second audio data are processed by using different adjusting parameters, wherein the adjusting parameters comprise preset volume gain and / or preset frequency gain.

[0126] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0127] According to a first audio type corresponding to the first audio data, a first preset adjusting parameter corresponding to the first audio data is determined; and according to a second audio type corresponding to the second audio data, a second preset adjusting parameter corresponding to the second audio data is determined, wherein the first preset adjusting parameter and the second preset adjusting parameter are different in size.

[0128] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0129] The first audio data and the second audio data are respectively subjected to decompression processing and decoding processing to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data.

[0130] The first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, so as to obtain first candidate audio data corresponding to the first intermediate audio data and second candidate audio data corresponding to the second intermediate audio data, respectively.

[0131] The first candidate audio data and the second candidate audio data are mixed to obtain the target audio data.

[0132] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0133] determine a volume size corresponding to the intermediate audio data;

[0134] process the volume of the intermediate audio data by using the preset volume gain, so as to obtain candidate audio data corresponding to the intermediate audio data.

[0135] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0136] determine a frequency size corresponding to the intermediate audio data;

[0137] process the frequency of the intermediate audio data by using the preset frequency gain, so as to obtain candidate audio data corresponding to the intermediate audio data.

[0138] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0139] determine a volume size and a frequency size corresponding to the intermediate audio data;

[0140] process the volume of the intermediate audio data by using the preset volume gain, and process the frequency of the intermediate audio data by using the preset frequency gain, so as to obtain candidate audio data corresponding to the intermediate audio data.

[0141] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0142] obtain a preset sampling rate;

[0143] According to the preset sampling rate, the first candidate audio data and the second candidate audio data are respectively processed by sampling, so as to obtain M pieces of sampling audio data corresponding to the first candidate audio data and M pieces of sampling audio data corresponding to the second candidate audio data, where M is an integer greater than or equal to 1.

[0144] Mix the M pieces of sampling audio data corresponding to the first candidate audio data and the M pieces of sampling audio data corresponding to the second candidate audio data to obtain the target audio data.

[0145] According to one or more embodiments of the present disclosure, the processing module 803 is specifically configured to:

[0146] determine M groups of audio data, wherein the i-th group of audio data includes the i-th piece of sampling audio data in the first candidate audio data and the i-th piece of sampling audio data in the second candidate audio data;

[0147] mix each piece of sampling audio data in each group of audio data to obtain M pieces of mixed audio data;

[0148] determine that the target audio data includes the M pieces of mixed audio data.

[0149] The audio data processing apparatus provided by the embodiments of the present disclosure can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects. Therefore, the details are not described here.

[0150] Figure 9 A structural schematic diagram of an electronic device is provided for the embodiments of the present disclosure. Please refer to Figure 9 which shows a structural schematic diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 includes a wearable device. Among them, the electronic device can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers (PAD), portable multimedia players (PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 9 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0151] As Figure 9As shown, the electronic device 900 can include a processing device (e.g., a central processor, a graphics processor, etc.) 901 that can perform various suitable actions and processes in accordance with programs stored in a Read Only Memory (ROM) 902 or loaded from a storage device 908 into a Random Access Memory (RAM) 903. Various programs and data required by the electronic device 900 for operation are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other by a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0152] Generally, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate wirelessly or wired with other devices to exchange data. Although Figure 9 The electronic device 900 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0153] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 909, or installed from the storage devices 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0154] It should be noted that the computer readable medium in the above disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0155] The computer readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.

[0156] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0157] The embodiment of the present disclosure provides a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the method as variously possible in the above embodiments is implemented.

[0158] The embodiment of the present disclosure provides a computer program product, which includes a computer program, and when a processor executes the computer program, the method as variously possible in the above embodiments is implemented.

[0159] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0160] The flow diagrams and the block diagrams in the drawings are meant as possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0161] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.

[0162] The functions described in this specification can be implemented in part or in whole by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0164] It should be noted that the modification of "one" or "multiple" mentioned in this disclosure is illustrative rather than limiting, and those skilled in the art should understand that "one" or "multiple" should be understood unless otherwise explicitly indicated in the context.

[0165] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0166] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0167] For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information. As an optional but not limited implementation manner, in response to receiving the active request of the user, the way of sending the prompt information to the user may, for example, be the way of pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0168] It can be understood that the above notification and obtaining of user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0169] It can be understood that the data (including but not limited to data itself, acquisition or use of data) involved in the technical solution should comply with the requirements of relevant laws, regulations and provisions. The data can include information, parameters and messages, etc., such as stream splitting indication information.

[0170] In a first aspect, the embodiments of the present disclosure provide an audio data processing method applied to a first electronic device, and the method comprises:

[0171] wirelessly receiving first audio data and second audio data from a second electronic device; wherein the first electronic device and the second electronic device are connected wirelessly through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type;

[0172] mixing the first audio data and the second audio data to obtain target audio data;

[0173] playing the target audio data.

[0174] According to one or more embodiments of the present disclosure, the first electronic device comprises a wearable device, and the second electronic device comprises a mobile terminal.

[0175] According to one or more embodiments of the present disclosure, the wearable device comprises at least one of a headset, smart glasses, a smart watch, and a smart bracelet; and the second electronic device comprises at least one of a mobile phone, a notebook computer, and a tablet computer.

[0176] According to one or more embodiments of the present disclosure, the first audio data comprises media audio data played by an audio / video player or call data for voice communication of the second electronic device; and the second audio data comprises assistant audio data, which comprises response audio data of the second electronic device to user voice data transmitted by the first electronic device.

[0177] According to one or more embodiments of the present disclosure, the first audio data is transmitted to the first electronic device through at least one of a hands-free call / hands-free operation telephone (HFP) protocol, a telephone call / operation telephone (HSP) protocol, and a Bluetooth audio transmission model agreement (A2DP) protocol; and the second audio data is encoded by the second electronic device in an Opus format.

[0178] According to one or more embodiments of the present disclosure, the assistant audio data comprises human voice.

[0179] According to one or more embodiments of the present disclosure, when the target audio data is played, the volume of the first audio type is different from the volume of the second audio type.

[0180] According to one or more embodiments of the present disclosure, the mixing processing of the first audio data and the second audio data obtains target audio data, including:

[0181] The first audio data and the second audio data are processed by using different sizes of adjustment parameters, and the adjustment parameters include preset volume gain and / or preset frequency gain

[0182] According to one or more embodiments of the present disclosure, the processing of the first audio data and the second audio data by using different sizes of adjustment parameters includes:

[0183] According to the first audio type corresponding to the first audio data, a first preset adjustment parameter corresponding to the first audio data is determined; according to the second audio type corresponding to the second audio data, a second preset adjustment parameter corresponding to the second audio data is determined, and the first preset adjustment parameter and the second preset adjustment parameter are different in size.

[0184] According to one or more embodiments of the present disclosure, the processing of the first audio data and the second audio data by using different sizes of adjustment parameters includes:

[0185] The first audio data and the second audio data are respectively processed by decompression and decoding to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data;

[0186] The first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, to obtain first selected audio data corresponding to the first intermediate audio data and second selected audio data corresponding to the second intermediate audio data, respectively;

[0187] The first selected audio data and the second selected audio data are mixed to obtain the target audio data.

[0188] According to one or more embodiments of the present disclosure, for any one of the first intermediate audio data and the second intermediate audio data; the preset adjustment parameter corresponding to the intermediate audio data includes the preset volume gain; the processing of the intermediate audio data by using the preset adjustment parameter includes:

[0189] The volume size corresponding to the intermediate audio data is determined;

[0190] The volume of the intermediate audio data is processed by using the preset volume gain to obtain selected audio data corresponding to the intermediate audio data.

[0191] According to one or more embodiments of the present disclosure, for any one of the first intermediate audio data and the second intermediate audio data; the preset adjustment parameter corresponding to the intermediate audio data includes the preset frequency gain; the processing of the intermediate audio data by using the preset adjustment parameter includes:

[0192] determining a frequency size corresponding to the intermediate audio data;

[0193] processing the frequency of the intermediate audio data by using the preset frequency gain to obtain the candidate audio data corresponding to the intermediate audio data.

[0194] According to one or more embodiments of the present disclosure, for any one of the first intermediate audio data and the second intermediate audio data; the preset adjustment parameter corresponding to the intermediate audio data includes the preset volume gain and the preset frequency gain; the processing of the intermediate audio data by using the preset adjustment parameter includes:

[0195] determining a volume size and a frequency size corresponding to the intermediate audio data;

[0196] processing the volume of the intermediate audio data by using the preset volume gain, and processing the frequency of the intermediate audio data by using the preset frequency gain, to obtain the candidate audio data corresponding to the intermediate audio data.

[0197] According to one or more embodiments of the present disclosure, the mixing processing of the first candidate audio data and the second candidate audio data to obtain the target audio data includes:

[0198] obtaining a preset sampling rate;

[0199] According to the preset sampling rate, performing sampling processing on the first candidate audio data and the second candidate audio data respectively to obtain M sampling audio data corresponding to the first candidate audio data and M sampling audio data corresponding to the second candidate audio data, where M is an integer greater than or equal to 1;

[0200] performing mixing processing on the M sampling audio data corresponding to the first candidate audio data and the M sampling audio data corresponding to the second candidate audio data to obtain the target audio data.

[0201] According to one or more embodiments of the present disclosure, the mixing processing of the M sampling audio data corresponding to the first candidate audio data and the M sampling audio data corresponding to the second candidate audio data to obtain the target audio data includes:

[0202] determining M groups of audio data, each i-th group of audio data including an i-th sample audio data in the first candidate audio data and an i-th sample audio data in the second candidate audio data;

[0203] mixing each sample audio data in each group of audio data respectively to obtain M mixed audio data;

[0204] determining that the target audio data includes the M mixed audio data.

[0205] In a second aspect, an audio data processing apparatus is provided, and the apparatus includes:

[0206] a receiving module configured to receive first audio data and second audio data from a second electronic device wirelessly, the second electronic device being wirelessly connected to the audio data processing apparatus through Bluetooth, the first audio data corresponding to a first audio type, and the second audio data corresponding to a second audio type;

[0207] a processing module configured to perform audio mixing on the first audio data and the second audio data to obtain target audio data;

[0208] a playing module configured to play the target audio data.

[0209] According to one or more embodiments of the present disclosure, the first electronic device includes a wearable device, and the second electronic device includes a mobile terminal.

[0210] According to one or more embodiments of the present disclosure, the wearable device includes at least one of a headset, smart glasses, a smart watch, and a smart bracelet; and the second electronic device includes at least one of a mobile phone, a notebook computer, and a tablet computer.

[0211] According to one or more embodiments of the present disclosure, the first audio data includes media audio data played by an audio / video player or call data obtained by voice communication of the second electronic device; and the second audio data includes assistant audio data, the assistant audio data including response audio data of the second electronic device to user voice data transmitted by the first electronic device.

[0212] According to one or more embodiments of the present disclosure, the first audio data is transmitted to the first electronic device through at least one of a hands-free telephony / Hands-Free Profile (HFP) protocol, a telephone call / Hands-Free Profile (HSP) protocol, and a Bluetooth Audio Transmission Model Agreement (A2DP) protocol; and the second audio data is encoded by the second electronic device in an Opus format.

[0213] According to one or more embodiments of the present disclosure, the assistant audio data includes human voice.

[0214] According to one or more embodiments of the present disclosure, the target audio data is played, and the volume of the first audio type is different from the volume of the second audio type.

[0215] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0216] The first audio data and the second audio data are processed by using different sizes of adjustment parameters, and the adjustment parameters include preset volume gain and / or preset frequency gain.

[0217] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0218] According to the first audio type corresponding to the first audio data, a first preset adjustment parameter corresponding to the first audio data is determined; and according to the second audio type corresponding to the second audio data, a second preset adjustment parameter corresponding to the second audio data is determined, and the first preset adjustment parameter and the second preset adjustment parameter are different in size.

[0219] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0220] The first audio data and the second audio data are respectively processed by decompression and decoding to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data.

[0221] The first intermediate audio data is processed by using the first preset adjustment parameter, and the second intermediate audio data is processed by using the second preset adjustment parameter, to obtain first selected audio data corresponding to the first intermediate audio data and second selected audio data corresponding to the second intermediate audio data.

[0222] The first selected audio data and the second selected audio data are mixed to obtain the target audio data.

[0223] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0224] The volume of the intermediate audio data is determined.

[0225] The volume of the intermediate audio data is processed by using the preset volume gain to obtain selected audio data corresponding to the intermediate audio data.

[0226] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0227] The frequency of the intermediate audio data is determined.

[0228] process the volume of the intermediate audio data by using the preset volume gain, and process the frequency of the intermediate audio data by using the preset frequency gain, to obtain the to-be-selected audio data corresponding to the intermediate audio data.

[0229] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0230] determine the volume size and the frequency size corresponding to the intermediate audio data;

[0231] process the volume of the intermediate audio data by using the preset volume gain, and process the frequency of the intermediate audio data by using the preset frequency gain, to obtain the to-be-selected audio data corresponding to the intermediate audio data.

[0232] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0233] obtain a preset sampling rate;

[0234] According to the preset sampling rate, perform sampling processing on the first to-be-selected audio data and the second to-be-selected audio data respectively, to obtain M groups of sampling audio data corresponding to the first to-be-selected audio data and M groups of sampling audio data corresponding to the second to-be-selected audio data, where M is an integer greater than or equal to 1;

[0235] perform mixing processing on the M groups of sampling audio data corresponding to the first to-be-selected audio data and the M groups of sampling audio data corresponding to the second to-be-selected audio data, to obtain the target audio data.

[0236] According to one or more embodiments of the present disclosure, the processing module is specifically configured to:

[0237] determine M groups of audio data, and the i th group of audio data includes the i th sampling audio data in the first to-be-selected audio data and the i th sampling audio data in the second to-be-selected audio data;

[0238] perform mixing processing on each sampling audio data in each group of audio data respectively, to obtain M mixed audio data;

[0239] determine that the target audio data includes the M mixed audio data.

[0240] In a third aspect, the present disclosure provides a chip, and the chip stores a computer program. When the computer program is executed by the chip, the method according to any one of the first aspect is implemented.

[0241] In a fourth aspect, the present disclosure provides a chip module, and the chip module stores a computer program. When the computer program is executed by the chip module, the method according to any one of the first aspect is implemented.

[0242] In a fifth aspect, an electronic device is provided, comprising:

[0243] at least one processor; and

[0244] a memory communicatively connected with the at least one processor; wherein

[0245] the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform the method of any one of the first aspect.

[0246] In a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the method of any one of the first aspect.

[0247] In a seventh aspect, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of any one of the first aspect.

[0248] The above description merely provides preferred embodiments of the present disclosure and a description of the technical principles of the application. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

[0249] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0250] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. An audio data processing method, characterized by, Applied to a first electronic device, the method comprises: wirelessly receiving first audio data and second audio data from a second electronic device; wherein the first electronic device and the second electronic device are connected wirelessly through Bluetooth, the first audio data corresponds to a first audio type, and the second audio data corresponds to a second audio type; mixing the first audio data and the second audio data to obtain target audio data; playing the target audio data.

2. The audio data processing method of claim 1, wherein, The first electronic device comprises a wearable device, and the second electronic device comprises a mobile terminal.

3. The audio data processing method of claim 2, wherein, The wearable device comprises at least one of a headset, smart glasses, a smart watch, and a smart bracelet; and the second electronic device comprises at least one of a mobile phone, a notebook computer, and a tablet computer.

4. The audio data processing method of claim 1, wherein, The first audio data comprises media audio data played by an audio-video player or call data for voice communication by the second electronic device; and the second audio data comprises assistant audio data, which comprises response audio data of the second electronic device to user voice data transmitted by the first electronic device.

5. The audio data processing method of claim 4, wherein, The first audio data is transmitted to the first electronic device through at least one of a hands-free call / hands-free operation telephone (HFP) protocol, a telephone call / operation telephone (HSP) protocol, and a Bluetooth audio transmission model agreement (A2DP) protocol; and the second audio data is encoded by the second electronic device in an Opus format.

6. The audio data processing method of claim 4, wherein, The assistant audio data comprises human voice.

7. The audio data processing method of claim 1, wherein, The volume of the first audio type is different from the volume of the second audio type when the target audio data is played.

8. The method of claim 1, wherein, Mixing the first audio data and the second audio data to obtain target audio data comprises: processing the first audio data and the second audio data respectively using different sizes of adjustment parameters, wherein the adjustment parameters comprise preset volume gain and / or preset frequency gain.

9. The method of claim 8, wherein, Processing the first audio data and the second audio data respectively using different sizes of adjustment parameters comprises: determining first preset adjustment parameters corresponding to the first audio data according to the first audio type corresponding to the first audio data, and determining second preset adjustment parameters corresponding to the second audio data according to the second audio type corresponding to the second audio data, wherein the first preset adjustment parameters and the second preset adjustment parameters are different in size.

10. The method of claim 9, wherein, Processing the first audio data and the second audio data respectively using different sizes of adjustment parameters comprises: respectively decompressing and decoding the first audio data and the second audio data to obtain first intermediate audio data corresponding to the first audio data and second intermediate audio data corresponding to the second audio data; processing the first intermediate audio data using the first preset adjustment parameters and processing the second intermediate audio data using the second preset adjustment parameters to respectively obtain first selected audio data corresponding to the first intermediate audio data and second selected audio data corresponding to the second intermediate audio data; and mixing the first selected audio data and the second selected audio data to obtain the target audio data. The first candidate audio data and the second candidate audio data are mixed to obtain the target audio data.

11. The method of claim 10, wherein, For any one of the first intermediate audio data and the second intermediate audio data, the preset adjustment parameter corresponding to the intermediate audio data includes the preset volume gain, and the processing of the intermediate audio data by using the preset adjustment parameter includes: determining a volume size corresponding to the intermediate audio data; processing the volume of the intermediate audio data by using the preset volume gain to obtain candidate audio data corresponding to the intermediate audio data.

12. The method of claim 10, wherein, For any one of the first intermediate audio data and the second intermediate audio data, the preset adjustment parameter corresponding to the intermediate audio data includes the preset frequency gain, and the processing of the intermediate audio data by using the preset adjustment parameter includes: determining a frequency size corresponding to the intermediate audio data; processing the frequency of the intermediate audio data by using the preset frequency gain to obtain candidate audio data corresponding to the intermediate audio data.

13. The method of claim 10, wherein, For any one of the first intermediate audio data and the second intermediate audio data, the preset adjustment parameter corresponding to the intermediate audio data includes the preset volume gain and the preset frequency gain, and the processing of the intermediate audio data by using the preset adjustment parameter includes: determining a volume size and a frequency size corresponding to the intermediate audio data; processing the volume of the intermediate audio data by using the preset volume gain and processing the frequency of the intermediate audio data by using the preset frequency gain to obtain candidate audio data corresponding to the intermediate audio data.

14. The method according to any one of claims 10 to 13, characterized in that, The processing of the first candidate audio data and the second candidate audio data to obtain the target audio data includes: obtaining a preset sampling rate; sampling the first candidate audio data and the second candidate audio data according to the preset sampling rate to obtain M groups of sampling audio data corresponding to the first candidate audio data and M groups of sampling audio data corresponding to the second candidate audio data, where M is an integer greater than or equal to 1; mixing the M groups of sampling audio data corresponding to the first candidate audio data and the M groups of sampling audio data corresponding to the second candidate audio data to obtain the target audio data.

15. The method of claim 14, wherein, The processing of the M groups of sampling audio data corresponding to the first candidate audio data and the M groups of sampling audio data corresponding to the second candidate audio data to obtain the target audio data includes: determining M groups of audio data, each of which includes an i-th sampling audio data in the first candidate audio data and an i-th sampling audio data in the second candidate audio data; mixing each sampling audio data in each group of audio data to obtain M mixed audio data; determining that the target audio data includes the M mixed audio data.

16. An audio data processing apparatus, characterized by comprising: The apparatus includes: receive a first audio data and a second audio data from a second electronic device wirelessly, the second electronic device being wirelessly connected to the audio data processing apparatus through Bluetooth, the first audio data corresponding to a first audio type, the second audio data corresponding to a second audio type; process the first audio data and the second audio data to obtain target audio data; play the target audio data.

17. An electronic device, comprising: comprise: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 15.

18. The electronic device of claim 17, wherein, The electronic device comprises a wearable device.

19. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, wherein the computer instructions are used to enable the computer to perform the method of any one of claims 1 to 15.

20. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 15.