Audio playing method, device, electronic device and storage medium
By combining stereo playback with multiple devices, the master device divides audio data and sends audio packets. The master device and slave device adjust the playback speed according to the target broadcast time, solving the problem of inconsistent audio playback between devices, achieving synchronous and continuous audio playback, and improving the user experience.
Patent Information
- Application Number
- CN202110839744.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-07-23
AI Technical Summary
In multi-device combination stereo playback, due to device hardware differences and network jitter, the speed of audio playback between devices is inconsistent, resulting in poor user experience.
The audio data stream to be played through the master device is divided into multiple frames of audio data, and the generated audio packet is sent to the slave device. The master device and the slave device adjust the playback speed of the audio data according to their respective target playback time to achieve synchronous playback of the audio data.
It effectively avoids the impact of hardware device differences or network jitter on playback speed, ensures the synchronization and continuity of combined audio playback of multiple devices, and improves stereo playback effect and user experience.
Smart Images

Figure CN113535115B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio playback, and in particular, to an audio playback method, apparatus, electronic device, and storage medium. Background Art
[0002] Combined stereo means that multiple smart devices are connected through a network to achieve the function of combined playback of multiple devices, enabling users to experience the effect of stereo playback. With the improvement and diversification of the functions of smart speakers, the combined stereo function has been applied. However, due to differences in device hardware, network jitter, etc., the playback speeds of devices are inconsistent, resulting in a poor user experience. Summary of the Invention
[0003] To overcome the problems in the related art, the present disclosure provides an audio playback method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an audio playback method, which is applied to a master device in a device combination, and the multi-device combination realizes audio synchronization for stereo playback. The method includes:
[0005] Obtain an audio data stream to be played, and divide the audio data stream into multiple frames of audio data;
[0006] Generate corresponding audio packets according to each frame of audio data, and send the generated multiple frames of audio packets to slave devices in the multi-device combination, and notify the slave devices to perform real-time audio playback;
[0007] In response to real-time audio playback, before playing the current frame of audio data, determine a first target playback time for the next frame of audio data, and adjust the playback speed of the next frame of audio data according to the first target playback time.
[0008] According to a second aspect of an embodiment of the present disclosure, there is provided an audio playback method, which is applied to a slave device in a device combination, and the multi-device combination realizes audio synchronization for stereo playback. The method includes:
[0009] Receive multiple frames of audio packets sent by a master device in the multi-device combination;
[0010] Parse each frame of audio packet to obtain the audio data in each frame of audio packet;
[0011] In response to the notification sent by the master device in the multi-devices to perform real-time audio playback, before playing the current frame of audio data, determine a second target playback time for the next frame of audio data, and adjust the playback speed of the next frame of audio data according to the second target playback time.
[0012] According to a third aspect of the embodiments of the present disclosure, there is provided an audio playback device, which is applied to a master device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The device includes:
[0013] A division processing module, configured to obtain an audio data stream to be played and divide the audio data stream into multiple frames of audio data;
[0014] A sending module, configured to generate corresponding audio packets according to each frame of audio data and send the generated multiple frames of audio packets to each slave device;
[0015] An obtaining module, configured to obtain the next frame of audio data before playing the current frame of audio data;
[0016] A determining module, configured to determine a first target playback time of the next frame of audio data;
[0017] An adjustment module, configured to adjust the playback speed of the next frame of audio data according to the first target playback time.
[0018] According to a fourth aspect of the embodiments of the present disclosure, there is provided an audio playback device, which is applied to a slave device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The device includes:
[0019] A receiving module, configured to receive multiple frames of audio packets sent by the master device in the multi-device combination;
[0020] An analysis module, configured to analyze each frame of audio packet to obtain the audio data in each frame of audio packet;
[0021] An obtaining module, configured to obtain the next frame of audio data before playing the current frame of audio data;
[0022] A determining module, configured to determine a second target playback time of the next frame of audio data;
[0023] An adjustment module, configured to adjust the playback speed of the next frame of audio data according to the second target playback time.
[0024] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio playback method described in the first aspect and / or the second aspect is implemented.
[0025] According to a sixth aspect of the embodiments of the present disclosure, a temporary computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the audio playing method described in the above first aspect and / or second aspect is implemented.
[0026] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The master device divides the audio data stream to be played into multiple frames of audio data, and sends the generated audio packets to each slave device, so that the synchronous playing of audio data can be realized through the transmission of audio data between the master device and the slave devices. In addition, both the master device and the slave devices correspondingly adjust the playing speed of their next frame of audio data according to their respective target playing times for the same next frame of audio data. In this way, the influence on the playing speed caused by differences in their respective hardware devices or network jitter can be avoided, so that the synchronization and continuity of multi-device combined audio playing can be ensured, and the stereo playing effect can be improved.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0029] Figure 1 is a flowchart of an audio playing method according to an embodiment of the present disclosure;
[0030] Figure 2 is an interaction diagram between a master device and slave devices according to an embodiment of the present disclosure;
[0031] Figure 3 is a flowchart of the master device adjusting the playing speed of the next frame of audio data according to an embodiment of the present disclosure;
[0032] Figure 4 is a schematic diagram of the principle of adjusting the playing speed for the next frame of audio data according to an embodiment of the present disclosure;
[0033] Figure 5 is a schematic diagram of the data expansion process according to an embodiment of the present disclosure;
[0034] Figure 6 is a schematic diagram of the data compression process according to an embodiment of the present disclosure;
[0035] Figure 7 is a flowchart of another method for the master device to adjust the playing speed of the next frame of audio data according to an embodiment of the present disclosure;
[0036] Figure 8It is a flowchart of another master device adjusting the playback speed of the next frame of audio data according to an embodiment of the present disclosure;
[0037] Figure 9 It is a flowchart of another audio playback method according to an embodiment of the present disclosure;
[0038] Figure 10 It is a flowchart of a slave device adjusting the playback speed of the next frame of audio data according to an embodiment of the present disclosure;
[0039] Figure 11 It is a flowchart of another slave device adjusting the playback speed of the next frame of audio data according to an embodiment of the present disclosure;
[0040] Figure 12 It is a flowchart of yet another slave device adjusting the playback speed of the next frame of audio data according to an embodiment of the present disclosure;
[0041] Figure 13 It is a block diagram of a structure of an audio playback device according to an embodiment of the present disclosure;
[0042] Figure 14 It is a block diagram of a structure of another audio playback device according to an embodiment of the present disclosure;
[0043] Figure 15 It is a block diagram of a structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0044] Here, exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0045] It should be noted that for the existing intelligent speaker stereo playback, a device combination for stereo playback is formed through WiFi interconnection, and information interaction is achieved through network transmission. The device combination for stereo playback includes a main speaker and one or more slave speakers. The main speaker obtains audio from the Internet and distributes it to other slave speakers in the group for playback while playing the audio. Ideally, after the initial synchronization, all devices only need to sequentially play their own audio. However, due to differences in device hardware, WiFi network jitter, and other situations, the playback progress of the devices may be different. Therefore, in order to ensure the synchronous playback of audio, each device needs to adjust the playback of the audio according to its own situation.
[0046] The existing adjustment solutions for the above problems are as follows: for devices with a fast playback progress, duplicate data is artificially inserted to delay the playback of the original data; for devices with a slow playback progress, some data is artificially discarded to enable the subsequent data to be played out as soon as possible. Based on this adjustment logic, the real-time synchronous playback of audio on each device is ensured. However, this solution simply adds or discards data at the connection of audio frames, which not only causes audio discontinuity but also introduces noise into the audio, thus reducing the user experience friendliness.
[0047] Based on the above problems, the present disclosure proposes an audio playback method, device, electronic device, and storage medium for audio synchronization in realizing stereo playback with a multi-device combination. Among them, the multi-device combination includes a master device and one or more slave devices, and the master device and the slave devices can be connected through transmission connection methods such as WiFi, 4G network, 5G network, and other networks (such as Bluetooth). In addition, the master device and the slave devices in the multi-device combination can be configured through the transmission between the devices. As an example, based on the configuration of the master-slave devices of a smart speaker, the user can select the stereo combination function on the configuration page of the master device and select the corresponding slave device from other smart speakers that can be recognized as slave devices according to the network connection; after the slave device receives the device combination request from the master device, it gives feedback to the master device. If the feedback information from the slave device is consent, it forms a device combination with the master device, that is, the configuration of the master device and the slave devices is completed. Next, for the convenience of understanding the solution, the audio playback method applied to the master device in the multi-device combination is first introduced.
[0048] Figure 1 It is a flowchart of an audio playback method proposed in an embodiment of the present disclosure. It should be noted that the audio playback method in the embodiment of the present disclosure is applied to the master device in the multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. In addition, the audio playback method in the embodiment of the present disclosure can be used in the audio playback device in the embodiment of the present disclosure, and the device can be configured in an electronic device. As Figure 1 shown, the audio playback method includes the following steps:
[0049] Step 101, obtain the audio data stream to be played and divide the audio data stream into multiple frames of audio data.
[0050] It can be understood that when using a multi-device combination to achieve the stereo playback effect, the master device is used to receive the audio data stream and send the corresponding audio data to the slave devices, so as to achieve synchronous playback between the master device and the slave devices.
[0051] In order to reduce the amount of audio data sent by the master device each time, and at the same time for subsequent monitoring of the playback progress and adjustment of the playback speed, in the embodiments of the present disclosure, the master device divides the acquired audio data stream to be played into multiple frames of audio data. That is to say, the master device divides the audio data stream to be played into continuous multiple segments of data, where each segment of data represents one frame of audio data. It should be noted that after the master device acquires the audio data stream to be played, it first needs to determine whether the current stereo playback function is in the enabled state, that is, first determine whether the master device uses a multi-device combination method for stereo playback. In response to the stereo playback function being in the enabled state, the audio data stream is divided into multiple frames of audio data.
[0052] As an example, when the stereo playback function of the master device is enabled, the master device obtains the audio data stream to be played through network transmission, such as mp3 audio data; after the master device obtains the audio data stream, it divides the audio data into frames based on a fixed length, such as dividing the audio into one frame every 20 ms, to obtain multiple frames of audio data.
[0053] Step 102, generate corresponding audio packets according to each frame of audio data, and send the generated multiple audio packets to the slave devices in the multi-device combination, and notify the slave devices to perform real-time audio playback.
[0054] It can be understood that in order to reduce the time delay during audio transmission, each frame of audio data can be sent separately, so that audio playback and audio transmission can be carried out simultaneously, avoiding playback waiting caused by too large a one-time transmission data volume.
[0055] In the embodiments of the present disclosure, in order to perform single-frame transmission on the multiple frames of audio data obtained after division, corresponding audio packets are generated according to each frame of audio data. The audio packet can include the time for dividing each audio frame, for example: the broadcast time of each frame of audio, the time length corresponding to each frame of audio, or the sequence number of each audio, etc.
[0056] As an example, the master device and the slave device negotiate a duration for audio packet sending and parsing. When the master device divides the audio stream data into frames, it calculates the broadcast time of each frame. The broadcast time of the first frame is the reception time plus the negotiated sending and parsing duration, and the broadcast time of the second frame is the broadcast time of the first frame plus the frame length, and so on. In this way, the master device generates corresponding audio packets according to each frame of audio data and the calculated broadcast time of each frame, so that the slave device can perform synchronous broadcast according to the corresponding broadcast time in the audio packet.
[0057] As another example, after the master device calculates the broadcast time of the first frame, it generates an audio packet for the first frame based on the audio data and the audio broadcast time of the first frame; the audio packets of other frames only contain the sequential numbers corresponding to the audio of each frame. For example, if the length of each frame is 20 ms, the first audio packet contains the broadcast time, the second audio packet contains its corresponding number "2", the third audio packet contains its corresponding number "3", and so on. In this way, the slave device can determine the broadcast time of each frame based on the corresponding number and the length of each frame, thereby achieving synchronous broadcast of audio.
[0058] To reduce the network load during the transmission of audio data packets, each audio data packet can be encoded and compressed to reduce the data stream. That is to say, each audio packet sent to the slave device can contain its corresponding division time and the encoded audio data. Here, the slave device refers to the slave device corresponding to the master device in the multi-device combination. The process of the master device sending audio packets to the slave device can be as follows: in response to the master device enabling the stereo function, obtain the corresponding list of slave devices in the configuration; send the generated multi-frame audio packets to each slave device in the list.
[0059] As Figure 2 shown, since there may be differences in the system clocks between each device, in order to ensure the synchronous playback of audio among the devices in the group, while the master device sends audio packets to the slave device, it will also interact with the slave device for clock information. This interaction of clock information can be understood as the time synchronization operation between the master device and the slave device to determine the clock difference between the master device and the slave device, and this interaction process can be achieved through WiFi, 4G network, 5G network, and other network transmission connection methods. To reduce the clock difference between the master device and the slave device, the interaction of clock information can be carried out in real time, for example, an interaction is performed at fixed intervals (1 s). Each time the clock interaction is carried out, the slave device calculates the difference τ between its system clock and the master device's clock according to the clock interaction information with the master device. 0 。
[0060] As an example, the master device and the slave device perform clock information interaction every 1 s, and each time the interaction is carried out, the slave device will obtain a system clock difference τ. 0 。For T 0 -moment audio, assuming that both the master device and the slave device broadcast after a fixed time (1 s), this fixed time refers to a delay time negotiated during the configuration of the master device and the slave device; for the master device, it only needs to broadcast the audio at T 0 + 1 s according to its own system clock, while for the slave speaker, it needs to consider the clock difference and play the audio at T 0 + τ 0 + 1 s.
[0061] Step 103: In response to real-time audio playback, before playing the current frame of audio data, determine the first target playback time of the next frame of audio data, and adjust the playback speed of the next frame of audio data according to the first target playback time.
[0062] It can be understood that, ideally, the master device and the slave device in the multi-device combination only need to play the audio in sequence according to multiple audio packets. However, due to differences in device hardware, network jitter, or software scheduling, etc., there may be differences in the audio playback progress of each device in the combination. That is to say, some devices in the combination may play audio faster, while some may play audio slower. Therefore, before playing the current frame of audio data, determine whether the audio data can be played on time according to the target playback time of the next frame of audio data. If it cannot be played on time, adjust the playback speed of the next frame of audio data to ensure the synchronous playback of the audio data of each device.
[0063] It should be noted that the master device and the slave device will pre-negotiate a playback response time. While the master device sends audio packets to the slave device, it also plays the audio data in sequence according to the playback time of the audio.
[0064] In the embodiments of the present disclosure, the first target playback time refers to the playback time corresponding to each frame of audio data when the master device divides the audio stream data into frames. In addition, to ensure the synchronization between the master device and the slave device, the master device and the slave device will pre-negotiate a playback response time. Optionally, the implementation manner for the master device to determine the first target playback time of the next frame of audio data may be: determine the playback response time pre-negotiated between the master device and the slave device; according to the start time of the next frame of audio data and the playback response time, determine the first target playback time of the next frame of audio data. For example, the first target playback time of the next frame of audio data may be the sum of the start time of the next frame of data and the playback response time.
[0065] It can be understood that the master device can predict the playback time of the next frame of audio data according to the current network conditions and hardware conditions. By comparing the predicted playback time of the next frame of audio data with the first target playback time, it can be determined whether the next frame of audio data can be played on time. If the predicted playback time is earlier than the first target playback time, it means that the current audio playback progress is fast, and the audio playback progress needs to be slowed down so that the next frame of audio data can be played on time. If the predicted playback time is later than the first target playback time, it means that the current audio playback progress is slow, and the audio playback progress needs to be accelerated so that the next frame of audio data can be played on time.
[0066] As an example, the playback speed of the next-frame audio data can be achieved by adjusting the amount of audio data at the connection between the current frame and the next frame. When the audio playback progress is slow, the amount of corresponding audio data is reduced; when the audio playback progress is fast, the amount of corresponding audio data is increased to adjust the audio playback progress. In fact, when adjusting the amount of data, it is necessary to ensure the continuity of the audio data as much as possible and also avoid introducing noise.
[0067] As another example, the playback speed of the next-frame audio data can be achieved by adjusting the amount of audio data between the current frame and the next frame. For example, the cubic spline interpolation method can be used to reduce or increase the amount of audio data between the current frame and the next frame to make the audio playback smoother.
[0068] According to the audio playback method of the embodiments of the present disclosure, the master device divides the audio data stream to be played into multiple frames of audio data and sends the generated audio packets to each slave device, so that the synchronous playback of the audio data can be realized through the transmission of the audio data between the master device and the slave devices. In addition, the master device adjusts the playback speed of the next-frame audio data according to the target playback time of the next-frame audio data. In this way, the influence on the playback speed caused by differences in hardware devices or network jitter can be avoided, so that the synchronization and continuity of the combined audio playback of multiple devices can be ensured, and the stereo playback effect can be improved.
[0069] Based on the above embodiments, the implementation method for adjusting the playback speed of the next-frame audio data will be further described below. Figure 3 This is a flowchart for adjusting the playback speed of the next-frame audio data in the embodiments of the present disclosure. As Figure 3 shown, the implementation process includes the following steps:
[0070] Step 310, determine the current time of the master device.
[0071] In the embodiments of the present disclosure, the current time of the master device refers to the time corresponding to the current system clock of the master device.
[0072] Step 320, predict the time required to write the next-frame audio data into the speaker in the master device and play it according to the hardware performance of the master device.
[0073] That is, according to the hardware performance of the master device, predict how long it takes from the current time to the playback of the next-frame audio data. It is necessary to determine the time required for the playback of the audio data of the current frame to be completed and the time required for the next-frame audio data to be written into the speaker in the master device and played according to the hardware performance of the master device.
[0074] Step 330, determine the predicted playback time of the next-frame audio data according to the current time and the required time.
[0075] It can be understood that the time from the current time to the time required for the predicted next-frame audio data to be broadcast, when added to the current time, can determine the predicted broadcast time of the next-frame audio data.
[0076] Step 340, if the predicted broadcast time is inconsistent with the first target broadcast time, then adjust the playback speed of the next-frame audio data.
[0077] It can be understood that the inconsistency between the predicted broadcast time and the first target broadcast time indicates that the next-frame audio cannot be broadcast at the first broadcast time, which will cause the problem of out-of-sync between the playback of the master device's audio data and the slave device. To avoid the occurrence of this problem, it is necessary to timely adjust the playback speed of the next-frame audio data.
[0078] Next, the adjustment methods for the playback speed of the next-frame audio data will be introduced respectively for the two cases where the predicted broadcast time is less than and greater than the first target broadcast time. Figure 4 This is the schematic diagram of the adjustment of the playback speed of the next-frame audio data in the embodiments of the present disclosure. As Figure 4 shown, if the predicted broadcast time is equal to the first target broadcast time, it means that the next-frame audio data can be played normally, and at this time, there is no need to adjust the playback speed, and the current-frame audio data can continue to be played. If the predicted broadcast time is less than the first target broadcast time, it means that the current audio playback progress is fast, and the audio playback speed needs to be slowed down. If the predicted broadcast time is greater than the first target broadcast time, it means that the current audio playback progress is slow, and the audio playback speed needs to be increased. Among them, the adjustment of the audio playback speed can be achieved by adjusting the data volume of the audio frame through cubic spline interpolation, or by adjusting the number of audio sampling points through linear predictive coding, or by adjusting the hardware sampling rate to adjust the playback speed, or other methods not mentioned in this application that can adjust the audio playback speed can be used for adjustment, and this application does not make any limitations in this regard. In the embodiments of this application, cubic spline interpolation will be used as an example to illustrate the adjustment of the audio playback speed. As Figure 3 shown, in response to the predicted broadcast time being less than the first target broadcast time, the speed adjustment process includes the following steps:
[0079] Step 341, if the predicted broadcast time is less than the first target broadcast time, then merge the current-frame audio data and the next-frame audio data.
[0080] Step 342, determine the first adjustment point number according to the difference between the predicted broadcast time and the first target broadcast time.
[0081] It can be understood that the predicted broadcast time being less than the first target broadcast time indicates that the current audio playback progress is fast, and data expansion is required to reduce the audio playback progress.
[0082] In the embodiments of the present disclosure, the first adjustment point number refers to the amount of audio data to be expanded, and its calculation method can be: subtracting the predicted broadcast time from the first target broadcast time to obtain a time difference t; obtaining the audio playback speed v (the amount of audio data played per unit time) of the current master device according to the current device hardware conditions and network conditions; multiplying the time difference t by the audio playback speed v of the current master device to obtain the first adjustment point number M = t×v.
[0083] Step 343: Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the first adjustment point number to obtain the audio data after data expansion.
[0084] It can be understood that expanding the audio data is to reduce the progress of audio playback so that the master device and the slave device can play the audio synchronously. Therefore, the expanded audio data needs to be continuous and smooth to ensure that the stereo effect is not affected.
[0085] In the embodiments of the present disclosure, to ensure the continuity of the audio data, the audio data of the current frame is merged with the audio data of the next frame, and interpolation calculation is performed on the sampling points corresponding to the merged data based on the cubic spline interpolation method. That is, interpolation is performed on the merged audio data of the current frame and the next frame to increase the number of sampling points by the first adjustment point number. In this way, compared with the method of only adding audio data at the connection between the current frame and the next frame, the continuity of the interpolated audio can be improved, so that the audio data can be played smoothly and the stereo playback effect can be improved.
[0086] The specific data expansion method can be as Figure 5 shown. Suppose the audio frame lengths of the current frame and the next frame are N, that is, the merged audio data has 2N sampling point data. Through the cubic spline interpolation method, the 2N-point data is changed to 2N + M-point data.
[0087] Step 344: Select the audio data corresponding to the first N interpolation nodes from the audio data after data expansion as the new current frame audio data, and use the remaining audio data in the audio data after data expansion as the new next frame audio data; where the value of N is the same as the number of sampling nodes of the current frame audio data.
[0088] In the embodiments of the present disclosure, when playing after data augmentation, the amount of data of the current frame audio data is still the number of original sampling nodes, and only the data of the sampling points is changed to the audio data corresponding to the first N difference nodes in the augmented audio data. The audio data corresponding to the remaining N+M difference nodes in the augmented audio data is used as the new next frame audio data. After the new current frame audio data is played, the new next frame audio data is played. In this way, the playing duration of the new next frame audio data will increase, so that the playing progress of the master device can be adjusted to be consistent with that of the slave device when the new next frame audio data is played.
[0089] As Figure 3 shown, if the predicted broadcast time is greater than the first target broadcast time, the speed adjustment process includes the following steps:
[0090] Step 345, if the predicted broadcast time is greater than the first target broadcast time, then merge the current frame audio data and the next frame audio data.
[0091] Step 346, determine the second adjustment point number according to the difference between the predicted broadcast time and the first target broadcast time.
[0092] It can be understood that the predicted broadcast time being greater than the first target broadcast time indicates that the current audio playing progress is fast, and data compression is required to accelerate the audio playing progress.
[0093] In the embodiments of the present disclosure, the second adjustment point number refers to the amount of audio data that needs to be compressed, and its calculation method can be: subtract the first target broadcast time from the predicted broadcast time to obtain a time difference t'; according to the current device hardware situation and network situation, obtain the audio playing speed v' of the current master device (the amount of audio data played per unit time); multiply the time difference t' by the audio playing speed v' of the current master device to obtain the second adjustment point number M = t'×v'.
[0094] Step 347, based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the second adjustment point number to obtain the audio data after data compression.
[0095] It can be understood that compressing the audio data is to accelerate the audio playing progress so that the master device and the slave device can play the audio synchronously. Therefore, the compressed audio data needs to be continuous and smooth to ensure that the stereo effect is not affected.
[0096] In the embodiments of the present disclosure, in order to ensure the continuity of audio data, the audio data of the current frame is merged with the audio data of the next frame, and interpolation calculation is performed on the adopted points corresponding to the merged data based on the cubic spline interpolation method. That is to say, interpolation is performed on the merged audio data of the current frame and the next frame, so that the number of sampling nodes in the merged audio data is reduced by the number of the second adjustment points. In this way, compared with the method of only discarding audio data at the connection between the current frame and the next frame, the continuity of the interpolated audio can be improved, so that the audio data can be played smoothly, and the stereo playback effect can be improved.
[0097] The specific data compression method can be as Figure 6 shown. Suppose the audio frame lengths of the current frame and the next frame are N, that is, the merged audio data has 2N sampling point data, and the number of the second adjustment points is M'. Through the cubic spline interpolation method, the 2N-point data is changed into 2N - M' point data.
[0098] Step 348: Select the audio data corresponding to the first N interpolation nodes from the audio data after data compression as the new current frame audio data, and use the remaining audio data in the audio data after data compression as the new next frame audio data; where the value of N is the same as the number of sampling nodes of the current frame audio data.
[0099] In the embodiments of the present disclosure, when playing after data compression, the data volume of the current frame audio data is still the number of the original sampling nodes, but the data of the sampling points is changed to the audio data corresponding to the first N interpolation nodes in the compressed audio data. Use the audio data corresponding to the remaining N - M' interpolation nodes in the compressed audio data as the new next frame audio data. After the new current frame audio data is played, play the new next frame audio data. In this way, the playing duration of the new next frame audio data will be reduced, so that the playing progress of the master device can be adjusted to be consistent with that of the slave device when the new next frame audio data is played.
[0100] According to the audio playing method proposed in the embodiments of the present disclosure, for the master device in a multi-device combination, by predicting the playing time of the next frame audio data before playing the current frame, and for the situation where the predicted playing time is inconsistent with the target playing time, adjusting the playing speed of the next frame audio data, it is possible to avoid the situation that the audio playing progress of the master device is slow or fast due to factors such as hardware device differences, and further ensure the synchronization of audio playing in the master device and each slave device. In addition, for the situation where the predicted playing time is inconsistent with the target playing time, the cubic spline interpolation method is introduced to expand or compress the data after merging the current frame and the next frame audio to adjust the playing speed, which improves the continuity of the audio data, so that the adjusted audio can still be played smoothly, and further ensures the stereo effect of multi-device combination audio playing and improves the user experience.
[0101] Since there are various ways to adjust the audio playback speed, the linear prediction coding method will be used to adjust the audio playback speed next.
[0102] Figure 7 This is a flowchart of another method for adjusting the playback speed of the next frame of audio data in the embodiments of the present disclosure. As Figure 7 shown, based on the above embodiments, the method further includes:
[0103] Step 741, if the predicted playback time is less than the first target playback time, determine the first adjustment point number according to the difference between the predicted playback time and the first target playback time.
[0104] It can be understood that the predicted playback time being less than the first target playback time indicates that the current audio playback progress is too fast, and data expansion is needed to reduce the audio playback progress.
[0105] In the embodiments of the present disclosure, the first adjustment point number refers to the amount of audio data that needs to be expanded, and its calculation method can be: subtract the predicted playback time from the first target playback time to obtain the time difference t; obtain the audio playback speed v (the amount of audio data played per unit time) of the current master device according to the current device hardware situation and network situation; multiply the time difference t by the audio playback speed v of the current master device to obtain the first adjustment point number M = t × v.
[0106] Step 742, based on the linear prediction coding method, predict the sampling point data corresponding to the first adjustment point number according to the current frame of audio data, and use the predicted sampling point data and the current frame of audio data as the new current frame of audio data for playback to adjust the playback speed of the next frame of audio data.
[0107] It can be understood that by also using the predicted sampling point data as the current frame of audio data, the data volume of the current frame of audio data can be expanded, thereby slowing down the audio playback speed and further adjusting the playback speed of the next frame of audio data.
[0108] It should be noted that the basic idea of linear prediction coding is that the current value of an audio sampling point can be approximated by a weighted linear combination of the past values of several audio sampling points.
[0109] As an example, based on the linear prediction coding method, the implementation method for predicting the sampling point data corresponding to the first adjustment point number according to the current frame of audio data can be as shown in formula (1):
[0110]
[0111] where x(n) is the sampling point at time n, T kis the prediction coefficient, P is the prediction order, usually p = 10 - 15, and x(n - k) is the sampling point at time n - k. Among them, the prediction coefficient T k is used to represent the weighting coefficient in the linear combination and can be obtained through the classical Levinson - Durbin iterative algorithm.
[0112] In the embodiments of the present disclosure, if the predicted number of sampling points is M and the number of sampling points of the current frame of audio is N, then the data corresponding to these M + N sampling points are all used as the new current frame of audio data for playback, thereby increasing the amount of data of the current frame of audio data, and then slowing down the audio playback speed so that the audio data can be played synchronously.
[0113] Step 743, if the predicted playback time is greater than the first target playback time, then reduce the number of sampling points in the audio frame according to the difference between the predicted playback time and the first target playback time to adjust the playback speed of the next frame of audio data.
[0114] As an example, it can be preset to remove one sampling point from each frame of audio; calculate the second adjustment point number M according to the predicted playback time and the first target playback time, that is, the number of sampling points M that need to be reduced, and evenly distribute the M sampling points that need to be reduced to the current frame and the M - 1 frames of audio after the current frame; randomly select one sampling point in each frame of audio for removal, and use the audio data after removal as the new audio data for playback. If there are N sampling points in the original audio frame, then there are N - 1 sampling points in each frame of the processed audio data. In this way, the audio playback speed is gradually adjusted by reducing the sampling points frame by frame.
[0115] The audio playback method proposed in the embodiments of the present disclosure is for the master device in the multi - device combination. For the situation where the predicted playback time is inconsistent with the target playback time, a linear prediction coding method is introduced to expand the data of the current frame to adjust the playback speed. Among them, the data of the predicted sampling points can ensure the continuity of the audio data, so that the adjusted audio can still be played smoothly, and further ensure the stereo effect of the multi - device combination audio playback, improving the user experience.
[0116] In addition, for the adjustment of the playback speed of the next frame of audio data, it can also be carried out by controlling the hardware settings. The present disclosure proposes another embodiment for this method. Figure 8 is a flowchart of another method for adjusting the playback speed of the next frame of audio data in the embodiments of the present disclosure. As Figure 8 shown, on the basis of the above - mentioned embodiments, the method further includes:
[0117] Step 841, if the predicted playback time is less than the first target playback time, then reduce the current playback sampling rate of the hardware driver in the master device to adjust the playback speed of the next frame of audio data.
[0118] Step 842, if the predicted broadcast time is greater than the first target broadcast time, increase the current playback sampling rate of the hardware driver in the master device to adjust the playback speed of the next frame of audio data.
[0119] For the adjustment of the audio playback speed, it can be achieved through the playback sampling rate of the hardware driver in the master device. As an example, if the predicted broadcast time is less than the first target broadcast time, that is, the current audio is playing too fast. Assuming that the current sampling rate of the master device speaker driver is 48 kHz, then increasing the playback sampling rate of the master device speaker driver can reduce the playback speed of the master device audio. For example, adjust the sampling rate of the speaker driver to 44.1 kHz; thereby adjusting the playback speed of the next frame of audio data to enable the master device to play audio synchronously with other slave devices. On the contrary, if the current audio playback speed is slow, the playback sampling rate of the master device hardware driver can be adjusted from 44.1 kHz to 48 kHz to adjust the playback speed of the next frame of audio data.
[0120] The audio playback method proposed by the embodiments of the present disclosure, for the master device in a multi-device combination, based on the adjustment of the playback speed of the next frame of audio when the predicted broadcast time is inconsistent with the target broadcast time, proposes to adjust the playback speed of the next frame of audio data by adjusting the playback sampling rate of the master device hardware driver, which is equivalent to proposing another way to adjust the audio data playback speed. It can not only meet the stereo effect of multi-device combination audio playback, but also improve the applicability of the method in practical applications.
[0121] Next, an audio playback method applied to the slave device in a multi-device combination will be introduced. Figure 9 It is a flowchart of another audio playback method proposed by the embodiments of the present disclosure. It should be noted that the audio playback method in the embodiments of the present disclosure is applied to the slave device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. As Figure 7 shown, the audio playback method includes the following steps:
[0122] Step 901, receive multiple frames of audio packets sent by the master device in the multi-device combination.
[0123] Step 902, parse each frame of audio packet to obtain the audio data in each frame of audio packet.
[0124] It can be understood that the slave device will immediately parse the audio packet sent by the master device. Since each frame of audio packet includes its corresponding audio data and division time, the slave device can obtain the audio data and its corresponding division time in each frame of audio packet after parsing each frame of audio packet. The slave device plays the audio data according to the obtained division time.
[0125] In addition, since the master device also interacts with the slave device for clock information, the slave device needs to consider the difference from the system clock of the master device when playing audio according to the divided time. For example, the master device and the slave device interact for clock information every 1 s, and each time the slave device obtains a system clock difference τ 0 . For the audio at time T 0 , assuming that both the master device and the slave device play the audio after a fixed time (1 s), where the fixed time is a delay time negotiated during the configuration of the master device and the slave device; for the master device, it only needs to play the audio at T 0 + 1 s according to its own system clock, while for the slave device, it needs to consider the clock difference and play the audio at T 0 + τ 0 + 1 s.
[0126] Step 903: In response to the notification sent by the master device among multiple devices to play audio in real time, before playing the current frame of audio data, determine the second target playback time of the next frame of audio data, and adjust the playback speed of the next frame of audio data according to the second target playback time.
[0127] It can be understood that ideally, the master device and the slave device in the multi-device combination only need to play the audio in sequence according to multiple audio packets. However, due to differences in device hardware, network jitter, or software scheduling, etc., there may be differences in the audio playback progress of each device in the combination. That is to say, some devices in the combination may play the audio faster, while some devices may play the audio slower. That is, according to the target playback time of the next frame of audio data, it is determined whether the next frame of audio data can be played on time. If it cannot be played on time, the playback speed of the next frame of audio data needs to be adjusted to ensure that the audio data of each slave device is played synchronously with the master device.
[0128] It should be noted that the master device and the slave device will pre-negotiate a playback response time. While the slave device receives the audio packet sent by the master device, it also plays the audio data in sequence according to the playback time of the audio.
[0129] In the embodiments of the present disclosure, the second target broadcast time refers to the broadcast time corresponding to each frame of audio data in the audio packets received by the device. In addition, to ensure the synchronization between the master device and the slave device, a broadcast response time is negotiated in advance between the master device and the slave device. Optionally, the implementation manner for the slave device to determine the second target broadcast time of the next frame of audio data may be: interacting with the master device for clock information, calculating the clock difference from the system clock information of the master device; determining the broadcast response time negotiated in advance between the master device and the slave device; and determining the second target broadcast time of the next frame of audio data according to the start time of the next frame of audio data, the broadcast response time, and the system clock difference. For example, the second target broadcast time of the next frame of audio data may be the sum of the start time of the next frame of data, the broadcast response time, and the system clock difference.
[0130] It can be understood that the slave device can predict the broadcast time of the next frame of audio data according to the current network conditions and hardware conditions, and by comparing the predicted broadcast time of the next frame of audio data with the first target broadcast time, it can be determined whether the next frame of audio data can be broadcast on time. If the predicted broadcast time is earlier than the first target broadcast time, it means that the current audio playback progress is fast, and the audio playback progress needs to be slowed down so that the next frame of audio data can be broadcast on time. If the predicted broadcast time is later than the first target broadcast time, it means that the current audio playback progress is slow, and the audio playback progress needs to be accelerated so that the next frame of audio data can be broadcast on time.
[0131] As an example, the playback speed of the next frame of audio data can be achieved by adjusting the data volume of the audio data at the connection between the current frame and the next frame. When the audio playback progress is slow, the data volume corresponding to the audio data is reduced, and when the audio playback progress is fast, the data volume corresponding to the audio data is increased to adjust the audio playback progress. In fact, the adjustment of the data volume needs to ensure the continuity of the audio data as much as possible and also avoid introducing noise as much as possible.
[0132] As another example, the playback speed of the next frame of audio data can be achieved by adjusting the data volume of the current frame and the next frame of audio data. For example, the cubic spline interpolation method is used to reduce or increase the data volume of the current frame and the next frame of audio data to make the audio playback smoother.
[0133] According to the audio playback method of the present disclosure embodiment, by receiving and parsing multi-frame audio data sent by the master device from the slave device, the audio data corresponding to each frame of audio packet is obtained, and the audio data is played according to the broadcast time, so that synchronous playback of audio data between the master device and the slave device can be realized. In addition, the slave device adjusts the playback speed of the next frame of audio data according to the target broadcast time of the next frame of audio data. In this way, the influence on the playback speed due to differences in hardware devices or network jitter can be avoided, so that the synchronization and continuity of multi-device combined audio playback can be ensured, and the stereo playback effect can be improved.
[0134] Based on the above embodiments, the implementation manner of adjusting the playback speed of the next frame of audio data will be further described below. Figure 10 It is a flowchart of the slave device adjusting the playback speed of the next frame of audio in the present disclosure embodiment. As Figure 10 shown, the implementation process includes the following steps:
[0135] Step 1010, determine the current time of the slave device.
[0136] In the present disclosure embodiment, the current time of the slave device refers to the time corresponding to the current system clock of the slave device.
[0137] Step 1020, predict the time required to write the next frame of audio data into the speaker of the slave device and broadcast it according to the hardware performance of the slave device.
[0138] That is to say, according to the hardware performance of the slave device, predict how long it takes from the current time to the broadcast of the next frame of audio data. Among them, according to the hardware performance of the slave device, the time required for the current frame of audio data to be played and the next frame of audio data to be written into the speaker of the slave device and broadcast is determined.
[0139] Step 1030, determine the predicted broadcast time of the next frame of audio data according to the current time and the required time.
[0140] It can be understood that the time required from the current time to the predicted broadcast of the next frame of audio data, added to the current time, can determine the predicted broadcast time of the next frame of audio data.
[0141] Step 1040, if the predicted broadcast time is inconsistent with the second target broadcast time, adjust the playback speed of the next frame of audio data.
[0142] It can be understood that the inconsistency between the predicted broadcast time and the second target broadcast time indicates that the next frame of audio cannot be broadcast at the second target broadcast time, which will cause the problem of out-of-sync between the playback of the audio data of this slave device and other devices. To avoid the occurrence of this problem, it is necessary to timely adjust the playback speed of the next frame of audio data.
[0143] Next, the adjustment method of the playing speed of the next-frame audio data will be introduced for two cases where the predicted playing time is less than and greater than the second target playing time respectively. As Figure 4 shown, if the predicted playing time is equal to the second target playing time, it means that the next-frame audio data can be played normally. At this time, there is no need to adjust the playing speed, and the current-frame audio data can be continued to be played. If the predicted playing time is less than the second target playing time, it means that the current audio playing progress is too fast, and the audio playing speed needs to be slowed down. If the predicted playing time is greater than the second target playing time, it means that the current audio playing progress is too slow, and the audio playing speed needs to be increased. Among them, the adjustment of the audio playing speed can be achieved by adjusting the data volume of audio frames through cubic spline interpolation, or by adjusting the number of audio sampling points through linear predictive coding, or by adjusting the hardware adoption rate to adjust the playing speed, or by using other methods not mentioned in this application to adjust the audio playing speed, and this application does not make any limitations in this regard. In the embodiments of this application, cubic spline interpolation will be used as an example to adjust the audio playing speed for illustration. As Figure 8 shown, in response to the predicted playing time being less than the second target playing time, the speed adjustment process of the slave device includes the following steps:
[0144] Step 1041, if the predicted playing time is less than the second target playing time, merge the current-frame audio data and the next-frame audio data.
[0145] Step 1042, determine the third adjustment point number according to the difference between the predicted playing time and the second target playing time.
[0146] It can be understood that the predicted playing time being less than the second target playing time indicates that the current audio playing progress is too fast, and data expansion is required to reduce the audio playing progress.
[0147] In the embodiments of the present disclosure, the third adjustment point number refers to the amount of audio data that needs to be expanded, and its calculation method can be: subtract the predicted playing time from the second target playing time to obtain the time difference t; according to the current device hardware situation and network situation, obtain the audio playing speed v (the amount of audio data played per unit time) of the current slave device; multiply the time difference t by the audio playing speed v of the current slave device to obtain the third adjustment point number M = t × v.
[0148] Step 1043, based on cubic spline interpolation, interpolate the sampling points corresponding to the merged audio data according to the third adjustment point number to obtain the audio data after data expansion.
[0149] It can be understood that expanding the audio data is to slow down the progress of audio playback so that the master device and the slave device can play the audio synchronously. Therefore, the expanded audio data needs to be continuous and smooth to ensure that the stereo effect is not affected.
[0150] In the embodiment of the present disclosure, to ensure the continuity of the audio data, the audio data of the current frame is merged with the audio data of the next frame, and interpolation calculation is performed on the adopted points corresponding to the merged data based on the cubic spline interpolation method. That is to say, interpolation is performed in the merged audio data of the current frame and the next frame, and the third adjustment number of sampling points is added. In this way, compared with the method of only adding audio data at the connection between the current frame and the next frame, the continuity of the interpolated audio can be improved, so that the audio data can be played smoothly and the stereo playback effect can be improved.
[0151] The specific data expansion method can be as Figure 5 shown. Suppose the audio frame lengths of the current frame and the next frame are N, that is, the merged audio data has 2N sampling point data. Through the cubic spline interpolation method, the 2N-point data is changed to 2N + M-point data.
[0152] Step 1044, select the audio data corresponding to the first N interpolation nodes from the expanded audio data as the new current frame audio data, and use the remaining audio data in the expanded audio data as the new next frame audio data; where the value of N is the same as the number of sampling nodes of the current frame audio data.
[0153] In the embodiment of the present disclosure, when playing after data expansion, the data volume of the current frame audio data is still the original number of sampling nodes, but the data of the sampling points is changed to the audio data corresponding to the first N interpolation nodes in the expanded audio data. Use the audio data corresponding to the remaining N + M interpolation nodes in the expanded audio data as the new next frame audio data. After the new current frame audio data is played, play the new next frame audio data. In this way, the playback duration of the new next frame audio data will increase, so that the playback progress of the slave device can be adjusted to be the same as that of other devices when the new next frame audio data is played.
[0154] As Figure 10 shown, if the predicted playback time is greater than the second target playback time, the speed adjustment process of the slave device includes the following steps:
[0155] Step 1045, if the predicted playback time is greater than the second target playback time, then merge the current frame audio data and the next frame audio data.
[0156] Step 1046, determine the fourth adjustment number according to the difference between the predicted playback time and the second target playback time.
[0157] It can be understood that if the predicted broadcast time is greater than the second target broadcast time, it means that the progress of the current audio playback is fast, and data compression is required to speed up the audio playback progress.
[0158] In the embodiment of the present disclosure, the fourth adjustment point number refers to the amount of audio data to be compressed, and its calculation method can be: taking the difference between the predicted broadcast time and the second target broadcast time to obtain a time difference t'; according to the current device hardware situation and network situation, obtaining the current audio playback speed v' of the slave device (the amount of audio data played per unit time); multiplying the time difference t' by the current audio playback speed v' of the slave device to obtain the fourth adjustment point number M = t'×v'.
[0159] Step 1047: Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the merged audio data according to the fourth adjustment point number to obtain the audio data after data compression.
[0160] It can be understood that compressing the audio data is to speed up the audio playback progress so that the slave device and other devices can play the audio synchronously. Therefore, the compressed audio data needs to be continuous and smooth to ensure that the stereo effect is not affected.
[0161] In the embodiment of the present disclosure, to ensure the continuity of the audio data, the audio data of the current frame is merged with the audio data of the next frame, and interpolation calculation is performed on the sampling points corresponding to the merged data based on the cubic spline interpolation method. That is, interpolation is performed on the merged audio data of the current frame and the next frame to reduce the number of sampling nodes by the fourth adjustment point number in the merged audio data. In this way, compared with the method of only discarding the audio data at the connection between the current frame and the next frame, the continuity of the interpolated audio can be improved, the audio data can be played smoothly, and the stereo playback effect can be improved.
[0162] The specific data compression method can be as Figure 6 shown. Suppose the audio frame lengths of the current frame and the next frame are N, that is, the merged audio data has 2N sampling point data, and the fourth adjustment point number is M'. By the cubic spline interpolation method, the 2N-point data is changed to 2N - M' point data.
[0163] Step 1048: Select the audio data corresponding to the first N interpolation nodes from the audio data after data compression as the new current frame audio data, and use the remaining audio data in the audio data after data compression as the new next frame audio data; where the value of N is the same as the number of sampling nodes of the current frame audio data.
[0164] In the embodiments of the present disclosure, when playing the compressed data, the amount of data of the current frame audio data is still the number of original sampling nodes, but the data of the sampling points is changed to the audio data corresponding to the first N interpolation nodes in the compressed audio data. The audio data corresponding to the remaining N - M' interpolation nodes in the compressed audio data is used as the new next frame audio data. After the new current frame audio data is played, the new next frame audio data is played. In this way, the playing duration of the new next frame audio data will be reduced, so that the playing progress of the slave device can be adjusted to be consistent with that of other devices when the new next frame audio data is played.
[0165] According to the audio playing method proposed by the embodiments of the present disclosure, for the slave device in a multi-device combination, by predicting the broadcast time of the next frame audio data before playing the current frame, and for the case where the predicted broadcast time is inconsistent with the target broadcast time, adjusting the playing speed of the next frame audio data, it is possible to avoid the situation that the audio playing progress of the slave device is slow or fast due to factors such as hardware device differences, and thus ensure the synchronization of audio playing between the slave device and other devices. In addition, for the case where the predicted broadcast time is inconsistent with the target broadcast time, the cubic spline interpolation method is introduced to expand or compress the merged data of the current frame and the next frame audio to adjust the playing speed, which improves the continuity of the audio data, so that the adjusted audio can still be played smoothly, and thus ensures the stereo effect of the multi-device combination audio playing and improves the user experience.
[0166] Since there are various ways to adjust the audio playing speed, the linear prediction coding method will be used to adjust the audio playing speed next.
[0167] Figure 11 It is a flowchart of another method for adjusting the playing speed of the next frame audio data in the embodiments of the present disclosure. As Figure 11 shown, on the basis of the above embodiments, the method further includes:
[0168] Step 1141, if the predicted broadcast time is less than the second target broadcast time, determine the third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time.
[0169] It can be understood that the predicted broadcast time being less than the second target broadcast time indicates that the current audio playing progress is fast, and data expansion is required to reduce the audio playing progress.
[0170] In the embodiments of the present disclosure, the third adjustment point number refers to the amount of audio data that needs to be expanded, and its calculation method can be as follows: subtract the predicted broadcast time from the second target broadcast time to obtain a time difference t'; according to the current device hardware conditions and network conditions, obtain the current audio playback speed v' of the slave device (the amount of audio data played per unit time); multiply the time difference t by the current audio playback speed v' of the slave device to obtain the first adjustment point number M' = t' × v'.
[0171] Step 1142, based on the linear predictive coding method, predict the sampling point data corresponding to the third adjustment point number according to the current frame audio data, and use the predicted sampling point data and the current frame audio data as the new current frame audio data for playback to adjust the playback speed of the next frame of audio data.
[0172] It can be understood that by also using the predicted sampling point data as the current frame audio data, the data volume of the current frame audio data is expanded, so that the audio playback speed can be slowed down, and then the playback speed of the next frame of audio data can be adjusted.
[0173] It should be noted that the basic idea of linear predictive coding is that the current value of an audio sampling point can be approximated by a weighted linear combination of the past values of several audio sampling points.
[0174] As an example, based on the linear predictive coding method, the implementation method of predicting the sampling point data corresponding to the third adjustment point number according to the current frame audio data can be as shown in the above formula (1).
[0175] In the embodiments of the present disclosure, if the predicted number of sampling points is M', and the number of sampling points of the current frame audio is N', then the data corresponding to the M'+N' sampling points are all used as the new current frame audio data for playback, thereby increasing the data volume of the current frame audio data, and then slowing down the audio playback speed, so that the audio data can be played synchronously.
[0176] Step 1143, if the predicted broadcast time is greater than the second target broadcast time, then reduce the number of sampling points in the audio frame according to the difference between the predicted broadcast time and the second target broadcast time to adjust the playback speed of the next frame of audio data.
[0177] As an example, it can be preset to remove one sampling point from each frame of audio; calculate the third adjustment point number M' according to the predicted broadcast time and the second target broadcast time, that is, the number of sampling points M' to be reduced, and evenly distribute the M' sampling points to be reduced to the current frame and the M'-1 frames of audio after the current frame; randomly select a sampling point in each frame of audio for removal, and use the audio data after removal as the new audio data for playback. If there are N' sampling points in the original audio frame, then there are N'-1 sampling points in each frame of the processed audio data. In this way, the audio playback speed is gradually adjusted by reducing the sampling points frame by frame.
[0178] The audio playback method proposed in the embodiments of the present disclosure is for the slave device in the multi-device combination. Based on the situation that the predicted broadcast time is inconsistent with the target broadcast time, a linear prediction coding method is introduced to expand the data of the current frame to adjust the playback speed. Among them, the data of the predicted sampling points can ensure the continuity of the audio data, so that the adjusted audio can still be played smoothly, thus ensuring the stereo effect of the multi-device combination audio playback and improving the user experience.
[0179] In addition, for the adjustment of the playback speed of the next frame of audio data, it can also be carried out by controlling the hardware settings. The present disclosure proposes another embodiment for this method. Figure 12 It is a flowchart for adjusting the playback speed of the next frame of audio data in the embodiments of the present disclosure. As Figure 12 shown, on the basis of the above embodiments, the method further includes:
[0180] Step 1241, if the predicted broadcast time is less than the second target broadcast time, then reduce the current playback sampling rate of the hardware driver in the slave device to adjust the playback speed of the next frame of audio data.
[0181] Step 1242, if the predicted broadcast time is greater than the second target broadcast time, then increase the current playback sampling rate of the hardware driver in the slave device to adjust the playback speed of the next frame of audio data.
[0182] For the adjustment of the audio playback speed, it can be achieved by the playback sampling rate of the hardware driver in the device. As an example, if the predicted playback time is less than the second target playback time, that is, the current audio is playing too fast. Assuming that the current sampling rate of the device speaker driver is 48 kHz, then increasing the playback sampling rate of the device speaker driver can reduce the playback speed of the device audio. For example, adjusting the sampling rate of the speaker driver to 44.1 kHz; thereby adjusting the playback speed of the next frame of audio data so that the device can play audio synchronously with other slave devices and the master device; on the contrary, if the current audio playback speed is slow, the playback sampling rate of the device hardware driver can be adjusted from 44.1 kHz to 48 kHz to adjust the playback speed of the next frame of audio data.
[0183] The audio playback method proposed in the embodiments of the present disclosure is for the slave device in a multi-device combination. Based on the adjustment of the playback speed of the next frame of audio when the predicted playback time is inconsistent with the target playback time, it is proposed to adjust the playback speed of the next frame of audio data by adjusting the playback sampling rate of the slave device hardware driver, which is equivalent to proposing another way to adjust the audio data playback speed. It can not only meet the stereo effect of multi-device combination audio playback, but also improve the applicability of the method in actual applications.
[0184] To implement the above embodiments, the present disclosure proposes an audio playback device.
[0185] Figure 13 It is a structural block diagram of an audio playback device proposed in the embodiments of the present disclosure. The device is applied to the master device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. As Figure 13 shown, the device includes:
[0186] A division processing module 1310, configured to obtain an audio data stream to be played and divide the audio data stream into multiple frames of audio data;
[0187] A sending module 1320, configured to generate corresponding audio packets according to each frame of audio data and send the generated multiple frames of audio packets to each slave device;
[0188] An obtaining module 1330, configured to obtain the next frame of audio data before playing the current frame of audio data;
[0189] A determining module 1340, configured to determine the first target playback time of the next frame of audio data;
[0190] An adjustment module 1350, configured to adjust the playback speed of the next frame of audio data according to the first target playback time.
[0191] In some embodiments of the present disclosure, the adjustment module 1350 includes:
[0192] A first determination unit 1351, configured to determine the current time of the master device;
[0193] A prediction unit 1352, configured to predict the required time when writing the next frame of audio data into the speaker in the master device and playing it out according to the hardware performance of the master device;
[0194] A second determination unit 1353, configured to determine the predicted playing time of the next frame of audio data according to the current time and the required time;
[0195] An adjustment unit 1354, configured to adjust the playing speed of the next frame of audio data when the predicted playing time is inconsistent with the first target playing time.
[0196] Furthermore, in some embodiments of the present disclosure, the adjustment unit 1354 is specifically configured to:
[0197] If the predicted playing time is less than the first target playing time, merge the current frame of audio data and the next frame of audio data;
[0198] Determine a first adjustment point number according to the difference between the predicted playing time and the first target playing time;
[0199] Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the first adjustment point number to obtain the audio data after data expansion;
[0200] Select the audio data corresponding to the first N interpolation nodes from the audio data after data expansion as the new current frame of audio data, and use the remaining audio data in the audio data after data expansion as the new next frame of audio data; where the value of N is the same as the number of sampling nodes of the current frame of audio data.
[0201] In other embodiments of the present disclosure, the adjustment unit 1354 is specifically configured to:
[0202] If the predicted playing time is greater than the first target playing time, merge the current frame of audio data and the next frame of audio data;
[0203] Determine a second adjustment point number according to the difference between the predicted playing time and the first target playing time;
[0204] Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the second adjustment point number to obtain the audio data after data compression;
[0205] Select the audio data corresponding to the first N interpolation nodes from the compressed audio data as the new current frame audio data, and use the remaining audio data in the compressed audio data as the new next frame audio data; where the value of N is the same as the number of sampling nodes of the current frame audio data.
[0206] In some other embodiments of the present disclosure, the adjustment unit 1354 is specifically configured to:
[0207] If the predicted broadcast time is less than the first target broadcast time, determine the first adjustment point number according to the difference between the predicted broadcast time and the first target broadcast time;
[0208] Based on the linear prediction coding method, predict the sampling point data corresponding to the number of the first adjustment point number according to the current frame audio data, and use the predicted sampling point data and the current frame audio data as the new current frame audio data for playing to adjust the playing speed of the next frame audio data.
[0209] In some other embodiments of the present disclosure, the adjustment unit 1354 is specifically configured to:
[0210] If the predicted broadcast time is less than the first target broadcast time, reduce the current playing sampling rate of the hardware driver in the master device to adjust the playing speed of the next frame audio data.
[0211] In some other embodiments of the present disclosure, the adjustment unit 1354 is specifically configured to:
[0212] If the predicted broadcast time is greater than the first target broadcast time, increase the current playing sampling rate of the hardware driver in the master device to adjust the playing speed of the next frame audio data.
[0213] In some embodiments of the present disclosure, the determination module 1340 is specifically configured to:
[0214] Determine the broadcast response time pre-negotiated between the master device and the slave device;
[0215] According to the start time of the next frame audio data and the broadcast response time, determine the first target broadcast time of the next frame audio data.
[0216] In addition, in some embodiments of the present disclosure, the sending module 1320 is specifically configured to:
[0217] While sending the audio packet to the slave device, also interact with the slave device for clock information.
[0218] Optionally, in some embodiments of the present disclosure, the audio playback device further includes:
[0219] A playback module 1360, configured to play the new next frame audio data after the current frame audio data is played.
[0220] According to the audio playback device of the embodiments of the present disclosure, the master device divides the audio data stream to be played into multiple frames of audio data, and sends the generated audio packets after division to each slave device, so that the synchronous playback of audio data can be realized through the transmission of audio data between the master device and the slave devices. In addition, the master device adjusts the playback speed of the next frame of audio data according to the target playback time of the next frame of audio data, so that the influence on the playback speed due to differences in hardware devices or network jitter can be avoided. In addition, the cubic spline interpolation method is introduced to expand or compress the merged data of the current frame and the next frame of audio to adjust the playback speed, improving the continuity of the audio data, so that the adjusted audio can still be played smoothly, thereby ensuring the stereo effect of multi-device combined audio playback and improving the user experience.
[0221] For the slave devices in the multi-device combination, the embodiments of the present disclosure propose another audio playback device.
[0222] Figure 10 It is a structural block diagram of another audio playback device proposed by the embodiments of the present disclosure. This device is applied to the slave devices in the multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. As Figure 10 shown, the device includes:
[0223] A receiving module 1410, configured to receive multiple frames of audio packets sent by the master device in the multi-device combination;
[0224] An analysis module 1420, configured to analyze each frame of audio packet to obtain the audio data in each frame of audio packet;
[0225] An obtaining module 1430, configured to obtain the next frame of audio data before playing the current frame of audio data;
[0226] A determining module 1440, configured to determine the second target playback time of the next frame of audio data;
[0227] An adjustment module 1450, configured to adjust the playback speed of the next frame of audio data according to the second target playback time.
[0228] In some embodiments of the present disclosure, the adjustment module 1450 includes:
[0229] A first determining unit 1451, configured to determine the current time of the slave device;
[0230] A prediction unit 1452, configured to predict the time required to write the next frame of audio data into the speaker of the slave device and be played according to the hardware performance of the slave device;
[0231] A second determination unit 1453, configured to determine a predicted broadcast time of the next frame of audio data according to the current time and the required time;
[0232] An adjustment unit 1454, configured to adjust the playback speed of the next frame of audio data in response to a difference between the predicted broadcast time and the second target broadcast time.
[0233] In some embodiments of the present disclosure, the adjustment unit 1454 is specifically configured to:
[0234] If the predicted broadcast time is less than the second target broadcast time, merge the current frame of audio data and the next frame of audio data;
[0235] Determine a third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time;
[0236] Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the third adjustment point number to obtain audio data with data expansion;
[0237] Select the audio data corresponding to the first N interpolation nodes from the audio data with data expansion as the new current frame of audio data, and use the remaining audio data in the audio data with data expansion as the new next frame of audio data; where the value of N is the same as the number of sampling nodes of the current frame of audio data.
[0238] In some other embodiments of the present disclosure, the adjustment unit 1454 is specifically configured to:
[0239] If the predicted broadcast time is greater than the second target broadcast time, merge the current frame of audio data and the next frame of audio data;
[0240] Determine a fourth adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time;
[0241] Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the fourth adjustment point number to obtain audio data with data compression;
[0242] Select the audio data corresponding to the first N interpolation nodes from the audio data with data compression as the new current frame of audio data, and use the remaining audio data in the audio data with data compression as the new next frame of audio data; where the value of N is the same as the number of sampling nodes of the current frame of audio data.
[0243] In some other embodiments of the present disclosure, the adjustment unit 1454 is specifically configured to:
[0244] If the predicted broadcast time is less than the second target broadcast time, determine a third adjustment point count according to the difference between the predicted broadcast time and the second target broadcast time;
[0245] Based on the linear prediction coding method, predict the sample point data corresponding to the third adjustment point count according to the current frame audio data, and use the predicted sample point data and the current frame audio data as the new current frame audio data for playback to adjust the playback speed of the next frame audio data.
[0246] In some other embodiments of the present disclosure, the adjustment unit 1454 is specifically configured to:
[0247] If the predicted broadcast time is less than the second target broadcast time, reduce the current playback sampling rate of the hardware driver in the slave device to adjust the playback speed of the next frame audio data.
[0248] In some other embodiments of the present disclosure, the adjustment unit 1454 is specifically configured to:
[0249] If the predicted broadcast time is greater than the second target broadcast time, increase the current playback sampling rate of the hardware driver in the slave device to adjust the playback speed of the next frame audio data.
[0250] In some embodiments of the present disclosure, the determination module 1440 is specifically configured to:
[0251] Interact with the master device for clock information, and calculate the system clock difference from the master device according to the system clock information of the master device;
[0252] Determine the broadcast response time pre-negotiated between the master device and the slave device;
[0253] According to the start time of the next frame audio data, the broadcast response time, and the system clock difference, determine the second target broadcast time of the next frame audio data.
[0254] In addition, in some embodiments of the present disclosure, the device further includes:
[0255] A playback module 1460, configured to play the new next frame audio data after the new current frame audio data is played.
[0256] According to the audio playback device proposed in the embodiments of the present disclosure, by receiving and parsing multiple frames of audio data sent by the slave device to the master device, the audio data corresponding to each frame of audio packet is obtained, and the audio data is played according to the broadcast time, so that synchronous playback of audio data between the master device and the slave device can be achieved. In addition, the slave device adjusts the playback speed of the next frame of audio data according to the target broadcast time of the next frame of audio data. In this way, the influence on the playback speed due to differences in hardware devices or network jitter can be avoided. In addition, in response to the situation where the predicted broadcast time is inconsistent with the target broadcast time, the cubic spline interpolation method is introduced to expand or compress the data merged from the current frame and the next frame of audio to adjust the playback speed, improving the continuity of the audio data, so that the adjusted audio can still be played smoothly, thereby ensuring the stereo effect of multi-device combined audio playback and improving the user experience.
[0257] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0258] Figure 15 FIG. is a block diagram of a device 1500 for audio playback according to an exemplary embodiment. For example, the device 1500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a speaker device, a personal digital assistant, etc.
[0259] Referring to Figure 15 , the device 1500 may include one or more of the following components: a processing component 1502, a memory 1504, a power component 1506, a multimedia component 1508, an audio component 1510, an input / output (I / O) interface 1512, a sensor component 1514, and a communication component 1516.
[0260] The processing component 1502 generally controls the overall operation of the device 1500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1502 may include one or more processors 1520 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1502 may include one or more modules to facilitate the interaction between the processing component 1502 and other components. For example, the processing component 1502 may include a multimedia module to facilitate the interaction between the multimedia component 1508 and the processing component 1502.
[0261] The memory 1504 is configured to store various types of data to support the operation of the device 1500. Examples of such data include instructions for any application or method operating on the device 1500, contact data, phone book data, messages, pictures, videos, and the like. The memory 1504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0262] The power component 1506 provides power to the various components of the device 1500. The power component 1506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 1500.
[0263] The multimedia component 1508 includes a screen that provides an output interface between the device 1500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1508 includes a front camera and / or a rear camera. When the device 1500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0264] The audio component 1510 is configured to output and / or input audio signals. For example, the audio component 1510 includes a microphone (MIC) that is configured to receive external audio signals when the device 1500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1504 or transmitted via the communication component 1516. In some embodiments, the audio component 1510 further includes a speaker for outputting audio signals.
[0265] The I / O interface 1512 provides an interface between the processing component 1502 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, and the like. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0266] The sensor assembly 1514 includes one or more sensors for providing a status assessment of various aspects of the device 1500. For example, the sensor assembly 1514 can detect the on / off state of the device 1500, the relative positioning of components, such as the display and keypad of the device 1500. The sensor assembly 1514 can also detect a change in the position of the device 1500 or a component of the device 1500, the presence or absence of user contact with the device 1500, the orientation or acceleration / deceleration of the device 1500, and the temperature change of the device 1500. The sensor assembly 1514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0267] The communication component 1516 is configured to facilitate communication between the device 1500 and other devices in a wired or wireless manner. The device 1500 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0268] In an exemplary embodiment, the device 1500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0269] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1504 including instructions, and the above instructions can be executed by a processor 1520 of the device 1500 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0270] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include known common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the invention are pointed out by the following claims.
[0271] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An audio playback method, characterized in that, the method is applied to the master device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The method includes: Obtaining an audio data stream to be played, and dividing the audio data stream into multiple frames of audio data; Generating corresponding audio packets according to each frame of audio data, and sending the generated multiple frames of audio packets to the slave devices in the multi-device combination, and notifying the slave devices to perform real-time playback of the audio; In response to real-time audio playback, before playing the current frame of audio data, determining the first target playback time of the next frame of audio data; Determining the current time of the master device; Predicting the time required to write the next frame of audio data into the speaker in the master device and play it according to the hardware performance of the master device; Determining the predicted playback time of the next frame of audio data according to the current time and the required time; If the predicted playback time is inconsistent with the first target playback time, then perform the following adjustments, including: If the predicted playback time is less than the first target playback time, then merge the current frame of audio data and the next frame of audio data, determine the first adjustment point number according to the difference between the predicted playback time and the first target playback time, and obtain the audio data after data expansion through interpolation; If the predicted playback time is greater than the first target playback time, then merge the current frame of audio data and the next frame of audio data, determine the second adjustment point number according to the difference between the predicted playback time and the first target playback time, and obtain the audio data after data compression through interpolation; Selecting the audio data corresponding to the first N interpolation nodes from the audio data after data expansion or data compression as the new current frame of audio data, and using the remaining audio data as the new next frame of audio data; Or, If the predicted playback time is less than the first target playback time, then determine the first adjustment point number according to the difference between the predicted playback time and the first target playback time, and based on the linear prediction coding method, predict the sampling point data corresponding to the number of the first adjustment point number according to the current frame of audio data, and use the predicted sampling point data and the current frame of audio data as the new current frame of audio data for playback.
2. The audio playback method according to claim 1, characterized in that, Based on the cubic spline interpolation method, interpolating the sampling points corresponding to the audio data obtained after merging according to the first adjustment point number or the second adjustment point number to obtain the audio data after data expansion or data compression; The value of N is the same as the number of sampling nodes of the current frame of audio data.
3. The audio playback method according to claim 2, characterized in that, further includes: After the new current frame of audio data is played, playing the new next frame of audio data.
4. The audio playback method according to claim 1, characterized in that, The determining the first target playback time of the next frame of audio data includes: Determine the broadcast response time pre-negotiated between the master device and the slave device; Determine the first target broadcast time of the next frame of audio data according to the start time of the next frame of audio data and the broadcast response time.
5. The audio playback method according to claim 1, characterized in that, while sending the audio packet to the slave device, the method further includes: interacting with the slave device for clock information.
6. An audio playback method, characterized in that, the method is applied to a slave device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The method includes: Receive multiple frames of audio packets sent by the master device in the multi-device combination; Parse each frame of audio packet to obtain the audio data in each frame of audio packet; In response to the notification sent by the master device in the multi-devices to play audio in real time, before playing the current frame of audio data, determine the second target broadcast time of the next frame of audio data; Determine the current time of the slave device; Predict the time required to write the next frame of audio data into the speaker of the slave device and be broadcast according to the hardware performance of the slave device; Determine the predicted broadcast time of the next frame of audio data according to the current time and the required time; If the predicted broadcast time is inconsistent with the second target broadcast time, the following adjustments are made, including: If the predicted broadcast time is less than the second target broadcast time, merge the current frame of audio data and the next frame of audio data, determine the third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and obtain the audio data after data expansion through interpolation; If the predicted broadcast time is greater than the second target broadcast time, merge the current frame of audio data and the next frame of audio data; determine the fourth adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and obtain the audio data after data compression through interpolation; Select the audio data corresponding to the first N interpolation nodes from the audio data after data expansion or data compression as the new current frame of audio data, and use the remaining audio data as the new next frame of audio data; Or, If the predicted broadcast time is less than the second target broadcast time, determine the third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and based on the linear prediction coding method, predict the sampling point data corresponding to the number of the third adjustment point number according to the current frame of audio data, and use the predicted sampling point data and the current frame of audio data as the new current frame of audio data for playback.
7. The audio playback method according to claim 6, characterized in that, Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the third adjustment point number or the fourth adjustment point number to obtain the audio data after data expansion or data compression; The value of N is the same as the number of sampling nodes of the current frame of audio data.
8. The audio playback method according to claim 7, Characterized in that, It further includes: After the new current frame audio data is played, play the new next frame audio data.
9. The audio playback method according to claim 6, Characterized in that, The determination of the second target broadcast time of the next frame audio data includes: Interact with the master device for clock information and calculate the system clock difference with the master device according to the system clock information of the master device; Determine the broadcast response time pre-negotiated between the master device and the slave device; According to the start time of the next frame audio data, the broadcast response time and the system clock difference, determine the second target broadcast time of the next frame audio data.
10. An audio playback device, Characterized in that, The device is applied to the master device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The device includes: A division processing module, configured to obtain an audio data stream to be played and divide the audio data stream into multiple frames of audio data; A sending module, configured to generate corresponding audio packets according to each frame of audio data and send the generated multiple frames of audio packets to each slave device; An obtaining module, configured to obtain the next frame of audio data before playing the current frame of audio data; A determination module, configured to determine the first target broadcast time of the next frame of audio data; An adjustment module, including: A first determination unit, configured to determine the current time of the master device; A prediction unit, configured to predict the time required to write the next frame of audio data into the speaker of the master device and be broadcast according to the hardware performance of the master device; A second determination unit, configured to determine the predicted broadcast time of the next frame of audio data according to the current time and the required time; An adjustment unit, configured to perform the following adjustments when the predicted broadcast time is inconsistent with the first target broadcast time, including: If the predicted broadcast time is less than the first target broadcast time, merge the current frame audio data and the next frame audio data; determine the first adjustment point number according to the difference between the predicted broadcast time and the first target broadcast time and obtain the audio data after data expansion through interpolation; If the predicted broadcast time is greater than the first target broadcast time, merge the current frame audio data and the next frame audio data, determine the second adjustment point number according to the difference between the predicted broadcast time and the first target broadcast time and obtain the audio data after data compression through interpolation; Select the audio data corresponding to the first N interpolation nodes from the audio data after data expansion or data compression as the new current frame audio data, and use the remaining audio data as the new next frame audio data; Or, If the predicted broadcast time is less than the first target broadcast time, then, according to the difference between the predicted broadcast time and the first target broadcast time, determine the first adjustment point number. Based on the linear prediction coding method, predict the sampling point data corresponding to the number of the first adjustment point number according to the current frame audio data, and use the predicted sampling point data and the current frame audio data as the new current frame audio data for playing.
11. The audio playback device according to claim 10, wherein, Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the first adjustment point number or the second adjustment point number to obtain the audio data after data expansion or data compression; The value of N is consistent with the number of sampling nodes of the current frame audio data.
12. The audio playback device according to claim 11, wherein, further comprising: A playback module, configured to play the new next frame audio data after the current frame audio data is played.
13. The audio playback device according to claim 10, wherein, The determining module is specifically configured to: Determine the broadcast response time pre-negotiated between the master device and the slave device; According to the start time of the next frame audio data and the broadcast response time, determine the first target broadcast time of the next frame audio data.
14. The audio playback device according to claim 10, wherein, The sending module is specifically configured to: While sending the audio packet to the slave device, also interact clock information with the slave device.
15. An audio playback device, wherein, The device is applied to a slave device in a multi-device combination, and the multi-device combination realizes audio synchronization for stereo playback. The device includes: A receiving module, configured to receive multiple frames of audio packets sent by a master device in the multi-device combination; An analysis module, configured to analyze each frame of audio packet to obtain the audio data in each frame of audio packet; An obtaining module, configured to obtain the next frame audio data before playing the current frame audio data; A determining module, configured to determine the second target broadcast time of the next frame audio data; An adjustment module, including: A first determining unit, configured to determine the current time of the slave device; A prediction unit, configured to predict the time required to write the next frame audio data into the speaker of the slave device and be broadcast according to the hardware performance of the slave device; A second determining unit, configured to determine the predicted broadcast time of the next frame audio data according to the current time and the required time; An adjustment unit, configured to perform the following adjustments when the predicted broadcast time is inconsistent with the second target broadcast time, including: If the predicted broadcast time is less than the second target broadcast time, then merge the current frame audio data and the next frame audio data, determine the third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and obtain the audio data after data expansion through interpolation; If the predicted broadcast time is greater than the second target broadcast time, merge the current frame audio data and the next frame audio data; determine the fourth adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and obtain the audio data after data compression through interpolation; Select the audio data corresponding to the first N interpolation nodes from the audio data after data augmentation or data compression as the new current frame audio data, and use the remaining audio data as the new next frame audio data; or, If the predicted broadcast time is less than the second target broadcast time, determine the third adjustment point number according to the difference between the predicted broadcast time and the second target broadcast time, and based on the linear predictive coding method, predict the sampling point data corresponding to the number of the third adjustment point number according to the current frame audio data, and use the predicted sampling point data and the current frame audio data as the new current frame audio data for playback.
16. The audio playback device according to claim 15, wherein, Based on the cubic spline interpolation method, interpolate the sampling points corresponding to the audio data obtained after merging according to the third adjustment point number or the fourth adjustment point number to obtain the audio data after data augmentation or data compression; The value of N is the same as the number of sampling nodes of the current frame audio data.
17. The audio playback device according to claim 16, wherein, further comprising: A playback module, configured to play the new next frame audio data after the new current frame audio data is played.
18. The audio playback device according to claim 15, wherein, The determining module is specifically configured to: Interact with the master device for clock information, and calculate the system clock difference from the master device according to the system clock information of the master device; Determine the broadcast response time pre-negotiated between the master device and the slave device; Determine the second target broadcast time of the next frame audio data according to the start time of the next frame audio data, the broadcast response time, and the system clock difference.
19. An electronic device, wherein, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, when the processor executes the computer program, implementing the audio playback method according to any one of claims 1 to 5, and / or implementing the audio playback method according to any one of claims 6 to 9.
20. A temporary computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by a processor, implementing the audio playback method according to any one of claims 1 to 5, and / or implementing the audio playback method according to any one of claims 6 to 9.
Citation Information
Patent Citations
Audio synchronous play method, audio synchronous play device, audio synchronous play system and terminal
CN106373600A
Audio synchronization method and system of bluetooth equipment
CN108111997A