Audio playing method, first terminal, medium and computing device

By obtaining the server's transmission latency and the delay duration of the audio playback on the terminal, the chorus time is corrected, which solves the problem of audio non-overlap caused by timer differences and optimizes the audio chorus effect.

CN116016979BActive Publication Date: 2026-01-27HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211565897.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-07
Publication Date
2026-01-27
Estimated Expiration
2042-12-07

AI Technical Summary

Technical Problem

Because the timer structures of different terminals are different, users start playing audio at different times, resulting in poor audio chorus effects. Users may mistakenly believe that the rhythm of their own singing is wrong.

Method used

By obtaining the server's transmission latency, the time difference between the terminal and the server, and the delay in the terminal's audio playback, the agreed-upon chorus time is corrected to optimize the chorus effect of the audio.

Benefits of technology

It effectively avoids audio mismatch issues caused by network latency, terminal playback delay, and inaccurate local time, thus improving the synchronization and effect of audio chorus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116016979B_ABST
    Figure CN116016979B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio playing method, a first terminal, a medium and a computing device. The method comprises: obtaining a first time point of a first terminal singing in chorus with one or more second terminals; determining a first time offset according to a local time of the first terminal and a standard time of a server; correcting the first time point to obtain a target time point according to a current transmission delay between the server and the first terminal, a delay duration of the first terminal playing sound and the first time offset; and in response to the current time point reaching the target time point, receiving and playing audio transmitted by the server. In the present disclosure, by obtaining the transmission delay of the server, the difference between the time of the terminal and the server, and the delay duration of the terminal playing audio, the agreed chorus time is corrected, thereby avoiding the problem that the audio sung by each user does not coincide due to factors such as network delay, playing delay of the terminal and inaccuracy of the local time of the terminal, and optimizing the chorus effect of the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the audio field, and more specifically, the embodiments of this disclosure relate to an audio playback method, a first terminal, a medium, and a computing device. Background Technology

[0002] This section is intended to provide background or context for embodiments of this disclosure. The description herein is not intended to imply that it is prior art simply because it is included in this section.

[0003] With the development of the internet, users can invite friends to sing together online anytime, anywhere.

[0004] The terminals of the users participating in the chorus will agree on the time for the audio chorus, and during the chorus, the terminals will transmit the audio of the user's singing to each other, so that one user can hear the audio of other users singing.

[0005] The agreed-upon chorus time for each terminal is based on the terminal's local time. However, different terminals have different internal timer structures, resulting in different recorded local times. Consequently, each terminal will start playing audio at different times, causing the audio heard by the user to differ from that of other users. In this situation, the user may mistakenly believe that their own singing rhythm is incorrect and adjust the rhythm, ultimately leading to a poor chorus effect. Summary of the Invention

[0006] This disclosure provides an audio playback method, apparatus, medium, and computing device for improving the chorus effect of audio.

[0007] In a first aspect of this disclosure, an audio playback method is provided, applied to a first terminal, comprising: acquiring a first time point in which the first terminal and one or more second terminals sing together; determining a first time offset based on the local time of the first terminal and the standard time of a server; correcting the first time point to obtain a target time point based on the current transmission delay between the server and the first terminal, the delay duration of the sound played by the first terminal, and the first time offset; and receiving and playing audio transmitted by the server in response to the current time point reaching the target time point, wherein the audio is the audio of the first terminal and the second terminals singing together.

[0008] In one embodiment of this disclosure, determining the first time offset based on the local time of the first terminal and the standard time of the server includes: determining the current time offset based on the standard time and the local time, and obtaining various historical time offsets; determining a time offset to be compared among the various historical time offsets, wherein the delay corresponding to the time offset to be compared is the minimum delay among the delays corresponding to the various historical time offsets; determining the time offset to be compared as the first time offset in response to the delay corresponding to the time offset to be compared being less than the current transmission delay; and determining the current time offset as the first time offset in response to the delay corresponding to the time offset to be compared being greater than the current transmission delay.

[0009] In another embodiment of this disclosure, the audio playback method further includes: playing a preset sound and determining the playback time point of the preset sound; in response to the acquisition of the played preset sound, acquiring the acquisition time point; and acquiring the delay duration of the sound playback by the first terminal based on the interval between the playback time point and the acquisition time point.

[0010] In another embodiment of this disclosure, the audio playback method further includes: sending a time acquisition request to a server; receiving a standard time sent by the server based on the time acquisition request; and acquiring the current transmission delay between the server and the first terminal based on the interval between the sending time of the time acquisition request and the receiving time of the standard time.

[0011] In another embodiment of this disclosure, after receiving and playing the audio transmitted by the server, the method further includes: acquiring a first sound sung by a first user associated with the first terminal based on the audio; sending the first sound to the second terminal, wherein the first sound is used to adjust the playback progress of the audio played by the second terminal; receiving a second sound sent by the second terminal, and adjusting the playback progress of the audio according to the second sound, wherein the second sound is the sound sung by a second user corresponding to the second terminal based on the audio.

[0012] In another embodiment of this disclosure, adjusting the playback progress of the audio based on the second sound includes: determining the difference between the time point when the second sound is sent and the time point when the first terminal receives the second sound; and adjusting the playback progress of the audio in response to the difference being greater than a preset difference.

[0013] In another embodiment of this disclosure, the audio playback method further includes: adjusting the playback progress of the audio in response to the difference being greater than a preset difference and the interval between the last time the playback progress of the audio was adjusted and the current local time being greater than a preset interval.

[0014] In another embodiment of this disclosure, sending the first sound to the second terminal includes: obtaining a first progress of the audio playback on the first terminal, and sending the first sound and the first progress to the second terminal; receiving a second sound sent by the second terminal and adjusting the playback progress of the audio according to the second sound includes: receiving a second sound sent by the second terminal and a second progress of the audio playback on the second terminal; and adjusting the playback progress of the audio in response to a progress difference between the first progress and the second progress being greater than a preset threshold.

[0015] In another embodiment of this disclosure, playing the audio transmitted by the server includes playing the accompaniment to the audio.

[0016] In another embodiment of this disclosure, the step of correcting the first time point based on the current transmission delay between the server and the first terminal, the delay duration of sound playback on the first terminal, and the first time offset to obtain the target time point includes:

[0017] Based on the current transmission delay between the server and the first terminal and the delay duration of the sound playback on the first terminal, the first time offset is corrected to obtain the target time offset.

[0018] The first time point is corrected based on the target time offset to obtain the target time point.

[0019] In a second aspect of this disclosure, a first terminal is also provided, comprising: a first acquisition module for acquiring a first time point in which the first terminal and one or more second terminals sing together; a determination module for determining a first time offset based on the local time of the first terminal and the standard time of a server; a correction module for correcting the first time point to obtain a target time point based on the current transmission delay between the server and the first terminal, the delay duration of sound playback by the first terminal, and the first time offset; and a first playback module for receiving and playing audio transmitted by the server in response to the current time point reaching the target time point, wherein the audio is the audio of the first terminal and the second terminals singing together.

[0020] In one embodiment of this disclosure, the determining module includes: a first determining unit, configured to determine the current time offset based on the standard time and the local time, and to acquire each historical time offset; the first determining unit is further configured to determine a time offset to be compared among each of the historical time offsets, wherein the delay corresponding to the time offset to be compared is the minimum delay among the delays corresponding to each of the historical time offsets; the first determining unit is further configured to determine the time offset to be compared as a first time offset in response to the delay corresponding to the time offset to be compared being less than the current transmission delay; the first determining unit is further configured to determine the current time offset as a first time offset in response to the delay corresponding to the time offset to be compared being greater than the current transmission delay.

[0021] In another embodiment of this disclosure, the first terminal further includes: a second playback module, configured to play a preset sound and determine the playback time of the preset sound; and a second acquisition module, configured to...

[0022] The first acquisition module is used to acquire the acquisition time point when the preset sound is acquired; the second acquisition module is also used to acquire the delay duration of the sound played by the first terminal based on the interval between the playback time point and the acquisition time point.

[0023] In another embodiment of this disclosure, the first terminal further includes: a first sending module, configured to send a time acquisition request to the server; and a first receiving module, configured to receive a time acquisition request from the server based on the request.

[0024] The time acquisition request is sent at a standard time; the third acquisition module is used to acquire the current transmission delay between the server and the first terminal based on the interval between the sending time of the time acquisition request and the receiving time of the standard time.

[0025] In another embodiment of this disclosure, the first terminal further includes: a collection module, configured to collect a first user base associated with the first terminal after receiving and playing the audio transmitted by the server.

[0026] The system comprises: a first sound from the audio performance; a second sending module for sending the first sound to the second terminal, wherein the first sound is used to adjust the playback progress of the audio on the second terminal; and a second receiving module for receiving a second sound from the second terminal and, based on the second sound...

[0027] The playback progress of the audio is adjusted by the sound, wherein the second sound is the sound sung by the second user corresponding to the second terminal based on the audio.

[0028] In another embodiment of this disclosure, the second receiving module includes: a second determining unit, configured to determine the difference between the time point at which the second sound is transmitted and the time point at which the first terminal receives the second sound; and a first adjusting unit, configured to adjust the playback progress of the audio in response to the difference being greater than a preset difference.

[0029] 5. In another embodiment of this disclosure, the first terminal further includes: an adjustment module, configured to respond to

[0030] If the difference is greater than a preset difference, and the interval between the last time the playback progress of the audio was adjusted and the current local time is greater than a preset interval, the playback progress of the audio is adjusted.

[0031] In another embodiment of this disclosure, the second sending module includes: an acquisition unit, configured to acquire a first progress of the audio played by the first terminal, and send the first sound and the first progress to the second terminal; the second receiving module includes: a receiving unit, configured to receive a second sound sent by the second terminal and a second progress of the audio played by the second terminal; and a second adjustment unit, configured to adjust the playback progress of the audio in response to a progress difference between the first progress and the second progress being greater than a preset threshold.

[0032] In another embodiment of this disclosure, the playback module includes: a playback unit for playing the accompaniment to the audio.

[0033] In another embodiment of this disclosure, the correction module includes: a correction unit, configured to correct the first time offset based on the current transmission delay between the server and the first terminal and the delay duration of the sound played by the first terminal, to obtain a target time offset; the correction unit is configured to correct the first time point based on the target time offset to obtain a target time point.

[0034] In a third aspect of this disclosure, a medium is also provided, comprising: computer execution instructions, which, when executed by a processor, are used to implement the method described above.

[0035] In a fourth aspect of this disclosure, a computing device is also provided, comprising:

[0036] Memory and processor;

[0037] The memory stores computer-executed instructions;

[0038] The processor executes computer execution instructions stored in the memory, causing the processor to perform the method described above.

[0039] In this embodiment, the agreed chorus time is corrected by obtaining the server's transmission latency, the time difference between the terminal and the server, and the latency of the terminal's audio playback. This avoids the problem of non-overlapping audio recordings sung by different users due to factors such as network latency, terminal playback latency, and inaccurate local time of the terminal, thereby optimizing the chorus effect of the audio. Attached Figure Description

[0040] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:

[0041] Figure 1 A schematic diagram illustrating an application scenario of the audio playback method according to embodiments of the present disclosure is shown.

[0042] Figure 2 A schematic flowchart according to an embodiment of the present disclosure is shown;

[0043] Figure 3 A schematic flowchart according to another embodiment of the present disclosure is shown;

[0044] Figure 4 A schematic flowchart according to yet another embodiment of the present disclosure is shown;

[0045] Figure 5 A schematic flowchart according to another embodiment of the present disclosure is shown;

[0046] Figure 6 A schematic diagram of a program product provided according to an embodiment of the present disclosure is shown;

[0047] Figure 7 A schematic diagram of the structure of a first terminal provided according to an embodiment of the present disclosure is shown.

[0048] Figure 8 A schematic diagram of the structure of a computing device provided according to an embodiment of the present disclosure is shown.

[0049] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0050] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0051] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0052] According to embodiments of this disclosure, an audio playback method, a first terminal, a medium, and a computing device are proposed.

[0053] Furthermore, the number of any elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.

[0054] In addition, the data involved in this disclosure may be data authorized by the user or fully authorized by all parties. The collection, dissemination and use of the data shall comply with the requirements of relevant national laws and regulations. The implementation methods / executives of this disclosure may be combined with each other.

[0055] The principles and spirit of this disclosure will be explained in detail below with reference to several representative embodiments. Invention Overview

[0057] The terminals of the users participating in the chorus will agree on the time for the audio chorus, and during the chorus, the terminals will transmit the audio of the user's singing to each other, so that one user can hear the audio of other users singing.

[0058] The inventors of this disclosure discovered that while the agreed-upon chorus time for each terminal is its local time, the different timer structures within each terminal result in different terminals starting to play audio at different times. This leads to a mismatch between the audio of other users and the user's own recording. In this situation, users may mistakenly believe their own recording is out of rhythm and adjust it, ultimately resulting in a poor chorus effect.

[0059] The inventors of this disclosure therefore conceived of a way to modify the agreed-upon chorus time by obtaining the server's transmission latency, the time difference between the terminal and the server, and the latency of the terminal's audio playback. This avoids the problem of non-overlapping audio recordings from different users due to factors such as network latency, terminal playback latency, and inaccurate local time on the terminal, thus optimizing the chorus effect of the audio.

[0060] Application Scenarios Overview

[0061] First refer to Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the audio playback method according to an embodiment of this disclosure. Terminal device 100 sends a chorus request to server 200. The chorus request includes the chorus time point, the chorus audio, and one or more terminal devices 200 participating in the chorus. Server 200 sends the chorus time point to terminal device 300. Terminal devices 100 and 300 can be mobile phones, tablets, or computers. Terminal devices 100 and 300 adjust the chorus time point based on their network latency, sound playback latency, and local time, thereby ensuring that the audio played by terminal devices 100 and 300 is synchronized, meaning that the audio heard by each user participating in the chorus is synchronized on their respective terminal devices.

[0062] Exemplary methods

[0063] The following is combined with Figure 1 Application scenarios, refer to Figures 2-5 This document describes an audio playback method according to exemplary embodiments of the present disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in any way. Rather, the embodiments of the present disclosure can be applied to any applicable scenario.

[0064] Reference Figure 2 , Figure 2 An exemplary flowchart of an embodiment of an audio playback method provided according to this disclosure is shown. The audio playback method includes:

[0065] Step S201: Obtain the first time point of the duet between the first terminal and one or more second terminals.

[0066] In this embodiment, multiple users can engage in audio chorus through their respective terminal devices. For example, terminal device a initiates a chorus request to the server. The chorus request includes the audio of the chorus, the timing of the chorus, and one or more terminal devices b participating in the chorus. Terminal device a or terminal device b acts as the executing entity; for ease of description, the executing entity is defined as the first terminal, and the other terminal devices participating in the chorus with the first terminal are defined as the second terminal.

[0067] The first terminal obtains the time point at which it sings a duet with one or more second terminals, and this time point is defined as the first time point. If the first terminal is the terminal device that initiated the duet request, it obtains the first time point input by the user; if the first terminal is not the terminal device that initiated the duet request, it receives the time point of the duet sent by the server as the first time point.

[0068] Step S202: Determine the first time offset based on the local time of the first terminal and the standard time of the server.

[0069] The times of the various terminal devices in the choir may differ. Both the first and second terminals need to correct the initial time point of the choir performance, and the time offset is one of several parameters used to correct this initial time point. Specifically, the first terminal obtains the server's standard time and its own local time simultaneously, and determines the initial time offset using both the local and standard times. For example, subtracting the local time from the standard time yields the initial time offset.

[0070] Step S203: Correct the first time point according to the current transmission delay between the server and the first terminal, the delay duration of the sound played by the first terminal, and the first time offset to obtain the target time point.

[0071] Transmission delay refers to round-trip time, which represents the total time elapsed from when the sender transmits data to when the sender receives an acknowledgment from the receiver (the receiver sends an acknowledgment immediately upon receiving the data). Round-trip delay is determined by three components: the propagation time of the link, the processing time of the end system, and the queuing and processing time in the router's buffer. The first two components are relatively fixed, while the queuing and processing time in the router's buffer varies with the overall network congestion. Since the audio played by the terminal devices participating in the chorus is transmitted from the server, the impact of network congestion on the audio chorus must be considered; that is, the timing of the chorus needs to be adjusted based on the transmission delay between the server and the first terminal.

[0072] Furthermore, when various terminal devices play audio at the same time, due to the differences in the audio playback components of each terminal device, the audio played by one terminal device will reach the user's ears faster, while the audio played by another terminal device will reach the user's ears slower. In other words, the delay time of audio playback varies among terminal devices. The delay time refers to the interval between the time when the terminal device plays the audio and the time when the user hears the audio. Therefore, when various terminal devices sing together, the difference in the delay time of audio playback among the terminal devices also needs to be taken into account.

[0073] In response, the first terminal obtains the current transmission latency between the server and the first terminal, and also obtains the delay duration for the first terminal to play sound.

[0074] In one example, the current transmission delay and latency duration are pre-stored, and the first terminal directly retrieves the current transmission delay and latency duration from the storage space.

[0075] In another example, the first terminal sends a time acquisition request to the server. After receiving the time acquisition request, the server sends its standard time to the first terminal. That is, the first terminal receives the standard time sent by the server based on the time acquisition request. The first terminal can obtain the current transmission delay between the server and the first terminal based on the interval between the time acquisition request sending time and the standard time receiving time. In other words, the interval between the time acquisition request sending time and the standard time receiving time is the current transmission delay.

[0076] The first terminal can obtain the target time point by correcting the first time point using the first time offset, the current transmission delay, and the delay duration.

[0077] In one example, the first terminal adds or subtracts the first time offset, the current transmission delay, and the delay duration from the first time point to obtain the target time point.

[0078] In another example, the first terminal corrects the first time offset based on the current transmission delay and the delay duration to obtain the target time offset, and then corrects the first time point using the target time offset to obtain the target time point. For example, the current transmission delay and the delay duration can be added to or subtracted from the first time offset to obtain the target time offset, and then the target time offset can be added to or subtracted from the first time point to obtain the target time point.

[0079] It should be noted that the second terminal also needs to correct the first time point of the chorus to obtain the corresponding target time point, and the process of the second terminal correcting the first time point is the same as that of the first terminal, so it will not be repeated here.

[0080] Step S204: In response to the current time point reaching the target time point, receive and play the audio transmitted by the server. The audio is the audio of the first terminal and the second terminal singing together.

[0081] The first terminal continuously monitors whether the current time has reached the target time. If the current time has reached the target time, it receives and plays the audio transmitted by the server. This audio is a duet between the first and second terminals. Furthermore, since the users corresponding to the first and second terminals are performing a duet, the server can send the accompaniment to both terminals, meaning the first terminal plays the accompaniment. Of course, the first terminal can also play the complete audio, not just the accompaniment. It should be noted that the current time refers to the first terminal's local time.

[0082] The following example illustrates the playback of audio on the first and second terminals.

[0083] The first time point agreed upon by the first and second terminals for the chorus is 9:00:00 am. The transmission delay between the first terminal and the server is 1 second; the local time of the first terminal is 3 seconds later than the standard time of the server, that is, the first time offset of the first terminal is 3 seconds; the delay time for the first terminal to play the sound is 1 second. The first terminal first corrects the first time point based on the first time offset, that is, changes the first time point to 9:00:03 am. The first terminal then modifies 9:00:3 am based on the current transmission delay, that is, delays 9:00:03 am by 1 second to obtain 9:00:04 am. Finally, the first terminal modifies 9:00:04 am based on the delay time, that is, delays 9:00:04 am by 1 second to obtain the target time point 9:00:05 am.

[0084] The transmission delay between the second terminal and the server is 2 seconds; the local time of the first terminal is 2 seconds later than the standard time of the server, which means the second time offset of the second terminal is 2 seconds; the delay for the second terminal to play sound is 3 seconds. The second terminal first corrects the first time point based on the second time offset, that is, it modifies the first time point to 9:00:02am. The second terminal then modifies 9:00:02am based on the current transmission delay, that is, it delays 9:00:02am by 2 seconds to get 9:00:04am. Finally, the second terminal modifies 9:00:04am based on the delay duration, that is, it delays 9:00:04am by 3 seconds to get the target time point 9:00:07am.

[0085] As can be seen from the above, the first terminal receives and plays the audio sent by the server when its local time is 9:00:05 am, while the second terminal receives and plays the audio sent by the server when its local time is 9:00:07 am.

[0086] In this embodiment, the agreed chorus time is corrected by obtaining the server's transmission latency, the time difference between the terminal and the server, and the latency of the terminal playing audio. This avoids the problem of non-overlapping audio recordings sung by different users due to factors such as network latency, terminal playback delay, and inaccurate local time of the terminal, thus optimizing the chorus effect of the audio.

[0087] Reference Figure 3 , Figure 3 An exemplary flowchart of another embodiment of the audio playback method provided according to embodiments of this disclosure is shown, based on... Figure 2 In the embodiment shown, step S202 includes:

[0088] Step S301: Determine the current time offset based on the standard time and the local time, and obtain the time offsets of each historical time.

[0089] When the first terminal plays audio in the past, it records the time offset between the local time and the standard time, and also records the transmission delay between the first terminal and the server. In other words, the first terminal stores multiple historical time offsets and the corresponding transmission delays.

[0090] The first terminal can calculate the current time offset based on the current local time and the current standard time, which is obtained by subtracting the current local time from the current standard time. The first terminal then obtains the time offsets from various historical data.

[0091] Step S302: Determine the time offset to be compared from each historical time offset, wherein the delay corresponding to the time offset to be compared is the minimum delay among the delays corresponding to each historical time offset.

[0092] The smaller the transmission latency between the server and the first terminal, the smaller the error between the local time of the first terminal and the standard time of the server. Therefore, the time offset corresponding to the minimum error needs to be used as the first time offset for the first terminal to correct the first time point. The first terminal determines the time offset to be compared from the historical time offsets based on the transmission latency corresponding to each historical time offset, that is, the historical offset with the minimum latency is used as the time offset to be compared.

[0093] Step S303: In response to the fact that the delay corresponding to the time offset to be compared is less than the current transmission delay, the time offset to be compared is determined as the first time offset.

[0094] Step S304: In response to the fact that the delay corresponding to the time offset to be compared is greater than the current transmission delay, the current time offset is determined as the first time offset.

[0095] The first terminal determines whether the time offset to be compared is greater than the current time offset. If the delay corresponding to the time offset to be compared is less than the current transmission delay, then the time offset to be compared is determined as the first time offset; if the delay corresponding to the time offset to be compared is greater than the current transmission delay, then the current time offset is determined as the first time offset.

[0096] In this embodiment, the first terminal finds the time offset corresponding to the minimum transmission delay as the first time offset, thereby accurately correcting the first time point and ensuring that the audio heard by each user is synchronized.

[0097] Reference Figure 4 , Figure 4 An exemplary flowchart of yet another embodiment of the audio playback method provided according to embodiments of the present disclosure is shown, based on... Figure 2 or Figure 3 In the embodiment shown, before step S203, the method further includes:

[0098] Step S401: Play the preset sound and determine the playback time of the preset sound.

[0099] In this embodiment, before playing audio, the first terminal tests the latency of its own sound playback. The first terminal plays a preset sound, which is any audio stored in the first terminal. While playing the preset sound, the first terminal records the playback time of the preset sound.

[0100] Step S402: In response to the acquisition of the preset sound being played, obtain the acquisition time point.

[0101] Before playing a preset audio file, the first terminal activates the sound acquisition module to collect sound in real time. After acquiring sound, the first terminal extracts the voiceprint features of the acquired sound and calculates the similarity between these features and the voiceprint features associated with the preset sound. If the similarity is greater than a preset threshold, the acquired sound is confirmed to be the preset sound. The sound acquisition module records the corresponding acquisition time point each time it acquires sound, and the first terminal uses this recorded time point to determine the acquisition time point of the preset sound.

[0102] Step S403: Obtain the delay duration of the sound played on the first terminal based on the interval between the playback time point and the acquisition time point.

[0103] The first terminal calculates the interval between the playback time point and the acquisition time point, which is the delay time for the first terminal to play the sound. The first terminal stores the delay time so that when the first terminal needs to play the chorus audio later, it can read the stored delay time to correct the timing of the chorus.

[0104] In this embodiment, the first terminal plays a preset sound and collects the preset sound. By using the playback time of the preset sound and the collection time, the delay duration of its own playback sound is obtained. Then, the timing of the chorus is corrected by the delay duration, thus eliminating the problem of audio asynchrony among users caused by the hardware of the first terminal itself.

[0105] Reference Figure 5 , Figure 5 An exemplary flowchart of another embodiment of the audio playback method provided according to embodiments of the present disclosure is shown, based on... Figures 2 to 4 In any of the embodiments shown, after step S203, the method further includes:

[0106] Step S501: Collect the first sound of the first user singing based on audio, which is associated with the first terminal.

[0107] Step S502: Send the first sound to the second terminal. The first sound is used to adjust the progress of the audio played on the second terminal.

[0108] In this embodiment, the first and second terminals in the chorus send playback progress to each other, allowing them to adjust their own audio playback progress based on the other's. To this end, the first terminal collects the first sound of a first user singing based on the audio. The first sound includes the audio played by the first terminal and the sound of the first user singing. The first sound is a data packet carrying the accompaniment, the user's singing, and the time of transmission. The first user is the user singing the audio played by the first terminal.

[0109] The first terminal sends the first sound to the second terminal. Since the first sound is actually a sound data packet carrying the accompaniment of the audio and the sound of the user singing, the second terminal can adjust the playback progress of the audio based on the first sound. In other words, the first sound is used to adjust the playback progress of the audio on the second terminal.

[0110] Step S503: Receive the second sound sent by the second terminal and adjust the playback progress of the audio according to the second sound, wherein the second sound is the sound sung by the second user corresponding to the second terminal based on the audio.

[0111] The second terminal also sends the second audio it collects to the first terminal. The second audio includes the audio played by the first terminal and the audio sung by the second user. The second audio is actually a data packet carrying the accompaniment, the transmission time, and the user's singing. The second user is the one singing the audio played by the second terminal. The first terminal adjusts the playback progress of the audio based on the second audio.

[0112] In addition, after the first terminal receives the second sound, it will play the second sound so that the first user can hear the second user's voice, thus achieving the effect of audio chorus.

[0113] In one example, the first terminal and the second terminal agree on the timing of audio transmission. For instance, after each terminal plays audio, they transmit the collected audio data to each other every 1 minute. The first terminal calculates the difference between the time the second audio is sent and the time the first terminal receives the second audio. If the difference is greater than a preset threshold, the playback progress of the audio is adjusted. The preset difference can be any suitable value, for example, 500ms. For example, if the difference is greater than the preset threshold, the playback progress of the audio is slowed down by an amount less than or equal to the difference.

[0114] In another example, frequent adjustments to the audio playback progress can disrupt the playback rhythm and negatively impact the user's singing performance. Therefore, the first terminal sets a preset interval. If the difference between the time the second sound is sent and the time the first terminal receives the second sound is greater than a preset difference, and the interval between the last time the audio playback progress was adjusted and the current local time is greater than the preset interval, then the audio playback progress is adjusted. This avoids frequent adjustments to the audio playback progress and improves the user experience.

[0115] In this embodiment, the first terminal sends the collected sound to the second terminal and receives the collected sound from the second terminal, thereby adjusting the audio playback progress based on the sound from the second terminal, so that the users of the first terminal and the second terminal can sing along to the audio synchronously, improving the user experience.

[0116] Exemplary media

[0117] After introducing the methods of exemplary embodiments of this disclosure, the following references are made. Figure 6 The storage medium of the exemplary embodiments of this disclosure will be described.

[0118] refer to Figure 6As shown, the storage medium 60 stores a program product for implementing the above-described method according to an embodiment of the present disclosure. This program product may be a portable compact disc read-only memory (CD-ROM) and includes computer-executable instructions for causing a computing device to execute the audio playback method provided in this disclosure. However, the program product of this disclosure is not limited thereto.

[0119] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0120] A readable signal medium may include data signals propagated in baseband or as part of a carrier wave, carrying computer-executed instructions. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.

[0121] Computer-executable instructions for performing the operations disclosed herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer-executable instructions can be executed entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).

[0122] Exemplary device

[0123] Having introduced the medium of exemplary embodiments of this disclosure, the following references are made to... Figure 7 The first terminal of the exemplary embodiment of this disclosure will be described. The first terminal is used to implement the method in any of the above-described audio playback method embodiments, and its implementation principle and technical effect are similar.

[0124] refer to Figure 7 , Figure 7 A schematic diagram of the structure of a first terminal provided according to an embodiment of the present disclosure is shown.

[0125] like Figure 7 As shown, the first terminal includes: a first acquisition module 710, used to acquire a first time point in which the first terminal and one or more second terminals sing together; a determination module 720, used to determine a first time offset based on the local time of the first terminal and the standard time of the server; a correction module 730, used to correct the first time point to obtain a target time point based on the current transmission delay between the server and the first terminal, the delay duration of the sound played by the first terminal, and the first time offset; and a first playback module 740, used to receive and play the audio transmitted by the server in response to the current time point reaching the target time point, wherein the audio is the audio of the first terminal and the second terminal singing together.

[0126] In one embodiment, the determining module 720 includes: a first determining unit, configured to determine the current time offset based on a standard time and a local time, and to acquire each historical time offset; the first determining unit is further configured to determine a time offset to be compared among the historical time offsets, wherein the delay corresponding to the time offset to be compared is the minimum delay among the delays corresponding to each historical time offset; the first determining unit is further configured to determine the time offset to be compared as a first time offset in response to the delay corresponding to the time offset to be compared being less than the current transmission delay; the first determining unit is further configured to determine the current time offset as the first time offset in response to the delay corresponding to the time offset to be compared being greater than the current transmission delay.

[0127] In one embodiment, the first terminal 700 further includes: a second playback module, configured to play a preset sound and determine the playback time of the preset sound; a second acquisition module, configured to acquire the acquisition time in response to the acquisition of the played preset sound; the second acquisition module is further configured to acquire the delay duration of the sound played by the first terminal based on the interval between the playback time and the acquisition time.

[0128] In one embodiment, the first terminal 700 further includes: a first sending module, configured to send a time acquisition request to a server; a first receiving module, configured to receive a standard time sent by the server based on the time acquisition request; and a third acquisition module, configured to acquire the current transmission delay between the server and the first terminal based on the interval between the sending time of the time acquisition request and the receiving time of the standard time.

[0129] In one embodiment, the first terminal 700 further includes: a acquisition module, configured to acquire a first sound sung by a first user associated with the first terminal based on the audio after receiving and playing the audio transmitted by the server; a second sending module, configured to send the first sound to a second terminal, the first sound being used to adjust the playback progress of the audio played by the second terminal; and a second receiving module, configured to receive a second sound sent by the second terminal and adjust the playback progress of the audio according to the second sound, wherein the second sound is the sound sung by the second user corresponding to the second terminal based on the audio.

[0130] In one embodiment, the second receiving module includes: a second determining unit, configured to determine the difference between the time point at which the second sound is transmitted and the time point at which the first terminal receives the second sound; and a first adjusting unit, configured to adjust the playback progress of the audio in response to the difference being greater than a preset difference.

[0131] In one embodiment, the first terminal further includes: an adjustment module, configured to adjust the playback progress of the audio in response to a difference greater than a preset difference and an interval between the last time the audio playback progress was adjusted and the current local time being greater than a preset interval.

[0132] In one embodiment, the second sending module includes: an acquisition unit, configured to acquire a first progress of audio played by the first terminal, and send the first sound and the first progress to the second terminal; the second receiving module includes: a receiving unit, configured to receive a second sound sent by the second terminal and a second progress of audio played by the second terminal; and a second adjustment unit, configured to adjust the playback progress of the audio in response to a progress difference between the first progress and the second progress exceeding a preset threshold.

[0133] In one embodiment, the playback module 740 includes a playback unit for playing accompaniment to audio.

[0134] In one embodiment, the correction module 730 includes: a correction unit, configured to correct the first time offset based on the current transmission delay between the server and the first terminal and the delay duration of the sound played by the first terminal, to obtain a target time offset; and a correction unit, configured to correct the first time point based on the target time offset to obtain a target time point.

[0135] Exemplary computing device

[0136] Having introduced the method, medium, and first terminal of exemplary embodiments of this disclosure, the following references... Figure 8 A computing device according to an exemplary embodiment of the present disclosure will be described.

[0137] Figure 8 The computing device 80 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein. Figure 8 As shown, the computing device 80 is presented in the form of a general-purpose computing device. The components of the computing device 80 may include, but are not limited to: at least one processing unit 801, at least one storage unit 802, and a bus 803 connecting different system components (including the processing unit 801 and the storage unit 802). The at least one storage unit 802 stores computer-executable instructions; the at least one processing unit 801 includes a processor that executes the computer-executable instructions to implement the methods described above.

[0138] The 803 bus includes a data bus, a control bus, and an address bus.

[0139] Storage unit 802 may include readable media in the form of volatile memory, such as random access memory (RAM) 8021 and / or cache memory 8022, and may further include readable media in the form of non-volatile memory, such as read-only memory (ROM) 8023.

[0140] Storage unit 802 may also include a program / utility 8025 having a set (at least one) program module 8024, such program module 8024 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0141] The computing device 80 can also communicate with one or more external devices 804 (e.g., keyboard, pointing device, etc.). This communication can be performed via input / output (I / O) interface 805. Furthermore, the computing device 80 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 806. Figure 8 As shown, network adapter 806 communicates with other modules of computing device 80 via bus 803. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 80, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0142] It should be noted that although several units / modules or sub-units / modules of the first terminal have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0143] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0144] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. An audio playback method, characterized in that, Applied to the first terminal, including: Obtain the first time point in which the first terminal sings in harmony with one or more second terminals; The first time offset is determined based on the local time of the first terminal and the standard time of the server; The target time point is obtained by correcting the first time point based on the current transmission delay between the server and the first terminal, the delay duration of the sound played by the first terminal, and the first time offset. In response to the current time point reaching the target time point, the system receives and plays the audio transmitted by the server, wherein the audio is the audio of the first terminal and the second terminal singing together; The step of determining the first time offset based on the local time of the first terminal and the standard time of the server includes: The current time offset is determined based on the standard time and the local time, and the time offsets of each historical time are obtained. A time offset to be compared is determined from each of the historical time offsets, wherein the delay corresponding to the time offset to be compared is the minimum delay among the transmission delays corresponding to each of the historical time offsets; the first terminal records the time offset between the local time and the standard time when playing audio in the past, as well as the transmission delay between the first terminal and the server, and the first terminal stores multiple historical time offsets and the transmission delays corresponding to the historical time offsets; In response to the fact that the transmission delay corresponding to the time offset to be compared is less than the current transmission delay, the time offset to be compared is determined as the first time offset; In response to the transmission delay corresponding to the time offset to be compared being greater than the current transmission delay, the current time offset is determined as the first time offset.

2. The audio playback method according to claim 1, characterized in that, The audio playback method further includes: Play a preset sound and determine the playback time of the preset sound; In response to the captured preset sound being played, the capture time point is obtained; The delay duration of the sound played by the first terminal is obtained based on the interval between the playback time point and the acquisition time point.

3. The audio playback method according to claim 1, characterized in that, The audio playback method further includes: Send a time retrieval request to the server; Receive the standard time sent by the server based on the time acquisition request; The current transmission delay between the server and the first terminal is obtained based on the interval between the sending time of the time acquisition request and the receiving time of the standard time.

4. The audio playback method according to claim 1, characterized in that, After receiving and playing the audio transmitted by the server, the method further includes: Collect the first voice sung by the first user associated with the first terminal based on the audio; The first sound is sent to the second terminal, and the first sound is used to adjust the progress of the audio played by the second terminal; The system receives a second sound sent by the second terminal and adjusts the playback progress of the audio based on the second sound, wherein the second sound is the sound sung by the second user corresponding to the second terminal based on the audio.

5. The audio playback method according to claim 4, characterized in that, The step of adjusting the playback progress of the audio according to the second sound includes: Determine the difference between the time point when the second sound was sent and the time point when the first terminal received the second sound; In response to the difference being greater than a preset difference, the playback progress of the audio is adjusted.

6. The audio playback method according to claim 5, characterized in that, The audio playback method further includes: In response to the difference being greater than a preset difference, and the interval between the last time the playback progress of the audio was adjusted and the current local time being greater than a preset interval, the playback progress of the audio is adjusted.

7. The audio playback method according to claim 4, characterized in that, Sending the first sound to the second terminal includes: Obtain the first progress of the audio played on the first terminal, and send the first sound and the first progress to the second terminal; The step of receiving the second sound sent by the second terminal and adjusting the playback progress of the audio according to the second sound includes: Receive the second sound sent by the second terminal and the second progress of the second terminal playing the audio; In response to the progress difference between the first progress and the second progress being greater than a preset threshold, the playback progress of the audio is adjusted.

8. The audio playback method according to any one of claims 1-7, characterized in that, Playing the audio transmitted by the server includes: Play the accompaniment to the audio.

9. The audio playback method according to any one of claims 1-7, characterized in that, The step of correcting the first time point based on the current transmission delay between the server and the first terminal, the delay duration of sound playback on the first terminal, and the first time offset to obtain the target time point includes: Based on the current transmission delay between the server and the first terminal and the delay duration of the sound playback on the first terminal, the first time offset is corrected to obtain the target time offset. The first time point is corrected based on the target time offset to obtain the target time point.

10. A terminal device, characterized in that, include: The first acquisition module is used to acquire the first time point when the first terminal and one or more second terminals sing together. The determining module is used to determine a first time offset based on the local time of the first terminal and the standard time of the server; The correction module is used to correct the first time point based on the current transmission delay between the server and the first terminal, the delay duration of the sound played by the first terminal, and the first time offset to obtain the target time point. The first playback module is used to receive and play the audio transmitted by the server in response to the current time point reaching the target time point, wherein the audio is the audio of the first terminal and the second terminal singing together; The determining module includes: The first determining unit is used to determine the current time offset based on the standard time and the local time, and to obtain the time offset of each historical time. The first determining unit is further configured to determine a time offset to be compared among the various historical time offsets, wherein the delay corresponding to the time offset to be compared is the minimum delay among the transmission delays corresponding to the various historical time offsets; the first terminal records the time offset between the local time and the standard time when playing audio in the past, as well as the transmission delay between the first terminal and the server, and the first terminal stores multiple historical time offsets and the transmission delays corresponding to the historical time offsets; The first determining unit is further configured to determine the time offset to be compared as a first time offset in response to the fact that the transmission delay corresponding to the time offset to be compared is less than the current transmission delay; The first determining unit is further configured to determine the current time offset as a first time offset in response to the transmission delay corresponding to the time offset to be compared being greater than the current transmission delay.

11. The terminal device according to claim 10, characterized in that, The first terminal also includes: The second playback module is used to play a preset sound and determine the playback time of the preset sound; The second acquisition module is used to acquire the acquisition time point in response to the acquisition of the preset sound being played. The second acquisition module is further configured to acquire the delay duration of the sound played by the first terminal based on the interval between the playback time point and the acquisition time point.

12. The terminal device according to claim 10, characterized in that, The first terminal also includes: The first sending module is used to send a time acquisition request to the server; The first receiving module is used to receive the standard time sent by the server based on the time acquisition request; The third acquisition module is used to acquire the current transmission delay between the server and the first terminal based on the interval between the sending time of the time acquisition request and the receiving time of the standard time.

13. The terminal device according to claim 10, characterized in that, The first terminal also includes: The acquisition module is used to acquire the first voice sung by the first user associated with the first terminal based on the audio after receiving and playing the audio transmitted by the server; The second sending module is used to send the first sound to the second terminal, and the first sound is used to adjust the progress of the audio played by the second terminal; The second receiving module is used to receive the second sound sent by the second terminal and adjust the playback progress of the audio according to the second sound, wherein the second sound is the sound sung by the second user corresponding to the second terminal based on the audio.

14. The terminal device according to claim 13, characterized in that, The second receiving module includes: The second determining unit is used to determine the difference between the time point when the second sound was sent and the time point when the first terminal received the second sound; The first adjustment unit is used to adjust the playback progress of the audio in response to the difference being greater than a preset difference.

15. The terminal device according to claim 14, characterized in that, The first terminal also includes: An adjustment module is used to adjust the playback progress of the audio in response to the difference being greater than a preset difference and the interval between the last time the playback progress of the audio was adjusted and the current local time being greater than a preset interval.

16. The terminal device according to claim 13, characterized in that, The second transmitting module includes: The acquisition unit is configured to acquire the first progress of the audio played by the first terminal, and send the first sound and the first progress to the second terminal. The second receiving module includes: The receiving unit is configured to receive the second sound sent by the second terminal and the second progress of the second terminal playing the audio; The second adjustment unit is used to adjust the playback progress of the audio in response to the progress difference between the first progress and the second progress being greater than a preset threshold.

17. The terminal device according to any one of claims 10-16, characterized in that, The playback module includes: The playback unit is used to play the accompaniment to the audio.

18. The terminal device according to any one of claims 10-16, characterized in that, The correction module includes: The correction unit is used to correct the first time offset based on the current transmission delay between the server and the first terminal and the delay duration of the sound played by the first terminal, so as to obtain the target time offset. The correction unit is used to correct the first time point according to the target time offset to obtain the target time point.

19. A medium, characterized in that, include: Computer execution instructions, when executed by a processor, are used to implement the method as described in any one of claims 1 to 9.

20. A computing device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Synchronous cooperation method for multiple intelligent electronic devices and multimedia playing system

    CN110392292A

  • Chorus singing method and device, electronic equipment and storage medium

    CN111028818A

  • Audio processing method and audio processing apparatus

    WO2022160669A1

  • KR20200117435A