Chorus method and device and storage medium

By receiving and processing multimedia streaming data in the live broadcast room and adding accurate timestamps, the problem of uneven delays in multiple choruses is solved, and the chorus effect and user experience are improved.

CN120238665APending Publication Date: 2025-07-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311866690.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the live broadcast room multiplayer chorus scene, due to network and device factors, the multimedia streaming data is not delayed accurately, and the prior art cannot accurately add synchronization timestamps, resulting in poor chorus effect.

Method used

By receiving the lead multimedia stream data, decoding and playing in turn, adding a time stamp to each frame of data, and collecting and processing data during the sergeant's singing process, the timestamp of the sergeant's multimedia data is determined based on the preset delay time to ensure the accuracy of the time stamp.

Benefits of technology

It improves the accuracy of multimedia streaming data alignment, and improves the effect and connection experience of multi-person chorus in the live broadcast room.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238665A_ABST
    Figure CN120238665A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a chorus method and device and a storage medium. The method comprises the following steps: receiving chorus multimedia stream data in a chat room; decoding and playing each frame of the singing multimedia data of the singing multimedia stream data in sequence; collecting and processing the current frame of auxiliary singing multimedia data in the auxiliary singing following process; and determining a second timestamp corresponding to the current frame of the auxiliary singing multimedia data according to the first timestamp in the currently decoded frame of the collar multimedia data and a preset delay time, and adding the second timestamp to the current frame of the auxiliary singing multimedia data, and transmitting the sub-singing multimedia stream data to other clients except the sub-singing role client participating in live broadcast microphone connection through the sub-singing multimedia stream data. According to the method and the device, the second timestamp which is as close as possible to the real timestamp of the current frame of sub-singing multimedia data can be determined, so that the accuracy of adding the timestamps is improved, the terrace multimedia stream data and the sub-singing multimedia stream data can be better aligned according to the timestamps, and the chorus effect is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical fields of computers and network communications, and in particular, to a chorus method, device, and storage medium. Background Art

[0002] Currently, in video live rooms or voice live rooms, etc., the host can conduct real-time live connection and interaction with guests by establishing a connection session, so that the audience in the live room can watch the connection and interaction content.

[0003] In one scenario, the host or guest in the live room can perform a multi-person chorus of a song. Due to factors such as network and device performance, there may be a certain delay in receiving the multimedia stream data of other parties by any party, and the delays may also be uneven. In order to facilitate the non-chorus role client to play the multimedia stream data of each chorus role client after alignment, the leading chorus role client can add synchronization timestamps to its leading chorus multimedia stream data according to the singing progress, and the backup chorus role client can also add the synchronization timestamps in the leading chorus multimedia stream data to the backup chorus multimedia stream data, so that the non-chorus role client can align the multimedia stream data of each chorus role client based on the synchronization timestamps.

[0004] However, when the backup chorus role client adds the synchronization timestamps in the leading chorus multimedia stream data to the backup chorus multimedia stream data, it usually cannot accurately add the synchronization timestamps, resulting in the offset of the synchronization timestamps and inability to align with the leading chorus multimedia stream data, resulting in a poor chorus effect and inability to meet the needs of multi-person chorus in the chat room. Summary of the Invention

[0005] Embodiments of the present disclosure provide a chorus method, device, and storage medium to improve the accuracy of adding timestamps to multimedia stream data in the chorus scenario of the live room and improve the chorus effect.

[0006] In a first aspect, embodiments of the present disclosure provide a chorus method applied to a backup chorus role client participating in a live connection. The method includes:

[0007] Receiving leading chorus multimedia stream data; and sequentially decoding and playing each frame of leading chorus multimedia data of the leading chorus multimedia stream data;

[0008] Collecting and processing the current frame of backup chorus multimedia data during the backup chorus's following singing;

[0009] Determining a second timestamp corresponding to the current frame of backup chorus multimedia data according to a first timestamp in a frame of leading chorus multimedia data currently decoded and a preset delay time, and adding the second timestamp to the current frame of backup chorus multimedia data;

[0010] Transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection and chorus except the backup singer role client.

[0011] In a second aspect, an embodiment of the present disclosure provides a chorus device, including:

[0012] A receiving unit, configured to receive lead singer multimedia stream data;

[0013] A processing unit, configured to decode and play each frame of lead singer multimedia data in the lead singer multimedia stream data in sequence;

[0014] An acquisition unit, configured to acquire the current frame of backup singer multimedia data during the process of the backup singer following the singing;

[0015] The processing unit is further configured to process the current frame of backup singer multimedia data; determine a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time, and add the second timestamp to the current frame of backup singer multimedia data; perform encoding processing on the current frame of backup singer multimedia data;

[0016] A sending unit, configured to transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection and chorus except the backup singer role client.

[0017] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;

[0018] The memory stores computer-executable instructions;

[0019] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the chorus method described in the first aspect above and various possible designs of the first aspect.

[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the chorus method described in the first aspect above and various possible designs of the first aspect is implemented.

[0021] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer-executable instructions, and when a processor executes the computer-executable instructions, the chorus method described in the first aspect above and various possible designs of the first aspect is implemented.

[0022] The chorus method, device, and storage medium provided by the embodiments of the present disclosure receive the lead singer multimedia stream data of the lead singer role client participating in the live connection. Decode and play each frame of the lead singer multimedia data in the lead singer multimedia stream data in sequence. Collect and process the current frame of backup singer multimedia data during the process of the backup singer following the singing. Determine the second time stamp corresponding to the current frame of backup singer multimedia data according to the first time stamp in the currently decoded frame of lead singer multimedia data and a preset delay time, and add the second time stamp to the current frame of backup singer multimedia data. Transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection except for the backup singer role client. When adding a time stamp to the current frame of backup singer multimedia data, based on the first time stamp in the currently decoded frame of lead singer multimedia data and in combination with the preset delay time, a second time stamp that is as close as possible to the real time stamp of the current frame of backup singer multimedia data can be determined and added to the current frame of backup singer multimedia data, so as to improve the accuracy of time stamp addition, enable the lead singer multimedia stream data and the backup singer multimedia stream data to be better aligned according to the time stamps, and ensure the chorus effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the following described drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1a It is a scenario example diagram of the chorus method provided by an embodiment of the present disclosure;

[0025] Figure 1b It is a scenario example diagram of the chorus method provided by an embodiment of the present disclosure;

[0026] Figure 1c It is a principle example diagram of the time stamp offset in a related technology;

[0027] Figure 2 It is a schematic flowchart of the chorus method provided by an embodiment of the present disclosure;

[0028] Figure 3 It is a structural block diagram of the chorus device provided by an embodiment of the present disclosure;

[0029] Figure 4 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without making creative efforts shall fall within the scope of protection of the present disclosure.

[0031] In one scenario, as Figure 1a shown, the host or guest in the live broadcast room can perform a multi-person chorus of a song. Among them, the client participating in the live connection can be divided into a chorus role client and a non-chorus role client from the perspective of the role. The chorus role client can include a lead singer (or also called the main singer) role client and a backup singer (or also called the following singer) role client. As Figure 1b shown, due to factors such as network and device performance, there may be a certain delay in any party receiving the multimedia stream data of other parties, and the delays may also be uneven. For example, the lead singer multimedia stream data is sent at time T1 and can reach the backup singer role client and the non-chorus role client at time T2, while the backup singer multimedia stream data after the backup singer follows the lead singer reaches the non-chorus role client at time T3. To facilitate the non-chorus role client to play the multimedia stream data of each chorus role client after alignment, the lead singer role client can add a synchronization timestamp to its lead singer multimedia stream data according to the singing progress, and the backup singer role client can also add the synchronization timestamp in the lead singer multimedia stream data to the backup singer multimedia stream data. In this way, the non-chorus role client can align the multimedia stream data of each chorus role client based on the synchronization timestamp, and the chorus effect can be presented after alignment and playback.

[0032] However, when the backup singer role client adds the synchronization timestamp in the lead singer multimedia stream data to the backup singer multimedia stream data, it is usually unable to accurately add the synchronization timestamp. When adding the synchronization timestamp, the backup singer role client adds the synchronization timestamp in a frame of decoded lead singer multimedia data to the currently captured frame of backup singer multimedia data, that is, it is assumed that the currently captured frame of backup singer multimedia data corresponds to a frame of decoded lead singer multimedia data. However, there are certain delays from the decoding to the playback of the lead singer multimedia data and then to the capture of the backup singer multimedia data, which results in the fact that the currently captured frame of backup singer multimedia data does not actually correspond to a frame of decoded lead singer multimedia data, as Figure 1cAs shown, the lead singer multimedia data of a frame is decoded at time T1 and played at time T2. The corresponding backup singer multimedia data of a frame is collected at time T3. However, the lead singer multimedia data of the frame currently being decoded at time T3 is no longer the one at time T1. Adding the synchronization timestamp of the lead singer multimedia data of the frame currently being decoded at time T3 to the currently collected backup singer multimedia data of the current frame will cause the synchronization timestamp of the current frame of backup singer multimedia data to shift, making it impossible to align with the lead singer multimedia stream data of a frame at the same singing progress. As a result, it is impossible to align the lead singer multimedia data and the backup singer multimedia data based on the synchronization timestamp later, and a good chorus effect cannot be presented, failing to meet the requirements of multi-person chorus in the live broadcast room.

[0033] To solve the above technical problems, an embodiment of the present disclosure provides a chorus method. The backup singer role client participating in the live connection receives the lead singer multimedia stream data, where each frame of the lead singer multimedia data in the lead singer multimedia stream data includes a first timestamp corresponding to the singing progress; each frame of the lead singer multimedia data in the lead singer multimedia stream data is decoded and played in sequence; the current frame of backup singer multimedia data is collected and processed during the backup singer's following of the singing; according to the first timestamp in the currently decoded frame of lead singer multimedia data and a preset delay time, a second timestamp corresponding to the current frame of backup singer multimedia data is determined, and the second timestamp is added to the current frame of backup singer multimedia data; the current frame of backup singer multimedia data is transmitted as backup singer multimedia stream data to other clients participating in the live connection except the backup singer role client.

[0034] Furthermore, after receiving the multimedia stream data of other clients participating in the live connection except the backup singer role client, the non-chorus role client participating in the live connection can cache and align the multimedia stream data of the chorus role client, ensuring that a chorus effect can be presented on the non-chorus role client participating in the connection, meeting the requirements of multi-person chorus in the live broadcast room and improving the connection experience in the live broadcast room.

[0035] The chorus method provided in the above embodiment is applied to a scenario as Figure 1a shown. The live broadcast room can be a video live broadcast room or a voice live broadcast room. From the perspective of roles, the live broadcast room can include clients participating in the connection and audience clients not participating in the connection. Among the clients participating in the connection, there are chorus role clients and non-chorus role clients. The chorus role client is the client participating in the chorus of the connection, which can include a lead singer (or main singer) role client and a backup singer role client. The backup singer role client participating in the live connection can respectively execute the corresponding chorus method described above.

[0036] It should be noted that the user information and data involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0037] The following will introduce the chorus method of the present disclosure in detail with specific embodiments.

[0038] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the chorus method provided by an embodiment of the present disclosure. In the chorus scenario of the live broadcast room, the live broadcast room can be a video live broadcast room or a voice live broadcast room. From the perspective of roles, the live broadcast room can include a client participating in the co-hosting and an audience client not participating in the co-hosting. Among them, the client participating in the co-hosting includes a chorus role client and a non-chorus role client. The chorus role client is the client participating in the co-hosting chorus, and can include a lead singer (or main singer) role client and a backup singer role client. From another perspective, the live broadcast room can include a host client and a guest client participating in the co-hosting, as well as an audience client not participating in the co-hosting. Among them, the host client and the guest client can be any of the above roles (where there is only one lead singer role client. If the host client or a certain guest client is the lead singer role client, the other clients participating in the co-hosting cannot be the lead singer role client).

[0039] The chorus method of this embodiment can be applied to the backup singer role client participating in the live co-hosting. The chorus method includes:

[0040] S201. Receive the lead singer multimedia stream data.

[0041] In this embodiment, each client participating in the live co-hosting can receive the multimedia stream data of other clients participating in the live co-hosting except itself. Among them, in the video live broadcast room, the multimedia stream data can be audio-video stream data; in the voice live broadcast room, the media stream data can be audio stream data.

[0042] Among them, since the backup singer role client needs to sing along with the multimedia stream data of the lead singer role client, there is a certain lag in the multimedia stream data of the backup singer role client relative to the multimedia stream data of the lead singer role client. For example, the multimedia stream data of the lead singer role client has reached the 10th second of the song, while the multimedia stream data of the backup singer role client 1 may only reach the 8th second of the song, and the singing progress of the multimedia stream data of different backup singer role clients may also be different. The multimedia stream data of another backup singer role client 2 may reach the 9th second of the song. In order to present a chorus effect on the non-chorus role client, the lead singer role client can add a timestamp to each frame of the lead singer multimedia data in the lead singer multimedia stream data according to the singing progress. Other backup singer role clients add timestamps to the local backup singer multimedia data collected according to the timestamps in each frame of the lead singer multimedia data in the lead singer multimedia stream data. In this way, the non-chorus role client can receive the multimedia stream data of the lead singer role client and the multimedia stream data of the backup singer role client, and can be played after being aligned based on the timestamp through caching, and the chorus effect can be presented.

[0043] When the lead singer role client sequentially collects each frame of the lead singer multimedia data, for the convenience of description, the added timestamp can be recorded as the first timestamp. Since the singing progress is advancing, the addition of the first timestamp in each frame of the lead singer multimedia data is also advancing. After a series of processes, each frame of the lead singer multimedia data is transmitted as the lead singer multimedia stream data to each other client except the lead singer role client participating in the live connection, which may include the non-chorus role client and the backup singer role client participating in the connection.

[0044] S202. Decode and play each frame of the lead singer multimedia data in the lead singer multimedia stream data in sequence.

[0045] In this embodiment, after the backup singer role client receives the lead singer multimedia stream data, it can sequentially decode and play each frame of the lead singer multimedia data in the lead singer multimedia stream data, and the first timestamp in each frame of the lead singer multimedia data can be obtained during decoding. The specific decoding process and playing process can adopt any known method and are not limited here.

[0046] S203. Collect and process the current frame of the backup singer multimedia data during the process of the backup singer following the singing.

[0047] In this embodiment, when playing each frame of the lead singer multimedia data of the lead singer multimedia stream, the user on the side of the backup singer client, that is, the backup singer, can sing along with the audio of the played lead singer multimedia data. The backup singer client can collect the backup singer multimedia data frame by frame during the process of the backup singer singing along. For the current frame of the backup singer multimedia data, other processing can be performed after collection. The specific processing process can adopt any known method and is not limited here.

[0048] S204. Determine the second timestamp corresponding to the current frame of the backup singer multimedia data according to the first timestamp in a frame of the lead singer multimedia data decoded currently and a preset delay time, and add the second timestamp to the current frame of the backup singer multimedia data.

[0049] In this embodiment, when it is necessary to add a timestamp to the current frame of the backup singer multimedia data, the first timestamp in a frame of the lead singer multimedia data decoded currently can be obtained. However, since there are certain delays from the decoding of the lead singer multimedia data to the playback and then to the collection of the backup singer multimedia data, the current frame of the backup singer multimedia data does not actually correspond to the frame of the lead singer multimedia data decoded currently. That is, the real timestamp of the current frame of the backup singer multimedia data should be earlier than the first timestamp of the frame of the lead singer multimedia data decoded currently. Therefore, a preset delay time can be obtained. Based on the first timestamp of the frame of the lead singer multimedia data decoded currently and in combination with the preset delay time, a second timestamp is determined, so that the second timestamp is as close as possible to the real timestamp of the current frame of the backup singer multimedia data. In this way, adding the second timestamp to the current frame of the backup singer multimedia data can improve the accuracy of adding timestamps to the backup singer multimedia data.

[0050] The preset delay time can be the delay time between the decoding moment of any frame of the lead singer multimedia data and the moment when a frame of the backup singer multimedia data at the same singing progress is completed (the processing here refers to the processing process before adding the timestamp). For example Figure 1c the deltaT in, that is, the delay time from T3 to T1, can be specifically obtained by measurement, can also be provided by the server, can also be determined according to experience, or determined by any other possible means. It is not limited in this embodiment.

[0051] And determining the second timestamp corresponding to the current frame of the backup singer multimedia data according to the first timestamp in a frame of the lead singer multimedia data decoded currently and the preset delay time can specifically be to determine the difference between the first timestamp in a frame of the lead singer multimedia data decoded currently and the preset delay time, and determine the difference as the second timestamp corresponding to the current frame of the backup singer multimedia data. That is, the second timestamp should actually be earlier than the first timestamp in a frame of the lead singer multimedia data decoded currently, and the required advance amount is the preset delay time.

[0052] S205. Transmit the current frame of backup vocal multimedia data as backup vocal multimedia stream data to other clients participating in the live connection in the live broadcast room except for the backup vocal role client.

[0053] In this embodiment, each frame of backup vocal multimedia data with a second timestamp added is sequentially processed and then transmitted as backup vocal multimedia stream data to other clients participating in the live connection in the live broadcast room except for the backup vocal role client, which may include non-chorus role clients, lead vocal role clients, and other backup vocal role clients participating in the connection. The processing process and transmission process involved can adopt any known method and are not limited in this embodiment.

[0054] The chorus method of this embodiment includes receiving lead vocal multimedia stream data; sequentially decoding and playing each frame of lead vocal multimedia data in the lead vocal multimedia stream data; collecting and processing the current frame of backup vocal multimedia data during the backup vocal's follow-up singing; determining the second timestamp corresponding to the current frame of backup vocal multimedia data according to the first timestamp in a currently decoded frame of lead vocal multimedia data and a preset delay time, and adding the second timestamp to the current frame of backup vocal multimedia data; transmitting the current frame of backup vocal multimedia data as backup vocal multimedia stream data to other clients participating in the live connection except for the backup vocal role client. When adding a timestamp to the current frame of backup vocal multimedia data, based on the first timestamp in a currently decoded frame of lead vocal multimedia data and in combination with the preset delay time, a second timestamp that is as close as possible to the true timestamp of the current frame of backup vocal multimedia data can be determined and added to the current frame of backup vocal multimedia data to improve the accuracy of timestamp addition, so that the lead vocal multimedia stream data and the backup vocal multimedia stream data can be better aligned according to the timestamps, ensuring the chorus effect.

[0055] Based on any of the above embodiments, since a preset delay time is required, it is also possible to determine the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when a frame of backup vocal multimedia data at the same singing progress is completed (the processing here refers to the processing process before adding the timestamp), and determine this delay time as the preset delay time.

[0056] In this embodiment, considering that the delay is caused by certain delays during the processes from the decoding of the lead singer's multimedia data to playback and then to the acquisition of the backup singer's multimedia data, it is possible to determine the delay time between the decoding moment of any frame of the lead singer's multimedia data and the moment when a frame of the backup singer's multimedia data with the same singing progress is completely acquired and processed (including any processing process before adding a timestamp to the backup singer's multimedia data). That is, the delay time between the moment when the first timestamp is obtained from the decoding of any frame of the lead singer's multimedia data and the moment when a timestamp is added to a frame of the backup singer's multimedia data with the same singing progress, which is the preset delay time.

[0057] Based on the above embodiment, determining the delay time between the decoding moment of any frame of the lead singer's multimedia data and the moment when a frame of the backup singer's multimedia data with the same singing progress is completely processed, and determining this delay time as the preset delay time may specifically include:

[0058] Determine a first delay time and a second delay time, where the first delay time is the delay time between the playback moment of any frame of the lead singer's multimedia data and the acquisition moment of a frame of the backup singer's multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer's multimedia data and the processing process of the backup singer's multimedia data; determine the sum of the first delay time and the second delay time as the preset delay time.

[0059] In this embodiment, the delay time can be divided into two parts, namely the first delay time and the second delay time. The first delay time is the delay time between the playback moment of any frame of the lead singer's multimedia data and the acquisition moment of a frame of the backup singer's multimedia data with the same singing progress, that is, the delay time from playback to acquisition, which can also be called the acquisition-playback delay; the second delay time is the remaining part of the delay time, mainly caused by the processing process of the multimedia data, including the delay time of the processing process of the lead singer's multimedia data (the processing process from after decoding to before playback, such as rendering, etc.) and the delay time of the processing process of the backup singer's multimedia data (the processing process from after acquisition to before adding a timestamp, such as encoding, etc.). Combining the first delay time and the second delay time in this way is the preset delay time.

[0060] The determination of the second delay time will be described first below. Since the second delay time is the delay caused by the processing of multimedia data, the delay time of each sub - processing process in the processing process of the lead - singer multimedia data and the processing process of the backup - singer multimedia data can be determined, and then the delay times of each sub - processing process are added together, and the sum is determined as the second delay time. Of course, if there are some consecutive sub - processing processes in the processing process, the consecutive sub - processing processes can also be regarded as an overall sub - processing process, and only one delay time for this overall sub - processing process needs to be obtained.

[0061] Among them, for any sub - processing process, the time interval between the input data time and the corresponding output data time can be obtained as the delay time of this sub - processing process. In actual implementation, it is not necessary to detect the delay time of each sub - processing process and add up the delay times of each sub - processing process. It is possible to only detect the delay time of each processing process in the processing link of one frame (or the average of multiple frames) of lead - singer multimedia data and the delay time of each processing process in the processing link of one frame (or the average of multiple frames) of backup - singer multimedia data when receiving the lead - singer multimedia stream data at the beginning, and then add them up to obtain a second delay time, and this second delay time is not updated in the subsequent process. It is also possible to periodically detect the delay time of each sub - processing process in the processing process of lead - singer multimedia data and the processing process of backup - singer multimedia data every preset time interval, so as to continuously update the second delay time to adapt to the change of the processing ability of the backup - singer role client and ensure the accuracy of timestamp addition.

[0062] Of course, the determination of the second delay time is not limited to the above method, and any other feasible method is also acceptable, which is not restricted here.

[0063] For the determination of the first delay time, it can be divided into the following different situations, specifically as follows:

[0064] Situation 1:

[0065] According to the first preset mapping relationship, determine whether the preset first delay time corresponding to the backup - singer role client can be obtained, where the first preset mapping relationship is the mapping relationship between different device information and the preset first delay time;

[0066] If it is determined that the preset first delay time corresponding to the backup - singer role client can be obtained, then the preset first delay time is determined as the first delay time.

[0067] In Case 1, considering that the processing performance of the same device during multimedia data playback and acquisition is the same or similar, the first delay time of various devices can be obtained in advance as the preset first delay time, and a mapping relationship between device information (such as device type, device model) and the preset first delay time, that is, the first preset mapping relationship, can be constructed. The acquisition method can be that developers play music on different devices and sing along, collect multimedia data during the singing-along process, and then detect the time delay between music playback and multimedia data acquisition during singing-along. Of course, other methods can also be used to determine the first delay time of different devices.

[0068] The first preset mapping relationship can be configured on the server side and can be updated regularly by developers (for example, when some new devices are added). The secondary singer role client can obtain the first preset mapping relationship from the server side and query the corresponding preset first delay time from the first preset mapping relationship according to the device information of its own side as the first delay time; or the secondary singer role client can also send a query request to the server side. The query request includes the device information of its own side, and the server side queries the preset first delay time corresponding to the secondary singer role client from the first preset mapping relationship and sends it to the secondary singer role client as the first delay time.

[0069] Case 2:

[0070] If it is determined that the preset first delay time corresponding to the secondary singer role client cannot be obtained and the playback mode of the lead singer multimedia data is the external playback mode, then the first delay time is determined based on the acoustic echo cancellation algorithm.

[0071] In Case 2, since it is impossible to cover all types of devices in the first preset mapping relationship, if the secondary singer role client cannot query the preset first delay time corresponding to the device information of the secondary singer role client from the first preset mapping relationship, other means can be used to obtain the first delay time. Specifically, if the secondary singer role client plays the lead singer multimedia data in the external playback mode (such as using a speaker, etc.), since the sound of the lead singer multimedia data played externally by the secondary singer role client will be collected again by the sound collection device of the secondary singer role client to form an acoustic echo, the acoustic echo cancellation algorithm (AEC) can be used in this embodiment to determine the first delay time. The acoustic echo cancellation algorithm (AEC) is to compare the signal collected by the microphone with the signal output by the speaker to estimate the echo signal, and then cancel the echo signal. The determination of the delay time of the echo signal is involved, that is, the delay time between the signal collected by the microphone and the signal output by the speaker. In this embodiment, the delay time can be directly used as the required first delay time.

[0072] It should be noted that when the backing vocal role client collects backing vocal multimedia data, an acoustic echo cancellation algorithm (AEC) is usually run for echo cancellation. Therefore, when determining the first delay time, it is not necessary to specifically run the acoustic echo cancellation algorithm (AEC), but the delay time can be directly extracted from the echo cancellation process as the first delay time.

[0073] Case 3:

[0074] If it is determined that the preset first delay time corresponding to the backing vocal role client cannot be obtained, and the playback mode of the lead vocal multimedia data is a non-external playback mode, then a query is made in the second preset mapping relationship according to the multimedia data acquisition method used by the backing vocal role client, where the second preset mapping relationship is the mapping relationship between different multimedia data acquisition methods and the preset first delay time;

[0075] The preset first delay time corresponding to the multimedia data acquisition method used by the backing vocal role client in the second preset mapping relationship is determined as the first delay time.

[0076] In Case 3, since it is impossible to cover all types of devices in the first preset mapping relationship, the backing vocal role client may not be able to query the preset first delay time corresponding to the device information of the backing vocal role client from the first preset mapping relationship. If the backing vocal role client does not play the lead vocal multimedia data in the external playback mode but uses headphones to play the lead vocal multimedia data, the method in Case 2 cannot be used to determine the first delay time either. In this embodiment, a second preset mapping relationship can be pre-configured, and the preset first delay time corresponding to different acquisition methods of multimedia data is configured in the second preset mapping relationship. Different acquisition methods of multimedia data mainly consider different acquisition methods of audio. For example, audio is acquired using OpenSL at the CS layer, or AudioRecord at the Java layer. The delay of OpenSL is relatively low, and the corresponding preset first delay time is relatively small. The first delay time of the backing vocal role client can be determined based on the second preset mapping relationship.

[0077] The second preset mapping relationship can be configured on the server side. The backing vocal role client can obtain the second preset mapping relationship from the server side and query the corresponding preset first delay time from the second preset mapping relationship according to the multimedia data acquisition method of its own side as the first delay time; or the backing vocal role client can also send a query request to the server side. The query request includes the multimedia data acquisition method of its own side, and the server side queries the preset first delay time corresponding to the backing vocal role client from the second preset mapping relationship and sends it to the backing vocal role client as the first delay time.

[0078] Based on the above embodiments, after determining the first delay time and the second delay time, the first delay time and the second delay time can be added together to obtain the above-mentioned preset delay time, that is, the delay time from the moment when the first timestamp is obtained by decoding any frame of the lead singer multimedia data to the moment when the timestamp is added to a frame of the backup singer multimedia data with the same singing progress.

[0079] On this basis, based on the first timestamp in a frame of the lead singer multimedia data being currently decoded, subtract the preset delay time, that is, determine the second timestamp corresponding to the current frame of the backup singer multimedia data, and add the second timestamp to the current frame of the backup singer multimedia data.

[0080] After a series of post-processing on the current frame of the backup singer multimedia data, it is transmitted to other clients participating in the live co-hosting connection in the form of stream data (that is, the backup singer multimedia stream data), except for the backup singer role client.

[0081] For non-chorus role clients participating in the live co-hosting connection, the multimedia stream data of each chorus role client participating in the co-hosting can be cached, including the multimedia stream data of the lead singer role client and the multimedia stream data of each backup singer role client. Then, according to the timestamps included in the multimedia stream data of each chorus role client, the multimedia stream data of each chorus role client is aligned and then played.

[0082] For other chorus role clients participating in the live co-hosting connection, including the lead singer role client and other backup singer role clients, the multimedia stream data of the backup singer role client can be played muted, that is, only the video is played, and the sound is not played, to avoid affecting the singing of the local user.

[0083] In addition, since the host client is also responsible for pushing the live stream to the audience client, regardless of which type of role client the host client belongs to, it is necessary to align the multimedia stream data of each chorus role client, and combine the aligned multimedia stream data of each chorus role client, the multimedia stream data of other non-chorus role clients participating in the co-hosting connection, and the local multimedia stream data (of course, when the host client belongs to different role clients, the local multimedia stream data may also be included in the previous several types of multimedia stream data), and transmit the combined multimedia stream data to the audience client, so that the audience client can also hear the chorus effect and at the same time hear the normal real-time interaction between non-chorus role clients.

[0084] Corresponding to the chorus method on the side of the backup singer role client in the above embodiment, Figure 3 This is the structural block diagram of the chorus device provided by the embodiments of the present disclosure. For the sake of simplicity, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 3, the chorus device 300 includes: a receiving unit 301, a processing unit 302, a collecting unit 303, and a transmitting unit 304.

[0085] Among them, the receiving unit 301 is configured to receive the lead singer multimedia stream data;

[0086] The processing unit 302 is configured to decode and play each frame of the lead singer multimedia data in the lead singer multimedia stream data in sequence;

[0087] The collecting unit 303 is configured to collect the current frame of backup singer multimedia data during the process of the backup singer following the singing;

[0088] The processing unit 302 is further configured to process the current frame of backup singer multimedia data; determine a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a frame of lead singer multimedia data decoded currently and a preset delay time, and add the second timestamp to the current frame of backup singer multimedia data; perform encoding processing on the current frame of backup singer multimedia data;

[0089] The transmitting unit 304 is configured to transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection and microphone connection except for the backup singer role client.

[0090] In one or more embodiments of the present disclosure, when determining the second timestamp corresponding to the current frame of backup singer multimedia data according to the first timestamp in a frame of lead singer multimedia data decoded currently and the preset delay time, the processing unit 302 is configured to:

[0091] Determine a difference between the first timestamp in a frame of lead singer multimedia data decoded currently and the preset delay time, and determine the difference as the second timestamp corresponding to the current frame of backup singer multimedia data.

[0092] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:

[0093] Determine a delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determine the delay time as the preset delay time.

[0094] In one or more embodiments of the present disclosure, when determining the delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determining the delay time as the preset delay time, the processing unit 302 is configured to:

[0095] Determine a first delay time and a second delay time, where the first delay time is the delay time between the playback time of any frame of lead singer multimedia data and the acquisition time of a frame of backup singer multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data;

[0096] Determine the sum of the first delay time and the second delay time as the preset delay time.

[0097] In one or more embodiments of the present disclosure, when determining the first delay time, the processing unit 302 is configured to:

[0098] According to a first preset mapping relationship, determine whether a preset first delay time corresponding to the backup singer role client can be obtained, where the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;

[0099] If it is determined that the preset first delay time corresponding to the backup singer role client can be obtained, then determine the preset first delay time as the first delay time.

[0100] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:

[0101] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained and the playback mode of the lead singer multimedia data is the external playback mode, then determine the first delay time based on the acoustic echo cancellation algorithm.

[0102] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:

[0103] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained and the playback mode of the lead singer multimedia data is a non-external playback mode, then query in a second preset mapping relationship according to the multimedia data acquisition method used by the backup singer role client, where the second preset mapping relationship is a mapping relationship between different multimedia data acquisition methods and the preset first delay time;

[0104] Determine the preset first delay time corresponding to the multimedia data acquisition method used by the backup singer role client in the second preset mapping relationship as the first delay time.

[0105] In one or more embodiments of the present disclosure, when determining the second delay time, the processing unit 302 is configured to:

[0106] Determine the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data;

[0107] The sum of the latency times of each sub - processing procedure is determined as the second latency time.

[0108] In one or more embodiments of the present disclosure, when the processing unit 302 determines the latency time of each sub - processing procedure in the processing procedure of the lead - singer multimedia data and the processing procedure of the backup - singer multimedia data, it is configured to:

[0109] For any sub - processing procedure, obtain the time interval between the input data time and the corresponding output data time as the latency time of this sub - processing procedure.

[0110] The device provided in this embodiment can be used to execute the technical solutions of the above - mentioned method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here.

[0111] Reference Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. The electronic device 400 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in - vehicle terminals (such as in - vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0112] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to the program stored in the read - only memory (ROM) 402 or the program loaded from the storage device 408 into the random - access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0113] Typically, the following devices can be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 can allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0114] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the above methods of the embodiments of the present disclosure are executed.

[0115] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0116] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0117] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0118] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0120] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0121] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] In a first aspect, according to one or more embodiments of the present disclosure, a chorus method is provided, which is applied to a backup singer role client participating in a live connection. The method includes:

[0124] Receiving the lead singer multimedia stream data;

[0125] Successively decoding and playing each frame of the lead singer multimedia data in the lead singer multimedia stream data;

[0126] Collecting and processing the current frame of backup singer multimedia data during the backup singer's following performance;

[0127] Determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time, and adding the second timestamp to the current frame of backup singer multimedia data;

[0128] Transmitting the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection except the backup singer role client.

[0129] According to one or more embodiments of the present disclosure, the determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time includes:

[0130] Determining the difference between the first timestamp in a currently decoded frame of lead singer multimedia data and the preset delay time, and determining the difference as the second timestamp corresponding to the current frame of backup singer multimedia data.

[0131] According to one or more embodiments of the present disclosure, the method further includes:

[0132] Determine the delay time between the decoding time of any frame of the lead singer multimedia data and the time when a frame of the backup singer multimedia data with the same singing progress is completely processed, and determine this delay time as the preset delay time.

[0133] According to one or more embodiments of the present disclosure, the determining the delay time between the decoding time of any frame of the lead singer multimedia data and the time when a frame of the backup singer multimedia data with the same singing progress is completely processed, and determining this delay time as the preset delay time includes:

[0134] Determine a first delay time and a second delay time, where the first delay time is the delay time between the playback time of any frame of the lead singer multimedia data and the acquisition time of a frame of the backup singer multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data;

[0135] Determine the sum of the first delay time and the second delay time as the preset delay time.

[0136] According to one or more embodiments of the present disclosure, the determining the first delay time includes:

[0137] According to a first preset mapping relationship, determine whether a preset first delay time corresponding to the backup singer role client can be obtained, where the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;

[0138] If it is determined that the preset first delay time corresponding to the backup singer role client can be obtained, then determine the preset first delay time as the first delay time.

[0139] According to one or more embodiments of the present disclosure, the method further includes:

[0140] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playback mode of the lead singer multimedia data is the external playback mode, then determine the first delay time based on the acoustic echo cancellation algorithm.

[0141] According to one or more embodiments of the present disclosure, the method further includes:

[0142] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playback mode of the lead singer multimedia data is the non-external playback mode, then query in a second preset mapping relationship according to the multimedia data acquisition method used by the backup singer role client, where the second preset mapping relationship is a mapping relationship between different multimedia data acquisition methods and the preset first delay time;

[0143] Determine the first delay time as the preset first delay time corresponding to the multimedia data acquisition method used by the background singer role client in the second preset mapping relationship.

[0144] According to one or more embodiments of the present disclosure, the determining the second delay time includes:

[0145] Determine the delay time of each sub - processing process in the processing process of the lead singer multimedia data and the processing process of the background singer multimedia data;

[0146] Determine the sum of the delay times of each sub - processing process as the second delay time.

[0147] According to one or more embodiments of the present disclosure, the determining the delay time of each sub - processing process in the processing process of the lead singer multimedia data and the processing process of the background singer multimedia data includes:

[0148] For any sub - processing process, obtain the time interval between the input data time and the corresponding output data time as the delay time of this sub - processing process.

[0149] In a second aspect, according to one or more embodiments of the present disclosure, a chorus device is provided, including:

[0150] A receiving unit, configured to receive lead singer multimedia stream data;

[0151] A processing unit, configured to decode and play each frame of lead singer multimedia data in the lead singer multimedia stream data in sequence;

[0152] An acquisition unit, configured to acquire the current frame of background singer multimedia data during the process of the background singer following the singing;

[0153] The processing unit is further configured to process the current frame of background singer multimedia data; determine the second time stamp corresponding to the current frame of background singer multimedia data according to the first time stamp in a frame of lead singer multimedia data decoded currently and a preset delay time, and add the second time stamp to the current frame of background singer multimedia data; perform encoding processing on the current frame of background singer multimedia data;

[0154] A sending unit, configured to transmit the current frame of background singer multimedia data as background singer multimedia stream data to other clients except the background singer role client participating in the live connection.

[0155] According to one or more embodiments of the present disclosure, when determining the second time stamp corresponding to the current frame of background singer multimedia data according to the first time stamp in a frame of lead singer multimedia data decoded currently and a preset delay time, the processing unit is configured to:

[0156] Determine the difference between the first timestamp in the currently decoded frame of lead singer multimedia data and the preset delay time, and determine the difference as the second timestamp corresponding to the currently decoded frame of backup singer multimedia data.

[0157] According to one or more embodiments of the present disclosure, the processing unit is further configured to:

[0158] Determine the delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determine the delay time as the preset delay time.

[0159] According to one or more embodiments of the present disclosure, when the processing unit determines the delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determines the delay time as the preset delay time, it is configured to:

[0160] Determine a first delay time and a second delay time, where the first delay time is the delay time between the playing moment of any frame of lead singer multimedia data and the acquisition moment of a frame of backup singer multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data;

[0161] Determine the sum of the first delay time and the second delay time as the preset delay time.

[0162] According to one or more embodiments of the present disclosure, when the processing unit determines the first delay time, it is configured to:

[0163] According to a first preset mapping relationship, determine whether a preset first delay time corresponding to the backup singer role client can be obtained, where the first preset mapping relationship is the mapping relationship between different device information and the preset first delay time;

[0164] If it is determined that the preset first delay time corresponding to the backup singer role client can be obtained, then determine the preset first delay time as the first delay time.

[0165] According to one or more embodiments of the present disclosure, the processing unit is further configured to:

[0166] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playing mode corresponding to the lead singer multimedia data is the external speaker mode, then determine the first delay time based on the acoustic echo cancellation algorithm.

[0167] According to one or more embodiments of the present disclosure, the processing unit is further configured to:

[0168] If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playback mode of the lead singer multimedia data is a non-external playback mode, then query in the second preset mapping relationship according to the multimedia data acquisition method used by the backup singer role client, where the second preset mapping relationship is the mapping relationship between different multimedia data acquisition methods and the preset first delay time;

[0169] Determine the preset first delay time corresponding to the multimedia data acquisition method used by the backup singer role client in the second preset mapping relationship as the first delay time.

[0170] According to one or more embodiments of the present disclosure, when determining the second delay time, the processing unit is configured to:

[0171] Determine the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data;

[0172] Determine the sum of the delay times of each sub-processing process as the second delay time.

[0173] According to one or more embodiments of the present disclosure, when determining the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data, the processing unit is configured to:

[0174] For any sub-processing process, obtain the time interval between the input data time and the corresponding output data time as the delay time of the sub-processing process.

[0175] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;

[0176] The memory stores computer-executable instructions;

[0177] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the chorus method as described in the first aspect and various possible designs of the first aspect above.

[0178] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the chorus method as described in the first aspect and various possible designs of the first aspect above is implemented.

[0179] Fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product including computer-executable instructions, which, when executed by a processor, implement the chorus method as described in the first aspect above and various possible designs of the first aspect.

[0180] The above description is only for the preferred embodiments of the present disclosure and the description of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0181] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0182] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A choral method, characterized in that, Applied to the client of the backup singer role participating in the live video call, the method includes: Receiving the lead singer multimedia stream data; and sequentially decoding and playing each frame of the lead singer multimedia data in the lead singer multimedia stream data; Collecting and processing the current frame of backup singer multimedia data during the backup singer's follow-up singing; Determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time, and adding the second timestamp to the current frame of backup singer multimedia data; Transmitting the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live video call except the backup singer role client.

2. The method according to claim 1, wherein The determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time includes: Determining a difference between the first timestamp in a currently decoded frame of lead singer multimedia data and the preset delay time, and determining the difference as the second timestamp corresponding to the current frame of backup singer multimedia data.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Determining a delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determining the delay time as the preset delay time.

4. The method according to claim 3, wherein The determining a delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determining the delay time as the preset delay time includes: Determining a first delay time and a second delay time, where the first delay time is the delay time between the playing moment of any frame of lead singer multimedia data and the collecting moment of a frame of backup singer multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data; Determining the sum of the first delay time and the second delay time as the preset delay time.

5. The method according to claim 4, characterized in that, The determining the first delay time includes: Determining whether a preset first delay time corresponding to the backup singer role client can be obtained according to a first preset mapping relationship, where the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time; If it is determined that a preset first delay time corresponding to the backup singer role client can be obtained, then determining the preset first delay time as the first delay time.

6. The method according to claim 5, wherein The method further includes: If it is determined that a preset first delay time corresponding to the backup singer role client cannot be obtained and the playing mode corresponding to the lead singer multimedia data is the external speaker mode, then determining the first delay time based on an acoustic echo cancellation algorithm.

7. The method according to claim 5, characterized in that, The method further includes: If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then query in the second preset mapping relationship according to the multimedia data acquisition method used by the backup singer role client, where the second preset mapping relationship is the mapping relationship between different multimedia data acquisition methods and the preset first delay time; Determine the preset first delay time corresponding to the multimedia data acquisition method used by the backup singer role client in the second preset mapping relationship as the first delay time.

8. The method according to claim 4, wherein The determining the second delay time includes: Determine the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data; Determine the sum of the delay times of each sub-processing process as the second delay time.

9. The method according to claim 8, wherein The determining the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data includes: For any sub-processing process, obtain the time interval between the input data time and the corresponding output data time as the delay time of this sub-processing process.

10. A choral device, characterized in that, Includes: A receiving unit, configured to receive lead singer multimedia stream data; A processing unit, configured to decode and play each frame of lead singer multimedia data in the lead singer multimedia stream data in sequence; An acquisition unit, configured to acquire the current frame of backup singer multimedia data during the backup singer's follow-up singing; The processing unit is further configured to process the current frame of backup singer multimedia data; determine the second time stamp corresponding to the current frame of backup singer multimedia data according to the first time stamp in the current decoded frame of lead singer multimedia data and the preset delay time, and add the second time stamp to the current frame of backup singer multimedia data; Perform encoding processing on the current frame of backup singer multimedia data; A sending unit, configured to transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection and microphone except the backup singer role client.

11. An electronic device, characterized in that, Includes: At least one processor and a memory; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the method according to any one of claims 1-9 is implemented.

13. A computer program product, characterized in that, Includes computer execution instructions, and when the processor executes the computer execution instructions, the method according to any one of claims 1-9 is implemented.