Choral singing method, device, and storage medium
By receiving the lead multimedia stream data on the sergeant role client, decoding and adding a timestamp, the problem of poor chorus effect in the live broadcast room due to inconsistent delays is solved, and the accurate alignment of the multimedia stream data and the improvement of the chorus effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/141419
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
In the live broadcast room multiplayer chorus scene, due to inconsistent delay caused by network and device factors, the client of the chorus role cannot accurately add a synchronous timestamp, resulting in poor chorus effect.
By receiving the lead multimedia stream data, decoding and playing in turn, adding a time stamp to each frame of data, and determining the time stamp of the auxiliary multimedia data based on the preset delay time, ensuring the accuracy of the time stamp.
Improve the accuracy of the timestamp addition of multimedia streaming data, ensure the alignment of multimedia streaming data of each chorus character client, and improve the effect and connection experience of multi-person chorus in the live broadcast room.
Smart Images

Figure CN2024141419_03072025_PF_FP_ABST
Abstract
Description
Chorus method, device and storage medium
[0001] This application claims priority to Chinese patent application No. 202311866690.4 filed on December 29, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a chorus method, device, and storage medium. Background Art
[0003] Currently, in live video broadcast rooms or voice live broadcast rooms, the host can establish a live broadcast session with the guests to conduct real-time live broadcast and live broadcast interaction, so that the audience in the live broadcast room can watch the live broadcast interaction content.
[0004] In one scenario, the host or guests in the live broadcast room can perform a multi-person chorus of a song. Due to factors such as the network and device performance, there may be a certain delay when any party receives the multimedia stream data of the other party, and the delay may also be uneven. In order to facilitate the non-chorus role client to align the multimedia stream data of each chorus role client and then play them, the lead role client can add a synchronization timestamp to its lead multimedia stream data according to the singing progress, and the chorus role client can also add the synchronization timestamp in the lead multimedia stream data to the chorus multimedia stream data. In this way, the non-chorus role client can align the multimedia stream data of each chorus role client based on the synchronization timestamp.
[0005] However, when the chorus role client adds the synchronization timestamp in the lead singer's multimedia stream data to the chorus multimedia stream data, it is usually unable to accurately add the synchronization timestamp, resulting in a synchronization timestamp offset and failure to align with the lead singer's multimedia stream data, resulting in the subsequent inability to present a good chorus effect and unable to meet the needs of multi-person chorus in the chat room. Summary of the Invention
[0006] The embodiments of the present disclosure provide a chorus method, device, and storage medium to improve the accuracy of adding timestamps to multimedia streaming data in a live broadcast chorus scene, thereby improving the chorus effect.
[0007] In a first aspect, an embodiment of the present disclosure provides a chorus method, which is applied to a client that plays a secondary vocal role in a live broadcast, and the method includes:
[0008] Receive the lead singer multimedia stream data; and decode and play each frame of the lead singer multimedia stream data in turn;
[0009] During the chorus singing, the multimedia data of the chorus in the current frame is collected and processed;
[0010] Determining a second timestamp corresponding to the current frame of chorus multimedia data based on a first timestamp in a currently decoded frame of lead chorus multimedia data and a preset delay time, and adding the second timestamp to the current frame of chorus multimedia data;
[0011] The current frame chorus multimedia data is transmitted to other clients participating in the live broadcast except the chorus role client in the form of chorus multimedia stream data.
[0012] In a second aspect, an embodiment of the present disclosure provides a chorus device, comprising:
[0013] A receiving unit, configured to receive the lead singer multimedia stream data;
[0014] A processing unit, configured to sequentially decode and play each frame of the lead singer multimedia stream data;
[0015] A collection unit, used for collecting the multimedia data of the chorus of the current frame during the process of the chorus following the singing;
[0016] The processing unit is further configured to process the current frame of chorus multimedia data; determine a second timestamp corresponding to the current frame of chorus multimedia data based on a first timestamp in a currently decoded frame of lead chorus multimedia data and a preset delay time, and add the second timestamp to the current frame of chorus multimedia data; and perform encoding processing on the current frame of chorus multimedia data;
[0017] The sending unit is used to transmit the current frame chorus multimedia data to other clients participating in the live broadcast except the chorus role client in the form of chorus multimedia stream data.
[0018] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0019] The memory stores computer-executable instructions;
[0020] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the chorus method described in the first aspect and various possible designs of the first aspect.
[0021] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the chorus method described in the first aspect and various possible designs of the first aspect is implemented.
[0022] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, comprising computer-executable instructions. When a processor executes the computer-executable instructions, the chorus method described in the first aspect and various possible designs of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, a brief introduction to the drawings required for the embodiments will be given below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] FIG1A is a diagram illustrating a scenario of a chorus method provided by an embodiment of the present disclosure;
[0025] FIG1B is a diagram illustrating a scenario of a chorus method provided by an embodiment of the present disclosure;
[0026] FIG1C is a diagram illustrating a principle of adding a timestamp offset in a related art;
[0027] FIG2 is a schematic flow chart of a chorus method provided by an embodiment of the present disclosure;
[0028] FIG3 is a structural block diagram of a chorus device provided in an embodiment of the present disclosure;
[0029] FIG4 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0031] In one scenario, as shown in FIG1A , the host or guest in the live broadcast room can perform a multi-person chorus of a song, wherein the clients participating in the live broadcast can be divided into chorus role clients and non-chorus role clients from the perspective of roles. The chorus role clients may include the lead singer (or lead singer) role client and the chorus (or follow-up singer) role client. As shown in Figure 1B, due to factors such as network and device performance, there may be a certain delay when any party receives the multimedia stream data of the other party, and the delay may also be uneven. For example, the lead singer's multimedia stream data is sent at time T1 and can reach the chorus role client and non-chorus role client at time T2, while the chorus multimedia stream data of the chorus role client after singing along arrives at the non-chorus role client at time T3. In order to facilitate the non-chorus role client to align the multimedia stream data of each chorus role client and play them, the lead singer role client can add a synchronization timestamp to its lead singer multimedia stream data according to the singing progress, and the chorus role client can also add the synchronization timestamp in the lead singer multimedia stream data to the chorus multimedia stream data. In this way, the non-chorus role client can align the multimedia stream data of each chorus role client based on the synchronization timestamp, and the chorus effect can be presented after playback after alignment.
[0032] However, when the chorus role client adds the synchronization timestamp in the lead singer multimedia stream data to the chorus multimedia stream data, it is usually unable to accurately add the synchronization timestamp. When adding the synchronization timestamp, the chorus role client adds the synchronization timestamp in the currently decoded frame of lead singer multimedia data to the collected current frame of chorus multimedia data, that is, the collected current frame of chorus multimedia data corresponds to the currently decoded frame of lead singer multimedia data by default. However, there is a certain delay from decoding the lead singer multimedia data to playing it, and then to collecting the chorus multimedia data. This results in the current frame of chorus multimedia data not actually corresponding to the currently decoded frame of lead singer multimedia data. As shown in Figure 1C, the decoding at time T1 is A frame of lead singing multimedia data is played at time T2, and the corresponding frame of chorus multimedia data is collected at time T3. However, the frame of lead singing multimedia data currently decoded at time T3 is no longer the frame of lead singing multimedia data at time T1. The synchronization timestamp of the current frame of chorus multimedia data decoded at time T3 is added to the synchronization timestamp of the current frame of lead singing multimedia data, which will cause the synchronization timestamp of the current frame of chorus multimedia data to be offset, and cannot be aligned with a frame of lead singing multimedia stream data with the same singing progress. As a result, the lead singing multimedia data and the chorus multimedia data cannot be aligned based on the synchronization timestamp subsequently, and a good chorus effect cannot be presented, which cannot meet the needs of multi-person chorus in the live broadcast room.
[0033] In order to solve the above technical problems, the embodiment of the present disclosure provides a chorus method, which receives lead singer multimedia stream data through a chorus role client participating in live broadcast and microphone connection, wherein each frame of the lead singer multimedia stream data includes a first timestamp corresponding to the singing progress; decodes and plays each frame of the lead singer multimedia stream data in turn; collects and processes the current frame of chorus multimedia data while the chorus follows the singing; determines the second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead singer multimedia data and a preset delay time, and adds the second timestamp to the current frame of chorus multimedia data; and transmits the current frame of chorus multimedia data to other clients participating in the live broadcast and microphone connection except the chorus role client as chorus multimedia stream data.
[0034] Furthermore, after receiving the multimedia stream data of other clients except the chorus role client participating in the live broadcast, the non-chorus role client participating in the live broadcast can cache and align the multimedia stream data of the chorus role client, thereby ensuring that the chorus effect is presented on the non-chorus role client participating in the live broadcast, meeting the needs of multi-person chorus in the live broadcast room and improving the connection experience of the live broadcast room.
[0035] The chorus method provided in the above embodiment is applied to the scenario shown in Figure 1A. The live broadcast room can be a video live broadcast room or a voice live broadcast room. From the perspective of the role, the live broadcast room may include clients participating in the live broadcast and audience clients not participating in the live broadcast. The clients participating in the live broadcast include chorus role clients and non-chorus role clients. The chorus role clients are clients participating in the live broadcast chorus, and may include lead singer (or main singer) role clients and secondary singer role clients. The secondary singer role clients participating in the live broadcast can respectively execute the above corresponding chorus methods.
[0036] It should be noted that the user information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0037] The chorus method disclosed herein will be described in detail below with reference to specific embodiments.
[0038] With reference to Figure 2, Figure 2 is a flow chart of the chorus method provided by an embodiment of the present disclosure. In the chorus scenario of the live broadcast room, the live broadcast room can be a video live broadcast room or a voice live broadcast room. From the perspective of the role, the live broadcast room may include clients participating in the microphone connection and audience clients not participating in the microphone connection, wherein the clients participating in the microphone connection include chorus role clients and non-chorus role clients, and the chorus role clients are clients participating in the microphone connection chorus, and may include lead singer (or main singer) role clients and secondary singer role clients; from another perspective, the live broadcast room may include host clients and guest clients participating in the microphone connection, as well as audience clients not participating in the microphone connection, wherein the host client and the guest client may be any of the above roles (wherein there is only one lead singer role client, if the host client or a guest client is the lead singer role client, then the other clients participating in the microphone connection cannot be the lead singer role client).
[0039] The chorus method of this embodiment can be applied to a client that plays a secondary vocal role in a live broadcast. The chorus method includes:
[0040] S201, receiving lead singer multimedia stream data.
[0041] In this embodiment, each client participating in the live broadcast and microphone connection can receive multimedia streaming data from other clients participating in the live broadcast and microphone connection. In the video live broadcast room, the multimedia streaming data can be audio and video streaming data; in the voice live broadcast room, the media streaming data can be audio streaming data.
[0042] Among them, since the chorus role client needs to sing along with the multimedia stream data of the lead role client, there is a certain lag in the multimedia stream data of the chorus role client relative to the multimedia stream data of the lead role client. For example, the multimedia stream data of the lead role client has been sung to the 10th second of the song, while the multimedia stream data of the chorus role client 1 may have only been sung to the 8th second of the song. The singing progress of the multimedia stream data of different chorus role clients may also be different. The multimedia stream data of another chorus role client 2 may be sung to the 9th second of the song. In order to present a chorus effect on non-chorus role clients, the lead role client can add a timestamp to each frame of the lead multimedia data according to the singing progress. Other chorus role clients also add a timestamp to the chorus multimedia data collected on their own end based on the timestamp in each frame of the lead multimedia data. In this way, the non-chorus role clients can receive the multimedia stream data of the lead role client and the multimedia stream data of the chorus role client, and can play them after caching and aligning based on the timestamps to present a chorus effect.
[0043] When the lead vocalist client sequentially collects each frame of lead vocal multimedia data, for ease of description, the timestamp added can be recorded as the first timestamp. Since the singing progresses, the first timestamp added to each frame of lead vocal multimedia data also progresses. After a series of processing, each frame of lead vocal multimedia data is transmitted as lead vocal multimedia stream data to all other clients participating in the live broadcast and microphone connection, except for the lead vocalist client, which may include non-chorus role clients and supporting vocalist role clients participating in the microphone connection.
[0044] S202: Decode and play each frame of the lead singer multimedia stream data in sequence.
[0045] In this embodiment, after the chorus client receives the lead singer multimedia stream data, it can sequentially decode and play each frame of the lead singer multimedia data in the lead singer multimedia stream data, wherein the first timestamp of each frame of the lead singer multimedia data can be obtained during decoding. The specific decoding and playing processes can adopt any known method and are not limited here.
[0046] S203, collecting and processing the current frame of chorus multimedia data during the chorus following the singing.
[0047] In this embodiment, when each frame of the lead singer multimedia data of the lead singer multimedia stream data is played, the user on the client side of the chorus role, that is, the chorus, can sing along with the audio of the played lead singer multimedia data. The chorus client can collect the chorus multimedia data frame by frame during the process of the chorus following the singing. For the current frame of the chorus multimedia data, other processing can be performed after collection. The specific processing process can adopt any known method, which is not limited here.
[0048] S204: Determine a second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead chorus multimedia data and a preset delay time, and add the second timestamp to the current frame of chorus multimedia data.
[0049] In this embodiment, when it is necessary to add a timestamp to the current frame of chorus multimedia data, the first timestamp in the currently decoded frame of lead chorus multimedia data can be obtained. However, since there is a certain delay from decoding the lead chorus multimedia data to playing it and then collecting the chorus multimedia data, the current frame of chorus multimedia data does not actually correspond to the currently decoded frame of lead chorus multimedia data. That is, the real timestamp of the current frame of chorus multimedia data should be earlier than the first timestamp of the currently decoded frame of lead chorus multimedia data. Therefore, a preset delay time can be obtained. Based on the first timestamp of the currently decoded frame of lead chorus multimedia data, a second timestamp is determined in combination with the preset delay time, so that the second timestamp is as close as possible to the real timestamp of the current frame of chorus multimedia data. In this way, adding the second timestamp to the current frame of chorus multimedia data can improve the accuracy of adding timestamps in the chorus multimedia data.
[0050] The preset delay time can be the delay time between the decoding moment of any frame of lead singing multimedia data and the moment when a frame of chorus multimedia data with the same singing progress is processed (the processing here refers to the processing process before adding the timestamp), such as deltaT in Figure 1C, that is, the delay time from T3 to T1, which can be obtained through measurement, provided by the server, determined based on experience, or determined by any other possible method, and is not limited in this embodiment.
[0051] Based on the first timestamp in the currently decoded frame of lead chorus multimedia data and the preset delay time, the second timestamp corresponding to the current frame of chorus multimedia data is determined. Specifically, the difference between the first timestamp in the currently decoded frame of lead chorus multimedia data and the preset delay time is determined, and the difference is determined as the second timestamp corresponding to the current frame of chorus multimedia data. That is, the second timestamp should actually be ahead of the first timestamp in the currently decoded frame of lead chorus multimedia data, and the required advance amount is the preset delay time.
[0052] S205: Transmit the chorus multimedia data of the current frame as chorus multimedia stream data to other clients participating in the live broadcast room except the chorus role client.
[0053] In this embodiment, each frame of chorus multimedia data with the second timestamp is sequentially processed and transmitted as chorus multimedia stream data to other clients in the live broadcast room participating in the live broadcast, except for the chorus role client, which may include non-chorus role clients participating in the live broadcast, the lead vocal role client, and other chorus role clients. The processing and transmission processes involved can adopt any known method and are not limited in this embodiment.
[0054] The chorus method of this embodiment receives lead singer multimedia stream data; decodes and plays each frame of lead singer multimedia data of the lead singer multimedia stream data in sequence; collects and processes the current frame of chorus multimedia data during the chorus singing; determines the second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead singer multimedia data and a preset delay time, and adds the second timestamp to the current frame of chorus multimedia data; transmits the current frame of chorus multimedia data to other clients participating in the live broadcast except the chorus role client as chorus multimedia stream data. When adding a timestamp to the current frame of chorus multimedia data, based on the first timestamp in the currently decoded frame of lead singer multimedia data and in combination with the preset delay time, a second timestamp as close as possible to the real timestamp of the current frame of chorus multimedia data can be determined and added to the current frame of chorus multimedia data to improve the accuracy of the timestamp addition, so that the lead singer multimedia stream data and the chorus multimedia stream data can be better aligned according to the timestamps, thereby ensuring the chorus effect.
[0055] Based on any of the above embodiments, since a preset delay time is required, the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when a frame of chorus multimedia data with the same singing progress is completed (the processing here refers to the processing process before adding the timestamp) can also be determined, and the delay time is determined as the preset delay time.
[0056] In this embodiment, considering that the delay is caused by a certain delay in the process from decoding the lead vocal multimedia data to playing and then collecting the chorus multimedia data, the delay time from the decoding moment of any frame of lead vocal multimedia data to the moment when a frame of chorus multimedia data with the same singing progress is collected and processed (including any processing process before adding a timestamp to the chorus multimedia data) can be determined, that is, the delay time from the moment when the first timestamp is obtained by decoding any frame of lead vocal multimedia data to the moment when a timestamp is added to a frame of chorus multimedia data with the same singing progress, which is the preset delay time.
[0057] On the basis of the above embodiment, the step of determining the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when the processing of a frame of chorus multimedia data with the same singing progress is completed, and determining the delay time as the preset delay time, may specifically include:
[0058] Determine a first delay time and a second delay time, wherein the first delay time is the delay time between the playback moment of any frame of lead vocal multimedia data and the collection moment of a frame of chorus multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead vocal multimedia data and the processing process of the chorus multimedia data; the sum of the first delay time and the second delay time is determined as the preset delay time.
[0059] In this embodiment, the delay time can be divided into two parts, namely the first delay time and the second delay time. The first delay time is the delay time from the playback moment of any frame of lead vocal multimedia data to the collection moment of a frame of chorus multimedia data with the same singing progress, that is, the delay time from playback to collection, which can also be called collection and broadcasting delay; the second delay time is the remaining delay time, which is mainly the delay caused by the processing of multimedia data, including the delay time of the processing process of the lead vocal multimedia data (the processing process from decoding to playback, such as rendering, etc.) and the delay time of the processing process of the chorus multimedia data (the processing process from collection to adding timestamps, such as encoding, etc.). In this way, the first delay time and the second delay time together are the preset delay time.
[0060] First, the determination of the second delay time is explained below. Since the second delay time is the delay caused by the processing process of the multimedia data, the delay time of each sub-processing process in the processing process of the lead multimedia data and the processing process of the chorus multimedia data can be determined, and then the delay time of each sub-processing process can be added up, and the sum can be determined as the second delay time; of course, if there are some continuous sub-processing processes in the processing process, the continuous sub-processing processes can also be regarded as a whole sub-processing process, and a delay time can be obtained for the whole sub-processing process.
[0061] Among them, for any sub-process, the time interval between the input data moment and the corresponding output data moment can be obtained as the delay time of the sub-process. In actual implementation, it is not necessary to detect the delay time of each sub-process and to sum the delay time of each sub-process. Only when the lead singer multimedia stream data is initially received, the delay time of each processing process of the processing link of one frame (or multiple frames averaged) of the lead singer multimedia data and the delay time of each processing process of the processing link of one frame (or multiple frames averaged) of the chorus multimedia data can be detected, and then summed to obtain a second delay time, which will not be updated in subsequent processes; or the delay time of each sub-process in the processing process of the lead singer multimedia data and the processing process of the chorus multimedia data can be periodically detected at a preset time interval, so as to continuously update the second delay time to adapt to changes in the processing capacity of the chorus role client and ensure the accuracy of the timestamp addition.
[0062] Of course, the determination of the second delay time is not limited to the above method, and any other feasible method may be used, which is not limited here.
[0063] The determination of the first delay time can be divided into the following different situations, as follows:
[0064] Case 1:
[0065] Determining whether a preset first delay time corresponding to the chorus role client can be obtained according to a first preset mapping relationship, wherein the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;
[0066] If it is determined that the preset first delay time corresponding to the chorus role client can be obtained, the preset first delay time is determined as the first delay time.
[0067] In case one, considering that the processing performance of the same device when playing and collecting multimedia data is the same or similar, the first delay time of various devices can be obtained in advance as the preset first delay time, and a mapping relationship between device information (such as device type, device model) and the preset first delay time can be constructed, that is, the first preset mapping relationship. The acquisition method can be that the developer plays music and sings along on different devices, collects multimedia data during the singing-along process, and then detects the time delay between playing music and collecting singing-along multimedia data. Of course, other methods can also be used to determine the first delay time of different devices.
[0068] The first preset mapping relationship can be configured on the server side and can be updated regularly by the developer (for example, when some new devices are added). The secondary role client can obtain the first preset mapping relationship from the server side and query the corresponding preset first delay time from the first preset mapping relationship based on the device information of the local side as the first delay time; or the secondary role client can also send a query request to the server side, and the query request includes the device information of the local side. The server side queries the preset first delay time corresponding to the secondary role client from the first preset mapping relationship and sends it to the secondary role client as the first delay time.
[0069] Case 2:
[0070] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is the external playback mode, the first delay time is determined based on the acoustic echo cancellation algorithm.
[0071] In case 2, since the first preset mapping relationship cannot cover all types of devices, if the secondary role client cannot query the preset first delay time corresponding to the device information of the secondary role client from the first preset mapping relationship, the first delay time can be obtained through other means. Specifically, if the secondary role client uses an external speaker mode (for example, using a speaker, etc.) to play the lead singer multimedia data, since the sound of the lead singer multimedia data played externally by the secondary role client will be collected again by the sound collection device of the secondary role client, forming an acoustic echo, an acoustic echo cancellation algorithm (AEC) can be used in this embodiment to determine the first delay time, wherein the acoustic echo cancellation algorithm (AEC) compares the signal collected by the microphone with the signal output by the speaker to estimate the echo signal and then eliminate the echo signal. This involves determining the echo signal delay time, that is, the delay time between the signal collected by the microphone and the signal output by the speaker. In this embodiment, the delay time can be directly used as the required first delay time.
[0072] It should be noted that when the chorus role client collects chorus multimedia data, it usually also runs the acoustic echo cancellation algorithm (AEC) for echo cancellation. Therefore, when determining the first delay time, it is not necessary to specifically run the acoustic echo cancellation algorithm (AEC), but the delay time can be directly extracted from the echo cancellation process as the first delay time.
[0073] Case 3:
[0074] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then querying in a second preset mapping relationship according to the multimedia data collection method used by the chorus role client, wherein the second preset mapping relationship is a mapping relationship between different multimedia data collection methods and the preset first delay time;
[0075] The preset first delay time corresponding to the multimedia data acquisition method used by the chorus role client in the second preset mapping relationship is determined as the first delay time.
[0076] In case three, since the first preset mapping relationship cannot cover all types of devices, the secondary role client may not be able to query the preset first delay time corresponding to the device information of the secondary role client from the first preset mapping relationship. If the secondary role client does not use the external speaker mode to play the lead singer multimedia data, but uses headphones to play the lead singer multimedia data, then the first delay time cannot be determined using the method in case two. In this embodiment, a second preset mapping relationship can be pre-configured, and the second preset mapping relationship is configured with preset first delay times corresponding to different multimedia data collection methods. The different multimedia data collection methods mainly consider different audio collection methods, such as using OpenSL at the CS layer for audio collection, or using AudioRecord at the Java layer for audio collection. The delay of OpenSL is relatively low, and the corresponding preset first delay time is relatively small. The first delay time of the secondary role client can be determined based on the second preset mapping relationship.
[0077] The second preset mapping relationship can be configured on the server side, and the secondary role client can obtain the second preset mapping relationship from the server side, and query the corresponding preset first delay time from the second preset mapping relationship according to the multimedia data acquisition method of this side, as the first delay time; or the secondary role client can also send a query request to the server side, and the query request includes the multimedia data acquisition method of this side, and the server side queries the corresponding preset first delay time of the secondary role client from the second preset mapping relationship and sends it to the secondary role client as the first delay time.
[0078] Based on the above embodiment, after determining the first delay time and the second delay time, the first delay time and the second delay time can be added to obtain the above-mentioned preset delay time, that is, the delay time between the moment when the first timestamp is obtained by decoding any frame of lead singing multimedia data and the moment when the timestamp is added to a frame of chorus multimedia data with the same singing progress.
[0079] On this basis, based on the first timestamp in the currently decoded frame of lead chorus multimedia data, the preset delay time is subtracted to determine the second timestamp corresponding to the current frame of chorus multimedia data, and the second timestamp is added to the current frame of chorus multimedia data.
[0080] After a series of post-processing, the current frame chorus multimedia data is transmitted in the form of streaming data (ie, chorus multimedia streaming data) to other clients participating in the live broadcast except the chorus role client.
[0081] For non-chorus role clients participating in the live broadcast, the multimedia stream data of each chorus role client participating in the live broadcast can be cached, including the multimedia stream data of the lead role client and the multimedia stream data of each supporting role client. Then, according to the timestamps included in the multimedia stream data of each chorus role client, the multimedia stream data of each chorus role client can be aligned and played.
[0082] For other chorus role clients participating in the live broadcast, including the lead singer role client and other secondary singer role clients, the multimedia streaming data of the secondary singer role client can be played silently, that is, only the picture is played without playing the sound, so as to avoid affecting the singing of the user on this end.
[0083] In addition, since the host client is also responsible for pushing the live broadcast room to the audience client, no matter which role client the host client belongs to, it is necessary to align the multimedia stream data of each chorus role client, and merge the aligned multimedia stream data of each chorus role client, the multimedia stream data of other non-chorus role clients participating in the connection, and the local multimedia stream data (of course, when the host client belongs to different role clients, the local multimedia stream data may also be included in the previous multimedia stream data), and transmit the merged multimedia stream data to the audience client, so that the audience client can also hear the chorus effect and the normal real-time interaction between the non-chorus role clients.
[0084] Corresponding to the chorus method for the chorus role client in the above embodiment, Figure 3 is a block diagram of the structure of the chorus device provided in the embodiment of the present disclosure. For ease of illustration, only the portions relevant to the embodiment of the present disclosure are shown. Referring to Figure 3, the chorus device 300 includes: a receiving unit 301, a processing unit 302, a collection unit 304, and a sending unit 305.
[0085] Wherein, the receiving unit 301 is used to receive the lead singer multimedia stream data;
[0086] The processing unit 302 is configured to sequentially decode and play each frame of the lead singer multimedia stream data;
[0087] The collecting unit 304 is used to collect the multimedia data of the chorus of the current frame during the chorus singing;
[0088] The processing unit 302 is further configured to process the current frame of chorus multimedia data; determine a second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead chorus multimedia data and a preset delay time, and add the second timestamp to the current frame of chorus multimedia data; and encode the current frame of chorus multimedia data;
[0089] The sending unit 305 is configured to transmit the chorus multimedia data of the current frame to other clients participating in the live broadcast except the chorus role client in the form of chorus multimedia stream data.
[0090] In one or more embodiments of the present disclosure, when determining the second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead vocal multimedia data and the preset delay time, the processing unit 302 is configured to:
[0091] The difference between a first timestamp in a currently decoded frame of lead vocal multimedia data and the preset delay time is determined, and the difference is determined as a second timestamp corresponding to the current frame of chorus multimedia data.
[0092] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:
[0093] Determine the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when a frame of chorus multimedia data with the same singing progress is completed, and determine the delay time as the preset delay time.
[0094] In one or more embodiments of the present disclosure, when the processing unit 302 determines the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when the processing of a frame of chorus multimedia data with the same singing progress is completed, and determines the delay time as the preset delay time, it is configured to:
[0095] Determining a first delay time and a second delay time, wherein the first delay time is the delay time between the playback time of any frame of lead vocal multimedia data and the acquisition time of a frame of chorus multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead vocal multimedia data and the processing process of the chorus multimedia data;
[0096] The sum of the first delay time and the second delay time is determined as the preset delay time.
[0097] In one or more embodiments of the present disclosure, when determining the first delay time, the processing unit 302 is configured to:
[0098] Determining whether a preset first delay time corresponding to the chorus role client can be obtained according to a first preset mapping relationship, wherein the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;
[0099] If it is determined that the preset first delay time corresponding to the chorus role client can be obtained, the preset first delay time is determined as the first delay time.
[0100] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:
[0101] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is the external playback mode, the first delay time is determined based on the acoustic echo cancellation algorithm.
[0102] In one or more embodiments of the present disclosure, the processing unit 302 is further configured to:
[0103] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then querying in a second preset mapping relationship according to the multimedia data collection method used by the chorus role client, wherein the second preset mapping relationship is a mapping relationship between different multimedia data collection methods and the preset first delay time;
[0104] The preset first delay time corresponding to the multimedia data acquisition method used by the chorus role client in the second preset mapping relationship is determined as the first delay time.
[0105] In one or more embodiments of the present disclosure, when determining the second delay time, the processing unit 302 is configured to:
[0106] Determine the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data;
[0107] The sum of the delay times of each sub-process is determined as the second delay time.
[0108] In one or more embodiments of the present disclosure, when determining the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data, the processing unit 302 is configured to:
[0109] For any sub-process, the time interval between the input data moment and the corresponding output data moment is obtained as the delay time of the sub-process.
[0110] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0111] Referring to FIG4 , a schematic diagram of the structure of an electronic device 400 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 400 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG4 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0112] As shown in Figure 4, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 are also stored in RAM 403. Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0113] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although FIG4 shows an electronic device 400 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0114] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the above-mentioned method of the embodiment of the present disclosure are performed.
[0115] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0116] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0117] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0118] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0120] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0121] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0122] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0123] In a first aspect, according to one or more embodiments of the present disclosure, a chorus method is provided, which is applied to a client that plays a secondary vocal role in a live broadcast, and the method includes:
[0124] Receive the lead singer multimedia stream data;
[0125] Decoding and playing each frame of the lead singing multimedia stream data in sequence;
[0126] During the chorus singing, the multimedia data of the chorus in the current frame is collected and processed;
[0127] Determining a second timestamp corresponding to the current frame of chorus multimedia data based on a first timestamp in a currently decoded frame of lead chorus multimedia data and a preset delay time, and adding the second timestamp to the current frame of chorus multimedia data;
[0128] The current frame chorus multimedia data is transmitted to other clients participating in the live broadcast except the chorus role client in the form of chorus multimedia stream data.
[0129] According to one or more embodiments of the present disclosure, determining, based on a first timestamp in a currently decoded frame of lead vocal multimedia data and a preset delay time, a second timestamp corresponding to the current frame of chorus multimedia data includes:
[0130] The difference between a first timestamp in a currently decoded frame of lead vocal multimedia data and the preset delay time is determined, and the difference is determined as a second timestamp corresponding to the current frame of chorus multimedia data.
[0131] According to one or more embodiments of the present disclosure, the method further includes:
[0132] Determine the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when a frame of chorus multimedia data with the same singing progress is completed, and determine the delay time as the preset delay time.
[0133] According to one or more embodiments of the present disclosure, determining the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when the processing of a frame of chorus multimedia data with the same singing progress is completed, and determining the delay time as the preset delay time includes:
[0134] Determining a first delay time and a second delay time, wherein the first delay time is the delay time between the playback time of any frame of lead vocal multimedia data and the acquisition time of a frame of chorus multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead vocal multimedia data and the processing process of the chorus multimedia data;
[0135] The sum of the first delay time and the second delay time is determined as the preset delay time.
[0136] According to one or more embodiments of the present disclosure, determining the first delay time includes:
[0137] Determining whether a preset first delay time corresponding to the chorus role client can be obtained according to a first preset mapping relationship, wherein the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;
[0138] If it is determined that the preset first delay time corresponding to the chorus role client can be obtained, the preset first delay time is determined as the first delay time.
[0139] According to one or more embodiments of the present disclosure, the method further includes:
[0140] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is the external playback mode, the first delay time is determined based on the acoustic echo cancellation algorithm.
[0141] According to one or more embodiments of the present disclosure, the method further includes:
[0142] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then querying in a second preset mapping relationship according to the multimedia data collection method used by the chorus role client, wherein the second preset mapping relationship is a mapping relationship between different multimedia data collection methods and the preset first delay time;
[0143] The preset first delay time corresponding to the multimedia data acquisition method used by the chorus role client in the second preset mapping relationship is determined as the first delay time.
[0144] According to one or more embodiments of the present disclosure, determining the second delay time includes:
[0145] Determine the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data;
[0146] The sum of the delay times of each sub-process is determined as the second delay time.
[0147] According to one or more embodiments of the present disclosure, determining the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data includes:
[0148] For any sub-process, the time interval between the input data moment and the corresponding output data moment is obtained as the delay time of the sub-process.
[0149] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a chorus device, comprising:
[0150] A receiving unit, configured to receive the lead singer multimedia stream data;
[0151] A processing unit, configured to sequentially decode and play each frame of the lead singer multimedia stream data;
[0152] A collection unit, used for collecting the multimedia data of the chorus of the current frame during the process of the chorus following the singing;
[0153] The processing unit is further configured to process the current frame of chorus multimedia data; determine a second timestamp corresponding to the current frame of chorus multimedia data based on a first timestamp in a currently decoded frame of lead chorus multimedia data and a preset delay time, and add the second timestamp to the current frame of chorus multimedia data; and perform encoding processing on the current frame of chorus multimedia data;
[0154] The sending unit is used to transmit the current frame chorus multimedia data to other clients participating in the live broadcast except the chorus role client in the form of chorus multimedia stream data.
[0155] According to one or more embodiments of the present disclosure, when the processing unit determines the second timestamp corresponding to the current frame of chorus multimedia data based on the first timestamp in the currently decoded frame of lead vocal multimedia data and the preset delay time, it is configured to:
[0156] The difference between a first timestamp in a currently decoded frame of lead vocal multimedia data and the preset delay time is determined, and the difference is determined as a second timestamp corresponding to the current frame of chorus multimedia data.
[0157] According to one or more embodiments of the present disclosure, the processing unit is further configured to:
[0158] Determine the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when a frame of chorus multimedia data with the same singing progress is completed, and determine the delay time as the preset delay time.
[0159] According to one or more embodiments of the present disclosure, when the processing unit determines the delay time between the decoding moment of any frame of lead vocal multimedia data and the moment when the processing of a frame of chorus multimedia data with the same singing progress is completed, and determines the delay time as the preset delay time, it is used to:
[0160] Determining a first delay time and a second delay time, wherein the first delay time is the delay time between the playback time of any frame of lead vocal multimedia data and the acquisition time of a frame of chorus multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead vocal multimedia data and the processing process of the chorus multimedia data;
[0161] The sum of the first delay time and the second delay time is determined as the preset delay time.
[0162] According to one or more embodiments of the present disclosure, when determining the first delay time, the processing unit is configured to:
[0163] Determining whether a preset first delay time corresponding to the chorus role client can be obtained according to a first preset mapping relationship, wherein the first preset mapping relationship is a mapping relationship between different device information and the preset first delay time;
[0164] If it is determined that the preset first delay time corresponding to the chorus role client can be obtained, the preset first delay time is determined as the first delay time.
[0165] According to one or more embodiments of the present disclosure, the processing unit is further configured to:
[0166] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is the external playback mode, the first delay time is determined based on the acoustic echo cancellation algorithm.
[0167] According to one or more embodiments of the present disclosure, the processing unit is further configured to:
[0168] If it is determined that the preset first delay time corresponding to the chorus role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then querying in a second preset mapping relationship according to the multimedia data collection method used by the chorus role client, wherein the second preset mapping relationship is a mapping relationship between different multimedia data collection methods and the preset first delay time;
[0169] The preset first delay time corresponding to the multimedia data acquisition method used by the chorus role client in the second preset mapping relationship is determined as the first delay time.
[0170] According to one or more embodiments of the present disclosure, when determining the second delay time, the processing unit is configured to:
[0171] Determine the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data;
[0172] The sum of the delay times of each sub-process is determined as the second delay time.
[0173] According to one or more embodiments of the present disclosure, when determining the delay time of each sub-process in the processing of the lead vocal multimedia data and the processing of the chorus multimedia data, the processing unit is configured to:
[0174] For any sub-process, the time interval between the input data moment and the corresponding output data moment is obtained as the delay time of the sub-process.
[0175] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0176] The memory stores computer-executable instructions;
[0177] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the chorus method described in the first aspect and various possible designs of the first aspect.
[0178] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the chorus method described in the first aspect and various possible designs of the first aspect is implemented.
[0179] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising computer-executable instructions. When a processor executes the computer-executable instructions, the chorus method as described in the first aspect and various possible designs of the first aspect is implemented.
[0180] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0181] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0182] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A chorus method, applied to the client of the backup singer role participating in the live connection by voice call, the chorus method comprising: Receiving the lead singer multimedia stream data; And successively decoding and playing each frame of the lead singer multimedia data in the lead singer multimedia stream data; Collecting and processing the current frame of backup singer multimedia data during the backup singer's follow-up singing; Determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time, and adding the second timestamp to the current frame of backup singer multimedia data; Transmitting the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection by voice call except the client of the backup singer role.
2. The choral method according to claim 1, wherein The determining a second timestamp corresponding to the current frame of backup singer multimedia data according to a first timestamp in a currently decoded frame of lead singer multimedia data and a preset delay time comprises: Determining a difference between the first timestamp in a currently decoded frame of lead singer multimedia data and the preset delay time, and determining the difference as the second timestamp corresponding to the current frame of backup singer multimedia data.
3. The chorus method according to claim 1 or 2, further comprising: Determining a delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determining the delay time as the preset delay time.
4. The choral method according to claim 3, wherein, The determining a delay time between the decoding moment of any frame of lead singer multimedia data and the moment when a frame of backup singer multimedia data with the same singing progress is completely processed, and determining the delay time as the preset delay time comprises: Determining a first delay time and a second delay time, wherein the first delay time is the delay time between the playing moment of any frame of lead singer multimedia data and the collecting moment of a frame of backup singer multimedia data with the same singing progress, and the second delay time is the total delay time of the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data; Determining the sum of the first delay time and the second delay time as the preset delay time.
5. The choral method according to claim 4, wherein, The determining the first delay time comprises: Determining whether a preset first delay time corresponding to the client of the backup singer role can be obtained according to a first preset mapping relationship, wherein the first preset mapping relationship is the mapping relationship between different device information and the preset first delay time; If it is determined that a preset first delay time corresponding to the client of the backup singer role can be obtained, then determining the preset first delay time as the first delay time.
6. The chorus method according to claim 5, further comprising: If it is determined that a preset first delay time corresponding to the client of the backup singer role cannot be obtained and the playing mode corresponding to the lead singer multimedia data is the external speaker mode, then determining the first delay time based on an acoustic echo cancellation algorithm.
7. The chorus method according to claim 5, further comprising: If it is determined that the preset first delay time corresponding to the backup singer role client cannot be obtained, and the playback mode corresponding to the lead singer multimedia data is a non-external playback mode, then query in the second preset mapping relationship according to the multimedia data acquisition method used by the backup singer role client, where the second preset mapping relationship is the mapping relationship between different multimedia data acquisition methods and the preset first delay time; Determine the preset first delay time corresponding to the multimedia data acquisition method used by the backup singer role client in the second preset mapping relationship as the first delay time.
8. The chorus method according to any one of claims 4-7, wherein, The determining of the second delay time includes: Determine the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data; Determine the sum of the delay times of each sub-processing process as the second delay time.
9. The choral method according to claim 8, wherein, The determining of the delay time of each sub-processing process in the processing process of the lead singer multimedia data and the processing process of the backup singer multimedia data includes: For any sub-processing process, obtain the time interval between the input data time and the corresponding output data time as the delay time of this sub-processing process.
10. A chorus device, comprising: A receiving unit configured to receive lead singer multimedia stream data; A processing unit configured to decode and play each frame of lead singer multimedia data in the lead singer multimedia stream data in sequence; An acquisition unit configured to acquire the current frame of backup singer multimedia data during the backup singer's follow-up singing; Wherein, the processing unit is further configured to process the current frame of backup singer multimedia data; determine the second time stamp corresponding to the current frame of backup singer multimedia data according to the first time stamp in a frame of lead singer multimedia data decoded currently and the preset delay time, and add the second time stamp to the current frame of backup singer multimedia data; perform encoding processing on the current frame of backup singer multimedia data; A sending unit configured to transmit the current frame of backup singer multimedia data as backup singer multimedia stream data to other clients participating in the live connection and mic except the backup singer role client.
11. An electronic device, comprising: At least one processor and a memory; wherein, The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the chorus method according to any one of claims 1-9.
12. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the chorus method according to any one of claims 1-9 is implemented.
13. A computer program product comprising computer-executable instructions, wherein, When the processor executes the computer execution instructions, the chorus method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Online real-time chorusing method and system
CN108600815A
Audio processing method and device
CN109859730A
Telephone transmitter-connected chorus method, system, device and storage medium
CN113596516A
Online chorus method and device based on microphone connection live broadcast and online chorus system
CN116962746A
System for synchronizing accompaniment with singing voice in online karaoke service and apparatus for performing same
WO2019117362A1