Real-time online and offline chorus method, device, and medium

By using the streaming and push service network to pull and progress mark real-time audio and video streams in offline singing scenes, and combining the merge function of the cloud convergence service network, real-time chorus between online users and offline users is realized, solving the problem of the inability to realize real-time chorus between online and offline users in the existing technology, and improving user interactivity.

WO2025112547A1PCT designated stage expired Publication Date: 2025-06-05TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/104758
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-07-10
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing technology cannot realize real-time chorus between online users and offline users. Online users can only watch offline concerts in real time through terminal devices, but cannot sing with offline users in real time.

Method used

By using the streaming and push service network to pull real-time audio and video streams in offline singing scenes from the content distribution network, mark them in progress, and push them to the real-time audio and video service network. The target chorus user terminal pulls and plays real-time audio and video streams from the real-time audio and video service network, and uses the current marking progress to mark the human voice and audio streams of local online chorus applications to obtain the tagged descendant voice and audio streams carrying progress information. The audio stream of the marker descendants is uploaded to the real-time audio and video service network, and is pulled to the cloud converged service network through the real-time audio and video streaming service network, merged with the real-time audio and video stream, and pushed to the content distribution network for distribution.

Benefits of technology

It realizes the real-time chorus effect between online users and offline users, solves the problem that online users cannot participate in chorus with offline users in real time, and improves user participation and interactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024104758_05062025_PF_FP_ABST
    Figure CN2024104758_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A real-time online and offline chorus method, a device, and a medium, relating to the technical field of communications. The method is applied to a cloud server, and comprises: a stream pull-to-push service network pulls an offline real-time audio and video stream and generates a current annotation progress, and a real-time audio and video service network issues the real-time audio and video stream pushed by the stream pull-to-push service network and the current annotation progress to a target chorus user terminal, such that the target chorus user terminal plays the real-time audio and video stream and uses the current annotation progress to annotate a local human voice audio stream so as to obtain an annotated human voice audio stream; and the real-time audio and video service network acquires the annotated human voice audio stream, and uses a real-time audio and video stream pulling service network to pull the real-time audio and video stream, the current annotation progress and the annotated human voice audio stream to a cloud confluence service network, such that the cloud confluence service network combines the real-time audio and video stream and the annotated human voice audio stream, and pushes the combined audio and video stream to a content distribution network for distribution. Real-time online and offline chorus can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Offline and online real-time chorus method, device and medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 27, 2023, with application number 202311602552.5 and invention name “A method, device and medium for offline and online real-time chorus”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the field of communication technologies, and in particular to a method, device, and medium for real-time offline and online chorus. Background Art

[0003] Currently, some chorus apps on the market allow users to participate in choruses online, but this means that participation is currently limited to online. For offline performances, such as celebrity concerts and streaming concerts, performers and audiences can participate offline, while online users can only watch the offline concert in real time through their devices. Therefore, a solution that enables real-time chorus between online and offline users is still lacking.

[0004] In summary, how to achieve real-time chorus between online and offline users is a problem that needs to be solved.

[0005] Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a method, device and medium for real-time offline and online chorus, which can realize real-time chorus between online and offline users. The specific scheme is as follows:

[0007] In a first aspect, the present application discloses a real-time offline and online chorus method, which is applied to a cloud server and includes:

[0008] Using a pull-stream-pushing service network to pull the real-time audio and video stream of the offline concert scene from the content distribution network, marking the progress of the real-time audio and video stream to obtain the current marking progress, and pushing the real-time audio and video stream and the current marking progress to the real-time audio and video service network;

[0009] The real-time audio and video stream and the current marking progress are sent to a target chorus user terminal using the real-time audio and video service network so that the target chorus user terminal can play the real-time audio and video stream, and the human voice audio stream acquired in real time by the local online chorus application is marked using the current marking progress to obtain a marked human voice audio stream carrying progress information;

[0010] The marked vocal audio stream uploaded by the target chorus user terminal is obtained through the real-time audio and video service network, and the real-time audio and video stream, the current marking progress and the marked vocal audio stream are pulled to the cloud-based merging service network using the real-time audio and video streaming service network, so as to merge the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and the progress information using the cloud-based merging service network, and push the merged audio and video stream to the content distribution network for distribution.

[0011] Optionally, the step of pulling the real-time audio and video stream of the offline concert scene from the content distribution network using the pull and push service network, and performing progress marking on the real-time audio and video stream to obtain the current marking progress, includes:

[0012] When the current time interval reaches a preset time interval threshold, a real-time audio and video stream corresponding to the preset time interval threshold is pulled from the current target audio and video stream using a stream pull and push service network based on the order of timestamps from front to back; the current target audio and video stream is an audio and video stream in the content distribution network that corresponds to the offline singing scene and has not been pulled yet;

[0013] A current progress sequence number corresponding to the current stream pulling operation is determined, and the current progress sequence number is used to perform progress marking on the real-time audio and video stream to obtain a current marking progress.

[0014] Optionally, when the current time interval reaches a preset time interval threshold, using the stream pull and push service network to pull the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the order of timestamps from front to back, includes:

[0015] Determining a target stream pulling timestamp corresponding to each stream pulling operation based on the first stream pulling timestamp and the preset time interval threshold; wherein the time when the stream pulling and forwarding service network first pulls the real-time audio and video stream in the offline concert scene from the content distribution network is used as the first stream pulling timestamp;

[0016] Whenever the current time reaches the target stream pulling timestamp, the stream pulling and pushing service network is used to pull the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the order of timestamps from front to back.

[0017] Optionally, determining a current progress sequence number corresponding to the current stream pulling operation, and using the current progress sequence number to perform progress marking on the real-time audio and video stream to obtain a current marking progress, includes:

[0018] Determining a current progress sequence number corresponding to the current stream pulling operation based on the historical marking progress corresponding to the previous stream pulling timestamp and the preset time interval threshold; wherein, if the current stream pulling operation is the first stream pulling operation, the current progress sequence number is a preset initial value;

[0019] The current progress sequence number is used as the current marking progress of the real-time audio and video stream.

[0020] Optionally, the step of pulling the real-time audio and video stream of the offline concert scene from the content distribution network using the pull and push service network, and performing progress marking on the real-time audio and video stream to obtain the current marking progress, includes:

[0021] The pull-stream and push-return service network is used to continuously pull real-time audio and video streams in offline singing scenes from the content distribution network, and when the current time interval reaches the preset time interval threshold, the real-time audio and video stream is progress-marked using the preset progress serial number to obtain the current marking progress.

[0022] Optionally, the offline and online real-time chorus method further includes:

[0023] Obtain online vocal audio streams uploaded by other chorus user terminals;

[0024] The real-time audio and video service network is used to forward the online vocal audio stream uploaded by the other chorus user terminals to the target chorus user terminal, so that the target chorus user terminal can simultaneously play the real-time audio and video stream and the online vocal audio stream uploaded by the other chorus user terminals.

[0025] Optionally, pushing the merged audio and video stream to the content distribution network for distribution includes:

[0026] The merged audio and video stream is pushed to each content distribution network node of the content distribution network, so that the content distribution network node sends the merged audio and video stream to the online audience user terminal after receiving the data pull request sent by the online audience user terminal; wherein, the online audience user terminal selects one content distribution network node from each content distribution network node based on the current network communication status and sends the data pull request to it.

[0027] In a second aspect, the present application discloses a real-time offline and online chorus method, which is applied to a target chorus user terminal, comprising:

[0028] Pulling and playing the real-time audio and video stream of the offline concert scene from the real-time audio and video service network in the cloud server, and obtaining the current marking progress of the real-time audio and video stream; wherein the cloud server uses the pull-stream-pushing service network to pull the real-time audio and video stream from the content distribution network, and performs progress marking on the real-time audio and video stream to obtain the current marking progress;

[0029] Using the current marking progress, marking the human voice audio stream acquired in real time by the local online chorus application, to obtain a marked human voice audio stream carrying progress information;

[0030] The marked human voice audio stream is uploaded to the real-time audio and video service network, so that the real-time audio and video streaming service network can send the real-time audio and video stream, the current marking progress and the marked human voice audio stream pulled from the real-time audio and video service network to the cloud-based merging service network, and merge the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and the progress information through the cloud-based merging service network, and push the merged audio and video stream to the content distribution network for distribution.

[0031] Optionally, the step of marking the human voice audio stream acquired in real time by the local online chorus application using the current marking progress includes:

[0032] A local audio acquisition device is started through a local online chorus application to collect local audio data through the local audio acquisition device, and the local audio data is subjected to denoising to obtain a human voice audio stream;

[0033] The human voice audio stream obtained through the local online chorus application is marked using the current marking progress.

[0034] In a third aspect, the present application discloses an electronic device, comprising:

[0035] Memory, used to store computer programs;

[0036] The processor is used to execute the computer program to implement the steps of the aforementioned offline and online real-time chorus method.

[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned offline and online real-time chorus method are implemented.

[0038] It can be seen that the present application discloses an offline and online real-time chorus method, which is applied to a cloud server, specifically comprising: utilizing a pull-stream-pushing service network to pull the real-time audio and video stream in the offline singing scene from the content distribution network, and marking the progress of the real-time audio and video stream to obtain the current marking progress, and pushing the real-time audio and video stream and the current marking progress to the real-time audio and video service network; utilizing the real-time audio and video service network to send the real-time audio and video stream and the current marking progress to the target chorus user terminal, so that the target chorus user terminal plays the real-time audio and video stream, and utilizing the current marking progress to the local online chorus The human voice audio stream obtained in real time by the singing application is marked to obtain a marked human voice audio stream carrying progress information; the marked human voice audio stream uploaded by the target chorus user terminal is obtained through the real-time audio and video service network, and the real-time audio and video stream pulling service network is used to pull the real-time audio and video stream, the current marking progress and the marked human voice audio stream to the cloud confluence service network, so as to use the cloud confluence service network to merge the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and the progress information, and push the merged audio and video stream to the content distribution network for distribution.

[0039] It can be seen that after the pull-stream-pushing service network pulls the real-time audio and video stream in the offline singing scene from the content distribution network, it will mark the progress of the real-time audio and video stream to obtain the current marking progress, and then push the real-time audio and video stream and the current marking progress to the real-time audio and video service network. That is, the pull-stream-pushing service in this application provides the functions of pulling, marking and pushing audio and video streams. Furthermore, the real-time audio and video service network sends the real-time audio and video stream and the current marking progress to the target chorus user terminal so that the target chorus user terminal plays the acquired real-time audio and video stream, and uses the current marking progress to mark the human voice audio stream acquired in real time by the local online chorus application, and obtains the marked human voice audio stream carrying progress information; it can be understood that by marking the locally collected human voice audio stream using the current marking progress, the offline real-time audio and video stream and the online human voice audio stream can be associated. After the target chorus user terminal uploads the marked vocal audio stream to the real-time audio and video service network, it then uses the real-time audio and video streaming service network to pull the real-time audio and video stream, the current marking progress, and the marked vocal audio stream to the cloud-based merging service network. Since the real-time audio and video stream and the vocal audio stream are now marked with the corresponding progress, the cloud-based merging service network accurately merges the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and progress information to obtain the merged audio and video stream, and finally pushes the merged audio and video stream back to the content distribution network for distribution. The above solution achieves the effect of real-time chorus of online and offline users. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0041] FIG1 is a schematic diagram of a system framework applicable to the offline and online real-time chorus solution disclosed in this application;

[0042] FIG2 is a flow chart of a cloud server-side offline and online real-time chorus method disclosed in this application;

[0043] FIG3 is a flowchart of a specific offline and online real-time chorus method disclosed in this application;

[0044] FIG4 is a flow chart of a method for real-time offline and online chorus of a chorus user terminal disclosed in this application;

[0045] FIG5 is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0047] Currently, in offline concerts, such as celebrity concerts and streaming concerts, performers and audiences can sing together offline, while online users can only watch the offline concert in real time through terminal devices. Therefore, there is still a lack of a solution to enable online and offline users to sing together in real time. To this end, the embodiments of the present application disclose a real-time offline and online chorus method, device, and medium that can enable online and offline users to sing together in real time.

[0048] In the offline and online real-time chorus solution of this application, the system framework adopted can be shown in Figure 1, which can specifically include a cloud server 01, a target chorus user terminal 02 and an online audience user terminal 03.

[0049] Among them, the cloud server 01 specifically includes a streaming and forwarding service network 011, a real-time audio and video service network 012, a real-time audio and video streaming service network 013 and a cloud-side converging service network 014.

[0050] The main function of the pull-stream and push-forward service network 011 is to provide content pulling and pushing, and this application also provides a progress marking function. It should be pointed out that audio and video acquisition equipment and push-stream equipment are installed in the offline singing scene, which are used to collect the audio and video streams in the offline singing scene in real time and push them to the content delivery network (Content Delivery Network, i.e. CDN). Therefore, the pull-stream and push-forward service network 011 specifically pulls the real-time audio and video streams in the offline singing scene from the content delivery network, and marks the progress of the real-time audio and video streams to obtain the current marking progress, and then pushes the real-time audio and video streams and the current marking progress to the real-time audio and video service network 012.

[0051] The real-time audio and video service network 012 is used to provide low-latency real-time audio and video services, including but not limited to Tencent Cloud TRTC (Tencent RTC, where RTC stands for Real Time Communication, real-time audio and video), Agora RTC, etc. Data is exchanged between the real-time audio and video service network 012, the streaming and forwarding service network 011, the target chorus user terminal 02, and the real-time audio and video streaming service network 013. Specifically, the real-time audio and video service network 012 is used to receive the real-time audio and video stream and the current marking progress pushed by the streaming and forwarding service network 011; the real-time audio and video service network 012 is also used to send the real-time audio and video stream and the current marking progress to the target chorus user terminal 02, and to obtain the marked human voice audio stream uploaded by the target chorus user terminal 02; the real-time audio and video service network 012 is also used to respond to the streaming operation of the real-time audio and video streaming service network 013.

[0052] The real-time audio and video streaming service network 013 is used to pull the real-time audio and video stream, current marking progress and marked human voice audio stream obtained from the real-time audio and video service network 012 to the cloud merging service network 014.

[0053] The cloud-based merging service network 014 is used to merge the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and the progress information carried by the marked vocal audio stream, and push the merged audio and video stream to the content distribution network for distribution, thereby achieving the effect that the online audience user terminal 03 can hear the real-time chorus of online and offline users.

[0054] The target chorus user terminal 02 is a user terminal that uses the online chorus application software to participate in the chorus online via the network. The terminal device held by the target chorus user can be a smartphone, tablet, computer, or other device. The target chorus user terminal 02 is used to obtain and play the offline real-time audio and video stream from the real-time audio and video service network 012, and use the current marking progress obtained from the real-time audio and video service network 012 to mark the human voice audio stream obtained in real time by the local online chorus application, obtaining a marked human voice audio stream carrying progress information, and then pushing the marked human voice audio stream to the real-time audio and video service network 012.

[0055] Online audience user terminal 03 is a user terminal that uses the online chorus application software to listen to the chorus via the Internet. The terminal device held by the audience user can be a smartphone, tablet, computer, etc. Online audience user terminal 03 is used to pull the combined audio and video stream from the content distribution network, thereby listening to the real-time chorus effect of online and offline users.

[0056] As shown in FIG2 , an embodiment of the present application discloses a real-time offline and online chorus method, which is applied to a cloud server. The method includes:

[0057] Step S11: Use the pull-stream push service network to pull the real-time audio and video stream in the offline singing scene from the content distribution network, mark the progress of the real-time audio and video stream to obtain the current marking progress, and push the real-time audio and video stream and the current marking progress to the real-time audio and video service network.

[0058] In this embodiment, after the pull-stream-pushing service network pulls the real-time audio and video stream of the offline concert scene from the content distribution network, it performs progress marking on the real-time audio and video stream to obtain the current marking progress, and then pushes the real-time audio and video stream and the current marking progress to the real-time audio and video service network. In one embodiment, the real-time audio and video stream and the current marking progress can be pushed to the real-time audio and video service network as two independent data. In another embodiment, the real-time audio and video stream can carry the current marking progress and then be pushed to the real-time audio and video service network as a whole.

[0059] It should be noted that the main function of the stream pulling and pushing service network is to provide content pulling and pushing. In the embodiment of the present application, the stream pulling and pushing service also provides a progress marking function for marking the progress of the pulled real-time audio stream.

[0060] In addition, it should be pointed out that audio and video acquisition equipment and streaming equipment are installed in the offline concert scene to collect the video images and sounds of the event in real time to obtain audio and video streams, which are then pushed to the content distribution network. In the relevant scheme, after the content distribution network obtains the real-time audio and video streams in the offline concert scene, it can then carry out subsequent content distribution, that is, push the real-time audio and video streams to each content distribution network node, so that the content distribution network node can send the real-time audio and video streams to the online audience user terminal after receiving the data pull request sent by the online audience user terminal, so that the online audience user terminal can watch the offline concert in real time.

[0061] In a specific embodiment, the method of using a stream pull and push service network to pull the real-time audio and video stream of an offline concert scene from a content distribution network and marking the real-time audio and video stream with a progress to obtain the current marking progress includes: using the stream pull and push service network to continuously pull the real-time audio and video stream of the offline concert scene from the content distribution network, and when the current time interval reaches a preset time interval threshold, marking the real-time audio and video stream with a pre-set progress sequence number to obtain the current marking progress. That is, in this embodiment, the stream pull and push service continuously pulls the real-time audio and video stream of the offline concert scene from the content distribution network, and the stream pull and push service also needs to mark the progress of the pulled real-time audio and video stream. Specifically, each time the current time interval reaches the preset time interval threshold, the currently pulled real-time audio and video stream is marked with a progress sequence number to obtain the current marking progress. For example, assuming that a frame of data is 20ms and the preset time interval threshold is set to 20ms, it is understood that the stream pull and push service performs a progress marking on the currently pulled real-time audio and video stream every 20ms. In a specific embodiment, the progress sequence number can also be set based on a preset time interval threshold. That is, assuming that the first time the real-time audio and video stream is pulled is set as the starting time 0, the progress is accumulated every 20ms, thereby obtaining progress sequence numbers of 0, 20, 40, 60, etc. In addition, the progress sequence number can also be set based on the number of marking operations. For example, each time a progress marking operation is performed, the progress sequence number is accumulated by 1, thereby obtaining progress sequence numbers of 1, 2, 3, etc. This embodiment does not limit the method for setting the progress sequence number.

[0062] Step S12: Using the real-time audio and video service network, the real-time audio and video stream and the current marking progress are sent to the target chorus user terminal so that the target chorus user terminal plays the real-time audio and video stream, and the current marking progress is used to mark the human voice audio stream obtained in real time by the local online chorus application to obtain a marked human voice audio stream carrying progress information.

[0063] In this embodiment, the real-time audio and video service network sends the real-time audio and video stream pushed by the pull-stream forwarding service and the current marking progress to the target chorus user terminal, so that the target chorus user terminal plays the obtained real-time audio and video stream, and uses the current marking progress to mark the human voice audio stream obtained in real time by the local online chorus application to obtain a marked human voice audio stream carrying progress information; it can be understood that by using the current marking progress to mark the locally collected human voice audio stream, the offline real-time audio and video stream and the online human voice audio stream can be associated.

[0064] Among them, the local online chorus application can specifically be application software such as Kugou or QQ Music, and users can click to join the chorus and exit the chorus at will.

[0065] Furthermore, the above method also includes: obtaining the online vocal audio stream uploaded by other chorus user terminals; forwarding the online vocal audio stream uploaded by the other chorus user terminals to the target chorus user terminal by using the real-time audio and video service network, so that the target chorus user terminal can simultaneously play the real-time audio and video stream and the online vocal audio stream uploaded by the other chorus user terminals. That is, in addition to playing the offline real-time audio and video stream, the target chorus user terminal will also simultaneously play the online vocal audio stream uploaded by other chorus user terminals, that is, the chorus users will hear the voices of offline chorus users and other online chorus users. In a specific embodiment, users participating in the chorus will pull and play the online vocal audio streams of other online choristers through the low-latency RTC channel of the real-time audio and video service network. That is, the terminal users participating in the chorus transmit audio streams to each other through the channel provided by the real-time audio and video service network.

[0066] Step S13: Obtain the marked vocal audio stream uploaded by the target chorus user terminal through the real-time audio and video service network, and use the real-time audio and video streaming service network to pull the real-time audio and video stream, the current marking progress and the marked vocal audio stream to the cloud-based merging service network, so as to use the cloud-based merging service network to merge the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and the progress information, and push the merged audio and video stream to the content distribution network for distribution.

[0067] In this embodiment, the target chorus user terminal will upload the marked human voice audio stream to the real-time audio and video service network, and then use the real-time audio and video streaming service network to pull the real-time audio and video stream, the current marking progress and the marked human voice audio stream to the cloud merging service network. Since the real-time audio and video stream and the human voice audio stream at this time have been marked with the corresponding progress, the cloud merging service network will accurately merge the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and progress information to obtain the merged audio and video stream, thereby achieving precise alignment of the audio and video streams. Finally, the cloud merging service pushes the merged audio and video stream back to the content distribution network for distribution, so that other viewers can hear the effect of singing together offline and online.

[0068] It can be seen that the present application discloses an offline and online real-time chorus method, which is applied to a cloud server, specifically comprising: utilizing a pull-stream-pushing service network to pull the real-time audio and video stream in the offline singing scene from the content distribution network, and marking the progress of the real-time audio and video stream to obtain the current marking progress, and pushing the real-time audio and video stream and the current marking progress to the real-time audio and video service network; utilizing the real-time audio and video service network to send the real-time audio and video stream and the current marking progress to the target chorus user terminal, so that the target chorus user terminal plays the real-time audio and video stream, and utilizing the current marking progress to the local online chorus The human voice audio stream obtained in real time by the singing application is marked to obtain a marked human voice audio stream carrying progress information; the marked human voice audio stream uploaded by the target chorus user terminal is obtained through the real-time audio and video service network, and the real-time audio and video stream pulling service network is used to pull the real-time audio and video stream, the current marking progress and the marked human voice audio stream to the cloud confluence service network, so as to use the cloud confluence service network to merge the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and the progress information, and push the merged audio and video stream to the content distribution network for distribution.

[0069] It can be seen that after the pull-stream-pushing service network pulls the real-time audio and video stream in the offline singing scene from the content distribution network, it will mark the progress of the real-time audio and video stream to obtain the current marking progress, and then push the real-time audio and video stream and the current marking progress to the real-time audio and video service network. That is, the pull-stream-pushing service in this application provides the functions of pulling, marking and pushing audio and video streams. Furthermore, the real-time audio and video service network sends the real-time audio and video stream and the current marking progress to the target chorus user terminal so that the target chorus user terminal plays the acquired real-time audio and video stream, and uses the current marking progress to mark the human voice audio stream acquired in real time by the local online chorus application, and obtains the marked human voice audio stream carrying progress information; it can be understood that by marking the locally collected human voice audio stream using the current marking progress, the offline real-time audio and video stream and the online human voice audio stream can be associated. After the target chorus user terminal uploads the marked vocal audio stream to the real-time audio and video service network, it then uses the real-time audio and video streaming service network to pull the real-time audio and video stream, the current marking progress, and the marked vocal audio stream to the cloud-based merging service network. Since the real-time audio and video stream and the vocal audio stream are now marked with the corresponding progress, the cloud-based merging service network accurately merges the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and progress information to obtain the merged audio and video stream, and finally pushes the merged audio and video stream back to the content distribution network for distribution. The above solution achieves the effect of real-time chorus of online and offline users.

[0070] As shown in FIG3 , the embodiment of the present application discloses a specific offline and online real-time chorus method. Compared with the previous embodiment, this embodiment further illustrates and optimizes the technical solution. Specifically, it includes:

[0071] Step S21: When the current time interval reaches the preset time interval threshold, the real-time audio and video stream corresponding to the preset time interval threshold is pulled from the current target audio and video stream based on the order of timestamps from front to back using the pull-stream push service network; the current target audio and video stream is the audio and video stream in the content distribution network that corresponds to the offline singing scene and has not been pulled yet.

[0072] In this embodiment, the pull-stream and forward-stream service can also intermittently pull the real-time audio and video stream of the offline concert scene from the content distribution network. However, to ensure real-time performance and reduce latency, the time interval threshold should not be set too large. For example, it can be set to pull from the content distribution network once every 20ms. Therefore, when the current time interval reaches the preset time interval threshold, the pull-stream and forward-stream service network pulls the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the timestamp order from the front to the back.

[0073] It is understood that to avoid pulling duplicate audio and video streams and ensure the continuity of the audio and video streams, this embodiment requires pulling real-time audio and video streams corresponding to the preset time interval threshold from the currently unpulled audio and video streams in the content distribution network in order of timestamps from the beginning to the end. In other words, the pull and push service pulls 20ms of audio and video streams from the content distribution network each time. Assuming that an audio and video stream is recorded in milliseconds as 0-100ms, the first pull will pull 0-19ms of content, the second pull will pull 20-39ms of content, and so on.

[0074] In a specific embodiment, when the current time interval reaches the preset time interval threshold, the pull-stream-pushing service network is used to pull the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the order of timestamps from front to back, including: determining the target pull-stream timestamp corresponding to each pull-stream operation based on the first pull-stream timestamp and the preset time interval threshold; wherein the time when the pull-stream-pushing service network first pulls the real-time audio and video stream in the offline singing scene from the content distribution network is used as the first pull-stream timestamp; whenever the current time reaches the target pull-stream timestamp, the pull-stream-pushing service network is used to pull the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the order of timestamps from front to back.

[0075] That is, since the pull-stream and forwarding service performs a pull-stream operation every preset time interval, after determining the time when the pull-stream and forwarding service network first pulls the real-time audio and video stream of the offline concert scene from the content distribution network (i.e., the first pull-stream timestamp), the target pull-stream timestamp corresponding to each pull-stream operation can be determined based on the first pull-stream timestamp and the preset time interval threshold. In this way, whenever the current time reaches the target pull-stream timestamp, the pull-stream and forwarding service network will use the timestamp order from the current target audio and video stream to pull the real-time audio and video stream corresponding to the preset time interval threshold.

[0076] Step S22: Determine the current progress sequence number corresponding to the current stream pulling operation, and use the current progress sequence number to mark the progress of the real-time audio and video stream to obtain the current marking progress, and push the real-time audio and video stream and the current marking progress to the real-time audio and video service network.

[0077] In this embodiment, after the stream pulling and pushing service performs a stream pulling operation and pulls the real-time audio and video stream, it uses the current progress serial number corresponding to the current stream pulling operation to mark the progress of the real-time audio and video stream to obtain the current marking progress, and then pushes the real-time audio and video stream and the current marking progress to the real-time audio and video service network.

[0078] In a specific embodiment, determining the current progress sequence number corresponding to the current stream pull operation and using the current progress sequence number to mark the progress of the real-time audio and video stream to obtain the current marked progress includes: determining the current progress sequence number corresponding to the current stream pull operation based on the historical marked progress corresponding to the previous stream pull timestamp and the preset time interval threshold; wherein, if the current stream pull operation is the first stream pull operation, the current progress sequence number is a preset initial value; and using the current progress sequence number as the current marked progress of the real-time audio and video stream. That is, this embodiment can set the marked progress sequence number based on the time of the stream pull. Specifically, the historical marked progress corresponding to the previous stream pull timestamp is determined, and then the preset time interval threshold is added to this to obtain the current progress sequence number corresponding to the current stream pull operation. For example, assuming the historical marked progress corresponding to the previous stream pull timestamp is 40 and the preset stream pull time interval threshold is 20, then the current progress sequence number corresponding to the current stream pull operation is 60. If the stream pull time interval threshold remains unchanged, the marked progress gradually increases at a rate of 20ms over time. If this is the first stream pulling operation, the current progress sequence number is a preset initial value, for example, it can be set to 0.

[0079] Step S23: Using the real-time audio and video service network, the real-time audio and video stream and the current marking progress are sent to the target chorus user terminal so that the target chorus user terminal plays the real-time audio and video stream, and the current marking progress is used to mark the human voice audio stream obtained in real time by the local online chorus application to obtain a marked human voice audio stream carrying progress information.

[0080] Step S24: Obtain the marked vocal audio stream uploaded by the target chorus user terminal through the real-time audio and video service network, and use the real-time audio and video streaming service network to pull the real-time audio and video stream, the current marking progress and the marked vocal audio stream to the cloud-based merging service network, so as to use the cloud-based merging service network to merge the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and the progress information.

[0081] Step S25: Push the merged audio and video stream to each content distribution network node of the content distribution network, so that the content distribution network node sends the merged audio and video stream to the online audience user terminal after receiving the data pull request sent by the online audience user terminal; wherein, the online audience user terminal selects one content distribution network node from each content distribution network node based on the current network communication status and sends the data pull request to it.

[0082] In this embodiment, after the cloud-based merging service network merges the real-time audio and video stream and the marked human voice audio stream, it pushes the merged audio and video stream to each content distribution network node of the content distribution network. After receiving the data pull request sent by the online viewer user terminal, the content distribution network node sends the merged audio and video stream to the online viewer user terminal. It can be understood that the online viewer user terminal selects a content distribution network node from each content distribution network node based on the current network communication status and sends a data pull request to it. In a specific embodiment, the online viewer user terminal will pull the merged audio and video stream through the content distribution network node closest to it, thereby achieving the effect that the online viewer can hear the offline user and the online user singing together.

[0083] For more specific processing procedures of the above steps S23 and S24, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0084] It can be seen that in the embodiment of the present application, the stream pulling and forwarding service network can complete a stream pulling and progress marking operation every preset time interval threshold. Specifically, since the stream pulling and forwarding service performs a stream pulling operation every preset time interval threshold, after determining the first stream pulling timestamp of the stream pulling and forwarding service network, the target stream pulling timestamp corresponding to each stream pulling operation can be determined based on the first stream pulling timestamp and the preset time interval threshold. In this way, whenever the current time reaches the target stream pulling timestamp, a stream pulling operation is performed. In this process, in order to avoid pulling repeated audio and video streams and ensure the continuity of the audio and video streams, this embodiment needs to pull the real-time audio and video streams corresponding to the preset time interval threshold from the audio and video streams that are not currently pulled in the content distribution network in the order of timestamps from front to back. In addition, the online audience user terminal will pull the merged audio and video stream through the content distribution network node closest to itself, thereby achieving the effect that the online audience can hear the offline user and the online user singing together.

[0085] As shown in FIG4 , an embodiment of the present application discloses a real-time offline and online chorus method, which is applied to a target chorus user terminal. The method includes:

[0086] Step S31: Pull and play the real-time audio and video stream in the offline singing scene from the real-time audio and video service network in the cloud server, and obtain the current marking progress of the real-time audio and video stream; wherein, the cloud server uses the pull-stream and push-stream service network to pull the real-time audio and video stream from the content distribution network, and performs progress marking on the real-time audio and video stream to obtain the current marking progress.

[0087] In this embodiment, the target chorus user terminal is the user terminal participating in the chorus online. The target chorus user terminal pulls and plays the real-time audio and video stream of the offline singing scene from the real-time audio and video service network in the cloud server, and obtains the current marking progress of the real-time audio and video stream. It should be noted that the cloud server specifically uses the pull-stream-pushing service network to pull the real-time audio and video stream from the content distribution network, and performs progress marking on the real-time audio and video stream to obtain the current marking progress. The pull-stream-pushing service then pushes the real-time audio and video stream and the current marking progress to the real-time audio and video service network.

[0088] Step S32: using the current marking progress to mark the human voice audio stream acquired in real time by the local online chorus application, to obtain a marked human voice audio stream carrying progress information.

[0089] In this embodiment, the target chorus user terminal uses the obtained current marking progress to mark the human voice audio stream collected by the local audio collection device and obtained by the local online chorus application, thereby obtaining a marked human voice audio stream carrying progress information. The local audio collection device can be a microphone.

[0090] In a specific embodiment, the aforementioned use of the current marking progress to mark the vocal audio stream captured in real time by the local online chorus application includes: enabling the local audio capture device through the local online chorus application to capture local audio data, and performing denoising on the local audio data to obtain a vocal audio stream; and marking the vocal audio stream captured by the local online chorus application using the current marking progress. It is understood that after the user confirms their participation in the chorus, the local online chorus application will enable the local audio capture device, i.e., a microphone, and then capture local audio data through the microphone. During this process, the local audio data may be denoised to remove ambient sound and noise to obtain a vocal audio stream. The vocal audio stream captured by the local online chorus application is then marked using the current marking progress, and the locally collected online vocal data is marked using the first progress information. By marking the locally collected vocal audio stream using the current marking progress, it is possible to associate the offline real-time audio and video stream with the online vocal audio stream.

[0091] Step S33: Upload the marked human voice audio stream to the real-time audio and video service network, so that the real-time audio and video streaming service network sends the real-time audio and video stream, the current marking progress and the marked human voice audio stream pulled from the real-time audio and video service network to the cloud-based merging service network, and merges the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and the progress information through the cloud-based merging service network, and pushes the merged audio and video stream to the content distribution network for distribution.

[0092] In this embodiment, the target chorus user terminal uploads the marked human voice audio stream to the real-time audio and video service network. Furthermore, the real-time audio and video streaming service network simultaneously pulls the real-time audio and video stream, the current marking progress and the marked human voice audio stream in the real-time audio and video service platform to the cloud-based merging service network. The cloud-based merging service network accurately merges the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and progress information to obtain a merged audio and video stream, thereby achieving precise alignment of the offline real-time audio and video stream and the displayed human voice audio stream. Finally, the cloud-based merging service network pushes the merged audio and video stream to the content distribution network for distribution, thereby achieving the effect of other viewers being able to hear the online and offline users singing together.

[0093] As can be seen, the target chorus user terminal is the user terminal participating in the chorus online. The target chorus user terminal pulls and plays the real-time audio and video stream of the offline singing scene from the real-time audio and video service network in the cloud server, and obtains the current labeling progress of the real-time audio and video stream. Furthermore, the target chorus user terminal uses the obtained current labeling progress to label the human voice audio stream collected by the local audio collection device and obtained by the local online chorus application, thereby obtaining a labeled human voice audio stream carrying progress information. Finally, the target chorus user terminal uploads the marked vocal audio stream to the real-time audio and video service network. Furthermore, the real-time audio and video streaming service network simultaneously pulls the real-time audio and video stream, the current marking progress and the marked vocal audio stream in the real-time audio and video service platform to the cloud-based merging service network. The cloud-based merging service network accurately merges the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and progress information to obtain the merged audio and video stream, thereby achieving precise alignment of the offline real-time audio and video stream and the displayed vocal audio stream. Finally, the cloud-based merging service network pushes the merged audio and video stream to the content distribution network for distribution, so that other viewers can hear the online and offline users singing together.

[0094] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Specifically, the device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the offline and online real-time chorus method performed by the electronic device as disclosed in any of the aforementioned embodiments.

[0095] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0096] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0097] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223, etc. The storage method can be temporary storage or permanent storage.

[0098] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, enabling the processor 21 to calculate and process the massive amount of data 223 in the memory 22. It can be Windows, Unix, Linux, etc. In addition to including computer programs capable of implementing the offline and online real-time chorus method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input and output interface 25.

[0099] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the offline and online real-time chorus method steps disclosed in any of the aforementioned embodiments are implemented.

[0100] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0101] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0102] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.

[0103] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0104] The above is a detailed introduction to the offline and online real-time chorus method, device and storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A real-time offline and online chorus method, characterized in that: Applied to cloud servers, including: Pull the real-time audio and video stream in the offline concert scene from the content distribution network by using the pull-stream-relay-push service network, mark the progress of the real-time audio and video stream to obtain the current marking progress, and push the real-time audio and video stream and the current marking progress to the real-time audio and video service network; The real-time audio and video stream and the current marking progress are sent to the target chorus user terminal by using the real-time audio and video service network, so that the target chorus user terminal plays the real-time audio and video stream, and the human voice audio stream acquired in real time by the local online chorus application is marked by using the current marking progress to obtain the marked human voice audio stream carrying the progress information; The marked vocal audio stream uploaded by the target chorus user terminal is obtained through the real-time audio and video service network, and the real-time audio and video stream, the current marking progress and the marked vocal audio stream are pulled to the cloud-based merging service network using the real-time audio and video streaming service network, so as to merge the real-time audio and video stream and the marked vocal audio stream based on the current marking progress and the progress information using the cloud-based merging service network, and push the merged audio and video stream to the content distribution network for distribution.

2. The offline and online real-time chorus method according to claim 1, characterized in that: The method of using the pull-stream-relay-pushing service network to pull the real-time audio and video stream in the offline singing scene from the content distribution network, and marking the progress of the real-time audio and video stream to obtain the current marking progress, includes: When the current time interval reaches a preset time interval threshold, the real-time audio and video stream corresponding to the preset time interval threshold is pulled from the current target audio and video stream by using the pull-stream forwarding service network based on the order of timestamps from front to back; the current target audio and video stream is the audio and video stream that is not currently pulled and corresponds to the offline singing scene in the content distribution network; A current progress sequence number corresponding to the current stream pulling operation is determined, and the real-time audio and video stream is progress-marked using the current progress sequence number to obtain a current marking progress.

3. The offline and online real-time chorus method according to claim 2, characterized in that: When the current time interval reaches the preset time interval threshold, the real-time audio and video stream corresponding to the preset time interval threshold is pulled from the current target audio and video stream by using the stream pull and push service network based on the order of timestamps from front to back, including: The target stream pulling timestamp corresponding to each stream pulling operation is determined based on the first stream pulling timestamp and the preset time interval threshold; wherein the time when the stream pulling and forwarding service network first pulls the real-time audio and video stream in the offline concert scene from the content distribution network is used as the first stream pulling timestamp; Whenever the current time reaches the target stream pulling timestamp, the stream pulling and pushing service network is used to pull the real-time audio and video stream corresponding to the preset time interval threshold from the current target audio and video stream based on the order of timestamps from front to back.

4. The offline and online real-time chorus method according to claim 2, characterized in that: The determining of the current progress sequence number corresponding to the current stream pulling operation, and using the current progress sequence number to perform progress marking on the real-time audio and video stream to obtain the current marking progress, includes: Determine the current progress sequence number corresponding to the current stream pulling operation based on the historical marking progress corresponding to the previous stream pulling timestamp and the preset time interval threshold; wherein, if the current stream pulling operation is the first stream pulling operation, the current progress sequence number is a preset initial value; The current progress sequence number is used as the current marking progress of the real-time audio and video stream.

5. The offline and online real-time chorus method according to claim 1, characterized in that: The method of using the pull-stream-relay-pushing service network to pull the real-time audio and video stream in the offline singing scene from the content distribution network, and marking the progress of the real-time audio and video stream to obtain the current marking progress, includes: The pull-stream-pushing service network is used to continuously pull real-time audio and video streams in offline singing scenes from the content distribution network, and when the current time interval reaches a preset time interval threshold, the real-time audio and video stream is progress-marked using a preset progress sequence number to obtain the current marking progress.

6. The offline and online real-time chorus method according to claim 1, characterized in that: Also includes: Obtain online vocal audio streams uploaded by other chorus user terminals; The real-time audio and video service network is used to forward the online vocal audio stream uploaded by the other chorus user terminals to the target chorus user terminal, so that the target chorus user terminal can simultaneously play the real-time audio and video stream and the online vocal audio stream uploaded by the other chorus user terminals.

7. The offline and online real-time chorus method according to any one of claims 1 to 6, characterized in that: The step of pushing the combined audio and video stream to the content distribution network for distribution includes: The merged audio and video stream is pushed to each content distribution network node of the content distribution network, so that the content distribution network node sends the merged audio and video stream to the online audience user terminal after receiving the data pull request sent by the online audience user terminal; wherein the online audience user terminal selects one content distribution network node from each content distribution network node based on the current network communication status and sends the data pull request to it.

8. A real-time offline and online chorus method, characterized in that: Applied to target chorus user terminals, including: Pull and play the real-time audio and video stream in the offline singing scene from the real-time audio and video service network in the cloud server, and obtain the current marking progress of the real-time audio and video stream; wherein the cloud server uses the pull-stream-pushing service network to pull the real-time audio and video stream from the content distribution network, and performs progress marking on the real-time audio and video stream to obtain the current marking progress; Using the current marking progress, marking the human voice audio stream acquired in real time by the local online chorus application, to obtain a marked human voice audio stream carrying progress information; The marked human voice audio stream is uploaded to the real-time audio and video service network, so that the real-time audio and video streaming service network can send the real-time audio and video stream, the current marking progress and the marked human voice audio stream pulled from the real-time audio and video service network to the cloud-based merging service network, and merge the real-time audio and video stream and the marked human voice audio stream based on the current marking progress and the progress information through the cloud-based merging service network, and push the merged audio and video stream to the content distribution network for distribution.

9. The offline and online real-time chorus method according to claim 8, characterized in that: The step of marking the human voice audio stream acquired in real time by the local online chorus application using the current marking progress includes: Opening a local audio acquisition device through a local online chorus application to collect local audio data through the local audio acquisition device, and performing denoising on the local audio data to obtain a human voice audio stream; The vocal audio stream obtained through the local online chorus application is marked using the current marking progress.

10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the offline and online real-time chorus method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the offline and online real-time chorus method as described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Chorusing method and system for online vocal concert

    CN105208039A

  • Live broadcast microphone connection method and device, terminal, server and storage medium

    CN113271470A

  • Live broadcast chorus interaction method, system and device and computer equipment

    CN114125480A

  • Online and offline real-time chorus method, equipment and medium

    CN117640987A

  • Method, device and system for synchronously playing message stream and audio-video stream

    US20200322670A1