Methods, apparatus, devices, media and products for synchronizing audio and video streams

CN122824929APending Publication Date: 2026-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353838.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0010]应当理解,该内容部分中所描述的内容并非旨在限定本公开的实施例的关键或重要特征,亦非用于限制本公开的范围。本公开的其它特征将通过以下的描述变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824929A_ABST
    Figure CN122824929A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, apparatus, device, medium and product for synchronizing an audio stream and a video stream. The method comprises determining a target audio frame in the audio stream to be transmitted. The method further comprises determining a set of time durations corresponding to a set of historical video frames in the video stream, the set of historical video frames being generated based on a set of historical audio frames in the audio stream. The method further comprises determining a delay time duration for the target audio frame based on the set of time durations. The method further comprises keeping the audio stream and the video stream synchronized by delaying the target audio frame by the delay time duration. By the method, the audio stream and the video stream are synchronized, the consistency of information is ensured, the interactive effect is optimized, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein generally relate to the field of audio and video processing, and specifically to methods, apparatuses, devices, media, and products for synchronizing audio streams and video streams. Background Technology

[0002] Currently, with the rapid development of real-time communication technology and multimedia streaming technology, there is a growing number of application scenarios where audio and video are processed separately at different source ends, merged, and then pushed to the audience. For example, in live streaming, virtual meetings, and remote interaction, video and audio usually come from different production ends. Audio may be uploaded by participants through real-time communication, while video may be generated through cloud rendering or remote acquisition.

[0003] In the process of audio and video processing, the links of audio and video processing may be different, but the data through these links can be integrated into a complete data stream and provided to the user. As a result, the user can see synchronized and complementary multimedia content through user terminal devices such as mobile phones, computers, and smart TVs, and obtain a comprehensive audio-visual experience. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, device, medium, and product for synchronizing audio and video streams.

[0005] According to a first aspect of this disclosure, a method for synchronizing an audio stream and a video stream is provided. The method includes determining a target audio frame in an audio stream to be transmitted. The method further includes determining a set of durations corresponding to a set of historical video frames in a video stream, the set of historical video frames being generated based on the set of historical audio frames in the audio stream. The method also includes determining a delay duration for the target audio frame based on the set of durations. The method further includes maintaining synchronization between the audio stream and the video stream by delaying the target audio frame by the delay duration.

[0006] According to a second aspect of this disclosure, an apparatus for synchronizing audio and video streams is provided. The apparatus includes a target audio frame determination module configured to determine a target audio frame in an audio stream to be transmitted; a set of duration determination module configured to determine a set of durations corresponding to a set of historical video frames in a video stream, the set of historical video frames being generated based on the set of historical audio frames in the audio stream; a delay duration determination module configured to determine a delay duration for the target audio frame based on the set of durations; and a synchronization module configured to maintain synchronization between the audio and video streams by delaying the target audio frame by the delay duration.

[0007] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this disclosure.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0010] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0012] Figure 1 The illustration shows a schematic diagram of an example environment in which some embodiments of the present disclosure may be implemented;

[0013] Figure 2 The illustration shows a schematic diagram of an example method for synchronizing audio and video streams according to some embodiments of the present disclosure;

[0014] Figure 3 The illustration shows a schematic diagram of an example method for updating duration according to some embodiments of the present disclosure;

[0015] Figure 4 The illustration shows a schematic diagram of an example method for a lip-sync generation chain according to some embodiments of the present disclosure;

[0016] Figure 5 The illustration shows another example method of a lip-sync generation chain according to some embodiments of the present disclosure;

[0017] Figure 6 The illustration shows a schematic block diagram of an apparatus for synchronizing audio and video streams according to some embodiments of the present disclosure;

[0018] Figure 7A schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure is illustrated. Detailed Implementation

[0019] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0021] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0022] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0023] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0026] In some current live streaming rooms where the audio and video processing chains are inconsistent, the audio comes from users who join the room in real time, such as the host or guests, while the video comes from cloud-rendered footage. The cloud-rendered footage relies on data obtained from audio processing on the client's local machine. While the host's or guests' audio stream is pushed to the audience in real time, the video stream experiences a delay after cloud rendering. This causes misalignment between the virtual character's lip movements and the audio, disrupting the audio-visual synchronization and resulting in a poor user experience.

[0027] For example, in a cloud gaming virtual character livestream scenario, the audio and video production processes are inconsistent. The audio streams from the host and guests can be pushed to the audience in real time, so the audio heard by the audience is, to a certain extent, latency-free. However, the lip movements of the virtual character generated in the cloud need to be rendered into corresponding visuals that match the audio, which takes time in the cloud. This causes the audio and video streams to be out of sync on the audience's end, thus significantly impacting the viewing experience.

[0028] Therefore, embodiments of this disclosure propose a method for synchronizing audio and video streams. In this method, a computing device determines a target audio frame in the audio stream to be transmitted. Then, the computing device can determine a set of durations corresponding to a set of historical video frames in the audio stream for a set of historical audio frames. Next, based on the determined durations of the historical video frames, a delay duration for the target audio frame is determined. Finally, the computing device delays the target audio frame by the calculated delay duration to keep the audio and video streams synchronized. By delaying the real-time audio data packets by a certain time, the arrival time of the audio data packets and video data packets at the user's viewing end is consistent, resulting in synchronized video and audio, achieving the expected improvement goal and enhancing the user experience.

[0029] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 The illustration shows an example environment in which the devices and / or methods of embodiments of the present disclosure may be implemented. In environment 100, computing device 102 may synchronize audio streams and video streams.

[0030] Examples of computing device 102 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0031] like Figure 1 The computing device 102 can acquire the audio stream 104. In scenarios where the audio stream 104 needs to be synchronized with the video stream 110, the computing device 102 can first determine the target audio frame 106 in the audio stream 104 to be sent. At this time, the computing device 102 will first obtain the required target audio frame in the audio stream 104 by collecting the user's voice.

[0032] After acquiring locally collected audio data, computing device 102 processes the audio data on the device to generate lip-sync data that can be used for rendering, such as lip-sync identifier data. Then, it sends the lip-sync data corresponding to the audio frames to the target server. Upon receiving the lip-sync data, the corresponding process on the target server generates a video stream 110 with lip-sync visuals corresponding to the audio frames. This video stream is then returned to computing device 102 or other devices via inter-process communication for transmission. For example, when live-streaming cloud gaming, the cloud gaming server can generate this video stream, which includes virtual characters whose lip-sync is driven by the audio stream, for example, determined by the lip-sync data corresponding to each audio frame in the audio stream. In one example, when computing device 102 is a device on the guest side of the live stream, the audio stream is from the guest. Therefore, audio stream 104 can also be provided to the live stream end to be merged and transmitted together with video stream 110. In another example, computing device 102 can be a device on the broadcaster side, in which case the audio stream can be from the broadcaster.

[0033] When sending an audio stream, computing device 102 can determine a set of durations 112 corresponding to a set of historical video frames 116 in video stream 110, wherein the set of historical video frames 116 is generated based on a set of historical audio frames in the audio stream. When determining each duration in this set, computing device 102 first determines a first timestamp of the lip-sync data corresponding to the historical video frame to be sent. For example, after determining the corresponding lip-sync data by processing the historical audio frames, the lip-sync data can be transmitted to a cloud gaming server, and the timestamp of the uploaded lip-sync data can be recorded as the first timestamp. This lip-sync data, after being uploaded to the server, can be used to generate lip-sync images in historical video frames in video stream 110, and then the lip-sync images are combined with the historical video frames and returned to computing device 102. Therefore, computing device 102 can further determine a second timestamp of the historical video frame receiving the lip-sync data. Finally, based on the first and second timestamps, computing device determines the duration corresponding to a historical video frame.

[0034] The difference between the second timestamp and the first timestamp represents the network round-trip time for the lip-sync data. For example, if the first timestamp of lip-sync data sent from the computing device is T1, and the second timestamp of the historical video frame containing the corresponding lip shape is T2, then the corresponding duration is T2-T1. Furthermore, this duration corresponding to the historical audio frame corresponds to the sum of the upload time of the lip-sync data for the historical audio frame, the generation time of the lip-sync image portion corresponding to the lip-sync data, and the downlink time of the historical video frame.

[0035] Through the above process, the temporal relationship between audio and video can be accurately determined, thereby enabling synchronous analysis of audio and video and providing higher precision control for the alignment of subsequent audio and video streams.

[0036] Subsequently, the computing device 102 determines the delay duration 108 for the target audio frame 106 based on a set of durations 112 of the aforementioned set of historical video frames 116. The durations within the set of durations 112 of the historical audio frames are not necessarily identical; the variations in these duration values ​​are influenced by the real-time round-trip latency (RTT) under network conditions. Higher RTT results in longer corresponding durations within the set of durations 112 of the historical audio frames. Therefore, to ensure the audio delay duration better reflects most results and reduces errors, the average duration within the set of durations 112 is first calculated when determining the delay duration 108 for the target audio frame 106, and then this average duration is used to determine the delay duration 108 for the target audio frame 106.

[0037] In some embodiments, when the computing device 102 is the device of a guest in a live broadcast, when the computing device 102 transmits audio data to the broadcaster's device, the computing device 102 can use the average duration plus the local computation time for generating lip-sync data for the audio frame, and then subtract the real-time communication time to the broadcaster's end to obtain the delay duration 108 for the audio frame. The local computation time and the real-time communication time can be predetermined or calculated according to certain rules.

[0038] In some embodiments, when the computing device 102 is the broadcaster's device in a live stream, it can also determine the delay duration 108 to be sent to the guest devices using the average duration. Additionally, the computing device 102 generates a combined audio stream for multiple users based on the audio stream including the target audio frame 106. These multiple users include guests participating in the audio input. Simultaneously, the computing device delays the combined audio stream by the calculated delay duration to obtain a delayed combined audio stream. After obtaining the combined audio stream, the computing device combines the corresponding audio stream with the video stream and sends it to the client for user viewing.

[0039] This set of durations can be updated over time, and a period for updating this set of durations is set for this purpose. If the period for updating a set of durations expires, the computing device 102 will continue to determine the second duration of the second video frame in the video stream, where the second video frame is generated based on the second audio frame in the audio stream. After determining the second duration of the video frame, the second duration will also be used to update a set of durations in the queue.

[0040] For example, when updating the duration of this set of times, to ensure the real-time nature and effectiveness of the audio latency, the computing device also maintains a queue storing this set of times to adjust the audio latency. Based on the first-in, first-out (FIFO) principle of the queue, the queue retains the round-trip latency results of the most recent threshold number of times, and dequeues latency times earlier than the threshold number. For example, after obtaining the latest round-trip latency, the earliest round-trip latency is dequeued, and then the average round-trip latency in the queue is used to determine the audio packet transmission latency.

[0041] Finally, the computing device 102 maintains synchronization between the audio and video streams at frame 114 by delaying the target audio frame 106 by calculating the delay duration. In one example, the delay handling process can differ for guests and hosts in a live stream. Since the guest's primary task is to send their own audio stream for merging, the guest's audio-video synchronization requirement is specific to their own real-time communication audio processing, ensuring that the locally emitted audio is aligned at the merging node. Therefore, the guest doesn't need to pay attention to other users' audio, so when a user is a guest, their real-time audio needs to be delayed. When a user is a host, they not only need to handle their own real-time communication audio-video synchronization but also ensure overall alignment of all audio and video streams at the merging node. The host's audio participates in the merging to generate the final mixed stream. To ensure synchronized audio and video processing for multiple users, the delay of the merging audio needs to be adjusted to ensure that all users' audio and video are synchronized in the merging result. At this point, the guest's audio and video are aligned with the audience's audio and video, providing a good viewing experience for both the audience and the guest.

[0042] The above combination Figure 1 The following is a schematic diagram illustrating an example environment in which some embodiments of this disclosure may be implemented, in conjunction with... Figure 2 A schematic diagram illustrating example methods for synchronizing audio and video streams according to some embodiments of the present disclosure. Figure 2 The method in can be derived from Figure 1 The computing device 102 or any suitable computing device in the system shall execute the test.

[0043] like Figure 2In example method 200, at block 202, computing device 102 determines the target audio frame in the audio stream to be sent. When computing device 102 acquires user audio signals, the user's audio is converted into an analog signal and then converted into a digital signal by the audio processing module of computing device 102. The continuous audio data is then divided into multiple fixed-length blocks, for example, each frame contains 10 milliseconds of audio data, and each block is one frame. In one embodiment, the computing device generates audio frame F1 at 100ms and audio frame F2 at 110ms. If the device is preparing to send audio frame F2, then F2 is the current target audio frame for the next operation, such as delaying for a specified duration.

[0044] At box 204, computing device 102 determines a set of durations corresponding to a set of historical video frames in the video stream. This set of historical video frames is generated based on a set of historical video frames in the audio stream. To achieve synchronization between the audio and video streams, a set of durations needs to be determined through the correspondence between historical audio and video frames. When processing audio frames, the computing device needs to extract feature information of the user's speech, such as the lip-sync data of the corresponding audio frame. Based on these speech features, it generates relevant data to describe the lip-sync and records the upload time of this data as a first timestamp. When the computing device receives a historical video frame corresponding to that lip-sync in the video stream, it records the time when the video frame was successfully received or processed as a second timestamp. The time difference between the first and second timestamps is calculated to determine the duration corresponding to the historical video frame.

[0045] Furthermore, the duration corresponding to historical video frames corresponds to the sum of the upload duration of lip-sync data, the generation duration of the lip-sync image portion of the lip-sync data, and the downlink duration of the historical video frames. The sum of the upload duration of the lip-sync data of historical audio frames, the generation duration of the lip-sync image portion of the lip-sync data, and the downlink duration of historical video frames represents the lip-sync generation link time, which dynamically changes in response to the Round-Trip Time (RTT). For example, in one embodiment, local audio frame processing takes 20ms, the network latency during upload is 50ms, video frame generation takes 100ms, and video frame transmission takes 40ms, resulting in a total duration of 210ms. Therefore, the lip-sync generation link time is 190ms. When the Real-Time Communications (RTC) transmission time is a fixed 10ms, the audio latency is 200ms. After adjustment, the audio stream and video stream are approximately aligned at the viewer's end.

[0046] At box 206, computing device 102 determines the latency for a target audio frame based on a set of durations. To ensure synchronization between the audio and video streams, the latency needs to be calculated for each target audio frame. The latency is determined by dynamically adjusting the latency based on analysis of historical video and audio frame duration sets to adapt to network fluctuations and processing delays.

[0047] After the computing device 102 obtains a duration queue of historical audio or video from multiple past frames, it calculates the average duration using the duration queue and then uses the obtained average duration to determine the audio delay duration. The obtained target video frame contains lip-sync information corresponding to the speech content of the target audio frame, ensuring that the lip movements in the video match the audio content. By calculating the average duration based on historical data, the delay time of the target audio frame can be dynamically adjusted to adapt to different network and processing environments.

[0048] When the computing device is the broadcasting client, in the case of multiple users' audio input, a combined audio stream needs to be generated for transmission to other viewers. This combined audio stream is generated by superimposing or mixing the audio data streams of multiple users in chronological order. When outputting the combined audio stream, a specified delay can be applied based on a calculated latency to maintain synchronization with the video. In one embodiment, in a live broadcast environment containing users A, B, and C, the computing device first acquires the audio streams of all three users and synthesizes them into a combined audio stream that includes all users' inputs. Then, the latency is calculated and applied to synchronize the audio and video, and the users' real-time video frames are combined into a video stream and synchronously sent to each client along with the combined audio stream. This ensures audio-video synchronization during client playback.

[0049] Additionally, the computing device 102 can use the extracted target audio frame to generate lip-sync data that matches the target audio frame, and then upload it to the cloud or server to generate the target video frame. Data processing and synchronization are achieved through interaction with the cloud or server. The server generates the target video frame based on the target lip-sync identifier data, including dynamic matching of mouth movements. The broadcaster receives the target video frame and synchronizes it with the target audio frame.

[0050] At box 208, the audio and video streams are synchronized by calculating the delay duration obtained from the target audio frame delay. During the synchronization process, dynamic updates are needed to adapt to changes in network conditions and processing performance. This dynamic update can also be achieved by updating a set of durations, as described below. Figure 3 The description is as follows.

[0051] The above methods can achieve accurate synchronization of audio and video streams in multi-user scenarios and dynamically adapt to network duration fluctuations and performance changes, providing users with a high-quality audio and video synchronization experience.

[0052] The above combination Figure 2 A schematic diagram illustrating example methods for synchronizing audio and video streams according to some embodiments of this disclosure is provided; the set of durations described above can be updated periodically. The following is in conjunction with... Figure 3 A schematic diagram illustrating an example method for updating duration according to some embodiments of the present disclosure. Figure 3 The method in can be derived from Figure 1 The computing device 102 or any suitable computing device in the system shall execute the test.

[0053] In Example 300, at box 302, computing device 102 determines whether a period for updating a set of durations has expired. For example, the system periodically checks whether a set of durations needs to be updated; this period can be adjusted manually or automatically in the background. For example, the period can be set to 180 seconds.

[0054] At box 304, computing device 102, in response to the expiration of a period, determines a second duration of a second video frame in the video stream, the second video frame being generated based on a second audio frame in the audio stream. After the set update period expires, a second duration associated with the audio frame and the generated video frame at that time can be further determined.

[0055] At box 306, computing device 102 updates a set of durations based on a second duration. After obtaining the second duration, computing device 102 can update the set of durations according to the second duration.

[0056] When updating a set of durations using the second duration, the computing device 102 can determine whether the number of durations in the set is less than a threshold number. For example, if the threshold number of durations in a set is 10, meaning a set of durations can hold a maximum of 10 durations, the computing device 102 can determine the number of durations in the set at this time and then compare that number with the threshold number.

[0057] If the number of durations in a set is less than the threshold number, the computing device 102 adds the second duration to the set. For example, if a set already has 8 durations, which is less than the threshold number of 10, the obtained duration can be directly added to the set.

[0058] If the computing device 102 determines that the number of durations in a set of durations is greater than or equal to a threshold number, it indicates that the number of durations in the set of data has met the requirement. At this point, the computing device 102 removes the earliest acquired duration from the set of durations, such as the earliest duration to enter the set of data. Then, the computing device 102 adds this second duration to the set of durations. Therefore, by updating the duration queue, the new video frame delay duration can be calculated.

[0059] For example, the duration of this group can be placed into a dynamically maintained RTT queue based on the first-in-first-out design principle, retaining the most recent N RTT measurement results; and the audio packet transmission delay of users using virtual characters can be dynamically updated with time T as the period; in the business targeting virtual characters, N can be set to 10 according to the evaluation, at which time the queue is rttQueue=[RTT1,RTT2,...,RTTn], and can be updated periodically.

[0060] The above combination Figure 3 A schematic diagram illustrating an example method for updating duration according to some embodiments of this disclosure is described below; in conjunction with... Figure 4 A schematic diagram illustrating an example process of a lip-sync generation link according to some embodiments of the present disclosure.

[0061] In Example 400, there are a cloud gaming server 402 and a live streaming room 404. The audio emitted by guest 412 in the live streaming RTC room 410 of live streaming room 404 is captured by a computing device. After local lip-sync data processing, the audio data is converted into lip-sync data suitable for rendering and then uploaded to the cloud gaming RTC room 408. At this point, the lip-sync data transmitted to the cloud gaming RTC room 408 is lip-sync identifier data, which is based on vowel mapping. The lip-sync data in the cloud gaming room generates lip-sync visuals in the game process 406 via inter-process communication and is then transmitted to the cloud gaming RTC room 408 via inter-process communication. The cloud gaming RTC room 408 pushes a video stream containing the lip-sync data to each guest in the live streaming RTC room 410 of live streaming room 404, namely guest 412, host 414, and guest 416 in the diagram. Simultaneously, the audio captured at the audio acquisition point is also delayed using the calculated audio delay duration. When the guest is a virtual character, only the RTC audio needs to be delayed after the audio capture. When the host is a virtual character, the host needs to delay all RTC audio and merged audio.

[0062] Therefore, as shown above, the lip-sync generation chain requires a packet delay for the RTC audio of the user using the virtual character to achieve alignment between the lip movements and the audio. The audio packet delay time plus the RTC audio transmission time should be approximately equal to the local lip-sync data processing time plus the lip-sync data uploading time plus the lip-sync image generation time plus the cloud gaming video downlink time, achieving near-alignment between the lip movements and audio. The RTC audio transmission time can be used as a fixed constant based on the average transmission time of the entire live streaming platform, while the lip-sync generation chain time needs to be dynamically calculated.

[0063] The above combination Figure 4 A schematic diagram illustrating an example method for a lip-sync generation chain according to some embodiments of the present disclosure is described below; in conjunction with... Figure 5 A schematic diagram illustrating another example process for generating link round-trip delay according to some embodiments of the present disclosure.

[0064] In Example 500, there is a cloud gaming server 502 and a live streaming room 504. Guest 512 in the live streaming RTC room 510 emits sound, which is first captured by the computing device in step 1 and then locally processed in step 2. In step 3, the lip-sync data from the processed data is uploaded to the cloud gaming server 502, for example, to the cloud gaming RTC room 508 of the cloud gaming server 502, and the current uplink timestamp t1 of the lip-sync data is recorded in step 3. Then, through inter-process communication with the cloud gaming RTC room 508, the lip-sync image is generated in step 4 in the game process 506. The video image with the lip-sync image is then transmitted to the cloud gaming RTC room 508 via inter-process communication. In step 5, the generated cloud gaming video stream is downlinked and supplementary enhancement information is added, including the timestamp information from step 3. Finally, the video stream carrying the lip-sync image is returned to guest 512, and the current timestamp t2 is recorded. Therefore, the round-trip time (RTT) value for this frame is t2-t1. The round-trip delay value is then placed in the RTT queue to prepare for the calculation and dynamic maintenance of the average round-trip delay value in the future.

[0065] Figure 6 The illustration shows a schematic block diagram of an apparatus for synchronizing audio and video streams according to some embodiments of the present disclosure. Figure 6 As shown, the apparatus 600 includes a target audio frame determination module 602, configured to determine a target audio frame in an audio stream to be transmitted; a set duration determination module 604, configured to determine a set duration corresponding to a set of historical video frames in a video stream, the set of historical video frames being generated based on a set of historical audio frames in the audio stream; a delay duration determination module 606, configured to determine a delay duration for the target audio frame based on the set duration; and a synchronization module 608, configured to keep the audio stream and video stream synchronized by delaying the target audio frame by a delay duration.

[0066] In some embodiments, the target audio frame determination module 602 includes: an acquisition module configured to acquire target audio frames in an audio stream by acquiring audio from the user's speech.

[0067] In some embodiments, a set of duration determination modules 604 includes: a first timestamp determination module configured to determine a first timestamp for transmitting lip-sync data corresponding to a historical audio frame in a set of historical audio frames; a first timestamp determination module configured to determine a second timestamp for receiving a historical video frame in a set of historical video frames that has lip-sync data corresponding to the lip-sync data; and a duration determination module configured to determine the duration corresponding to the historical video frame based on the first timestamp and the second timestamp.

[0068] In some embodiments, the duration determination module includes: a difference determination module configured to determine a timestamp difference between a second timestamp and a first timestamp; and a first duration determination module configured to determine the timestamp difference as a duration corresponding to a historical video frame.

[0069] In some embodiments, the duration corresponds to the sum of the upload duration of the lip-sync data, the generation duration of the lip-sync image portion corresponding to the lip-sync data, and the downlink duration of the historical video frames.

[0070] In some embodiments, the delay duration determination module 606 includes: an average duration determination module configured to determine an average duration of a set of durations based on a set of durations; and a first delay duration determination module configured to determine the average duration as the delay duration for a target audio frame.

[0071] In some embodiments, the target video frame includes lip movements corresponding to the target audio frame.

[0072] In some embodiments, the apparatus 600 further includes: a combined audio stream generation module configured to generate a combined audio stream for multiple users based on an audio stream including target audio frames; a delayed audio stream module configured to obtain a delayed combined audio stream by delaying the combined audio stream by a certain duration; and a sending module configured to combine the delayed combined audio stream and a combined video stream formed by the video stream to send to a client.

[0073] In some embodiments, the apparatus 600 further includes: a lip-shape generation module configured to generate target lip-shape data corresponding to a target audio frame based on the target audio frame; a lip-shape data transmission module configured to transmit the target lip-shape data to a target server; and a video frame receiving module configured to receive a target video frame from the target server that includes lip shapes corresponding to the target lip-shape data.

[0074] In some embodiments, the target audio frame is a first audio frame, and the apparatus 600 further includes: a period detection module configured to determine whether a period for updating a set of durations has expired; a second duration determination module configured to determine a second duration of a second video frame in a video stream in response to the expiration of the period, the second video frame being generated based on the second audio frame in an audio stream; and an update module configured to update a set of durations based on the second duration.

[0075] In some embodiments, the update module includes: a number determination module configured to determine whether the number of durations in a set of durations is less than a threshold number; and a first addition module configured to add a second duration to a set of durations in response to the number of durations in a set of durations being less than the threshold number.

[0076] In some embodiments, the update module further includes: a removal module configured to remove the earliest obtained duration in a set of durations in response to determining that the number of durations in a set of durations is greater than or equal to a threshold number; and a second addition module configured to add the second duration to a set of durations.

[0077] In some embodiments, the audio stream comes from a streamer or guest performing a live cloud game, the video stream is the video stream of the cloud game, and the lip movements of the virtual characters in the cloud game are driven by the audio stream.

[0078] Figure 7 A schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure is shown. Figure 1 The computing device 102 can be implemented using device 700. As shown, device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. The RAM 703 can also store various programs and data required for the operation of device 700. The CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0079] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage page 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0080] The various processes and procedures described above, such as methods 200 and 300, can be executed by processing unit 701. For example, in some embodiments, methods 200 and 300 can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of the example methods 200 and 300 described above can be performed.

[0081] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0082] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0083] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0084] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0085] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0086] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0087] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0089] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for synchronizing audio streams and video streams, comprising: Determine the target audio frame in the audio stream to be sent; Determine a set of durations corresponding to a set of historical video frames in the video stream, the set of historical video frames being generated based on a set of historical audio frames in the audio stream; Based on the set of durations, determine the delay duration for the target audio frame; as well as The audio stream and the video stream are kept synchronized by delaying the target audio frame by the specified delay duration.

2. The method of claim 1, wherein determining the target audio frame in the audio stream to be transmitted comprises: The target audio frame in the audio stream is obtained by capturing the user's voice.

3. The method of claim 1, wherein determining a set of durations corresponding to a set of historical video frames in the video stream comprises: Determine a first timestamp for sending lip-sync data corresponding to a historical audio frame in the set of historical audio frames; A second timestamp is determined for the historical video frame in the set of historical video frames that has a lip shape corresponding to the lip shape data; as well as Based on the first timestamp and the second timestamp, the duration corresponding to the historical video frame is determined.

4. The method according to claim 3, wherein determining the duration corresponding to the historical video frame based on the first timestamp and the second timestamp includes: Determine the timestamp difference between the second timestamp and the first timestamp; as well as The timestamp difference is determined as the duration corresponding to the historical video frame.

5. The method according to claim 4, wherein the duration corresponds to the sum of the upload duration of the lip-sync data, the generation duration of the lip-sync image portion corresponding to the lip-sync data, and the downlink duration of the historical video frames.

6. The method of claim 1, wherein determining the delay duration for the target audio frame based on the set of durations comprises: Based on the set of durations, determine the average duration of the set of durations; as well as The average duration is determined as the delay duration for the target audio frame.

7. The method according to claim 1, wherein the target video frame includes lip movements corresponding to the target audio frame.

8. The method according to claim 1, further comprising: A combined audio stream for multiple users is generated based on the audio stream including the target audio frame; The delayed combined audio stream is obtained by extending the combined audio stream by the delay duration; as well as The delayed combined audio stream and the combined video stream formed by the video stream are combined and sent to the client.

9. The method according to claim 1, further comprising: Based on the target audio frame, generate target lip shape data corresponding to the target audio frame; Send the target lip shape data to the target server; as well as Receive the target video frame from the target server, which includes the lip shape corresponding to the target lip shape data.

10. The method of claim 1, wherein the target audio frame is a first audio frame, and the method further comprises: Determine whether the period used to update the set of durations has expired; In response to the expiration of the period, a second duration of a second video frame in the video stream is determined, the second video frame being generated based on a second audio frame in the audio stream; as well as Update the set of durations based on the second duration.

11. The method of claim 10, wherein updating the set of durations based on the second duration comprises: Determine whether the number of durations in the set of durations is less than a threshold number; as well as In response to the fact that the number of durations in the set of durations is less than a threshold number, the second duration is added to the set of durations.

12. The method of claim 11, wherein updating the set of durations based on the second duration further comprises: In response to determining that the number of durations in the set of durations is greater than or equal to a threshold number, the earliest duration obtained in the set of durations is removed; as well as Add the second duration to the set of durations.

13. The method of claim 1, wherein the audio stream originates from a streamer or guest performing a live broadcast of a cloud game, the video stream is the video stream of the cloud game, and the lip movements of the virtual characters in the cloud game are driven by the audio stream.

14. An apparatus for synchronizing audio streams and video streams, comprising: The target audio frame determination module is configured to determine the target audio frame in the audio stream to be sent. A set of duration determination modules is configured to determine a set of durations corresponding to a set of historical video frames in a video stream, the set of historical video frames being generated based on a set of historical audio frames in an audio stream; The delay duration determination module is configured to determine the delay duration for the target audio frame based on the set of durations; as well as A synchronization module is configured to keep the audio stream and the video stream synchronized by delaying the target audio frame by the delay duration.

15. An electronic device comprising: At least one processor; as well as A storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, the computer program implementing the method according to any one of claims 1-13 when executed by a processor.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-13.