Audio and video processing method and device and electronic equipment
By adjusting the starting timestamp and global offset processing of audio and video clips, the audio and video out-synchronization problem caused by insufficient timestamp offset and synchronization mechanism in real-time audio and video processing services is solved, and the precise synchronization of audio and video streams and high-quality playback are achieved.
Patent Information
- Application Number
- CN202510547359.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-11
AI Technical Summary
The problem of audio and video dissynchronization due to time stamp offset and insufficient synchronization mechanism in the prior art in real-time audio and video processing services is difficult to effectively solve especially in complex network environments.
By receiving audio and video clips, adjusting the start timestamp of sub-audio and video clips, determining the target time, and performing global offset processing to correct the time deviation and ensure that the audio and video streams are played consistently and synchronously under the global time reference.
It realizes accurate synchronization of audio and video streams, significantly improves the playback quality of real-time audio and video services, solves the problem of audio and video out-of-synchronization, and enhances the applicability and flexibility of the system, especially in scenarios with high requirements for audio and video synchronization.
Smart Images

Figure CN120302101A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio - video processing and transmission. Specifically, it relates to an audio - video processing method, apparatus, and electronic device. Background Art
[0002] In the modern digital communication and entertainment fields, real - time audio - video processing services have become an indispensable part and are widely used in multiple scenarios such as online education, remote conferencing, and live interactive sessions. With the popularization of high - quality media formats such as high - definition video and surround sound, users' requirements for audio - video synchronization are becoming increasingly strict. However, when dealing with real - time audio - video streams, existing technologies often encounter the challenge of audio - video asynchronization, which is mainly due to timestamp offsets and limitations of synchronization mechanisms.
[0003] Timestamp offsets result from the tiny time differences generated during the acquisition, encoding, transmission, etc. of audio - video data. Especially, the processing speeds of audio and video data are inconsistent, leading to timestamp misalignment at the playback end, thus affecting the synchronization effect. In addition, the design of synchronization mechanisms often focuses on one direction, such as relying solely on timestamps or caching strategies. In a complex and variable network environment, these mechanisms are difficult to fully handle the sudden delays or losses of audio - video data, resulting in more serious audio - video asynchronization.
[0004] Regarding the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This application provides an audio - video processing method, apparatus, and electronic device to at least solve the technical problem of audio - video asynchronization caused by timestamp offsets and insufficient synchronization mechanisms in real - time audio - video processing services in the prior art.
[0006] According to one aspect of this application, an audio - video processing method is provided, including: receiving a target audio - video segment, where the target audio - video segment includes M sub - audio - video segments, and each sub - audio - video segment includes S video frames of information and N audio frames of information, where M, N, and S are all integers greater than or equal to 1; adjusting the start timestamp of each of the M sub - audio - video segments to obtain M adjusted first sub - audio - video segments; determining a target time based on the start timestamp of each first sub - audio - video segment, and adjusting the timestamps of the S video frames of information and N audio frames of information in each first sub - audio - video segment according to the target time to obtain M second sub - audio - video segments, where the target time is the timestamp among the start timestamps corresponding to the M first sub - audio - video segments that is less than a preset threshold; performing global offset processing on the M second sub - audio - video segments, and pushing each processed second sub - audio - video segment, where the global offset processing is used to correct the time deviation generated during the transmission of different sub - audio - video segments.
[0007] Optionally, adjust the start timestamps of each of the M sub-audio-video segments to obtain M adjusted first sub-audio-video segments, including: determining the start timestamp of the first video information in the first sub-audio-video segment among the M sub-audio-video segments, and using the start timestamp of the first video information in the first sub-audio-video segment as the first time; determining the start timestamp of the first audio information in the first sub-audio-video segment among the M sub-audio-video segments, and using the start timestamp of the first audio information in the first sub-audio-video segment as the second time; when the first time is greater than or equal to the second time, using the first time as the reference time, where the reference time is used to adjust the start timestamps of each sub-audio-video segment in the target audio-video segment; when the first time is less than the second time, using the second time as the reference time; adjusting the start timestamps of each sub-audio-video segment based on the reference time to obtain M adjusted first sub-audio-video segments.
[0008] Optionally, adjust the start timestamps of each of the M sub-audio-video segments based on the reference time to obtain M adjusted first sub-audio-video segments, including: taking the absolute difference between the start timestamp of the first video information in each sub-audio-video segment and the reference time as the first offset, to obtain M first offsets; taking the absolute difference between the start timestamp of the first audio information in each sub-audio-video segment and the reference time as the second offset, to obtain M second offsets; adjusting the start timestamp of each sub-audio-video segment according to the corresponding first offset and second offset of each sub-audio-video segment to obtain M adjusted first sub-audio-video segments.
[0009] Optionally, determine a target time based on the start timestamps of each of the M first sub-audio-video segments, and adjust the timestamps of the S-frame video information and N-frame audio information in each of the M first sub-audio-video segments according to the target time to obtain M second sub-audio-video segments, including: determining a target time based on the start timestamps of each of the M first sub-audio-video segments, and detecting whether the target time is negative; if the target time is negative, determining a third offset according to the target time, where the third offset is used to adjust the negative timestamps in each of the M first sub-audio-video segments to positive timestamps; if the target time is not negative, determining the third offset to be 0; adjusting the timestamps of the S-frame video information and N-frame audio information in each of the M first sub-audio-video segments according to the third offset to obtain M second sub-audio-video segments.
[0010] Optionally, before performing global offset processing on the M second sub-audio-video segments and pushing each processed second sub-audio-video segment, the method further includes: constructing J video data packets and K audio data packets according to the M second sub-audio-video segments, where J and K are integers greater than or equal to 1; pushing the J video data packets and K audio data packets into the audio-video queue.
[0011] Optionally, perform global offset processing on the M second sub-audio-video segments, and push each processed second sub-audio-video segment, including: setting a global playback time, where the global playback time is used to represent the starting playback moment of the target audio-video segment; reading J video data packets and K audio data packets composed of the M second sub-audio-video segments from the audio-video queue; determining a third time according to the J video data packets and the K audio data packets, where the third time is the time stamp less than the second preset threshold among the time stamps corresponding to the J video data packets and the K audio data packets; determining a target offset according to the global playback time and the third time, where the target offset is used to represent the time deviation generated when each sub-audio-video segment is transmitted; perform global offset processing on the time stamps corresponding to the J video data packets and the K audio data packets according to the target offset, and push the processed J video data packets and K audio data packets.
[0012] Optionally, pushing the processed J video data packets and K audio data packets includes: if there are time stamps less than or equal to the global playback time in the processed J video data packets and K audio data packets, re-perform global offset processing on the J video data packets and K audio data packets; if the time stamps of the processed J video data packets and K audio data packets are all greater than the global playback time, push the processed J video data packets and K audio data packets.
[0013] Optionally, pushing the processed J video data packets and K audio data packets includes: sorting the time stamps corresponding to the processed J video data packets and K audio data packets to obtain a sorting result; pushing the J video data packets and K audio data packets according to the sorting result.
[0014] According to another aspect of the present application, there is also provided an audio-video processing device, including: a receiving unit, configured to receive a target audio-video segment, where the target audio-video segment includes M sub-audio-video segments, and each sub-audio-video segment includes S-frame video information and N-frame audio information, where M, N, and S are all integers greater than or equal to 1; a first adjustment unit, configured to adjust the start timestamps of each of the M sub-audio-video segments to obtain M adjusted first sub-audio-video segments; a first adjustment unit, configured to determine a target time based on the start timestamps of each of the first sub-audio-video segments, and adjust the timestamps of the S-frame video information and the N-frame audio information in each of the first sub-audio-video segments according to the target time to obtain M second sub-audio-video segments, where the target time is the timestamp among the start timestamps corresponding to the M first sub-audio-video segments that is less than a preset threshold; a pushing unit, configured to perform global offset processing on the M second sub-audio-video segments, and push each of the processed second sub-audio-video segments, where the global offset processing is used to correct the time deviation generated during the transmission of different sub-audio-video segments.
[0015] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned audio-video processing method.
[0016] In the present application, first, a target audio-video segment is received, where the target audio-video segment includes M sub-audio-video segments, and each sub-audio-video segment includes S-frame video information and N-frame audio information, where M, N, and S are all integers greater than or equal to 1. Then, the start timestamps of each of the M sub-audio-video segments are adjusted to obtain M adjusted first sub-audio-video segments. Then, a target time is determined based on the start timestamps of each of the first sub-audio-video segments, and the timestamps of the S-frame video information and the N-frame audio information in each of the first sub-audio-video segments are adjusted according to the target time to obtain M second sub-audio-video segments, where the target time is the timestamp among the start timestamps corresponding to the M first sub-audio-video segments that is less than a preset threshold. Finally, global offset processing is performed on the M second sub-audio-video segments, and each of the processed second sub-audio-video segments is pushed, where the global offset processing is used to correct the time deviation generated during the transmission of different sub-audio-video segments. That is, by means of global time reference unification and dynamic timestamp adjustment, the purpose of ensuring the temporal consistency and synchronous playback of the audio-video stream under the global time reference is achieved, thereby realizing the precise synchronization of the audio-video stream and significantly improving the playback quality of the real-time audio-video service, and further solving the technical problem of audio-visual asynchrony caused by timestamp offset and insufficient synchronization mechanism in the existing real-time audio-video processing service. Description of the Drawings
[0017] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is a flowchart of an optional audio - video processing method according to an embodiment of the present application;
[0019] Figure 2 is a schematic diagram of an optional audio - video processing method according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of an optional audio - video processing apparatus according to an embodiment of the present application. Detailed implementation manners
[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above - mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0023] It should also be noted that the information and data collected in this application are information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant regions, necessary confidentiality measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject. For example, an interface is set between this system and relevant users or institutions. Before obtaining relevant information, a request for obtaining needs to be sent to the aforementioned user or institution through the interface, and after receiving the consent information feedback from the aforementioned user or institution, the relevant information is obtained.
[0024] According to an embodiment of the present application, a method embodiment of an audio-video processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0025] It should be noted that an intelligent processing system can be used as the execution subject of the audio-video processing method in the embodiment of the present application. It can be understood that the audio-video processing method provided in the embodiment of the present application can also be executed by other systems or devices as the execution subject, and the embodiment of the present application does not make specific limitations on this.
[0026] Figure 1 is a flowchart of an optional audio-video processing method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0027] Step S101, receiving a target audio-video segment.
[0028] In step S101, the target audio-video segment includes M sub-audio-video segments, and each sub-audio-video segment includes S-frame video information and N-frame audio information.
[0029] In step S101, M, N, and S are all integers greater than or equal to 1.
[0030] Optionally, the target audio-video segment refers to the audio-video data unit received by the intelligent processing system in the real-time synthesis and push stream of audio and video, and each data unit is composed of M sub-segments.
[0031] Optionally, the sub-audio-video segment refers to the basic component unit in the target segment, and each sub-segment contains S-frame video information and N-frame audio information.
[0032] Optionally, receiving the audio-video segment is the starting point of the synchronous push stream process. The intelligent processing system obtains a series of audio-video segments from the audio-video data source to provide a data basis for subsequent synchronous processing.
[0033] Step S102: Adjust the start timestamps of each of the M sub audio-video segments to obtain M adjusted first sub audio-video segments.
[0034] Optionally, the start timestamp refers to the playback time identifier of the first frame of video information and the first frame of audio information of each sub-segment.
[0035] Optionally, the first sub audio-video segment refers to the sub-segment after the start timestamp adjustment.
[0036] Optionally, this step is to ensure the uniformity of the audio-video data on the time basis. By adjusting the start timestamps of the sub-segments, the time basis differences between the segments are eliminated, providing a unified time reference for subsequent audio-video synchronization processing.
[0037] Step S103: Determine the target time based on the start timestamps of each first sub audio-video segment, and adjust the timestamps of the S-frame video information and the N-frame audio information in each first sub audio-video segment according to the target time to obtain M second sub audio-video segments.
[0038] In step S103, the target time is the timestamp among the start timestamps corresponding to the M first sub audio-video segments that is less than the preset threshold.
[0039] Optionally, the target time: Among the start timestamps of all the first sub audio-video segments, select a timestamp less than the preset threshold as the target time for subsequent timestamp adjustment. In this embodiment, the preset threshold can be 0 for determining negative timestamps.
[0040] Optionally, if the target time is a negative timestamp, the intelligent processing system will correct the timestamp based on the target time to ensure that all timestamps are positive, avoiding playback errors caused by negative timestamps.
[0041] Optionally, the intelligent processing system corrects each frame of information in the first sub audio-video segment in terms of time by determining the target time, further enhancing the synchronization of the audio-video stream.
[0042] Step S104: Perform global offset processing on the M second sub audio-video segments, and push each processed second sub audio-video segment.
[0043] In step S104, the global offset processing is used to correct the time deviation generated during the transmission of different sub audio-video segments.
[0044] Optionally, the global offset processing is performed on the second sub-audio-video segment to correct the additional time deviations that may occur during the transmission of different sub-segments, which may be caused by network latency, packet loss, or processing time differences. The global offset processing ensures the synchronization of all sub-segments on the global timeline by adjusting the timestamps of each frame.
[0045] Optionally, the global offset processing is the last step of the audio-video stream synchronization. It corrects the time deviations caused by factors such as network conditions or processing delays, ensuring the timing and playback synchronization of the audio-video stream. Subsequently, the push operation ensures that the adjusted audio-video stream can be transmitted to the user in a timely and accurate manner, achieving high-quality audio-video synchronous playback.
[0046] Optionally, the intelligent processing system first reads multiple audio-video segments and uniformly adjusts the time bases of each read audio-video segment to ensure that the start times of the audio and video segments are the same. This rebase operation can effectively avoid the situation of audio-video desynchronization, especially when there are differences in the time bases of different segments. By adjusting the base, it lays the foundation for synchronous playback; after the rebase operation, an overall time offset is performed on the current audio-video segment to repair any negative timestamp situations and ensure that the start times of all segments are positive. This step can solve the problem that the start frame time of some segments is negative during playback and ensure the correct alignment of the segments on the global timeline; after the overall offset is completed, a global time offset adjustment is performed on each frame of the audio and video data to further ensure their alignment on the global timeline. This global offset operation is crucial for real-time playback and the continuity of the stream, ensuring that the data stream does not have a time misalignment; after the rebase and offset processing, the audio and video segments are respectively pushed into the global audio-video queue. Before the push, the audio stream will be further overall offset according to the global time (global offset processing) to ensure its consistency with the video stream on the timeline; finally, the global push thread simultaneously processes the data in the two queues of the audio and video, sorts them according to the timestamps of the current frames, and preferentially pushes the stream segments with smaller times. Through this timestamp comparison mechanism, accurate synchronous playback of the audio-video segments is achieved.
[0047] Optionally, the intelligent processing system repairs the negative timestamps and ensures the alignment of the audio and video on the timeline through the segment rebase operation and overall offset, providing guarantee for the synchronization of subsequent pushes. The combination of the global offset and the priority push mechanism ensures the timing synchronization and smooth playback of the audio-video stream. This technical solution has made an innovative improvement on the existing audio-video synchronization technology, can effectively solve the problem of audio-video desynchronization, and improve the quality of the real-time streaming service and the user experience.
[0048] Optionally, the intelligent processing system successfully reduces the time difference between audio and video streams through precise audio-visual clip rebase operations, overall offset adjustments, and global time offset processing (global offset processing), achieving precise synchronization. The optimized push mechanism makes the timing management of audio and video streams more efficient, ensuring the smoothness of real-time playback and thus avoiding user discomfort caused by out-of-sync audio and video.
[0049] Optionally, the technical solution of this embodiment is not only applicable to the push of a single audio-visual stream but also capable of handling the push scenario of multiple video files. By dynamically calculating the starting offset based on the played audio-visual timestamps, this embodiment can ensure the audio-visual synchronization relationship between different video files, significantly enhancing the applicability and flexibility of the system, especially in scenarios with high requirements for audio-visual synchronization such as digital human systems, remote conferences, and online education.
[0050] Based on the content of the above steps S101 to S104, in this application, first, a target audio-visual clip is received, where the target audio-visual clip includes M sub-audio-visual clips, and each sub-audio-visual clip includes S frames of video information and N frames of audio information, where M, N, and S are all integers greater than or equal to 1. Then, the starting timestamp of each of the M sub-audio-visual clips is adjusted to obtain M adjusted first sub-audio-visual clips. Then, the target time is determined based on the starting timestamp of each first sub-audio-visual clip, and the timestamps of the S frames of video information and N frames of audio information in each first sub-audio-visual clip are adjusted according to the target time to obtain M second sub-audio-visual clips, where the target time is the timestamp among the starting timestamps corresponding to the M first sub-audio-visual clips that is less than the preset threshold. Finally, global offset processing is performed on the M second sub-audio-visual clips, and each processed second sub-audio-visual clip is pushed, where the global offset processing is used to correct the time deviation generated during the transmission of different sub-audio-visual clips. That is, through the method of unified global time reference and dynamic timestamp adjustment, the purpose of ensuring the timing consistency and synchronous playback of audio and video streams under the global time reference is achieved, thereby realizing the precise synchronization of audio and video streams and significantly improving the playback quality of real-time audio and video services, and further solving the technical problem of out-of-sync audio and video caused by timestamp offset and insufficient synchronization mechanism in the existing real-time audio and video processing services.
[0051] In an alternative embodiment, the intelligent processing system first determines the start timestamp of the first frame of video information in the first sub-audio-video segment among the M sub-audio-video segments, and uses the start timestamp of the first frame of video information in the first sub-audio-video segment as the first time. Then, it determines the start timestamp of the first frame of audio information in the first sub-audio-video segment among the M sub-audio-video segments, and uses the start timestamp of the first frame of audio information in the first sub-audio-video segment as the second time. When the first time is greater than or equal to the second time, the first time is used as the reference time, where the reference time is used to adjust the start timestamps of each sub-audio-video segment in the target audio-video segment. When the first time is less than the second time, the second time is used as the reference time. Finally, based on the reference time, the start timestamps of each sub-audio-video segment are adjusted to obtain M adjusted first sub-audio-video segments.
[0052] Optionally, the intelligent processing system first reads and analyzes the first segment among the M sub-audio-video segments, determines the start timestamp of the first frame of video information in the video stream of this segment, and marks it as the first time. Subsequently, the system determines the start timestamp of the first frame of audio information in the audio stream of the same segment and marks it as the second time. Then, the system compares the magnitudes of the first time and the second time. If the first time (the start timestamp of the video stream) is greater than or equal to the second time (the start timestamp of the audio stream), the first time is used as the reference time. Conversely, if the first time is less than the second time, the second time is used as the reference time. The selection of the reference time ensures that subsequent time adjustments can be based on the one with the larger start timestamp between the audio and the video, thus avoiding the occurrence of the initial stage of audio-visual out-of-sync phenomenon.
[0053] Optionally, after determining the reference time, the intelligent processing system adjusts the start timestamp of each of the M sub-audio-video segments based on this reference time. The purpose of the adjustment is to align all segments relative to the reference time, eliminate the initial time difference between the video and audio streams, and ensure that the starting points of audio-visual synchronization are consistent. The adjustment method can be to use the difference between the start timestamp of each segment and the reference time as the time offset and apply it to all frame data of this segment, thereby achieving the unification of the time reference. After the above time adjustment, the system generates M adjusted first sub-audio-video segments, and the start timestamp of each segment is aligned with the reference time, ensuring the synchronization of the audio-visual streams on the global time axis.
[0054] As can be seen from the above, by determining the reference time and adjusting the timestamps, the intelligent processing system effectively solves the problem of the initial time difference of audio-visual out-of-sync during the real-time synthesis and push stream of audio and video. The specific technical effects include: by accurately determining the reference time, this embodiment ensures that the audio-visual streams can achieve high-precision synchronization at the starting stage of synchronous push stream, improving the reliability of the synchronization effect; the determination of the reference time and the adjustment of the sub-segment timestamps eliminate the time difference between the audio and video streams at the starting stage, avoiding the early occurrence of audio-visual out-of-sync phenomena and enhancing the user experience; the adjusted first sub-audio-visual segment has a unified time reference, providing time alignment for subsequent audio-visual synthesis and push stream, and ensuring the synchronization of the audio-visual streams during transmission.
[0055] In an alternative embodiment, the intelligent processing system takes the absolute difference between the start timestamp of the first-frame video information in each sub-audio-visual segment and the reference time as the first offset, obtaining M first offsets, and takes the absolute difference between the start timestamp of the first-frame audio information in each sub-audio-visual segment and the reference time as the second offset, obtaining M second offsets. Then, according to the first offset and the second offset corresponding to each sub-audio-visual segment, it adjusts the start timestamp of the sub-audio-visual segment to obtain M adjusted first sub-audio-visual segments.
[0056] Optionally, first, the intelligent processing system receives a series of target audio-visual segments from the audio-visual source. Each segment is composed of multiple sub-audio-visual segments, and each sub-segment contains S frames of video information and N frames of audio information, where S and N are integers greater than or equal to 1. Then, the system calculates the absolute difference between the start timestamp of the first-frame video information in each sub-audio-visual segment and the reference time to obtain the first offset corresponding to each sub-segment. The first offset reflects the time difference between the video information and the reference time and is an important basis for adjusting the time reference of the video segment. Similarly, the intelligent processing system calculates the absolute difference between the start timestamp of the first-frame audio information in each sub-audio-visual segment and the reference time as the second offset corresponding to each sub-segment. The second offset reveals the deviation degree of the audio information relative to the reference time and is used for adjusting the audio time reference.
[0057] Optionally, the intelligent processing system adjusts the start timestamp of each sub-audio-visual segment according to the calculated M first offsets and M second offsets. The adjustment principle is to add the corresponding first offset to the start timestamp of the video information and add the corresponding second offset to the start timestamp of the audio information, thereby obtaining M adjusted first sub-audio-visual segments to ensure the temporal alignment of the audio-visual information within each segment.
[0058] Optionally, the intelligent processing system reads audio-visual segments from the audio-visual data source, including video streams and audio streams. It performs a unified adjustment of the time base for each audio-visual segment to ensure that the start times of the video and audio segments are consistent.
[0059] The specific operations are as follows:
[0060] 1) Read the start timestamp of the audio-visual segment. Assuming the start audio-visual segment number is i, the start time of the video stream in the start audio-visual segment i is T v,i , and the start time of the audio stream is T a,i . Then the reference time T base,i is as shown in formula (1):
[0061] T base,i = max(T v,i , T a,i ) (1)
[0062] 2) Calculate the time base offset of the segment. Taking the first frame of the video stream and the audio stream in the start audio-visual segment as an example, the time base offset of the video stream ΔT v,i and the time base offset of the audio stream ΔT a,i are respectively as shown in formula (2) and formula (3):
[0063] ΔT v,i = T base,i - T v,i (2)
[0064] ΔT a,i = T base,i - T a,i (3)
[0065] 3) After determining the offsets corresponding to all frame data in the segment, perform a time base adjustment on each frame of data in the segment (still taking the first frame of the video stream and the audio stream in the start audio-visual segment as an example) to ensure that the start time of the segment is a unified time base. The timestamp t' v of the first frame of the video stream in the adjusted start audio-visual segment and the timestamp t' a of the first frame of the audio stream in the adjusted start audio-visual segment are respectively as shown in formula (4) and formula (5):
[0066] t' v = T v,i + ΔT v,i (4)
[0067] t' a = T a,i + ΔT a,i (5)
[0068] As can be seen from the above, by introducing the method of intelligent calculation of the first offset and the second offset, the precise adjustment of the time reference of the audio-visual segment is achieved. This method fully considers the individual differences between the audio-visual information inside the segment and the reference time. Through intelligent operations, it effectively solves the synchronization error caused by inconsistent time references during the push stream process of the audio-visual stream, and greatly improves the accuracy and smoothness of audio-visual synchronization. The specific technical effects are reflected in: by accurately calculating the difference between the first frame of video and audio information and the reference time, the fine-tuning of the start timestamp of the segment is achieved, reducing the synchronization deviation of the entire audio-visual stream to an unprecedented level; since the time reference is adjusted at the segment level, this method reduces the latency of global data processing and enhances the smoothness and coherence of real-time push stream of audio-visual.
[0069] In an alternative embodiment, the intelligent processing system determines a target time based on the start timestamp of each first sub-audio-visual segment and detects whether the target time is negative. If the target time is negative, a third offset is determined according to the target time, where the third offset is used to adjust the negative timestamps in each first sub-audio-visual segment to positive timestamps. If the target time is not negative, the third offset is determined to be 0, and then the timestamps of the S-frame video information and N-frame audio information in each first sub-audio-visual segment are adjusted according to the third offset to obtain M second sub-audio-visual segments.
[0070] Optionally, first, the intelligent processing system receives M first sub-audio-visual segments, each segment containing S-frame video information and N-frame audio information. The system analyzes the start timestamp of each segment to determine a common target time point. The target time is determined by comparing the start timestamps of all segments and selecting the smallest timestamp as the target time. After determining the target time, the intelligent processing system checks whether the timestamp is negative. A negative timestamp indicates that the start time of the segment is earlier than the system's time reference, which is unreasonable in audio-visual synchronization processing because it may cause chaos in audio-visual playback.
[0071] Optionally, if the target time is negative, the intelligent processing system determines the third offset according to the absolute value of the target time. The role of the third offset is to adjust the negative timestamps in all segments to positive timestamps to ensure that the start time of each segment meets the requirements of synchronous playback. If the target time is not negative, the third offset is determined to be 0, and there is no need to adjust the positive and negative of the timestamp.
[0072] Optionally, once the third offset is determined, the intelligent processing system adjusts the timestamps of the S-frame video information and N-frame audio information in each first sub-audio-video segment to ensure that all segments can be correctly aligned on the timeline. After the above adjustments, the intelligent processing system will obtain M second sub-audio-video segments. The timestamps of these segments are unified and corrected, ensuring the synchronization of audio-video segments during processing and streaming, providing an optimized basis for subsequent global offset processing and synchronous pushing.
[0073] Optionally, the intelligent processing system makes an overall offset adjustment to the audio-video segments, fixes the negative timestamp problem in the segments, and ensures that the start time of all segments is positive. Let the adjusted start time of each segment be t′ start,i , that is, t′ start,i =t′ v =t′ a , which means that the start times of the video and audio in the current segment are the same. Let the minimum start time be T min =min({t′ start,i}), that is, the minimum start time among all segments. Then the overall offset is β = max(0, -T min ), so the corrected timestamp corresponding to each segment is t″ = t′ start,i +β, that is, the timestamp of each frame of video information and audio information in each segment should be added with the overall offset.
[0074] As can be seen from the above, the intelligent processing system effectively solves the negative timestamp problem that may occur during the real-time streaming of audio and video. By determining and applying the third offset, it ensures that the timestamps of all audio-video segments are positive, thus accurately aligning them on the timeline. The implementation of this technical solution significantly improves the accuracy and reliability of audio-video stream synchronization and optimizes the real-time playback experience. By eliminating negative timestamps, it avoids the misalignment of segments on the global timeline, enhances the real-time performance and smoothness of audio-video synthesis and streaming, and finally achieves a high-quality audio-video synchronous playback effect.
[0075] In an optional embodiment, the intelligent processing system constructs J video data packets and K audio data packets based on the M second sub-audio-video segments, where J and K are integers greater than or equal to 1, and then pushes the J video data packets and K audio data packets into the audio-video queue.
[0076] Optionally, after variable base and overall offset adjustments, the audio-video segments are respectively pushed by the intelligent processing system into the global audio-video queue, that is, the video queue is Q u =∪{t′ v ′}, and the audio queue is Q a =∪{t′ a ′}, where t′v ′ is used to represent the video data packet after overall offset, t′ a ′ is used to represent the audio data packet after overall offset.
[0077] Optionally, the J video data packets and K audio data packets refer to a set of data packets reconstructed by extracting video information and audio information respectively from M sub - segments. J and K are integers greater than or equal to 1, representing the number of data packets.
[0078] Optionally, the intelligent processing system first performs data segmentation on the M second sub - audio - video segments, separates the video information and audio information in each segment, and processes them separately. The system constructs J video data packets and K audio data packets according to the separated video and audio information. Each data packet may contain information from multiple sub - segments, and the construction basis can be the size of the segment, transmission efficiency, or playback requirements, etc.
[0079] Optionally, the audio - video queue refers to a cache queue inside the intelligent processing system used to store and process audio - video data packets, and it is a relay point for the audio - video data to be transmitted to the push - stream service or the playback end.
[0080] Optionally, the intelligent processing system pushes the constructed J video data packets and K audio data packets into the audio - video queue respectively, waiting for further processing and transmission.
[0081] Optionally, the intelligent processing system manages the audio - video queue, such as sorting the queue according to the timestamp information of the data packets, ensuring that the audio - video data packets can be processed and transmitted in the correct time order.
[0082] As can be seen from the above, the intelligent processing system realizes the efficient management and optimized transmission of audio - video data by constructing video and audio data packets and pushing them into the audio - video queue in an orderly manner. At the same time, it enhances the stability of audio - video synchronization, is applicable to real - time audio - video push - stream services, and significantly improves the playback experience and synchronization quality. This solution focuses on the construction of data packets and queue management, and effectively avoids the problem of out - of - sync audio and video by optimizing the data processing and transmission process, ensuring the high - quality transmission and playback of real - time audio - video streams.
[0083] In an alternative embodiment, the intelligent processing system sets a global playback time, where the global playback time is used to represent the starting playback moment of the target audio-video segment. Then, it reads J video data packets and K audio data packets composed of M second sub-audio-video segments from the audio-video queue. Next, it determines a third time based on the J video data packets and K audio data packets, where the third time is the timestamp among the timestamps corresponding to the J video data packets and K audio data packets that is less than a second preset threshold. Then, it determines a target offset based on the global playback time and the third time, where the target offset is used to represent the time deviation generated during the transmission of each sub-audio-video segment. Finally, it performs global offset processing on the timestamps corresponding to the J video data packets and K audio data packets according to the target offset, and pushes the processed J video data packets and K audio data packets.
[0084] Optionally, the global playback time serves as a reference for the starting playback moment of the target audio-video segment and is a core time reference point in the intelligent processing system. After receiving the target audio-video segment, the intelligent processing system sets a global playback time according to the current playback environment and existing time information. This time point is crucial for subsequent data packet reading and timestamp adjustment. It serves as the starting point for the playback of all audio-video data, ensuring global consistency in audio-video synchronization.
[0085] Optionally, from the global audio-video queue, the intelligent processing system reads J video data packets and K audio data packets composed of M second sub-audio-video segments. Here, the "J video data packets" and "K audio data packets" represent the audio-video information organized in the form of data packets after preliminary processing of the audio-video segments, ready for further time synchronization and streaming. Then, the system determines a third time based on the timestamp information of the read J video data packets and K audio data packets. The selection criterion for the third time is to select one or more timestamps less than a second preset threshold among the timestamps corresponding to all data packets as the third time. The third time can be the smallest timestamp among the timestamps corresponding to all data packets. The setting of the second preset threshold is to identify and handle potential time offsets, ensuring that the timestamp adjustment is carried out within a reasonable range and avoiding synchronization problems caused by excessive time offsets.
[0086] Optionally, after determining the third time, the intelligent processing system calculates the target offset based on the global playback time and the third time. The role of the target offset is to represent the slight time deviation generated during the transmission of the audio-visual segment, and it is an important parameter for precise timestamp adjustment. By calculating the difference between the global playback time and the third time, the system can identify and correct the time differences in the transmission and processing of the audio-visual stream, ensuring the consistency of audio-visual synchronization. Finally, the intelligent processing system performs global offset processing on the timestamps corresponding to the J video data packets and the K audio data packets according to the calculated target offset. The adjusted packet timestamps are aligned with the global playback time, eliminating the time deviation generated during transmission and ensuring the synchronized playback of the audio-visual stream. After the processing is completed, the system pushes the J video data packets and the K audio data packets with adjusted timestamps in real time, transmits the audio-visual stream to the user side, and realizes a high-quality synchronized playback experience.
[0087] Optionally, the intelligent processing system first sets the global playback start time T play (to ensure their alignment on the global timeline), then reads the audio-visual data packets from the global audio-visual queue, and determines that the timestamp of the first frame (the third time) is T first = min(Q u ∪Q a ), where Q u represents the video queue, Q a represents the audio queue, and T first represents the timestamp of the packet with the smallest value among the corresponding timestamps in the video queue and the audio queue. Then the global offset (target offset) is γ = T play - T first , and then adjusts all the packets according to the global offset. The globally aligned timestamp is t global = t″ + γ, where t″ is used to represent the timestamp corresponding to the packet after the rebase operation and the overall offset adjustment, and t global represents the timestamp corresponding to the packet after the global offset processing, that is, each packet in the queue needs to be added with this global offset.
[0088] As can be seen from the above, by setting the global playback time and determining the target offset, the intelligent processing system effectively solves the time deviation problem caused by transmission and processing during the real-time push of audio and video streams. The specific technical effects include: the global playback time serves as the starting point for audio and video synchronization, providing a unified time reference and laying the foundation for subsequent adjustment of the packet timestamps; the determination of the third time and the calculation of the target offset enable the system to accurately identify and correct the time deviation generated during the transmission of the audio and video streams, ensuring the synchronized playback of the audio and video data; through global offset processing, the timestamps of the audio and video streams are uniformly adjusted, improving the accuracy and smoothness of audio and video synchronization and providing users with a better real-time audio and video playback experience.
[0089] In an optional embodiment, if there are timestamps less than or equal to the global playback time among the processed J video packets and K audio packets, the intelligent processing system re-performs global offset processing on the J video packets and K audio packets; if the timestamps of the processed J video packets and K audio packets are all greater than the global playback time, the intelligent processing system pushes the processed J video packets and K audio packets.
[0090] Optionally, the intelligent processing system checks whether there are timestamps less than or equal to the current global playback time among the J video packets and K audio packets. This is to ensure that the timestamps of the packets are reasonable on the playback timeline and avoid audio and video out-of-sync problems caused by improper adjustment of the time reference. If there are packets with timestamps less than or equal to the global playback time, the intelligent processing system will re-perform global offset processing on these video and audio packets. The purpose of this processing is to correct the deviation between the segment time reference and the actual playback time caused by network latency or other factors, ensuring that the timestamps of all packets are after the global playback time, so that the audio and video streams can be played in the correct time order and avoid phenomena such as audio-visual out-of-sync or packet omission.
[0091] Optionally, after the global offset processing is completed, if all the timestamps of the J video packets and K audio packets are greater than the global playback time, then the intelligent processing system will push these packets. The push mechanism ensures that the packets can be correctly transmitted to the playback end in the order of the adjusted timestamps, realizing the real-time synchronous playback of audio and video, that is, when t global >T play the adjusted audio and video packets are pushed to the push service to ensure the timing synchronization of the audio and video streams.
[0092] As can be seen from the above, by intelligently checking and dynamically adjusting the timestamps of audio and video data packets, the intelligent processing system ensures the rationality and sequencing of the data packets on the global playback timeline, effectively avoiding audio and video out-of-sync problems caused by improper adjustment of the time reference or network latency. The specific technical effects are as follows: By real-time monitoring and adjusting the timestamps of audio and video data packets, this solution can handle time delays or data packet timestamp offsets in different network environments, enhancing the stability and robustness of audio and video synchronization; reprocessing data packets with timestamps less than or equal to the global playback time to ensure that all audio and video data packets are correctly pushed in the adjusted time order, avoiding situations such as chaotic playback order or audio-visual out-of-sync; after global offset processing, the data packets are consistent with or after the global playback time in terms of timestamps, ensuring the real-time playback effect of the audio and video stream and avoiding audio and video out-of-sync problems that users may encounter in scenarios such as watching live broadcasts, online education, or remote meetings, greatly improving the user experience of real-time audio and video services.
[0093] In an alternative embodiment, the intelligent processing system sorts the timestamps corresponding to the J processed video data packets and K audio data packets to obtain a sorting result, and then pushes the J video data packets and K audio data packets according to the sorting result.
[0094] Optionally, for timestamp sorting: the intelligent processing system sorts the timestamps of the J video data packets and K audio data packets respectively to ensure that they are arranged in chronological order.
[0095] Optionally, the sorting is performed in the order of the corresponding start timestamps in each video data packet and audio-video packet.
[0096] Optionally, the intelligent processing system collects the timestamps of all the J video data packets and K audio data packets to be pushed, and these timestamps have been adjusted in the global offset processing. The system first sorts the timestamps of the J video data packets, and then sorts the timestamps of the K audio data packets. The sorting process is carried out in ascending order of the timestamp values, ensuring that the data packets that should be played earliest are arranged in the front. After the sorting is completed, the system generates two sorted lists, namely the sorted list of video data packets and the sorted list of audio data packets, and each list is arranged in ascending order of the timestamps.
[0097] Optionally, the intelligent processing system first retrieves the data packets with the smallest timestamps from the two sorted lists of video and audio, and compares the magnitudes of their timestamp values. If the timestamp of the video data packet is less than or equal to the timestamp of the audio data packet, the system preferentially pushes the video data packet; otherwise, it preferentially pushes the audio data packet. This timestamp-based preferential push strategy ensures the temporal synchronization of the audio-visual stream. The system repeatedly performs the timestamp comparison and push operations on the data packets until all J video data packets and K audio data packets have been pushed.
[0098] Optionally, in this embodiment, by optimizing the processing and push mechanism of audio-visual data, low-latency real-time playback is achieved. This low-latency performance significantly improves the user's viewing experience, especially in application scenarios that require real-time feedback and high fidelity, such as the interactive experience and content display in a digital human system, where the effect is particularly remarkable.
[0099] Optionally, retrieve the audio-visual data packets from the audio-visual queue. Let the video packet time be represented as t global,v , and the audio packet time be represented as t global,a . Compare the playback times of the audio-visual data packets, calculate the data packets to be preferentially pushed, and ensure the temporal synchronization of the audio-visual stream. The condition for preferential push is to select min(t global,v , t global,a ), that is, as shown in formula (6):
[0100]
[0101] where PushOrder represents the push order. If the video packet time is less than or equal to the audio packet time, the video packet corresponding to the video packet time is preferentially pushed; otherwise, if the audio packet time is less than the video packet time, the audio packet corresponding to the audio packet time is preferentially pushed.
[0102] Optionally, based on the timestamp-based preferential push mechanism, the intelligent processing system can ensure the temporal synchronization of the audio-visual stream and avoid audio-visual desynchronization caused by out-of-order data packets.
[0103] Optionally, Figure 2 is a schematic diagram of an optional audio-visual processing method according to an embodiment of the present application. As Figure 2 shown, first obtain the audio-visual segments, then perform rebase operations and overall offset processing on each audio-visual segment to obtain the corresponding video sequence and audio sequence for each segment, transmit the video sequence and audio sequence to the video stream data packet queue and the audio stream data packet queue respectively, then perform global data packet offset processing on all the audio-visual data packets in the queue, and after the processing is completed, transmit them to the data reception queue, and finally send them to the push media server for pushing.
[0104] Optionally, this embodiment solves the problem of out-of-order data packets caused by cross-receiving audio stream and video stream data packets in the push queue. When playing an audio-video stream, the audio-video (A-V) display value (which is an important indicator for measuring the synchronization of audio-video stream playback) usually remains within 0.05 seconds. This indicates that the time difference in the audio-video synchronization state is within 0.05 seconds, ensuring a high-quality real-time playback experience, which is significantly better than the prior art where the A-V deviation may reach 0.1 second or higher.
[0105] As can be seen from the above, by sorting the timestamps of the processed video and audio data packets and pushing them accordingly, the synchronization and smoothness of the real-time audio-video push service are effectively improved. The specific technical effects are as follows: Through timestamp sorting, the intelligent processing system can accurately identify and push the data packets according to the actual playback timing of the audio-video stream, avoiding the problem of out-of-sync audio and video, and ensuring playback synchronization; The push mechanism based on the sorting result improves the efficiency of data packet transmission, reduces unnecessary waiting time, makes the push of the audio-video stream more timely, and enhances the real-time playback experience; The application of this technical solution significantly reduces the A-V display value, that is, the audio-video synchronization deviation, so that users can hardly feel the delay when watching audio-video content, greatly improving the user experience.
[0106] The embodiment of the present application also provides an audio-video processing device. It should be noted that the audio-video processing device of the embodiment of the present application can be used to execute the audio-video processing method provided by the embodiment of the present application. The audio-video processing device provided by the embodiment of the present application will be introduced below.
[0107] According to the embodiment of the present application, there is also provided a device for implementing the above audio-video processing method. Figure 3 is a schematic diagram of an optional audio-video processing device according to the embodiment of the present application, as Figure 3 shown, including: a receiving unit 301, a first adjustment unit 302, a second adjustment unit 303, and a pushing unit 304.
[0108] Optionally, a receiving unit 301 is configured to receive a target audio-visual segment, where the target audio-visual segment includes M sub-audio-visual segments, each sub-audio-visual segment includes S frames of video information and N frames of audio information, and M, N, and S are all integers greater than or equal to 1; a first adjustment unit 302 is configured to adjust the start timestamps of each of the M sub-audio-visual segments to obtain M adjusted first sub-audio-visual segments; a second adjustment unit 303 is configured to determine a target time based on the start timestamps of each of the first sub-audio-visual segments, and adjust the timestamps of the S frames of video information and the N frames of audio information in each of the first sub-audio-visual segments according to the target time to obtain M second sub-audio-visual segments, where the target time is the timestamp less than a preset threshold among the start timestamps corresponding to the M first sub-audio-visual segments; a pushing unit 304 is configured to perform global offset processing on the M second sub-audio-visual segments, and push each of the processed second sub-audio-visual segments, where the global offset processing is used to correct the time deviation generated during the transmission of different sub-audio-visual segments.
[0109] Optionally, the first adjustment unit 302 includes: a first determination subunit, a second determination subunit, a third determination subunit, a fourth determination subunit, and a first adjustment subunit. Among them, the first determination subunit is configured to determine the start timestamp of the first frame of video information in the first sub-audio-visual segment among the M sub-audio-visual segments, and use the start timestamp of the first frame of video information in the first sub-audio-visual segment as the first time; the second determination subunit is configured to determine the start timestamp of the first frame of audio information in the first sub-audio-visual segment among the M sub-audio-visual segments, and use the start timestamp of the first frame of audio information in the first sub-audio-visual segment as the second time; the third determination subunit is configured to use the first time as the reference time when the first time is greater than or equal to the second time, where the reference time is used to adjust the start timestamps of each of the sub-audio-visual segments in the target audio-visual segment; the fourth determination subunit is configured to use the second time as the reference time when the first time is less than the second time; the first adjustment subunit is configured to adjust the start timestamps of each of the sub-audio-visual segments based on the reference time to obtain M adjusted first sub-audio-visual segments.
[0110] Optionally, the first adjustment subunit includes: a first determination module, a second determination module, and a first adjustment module. Among them, the first determination module is configured to use the absolute difference between the start timestamp of the first-frame video information in each sub-audio-video segment and the reference time as the first offset, and obtain M first offsets; the second determination module is configured to use the absolute difference between the start timestamp of the first-frame audio information in each sub-audio-video segment and the reference time as the second offset, and obtain M second offsets; the first adjustment module is configured to adjust the start timestamp of each sub-audio-video segment according to the corresponding first offset and second offset of each sub-audio-video segment, and obtain M adjusted first sub-audio-video segments.
[0111] Optionally, the second adjustment unit 303 includes: a fifth determination subunit, a sixth determination subunit, a seventh determination subunit, and a second adjustment subunit. Among them, the fifth determination subunit is configured to determine a target time based on the start timestamp of each first sub-audio-video segment and detect whether the target time is negative; the sixth determination subunit is configured to, if the target time is negative, determine a third offset according to the target time, where the third offset is used to adjust the negative timestamp in each first sub-audio-video segment to a positive timestamp; the seventh determination subunit is configured to, if the target time is not negative, determine the third offset to be 0; the second adjustment subunit is configured to adjust the timestamps of the S-frame video information and the N-frame audio information in each first sub-audio-video segment according to the third offset, and obtain M second sub-audio-video segments.
[0112] Optionally, the audio-video processing device further includes: a construction unit and a first push unit. Among them, the construction unit is configured to construct J video data packets and K audio data packets according to the M second sub-audio-video segments, where J and K are integers greater than or equal to 1; the first push unit is configured to push the J video data packets and the K audio data packets into the audio-video queue.
[0113] Optionally, the pushing unit 304 includes: a first setting subunit, a first reading subunit, an eighth determining subunit, a ninth determining subunit, and a first pushing subunit. Among them, the first setting subunit is configured to set a global playing time, where the global playing time is used to represent the starting playing moment of the target audio-visual segment; the first reading subunit is configured to read J video data packets and K audio data packets composed of M second sub-audio-visual segments from the audio-visual queue; the eighth determining subunit is configured to determine a third time according to the J video data packets and the K audio data packets, where the third time is the time stamp less than a second preset threshold among the time stamps corresponding to the J video data packets and the K audio data packets; the ninth determining subunit is configured to determine a target offset according to the global playing time and the third time, where the target offset is used to represent the time deviation generated when each sub-audio-visual segment is transmitted; the first pushing subunit is configured to perform global offset processing on the time stamps corresponding to the J video data packets and the K audio data packets according to the target offset, and push the processed J video data packets and K audio data packets.
[0114] Optionally, the first pushing subunit includes: a first processing module and a first pushing module. Among them, the first processing module is configured to re-perform global offset processing on the J video data packets and the K audio data packets if there are time stamps less than or equal to the global playing time among the processed J video data packets and K audio data packets; the first pushing module is configured to push the processed J video data packets and K audio data packets if the time stamps of the processed J video data packets and K audio data packets are all greater than the global playing time.
[0115] Optionally, the first pushing module includes: a first sorting sub-module and a first pushing sub-module. Among them, the first sorting sub-module is configured to sort the time stamps corresponding to the processed J video data packets and K audio data packets to obtain a sorting result; the first pushing sub-module is configured to push the J video data packets and K audio data packets according to the sorting result.
[0116] According to another aspect of the present application, an electronic device is further provided, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above audio-visual processing method.
[0117] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0118] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0119] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed among each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0120] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0121] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0122] If the above-mentioned integrated units are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0123] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An audio-video processing method, characterized in that Including: Receiving a target audio - video segment, where the target audio - video segment includes M sub - audio - video segments, and each sub - audio - video segment includes S - frame video information and N - frame audio information, where M, N, and S are all integers greater than or equal to 1; Adjusting the start timestamp of each of the M sub - audio - video segments to obtain M adjusted first sub - audio - video segments; Determining a target time based on the start timestamp of each first sub - audio - video segment, and adjusting the timestamps of the S - frame video information and N - frame audio information in each first sub - audio - video segment according to the target time to obtain M second sub - audio - video segments, where the target time is the timestamp among the start timestamps corresponding to the M first sub - audio - video segments that is less than a preset threshold; Performing global offset processing on the M second sub - audio - video segments, and pushing each processed second sub - audio - video segment, where the global offset processing is used to correct the time deviation generated during the transmission of different sub - audio - video segments.
2. The audio-video processing method according to claim 1, wherein Adjusting the start timestamp of each of the M sub - audio - video segments to obtain M adjusted first sub - audio - video segments, including: Determining the start timestamp of the first - frame video information in the first sub - audio - video segment among the M sub - audio - video segments, and taking the start timestamp of the first - frame video information in the first sub - audio - video segment as the first time; Determining the start timestamp of the first - frame audio information in the first sub - audio - video segment among the M sub - audio - video segments, and taking the start timestamp of the first - frame audio information in the first sub - audio - video segment as the second time; When the first time is greater than or equal to the second time, taking the first time as the reference time, where the reference time is used to adjust the start timestamp of each sub - audio - video segment in the target audio - video segment; When the first time is less than the second time, taking the second time as the reference time; Adjusting the start timestamp of each sub - audio - video segment based on the reference time to obtain M adjusted first sub - audio - video segments.
3. The audio-video processing method according to claim 2, wherein Adjusting the start timestamp of each sub - audio - video segment based on the reference time to obtain M adjusted first sub - audio - video segments, including: Taking the absolute difference between the start timestamp of the first - frame video information in each sub - audio - video segment and the reference time as the first offset amount to obtain M first offset amounts; Taking the absolute difference between the start timestamp of the first - frame audio information in each sub - audio - video segment and the reference time as the second offset amount to obtain M second offset amounts; Adjusting the start timestamp of each sub - audio - video segment according to the corresponding first offset amount and second offset amount of each sub - audio - video segment to obtain M adjusted first sub - audio - video segments.
4. The audio and video processing method according to claim 1, wherein Determining a target time based on the start timestamp of each first sub - audio - video segment, and adjusting the timestamps of the S - frame video information and N - frame audio information in each first sub - audio - video segment according to the target time to obtain M second sub - audio - video segments, including: Determine a target time based on the start timestamp of each of the first sub-audio-video segments, and detect whether the target time is negative; If the target time is negative, determine a third offset according to the target time, where the third offset is used to adjust the negative timestamps in each of the first sub-audio-video segments to positive timestamps; If the target time is not negative, determine the third offset to be 0; Adjust the timestamps of the S-frame video information and the N-frame audio information in each of the first sub-audio-video segments according to the third offset to obtain M second sub-audio-video segments.
5. The audio and video processing method according to claim 1, wherein Before performing global offset processing on the M second sub-audio-video segments and pushing each processed second sub-audio-video segment, the method further includes: Construct J video data packets and K audio data packets according to the M second sub-audio-video segments, where J and K are integers greater than or equal to 1; Push the J video data packets and the K audio data packets into an audio-video queue.
6. The audio-video processing method according to claim 5, wherein Performing global offset processing on the M second sub-audio-video segments and pushing each processed second sub-audio-video segment includes: Set a global playback time, where the global playback time is used to represent the start playback moment of the target audio-video segment; Read the J video data packets and the K audio data packets composed of the M second sub-audio-video segments from the audio-video queue; Determine a third time according to the J video data packets and the K audio data packets, where the third time is the timestamp less than a second preset threshold among the timestamps corresponding to the J video data packets and the K audio data packets; Determine a target offset according to the global playback time and the third time, where the target offset is used to represent the time deviation generated when each sub-audio-video segment is transmitted; Perform global offset processing on the timestamps corresponding to the J video data packets and the K audio data packets according to the target offset, and push the processed J video data packets and the K audio data packets.
7. The audio and video processing method according to claim 6, wherein Pushing the processed J video data packets and the K audio data packets includes: If there are timestamps less than or equal to the global playback time in the processed J video data packets and the K audio data packets, re-perform the global offset processing on the J video data packets and the K audio data packets; If the timestamps of the processed J video data packets and the K audio data packets are all greater than the global playback time, push the processed J video data packets and the K audio data packets.
8. The audio-video processing method according to claim 7, wherein Pushing the processed J video data packets and the K audio data packets includes: Sort the timestamps corresponding to the processed J video data packets and the K audio data packets to obtain a sorting result; Push the J video data packets and the K audio data packets according to the sorting result.
9. An audio-video processing device, characterized in that, Include: A receiving unit, configured to receive a target audio-video segment, where the target audio-video segment includes M sub-audio-video segments, and each sub-audio-video segment includes S-frame video information and N-frame audio information, where M, N, and S are all integers greater than or equal to 1; A first adjustment unit, configured to adjust the start timestamps of each of the M sub-audio-video segments to obtain M adjusted first sub-audio-video segments; A second adjustment unit, configured to determine a target time based on the start timestamps of each of the first sub-audio-video segments, and adjust the timestamps of the S-frame video information and the N-frame audio information in each of the first sub-audio-video segments according to the target time to obtain M second sub-audio-video segments, where the target time is the timestamp among the start timestamps corresponding to the M first sub-audio-video segments that is less than a preset threshold; A pushing unit, configured to perform global offset processing on the M second sub-audio-video segments, and push each of the processed second sub-audio-video segments, where the global offset processing is used to correct the time deviation generated during the transmission of different sub-audio-video segments.
10. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory is configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the audio-video processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Steel rail flaw detection audio and video automatic synchronization merging and filing method, equipment and medium
CN121985165A