Audio and picture synchronization method, device and equipment and computer storage medium
By obtaining the time offset information in the video file and modifying the time information of the audio file, the problem of audio and video out of synchronization during the live broadcast is solved, synchronous playback of audio and video is realized, and the playback effect of video resources is improved.
Patent Information
- Application Number
- CN202311485682.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-13
AI Technical Summary
During the live broadcast process, due to the different decoding order and display order of each video frame in the video file, the display time of the audio file and the video file is deviated, which leads to the problem of out-of-synchronization of the pronunciation.
By obtaining the time offset information in the video file, the information includes the deviation between the video display time and the decoding time of the video frame, and modifying the audio time information in the audio file based on the information to ensure that the audio display time of the first frame of the audio file is the same as the video display time of the first frame of the video file.
The audio and video synchronization during video playback during live broadcast is realized, and the audio and video out-synchronization problem is avoided due to inconsistent decoding order and display order in the video file, and the playback effect of video resources is improved.
Smart Images

Figure CN119996742A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device and equipment for synchronizing audio and video, and a computer storage medium. Background Art
[0002] As the Internet becomes increasingly popular, more and more people are participating in the Internet world. Among them, live broadcast technology has been widely welcomed because of its advantage of being able to quickly broadcast a scene to any screen in real time.
[0003] For live broadcast technology, in addition to ensuring the stability of real-time data transmission, it is also necessary to ensure that the audio and video in the transmitted data can be played synchronously (hereinafter referred to as audio and video synchronization).
[0004] Usually, in order to achieve audio and video synchronization during the playback of video resources, each frame of data in the video file and the audio file will record its corresponding display time, so as to ensure that the corresponding audio frames and video frames in the audio and video files can be played synchronously.
[0005] However, in order to reduce the pressure of data transmission during live broadcast, a video encoding format with low data storage capacity is used to encode and transmit the video file. This will lead to the phenomenon that the playback order and decoding order of each video frame in the video file are different in such video encoding. For example, Figure 1 As shown, for the 1st, 2nd, 3rd, and 4th video frames, the playback order is 1, 2, 3, and 4, but the decoding order is 1, 3, 4, and 2. Therefore, when playing the 2nd video frame, there will be a waiting process for decoding, and the waiting time is the corresponding offset time (Composition Time, CTS).
[0006] The transmission format requires that, for a video file using this encoding format, the display time of each video frame is the sum of the decoding time and the offset time.
[0007] This results in a deviation between the display time of the first frame of the video file and the display time of the first frame of the audio file, because the decoding order of each audio frame in the audio file is the same as the display order (in other words, the display time of the audio frame is the decoding time of each frame). When the offset time is too large, it is easy to cause the audio and video to be out of sync.
[0008] Therefore, a new method for synchronizing audio and video is urgently needed to achieve audio and video synchronization during video playback during live broadcast. Summary of the invention
[0009] The present application provides a method, device, equipment and computer storage medium for synchronizing audio and video, so as to achieve audio and video synchronization during video playback during live broadcast.
[0010] In a first aspect, the present application provides a method for synchronizing audio and video, which is applied to a network device, comprising:
[0011] Obtain a video file and an audio file corresponding to the video file;
[0012] Based on the video file, corresponding time offset information is obtained; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file;
[0013] Based on the time offset information, the audio time information in the audio file is modified to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame included in the audio file; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file;
[0014] Based on the video display time corresponding to each of the video frames and the target time information, the video file and the audio file are synchronized with each other.
[0015] In a second aspect, the present application provides a method for synchronizing audio and video, which is applied to a video acquisition end, comprising:
[0016] Collecting video data and audio data corresponding to the video data;
[0017] Based on the video data, corresponding time offset information is obtained; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame in the video data;
[0018] Based on the time offset information, the audio time information corresponding to the audio data is modified to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video data;
[0019] Based on the video display time corresponding to each of the video frames and the target time information, the video data and the audio data are processed in synchronization with each other.
[0020] In a third aspect, the present application provides a device for synchronizing audio and video, which is applied to a network device, and the device includes:
[0021] The first acquisition module is used to acquire a video file and an audio file corresponding to the video file; based on the video file, acquire corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file;
[0022] A first processing module is configured to modify the audio time information in the audio file based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame included in the audio file; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file;
[0023] The first synchronization module is used to perform audio and video synchronization processing on the video file and the audio file based on the video display time corresponding to each video frame and the target time information.
[0024] In a possible implementation manner, the first processing module is used to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0025] The deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is respectively integrated into the audio display time and the audio decoding time corresponding to each audio frame in the audio file to obtain the corresponding target time information;
[0026] The target time information includes: a target audio display time and a target audio decoding time corresponding to each audio frame in the audio file; each target audio display time is the same as a video display time of a corresponding video frame.
[0027] In a possible implementation manner, the deviations between the video display time and the video decoding time corresponding to each video frame in the video file are the same;
[0028] The first processing module is used to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically for:
[0029] Obtaining a deviation value between a video display time and a video decoding time corresponding to a first video frame in the time offset information;
[0030] Based on the deviation value, all audio display times and audio decoding times contained in the audio file are modified to obtain corresponding target audio display times and target audio decoding times; wherein, the target audio display time and the corresponding audio display time are positively correlated with the deviation value, and the target audio decoding time and the corresponding audio decoding time are positively correlated with the deviation value.
[0031] In a possible implementation, the video file and the audio file are both packaged in a streaming media package format;
[0032] Then the first acquisition module is used to acquire the corresponding time offset information based on the video file, specifically to:
[0033] Based on the tag headers corresponding to the respective video frames in the video file, obtaining corresponding deviation values;
[0034] The time offset information is obtained based on the deviation values corresponding to the respective video frames.
[0035] In a possible implementation manner, the first processing module is used to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0036] If it is determined that there is a bidirectional predictive interpolation coding frame in the video file, then based on the time offset information, the audio time information in the audio file is modified to obtain corresponding target time information;
[0037] If it is determined that there is no bidirectional predictive interpolation coding frame in the video file, the deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is determined to be zero, and the audio time information is used as the target time information.
[0038] In a possible implementation manner, the first processing module is used to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0039] Based on a preset encoding format, the audio file is re-encoded to obtain a corresponding intermediate file; based on the time offset information, the audio time information in the intermediate file is modified to obtain corresponding target time information.
[0040] In a fourth aspect, the present application provides a device for synchronizing audio and video, which is applied to a video acquisition end, including:
[0041] The second acquisition module is used to collect video data and audio data corresponding to the video data; based on the video data, obtain corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame in the video data;
[0042] A second processing module is configured to modify the audio time information corresponding to the audio data based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video data;
[0043] The second synchronization module is used to perform audio and video synchronization processing on the video data and the audio data based on the video display time corresponding to each video frame and the target time information.
[0044] In a possible implementation manner, when the second acquisition module is used to acquire the corresponding time offset information based on the video data, it is specifically used to:
[0045] Encoding the video data based on a streaming media encapsulation format to obtain a corresponding video file;
[0046] Based on the tag headers corresponding to the respective video frames in the video file, a corresponding deviation value is obtained; the deviation value is used to indicate the deviation between the video display time and the video decoding time;
[0047] The time offset information is obtained based on the deviation values corresponding to the respective video frames.
[0048] In a fifth aspect, the present application provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor implements the steps of any of the above methods.
[0049] In a sixth aspect, the present application also provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to enable the electronic device to execute the steps of any of the above methods.
[0050] In a seventh aspect, the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when the computer program is executed by a processor.
[0051] The beneficial effects of this application are as follows:
[0052] In this solution, by using the time offset information in the video file, the audio time information in the audio file is modified, and the overall audio display time of the audio file is pushed back, ensuring that the audio display time of the first frame of the audio file is the same as the video display time of the first frame of the video file, eliminating the problem of audio and video asynchrony that may occur during the playback of video resources. In the subsequent audio and video synchronization process, the synchronous playback of video and audio can be completed through the same display time corresponding to each frame of the audio file and the video file, without discarding or fast-forwarding the audio frames in the audio file, which maximizes the integrity of the audio and video files and improves the playback effect of video resources.
[0053] On the other hand, by implementing the audio and video synchronization method on the network device side, there is no need for video acquisition equipment or video playback equipment to perform separate audio and video synchronization processing, which reduces the cost of popularizing the method. The video playback quality is improved without the user having to perform additional operations, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram showing the difference between a decoding order and a display order of a video file;
[0055] Figure 2 A schematic diagram of the difference between a decoding order and a display order of a video file containing B frames provided in an embodiment of the present application;
[0056] Figure 3 A schematic diagram of a possible application scenario provided for an embodiment of the present application;
[0057] Figure 4 A flowchart of a method for synchronizing audio and video provided in an embodiment of the present application;
[0058] Figure 5 A schematic diagram of the basic stages of a live broadcast technology provided in an embodiment of the present application;
[0059] Figure 6 A schematic diagram of a video file in FLV packaging format provided in an embodiment of the present application;
[0060] Figure 7 A flowchart of a method for obtaining time offset information provided in an embodiment of the present application;
[0061] Figure 8 A logical schematic diagram of a method for obtaining time offset information provided in an embodiment of the present application;
[0062] Fig. 9 A schematic diagram showing the deviation of display time of a video file and an audio file provided in an embodiment of the present application;
[0063] Fig.10 A flowchart of a method for modifying audio time information provided by an embodiment of the present application;
[0064] Fig.11 A schematic diagram of a timing for modifying audio time information provided by an embodiment of the present application;
[0065] Fig.12 A logical schematic diagram of a method for modifying audio time information provided by an embodiment of the present application;
[0066] Fig.13 A flowchart of a method for modifying audio time information provided by an embodiment of the present application;
[0067] Fig.14 A flowchart of a method for synchronizing audio and video provided in an embodiment of the present application;
[0068] Fig.15A A logic diagram of a method for synchronizing audio and video provided in an embodiment of the present application;
[0069] Fig. 15B A schematic diagram of an audio-visual synchronization effect provided by an embodiment of the present application;
[0070] Fig.16 A schematic diagram of the structure of a device for synchronizing audio and video provided in an embodiment of the present application;
[0071] Fig.17 A schematic diagram of the structure of another device for synchronizing audio and video provided in an embodiment of the present application;
[0072] Fig.18 A schematic diagram of a hardware structure of an electronic device in an embodiment of the present application;
[0073] Fig.19 A schematic diagram of the hardware composition structure of another electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0075] It is understandable that in the following specific implementations of the present application, when the relevant data such as audio and video synchronization is involved, when the various embodiments of the present application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, when it is necessary to obtain relevant data, relevant volunteers can be recruited and relevant agreements for volunteer authorization data can be signed, and then the data of these volunteers can be used for implementation; or, by implementing within the scope of the authorized organization, the following implementation method is implemented by using the data of internal members of the organization to identify internal members; or, the relevant data used in the specific implementation are all simulated data, such as simulated data generated in a virtual scene.
[0076] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:
[0077] Content Delivery Network (CDN): It is a distributed network consisting of servers in different regions and built on and covered by the bearer network. Its function is to cache the source station resources to edge servers across the country for users to obtain nearby and reduce the pressure on the source station. Its basic idea is to avoid bottlenecks and links on the Internet that may affect the speed and stability of data transmission as much as possible, so as to make content transmission faster and more stable. By placing node servers in various parts of the network to form a layer of intelligent virtual network on the basis of the existing Internet, the CDN system can redirect the user's request to the service node closest to the user in real time based on comprehensive information such as network traffic and the connection and load status of each node, as well as the distance to the user and the response time. Its purpose is to enable users to obtain the required content nearby, solve the problem of Internet network congestion, and improve the response speed of users visiting websites.
[0078] Intra-coded picture (I frame), also known as key frame. I frame is the first frame of a group of pictures (GOP). It retains all the information of a scene and only uses the spatial correlation within a single frame for encoding, without using temporal correlation. It can be understood as the complete preservation of this frame, and only the data of this frame is needed for decoding.
[0079] A forward predictive-frame (P-frame) is a coded image that compresses the amount of data transmitted by fully utilizing the temporal redundancy information that is lower than that of the previously coded frames in the image sequence. It is also called a predicted frame. A P-frame represents the difference between this frame and the previous key frame (or P-frame). When decoding, the previously cached picture needs to be superimposed with the difference defined by this frame to generate the final picture.
[0080] Bidirectional interpolated prediction frame (B frame) is a coded image that takes into account the temporal redundant information between the previous coded frames in the source image sequence and the subsequent coded frames in the source image sequence to compress the amount of transmitted data. It is also called a bidirectional difference frame. In short, the B frame records the difference between the current frame and the previous and next frames, and has a high compression rate.
[0081] Presentation Time Stamps (PTS): This timestamp is used to tell the player when to display the data of this frame.
[0082] Decoder Time Stamps (DTS): The purpose of this timestamp is to tell the player when to decode the data of this frame.
[0083] Real Time Messaging Protocol (RTMP): A network protocol designed for real-time data communication. It is mainly used for audio, video and data communication between Flash / AIR platform and streaming media / interactive servers that support RTMP protocol.
[0084] Streaming media container (Flash Video, FLV): FLV streaming media format is a video format developed with the launch of Flash MX. Because it forms extremely small files and loads very quickly, it makes it possible to watch video files on the Internet. Its appearance effectively solves the problem that after the video files are imported into Flash, the exported files become large in size and cannot be used well on the Internet. This container format is simply encapsulated and is currently the most commonly used container format encapsulation protocol for live streaming viewing.
[0085] The following is a brief introduction to the design concept of the embodiment of the present application:
[0086] For video playback, audio and video synchronization is a technical indicator that needs to be guaranteed as much as possible. To this end, during the playback of video resources, the corresponding display time of the video files and audio files in the video resources is usually used to ensure that the corresponding video frames and audio frames in the video and audio files are played synchronously.
[0087] However, the decoding logic of audio files and video files may be different in some cases. Generally speaking, the decoding order of audio files is linear, that is, in an audio file, the decoding time and display time corresponding to the same audio frame are equal, and any audio frame can be displayed immediately after decoding. For video files, the decoding order needs to be determined based on whether there are B frames in the video file to determine whether it is a linear order.
[0088] From the above explanations, it can be seen that after collecting video data, in order to reduce the size of the memory occupied by the encoded video file, it is generally not adopted to encode all the video data in the I frame format, but it is chosen to encode and save part of the video data in the P frame or B frame format.
[0089] For a video file, when all the video frames it includes are encoded as I frames and P frames, then when decoding the video file, the decoding time corresponding to each video frame can be the same as its corresponding display time. This is because, for I frames, their decoding can be completed only based on the content they contain, and for P frames, their decoding only needs to be completed based on the content they contain and the content of the previous frame that has been decoded. In this way, each video frame in the video file can be decoded in the order of display.
[0090] However, if there are B frames in the video file, due to the special nature of B frames in decoding, the decoding order of the video file containing B frames will be different from the display order, and the corresponding decoding time will also be different from the display time. This is because the decoding of B frames needs to be based on the content of the two frames before and after the B frame to complete. For example, Figure 2 As shown, assuming that the 1st to 4th video frames in a video file are I frame, B frame, B frame, and P frame respectively, then, since the decoding of B frame can only be realized based on the content of the two frames before and after, when decoding the B frame, it is necessary to wait for the video frame after it to be decoded before the decoding of the B frame can be completed. In other words, the display order of I frame, B frame, B frame, and P frame is 1, 2, 3, 4, but the corresponding decoding order is 1, 3, 4, 2. This results in that when the video file is played, there will be a corresponding offset time to ensure that when each video frame is displayed, the content of the decoded video frame can be obtained.
[0091] In the field of live broadcasting, in order to ensure the real-time transmission of live broadcast content and reduce the data transmission pressure during the live broadcasting process, when encoding and transmitting videos, most of them will use B frames to improve the video compression rate and reduce the data transmission bandwidth occupied by video files. The transmission format requires that for video files encoded in the B frame format, the display time of each video frame is the sum of the decoding time and the offset time. This results in that when the video is played on the playback end, the display time of the first video frame of the video file is compared with the display time of the first audio frame of the audio file. There is a difference in offset time. When the offset time is large, the phenomenon of audio and video being out of sync will be more obvious.
[0092] In view of this, an embodiment of the present application provides a method for synchronizing audio and video, which utilizes the time offset information of a video file in a video resource to modify the audio time information of an audio file corresponding to the video file, so that the audio display time of each audio frame in the audio file can be the same as the video display time of the corresponding video frames in the video file, thereby performing subsequent audio and video synchronization processing.
[0093] Specifically, after obtaining the video file and the audio file corresponding to the video file, the corresponding time offset information is obtained from the video file, wherein the time offset information includes the deviation between the video display time and the video decoding time corresponding to each video frame in the video file; after completing the acquisition of the offset information, the audio time information in the audio file can be modified according to the time offset information to obtain the corresponding target time information, wherein the audio time information includes the audio display time and the audio decoding time corresponding to each audio frame in the audio file, and the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file. In this way, after completing the modification of the audio time information of the audio file, the video file and the audio file can be processed for audio and video synchronization based on the video display time corresponding to each video frame and the target time information corresponding to the audio file.
[0094] The following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0095] See also Figure 3 , is a schematic diagram of a possible application scenario provided in an embodiment of the present application, in which a terminal device 301 and a server 302 may be included.
[0096] The terminal device 301 may be a mobile phone, a tablet computer (PAD), a personal computer (PC), a wearable device, a vehicle-mounted terminal, or a camera, a camcorder, or the like.
[0097] Server 302 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0098] The server 302 may include one or more processors 3021, a memory 3022, and an I / O interface 3023 for interacting with a terminal. In addition, the server 302 may also be configured with a database 3024, which may be used to store video resources such as video files and audio files. The memory 3022 of the server 302 may also store program instructions of the method for synchronizing audio and video provided in the embodiment of the present application, which, when executed by the processor 3021, may be used to implement the steps of synchronizing audio and video provided in the embodiment of the present application, so as to achieve audio and video synchronization when playing video resources.
[0099] The terminal device 301 and the server 302 may be directly or indirectly connected to each other through one or more communication networks 303. The communication network 303 may be a wired network or a wireless network, for example, a mobile cellular network or a wireless fidelity (Wireless-Fidelity, WIFI) network, or other possible networks, which are not limited in the present application.
[0100] It should be noted that each method in the embodiments of the present application can be executed by an electronic device, which may be a terminal device 301 or a server 302, that is, each method can be executed separately by the terminal device 301 or the server 302.
[0101] For example, when the terminal device 301 independently executes the audio and video synchronization method provided in the present application, the terminal device 301 can obtain the time offset information corresponding to each video frame after the video data is encoded when encoding the video data and the audio data, and then based on the time offset information, modify the display time and decoding time corresponding to each audio frame in the audio file during the encoding of the audio data. In this way, when the video resource is subsequently played, the audio and video synchronization processing can be performed according to the display time of the video file and the audio file.
[0102] For another example, when the server 302 executes the audio and video synchronization method provided by the present application, the server 302 can decode the acquired video files and audio files, re-encode and re-package them, and in the process of re-encoding and re-packaging, use the time offset information in the video file to modify the audio time information in the audio file to obtain the corresponding target time information, and perform audio and video synchronization processing on the video file and the audio file according to the video display time corresponding to each video frame in the video file and the target time information.
[0103] It should be noted that Figure 3 What is shown is just an example. In fact, the number and communication mode of terminal devices and servers are not limited and are not specifically limited in the embodiments of the present application.
[0104] The following describes the audio and video synchronization method provided by the exemplary embodiment of the present application in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this regard.
[0105] See also Figure 4 , which is a flow chart of a method for synchronizing audio and video provided in an embodiment of the present application. The method is applied to a network device, which may be the server proposed above, or a switching device, etc., wherein the server may include an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server, etc., and the switching device may include switches, routers, firewalls, bridges, hubs, gateways and other devices, and the present application does not limit the specific contents thereof.
[0106] In order to facilitate the subsequent description of the method, the network device will be used as the execution subject of any method. Figure 4 The implementation steps of the method are introduced in this paper. Figure 4 As shown, the specific implementation steps of this method are as follows:
[0107] Step S401: Acquire a video file and an audio file corresponding to the video file.
[0108] When the server obtains the video file and the corresponding audio file, it can obtain the corresponding video file and audio file from the video resources stored in the server, or obtain the video file and audio file transmitted in real time from the terminal device connected to the server. This application does not elaborate on the process of obtaining the video file and audio file from the internal storage, but the process of obtaining the real-time video resource from the terminal device requires some corresponding introduction.
[0109] From the foregoing, it can be seen that the method provided in the embodiment of the present application can solve the problem of audio and video asynchrony caused by B frames during the video playback stage during the live broadcast process. First, the live broadcast scenario will be described:
[0110] like Figure 5 As shown, the live broadcast process can be roughly divided into the push stream stage and the pull stream stage. The push stream stage refers to the process in which the object initiating the live broadcast sends the obtained video resources to the network device after completing the acquisition, encoding and packaging of the video resources; and the corresponding live broadcast pull stream stage refers to the process in which the object watching the live broadcast obtains the corresponding live broadcast content from the network device. When it is necessary to explain, the network device can be a content distribution network or other forms such as distributed cloud storage.
[0111] After clarifying the basic process of live broadcast, it can be known that the video files and audio files obtained by the server are the video files and audio files obtained by the party conducting the live broadcast after collecting relevant data through its video equipment and audio equipment and encoding and packaging them.
[0112] Step S402: Based on the obtained video file, corresponding time offset information is obtained, wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file.
[0113] After the server obtains the video file, the corresponding time offset information can be obtained from the video file. The time offset information includes the deviation between the poem display time corresponding to each video frame contained in the video file and the video decoding time.
[0114] Optionally, in a live broadcast scenario, when the communication protocol used in the live broadcast streaming process is RTMP, both the video file and the audio file are encapsulated in a streaming media encapsulation format (FLV container). Figure 6 As shown, the encapsulated file of the obtained video data consists of a file header (flie header) and a file body (file body). Among them, the file body of FLV contains multiple tags. Tags can generally be divided into three types: script (frame) data type, audio data type, and video data type. In this application, the video data is mainly introduced, so the tag of the video data type will be introduced below.
[0115] In the tag of the video data type, the tag header and tag data, the tag header includes the tag of the video data type of the tag, and the timestamp information corresponding to the frame of video data, the timestamp information includes the time offset value corresponding to the video frame (ie, CTS).
[0116] Therefore, in this case, when the server performs the acquisition of time offset information in step S402, it may perform the following operations:
[0117] See also Figure 7 , is a flow chart of a method for obtaining time offset information provided by an embodiment of the present application, such as Figure 7 As shown, the specific implementation steps of this method are as follows:
[0118] Step S701: based on the tag headers corresponding to the respective video frames in the video file, obtain the corresponding deviation values.
[0119] Step S702: obtaining time offset information based on the obtained deviation values corresponding to each video frame.
[0120] like Figure 8 As shown, a video file may contain multiple video frames. Therefore, when obtaining the time offset information, the server can obtain the corresponding offset value according to the time offset value included in the tag header corresponding to each video frame. Then, the obtained offset values are combined and used as the above-mentioned time offset information for subsequent modification of the audio time information.
[0121] It should be noted that although the time offset value is directly obtained from the tag header in the video file, there are several different ways to calculate the time offset value (cts). One is to calculate it through the display timestamp and decoding timestamp corresponding to the video frame. The specific formula is: cts = (pts-dts) / 90, where the unit of cts is milliseconds; the second is to calculate it through the interval (minigop) between two P frames in the video file. The specific formula of this calculation method is: cts = mingop*1000 / FPS, where FPS represents the video frame rate of the video file.
[0122] After completing the acquisition of the time offset information, the server can perform the following operations on the audio time information of the audio file according to the time offset information:
[0123] Step S403: Based on the time offset information, modify the audio time information in the audio file to obtain corresponding target time information, wherein the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame contained in the audio file, and the audio display time of the first audio frame contained in the target time information is the same as the video display time of the first video frame in the video file.
[0124] As mentioned above, since the video file may contain B-frame video frames, and the video display time (pts) of the video file is the sum of the video decoding time (dts) and the time offset value (cts), therefore, for the corresponding video file and audio file, it may appear, such as Fig. 9 As shown, there is an error of the time offset value between the video display time (0+cts) corresponding to the first video frame and the audio display time (0) corresponding to the first audio frame. Generally speaking, the display time of the entire audio file will be one cts earlier than the display time of the video file.
[0125] Therefore, in order to avoid the above problem, the present application proposes in step S403 that the audio time information in the audio file can be modified according to the above obtained time offset information, so that in the modified audio file, the audio display time of the first audio frame is the same as the video display time of the first video frame in the video file. In this way, the display time of the audio file is shifted backward by a cts unit as a whole, so that the display time of the audio file and the video file are aligned as much as possible.
[0126] Optionally, regarding the timing of modifying the audio time information of the audio file, the server may choose to perform the modification during the re-encoding and packaging process after decoding the acquired video file and audio file. Specifically, the server completes the modification of the audio time information by performing the following operations.
[0127] See also Fig.10 , is a flowchart of a method for modifying audio time information provided by an embodiment of the present application, such as Fig.10 As shown, the specific implementation steps of this method are as follows:
[0128] Step S1001: re-encoding the audio file based on a preset encoding format to obtain a corresponding intermediate file;
[0129] Step S1002: based on the time offset information, modify the audio time information in the intermediate file to obtain corresponding target time information.
[0130] Among them, the preset encoding format can be the video encoding format required by the server when it is configured to save video resources by itself, or it can be the video encoding format used for data transmission agreed upon by the server and the live streaming end. This application does not impose any restrictions on this.
[0131] When re-encoding an audio file, the server first needs to decode the audio file based on the encoding format of the acquired audio file to obtain the corresponding audio data, and then re-encode the audio data according to the preset encoding format to obtain the corresponding intermediate file. In the process of obtaining the intermediate file, the server can modify the audio time information in the intermediate file at this stage using the time offset information obtained above, and finally obtain the corresponding target time information. And it should be noted that after completing the modification of the audio time information, the server also needs to encapsulate it so that the subsequent data transmission operation can be carried out smoothly, and the recipient can also perform the corresponding playback operation. Therefore, if Fig.11 As shown, the timing of modifying the audio time information in the audio file is the process of re-encoding and packaging the audio file after decoding.
[0132] After clarifying the modification timing of the audio time information, the specific modification operation of the audio time information will be introduced next.
[0133] Optionally, when modifying the audio time information of the audio file, the server can fuse the deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information to the audio display time and audio decoding time corresponding to each audio frame in the audio file, so as to obtain the corresponding target time information. At this time, the target time information includes: the target audio display time and the target audio decoding time corresponding to each audio frame in the audio file, and each target audio display time is the same as the video display time of the corresponding video frame.
[0134] For example, the understanding of the above method can be combined with Fig.12 If Fig.12 As shown, for a video resource, each frame will correspond to a pair of video frames and audio frames, and each video frame and audio frame will correspond to their respective video display time (v-pts), video decoding time (v-dts) and audio display time (m-pts), audio decoding time (m-dts). Since the value of the video display time of each video frame is the sum of the video decoding time and the time deviation value, and the value of the audio display time of the audio frame is the same as the audio decoding time, when modifying the audio time information of the audio file, it is necessary to merge the time deviation values corresponding to each video frame into the audio display time and audio decoding time of each corresponding audio frame, so as to obtain the corresponding target audio display time and target audio decoding time; in this way, the target audio display time corresponding to each audio frame is the same as the video display time corresponding to each video frame.
[0135] For some video files, the deviations between the video display time and the video decoding time corresponding to each video frame contained in a video file are the same. Therefore, in this case, when modifying the audio time information in the audio file, the following method can be used to modify the audio time information:
[0136] See also Fig.13 , is a flowchart of a method for modifying audio time information provided by an embodiment of the present application, such as Fig.13 As shown, the specific implementation steps of this method are as follows:
[0137] Step S1301: Obtain the deviation value between the video display time and the video decoding time corresponding to the first video frame in the time deviation information;
[0138] Step S1302: Based on the deviation value, all audio display times and audio decoding times contained in the audio file are modified to obtain corresponding target audio display times and target audio decoding times.
[0139] Among them, the target audio display time is positively correlated with the corresponding audio display time and the deviation value, and the target audio decoding time is also positively correlated with the corresponding audio decoding time and the deviation value.
[0140] For video files with the same deviation value for each video frame, the server can directly obtain the deviation value between the video display time and the video decoding time corresponding to the first video frame in the video file, and use the deviation value to modify the audio display time and audio decoding time in the audio file. In this way, the corresponding target audio display time and target audio decoding time can be obtained.
[0141] As for the modification method, a simple and feasible method is to directly add the deviation value corresponding to the video file to the value of the audio display time, so as to ensure that the target audio display time obtained is the same as the video display time of each corresponding video frame in the video file. At the same time, since the display time and decoding time of each audio frame in the audio file are the same, when modifying the audio display time, the audio decoding time also needs to be modified in the same way.
[0142] The above describes that the server modifies the audio file directly based on the time offset information in the video file. In the above method, regardless of whether there is a B-frame video frame in the video file, the server can merge the offset value in the video file into the audio time information of the audio file. The only difference is that when there is no B-frame video frame, the offset value is 0, and when there is a B-frame video frame, the offset value is the offset between the video display time and the video decoding time.
[0143] Optionally, when the server distinguishes based on whether there is a B-frame video frame in the video file, the server may perform the following operations based on different situations:
[0144] When the server determines that there is a B frame (i.e., a bidirectional predictive interpolation coded frame) in the video file, it is necessary to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information;
[0145] When the server determines that there are no B frames (i.e., bidirectionally predictive interpolated coded frames) in the video file, it can be determined that the deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information in the video file is zero. At this time, the server can directly use the audio time information as the target time information without making any further modifications to the audio time information.
[0146] In this way, by distinguishing whether there is a B frame in the video file, the server can use different methods to modify the audio time information, and then complete the subsequent audio and video synchronization operation.
[0147] Step S404: Based on the video display time corresponding to each video frame and the target time information, perform audio and video synchronization processing on the video file and the audio file.
[0148] After obtaining the matching audio display time and video display time, the server can perform audio and video synchronization processing on the video file and the audio file based on this information.
[0149] In this solution, by using the time offset information in the video file, the audio time information in the audio file is modified, and the overall audio display time of the audio file is pushed back, ensuring that the audio display time of the first frame of the audio file is the same as the video display time of the first frame of the video file, eliminating the problem of audio and video asynchrony that may occur during the playback of video resources. In the subsequent audio and video synchronization process, the synchronous playback of video and audio can be completed through the same display time corresponding to each frame of the audio file and the video file, without discarding or fast-forwarding the audio frames in the audio file, which maximizes the integrity of the audio and video files and improves the playback effect of video resources.
[0150] On the other hand, by implementing the audio and video synchronization method on the network device side, there is no need for video acquisition equipment or video playback equipment to perform separate audio and video synchronization processing, which reduces the cost of popularizing the method. The video playback quality is improved without the user having to perform additional operations, thereby improving the user experience.
[0151] The above introduces a method for performing audio and video synchronization processing on audio files and video files on the network device side. Based on the same inventive concept, the present application also provides a method for audio and video synchronization applied to the video acquisition end. The execution subject of the method can be performed by different terminal devices, such as mobile phones, tablet computers, personal computers, etc. For the sake of convenience of explanation, the method will be introduced below by taking the execution subject as an example of a terminal device.
[0152] See also Fig.14 , is a flow chart of a method for synchronizing audio and video provided in an embodiment of the present application, such as Fig.14 As shown, the specific steps of this method are as follows:
[0153] Step S1401: Collect video data and audio data corresponding to the video data.
[0154] When the terminal device executes step S1401, the terminal device can complete the acquisition of video resources through the camera and recording equipment contained in itself, or complete the acquisition of video data and audio data through an external device connected to itself in communication, and this application does not limit this. For example, the terminal device can complete the shooting of the screen through a camera, and then use these shooting data as video data, or obtain the interface displayed on the screen of the terminal device itself as corresponding video data, etc. The corresponding audio data can also be obtained in a similar manner, and this application does not limit this.
[0155] Step S1402: Based on the video data, corresponding time offset information is obtained, wherein the time offset information includes: a deviation between a video display time and a video decoding time corresponding to each video frame in the video data.
[0156] After completing the acquisition of the video data and the audio data, the terminal device can perform corresponding audio and video synchronization operations on these data to improve the playback effect of subsequent video playback.
[0157] To this end, the terminal device may first obtain corresponding time offset information based on the video data, where the time offset information includes the deviation between the video display time and the video decoding time corresponding to each video frame in the video data.
[0158] Optionally, in order to obtain the above time offset information, the terminal device may perform the following operations:
[0159] Specifically, after obtaining the video data, the terminal device can encode the video data based on the streaming media encapsulation format to obtain the corresponding video file. The specific format of the video file can be found in the introduction of the video file using the streaming media encapsulation format by the above-mentioned network device, and no further details will be given here.
[0160] Then, the terminal device can obtain the deviation value corresponding to each video frame based on the tag header (TagHeader) corresponding to each video frame in the video file, and the deviation value indicates the deviation between the video display time and the video decoding time. It should be noted that although the video display time of the video frame is used to guide the time requirement for the video playback end to display the video frame, its generation process is performed when the video data is encoded. Therefore, after the above encoding process of the video data, each video frame also obtains its corresponding video display time and video encoding time, as well as its corresponding deviation value.
[0161] Finally, by obtaining the deviation values corresponding to the respective video frames, the terminal device can correspondingly obtain the required time offset information.
[0162] Step S1403: Based on the time offset information, the audio time information corresponding to the audio data is modified to obtain corresponding target time information, wherein the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame contained in the target time information is the same as the video display time of the first video frame in the video data.
[0163] After the time offset information is obtained, the modification of the audio time information in the audio data is similar to the modification method applied to the network device. The specific operation can refer to the above content and will not be repeated here.
[0164] Step S1404: Based on the video display time corresponding to each video frame and the target time information, perform audio and video synchronization processing on the video data and the audio data.
[0165] In this way, when the video data and audio data are collected and encoded at the video acquisition end, the modification of the audio time information can be completed synchronously without the need for subsequent processing on the network device side or the video playback end, thereby achieving the best audio and video synchronization effect.
[0166] After completing the audio and video synchronization processing of the video data and audio data, the terminal device can send the processed video data and audio data to the network device to complete the push streaming stage of the live broadcast process. After receiving the video files and audio files that have completed the above audio and video synchronization processing, the subsequent content distribution network does not need to perform the above process again. Figure 4 Instead of the audio and video synchronization operation shown, the corresponding video file and audio file are sent to the video player directly based on the stream pulling request of the video player.
[0167] The above introduces the methods of audio and video synchronization applied to network devices and applied to video acquisition terminals respectively. It should be noted that the above-mentioned audio and video synchronization operations can also be applied to the video acquisition terminal. At the video acquisition terminal, the method of audio and video synchronization is the same as the method of audio and video synchronization applied in the network device. Therefore, this application will not go into details about this.
[0168] After completing the above introduction of each method, the following will introduce the audio and video synchronization method applied to the network device in combination with a specific application scenario.
[0169] For example, Fig.15A As shown, in a live broadcast scenario, after the live broadcast host completes the acquisition of the video recording of the live broadcast scene through the camera, or after the live broadcast host completes the acquisition of the screen content, the terminal device can encode and package these video data and audio data, and then upload these live contents to the network device through the RTMP communication protocol.
[0170] At this time, the network device can use the time offset information contained in the video file to modify the audio time information in the audio file during the re-encoding and packaging process after decoding these live contents, so that the display time of each audio frame in the audio file is shifted backward as a whole. Fig. 15B As shown, the display time of the first video frame of the video file can be the same as the display time of the first audio frame of the audio file, avoiding the black screen or audio and video asynchronism that may occur at the beginning of the playback when playing live content.
[0171] After the video player sends a live streaming request to the network device, the network device can transmit the corresponding video file and audio file after audio and video synchronization processing to the video player so that the video player can complete the playback of the live content.
[0172] Based on the same inventive concept, the present application also provides a device for synchronizing audio and video, see Fig.16 , which is a structural diagram of a device for audio and video synchronization provided in an embodiment of the present application. The device may be the above-mentioned server or a chip or integrated circuit therein, etc. The device includes a module / unit / technical means for executing the method executed by the server in the above-mentioned method embodiment.
[0173] Exemplarily, the device 1600 includes:
[0174] The first acquisition module 1601 is used to acquire a video file and an audio file corresponding to the video file; based on the video file, acquire corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file;
[0175] The first processing module 1602 is used to modify the audio time information in the audio file based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame included in the audio file; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file;
[0176] The first synchronization module 1603 is used to perform audio and video synchronization processing on the video file and the audio file based on the video display time corresponding to each video frame and the target time information.
[0177] In a possible implementation manner, the first processing module 1602 is configured to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0178] The deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is respectively integrated into the audio display time and the audio decoding time corresponding to each audio frame in the audio file to obtain the corresponding target time information;
[0179] The target time information includes: a target audio display time and a target audio decoding time corresponding to each audio frame in the audio file; each target audio display time is the same as a video display time of a corresponding video frame.
[0180] In a possible implementation manner, the deviations between the video display time and the video decoding time corresponding to each video frame in the video file are the same;
[0181] The first processing module 1602 is used to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically for:
[0182] Obtaining a deviation value between a video display time and a video decoding time corresponding to a first video frame in the time offset information;
[0183] Based on the deviation value, all audio display times and audio decoding times contained in the audio file are modified to obtain corresponding target audio display times and target audio decoding times; wherein, the target audio display time and the corresponding audio display time are positively correlated with the deviation value, and the target audio decoding time and the corresponding audio decoding time are positively correlated with the deviation value.
[0184] In a possible implementation, the video file and the audio file are both packaged in a streaming media package format;
[0185] Then, when the first acquisition module 1601 is used to acquire the corresponding time offset information based on the video file, it is specifically used to:
[0186] Based on the tag headers corresponding to the respective video frames in the video file, obtaining corresponding deviation values;
[0187] The time offset information is obtained based on the deviation values corresponding to the respective video frames.
[0188] In a possible implementation manner, the first processing module 1602 is configured to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0189] If it is determined that there is a bidirectional predictive interpolation coding frame in the video file, then based on the time offset information, the audio time information in the audio file is modified to obtain corresponding target time information;
[0190] If it is determined that there is no bidirectional predictive interpolation coding frame in the video file, the deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is determined to be zero, and the audio time information is used as the target time information.
[0191] In a possible implementation manner, the first processing module 1602 is configured to modify the audio time information in the audio file based on the time offset information to obtain the corresponding target time information, specifically to:
[0192] Based on a preset encoding format, the audio file is re-encoded to obtain a corresponding intermediate file; based on the time offset information, the audio time information in the intermediate file is modified to obtain corresponding target time information.
[0193] Based on the same inventive concept, the present application also provides a device for synchronizing audio and video, see Fig.17 , which is a structural diagram of a device for audio and video synchronization provided in an embodiment of the present application. The device may be the above-mentioned terminal device or a chip or integrated circuit therein, etc. The device includes a module / unit / technical means for executing the method executed by the server in the above-mentioned method embodiment.
[0194] Exemplarily, the device 1700 includes:
[0195] The second acquisition module 1701 is used to collect video data and audio data corresponding to the video data; based on the video data, obtain corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame in the video data;
[0196] The second processing module 1702 is used to modify the audio time information corresponding to the audio data based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video data;
[0197] The second synchronization module 1703 is used to perform audio and video synchronization processing on the video data and the audio data based on the video display time corresponding to each video frame and the target time information.
[0198] In a possible implementation manner, when the second acquisition module 1701 is used to acquire the corresponding time offset information based on the video data, it is specifically used to:
[0199] Encoding the video data based on a streaming media encapsulation format to obtain a corresponding video file;
[0200] Based on the tag headers corresponding to the respective video frames in the video file, a corresponding deviation value is obtained; the deviation value is used to indicate the deviation between the video display time and the video decoding time;
[0201] The time offset information is obtained based on the deviation values corresponding to the respective video frames.
[0202] Based on the same inventive concept, the embodiment of the present application also provides an electronic device. In a possible implementation, the electronic device may be a server, such as Figure 3 In this embodiment, the structure of the electronic device 1800 is as follows: Fig.18 As shown, it may include at least a memory 1801 , a communication module 1803 , and at least one processor 1802 .
[0203] The memory 1801 is used to store computer programs executed by the processor 1802. The memory 1801 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0204] The memory 1801 may be a volatile memory, such as a random-access memory (RAM); the memory 1801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1801 may be any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1801 may be a combination of the above memories.
[0205] The processor 1802 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1802 is used to implement the above-mentioned audio and video synchronization method when calling the computer program stored in the memory 1801.
[0206] The communication module 1803 is used to communicate with terminal devices and other servers.
[0207] The specific connection medium between the memory 1801, the communication module 1803 and the processor 1802 is not limited in the embodiment of the present application. Fig.18 In the embodiment, the memory 1801 and the processor 1802 are connected via a bus 1804. The bus 1804 is Fig.18 The connections between the other components are only for illustration and are not intended to be limiting. The bus 1804 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig.18 The diagram shows that only one thick line is used, but this does not mean that there is only one bus or only one type of bus.
[0208] The memory 1801 stores a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to implement the audio and video synchronization method of the embodiment of the present application. The processor 1802 is used to execute the above audio and video synchronization method.
[0209] In another embodiment, the electronic device may also be other electronic devices, such as Figure 3 The terminal device 301 is shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Fig.19As shown, it includes: a communication component 1910, a memory 1920, a display unit 1930, a camera 1940, a sensor 1950, an audio circuit 1960, a Bluetooth module 1970, a processor 1980 and other components.
[0210] The communication component 1910 is used to communicate with the server. In some embodiments, a wireless fidelity (WiFi) module may be included. The WiFi module belongs to a short-range wireless transmission technology. The electronic device can help the object to send and receive information through the WiFi module.
[0211] The memory 1920 can be used to store software programs and data. The processor 1980 executes various functions and data processing of the terminal device 301 by running the software programs or data stored in the memory 1920. The memory 1920 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1920 stores an operating system that enables the terminal device 301 to run. In the present application, the memory 1920 can store an operating system and various application programs, and can also store a computer program for executing the method for matching a target vehicle in an embodiment of the present application.
[0212] The display unit 1930 can also be used to display information input by the object or information provided to the object and a graphical user interface (GUI) of various menus of the terminal device 301. Specifically, the display unit 1930 may include a display screen 1932 disposed on the front of the terminal device 301. The display screen 1932 may be configured in the form of a liquid crystal display, a light emitting diode, etc. The display unit 1930 may be used to display a defect detection interface, a model training interface, etc. in the embodiment of the present application.
[0213] The display unit 1930 can also be used to receive input digital or character information, and generate signal input related to the object setting and function control of the terminal device 301. Specifically, the display unit 1930 may include a touch screen 1931 set on the front of the terminal device 301, which can collect touch operations of objects on or near it, such as clicking a button, dragging a scroll box, etc.
[0214] The touch screen 1931 can be covered on the display screen 1932, or the touch screen 1931 and the display screen 1932 can be integrated to realize the input and output functions of the physical terminal device 301. The integrated touch screen can be referred to as a touch display screen. In this application, the display unit 1930 can display the application and the corresponding operation steps.
[0215] The camera 1940 can be used to capture static images, and the subject can publish the images taken by the camera 1940 through the application. The camera 1940 can be one or more. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1980 to convert it into a digital image signal.
[0216] The physical terminal device may also include at least one sensor 1950, such as an acceleration sensor 1951, a distance sensor 1952, a fingerprint sensor 1953, and a temperature sensor 1954. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0217] The audio circuit 1960, the speaker 1961, and the microphone 1962 can provide an audio interface between the object and the terminal device 301. The audio circuit 1960 can transmit the electrical signal converted from the received audio data to the speaker 1961, which is converted into a sound signal for output. The physical terminal device 301 can also be configured with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 1962 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1960 and converted into audio data, and then the audio data is output to the communication component 1930 to be sent to, for example, another physical terminal device 301, or the audio data is output to the memory 1920 for further processing.
[0218] The Bluetooth module 1970 is used to exchange information with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the physical terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1970 to exchange data.
[0219] The processor 1980 is the control center of the physical terminal device. It uses various interfaces and lines to connect various parts of the entire terminal. It executes various functions of the terminal device and processes data by running or executing software programs stored in the memory 1920 and calling data stored in the memory 1920. In some embodiments, the processor 1980 may include one or more processing units; the processor 1980 may also integrate an application processor and a baseband processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the baseband processor mainly processes wireless communications. It is understandable that the above-mentioned baseband processor may not be integrated into the processor 1980. In the present application, the processor 1980 can run the operating system, application programs, user interface display and touch response, as well as the method for matching the target vehicle in the embodiment of the present application. In addition, the processor 1980 is coupled to the display unit 1930.
[0220] In some possible implementations, various aspects of the method for matching a target vehicle provided in the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the method for matching a target vehicle according to various exemplary embodiments of the present application described above in this specification.
[0221] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0222] The program product of the embodiment of the present application may adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the program product of the present application is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which can be used by or in combination with a command execution system, apparatus, or device.
[0223] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0224] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0225] The computer program for performing the operation of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages. The computer program can be executed entirely on the user electronic device, partially on the user electronic device, as a separate software package, partially on the user electronic device and partially on a remote electronic device, or entirely on the remote electronic device. In the case of a remote electronic device, the remote electronic device can be connected to the user electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (for example, using an Internet service provider to connect through the Internet).
[0226] It should be noted that, although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into multiple units to be embodied.
[0227] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0228] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0229] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0230] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0231] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0232] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for synchronizing audio and video, characterized in that: Applied to network equipment, including: Obtain a video file and an audio file corresponding to the video file; Based on the video file, corresponding time offset information is obtained; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file; Based on the time offset information, the audio time information in the audio file is modified to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame included in the audio file; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file; Based on the video display time corresponding to each of the video frames and the target time information, the video file and the audio file are synchronized with each other.
2. The method according to claim 1, characterized in that The step of modifying the audio time information in the audio file based on the time offset information to obtain corresponding target time information includes: The deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is respectively integrated into the audio display time and the audio decoding time corresponding to each audio frame in the audio file to obtain the corresponding target time information; The target time information includes: a target audio display time and a target audio decoding time corresponding to each audio frame in the audio file; each target audio display time is the same as a video display time of a corresponding video frame.
3. The method according to claim 1, characterized in that The deviations between the video display time and the video decoding time corresponding to each video frame in the video file are the same; Then, based on the time offset information, the audio time information in the audio file is modified to obtain the corresponding target time information, including: Obtaining a deviation value between a video display time and a video decoding time corresponding to a first video frame in the time offset information; Based on the deviation value, all audio display times and audio decoding times contained in the audio file are modified to obtain corresponding target audio display times and target audio decoding times; wherein, the target audio display time and the corresponding audio display time are positively correlated with the deviation value, and the target audio decoding time and the corresponding audio decoding time are positively correlated with the deviation value.
4. The method according to any one of claims 1 to 3, characterized in that: The video file and the audio file are both packaged in a streaming media packaging format; Then, based on the video file, obtaining corresponding time offset information includes: Based on the tag headers corresponding to the respective video frames in the video file, obtaining corresponding deviation values; The time offset information is obtained based on the deviation values corresponding to the respective video frames.
5. The method according to any one of claims 1 to 3, characterized in that: The step of modifying the audio time information in the audio file based on the time offset information to obtain corresponding target time information includes: If it is determined that there is a bidirectional predictive interpolation coding frame in the video file, then based on the time offset information, the audio time information in the audio file is modified to obtain corresponding target time information; If it is determined that there is no bidirectional predictive interpolation coding frame in the video file, the deviation between the video display time and the video decoding time corresponding to each video frame in the time offset information is determined to be zero, and the audio time information is used as the target time information.
6. The method according to any one of claims 1 to 3, characterized in that: The step of modifying the audio time information in the audio file based on the time offset information to obtain corresponding target time information includes: Based on a preset encoding format, re-encoding the audio file to obtain a corresponding intermediate file; Based on the time offset information, the audio time information in the intermediate file is modified to obtain corresponding target time information.
7. A method for synchronizing audio and video, characterized in that: Applied to video acquisition end, including: Collecting video data and audio data corresponding to the video data; Based on the video data, corresponding time offset information is obtained; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame in the video data; Based on the time offset information, the audio time information corresponding to the audio data is modified to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video data; Based on the video display time corresponding to each of the video frames and the target time information, the video data and the audio data are processed in synchronization with each other.
8. The method according to claim 7, characterized in that The acquiring corresponding time offset information based on the video data includes: Encoding the video data based on a streaming media encapsulation format to obtain a corresponding video file; Based on the tag headers corresponding to the respective video frames in the video file, a corresponding deviation value is obtained; the deviation value is used to indicate the deviation between the video display time and the video decoding time; The time offset information is obtained based on the deviation values corresponding to the respective video frames.
9. A device for synchronizing audio and video, characterized in that: Applied to network equipment, the device comprises: The first acquisition module is used to acquire a video file and an audio file corresponding to the video file; based on the video file, acquire corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame included in the video file; A first processing module is configured to modify the audio time information in the audio file based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame included in the audio file; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video file; The first synchronization module is used to perform audio and video synchronization processing on the video file and the audio file based on the video display time corresponding to each video frame and the target time information.
10. A device for synchronizing audio and video, characterized in that: Applied to a video player, the device comprises: The second acquisition module is used to collect video data and audio data corresponding to the video data; based on the video data, obtain corresponding time offset information; wherein the time offset information includes: the deviation between the video display time and the video decoding time corresponding to each video frame in the video data; A second processing module is configured to modify the audio time information corresponding to the audio data based on the time offset information to obtain corresponding target time information; the audio time information includes: the audio display time and audio decoding time corresponding to each audio frame in the audio data; the audio display time of the first audio frame included in the target time information is the same as the video display time of the first video frame in the video data; The second synchronization module is used to perform audio and video synchronization processing on the video data and the audio data based on the video display time corresponding to each video frame and the target time information.
11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
12. A computer-readable storage medium, characterized in that: The method comprises a program code, and when the program code is run on a computing device, the program code is used to make the computing device execute the steps of the method according to any one of claims 1 to 8.
13. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 8 when being executed by a processor.