Audio and video stream processing method, device and equipment for cross-platform live voice chat and medium

By automatically analyzing audio and video stream playback anomalies through a cross-platform live streaming system and generating videos in preset formats, the problem of low efficiency in determining the causes of audio and video stream playback anomalies in online live streaming has been solved, achieving rapid and automatic anomaly identification.

CN119450119BActive Publication Date: 2025-12-19GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310989728.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-07
Publication Date
2025-12-19
Estimated Expiration
2043-08-07

AI Technical Summary

Technical Problem

During live streaming, audio and video stream playback anomalies (such as playback desynchronization, stuttering, frame drops, etc.) make it difficult to determine the cause of the anomalies using existing technologies, requiring manual intervention.

Method used

By using a cross-platform live streaming system, information on abnormal events in real-time audio and video stream playback is obtained. Combined with the replay video of the live stream, information on several audio and video frames is analyzed to generate a video in a preset format and automatically analyze the cause of the abnormality.

Benefits of technology

It can quickly identify the cause of abnormal audio and video stream playback without manual intervention, thus improving the efficiency of identifying the cause of the abnormality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119450119B_ABST
    Figure CN119450119B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network live broadcast, and discloses a cross-platform live broadcast microphone connection audio and video stream processing method and device, electronic equipment and a storage medium, the method saves an audio and video file in a second live broadcast platform server, extracts audio frame information and video frame information from the audio and video file according to real-time microphone connection audio and video stream playing abnormal event information, and the cause of the real-time microphone connection audio and video stream playing abnormal event can be automatically and quickly determined according to the audio frame information and the video frame information, without manual participation, so that the cause determination efficiency of the audio and video stream playing abnormality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the network live broadcast technical field, and particularly relate to an audio and video stream processing method and device for cross-platform live broadcast and microphone connection, electronic equipment and storage medium. BACKGROUND

[0002] With the progress of network communication technology, network live broadcast has become a new network interactive mode. Network live broadcast refers to a technology in which a host shares live audio and video streams on a network with audiences through a network live broadcast platform. Network live broadcast is a new network industry, which embodies the characteristics of internet openness and sharing, and enables every ordinary person to have the opportunity to show their talents on the network.

[0003] Currently, in the process of network live broadcast, the live audio and video streams have the problem of abnormal playing, such as asynchronous playing, lagging, frame loss and the like. In the prior art, the test of the abnormal playing of the audio and video streams often needs to rely on subjective judgment and artificial operation of people, resulting in low efficiency of determining the cause of the abnormal playing of the audio and video streams. SUMMARY

[0004] Embodiments of the present application provide an audio and video stream processing method and device for cross-platform live broadcast and microphone connection, electronic equipment and storage medium, which can improve the efficiency of determining the cause of the abnormal playing of the audio and video streams. The technical solution is as follows:

[0005] In a first aspect, the embodiments of the present application provide an audio and video stream processing method for cross-platform live broadcast and microphone connection, which is applied to a cross-platform live broadcast and microphone connection system. The cross-platform live broadcast and microphone connection system includes a first live broadcast platform server and a second live broadcast platform server. The first live broadcast platform server and the second live broadcast platform server are in communication connection. The first live broadcast platform server pushes the audio and video stream of a first host client to the second live broadcast platform server. The second live broadcast platform server delivers the audio and video stream to a second host client. The method includes the following steps:

[0006] Obtaining real-time microphone connection audio and video stream abnormal event information of the second host client in the microphone connection live broadcast process; wherein the real-time microphone connection audio and video stream abnormal event information includes a start time and an end time of a real-time microphone connection audio and video stream abnormal event; wherein the real-time microphone connection audio and video stream is the audio and video stream that the second host client receives from the second live broadcast platform server in real time;

[0007] Obtaining a microphone connection live broadcast playback video of the second host client; wherein the microphone connection live broadcast playback video is a video obtained after the second host client receives and stores the audio and video stream;

[0008] If the live broadcast playback video is played normally, a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time are determined according to the audio and video file obtained from the second live broadcast platform server; the audio and video file is a file obtained by the second live broadcast platform server receiving and decoding the audio and video stream pushed by the first live broadcast platform server;

[0009] The video in the preset format is obtained according to the plurality of audio frame information and the plurality of video frame information; and if the video in the preset format is synchronized with the video picture of the live broadcast playback video, the video in the preset format is parsed.

[0010] In a second aspect, an audio and video stream processing device for cross-platform live broadcast and live broadcast is provided, and the device comprises:

[0011] An event information obtaining module is configured to obtain real-time live broadcast audio and video stream playing abnormal event information of the second anchor client during the live broadcast; the real-time live broadcast audio and video stream playing abnormal event information comprises a start time and an end time of a real-time live broadcast audio and video stream playing abnormal event; the real-time live broadcast audio and video stream is an audio and video stream received by the second anchor client from the second live broadcast platform server in real time;

[0012] A playback video obtaining module is configured to obtain a live broadcast playback video of the second anchor client; the live broadcast playback video is a video obtained by the second anchor client receiving the audio and video stream and storing the audio and video stream;

[0013] A frame information determining module is configured to, if the live broadcast playback video is played normally, determine a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time according to an audio and video file obtained from the second live broadcast platform server; the audio and video file is a file obtained by the second live broadcast platform server receiving and decoding the audio and video stream pushed by the first live broadcast platform server;

[0014] A video parsing module is configured to obtain a video in a preset format according to the plurality of audio frame information and the plurality of video frame information; and if the video in the preset format is synchronized with the video picture of the live broadcast playback video, the video in the preset format is parsed.

[0015] In a third aspect, an electronic device is provided, which comprises a processor, a memory, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the method of the first aspect are implemented.

[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program; when the computer program is executed by a processor, the steps of the method of the first aspect are implemented.

[0017] The embodiment of the application obtains real-time microphone live audio and video stream playing abnormal event information of a second anchor client in a microphone live broadcast process, wherein the real-time microphone live audio and video stream playing abnormal event information comprises a start time and an end time of a real-time microphone live audio and video stream playing abnormal event, wherein the real-time microphone live audio and video stream is an audio and video stream received by the second anchor client from a second live broadcast platform server in real time; a microphone live broadcast playback video of the second anchor client is obtained, wherein the microphone live broadcast playback video is a video obtained after the second anchor client receives and stores an audio and video stream; if the microphone live broadcast playback video is played without any abnormality, a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time are determined according to an audio and video file obtained from the second live broadcast platform server, wherein the audio and video file is a file obtained by the second live broadcast platform server receiving and decoding an audio and video stream pushed by a first live broadcast platform server; a video in a preset format is obtained according to the plurality of audio frame information and the plurality of video frame information; if the video in the preset format is synchronized with a video picture of the microphone live broadcast playback video, the video in the preset format is parsed. The application saves an audio and video file in the second live broadcast platform server, extracts audio frame information and video frame information from the audio and video file according to real-time microphone live audio and video stream playing abnormal event information, obtains a video in a preset format according to the audio frame information and the video frame information, and parses the video in the preset format, so that the cause of the real-time microphone live audio and video stream playing abnormal event can be automatically and quickly determined without human intervention, and the determination efficiency of the cause of the audio and video stream playing abnormality is improved.

[0018] In order to better understand and implement, the technical solutions of the application are described in detail below with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 An application scenario diagram of the audio and video stream processing method for cross-platform live broadcast microphone live broadcast provided by the embodiment of the application is shown.

[0020] Figure 2 A flowchart of the audio and video stream processing method for cross-platform live broadcast microphone live broadcast provided by one embodiment of the application is shown.

[0021] Figure 3 A structural diagram of the audio and video stream processing device for cross-platform live broadcast microphone live broadcast provided by one embodiment of the application is shown.

[0022] Figure 4 A structural diagram of an electronic device provided by one embodiment of the application is shown. DETAILED DESCRIPTION

[0023] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The description herein addresses the exemplary embodiments and is not intended to represent all embodiments in accordance with this application. Rather, they are merely examples in accordance with some aspects of this application as detailed in the appended claims.

[0024] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in this application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any or all possible combinations of one or more of the associated listed items.

[0025] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a temporal sequence. Rather, these terms are used only as distinguishable to refer to one from another instance of a same type. For example, without departing from the scope of this application, a first information can also be termed a second information, and similarly, a second information can also be termed a first information. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to a determination."

[0026] Those skilled in the art can understand that the "client", "terminal", "terminal device" used in the present application includes not only the device with a wireless signal receiver, but also the device with receiving and transmitting hardware, which can communicate bidirectionally on a bidirectional communication link. Such devices can include cellular or other communication devices with or without a multi-line display, a single-line display, or no display; Personal Communications Service (PCS) devices that can combine a voice, data processing, facsimile, and / or data communications capabilities; Personal Digital Assistants (PDAs) that can include a radio frequency receiver, a pager, Internet / Intranet access, a Web browser, a calendar, and / or a Global Positioning System (GPS) receiver; conventional laptop and / or palmtop computers or other devices that have a radio frequency receiver; and / or any other device that has a radio frequency receiver. The "client", "terminal", "terminal device" used herein can be portable, transportable, installed in a vehicle (air, marine, and / or land), or suitable and / or configured to run locally, and / or in a distributed manner, on Earth and / or any other location in space. The "client", "terminal", "terminal device" used herein can also be a communication terminal, an Internet terminal, a music / video playing terminal, such as a PDA, a Mobile Internet Device (MID), and / or a mobile phone with music / video playing function, a smart television, a set-top box, and the like.

[0027] The "server", "client", "service node", and the like referred to in the present application essentially refer to a computer device with the equivalent capability of a personal computer, which is a hardware device with necessary components disclosed in the Von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. A computer program is stored in the memory, the central processing unit calls the program stored in the external memory into the memory to run, executes the instructions in the program, and interacts with the input and output devices, thereby completing a specific function.

[0028] It should be noted that the concept of "server" in the present application can also be extended to the case of a server cluster. According to the principle of network deployment understood by those skilled in the art, the servers should be logically divided, and in physical space, these servers can be independent of each other but can be called through an interface, or can be integrated into a physical computer or a computer cluster. Those skilled in the art should understand this variation and should not be restricted by the implementation of the network deployment of the present application.

[0029] Please refer to Figure 1 , Figure 1 The application scenario of the cross-platform live interaction method provided by the embodiments of the present application is shown in the figure, which includes the first live platform server 101 provided by the embodiments of the present application, and the first anchor client 102 and the first audience client 103 logged in the first live platform server 101. The first anchor client 102 and the first audience client 103 interact through the first live platform server 101; the second live platform server 201 and the second anchor client 202 and the second audience client 203 logged in the second live platform server 201, the second anchor client 202 and the second audience client 203 interact through the second live platform server 201.

[0030] Among them, the first anchor client 102 and the second anchor client 202 refer to the end of sending network live video, which is usually the client used by the anchor (i.e. live anchor user) in the network live.

[0031] The first audience client 103 and the second audience client 203 refer to the end of receiving and watching network live video, which is usually the client used by the audience (i.e. live audience user) watching the video in the network live.

[0032] The hardware pointed to by the client essentially refers to a computer device, specifically, it can be a smart phone, a smart interactive tablet and a personal computer and other types of computer devices. Each client can access the Internet through a known network access method and establish a data communication link with the server.

[0033] The first live platform server 101 and the second live platform server 201 can be business servers, each of which includes or is connected to the related audio data server, video stream server and other servers providing related support of the first live platform and the second live platform, etc., to form a logically associated server cluster to provide services for related terminal devices, such as each client shown in Figure 1 .

[0034] In the embodiments of the present application, the anchor client and the audience client of the same live broadcast platform can join the live broadcast room (i.e., live broadcast channel) created by the corresponding live broadcast platform server. The live broadcast room is a chat room realized by Internet technology, and generally has audio and video control functions. The anchor user performs live broadcast in the live broadcast room through the anchor client, and the audience of the audience client can log in to the live broadcast platform server to enter the live broadcast room to watch the live broadcast.

[0035] In the live broadcast room, the anchor and the audience can interact through voice, video, text and other known online interaction methods. Generally, the anchor user performs a program for the audience in the form of an audio and video stream, and behaviors can also be generated in the interaction process. Of course, the application form of the live broadcast room is not limited to online entertainment, but can also be extended to other related scenarios, such as video conference scenarios, product promotion and sales scenarios, and other scenarios that require similar interaction.

[0036] Specifically, the process of the audience watching the live broadcast is as follows: the audience can click to access the live broadcast application installed on the audience client, and select to enter any live broadcast room, triggering the audience client to load the live broadcast room interface for the audience, and to perform various online interactions.

[0037] The first live broadcast platform server 101 and the second live broadcast platform server 201 are two independent network live broadcast platforms, such as the servers of the first network live broadcast platform and the second network live broadcast platform.

[0038] The first live broadcast platform server 101 and the second live broadcast platform server 201 establish a network connection and perform data interaction through the network. The first live broadcast platform server 101 creates a first virtual anchor client 104 and a first virtual anchor account corresponding to the first virtual anchor client 104 locally. The first virtual anchor client 104 is simulated by the first live broadcast platform server 101 through calculation, and the first virtual anchor account corresponding to the first virtual anchor client 104 is an anchor account that can log in to the first live broadcast platform server 101.

[0039] In the implementation of the anchor cross-platform live broadcast, the first live platform server 101 acquires the live media stream uploaded by the first anchor client 102 logged in the platform. And the first live platform server 101 acquires the second anchor account corresponding to the second anchor client 202 logged in the second live platform server 201, and establishes a binding relationship between the second anchor account and the first virtual anchor account. The first live platform server 101 acquires the live media stream uploaded by the second anchor client 202 from the second live platform server 201 through the binding relationship, and maps and converts the live media stream uploaded by the second anchor client 202 into the media stream of the first virtual anchor client 104 created locally, that is, the second live media stream, so that the second anchor client 202 logged in the second live platform server 201 can live broadcast in the first live platform server 101 as the first virtual anchor client 104. The live media stream of the first virtual anchor client 104 can be set by the first live platform server 101 to be individually viewable / searchable or not individually viewable / searchable. In this embodiment, the first live platform server 101 sets the live media stream of the first virtual anchor client 104 to be not individually viewable / searchable, that is, the live media stream of the first virtual anchor client 104 can only be used for live interaction.

[0040] When the live interaction is enabled, the first live platform server 101 creates a live interaction room, and adds the first virtual anchor client 104 and the first anchor client 102 to the live interaction room, so that the first live media stream corresponding to the first anchor client 102 and the second live media stream corresponding to the first virtual anchor client 104 are both displayed in the live interaction room.

[0041] In one embodiment, the first live platform server 101 acquires the first live media stream corresponding to the first anchor account and the second live media stream corresponding to the first virtual anchor account, mixes the first live media stream and the second live media stream to obtain a mixed live media stream, and delivers the mixed live media stream to the clients that have joined the live interaction room. The clients that have joined the live interaction room created by the first live platform server 101 include the first live client 102 and the first audience client 103 logged in the first live platform.

[0042] The second live platform server 201 implements cross-platform live broadcast in at least two ways:

[0043] The first mode is that the second live broadcast platform server 201 also creates a second virtual anchor client 204 and a second virtual anchor account corresponding to the second virtual anchor client 204 locally. The second virtual anchor client 204 is simulated by the second live broadcast platform server 201 through calculation, and the corresponding second virtual anchor account is an anchor account that can be logged in on the second live broadcast platform server 201. The second live broadcast platform server 201 obtains a first anchor account corresponding to the first anchor client 102 logged in on the first live broadcast platform server 101, and establishes a binding relationship between the first anchor account and the second virtual anchor account. The second live broadcast platform server 201 obtains the live media stream uploaded by the first anchor client 102 from the first live broadcast platform server 101 through the binding relationship, and maps and converts the live media stream uploaded by the first anchor client 102, that is, the first live media stream, to the media stream of the locally created second virtual anchor client 204, so that the first anchor client 102 logged in on the first live broadcast platform server 101 can live broadcast in the second live broadcast platform server 201 as the second virtual anchor client 204, and can also perform a call-in operation. The operation of the second virtual anchor client 204 in the second live broadcast platform server 201 is similar to that of the first virtual anchor client 104 in the first live broadcast platform server 101, and will not be described here.

[0044] The second mode is that the second live broadcast platform server 201 directly obtains the first live media stream and the second live media stream from the first live broadcast platform server 101 through network connection with the first live broadcast platform server 101, mixes the first live media stream and the second live media stream to form a mixed live media stream, and then sends the mixed live media stream to the client that has joined the call-in live broadcast room created by the second live broadcast platform server 201. The client that has joined the call-in live broadcast room created by the second live broadcast platform server 201 includes the second live broadcast client 202 and the second audience client 203 logged in on the second live broadcast platform.

[0045] In the prior art, anchors can only perform voice or video call-in on the same live broadcast platform, the call-in form is relatively single, and the interactive experience of users in the network live broadcast process is affected. For example, in the live broadcast call-in process, anchors of a self-developed live broadcast platform can only perform call-in interaction with anchors of the same self-developed platform, and cannot perform call-in interaction with anchors of other third-party live broadcast platforms. In addition, in the process of call-in on the same live broadcast platform, the audio and video stream playback of live broadcast has problems such as asynchronization, lag, and frame loss.

[0046] In the prior art, the test of audio and video stream playback abnormalities often needs to rely on subjective judgment and artificial operation of people, resulting in low efficiency of determining the cause of audio and video stream playback abnormalities.

[0047] To this end, the embodiment of the present application provides an audio and video stream processing method for cross-platform live broadcast and microphone connection.

[0048] Please refer to Figure 2 , Figure 2 The flowchart of the audio and video stream processing method for cross-platform live broadcast and microphone connection provided by an embodiment of the present application is shown in the figure. The method is applied to a cross-platform live broadcast and microphone connection system, which includes a first live broadcast platform server and a second live broadcast platform server. The first live broadcast platform server and the second live broadcast platform server are in communication connection. The first live broadcast platform server pushes the audio and video stream of a first anchor client to the second live broadcast platform server. The second live broadcast platform server issues the audio and video stream to a second anchor client. The method includes the following steps:

[0049] S10: Real-time microphone connection audio and video stream playing abnormal event information of the second anchor client in the microphone connection live broadcast process is acquired. The real-time microphone connection audio and video stream playing abnormal event information includes the start time and the end time of the real-time microphone connection audio and video stream playing abnormal event. The real-time microphone connection audio and video stream is the audio and video stream that the second anchor client receives from the second live broadcast platform server in real time.

[0050] The audio and video stream is the live broadcast media stream uploaded by the first anchor client to the first live broadcast platform server, which includes an audio stream and a video stream. Specifically, the first anchor client collects audio data of the anchor and encodes the audio data to obtain the audio stream. The first anchor client collects video data of the anchor and encodes the video data to obtain the video stream. The audio stream includes a plurality of audio frames, and the video stream includes a plurality of video frames. The first anchor client marks each audio frame and each video frame with a collection timestamp.

[0051] In the microphone connection live broadcast process, the process in which the second live broadcast platform server issues the audio and video stream to the second anchor client includes that the second live broadcast platform server is built-in with a first SDK and a second SDK. The first SDK is used to decode the audio and video stream to obtain decoded audio frames and video frames. The first SDK sends the decoded audio frames and video frames to the second SDK. The second SDK re-encodes the decoded audio frames and video frames according to a preset encoding protocol to obtain encoded audio and video streams. The encoded audio and video streams are sent to the second anchor client again. The first SDK can be an SDK matched with the first anchor client, which is used to decode the audio and video stream pushed by the first live broadcast platform server. The second SDK can be an SDK matched with the second anchor client, which is used to re-encode the audio and video stream decoded by the first SDK for the second anchor client to play.

[0052] After receiving the audio and video stream issued by the second live broadcast platform server, the second anchor client needs to play the audio and video stream in real time. In order to ensure low latency and smoothness of the audio and video stream playing, the second anchor client will preset a small cache space. Specifically, the second live broadcast platform server issues the audio and video stream to the cache space of the second anchor client, and the cache space can cache a certain number of audio frames and video frames. After the cache space is occupied by the audio frames and the video frames, the audio frames and the video frames are played frame by frame according to the order of arrival time.

[0053] The real-time audio and video stream playing abnormal event includes that the audio and video stream is played out of sync, lagging and frame loss during the live broadcast process. The reasons for the audio and video stream playing out of sync include that the collection time stamps of the audio frames and the video frames are inconsistent, and the time intervals at which the audio frames and the video frames arrive at the cache space of the second anchor client are inconsistent. For example, the cache space can cache 10 audio frames and 10 video frames. The difference between the time at which a certain video frame arrives at the cache space and the time at which the previous video frame arrives at the cache space is greater than the sum of the time intervals of the 10 video frames, while the time intervals at which each audio frame arrives at the cache space are the same. Therefore, when the second anchor client plays the audio frames, the video frame corresponding to a certain audio frame will be displayed lagging, resulting in audio and video playing out of sync. The reasons for the audio and video stream playing lagging include that the audio and video stream is interrupted, the audio and video frames are lost, the collection time stamps of the audio and video frames are out of order, and the time intervals at which the audio and video frames arrive at the cache space of the second anchor client are inconsistent. The reasons for the audio and video stream losing frames include that the audio and video frames are lost when arriving at the second live broadcast platform server, and the audio and video frames are lost when arriving at the cache space of the second anchor client.

[0054] In the embodiments of the present application, the second anchor client can detect whether the real-time audio and video stream playing is abnormal. If it is detected that the playing is abnormal, the real-time audio and video stream playing abnormal event information is generated. Specifically, the second anchor client will sample the real-time audio and video stream at regular intervals or periodically when playing the real-time audio and video stream, and generate sampling log data. The sampling log data includes the collection time stamp of each audio frame or video frame in the real-time audio and video stream. When detecting that the real-time audio and video stream playing is abnormal, a time point before a preset time period before the playing abnormal time point can be selected as the start time point of the real-time audio and video stream playing abnormal event, and a time point after a preset time period after the playing abnormal time point can be selected as the end time point of the real-time audio and video stream playing abnormal event. For example, if the playing abnormal time point is 12:00:00, 11:58:00 can be selected as the start time point and 12:03:00 can be selected as the end time point.

[0055] S20: Obtain the live broadcast playback video of the second anchor client; wherein the live broadcast playback video is a video obtained after the second anchor client receives and stores the audio and video stream.

[0056] In the embodiment of the present application, the second anchor client is preconfigured with an offline cache space, which can store a large number of audio frames and video frames. After the audio and video stream sent by the second live broadcast platform server reaches the offline cache space, the second anchor client aligns and synthesizes each audio frame and each video frame according to the time sequence of reaching the offline cache space, and obtains the live broadcast playback video. The live broadcast playback video is obtained from the second anchor client.

[0057] In an optional embodiment, a real-time audio and video service network RTN can also be arranged between the SDK2 and the second anchor client. The audio and video stream sent by the SDK2 is recorded through the audio and video service network RTN to obtain the live broadcast playback video. The live broadcast playback video is obtained from the audio and video service network RTN.

[0058] S30: If the live broadcast playback video is played without exception, determine the audio frame information and the video frame information corresponding to the start time and the end time from the audio and video file obtained from the second live broadcast platform server; wherein the audio and video file is a file obtained by the second live broadcast platform server receiving and decoding the audio and video stream pushed by the first live broadcast platform server.

[0059] The audio and video file is a file saved by the second live broadcast platform server after decoding the audio and video stream. The audio and video stream includes audio and video frames, private encoding protocols of the audio and video frames, and private data of the audio and video frames. The private encoding protocols include but are not limited to H264 and H265 encoding protocols, and the private data includes but is not limited to the collection timestamp of the audio and video frames, the encoding type of the audio and video frames, the frame header information of the audio and video frames, the frame data length of the audio and video frames, and the frame data content.

[0060] In the embodiment of the present application, if the audio and video in the live broadcast playback video play out of sync, it indicates that the collection time stamps of the audio frame and the video frame are inconsistent, and it is necessary to further determine the link that causes the collection time stamps of the audio frame and the video frame to be inconsistent. It can be that the first anchor client inconsistently stamps the collection time stamps on each audio frame and each video frame, or it can be that the second SDK of the second live broadcast platform server reconstructs the collection time stamps for each audio frame and each video frame. If the audio and video in the live broadcast playback video play in sync, it indicates that the collection time stamps of the audio frame and the video frame are consistent, and it is necessary to further determine the link that causes the time interval of the audio frame and the video frame to be inconsistent when reaching the cache space of the second anchor client. It can be that the time interval of the audio frame and the video frame reaching the first SDK is inconsistent, or it can be that the time interval of the audio frame and the video frame reaching the first SDK is consistent, and the time interval of the audio and video frames reaching the cache space of the second anchor client is different due to processing or link transmission time consumption in the process of the second SDK sending the audio frame and the video frame to the second anchor client.

[0061] If the playing abnormality is audio and video playing lag, it is necessary to determine the reason for the playing abnormality, which is audio and video flow interruption, audio and video frame loss, audio and video collection time stamp disorder, and inconsistent time interval of the audio frame and the video frame reaching the cache space of the second anchor client. If the live broadcast playback video plays with lag, the existing technology is used to troubleshoot the reason from audio and video flow interruption, audio and video frame loss, and audio and video collection time stamp disorder. If the live broadcast playback video plays without lag, it indicates that the time interval of the audio frame and the video frame reaching the cache space of the second anchor client is inconsistent. It is necessary to further determine the link that causes the time interval of the audio frame and the video frame reaching the cache space of the second anchor client to be inconsistent. It can be that the time interval of the audio frame and the video frame reaching the first SDK is inconsistent, or it can be that the time interval of the audio frame and the video frame reaching the first SDK is consistent, and the time interval of the audio and video frames reaching the cache space of the second anchor client is different due to processing or link transmission time consumption in the process of the second SDK sending the audio frame and the video frame to the second anchor client.

[0062] If the playing abnormality is audio and video frame loss, it is necessary to determine the reason for the playing abnormality, which is real-time network fluctuation in the process of the second anchor client, audio and video frame loss when reaching the second live broadcast platform server, and audio and video frame loss when being sent from the second live broadcast platform server to the second anchor client.

[0063] The audio frame information and the video frame information corresponding to the duration between the start time and the end time can be extracted from the audio and video file. The audio frame information includes each audio frame in the duration and the actual time of each audio frame reaching the second live broadcast platform server. The video frame information includes each video frame in the duration and the actual time of each video frame reaching the second live broadcast platform server. Specifically, the actual time of the audio frame reaching the second live broadcast platform server can be the time when the audio frame is decoded, and the actual time of the video frame reaching the second live broadcast platform server can be the time when the video frame is decoded.

[0064] S40: According to the audio frame information and the video frame information, a video in a preset format is obtained. If the video in the preset format is synchronized with the video picture of the mic live broadcast playback video, the video in the preset format is parsed.

[0065] The video in the preset format can be a video that can be opened by a general playback tool or playback software, including but not limited to MP4, AVI, and WMV.

[0066] In the embodiments of the present application, the audio frame information and the video frame information corresponding to the duration between the start time and the end time can be combined into an MP4. Whether the video picture of the corresponding duration in the MP4 and the mic live broadcast playback video is consistent is compared. If they are consistent, the MP4 is parsed to obtain the audio frame information and the video frame information. According to the decoding time stamp of the audio frame and the video frame in the audio frame information and the video frame information, the link where the playback abnormality occurs can be determined.

[0067] If they are not consistent, it means that the video picture of the corresponding duration in the mic live broadcast playback video has frame loss, and the reason for the frame loss needs to be further investigated. Specifically, according to the collection time stamp of the audio frame and the video frame in the audio frame information and the video frame information, if the time interval of the collection time stamp of the adjacent audio frame and the adjacent video frame is fixed, it means that there is no frame loss when the audio and video frames reach the second live broadcast platform server, and there is frame loss when the audio and video frames are sent from the second live broadcast platform server to the second anchor client. If the time interval of the collection time stamp of the adjacent audio frame or the adjacent video frame is not fixed, it means that there is frame loss when the audio and video frames reach the second live broadcast platform server.

[0068] The application embodiment is applied to acquire real-time microphone live audio and video stream playing abnormal event information of the second anchor client in the microphone live broadcast process, wherein the real-time microphone live audio and video stream playing abnormal event information includes a start time and an end time of the real-time microphone live audio and video stream playing abnormal event, the real-time microphone live audio and video stream is an audio and video stream received by the second anchor client from the second live broadcast platform server in real time, the microphone live broadcast playback video of the second anchor client is acquired, the microphone live broadcast playback video is a video obtained after the second anchor client receives and stores the audio and video stream, if the microphone live broadcast playback video is played without any abnormality, a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time are determined according to the audio and video file acquired from the second live broadcast platform server, the audio and video file is a file obtained by the second live broadcast platform server receiving and decoding the audio and video stream pushed by the first live broadcast platform server, a video in a preset format is obtained according to the plurality of audio frame information and the plurality of video frame information, if the video in the preset format is synchronized with the video screen of the microphone live broadcast playback video, the video in the preset format is analyzed. The application saves the audio and video file in the second live broadcast platform server, extracts the audio frame information and the video frame information from the audio and video file according to the real-time microphone live audio and video stream playing abnormal event information, obtains the video in the preset format according to the audio frame information and the video frame information, and analyzes the video in the preset format, so that the cause of the real-time microphone live audio and video stream playing abnormal event can be automatically and quickly determined without manual participation, and the determination efficiency of the audio and video stream playing abnormal cause is improved.

[0069] In one embodiment, the audio and video file includes an audio file and a video file, and before the step S30 of determining the plurality of audio frame information and the plurality of video frame information corresponding to the start time and the end time from the audio and video file acquired from the second live broadcast platform server, steps S301-S302 are further included, which are specifically as follows:

[0070] S301: The second live broadcast platform server decodes the audio and video stream to obtain a plurality of audio frames, a plurality of video frames, a collection time stamp of each audio frame, and a collection time stamp of each video frame, and generates a decoding time stamp of each audio frame and a decoding time stamp of each video frame.

[0071] In the embodiment of the present application, after the first SDK of the second live platform server receives the audio and video stream pushed by the first live platform server, the audio and video stream is decoded to obtain a plurality of audio frames, a plurality of video frames, a collection time stamp of each audio frame, a collection time stamp of each video frame, an encoding type of each audio frame, frame header information, frame data length and frame data content, an encoding type of each video frame, frame header information, frame data length and frame data content, and a frame sequence number and a decoding time stamp are generated for each audio frame and video frame. The frame sequence number is used to uniquely identify the audio and video frame, the decoding time stamp of each audio frame can be regarded as the actual time when each audio frame reaches the second live platform server, and the decoding time stamp of each video frame can be regarded as the actual time when each video frame reaches the second live platform server.

[0072] S302: The second live platform server saves the plurality of audio frames, the collection time stamp of each audio frame and the decoding time stamp of each audio frame as an audio file, and saves the plurality of video frames, the collection time stamp of each video frame and the decoding time stamp of each video frame as a video file.

[0073] In the embodiment of the present application, the audio file can be written in the order of the collection time stamp, the decoding time stamp, the frame sequence number, the encoding type of each audio frame, the frame header information, the frame data length and the frame data content of each audio frame. The video file can be written in the order of the collection time stamp, the decoding time stamp, the frame sequence number, the encoding type of each video frame, the frame header information, the frame data length and the frame data content of each video frame. The second live platform server saves the audio file and the video file in the storage space. By decoding the audio and video stream, the audio and video file can be automatically and quickly obtained.

[0074] In one embodiment, the step of determining a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time from the audio and video file obtained from the second live platform server in step S30 includes steps S311-S312, which are as follows:

[0075] S311: Match a first audio frame whose collection time stamp indicates the same time as the start time and a second audio frame whose collection time stamp indicates the same time as the end time from the audio file, and take the first audio frame, the second audio frame and all the audio frames between the first audio frame and the second audio frame as a plurality of audio frames corresponding to the time within the range from the start time to the end time; obtain the collection time stamp and the decoding time stamp of the plurality of audio frames from the audio file, and determine the plurality of audio frames and the collection time stamp and the decoding time stamp of each audio frame as the plurality of audio frame information;

[0076] S312: match the first video frame consistent with the starting time from the video file indicated by the collection timestamp, match the second video frame consistent with the ending time from the video file indicated by the collection timestamp, take the first video frame, the second video frame, and all video frames between the first video frame and the second video frame as a plurality of video frames corresponding to the time range from the starting time to the ending time; obtain the collection timestamp and the decoding timestamp of the plurality of video frames from the video file, and determine the plurality of video frames and the collection timestamp and the decoding timestamp of each video frame as a plurality of video frame information.

[0077] In the embodiment of the application, the audio and video files can be obtained from the storage space of the second live platform server, and the audio files and video files in the time range between the starting time and the ending time are extracted from the audio and video files. Specifically, by matching the starting time, the ending time, and the collection timestamp in the audio and video files, the first audio frame corresponding to the starting time and the first video frame are determined, and the second audio frame corresponding to the ending time and the second video frame are determined. After determining the first audio frame and the second audio frame, according to the frame serial number, all audio frames between the first audio frame and the second audio frame are obtained by traversing the audio file, and all video frames between the first video frame and the second video frame are obtained by traversing the video file. Then, the collection timestamp and the decoding timestamp corresponding to each audio frame and video frame are obtained.

[0078] By matching the starting time and the ending time with the collection timestamp in the audio and video files respectively, a plurality of audio frame information and a plurality of video frame information corresponding to the starting time and the ending time can be automatically and quickly obtained.

[0079] In one embodiment, the step of obtaining a video of a preset format according to the plurality of audio frame information and the plurality of video frame information in step S40 includes steps S401-S403, which are as follows:

[0080] S401: insert the collection timestamp and the decoding timestamp of each audio frame into the corresponding audio frame supplemental enhancement information to obtain a plurality of target audio frames;

[0081] S402: insert the collection timestamp and the decoding timestamp of each video frame into the corresponding video frame supplemental enhancement information to obtain a plurality of target video frames;

[0082] S403: combine the plurality of target audio frames and the plurality of target video frames into a video of a preset format.

[0083] The supplemental enhancement information (Supplemental Enhancement Information, SEI) provides a method of adding information to the audio stream and the video stream.

[0084] In the embodiment of the present application, the preset format of the video is mp4, the supplementary enhancement information is taken as a field of the frame header in the audio frame and the video frame, the ffmpeg interface is called in sequence, and each target audio frame and each target video frame is written into the mp4. By writing the collection timestamp and the decoding timestamp into the supplementary enhancement information, the collection timestamp and the decoding timestamp can be saved in the mp4, which facilitates the determination of the cause of the abnormal event of the real-time live audio and video stream playback.

[0085] In one embodiment, after step S40, steps S411-S414 are included, and specifically as follows:

[0086] S411: Obtain the decoding timestamp of each video frame and the decoding timestamp of each audio frame, wherein the each video frame and the each audio frame are obtained by analyzing the video in the preset format;

[0087] S412: Calculate the first time difference between the time instants indicated by the decoding timestamps of adjacent audio frames, and calculate the second time difference between the time instants indicated by the decoding timestamps of adjacent video frames;

[0088] S413: If at least one first time difference is greater than a preset threshold, and / or at least one second time difference is greater than a preset threshold, it is determined that the cause of the abnormal event of the real-time live audio and video stream playback is that the time of the audio stream transmitted from the first live streaming platform server to the second live streaming platform server is not synchronized;

[0089] S414: If each first time difference is less than or equal to a preset threshold, and each second time difference is less than or equal to a preset threshold, it is determined that the cause of the abnormal event of the real-time live audio and video stream playback is that the time of the audio stream transmitted from the second live streaming platform server to the second anchor client is not synchronized.

[0090] In the embodiment of the present application, the multimedia analysis tool can be used to analyze the video in the preset format. Since the supplementary enhancement information is carried in the video in the preset format, the decoding timestamp of the audio and video frame can be obtained by analyzing the video in the preset format. By calculating the time interval between the decoding timestamps of adjacent audio frames and adjacent video frames, the cause of the abnormal event of the real-time live audio and video stream playback can be determined.

[0091] Specifically, if each of the first time differences is less than or equal to the preset threshold value, and each of the second time differences is less than or equal to the preset threshold value, it indicates that the time interval of each audio frame and each video frame from the first live platform server to the first SDK of the second live platform server is regular, and the time interval of each audio frame and each video frame from the second SDK of the second live platform server to the cache space of the second anchor client is no longer regular. It is possible that, in the process of sending the audio frame and the video frame by the second SDK to the second anchor client, due to processing or link transmission time consumption, the time interval of the audio and video frames to the cache space of the second anchor client is different. If at least one of the first time differences is greater than the preset threshold value, and / or at least one of the second time differences is greater than the preset threshold value, it indicates that the time interval of each audio frame and each video frame from the first live platform server to the first SDK of the second live platform server is irregular, and it is possible that there is network fluctuation in the process of transmitting the audio and video stream from the first live platform server to the second live platform server.

[0092] By comparing the time interval between the decoding time stamps of adjacent audio frames or adjacent video frames with the preset threshold value, the cause of the real-time mic-in audio and video stream playing abnormal event can be automatically and quickly determined.

[0093] Referring to Figure 3 A structural schematic diagram of an audio and video stream processing device for cross-platform live mic-in is provided for an embodiment of the present application. The device can be realized by software, hardware, or a combination of both to become all or part of a server. The device 5 includes:

[0094] An event information acquisition module 51 is configured to acquire real-time mic-in audio and video stream playing abnormal event information of the second anchor client in the live mic-in process. The real-time mic-in audio and video stream playing abnormal event information includes a start time and an end time of the real-time mic-in audio and video stream playing abnormal event. The real-time mic-in audio and video stream is an audio and video stream received by the second anchor client from the second live platform server in real time.

[0095] A playback video acquisition module 52 is configured to acquire a live mic-in playback video of the second anchor client. The live mic-in playback video is a video obtained by the second anchor client after receiving and storing the audio and video stream.

[0096] A frame information determination module 53 is configured to, if the live mic-in playback video plays normally, determine a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time according to an audio and video file acquired from the second live platform server. The audio and video file is a file obtained by the second live platform server after receiving and decoding the audio and video stream pushed by the first live platform server.

[0097] The video parsing module 54 is used to obtain a video in a preset format based on several audio frame information and several video frame information; if the video in the preset format is synchronized with the video screen of the live broadcast replay video, the video in the preset format is parsed.

[0098] It should be noted that the cross-platform live streaming audio and video stream processing device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the cross-platform live streaming audio and video stream processing method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the cross-platform live streaming audio and video stream processing device and the cross-platform live streaming audio and video stream processing method provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0099] Please see Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Figure 4 As shown, the computer device 21 may include: a processor 210, a memory 211, and a computer program 212 stored in the memory 211 and capable of running on the processor 210, such as a cross-platform live streaming audio and video stream processing program; when the processor 210 executes the computer program 212, it implements the steps in the above embodiments.

[0100] The processor 210 can include one or more processing cores. The processor 210 connects various parts within the computer device 21 by various interfaces and lines, executes various functions of the computer device 21 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 211, and calling data in the memory 211. Optionally, the processor 210 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programable logic array (PLA). The processor 210 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 210, but can be implemented by a separate chip.

[0101] The memory 211 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory 211 includes a non-transitory computer-readable storage medium. The memory 211 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 211 can include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area can store data related to the above-mentioned various method embodiments, etc. The memory 211 can also be at least one storage device located away from the above-mentioned processor 210.

[0102] The embodiments of the present application also provide a computer storage medium, which can store a plurality of instructions. The instructions are suitable for being loaded and executed by a processor to execute the method steps of the above-mentioned embodiments. The specific execution process can be referred to the specific description of the above-mentioned embodiments, which will not be described here.

[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0104] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0105] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0106] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the above-described apparatus / terminal device embodiments are only schematic. The division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0107] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0108] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0109] If the integrated module / unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form.

[0110] The present application is not limited to the above-described embodiments, and various modifications or changes can be made to the present application without departing from the spirit and scope of the present application. If the modifications and changes belong to the scope of the claims of the present application and the equivalent technical scope, the present application also intends to include the modifications and changes.

Claims

1. An audio and video stream processing method for cross-platform live streaming and microphone connection, applied to a cross-platform live streaming and microphone connection system, the cross-platform live streaming and microphone connection system comprising a first live streaming platform server and a second live streaming platform server, the first live streaming platform server and the second live streaming platform server being communicatively connected, the first live streaming platform server pushing an audio and video stream of a first host client to the second live streaming platform server, and the second live streaming platform server delivering the audio and video stream to a second host client; characterized in that, comprising the following steps: obtaining real-time live audio and video stream playing abnormal event information of the second anchor client in the live audio and video process; wherein the real-time live audio and video stream playing abnormal event information includes the start time and the end time of the real-time live audio and video stream playing abnormal event; wherein the real-time live audio and video stream is the audio and video stream received by the second anchor client from the second live platform server in real time; obtaining the live audio and video playback video of the second anchor client; wherein the live audio and video playback video is the video obtained by the second anchor client after receiving and storing the audio and video stream; if the live audio and video playback video plays normally, determining a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time according to the audio and video file obtained from the second live platform server; wherein the audio and video file is the file obtained by the second live platform server after receiving and decoding the audio and video stream pushed by the first live platform server; obtaining a video in a preset format according to a plurality of audio frame information and a plurality of video frame information; if the video in the preset format is synchronized with the video picture of the live audio and video playback video, analyzing the video in the preset format to obtain the decoding time stamp of each video frame and the decoding time stamp of each audio frame; wherein each video frame and each audio frame is obtained by analyzing the video in the preset format; calculating the first time difference between the time indicated by the decoding time stamp of adjacent audio frames, and calculating the second time difference between the time indicated by the decoding time stamp of adjacent video frames; if at least one of the first time difference is greater than a preset threshold, and / or at least one of the second time difference is greater than the preset threshold, it is determined that the reason for the real-time live audio and video stream playing abnormal event is that the time of the audio and video stream from the first live platform server to the second live platform server is not synchronized.

2. The cross-platform live audio and video stream processing method of claim 1, wherein: the audio and video file includes an audio file and a video file; before the step of determining a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time according to the audio and video file obtained from the second live platform server, comprising: the second live platform server decodes the audio and video stream to obtain a plurality of audio frames, a plurality of video frames, a collection time stamp of each audio frame, and a collection time stamp of each video frame, and generates a decoding time stamp of each audio frame and a decoding time stamp of each video frame; the second live platform server saves a plurality of audio frames, a collection time stamp of each audio frame, and a decoding time stamp of each audio frame as the audio file, and saves a plurality of video frames, a collection time stamp of each video frame, and a decoding time stamp of each video frame as the video file.

3. The cross-platform live audio and video stream processing method of claim 2, wherein: The step of determining a plurality of audio frame information and a plurality of video frame information corresponding to the start time and the end time according to the audio and video file obtained from the second live platform server comprises: matching a first audio frame in which the time indicated by the collection timestamp is consistent with the start time and a second audio frame in which the time indicated by the collection timestamp is consistent with the end time from the audio file, taking the first audio frame, the second audio frame, and all audio frames between the first audio frame and the second audio frame as a plurality of audio frames corresponding to the time within the range from the start time to the end time; obtaining the collection timestamp and the decoding timestamp of a plurality of audio frames from the audio file, and determining the plurality of audio frames and the collection timestamp and the decoding timestamp of each audio frame as the plurality of audio frame information; matching a first video frame in which the time indicated by the collection timestamp is consistent with the start time and a second video frame in which the time indicated by the collection timestamp is consistent with the end time from the video file, taking the first video frame, the second video frame, and all video frames between the first video frame and the second video frame as a plurality of video frames corresponding to the time within the range from the start time to the end time; obtaining the collection timestamp and the decoding timestamp of a plurality of video frames from the video file, and determining the plurality of video frames and the collection timestamp and the decoding timestamp of each video frame as the plurality of video frame information.

4. The audio and video stream processing method for cross-platform live connection according to claim 2, characterized in that: The step of obtaining a video in a preset format according to the plurality of audio frame information and the plurality of video frame information comprises: inserting the collection timestamp and the decoding timestamp of each audio frame into the supplemental enhancement information of the corresponding audio frame to obtain a plurality of target audio frames; inserting the collection timestamp and the decoding timestamp of each video frame into the supplemental enhancement information of the corresponding video frame to obtain a plurality of target video frames; combining the plurality of target audio frames and the plurality of target video frames into a video in a preset format.

5. The audio and video stream processing method for cross-platform live connection according to claim 2, characterized in that: after the step of calculating the first time difference between the time indicated by the decoding timestamp of adjacent audio frames and the second time difference between the time indicated by the decoding timestamp of adjacent video frames, the method further comprises: if the first time difference is less than or equal to a preset threshold value and the second time difference is less than or equal to the preset threshold value, it is determined that the cause of the real-time live connection audio and video stream playing abnormal event is that the time of the audio and video stream from the second live platform server to the second anchor client is not synchronized.

6. The audio and video stream processing method for cross-platform live connection according to claim 1, characterized in that: The step of obtaining the real-time live connection audio and video stream playing abnormal event information of the second anchor client in the live connection process comprises: According to the sampling log data in the second anchor client, a time point in a preset time period before a playing abnormal time point is selected as a start time point of a real-time mic-in audio and video stream playing abnormal event, and a time point in a preset time period after the playing abnormal time point is selected as an end time point of the real-time mic-in audio and video stream playing abnormal event; wherein the second anchor client performs timed or periodic sampling on the real-time mic-in audio and video stream to generate the sampling log data when playing the real-time mic-in audio and video stream, and the sampling log data includes a collection time stamp of each audio frame or video frame in the real-time mic-in audio and video stream.

7. The audio and video stream processing method for cross-platform live streaming mic-in according to claim 1, characterized in that: if the playing abnormality of the mic-in live streaming playback video is audio and video playing lag, it is determined that the reason for the playing abnormality is audio and video flow interruption, audio and video frame loss, unordered collection time stamps of audio and video, or inconsistent time intervals of audio frames and video frames reaching the cache space of the second anchor client; if the playing abnormality of the mic-in live streaming playback video is audio and video frame loss, it is determined that the reason for the playing abnormality is real-time network fluctuation in the mic-in process of the second anchor client, audio and video frame loss when the audio and video frames reach the second live streaming platform server, or audio and video frame loss when the audio and video frames are sent from the second live streaming platform server to the second anchor client.

8. An audio and video stream processing device for cross-platform live co-mic, characterized in that, including: an event information acquisition module configured to acquire real-time mic-in audio and video stream playing abnormal event information of a second anchor client in a mic-in live streaming process; wherein the real-time mic-in audio and video stream playing abnormal event information includes a start time point and an end time point of a real-time mic-in audio and video stream playing abnormal event; wherein the real-time mic-in audio and video stream is an audio and video stream received by the second anchor client from a second live streaming platform server in real time; a playback video acquisition module configured to acquire a mic-in live streaming playback video of the second anchor client; wherein the mic-in live streaming playback video is a video obtained by the second anchor client after receiving and storing the audio and video stream; a frame information determination module configured to, if the mic-in live streaming playback video plays normally, determine a plurality of audio frame information and a plurality of video frame information corresponding to the start time point and the end time point according to an audio and video file acquired from the second live streaming platform server; wherein the audio and video file is a file obtained by the second live streaming platform server after receiving and decoding the audio and video stream pushed by a first live streaming platform server; The video analysis module is configured to obtain a video in a preset format according to the audio frame information and the video frame information, analyze the video in the preset format to obtain a decoding timestamp of each video frame and a decoding timestamp of each audio frame if the video in the preset format is synchronized with a video picture of the live streaming playback video, calculate a first time difference between time instants indicated by decoding timestamps of adjacent audio frames, and calculate a second time difference between time instants indicated by decoding timestamps of adjacent video frames; and determine that a cause of the real-time live streaming audio / video stream playing abnormal event is that a time of transmitting the audio / video stream from the first live streaming platform server to the second live streaming platform server is not synchronized if at least one of the first time differences is greater than a preset threshold and / or at least one of the second time differences is greater than the preset threshold.

9. An electronic device comprising: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Live-streaming microphone-connecting delay control method and device, electronic equipment and storage medium

    CN109413469A

  • Microphone connection live broadcast method, device and system

    CN111836074A