A method for processing multi-view video and related apparatus

By using audio alignment algorithms and timestamp modification technology, the problems of low alignment accuracy and high cost in multi-view video live streaming have been solved, achieving efficient synchronization across different devices.

CN119676484BActive Publication Date: 2025-12-16E SURFING VISION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918685.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-12-16
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing technologies for multi-view video live streaming have low alignment accuracy and high cost, mainly because differences in the performance and hardware configuration of video acquisition devices make time synchronization difficult to achieve, and require the replacement of high-cost equipment.

Method used

The audio alignment algorithm processes multiple audio data streams, extracts and matches audio features, aligns video data based on the audio alignment results, and modifies the display timestamps of the video data to achieve synchronization of multi-view videos.

Benefits of technology

It eliminates the need for time synchronization with video capture equipment, improving the alignment accuracy of multi-view videos and reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119676484B_ABST
    Figure CN119676484B_ABST
Patent Text Reader

Abstract

The application provides a multi-view video processing method and related device. The method realizes audio alignment, and then marks each video with display timestamp information, and encodes multiple videos with different resolutions, and transmits the videos through a multicast protocol in the same or different IP and different ports, and a live terminal obtains corresponding live video streams according to its own playing capacity, and realizes multi-view video decoding and synchronous playing. The application realizes the audio synchronization technology based on feature recognition, multicast transmission, terminal capacity matching, and solves the audio and picture synchronization problem of multi-view video, and meets the demand of multi-view video playing of different capacity terminals. The application is suitable for multicast protocol to realize multi-view video live broadcast, and provides a solution for the multicast service scene in the field of visual networking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vision Internet, and particularly relates to a multi-view video processing method and related device. BACKGROUND

[0002] With the development and popularization of vision Internet, in scenes such as live broadcast of tourist attractions, live broadcast of network games, live broadcast of e-commerce platforms and the like, multi-view video live broadcast has become one of the popular live broadcast modes. Among them, multi-view video live broadcast refers to a live broadcast mode in which multiple video collection devices simultaneously collect audio and video from different angles and align the multi-view video collected according to the time dimension. The audience in the live broadcast room can watch the live broadcast content of different angles at the same time through the aligned multi-view video, thereby enriching the watching experience and enhancing the content visualization.

[0003] At present, the prior art usually aligns the multi-view video according to the collection time stamp of the live broadcast video, which requires strict time synchronization between the video collection devices and / or requires additional processing of the collected audio and video streams by the video collection devices.

[0004] However, considering that each video collection device may have differences in device performance, functional configuration, hardware configuration and the like, in some cases, the time synchronization between the video collection devices cannot meet the requirements and / or the audio and video streams cannot be processed additionally, thereby reducing the alignment accuracy of the multi-view video. Or, it is necessary to replace the video collection device with high cost for collection, which has the technical defect of high cost. SUMMARY

[0005] The present application aims to at least solve one of the above technical defects, in particular, the technical defects of low alignment accuracy and high cost of multi-view video in the prior art.

[0006] In a first aspect, an embodiment of the present application provides a multi-view video processing method, comprising:

[0007] Obtaining multi-view video; wherein the multi-view video includes N-way live audio and video data collected at different angles, each way of the live audio and video data includes audio data and video data, N is a positive integer and N≥2;

[0008] Processing the multi-way audio data by using an audio alignment algorithm, and obtaining an audio alignment result;

[0009] Aligning the multi-way video data based on the audio alignment result, and obtaining a video alignment result;

[0010] According to the video alignment result, display timestamps corresponding to each of the video data are modified respectively, so that video frames at the same position correspond to the same display timestamp.

[0011] In some embodiments, the audio alignment algorithm comprises an audio feature extraction algorithm and an audio feature matching algorithm.

[0012] The audio alignment algorithm is used to process the multiple pieces of audio data, and an audio alignment result is obtained.

[0013] The audio feature extraction algorithm is used to extract audio features corresponding to each of the audio data respectively.

[0014] The audio feature matching algorithm is used to match the audio features corresponding to each of the audio data, and the audio alignment result is obtained based on the matching result.

[0015] The audio alignment result comprises N similar features and audio position information of each similar feature in corresponding audio data, and the N similar features belong to N pieces of audio data.

[0016] In some embodiments, the video alignment result comprises N positioning video frames and video position information of each positioning video frame in corresponding video data.

[0017] The video data is aligned based on the audio alignment result, and a video alignment result is obtained.

[0018] For each similar feature, the live audio-video data to which the similar feature belongs is taken as target audio-video data, the audio position information corresponding to the similar feature is taken as target position information, according to the audio-video correspondence relationship of the target audio-video data, a video frame corresponding to the target position information in the video data of the target audio-video data is taken as a positioning video frame, and the video position information of the positioning video frame is determined.

[0019] According to the video alignment result, display timestamps corresponding to each of the video data are modified respectively, so that video frames at the same position correspond to the same display timestamp.

[0020] According to the server timestamp and the video position information corresponding to the N positioning video frames, display timestamps corresponding to each of the video data are modified respectively.

[0021] In some embodiments, the method further comprises:

[0022] For each of the N pieces of aligned live audio-video data, the live audio-video data is encoded according to a plurality of preset encoding specifications, and a plurality of live encoded data is obtained; wherein the encoding specifications comprise resolution and frame rate.

[0023] According to the various live coding data, live video is pushed to the live terminal.

[0024] In some embodiments, the pushing of the live video to the live terminal according to the various live coding data comprises:

[0025] Based on the predetermined multicast address information, the various live coding data are multicast transmitted to push the live video to the live terminal.

[0026] In some embodiments, the pushing of the live video to the live terminal according to the various live coding data further comprises:

[0027] A live play request sent by the live terminal is received, wherein the live play request carries decoding capability information of the live terminal.

[0028] A target coding specification is determined according to the decoding capability information, and a request response information is generated based on target multicast address information corresponding to the target coding specification.

[0029] The request response information is returned to the live terminal.

[0030] In some embodiments, the obtaining of the multi-view video comprises:

[0031] According to a plurality of preset paths, live audio and video data collected by a plurality of video collection devices are respectively received, wherein the plurality of preset paths correspond one-to-one to the plurality of live audio and video data.

[0032] In a second aspect, the embodiments of the present application provide a multi-view video processing apparatus, comprising:

[0033] A video obtaining module is configured to obtain a multi-view video, wherein the multi-view video comprises N live audio and video data collected at different angles, each of the live audio and video data comprises audio data and video data, and N is a positive integer and N≥2.

[0034] An audio alignment result obtaining module is configured to process the plurality of audio data by using an audio alignment algorithm, and obtain an audio alignment result.

[0035] A video alignment module is configured to align the plurality of video data based on the audio alignment result, and obtain a video alignment result.

[0036] A timestamp modifying module is configured to modify display timestamps corresponding to the plurality of video data respectively according to the video alignment result, so that video frames at the same position correspond to the same display timestamp.

[0037] In a third aspect, an embodiment of the present application provides a storage medium, which stores computer readable instructions. When the computer readable instructions are executed by one or more processors, the one or more processors perform the steps of the multi-view video processing method according to any of the embodiments.

[0038] In a fourth aspect, an embodiment of the present application provides a live video server, which comprises one or more processors and a memory.

[0039] The memory stores computer readable instructions. When the computer readable instructions are executed by the one or more processors, the one or more processors perform the steps of the multi-view video processing method according to any of the embodiments.

[0040] In the multi-view video processing method and the related device provided by some embodiments of the present application, the live video server can separate the audio data from each of the live audio and video data of the multi-view video, and perform audio synchronization on the audio data by using an audio alignment algorithm, so as to realize audio alignment of the multi-view video. Based on the audio alignment result, the live video server can align the video data in the time dimension, and modify the display time stamp corresponding to each of the video data based on the video alignment result, so that the video frames at the same position correspond to the same display time stamp. In this way, the multi-view video alignment can be realized. As can be seen, the present application does not require the video acquisition device to perform additional processing on the acquired audio and video data, and does not rely on the time synchronization between the multiple video acquisition devices. Even if the video acquisition device does not meet the time synchronization requirement, the present application can accurately realize the alignment and synchronization of the multi-view video, thereby improving the alignment accuracy of the multi-view video and reducing the cost. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0042] Figure 1 For one embodiment, a system schematic diagram of a multi-view video processing system is shown.

[0043] Figure 2 For one embodiment, a flowchart of a multi-view video processing method is shown.

[0044] Figure 3 For one embodiment, a system schematic diagram of a multi-view video processing system is shown.

[0045] Figure 4 For an embodiment, a functional module diagram of a multi-view video processing system is shown in FIG. 1.

[0046] Figure 5 For an embodiment, a flowchart of a multi-view video processing method is shown in FIG. 2.

[0047] Figure 6 For an embodiment, a structural diagram of a multi-view video processing device is shown in FIG. 3.

[0048] Figure 7 For an embodiment, an internal structure diagram of a live video server is shown in FIG. 4. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0050] In some embodiments, the multi-view video processing method provided by the present application can be applied to a multi-view video processing system. It can be understood that the multi-view video processing system can be used to provide multi-view video live broadcast services in the field of video networking, and / or provide multi-view video live broadcast services in other video scenarios.

[0051] Among them, the live video (Live Video, abbreviated as LV) refers to a technology that captures live scenes through a video capture device, and pushes audio data and video data to the network in real time through network transmission, so that users can receive and watch videos through terminals such as browsers or application programs APP. Multi-view video can include N live audio and video data, and any two live audio and video data can correspond to different shooting angles, where N is a positive integer and N≥2.

[0052] It can be understood that the number of N can be determined according to actual conditions, which is not specifically limited herein, for example, it can be 2, 3, 4, 5, 8, etc. For ease of description, some embodiments of the present application take N=4 as an example for description, that is, the multi-view video is a 4-view video, which includes 4 live audio and video data.

[0053] For example, as shown in FIG. 5, a 4-view video includes four live audio and video data, which correspond to four different shooting angles, respectively. Figure 1As shown, the multi-view video processing system can include a video collection device 102, a live video server 104 and a live terminal 106. The video collection device 102 refers to a device for collecting live audio and video data, which can be, but is not limited to, a camera, a camera, etc. The live terminal 106 refers to a terminal device for playing multi-view video, which can be, but is not limited to, a smart phone, a wearable device, a household appliance, an Internet of Things device, a tablet computer, a notebook computer, a desktop computer, etc. The live video server 104 is used to align multiple live audio and video data in the multi-view video in the time dimension, so that the aligned multi-view video can display the audio and video content corresponding to different angles at the same time.

[0054] It can be understood that the number of video collection devices 102 can be multiple, and the number of live terminals 106 can be one or more. The specific number of video collection devices 102 and live terminals 106 can be determined according to actual conditions, which is not limited herein.

[0055] Specifically, each video collection device 102 can collect live audio and video data of a corresponding angle, and transmit the collected live audio and video data to the live video server 104 in real time. The live video server 104 can receive the live audio and video data uploaded by each video collection device 102 respectively, so as to obtain a multi-view video. The live video server 104 can also process the multi-view video according to the multi-view video processing method provided in the following embodiments to align each live audio and video data in the time dimension. In this way, the collection ability / processing ability of the video collection device 102 can not be changed, and the time synchronization of each video collection device 102 is not strictly required, and the audio alignment synchronization technology is introduced to complete the alignment and synchronization of the multi-view video in the live video server 104.

[0056] The live terminal 106 can report the number of simultaneous decoding supported by the terminal, the resolution / frame rate of each channel, and other capability information to the live video server 104. The live video server 104 can also push the aligned multi-view video to the live terminal 106 based on the received information, and the multi-view video pushed by the live video server 104 matches the terminal decoding capability of the live terminal 106. After receiving the live audio and video data issued by the live video server 104, the live terminal 106 can decode the live audio and video data, and perform video rendering according to the display timestamp to synchronously play the multi-view video.

[0057] The multi-view video processing method provided by the present application is described below.

[0058] In some embodiments, the present application provides a multi-view video processing method, and the following embodiments are applied to the method. Figure 1The following explanation uses a live video server as an example. Figure 2 As shown, the method may include the following steps:

[0059] S202: Acquire multi-view video; where multi-view video includes N channels of live audio and video data acquired from different perspectives, each channel of live audio and video data includes audio data and video data, N is a positive integer and N≥2;

[0060] S204: The audio alignment algorithm is used to process the multi-channel audio data and obtain the audio alignment result;

[0061] S206: Align multiple video data streams based on the audio alignment results and obtain the video alignment results;

[0062] S208: Based on the video alignment results, modify the display timestamps corresponding to each video data stream so that video frames at the same location correspond to the same display timestamps.

[0063] In this application, the live video server can acquire N channels of live audio and video data collected by multiple video capture devices. It is understood that in the multiple channels of live audio and video data, the shooting angles corresponding to every two channels are different, allowing the multiple channels of live audio and video data to present live content from different angles. Each channel of live audio and video data can include both audio and video data, and the audio and video data within the same channel have a temporal correspondence, enabling audio-visual synchronization within a single channel of live audio and video data.

[0064] Given N channels of live audio and video data, the live video server can extract audio from each channel to separate N channels of audio data. For example, when N=2, the live video server can extract audio from the first and second channels of live audio and video data, obtaining the audio data corresponding to the first and second channels respectively, for a total of two audio data channels.

[0065] Given N audio streams, a live video server can process them using an audio alignment algorithm to obtain the alignment result. The audio alignment algorithm refers to the algorithm used to determine the alignment relationship of multiple audio streams on the time axis, and may include time series analysis algorithms, dynamic time warping algorithms, etc. The audio alignment result describes the alignment relationship of the N audio streams on the time axis, for example, it can be the time difference between each target audio stream and a reference audio stream, where the reference audio stream is one of the N audio streams, and the target audio streams are the N-1 audio streams excluding the reference audio stream.

[0066] It can be understood that the audio extraction can be implemented in any manner, and the specific algorithm type and specific implementation manner of the audio extraction algorithm can be determined according to actual conditions, and the present application does not make a specific limitation. Similarly, any audio alignment algorithm can be used to obtain the audio alignment result, and the present application does not make a specific limitation.

[0067] Since the audio and video data of each live audio and video data has a time corresponding relationship, and the alignment relationship of the N audio data in the time dimension can be reflected by the audio alignment result, after determining the audio alignment result of the N audio data, the N video data in the time dimension can be aligned according to the audio alignment result, and a video alignment result is obtained. The video alignment result is used to describe the alignment relationship of the N video data on the time axis.

[0068] After determining the alignment relationship of the N video data, the present application can modify the corresponding PTS (Presentation Time Stamp, display timestamp) of the N video data according to the video alignment result, so that the video frames at the same time position correspond to the same display timestamp.

[0069] For example, when N=2, if the first video frame (hereinafter referred to as video frame 1) of the first video data is time-aligned with the second video frame (hereinafter referred to as video frame 2) of the second video data, that is, the video frame 1 and the video frame 2 correspond to the same time position, the PTS of the video frame 1 and the PTS of the video frame 2 can be modified to the same timestamp, so that the video frame 1 and the video frame 2 correspond to the same display timestamp. Other video frames of the first video data and other video frames of the second video data can be modified according to the foregoing manner to modify the corresponding PTS of the video frames, so that the video frames at the same time position correspond to the same display timestamp.

[0070] As can be seen, the present application does not require the video capture device to perform additional processing on the captured audio and video data, and does not rely on the time synchronization between multiple video capture devices. Even in the case where the video capture device does not meet the time synchronization requirement, the present application can accurately realize the alignment and synchronization of multi-view video, thereby improving the alignment accuracy of multi-view video and reducing the cost.

[0071] In some embodiments, the audio alignment algorithm includes an audio feature extraction algorithm and an audio feature matching algorithm. The audio feature extraction algorithm refers to an algorithm for extracting audio features from audio data. The audio feature matching algorithm refers to an algorithm for performing feature matching on the audio features corresponding to the multi-channel audio data. Further, the audio feature described in the present application can be a result of digitally representing a certain attribute of the audio data, such as time domain features, frequency domain features, mel spectrum, etc.

[0072] The audio alignment algorithm is used to process multiple audio data, and the audio alignment result is obtained, including:

[0073] Step A1: An audio feature extraction algorithm is used to extract the audio features corresponding to each audio data respectively.

[0074] Step A3: An audio feature matching algorithm is used to match the audio features corresponding to each audio data, and the audio alignment result is obtained based on the matching result.

[0075] The audio alignment result includes N similar features and audio position information of each similar feature in the corresponding audio data, and the N similar features belong to N audio data.

[0076] In this embodiment, the live video server can use an audio feature extraction algorithm to obtain the audio features corresponding to each audio data respectively, and use an audio feature matching algorithm to match the audio features corresponding to N audio data to identify similar audio features possessed by N audio data, and then obtain the audio alignment result. In this way, the audio alignment result can be obtained through the audio synchronization technology based on feature recognition, thereby further improving the accuracy of alignment.

[0077] Specifically, the live video server can use an audio feature extraction algorithm to perform feature recognition processing on N audio data respectively to obtain the audio features corresponding to each audio data. When the audio features corresponding to N audio data are obtained, the live video server can use an audio feature matching algorithm to match the N audio features, and obtain N similar features and audio position information of each similar feature in the corresponding audio data according to the feature matching result.

[0078] The N similar features are audio features in N audio data that reflect the same audio content. In some examples, the N similar features can be understood as common audio features of N audio data. Further, the live video server can align according to a preset feature matching threshold, so that the feature similarity between the N similar features satisfies the preset feature matching condition.

[0079] The audio position information of each similar feature in the corresponding audio data refers to the position information of the similar feature in the audio data to which it belongs. For example, when the first similar feature is the audio feature of the first audio data and the second similar feature is the audio feature of the second audio data, the audio position information corresponding to the first similar feature refers to the position information of the first similar feature in the first audio data, and the audio position information corresponding to the second similar feature refers to the position information of the second similar feature in the second audio data.

[0080] It should be noted that the position information corresponding to the similar feature can be described by different forms of parameters, and the present application does not make a specific limitation thereon, as long as the position information can describe the time position of the similar feature in the corresponding audio data. For example, the present application can take the audio frame sequence number corresponding to the similar feature as the position information, or take the time stamp corresponding to the similar feature in the corresponding audio data as the position information.

[0081] In some embodiments, the video alignment result includes N positioning video frames and video position information of each positioning video frame in the corresponding video data. Wherein, the N positioning video frames belong to the N video data respectively.

[0082] Based on the audio alignment result, the multiple video data are aligned, and a video alignment result is obtained, including:

[0083] For each similar feature, the live audio and video data to which the similar feature belongs is taken as target audio and video data, the audio position information corresponding to the similar feature is taken as target position information, according to the audio and video correspondence relationship of the target audio and video data, the video frame corresponding to the target position information in the video data of the target audio and video data is taken as a positioning video frame, and the video position information of the positioning video frame is determined;

[0084] According to the video alignment result, the display time stamp corresponding to each video data is modified respectively, including:

[0085] According to the server time stamp and the video position information corresponding to the N positioning video frames, the display time stamp corresponding to each video data is modified respectively.

[0086] In the present embodiment, since the N similar features of the audio alignment result are audio features in the N audio data for reflecting the same audio content, therefore, according to the corresponding relationship between the same audio and video data, according to the audio position information corresponding to the N similar features, the positioning video frame corresponding to the same audio content in each video data is determined respectively, and N positioning video frames and the video position information corresponding to each positioning video frame are obtained.

[0087] The following describes the determination of a positioning video frame in one channel of video data. It is understood that the positioning video frame of N channels of video data can be determined in the following manner. Assume that one of the N similar features (hereinafter referred to as the first similar feature) belongs to the first channel of audio data, and the first channel of audio data and the first channel of video data are extracted from the first channel of live audio-video data. In this case, the present application can determine the video frame that satisfies the aforementioned audio-video correspondence relationship in the first channel of video data as the positioning video frame according to the audio position information of the first similar feature in the first channel of audio data and the audio-video correspondence relationship between the first channel of audio data and the first channel of video data. It is understood that the positioning video frame and the first similar feature satisfy the audio-video correspondence relationship between the first channel of live audio-video data.

[0088] The live video server can determine N positioning video frames and the video position information of each positioning video frame in the corresponding video data in the N channels of video data in the above manner. The N positioning video frames belong to the N channels of video data, and the N positioning video frames correspond to the same audio content. Thereafter, the live video server can modify the display timestamps corresponding to each channel of video data according to the server timestamp and the video position information corresponding to the N positioning video frames, so that the N channels of video data can be aligned in the time dimension.

[0089] For example, when N = 3, the first similar feature corresponds to the first position, and the first similar feature belongs to the first channel of audio data. Similarly, the second similar feature corresponds to the second position, and the second similar feature belongs to the second channel of audio data. The third similar feature corresponds to the third position, and the third similar feature belongs to the third channel of audio data. In this case, the PTS of the first channel of video data at the first position, the PTS of the second channel of video data at the second position, and the PTS of the third channel of video data at the third position are the same, for example, all can be the current timestamp of the live video server. After determining the PTS of the first channel of video data at the first position, the live video server can modify the PTS of the first channel of video data at other positions based on this. Similarly, the live video server can modify the PTS of the second channel of video data at other positions based on the PTS of the second channel of video data at the second position. The live video server can modify the PTS of the third channel of video data at other positions based on the PTS of the third channel of video data at the third position.

[0090] In some embodiments, after aligning each channel of live audio-video data based on the audio alignment result, the method further comprises:

[0091] Step C1: encoding each of the N pieces of aligned live audio and video data according to a preset plurality of encoding specifications, and obtaining a plurality of live encoding data; wherein the encoding specifications include resolution and frame rate;

[0092] Step C3: pushing the live video to the live terminal according to the various live encoding data.

[0093] In this embodiment, considering that live terminals are various and different live terminals have different decoding capabilities, the live video server can provide multi-view video of a plurality of encoding specifications to meet the playing requirements of live terminals with different decoding capabilities for multi-view video, and thus can provide low-latency multi-view video to the live terminal and improve the playing fluency of multi-view video.

[0094] In this embodiment, after aligning the N pieces of live audio and video data, the live video server can perform video encoding according to a preset M encoding rules, and obtain a plurality of live encoding data. M is a positive integer greater than 1.

[0095] It can be understood that each piece of live audio and video data corresponds to M pieces of live encoding data, and the M pieces of live encoding data correspond to the M encoding rules one by one. After encoding processing of the N pieces of live audio and video data, NXM pieces of live encoding data can be obtained. For example, if N=2, M=3, and the first encoding rule is: resolution 720P / frame rate 30, the second encoding specification is resolution 1080P / frame rate 50, and the third encoding rule is: resolution 2160P / frame rate 60, the first piece of live audio and video data can be encoded according to the first encoding rule, the second encoding rule and the third encoding rule, and three pieces of live encoding data are obtained. Similarly, the second piece of live audio and video data can be encoded according to the first encoding rule, the second encoding rule and the third encoding rule, and three pieces of live encoding data are obtained.

[0096] Further, in one example, in the encoding process, the live video server can write the timestamp of the live video server into the PES packet header PTS corresponding to the audio and video frame, and the live terminal can realize synchronous playing of multi-view pictures by analyzing the PTS. Wherein, PES refers to the grouping basic stream obtained by packing and grouping the basic audio and video stream according to the specification of PS / TS, and a single PES includes a PES packet header and basic stream data.

[0097] The live video server can push live video to the live terminal according to the capability of the live terminal, so that the live terminal can receive N live audio and video data of different perspectives. It should be noted that for N live audio and video data received by the same live terminal, the encoding specifications corresponding to each two live audio and video data can be the same or different, and the present application does not make a specific limitation.

[0098] In some embodiments, the live video is pushed to the live terminal according to various live encoding data, including:

[0099] Based on the predetermined multicast address information, the various live encoding data are transmitted by multicast to push the live video to the live terminal.

[0100] Specifically, the prior art usually implements the push of multi-perspective video in the form of on-demand, and some schemes use video slice transmission protocols such as HLS (HTTP Live Streaming, a kind of adaptive bit rate streaming media transmission protocol based on HTTP) to complete, resulting in a delay of more than 10 seconds in live broadcast. To solve this problem, the present embodiment uses the multicast mode to live broadcast multi-perspective video, so that the transmission bandwidth can be saved on the basis of meeting the decoding capability diversity of the live terminal, and then the live delay can be reduced to realize low-delay live broadcast.

[0101] The live video server can transmit video encoding data of different perspectives and different encoding rules, video encoding data of different perspectives and same encoding rules, and video encoding data of same perspective and different encoding rules to the network through multicast protocols and different multicast address information to realize live video push. The multicast address information can include IP (Internet Protocol) address and port. In one example, the multicast address can be numbered with fixed port intervals.

[0102] In some embodiments, the live video is pushed to the live terminal according to various live encoding data, further including:

[0103] Step D1: receiving a live play request sent by a live terminal; wherein the live play request carries decoding capability information of the live terminal;

[0104] Step D3: determining a target encoding specification according to the decoding capability information, and generating request response information based on target multicast address information corresponding to the target encoding specification;

[0105] Step D5: returning the request response information to the live terminal.

[0106] In this embodiment, considering that the live terminal does not know the correlation between videos when joining the multicast, the live terminal itself cannot determine the multicast to be joined for simultaneously watching multiple videos of different perspectives. Based on this, the live video server can match the playing capability of the live terminal and provide the live terminal with the multicast playing function of the multi-perspective video, while meeting the service requirements of a large number of user accesses, low-latency live broadcast, and interactive real-time performance.

[0107] Specifically, the live terminal can report the decoding capability information of the live terminal to the live video server through a live playing request. The live video server can extract the decoding capability information corresponding to the live terminal from the live playing request, determine a target coding specification matched with the capability of the live terminal according to the decoding capability information, and determine target multicast address information according to the target coding specification, so that the live terminal can receive the multi-perspective video corresponding to the target coding specification based on the target multicast address information.

[0108] Further, after obtaining the multi-perspective video corresponding to the target coding specification, the live terminal can perform decoding processing on the received multi-perspective video, and play according to the PTS of each video, so as to display the time-aligned multi-perspective video.

[0109] In this way, the live video server does not need to perform mixstreaming processing on the multi-channel collected audio and video, the live terminal can obtain the live video stream matched with the playing capability of the live terminal through the live playing request, and can realize the synchronous playing of the multi-perspective video by analyzing the PTS information.

[0110] In some embodiments, the multi-perspective video is obtained, including:

[0111] According to a plurality of preset paths, live audio and video data of each channel is read respectively; wherein the plurality of preset paths correspond to the multi-channel audio and video data one by one.

[0112] In this embodiment, different video collection devices can transmit the live audio and video data collected by the device to the live video server in real time according to the fixed path negotiated with the live video server. In this way, the live video server can read the multi-channel live audio and video streams collected in real time according to the negotiated fixed path, and realize the acquisition of the original multi-perspective video.

[0113] In order to facilitate understanding of the scheme of the present application, a specific example is used for illustration.

[0114] In one example, as shown in Figure 3 , Figure 4 and Figure 5 , the multi-perspective video processing method of the present example can include the following steps:

[0115] S501: Four video acquisition devices respectively acquire four live audio and video data, and transmit the acquired live audio and video data to a live video server in real time according to an agreed path; further, the live audio and video data can be transmitted by using an RTP (Real-time Transport Protocol).

[0116] S502: The live video server receives and stores four live audio and video stream data, and separates four audio data from the four live audio and video stream data.

[0117] S503: The live video server analyzes the four audio data, identifies the same / similar audio features from the four audio data, and writes a server timestamp into a PTS of a PES packet header. In this step, after the same / similar audio features are identified, the live video server takes the same / similar audio features as a starting time point, marks the PES packet header information at the aligned audio position and the video frame synchronized with the audio, and writes the server timestamp into the PTS. After that, the PTS of each packet header is filled with the server timestamp.

[0118] S504: The live video server encodes and transcodes according to different encoding rules to form video streams of different resolutions. In this step, three encoding specifications can be used for encoding, for example, specification 1: resolution 720P / frame rate 30, specification 2: resolution 1080P / frame rate 50, and specification 3: resolution 2160P / frame rate 60.

[0119] S505: The live video server transmits the multi-path and multi-specification audio and video streams to the network. For example, the multicast IP address of the first live video stream data is 239.45.3.41, and the video streams of three encoding specifications are transmitted to ports 5100 / 5102 / 5104; the multicast IP address of the second video stream data is 239.45.3.42, and the streams of three specifications are transmitted to ports 5100 / 5102 / 5104; and so on.

[0120] S506: The live terminal sends a live play request carrying capability information. For example, the live play request sent by using HTTP (Hypertext Transfer Protocol) can deliver parameters “name”: “frame1”, “capability”: “2160P / 60fps”, “name”: “frame2”, “capability”: “720P / 30fps”.

[0121] S507: The live video server matches the decoding capability of the terminal and returns the target multicast address information to the live terminal. For example, after receiving the live play request described above, the HTTP response returned by the live video server to the live terminal can pass the parameters "view_num": 4, "view_angle": "0 / 90 / 180 / 270", "view_IP": "239.45.3.41 / 239.45.3.42 / 239.45.3.43 / 239.45.3.44", "view_Port": "frame1-5104 / frame2-5100".

[0122] S508: The live terminal selects a view angle, joins the corresponding multicast address, and parses the PTS to synchronously present the multi-view picture. For example, the live terminal can join the multicast according to the multicast addresses igmp: / / 239.45.3.41 / 5104 and igmp: / / 239.45.3.42 / 5100 to obtain a 4K video stream at 0-degree view angle and a 720P video stream at 90-degree view angle.

[0123] The present application provides a method and system for providing multicast play for multi-view video. The method realizes audio alignment and marking of PTS timestamp information on each video stream, and multi-video encoding of different resolutions, and transmission through multicast protocol in the same or different IP and different port modes, and terminal access to corresponding live video stream according to its own play capability, and multi-view video decoding and synchronous play. The present application realizes audio synchronization technology based on feature recognition, multicast transmission, terminal capability matching, and solves the audio and picture synchronization problem of multi-view video, and meets the demand of multi-view video play of different capability terminals of different types. The present application is suitable for live broadcast of multi-view video through multicast protocol, and provides a solution for multicast service scene in the field of vision internet.

[0124] The multi-view video processing device provided by the embodiments of the present application is described below, and the multi-view video processing device described below can be correspondingly referred to the multi-view video processing method described above.

[0125] In some embodiments, as shown in Figure 6 The multi-view video processing device 600 provided by the embodiments of the present application includes:

[0126] The video acquisition module 602 is configured to acquire multi-view video; wherein the multi-view video includes N live audio and video data collected at different view angles, each live audio and video data includes audio data and video data, and N is a positive integer and N≥2.

[0127] The audio alignment result acquisition module 604 is configured to process the multi-channel audio data by using an audio alignment algorithm, and obtain an audio alignment result.

[0128] The video alignment module 606 is configured to align the multi-channel video data based on the audio alignment result, and obtain a video alignment result.

[0129] The timestamp modification module 608 is configured to modify the display timestamps corresponding to the video data of each channel respectively according to the video alignment result, so that the video frames at the same position correspond to the same display timestamp.

[0130] In some embodiments, the audio alignment algorithm includes an audio feature extraction algorithm and an audio feature matching algorithm. The audio alignment result acquisition module 604 of the present application includes:

[0131] The feature extraction unit is configured to extract the audio features corresponding to each of the audio data respectively by using the audio feature extraction algorithm.

[0132] The feature matching unit is configured to match the audio features corresponding to each of the audio data by using the audio feature matching algorithm, and obtain the audio alignment result based on the matching result. The audio alignment result includes N similar features and the audio position information of each similar feature in the corresponding audio data, and the N similar features belong to N audio data.

[0133] In some embodiments, the video alignment result includes N positioning video frames and the video position information of each positioning video frame in the corresponding video data. The video alignment module 606 of the present application includes:

[0134] The positioning video frame determination unit is configured to, for each similar feature, take the live audio-video data to which the similar feature belongs as target audio-video data, take the audio position information corresponding to the similar feature as target position information, and according to the audio-video correspondence relationship of the target audio-video data, take the video frame corresponding to the target position information in the video data of the target audio-video data as a positioning video frame, and determine the video position information of the positioning video frame.

[0135] The timestamp modification module 608 of the present application includes:

[0136] The PTS modification unit is configured to modify the display timestamps corresponding to the video data of each channel respectively according to the server timestamp and the video position information corresponding to the N positioning video frames.

[0137] In some embodiments, the multi-view video processing apparatus 600 of the present application further includes:

[0138] The encoding module is configured to encode each of the aligned N live audio and video data according to a preset plurality of encoding specifications, and obtain a plurality of live encoding data; wherein the encoding specifications include resolution and frame rate.

[0139] The pushing module is configured to push the live video to the live terminal according to the live encoding data.

[0140] In some embodiments, the pushing module of the present application comprises:

[0141] The multicasting unit is configured to multicast the live encoding data based on predetermined multicast address information, and push the live video to the live terminal.

[0142] In some embodiments, the pushing module of the present application further comprises:

[0143] The request receiving unit is configured to receive a live play request sent by the live terminal; wherein the live play request carries decoding capability information of the live terminal.

[0144] The target address determining unit is configured to determine a target encoding specification according to the decoding capability information, and generate request response information based on target multicast address information corresponding to the target encoding specification.

[0145] The response returning unit is configured to return the request response information to the live terminal.

[0146] In some embodiments, the video obtaining module 602 of the present application comprises:

[0147] The data reading unit is configured to read each of the live audio and video data according to a plurality of preset paths; wherein the plurality of preset paths correspond to the plurality of live audio and video data one by one.

[0148] In one embodiment, the present application further provides a storage medium, which stores computer readable instructions. When the computer readable instructions are executed by one or more processors, the one or more processors perform the steps of the multi-view video processing method in any embodiment.

[0149] In one embodiment, the present application further provides a live video server, which stores computer readable instructions. When the computer readable instructions are executed by one or more processors, the one or more processors perform the steps of the multi-view video processing method in any embodiment.

[0150] Schematically, Figure 7An internal structure diagram of a live video server is provided in the embodiments of the present application. In one example, the live video server can be a server. Referring to Figure 7 The live video server 900 includes a processing component 902, which further includes one or more processors, and a memory resource represented by a memory 901 for storing instructions, such as application programs, executable by the processing component 902. The application programs stored in the memory 901 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 902 is configured to execute the instructions to perform the steps of the multi-view video processing method described in any of the embodiments above.

[0151] The live video server 900 can further include a power supply component 903 configured to perform power management of the live video server 900, a wired or wireless network interface 904 configured to connect the live video server 900 to a network, and an input / output (I / O) interface 905. The live video server 900 can operate based on an operating system stored in the memory 901, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM, or the like.

[0152] Those skilled in the art can understand that the internal structure of the live video server shown in the present application is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the live video server to which the scheme of the present application is applied. The specific live video server can include more or less components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0153] Finally, it should be noted that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element. In this document, "a", "one", "said", "the" and "it" can also include plural forms, unless the context clearly indicates otherwise. A plurality means at least two, such as 2, 3, 5 or 8, etc. "And / or" includes any and all combinations of the related listed items.

[0154] The various embodiments described in this specification are presented for the purpose of illustration and description. Each of the embodiments described in this specification can be implemented alone or in combination with others. The same, similar or identical components shown in different embodiments are identified by the same reference numerals and will not be described repeatedly.

[0155] The above description of disclosed embodiments is intended to be illustrative and not restrictive. Many modifications and variations to the described embodiments will be apparent to those skilled in the art from this disclosure. It is therefore contemplated to embrace within the scope of claims and their equivalents whatever falls within the scope of the disclosure to achieve the advantages of the application.

Claims

1. A method for processing multi-view video, characterized in that, include: Acquire multi-view video; wherein, the multi-view video includes N channels of live audio and video data acquired from different perspectives, each channel of live audio and video data includes audio data and video data, N is a positive integer and N≥2; An audio alignment algorithm is used to process the multiple audio data streams and obtain the audio alignment result. Based on the audio alignment result, the multiple video data streams are aligned to obtain the video alignment result; Based on the video alignment results, the display timestamps corresponding to each video data stream are modified so that video frames at the same position correspond to the same display timestamp. For each of the N aligned live audio and video data streams, the live audio and video data stream is encoded according to multiple preset encoding specifications to obtain multiple live encoded data streams; wherein, the encoding specifications include resolution and frame rate; Based on the pre-determined multicast address information, various live broadcast encoded data are multicast transmitted to push live video to the live broadcast terminal. Receive a live streaming playback request sent by the live streaming terminal; wherein the live streaming playback request carries the decoding capability information of the live streaming terminal; The target encoding specification is determined based on the decoding capability information, and a request response information is generated based on the target multicast address information corresponding to the target encoding specification. The request response information is returned to the live streaming terminal.

2. The method according to claim 1, characterized in that, The audio alignment algorithm includes an audio feature extraction algorithm and an audio feature matching algorithm; The process of using an audio alignment algorithm to process multiple audio data streams and obtain audio alignment results includes: The audio feature extraction algorithm is used to extract the audio features corresponding to each audio data stream. The audio feature matching algorithm is used to match the audio features corresponding to each audio data stream, and the audio alignment result is obtained based on the matching result. The audio alignment result includes N similar features and the audio position information of each similar feature in the corresponding audio data. The N similar features belong to N audio data streams.

3. The method according to claim 2, characterized in that, The video alignment result includes N positioning video frames and the video position information of each positioning video frame in the corresponding video data; The step of aligning multiple video data streams based on the audio alignment result to obtain a video alignment result includes: For each of the similar features, the live audio and video data to which the similar feature belongs is taken as the target audio and video data, and the audio location information corresponding to the similar feature is taken as the target location information. According to the audio and video correspondence of the target audio and video data, the video frame in the video data of the target audio and video data that corresponds to the target location information is taken as a positioning video frame, and the video location information of the positioning video frame is determined. The step of modifying the display timestamps corresponding to each video data stream according to the video alignment result includes: Based on the server timestamp and the video position information corresponding to the N positioning video frames, the display timestamps corresponding to each of the video data streams are modified respectively.

4. The method according to any one of claims 1 to 3, characterized in that, The acquisition of multi-view video includes: According to multiple preset paths, live audio and video data collected by multiple video acquisition devices are received respectively; wherein, the multiple preset paths correspond one-to-one with the multiple live audio and video data streams.

5. A multi-view video processing apparatus, characterized in that, include: The video acquisition module is used to acquire multi-view videos; wherein, the multi-view videos include N channels of live audio and video data acquired from different perspectives, each channel of live audio and video data includes audio data and video data, where N is a positive integer and N≥2; The audio alignment result acquisition module is used to process the multiple audio data using an audio alignment algorithm and obtain the audio alignment result; The video alignment module is used to align multiple video data streams based on the audio alignment result and obtain a video alignment result. The timestamp modification module is used to modify the display timestamps corresponding to each video data stream according to the video alignment results, so that video frames at the same position correspond to the same display timestamp. The encoding module is used to encode each live audio and video data in N aligned live audio and video data according to a variety of preset encoding specifications, and obtain a variety of live encoded data; wherein, the encoding specifications include resolution and frame rate; The push module is used to multicast various live broadcast encoded data based on pre-determined multicast address information in order to push live video to the live broadcast terminal. The push module is further configured to receive a live playback request sent by the live streaming terminal; wherein the live playback request carries decoding capability information of the live streaming terminal; determine the target encoding specification based on the decoding capability information, and generate request response information based on the target multicast address information corresponding to the target encoding specification; and return the request response information to the live streaming terminal.

6. A storage medium, characterized in that, The storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the multi-view video processing method as described in any one of claims 1 to 4.

7. A live video server, characterized in that, include: One or more processors, and memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, perform the steps of the multi-view video processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-video synchronization method and device

    CN112995708A

  • Directing method, device and equipment and computer storage medium

    CN114339302A