Multi-channel live video processing method and system based on dynamic resource allocation
The multi-channel live video processing method with dynamic resource allocation solves the problems of cumbersome operation and strong hardware dependence of multi-channel live video processing, and realizes efficient and portable multi-channel live video processing.
Patent Information
- Application Number
- CN202510738921.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
Existing multi-channel video live broadcast solutions are cumbersome to operate, highly dependent on hardware, and cannot be adapted to mobile shooting scenarios.
A multi-channel live video processing method based on dynamic resource allocation is adopted to realize intelligent directing through the client or multi-channel input director console, dynamically allocate acquisition, processing and encoding resource pools, reduce manual operations, and reduce equipment size and power consumption.
It improves operational efficiency, reduces the probability of errors, adapts to mobile shooting scenes, and realizes intelligent processing of multi-channel input data.
Smart Images

Figure CN120640029A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video directing technology, and in particular to a multi-channel video live broadcast processing method and system based on dynamic resource allocation. Background Art
[0002] As live streaming becomes more professional and mobile, multi-channel video live streaming has become a core requirement for live streaming scenarios such as large-scale events, sports events, and outdoor live streaming. In existing technologies, multi-channel video live streaming mainly relies on computer-side director software. Typical solutions include manually switching the screen layout of multiple input sources or performing simple splicing through the software interface. However, such traditional solutions have the following systemic defects: (1) Cumbersome operation: Users need to manually adjust the input priority and screen layout of multiple video channels, especially in complex splicing modes (such as multi-screen split screen), which is inefficient and prone to errors; (2) Strong hardware dependence and limited mobility: Existing solutions rely on high-performance computers to run director software, resulting in large equipment size, high power consumption, and inability to adapt to mobile shooting scenarios. Summary of the Invention
[0003] Based on this, it is necessary to address the technical problems of the existing multi-channel video live broadcast solution, which has cumbersome operation, strong hardware dependence and limited mobility. A multi-channel video live broadcast processing method based on dynamic resource allocation is proposed.
[0004] In a first aspect, a method for processing multi-channel video live broadcasts based on dynamic resource allocation is provided, the method comprising: the method is used to control a multi-channel input director station, the method comprising: Get live broadcast signal; In response to the live broadcast signal, obtaining a picture splicing mode, wherein the picture splicing mode is data input by a user through a client or the multi-channel input director station; Based on the dynamic resource allocation strategy, collection pool, processing pool and encoding pool for the multi-input director station, data collection, data splicing and encoding output are performed according to the picture splicing mode.
[0005] In the second aspect, a multi-channel video live broadcast processing system based on dynamic resource allocation is provided, the system comprising: a multi-channel input director station and a client, the client being communicatively connected to the multi-channel input director station, and the multi-channel input director station being configured to implement the multi-channel video live broadcast processing method based on dynamic resource allocation described in the first aspect.
[0006] The multi-channel video live broadcast processing method and system based on dynamic resource allocation of the present application, the user inputs the picture splicing mode through the client or the multi-channel input director station, and the user no longer needs to manually switch the picture layout of the multi-channel input source or perform simple splicing through the software interface, thereby solving the problem of cumbersome operation, greatly improving operation efficiency and reducing the probability of error; the technical solution of the present application is used to control the multi-channel input director station, and no longer relies on a high-performance computer to run the director software, getting rid of the problems of large equipment size and high power consumption, and the multi-channel input director station is easy to carry, so it can be well adapted to mobile shooting scenes; through the dynamic resource allocation strategy and modular resource pool architecture for the multi-channel input director station (that is, the three-level resource pool formed by the acquisition pool, processing pool and encoding pool), intelligent director processing of multi-channel input data is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0008] in: Figure 1 This is an application environment diagram of a multi-channel video live broadcast processing method based on dynamic resource allocation in one embodiment; Figure 2 Flowchart of a multi-channel video live broadcast processing method based on dynamic resource allocation in one embodiment; Figure 3 This is a schematic diagram of the structure of the multi-input director station of this application; Figure 4 This is a structural block diagram of a multi-channel video live broadcast processing system based on dynamic resource allocation in one embodiment. DETAILED DESCRIPTION
[0009] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0010] The multi-channel video live broadcast processing method based on dynamic resource allocation provided by the embodiment of the present invention can be applied in the following situations: Figure 1The application environment includes: at least two acquisition devices 3, a multi-channel input director station 1, and a client 2. The client 2 is connected to the multi-channel input director station 1 via wired communication technology and / or wireless communication technology. The acquisition device 3 is connected to the multi-channel input director station 1 via wired communication technology and / or wireless communication technology.
[0011] Optionally, the multi-channel input director station 1 is configured to implement the multi-channel video live broadcast processing method based on dynamic resource allocation of the present application, specifically including: obtaining a live broadcast signal; obtaining a picture splicing mode in response to the live broadcast signal, wherein the picture splicing mode is data input by the user through the client 2 or the multi-channel input director station 1; based on the dynamic resource allocation strategy, collection pool, processing pool and encoding pool for the multi-channel input director station 1, data collection, data splicing and encoding output are performed according to the picture splicing mode.
[0012] Optionally, the client 2 is configured to implement the multi-channel video live broadcast processing method based on dynamic resource allocation of the present application, specifically including: obtaining a live broadcast signal; responding to the live broadcast signal to obtain a picture splicing mode, wherein the picture splicing mode is data input by the user through the client 2 or the multi-channel input director station 1; based on the dynamic resource allocation strategy, collection pool, processing pool and encoding pool for the multi-channel input director station 1, controlling the multi-channel input director station 1 to perform data collection, data splicing and encoding output according to the picture splicing mode.
[0013] The user inputs the picture splicing mode through the client 2 or the multi-channel input director station 1, and the user no longer needs to manually switch the picture layout of the multi-channel input source or perform simple splicing through the software interface, thereby solving the problem of cumbersome operation, greatly improving operation efficiency and reducing the probability of error; the technical solution of the present application is used to control the multi-channel input director station 1, and no longer relies on a high-performance computer to run the director software, getting rid of the problems of large equipment size and high power consumption, and the multi-channel input director station 1 is easy to carry, so it can be well adapted to mobile shooting scenes; through the dynamic resource allocation strategy and modular resource pool architecture for the multi-channel input director station 1 (that is, the three-level resource pool formed by the acquisition pool, processing pool and encoding pool), intelligent director processing of multi-channel input data is realized.
[0014] The client 2 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices.
[0015] The multi-input broadcast director station 1 includes: a communication component, a memory, a processor, and a computer program stored in the memory and runnable on the processor. The processor is electrically connected to the communication component. When the processor executes the computer program, the steps of the multi-channel video live broadcast processing method based on dynamic resource allocation are implemented.
[0016] The communication components include: a wireless communication module and / or a wired communication module. The wireless communication module uses wireless communication technology to communicate wirelessly with devices other than the multi-input director station 1. The wired communication module uses wired communication technology to communicate with devices other than the multi-input director station 1. The wireless communication module can be selected from existing technologies and will not be described in detail here. The wired communication module can be selected from existing technologies and will not be described in detail here.
[0017] See also Figure 3 Optionally, the first side 19 of the multi-channel input control station 1 is provided with a first HD input interface 11, a second HD input interface 12, a HD output interface 13, and a TYPE-C (USB interface standard) output interface 14. The second side 18 of the multi-channel input control station 1 is provided with a microphone input interface 15. The multi-channel input control station 1 obtains HD or SD video data input by the camera 31 (i.e., the acquisition device 3) from the first HD input interface 11 via a wired communication module. The multi-channel input control station 1 obtains HD or SD video data input by the camera 32 (i.e., the acquisition device 3) from the second HD input interface 12 via a wired communication module. The multi-channel input control station 1 outputs video data from the HD output interface to a display device 41 (i.e., a receiving device), such as a display screen or a laptop computer, via a wired communication module. The multi-channel input control station 1 outputs video data from the TYPE-C output interface 14 to a desktop computer 42 (i.e., a receiving device) via a wired communication module. The multi-channel input director station 1 obtains the audio data input by the microphone (that is, the acquisition device 3) through the microphone interface 15.
[0018] Optionally, the multi-input control station 1 supports video capture in formats such as YUY2 and MJPEG. The maximum resolution of the video captured by the multi-input control station 1 is 1080P, and the refresh rate is 60Hz. The multi-input control station 1 supports HD and 1080P loop-out (i.e., encoded video output). The audio captured by the multi-input control station 1 has a sampling rate of 48kHz. YUY2 is a common video data format, a sampling format in the YUV (luminance-chrominance) color space. MJPEG (Motion-JPEG), also known as dynamic JPEG, is a video encoding format.
[0019] It can be understood that the multi-input director station 1 is provided with one or more buttons, and the user inputs signals (such as live broadcast signals, switching signals, etc.) to the processor by pressing the buttons.
[0020] Optionally, the multi-input director station 1 sets a separate button for each picture splicing mode, and there is no need to reuse one button for multiple modes, which helps to simplify user operations.
[0021] Optionally, the button of the multi-channel input director station 1 corresponds to at least one signal. By multiplexing the buttons, the number of buttons to be set is reduced, and the cost of the multi-channel input director station 1 is reduced.
[0022] Optionally, the multi-input director station 1 is provided with a touch screen, and the user inputs signals and picture splicing modes to the processor by operating on the touch screen.
[0023] It can be understood that the data collected by the multi-channel input director station 1 includes at least two channels of video data and at least one channel of audio data.
[0024] The present invention is described in detail below through specific examples.
[0025] See also Figure 2 As shown, Figure 2 A flowchart of a multi-channel video live broadcast processing method based on dynamic resource allocation provided by an embodiment of the present invention is provided. The method is used to control a multi-channel input director station, and the method includes the following steps: S1: Get live broadcast signal; The live broadcast signal is a signal to start live broadcast.
[0026] Specifically, it can be data input by the user through the client or the multi-channel input director station, or it can be a live broadcast signal automatically triggered after the multi-channel input director station is turned on and initialized.
[0027] S2: Responding to the live broadcast signal, obtaining a picture splicing mode, wherein the picture splicing mode is data input by a user through a client or the multi-channel input director station; Specifically, when the live broadcast signal is obtained, the picture splicing mode input in real time by the user through the client or the multi-channel input director station can be obtained, and the picture splicing mode stored in advance by the user can also be obtained.
[0028] The image splicing mode includes image frame splicing description data and audio splicing description data. The image frame splicing description data describes how to splice image frames from multiple video channels. The audio splicing description data describes how to combine multiple audio channels. The image frame splicing description data records the splicing method and the image splicing ratio.
[0029] For example, when there are two channels of video input to the multi-channel input director station, the pictures of the two channels of video (that is, the image frames in the video) are spliced in a single-channel output, equal left and right ratios, unequal left and right ratios, equal top and bottom ratios, unequal top and bottom ratios, picture-in-picture, etc.
[0030] Audio splicing description data: highlight one audio channel and block the other audio channel, or block all audio channels, or both audio channels exist at the same time.
[0031] S3: Based on the dynamic resource allocation strategy, acquisition pool, processing pool and encoding pool for the multi-input director station, data acquisition, data splicing and encoding output are performed according to the picture splicing mode.
[0032] Specifically, based on the dynamic resource allocation strategy for the multi-input director station, according to the actual description data of the multi-input director station (for example, one or more of the acquisition performance data, output performance data and output mode data), the resource allocation data (for example, thread resources) corresponding to the acquisition pool, processing pool and encoding pool are allocated respectively. Based on the target control data, according to the resource allocation data corresponding to the acquisition pool, the acquisition pool is controlled to perform data acquisition; based on the target control data, according to the resource allocation data corresponding to the processing pool, the processing pool is controlled to perform data splicing; based on the target control data, according to the resource allocation data corresponding to the encoding pool, the encoding pool is controlled to perform encoding output.
[0033] The dynamic resource allocation strategy is an algorithm that adjusts the allocation of system resources (CPU, memory, and bandwidth) in real time. Based on the current input load (such as the number of video channels and resolution), output requirements (such as encoding complexity), and user configuration (such as latency threshold), it dynamically optimizes the proportion of resources in the three stages of acquisition, processing, and encoding to ensure efficient and low-latency broadcast processing.
[0034] Optionally, the dynamic resource allocation strategy uses a weighted priority algorithm to calculate the resource allocation weights corresponding to the acquisition pool, processing pool, and encoding pool in real time based on the input load (number of video channels, resolution, frame rate, and audio parameters) and output requirements (coding complexity, delay threshold). The resource allocation weight formula is: Resource allocation weight = α × input load + β × output demand + γ × delay threshold Among them, α, β, and γ are adjustable coefficients. The specific values of α, β, and γ can be determined by fitting based on multiple sets of test data.
[0035] The acquisition pool is a resource pool responsible for collecting and temporarily storing multiple channels of video and audio. Its core functions include data segmentation, timestamp alignment, format conversion (such as YUY2 to RGB), and data buffering and asynchronous processing through queue management.
[0036] The processing pool is responsible for image stitching, image quality optimization, and audio and video synchronization. By invoking image processing algorithms (such as interpolation and color correction) and audio processing modules (such as mixing and noise reduction), the processing pool integrates multiple input data streams into a single output stream (that is, audio and video data) that conforms to the user's stitching mode.
[0037] The encoding pool compresses and encodes the processed audio and video data according to the target format (such as H.264) and pushes it to the resource pool of the output device. The encoding pool supports dynamic bitrate adjustment, multi-protocol streaming (such as RTMP and SRT), and error recovery mechanisms.
[0038] Among them, data collection is to process the data input into the multi-channel director station according to the preset processing requirements (for example, format conversion, segmentation) and then cache it according to the preset caching method (for example, the data segments corresponding to the same channel of input data are cached in the same queue).
[0039] Data splicing is to splice the results of data collection into images and combine audio and video to form a new audio segment, and cache the new audio segment.
[0040] Encoding output is to encode the new audio segment according to the required encoding format, and then transmit the encoded data to a designated device (for example, a laptop) or to a designated live broadcast application.
[0041] The user inputs the picture splicing mode through the client or the multi-channel input director station, and the user no longer needs to manually switch the picture layout of the multi-channel input source or perform simple splicing through the software interface, thereby solving the problem of cumbersome operation, greatly improving operation efficiency and reducing the probability of error; the technical solution of the present application is used to control the multi-channel input director station, and no longer relies on a high-performance computer to run the director software, getting rid of the problems of large equipment size and high power consumption, and the multi-channel input director station is easy to carry, so it can be well adapted to mobile shooting scenes; through the dynamic resource allocation strategy and modular resource pool architecture for the multi-channel input director station (that is, the three-level resource pool formed by the acquisition pool, processing pool and encoding pool), intelligent director processing of multi-channel input data is realized.
[0042] In one embodiment, the steps of performing data acquisition, data splicing, and encoding output according to the picture splicing mode based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool for a multi-channel input director station include: S31: Acquire target control data and resource allocation data, wherein the target control data includes: acquisition sub-data, picture processing sub-data, and encoding output sub-data; S32: performing data collection, data splicing, and encoding output according to the target control data, the resource allocation data, the collection pool, the processing pool, the encoding pool, and the picture splicing mode; The target control data is determined by the following steps: obtaining an analysis signal, obtaining acquisition performance data, output performance data, output mode data, and delay configuration data in response to the analysis signal, and determining the target control data based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data, wherein the delay configuration data is data input by a user through a client or the multi-input director console; The resource allocation data is determined by the following steps: responding to the analysis signal, determining the resource allocation data corresponding to the acquisition pool, the processing pool, and the encoding pool respectively according to the dynamic resource allocation strategy based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data.
[0043] Specifically, the hardware interface or software module monitors user operations or system status changes (such as live broadcast startup, mode switching, etc.) in real time to trigger analysis signals. This analysis signal may come from user key operations, client instructions, or timed tasks of the program implementing this application.
[0044] The acquisition performance data includes the hardware capability parameters of the multi-input director station for acquiring input data (such as the maximum number of supported input channels, resolution, frame rate, encoding format, etc.).
[0045] The output performance data describes the performance of the external output data of the multi-channel input director station to the multi-channel input director station, such as the resolution, bit rate, network transmission delay and other parameters of the encoded output.
[0046] Output mode data: describes the configuration requirements for the output data of the multi-input director station, including: (1) Output protocol and transmission method: the output protocol (such as RTMP, SRT, HLS) and transmission method (real-time streaming, scheduled push) set according to the receiving device or application type; (2) Output quality parameters: user-defined or pre-stored target resolution (such as 1080P), bit rate upper limit (such as 8Mbps), frame rate (such as 60fps), audio encoding format (such as AAC); (3) Terminal adaptation rules: dynamically adjust encoding parameters (such as resolution degradation, bit rate adaptation) for different receiving devices (such as mobile terminals, PC terminals).
[0047] The delay configuration data is the minimum and maximum allowable delay thresholds set by the user through the client or the multi-channel input director, which is used to dynamically adjust the data processing rhythm. Delay refers to the time interval between acquisition and output encoding.
[0048] Optionally, the target control data is determined by using a table lookup method based on the acquisition performance data, the output performance data, the output mode data and the delay configuration data.
[0049] Optionally, a regular expression is used to calculate the indicator value of each first indicator based on the acquisition performance data, the output performance data, the output mode data and the delay configuration data, and the target control data is determined by a table lookup method based on each indicator value.
[0050] The selection range of the first indicator includes but is not limited to: input load indicator, processing capacity indicator, output demand indicator and delay tolerance indicator.
[0051] Input load metrics are calculated based on the acquisition performance data. Input load metrics include: (1) Number of input channels: matching the number of input interfaces; (2) Average resolution: extracting a value from the resolution of the video input to the multi-channel input station; (3) Total frame rate: summing the frame rates of all videos input to the multi-channel input station; and (4) Peak bit rate: extracting the maximum value from the input bit rate of the data input to the multi-channel input station.
[0052] Processing capacity indicators are calculated based on collected performance data and output performance data. Processing capacity indicators include: (1) CPU utilization: obtained through the system monitoring interface in real time; (2) memory usage: matching the percentage of memory usage; (3) current processing delay: calculated based on timestamp difference.
[0053] Output demand indicators are calculated based on output performance data and output mode data. Output demand indicators include: (1) target bit rate: the upper limit of the bit rate extracted from the output mode data; (2) coding complexity: graded according to the coding format (e.g., H.264 = 1, H.265 = 2); (3) audio synchronization level: extracted from the audio and video combination data.
[0054] The delay tolerance metric is calculated based on the delay configuration data. The delay tolerance metric includes: (1) Maximum allowed delay; (2) Real-time priority: mapped to a level based on the delay threshold set by the user, for example, delay <100ms = high, delay 100-200ms = medium, and delay >200ms = low.
[0055] Optionally, the acquisition performance data, the output performance data, the output mode data and the delay configuration data are input into a pre-trained first model for classification prediction, and a vector element with the largest value is selected from the predicted vector, and the control data corresponding to the classification category corresponding to the selected vector element is used as the target control data.
[0056] The first model is a pre-trained multi-classification model. The model structure and model training method of the first model can be selected from existing technologies.
[0057] The collected sub-data includes but is not limited to dynamically calculating the segmentation interval based on the resolution, frame rate, and bit rate of the input video in combination with the delay configuration data.
[0058] The image processing sub-data includes but is not limited to defining the splicing mode (such as the left-right split screen ratio) and image quality processing parameters (such as the color correction matrix and the interpolation algorithm).
[0059] Encoding output sub-data includes but is not limited to setting encoding format, bit rate and audio synchronization strategy.
[0060] Based on the dynamic resource allocation strategy, system resources (such as CPU threads, memory, and bandwidth) are allocated to three-level resource pools according to priority to determine the resource allocation data of each resource pool (i.e., acquisition pool, processing pool, and encoding pool): (1) Acquisition pool: allocates resources for real-time acquisition and segmented storage of multiple videos; (2) Processing pool: allocates resources for image splicing, image quality optimization, and audio and video synchronization; (3) Encoding pool: allocates resources for final encoding and output stream push.
[0061] Specifically, based on the acquisition sub-data, according to the resource allocation data corresponding to the acquisition pool, the acquisition pool is controlled to perform data acquisition; based on the picture processing sub-data, according to the resource allocation data corresponding to the processing pool, the processing pool is controlled to perform data splicing; based on the encoding output sub-data, according to the resource allocation data corresponding to the encoding pool, the encoding pool is controlled to perform encoding output.
[0062] Among them, the collection pool is controlled by allocating resource data corresponding to the collection pool, and each input video is divided into time-stamp aligned data segments (such as 50ms segments) according to the segmentation interval of the collected sub-data. According to the delay configuration data, the segmentation granularity is dynamically adjusted (such as increasing the segment length and reducing the processing frequency when high delay tolerance occurs), and the segmented data is stored in the first storage area of the collection pool. Independent queue management is established according to the number of input channels.
[0063] Through the resource allocation data corresponding to the processing pool, the processing pool is controlled to extract the video segments with the same timestamp from the first storage area, perform pixel-level splicing according to the picture splicing mode (such as left and right split screen), and then perform image quality normalization processing, and finally perform audio and video synchronization processing.
[0064] The image quality normalization processing specifically includes: (1) color correction: dynamically matching the color space of each video segment (such as sRGB to BT.709); (2) brightness compensation: interpolating the edge areas of the picture to eliminate the brightness difference at the splicing point; (3) resolution mapping: uniformly scaling inputs of different resolutions to the target resolution (such as 1080P); (4) audio and video synchronization processing: based on the audio and video combination data, selecting the main audio stream or mixed output, and strictly aligning it with the picture timestamp.
[0065] Through the resource allocation data corresponding to the encoding pool, the encoding pool is controlled to compress the spliced image and audio according to the encoding output sub-data (such as H.264 encoding, with a bit rate of 8Mbps). Adaptive bit rate control technology is used to dynamically adjust the output bit rate according to the network bandwidth, and push it to the designated device (such as the live broadcast server, local display) or store it as a file.
[0066] This embodiment determines target control data and resource allocation data based on acquisition performance data, output performance data, output mode data, and delay configuration data, and then performs data acquisition, splicing, and encoding output based on this data, the acquisition pool, the processing pool, the encoding pool, and the image splicing mode, thereby efficiently and flexibly processing the multi-channel data input to the multi-channel input director station. This method can be adaptively adjusted based on various factors such as different device performance, output requirements, and user-set delays to ensure the accuracy of data acquisition, the rationality of data splicing, and the effectiveness of encoding output, thereby improving the adaptability and operational efficiency of the entire director process in different application scenarios and reducing problems such as performance degradation or data processing errors caused by unreasonable resource allocation or mismatched settings.
[0067] In one embodiment, the step of performing data acquisition, data splicing, and encoding output according to the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the picture splicing mode includes: S3211: Based on the resource allocation data corresponding to the collection pool and the collection sub-data of the target control data, segment the input data of each channel of the multi-channel input director station using the same segmentation timestamp, generate data segments that conform to the storage structure of the first storage area of the collection pool as first data segments, and store each first data segment in the first storage area; The dynamic determination of the acquired sub-data includes calculating a segmentation time interval based on the resolution, frame rate, and bit rate of the video input data as an initial interval, adjusting the initial interval based on a maximum allowable delay threshold corresponding to the delay configuration data as a target interval, determining a segmentation granularity based on the acquisition performance data and the delay configuration data, and using the target interval and segmentation granularity as the acquired sub-data. By calculating the target interval and segmentation granularity in real time, a balance between delay and computational overhead is achieved.
[0068] The first storage area is divided into independent queues according to the number of input paths, and each queue stores data segments obtained by segmented processing of the corresponding input data.
[0069] Segmentation processing is based on a segmentation method that strictly aligns the segment timestamps of multiple input data. The continuous input video / audio stream is divided into discrete data segments at fixed time intervals to facilitate subsequent splicing and synchronization.
[0070] Segment granularity, the minimum time unit of a segment (e.g., 10ms), determines the frequency and accuracy of data processing. Low-latency scenarios require a small segment granularity to achieve high-frequency processing, while high-throughput scenarios can use a larger segment granularity.
[0071] Among them, initial interval = frame rate × resolution weight / (bit rate × complexity coefficient), the complexity coefficient is determined by a lookup table method based on the encoding format, and the resolution weight is data determined by a lookup table method based on the average resolution of the input data of the video type.
[0072] The initial interval is limited according to the maximum delay threshold. For example, if the initial interval is 9ms but the maximum allowed delay is 200ms, the target interval is adjusted to min(9ms,200ms / N), where N is the number of input channels of the multi-input director.
[0073] According to the acquisition performance data and the delay configuration data, a segmentation granularity is determined by using a table lookup method.
[0074] It can be understood that the collected sub-data include: target interval and segmentation granularity.
[0075] Optionally, the collected sub-data further includes queue slices. Queue slices are pre-set data, and each channel of input data corresponds to a fixed queue slice.
[0076] Specifically, a hardware clock or NTP protocol (Network Time Protocol) is used to synchronize the timestamps of the multi-channel input data of the multi-channel input director station to ensure the alignment of all input data; at the end time of each interval period of the target interval of the collected sub-data, each input data of the multi-channel input director station is segmented according to the segmentation granularity of the collected sub-data using the same segmentation timestamp, and data segments with a length equal to the segmentation granularity are segmented (if the input data is insufficient, it can also be a data segment smaller than the segmentation granularity). If the format of the data segment meets the requirements of the storage structure of the first storage area of the collection pool, the data segment is used as the first data segment; if the format of the data segment does not meet the requirements of the storage structure of the first storage area of the collection pool, the data segment is format-converted and used as the first data segment; the first data segment is stored in the queue slice corresponding to the input data corresponding to the first data segment.
[0077] It can be understood that the first storage area stores the first data segments in the order of timestamps.
[0078] Queue sharding achieves resource isolation, and the independent queue design avoids data competition between multiple inputs and improves stability.
[0079] This embodiment calculates the initial interval based on the resolution, frame rate, and bit rate of the video-type input data, adjusts it to the target interval based on the delay configuration data, and simultaneously determines the segmentation granularity based on the acquisition performance data to generate the acquired sub-data, which is then segmented using the same segmentation timestamp. This approach ensures that the generated data segments conform to the storage structure of the first storage area of the acquisition pool, facilitates efficient storage of the first data segments, and optimizes data acquisition and storage. This refined processing based on multiple data types can better adapt to different input data characteristics and system requirements, improve the rationality and effectiveness of data acquisition and storage, and reduce confusion and errors in the data processing process.
[0080] In one embodiment, the step of performing data acquisition, data splicing, and encoding output according to the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the picture splicing mode further includes: S3221: Based on the resource allocation data corresponding to the processing pool, and in accordance with the picture splicing mode and the splicing configuration of the splicing sub-data of the target control data, splicing picture data with the same timestamp for each of the first data segments in the first storage area to obtain a first picture; The stitching configuration describes the geometric alignment rules (such as coordinate offset and scaling) of multiple video channels.
[0081] Specifically, all first data segments with the same timestamp (such as 0-50ms) are extracted from the first storage area, the position and proportion of each video are determined according to the screen stitching mode, the resolution of the image frame of each input video is adjusted according to the stitching configuration, and according to the stitching configuration, multiple image frames with adjusted resolutions (that is, screen data) are stitched onto the same canvas according to the coordinate offset, the image on the canvas is used as the first screen, and the first screen is stored in a temporary buffer corresponding to the processing pool for subsequent image quality normalization processing.
[0082] S3222: Based on the resource allocation data corresponding to the processing pool and according to the quality processing data in the spliced sub-data, perform image quality normalization processing on the first image to obtain a target image; The quality processing data describes a set of parameters used to improve the visual quality of the spliced image, including color correction matrix, brightness compensation algorithm, resolution mapping rules, etc.
[0083] Image quality normalization is a process that adjusts the color, brightness, and resolution of the image frame of the first screen to a consistent standard, eliminating visual discontinuities caused by device differences.
[0084] The first image may have color inconsistencies, brightness differences, and resolution mismatches. To address these issues, the color correction matrix in the quality-processed data is first applied based on the color space of the input video (e.g., sRGB, BT.709) to unify each image frame in the first image to the target color standard (e.g., BT.709) for color correction. Then, based on the brightness compensation algorithm in the quality-processed data, the spliced area of the color-corrected first image (e.g., the center seam of the left and right split screens) is detected and the brightness difference between the left and right images is interpolated and smoothed to achieve brightness compensation. Finally, based on the resolution mapping rules in the quality-processed data, the low-resolution input image is interpolated and upscaled to the target resolution, or the high-resolution image is downsampled to the target resolution to achieve unification to the target resolution (i.e., scaled image). An adaptive sharpening filter is applied to the scaled image to enhance detail clarity, and the image processed by the adaptive sharpening filter is used as the target image.
[0085] S3223: Based on the resource allocation data corresponding to the processing pool, according to each of the target images and each of the first data segments of the audio type in the first storage area, the image and audio are matched according to the audio-visual combination data in the splicing sub-data to obtain a second data segment, and the second data segment is stored in the second storage area of the processing pool.
[0086] The audio and video combination data describes the synchronization rules between audio and picture, such as main audio stream selection, audio delay compensation, mixing strategy, etc.
[0087] The second data segment is a complete data unit after audio and video synchronization, including the target picture and the corresponding audio stream.
[0088] The second storage area is a buffer area for temporarily storing the spliced and optimized second data segments. The data segments in the second storage area are arranged in chronological order for the encoding pool to extract as needed.
[0089] Specifically, audio is first extracted and processed based on the audio-visual combination data. The audio segment corresponding to the timestamp of the target image (i.e., the first data segment of audio type) is extracted from the first storage area. If the audio-visual combination data is the primary audio stream (e.g., input channel 1), other audio is masked. If the audio-visual combination data describes a mixed audio stream, multiple audio channels are mixed according to the proportions specified in the audio-visual combination data. Then, each target image and the processed audio stream are encapsulated into a composite data packet (i.e., the second data segment) based on the timestamp. For example, this may be encapsulated using the MPEG-TS (a video encapsulation format based on the MPEG-2 standard) format, with the video track being H.264 (one of the ITU-T video codec technology standards named in the H.26x series) and the audio track being AAC (Advanced Audio Coding). Finally, the second data segment is pushed into the second storage area in chronological order.
[0090] This embodiment first splices the picture data of the first data segment with the same timestamp according to the picture splicing mode and the splicing configuration of the splicing sub-data in the target control data to obtain the first picture, thereby realizing the preliminary integration of the picture. Then, the first picture is subjected to picture quality normalization processing based on the quality processing data to obtain the target picture, which helps to improve the overall quality of the picture and make it more in line with the requirements. Finally, the target picture and the audio are matched according to the audio-visual combination data to obtain the second data segment and stored in the second storage area of the processing pool. This corresponding storage of the picture and the audio not only makes the data more organized in subsequent processing, but also optimizes the process from data acquisition to preliminary integration processing as a whole, improves the accuracy and orderliness of data processing, can better meet the needs of picture and audio processing in various application scenarios, and at the same time improves the data processing efficiency of the entire system.
[0091] In one embodiment, the step of performing image quality normalization processing on the first image according to the quality processing data in the splicing sub-data to obtain the target image includes: S32221: Performing dynamic correction of a color correction matrix, brightness cross-region interpolation compensation, and resolution adaptive mapping on the first picture according to the quality processing data in the spliced sub-data to obtain a second picture; Specifically, first, according to the color space of the input video (such as sRGB, BT.709), the color correction matrix in the quality processing data is applied to unify the image frames in the first picture to the target color standard (such as BT.709) to achieve color correction; then, according to the brightness compensation algorithm of the quality processing data, the splicing area of the first picture after color correction (such as the middle seam of the left and right split screens) is detected, and the brightness difference between the left and right pictures is interpolated and smoothed to achieve brightness compensation and eliminate the sense of visual fragmentation; finally, according to the resolution mapping rule of the quality processing data, the low-resolution input picture is interpolated and amplified, or the high-resolution picture is downsampled to unify it to the target resolution (that is, the scaled picture), and the first picture at this time is used as the second picture.
[0092] S32222: Use the second picture as the target picture, or identify a specific object area on the second picture. If a specific object area is identified, apply dynamic interpolation compensation to the specific object area according to the quality processing data to obtain the target picture.
[0093] The specific object can be a human body or a human face, an animal, or a certain type of object.
[0094] After color correction, brightness compensation, and resolution mapping, the intermediate image (also known as the second image) has eliminated basic image quality issues, but has not yet been optimized for image regions corresponding to specific objects. To further improve image quality, computer vision algorithms (such as YOLO and OpenCV Haar cascade classifiers) are first used to detect image regions corresponding to specific objects in the image (referred to as specific object regions). For these identified specific object regions, data is processed based on the quality, and a higher-precision interpolation algorithm (such as bicubic interpolation) is applied to improve detail clarity and dynamic range. The final output is an optimized image (also known as the target image). The optimized image has the advantages of consistent color, uniform brightness, and adaptive resolution, and the details of the image regions corresponding to specific objects are enhanced. By applying high-computation processing only to image regions corresponding to specific objects, we balance performance and quality.
[0095] This embodiment obtains a second image by dynamically correcting the color correction matrix, performing cross-region brightness interpolation compensation, and adaptively mapping the resolution of the first image. This series of operations can effectively improve the color accuracy, brightness uniformity, and resolution adaptability of the image, thereby optimizing the image in multiple dimensions. Furthermore, the second image is used as the target image, or when a specific object area is identified, dynamic interpolation compensation is applied to the specific object area based on quality processing data to obtain the target image. This measure, while maintaining overall image optimization, can perform special processing on specific object areas, enhancing the visual effects of specific object areas and providing better visual presentation when focusing on key areas. This makes the image more consistent with processing requirements in terms of both visual quality and processing of specific areas, thereby improving the overall image quality.
[0096] In one embodiment, obtaining a live broadcast signal includes: S11: Acquire a switching signal, and acquire the live broadcast signal in response to the switching signal, wherein the switching signal is a signal input by a user through a client or the multi-channel input director station; The switching signal is a command triggered by the user through the client or multi-channel input director console, which is used to start live broadcast, switch screen layout, etc.
[0097] The live broadcast signal is a control instruction generated after receiving the switching signal to start the live broadcast process, triggering the entire process of data collection, splicing and encoding output.
[0098] Specifically, the multi-input director station monitors the physical button signals (such as the "start live broadcast" button) of the multi-input director station through the GPIO (General-purpose input / output) interface, and / or the client sends a switching signal to the multi-input director station.
[0099] The live broadcast process is initialized by obtaining the live broadcast signal in response to the switching signal.
[0100] The step of using the second picture as the target picture, or identifying a specific object area on the second picture within a preset time period after the switching signal is generated, and applying dynamic interpolation compensation to the specific object area based on the quality processing data to obtain the target picture if the specific object area is identified, includes: S322221: performing a smooth transition process on the second picture according to the designated picture to obtain the target picture, or performing a specific object area recognition on the second picture. If a specific object area is recognized, applying dynamic interpolation compensation and a smooth transition process to the specific object area according to the quality processing data and the designated picture to obtain the target picture. Otherwise, performing a smooth transition process on the second picture according to the designated picture to obtain the target picture. The designated picture is the second picture that is closest to the generation time of the switching signal in history.
[0101] A circular buffer is maintained in the multi-input director station to save the image data of the last N seconds (such as 5 seconds), and quickly retrieve it by timestamp. In other words, a specified image can be obtained from the circular buffer.
[0102] It is understandable that if the ring buffer is empty, a black screen or a static background is used as the designated screen by default.
[0103] Specifically, the method of smooth transition processing is to perform Alpha blending on the two images. Alpha blending is an image processing technology used to achieve transparency and blending effects of images.
[0104] Alpha blending of two images (the designated image and the second image) is performed through a time window (within a preset time period after the generation time of the switching signal) to eliminate image jumps and improve viewing smoothness.
[0105] This embodiment first obtains a live signal by responding to a switching signal. This mechanism can flexibly obtain live content based on the signals input by the user through the client or the multi-input director, increasing operational flexibility and user controllability. Then, when performing image processing within a preset period of time after the switching signal is generated, whether it is smooth transition processing of the second image based on a specified image (i.e., the second image that is historically closest to the generation time of the switching signal) to obtain the target image, or when a specific object area is identified, dynamic interpolation compensation and smooth transition processing are applied to the specific object area based on the quality processing data and the specified image to obtain the target image, these operations can ensure the continuity and stability of the image during signal switching. Smooth transition processing avoids the abrupt feeling that may occur when the image switches, while special processing for specific object areas enhances the visual effect of the specific object area while ensuring the overall transition of the image is natural, so that the image can provide better visual presentation when focusing on the key area, improving the user's experience of watching live broadcasts.
[0106] In one embodiment, the step of splicing picture data with the same timestamp for each of the first data segments in the first storage area according to the picture splicing mode and the splicing configuration of the splicing sub-data of the target control data to obtain the first picture includes: S32211: Performing a missing type detection on the picture data of the same timestamp on each of the first data segments of the video type in the first storage area according to the picture splicing mode and the splicing configuration, to obtain a first result; The first result is the output of the missing type detection, marking the data integrity status (such as "complete" or "missing 0-50ms data of input path 2").
[0107] Specifically, first, extract all the first data segments with the same timestamp from the first storage area, and generate a list to be detected according to the number of input paths; traverse the list to be detected and check whether each input path contains the data segment with the current timestamp; if data exists on all input paths, output {"status":"complete"}; if there is missing data, output {"status":"incomplete","missing routes":missing_routes,"missing timestamp":current_ts}.
[0108] S32212: If the first result is incomplete, splicing the first data segments in the first storage area with the same timestamp and the incomplete first result according to the splicing configuration of the splicing sub-data of the target control data based on an exception handling strategy to obtain the first picture, wherein the exception handling strategy is determined based on the first result and the picture splicing mode; Exception handling strategies are predefined rules based on the missing data type and image stitching mode, used to repair or compensate for missing data. Exception handling strategies include, but are not limited to, interpolation and degradation. Interpolation fills missing segments with the previous frame's data or a solid color. Degradation ignores missing paths and stitches only available inputs.
[0109] Specifically, if the first result is incomplete, the first data segment of the previous timestamp on the missing path is extracted for the timestamp corresponding to the first result. A motion compensation algorithm (such as optical flow) is applied to generate an approximate current frame, or a preset solid color is used as the approximate current frame. The approximate current frame is then added to the first data segment. Based on the splicing configuration of the splicing sub-data of the target control data, each first data segment (compensated data segments and data segments not requiring compensation) is geometrically spliced according to the splicing configuration to generate the first image. Through missing data detection and dynamic compensation, the live stream can continue to be output even when some input is abnormal.
[0110] S32213: If the first result is complete, splicing the picture data with the same timestamp for each of the first data segments in the first storage area according to the picture splicing mode and the splicing configuration to obtain the first picture.
[0111] Specifically, if the first result is complete, there is no need to generate an approximate current frame for the timestamp corresponding to the first result. The picture data with the same timestamp of each first data segment in the first storage area is directly spliced according to the picture splicing mode and the splicing configuration to obtain the first picture.
[0112] It may contain picture holes or misalignments caused by missing data, which requires further optimization. In order to solve this problem, this embodiment first performs a missing type detection on the picture data with the same timestamp, which can accurately grasp the integrity of the picture data in the first data segment of the video type. When the detection result is incomplete, picture splicing is performed based on the exception handling strategy determined according to the detection result and the picture splicing mode. This measure can flexibly deal with the situation of incomplete data, ensuring that a complete picture is still obtained as the first picture when there is missing data, avoiding the situation where the whole picture has picture holes or misalignments. When the detection result is complete, the picture splicing is performed according to the picture splicing mode and splicing configuration, which can efficiently and accurately complete the splicing of the picture data with the same timestamp when the data is complete, and obtain the first picture, which improves the accuracy, stability and adaptability of the picture splicing to different data states as a whole.
[0113] In one embodiment, the method further comprises: S41: Based on the resource allocation data corresponding to the collection pool, monitor the timestamp differences and audio delay parameters of the input data of the multi-channel input director station in real time to generate a synchronization offset set; A synchronization offset set records the timestamp differences between multiple input video / audio channels (for example, stream A is 50ms slower than stream B) and the audio / video synchronization errors. The synchronization offset set includes video timestamp differences and audio delay parameters.
[0114] Video timestamp difference: the relative delay between multiple input videos (unit: ms).
[0115] Audio delay parameter: the time difference between the audio and the corresponding video (for example, the audio lags behind the video by 20ms).
[0116] Optionally, synchronize the system time of all acquisition devices using the NTP protocol or hardware clock (such as PTP) to reduce initial clock deviation.
[0117] Specifically, timestamps are extracted from each video input, and the relative delay between the multiple inputs is calculated as the video timestamp difference. The timestamps of each audio input (such as the audio PTS) are extracted and compared with the timestamps of the corresponding video input data to obtain the audio delay parameter.
[0118] S42: Based on the resource allocation data corresponding to the acquisition pool, based on the adaptive clock synchronization technology and the dynamic interpolation technology, according to the synchronization offset set, dynamically align the first data segment of the video type in the first storage area.
[0119] Adaptive clock synchronization technology is an algorithm that dynamically adjusts the clock frequency or timestamp of input streams to eliminate time differences between multiple inputs. Timeline alignment is achieved by analyzing synchronization offsets, controlling the data acquisition rate, or performing frame processing (inserting / deleting redundant frames).
[0120] Specifically, based on the resource allocation data corresponding to the acquisition pool and adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set, and the timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with large delays (i.e., the analyzed offset is greater than the acquisition offset threshold) (e.g., input path 2 is 50ms slower), the acquisition rate is increased (e.g., from 30fps to 31fps), gradually reducing the time difference. For the first data segments of the video type in the first storage area, generated frames are inserted in the first data segments with positive offsets to catch up with other streams, and redundant frames are deleted in the first data segments with negative offsets to slow down the progress.
[0121] A positive offset means that the timestamp of a video is delayed (lags behind) relative to the reference timeline.
[0122] A negative offset means that the timestamp of a video stream is ahead of the reference timeline.
[0123] The reference time axis takes the timestamp of the earliest arriving input data as the reference time axis.
[0124] This embodiment generates a synchronization offset set by real-time monitoring of the timestamp differences and audio delay parameters of the multi-channel input director station input data based on the resource allocation data corresponding to the acquisition pool, and can accurately record the timestamp differences and other situations between the multi-channel input video / audio streams. This synchronization offset set provides a key basis for subsequent processing, so that the insertion or deletion of frames can be used to achieve accurate synchronization of multiple data channels. Then, based on the resource allocation data corresponding to the acquisition pool, adaptive clock synchronization technology and dynamic insertion technology are used to dynamically align the first data segment of the video type in the first storage area according to the synchronization offset set, effectively solving the problem of time asynchrony of the multi-channel input data, improving the synchronization and accuracy of the video data, thereby improving the coordination and stability of the entire process when processing multi-channel input data, and optimizing the final output effect.
[0125] In one embodiment, the step of dynamically aligning the first data segment of the video type in the first storage area according to the synchronization offset set based on the adaptive clock synchronization technology and the dynamic frame insertion technology includes: S421: Perform resolution detection on each of the first data segments in the first storage area that corresponds to the same timestamp and is of video type, and use the resolution as input resolution; Specifically, all video data segments (ie, the first data segment of the video type) with the same timestamp (eg, 0-50ms) are extracted from the first storage area, and the resolution parameter in the video segment header information is parsed to obtain the input resolution.
[0126] S422: When the input resolution is lower than a first threshold, using a nearest neighbor interpolation mode, determining an offset of each data segment on the time axis according to the synchronization offset set, and performing preliminary alignment on the first data segments of the video type in the first storage area by inserting or deleting frames at corresponding positions according to the offset; The nearest neighbor interpolation mode is a simple interpolation algorithm that directly uses the nearest pixel value to fill in new pixels. It has a small amount of calculation but may produce jagged edges.
[0127] The offset on the timeline is the delay or advance of a certain input video relative to the reference timeline.
[0128] Preliminary alignment means quickly correcting large-scale time offsets by inserting or deleting frames, but this may sacrifice some image quality.
[0129] Specifically, the offset of each input path at the same timestamp is read from the synchronization offset set. For the first data segment of the video type in the first storage area, if the offset is positive, the number of frames to be inserted is calculated, and the average value of the previous frame and the next frame is used as the inserted frame; if the offset is negative, the redundant frame is deleted.
[0130] S423: When the input resolution is between the first threshold and the second threshold, a bilinear interpolation mode and an adaptive clock synchronization technology are used to dynamically align the first data segment of the video type in the first storage area according to the synchronization offset set. The bilinear interpolation mode is a linear interpolation algorithm based on 4 adjacent pixels. It has better smoothing effect than the nearest neighbor and is suitable for medium resolution.
[0131] Dynamic alignment processing, a comprehensive alignment solution combining clock synchronization and high-precision interpolation, takes into account both timing accuracy and image quality.
[0132] Specifically, based on the resource allocation data corresponding to the acquisition pool and adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set, and the timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with large delays (i.e., the analyzed offset is greater than the acquisition offset threshold) (e.g., input path 2 is 50ms slower), the acquisition rate is increased (e.g., from 30fps to 31fps), gradually reducing the time difference. For the first data segment of the video type in the first storage area, if the offset is positive, bilinear interpolation is used to generate intermediate frames for the frames to be inserted, thereby achieving motion-compensated interpolation. S424: When the input resolution is higher than a second threshold, a bicubic interpolation mode and an adaptive clock synchronization technology are used to dynamically align the first data segment of the video type in the first storage area according to the synchronization offset set.
[0133] The bicubic interpolation mode uses a cubic polynomial interpolation algorithm based on 16 adjacent pixels, which retains more details and is suitable for high-resolution videos.
[0134] Specifically, based on the resource allocation data corresponding to the acquisition pool and adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set, and the timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with significant delays (i.e., the analyzed offset is greater than the acquisition offset threshold) (e.g., input path 2 is 50ms slower), the acquisition rate is increased (e.g., from 30fps to 31fps), gradually reducing the time difference. For the first video data segment in the first storage area, if the offset is positive, bicubic interpolation is used to generate intermediate frames for the frames to be inserted, thereby achieving motion-compensated interpolation.
[0135] That is, for low-resolution videos, this embodiment uses nearest neighbor interpolation with less computational effort to reduce latency; and for high-resolution videos, uses bicubic interpolation to retain more details.
[0136] Optionally, bicubic interpolation is only enabled at high resolutions and sufficient system resources, otherwise it degrades to bilinear interpolation.
[0137] System resource adequacy is determined by real-time monitoring of the following parameters: (1) CPU utilization: The CPU utilization of the current processing pool is lower than the preset threshold (e.g., ≤70%); (2) Available memory: The remaining memory percentage of the processing pool is higher than the preset threshold (e.g., ≥30%); (3) Processing delay: The current frame processing delay is less than 50% of the maximum allowed delay threshold; (4) Thread idle rate: The number of idle threads in the processing pool accounts for ≥ 20%.
[0138] Optionally, when the following conditions are met simultaneously, the system is considered to have sufficient resources and bicubic interpolation is enabled. The conditions include: CPU utilization ≤ 70%, available memory ≥ 30%, processing delay < maximum allowed delay × 0.5, and number of idle threads ≥ total number of threads × 20%. If any of the judgment conditions is not met, it will automatically degrade to bilinear interpolation.
[0139] This embodiment first performs resolution detection to obtain the input resolution, providing a basis for determining different processing modes. When the input resolution is below a first threshold, a nearest neighbor interpolation mode combined with adaptive clock synchronization technology is used to determine the offset based on a synchronization offset set. Initial alignment of the first video data segment is achieved by inserting or deleting frames. This approach provides a simple and efficient solution for aligning data segments on the time axis at low resolutions. When the input resolution is between the first and second thresholds, a bilinear interpolation mode and adaptive clock synchronization technology are used to dynamically align the video data based on the synchronization offset set, enabling more precise alignment of the video data at medium resolutions and improving alignment accuracy. When the input resolution is above the second threshold, a bicubic interpolation mode and adaptive clock synchronization technology are used to dynamically align the video data based on the synchronization offset set. This allows for more precise and high-quality dynamic alignment of video data segments at high resolutions, adapting to the alignment requirements of video data at different resolutions. Overall, this improves the flexibility, accuracy, and effectiveness of dynamic alignment of video data at different resolutions, optimizing video data processing.
[0140] See also Figure 4 As shown, in one embodiment, a multi-channel video live broadcast processing system based on dynamic resource allocation is provided, the system comprising: a multi-channel input director station 1 and a client 2, the client 2 being communicatively connected to the multi-channel input director station 1, and the multi-channel input director station 1 being configured to implement the above-mentioned multi-channel video live broadcast processing method based on dynamic resource allocation.
[0141] The user inputs the picture splicing mode through the client or the multi-channel input director station, and the user no longer needs to manually switch the picture layout of the multi-channel input source or perform simple splicing through the software interface, thereby solving the problem of cumbersome operation, greatly improving operation efficiency and reducing the probability of error; the technical solution of the present application is used to control the multi-channel input director station, and no longer relies on a high-performance computer to run the director software, getting rid of the problems of large equipment size and high power consumption, and the multi-channel input director station is easy to carry, so it can be well adapted to mobile shooting scenes; through the dynamic resource allocation strategy and modular resource pool architecture for the multi-channel input director station (that is, the three-level resource pool formed by the acquisition pool, processing pool and encoding pool), intelligent director processing of multi-channel input data is realized.
[0142] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0143] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by signaling related hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0144] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0145] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A multi-channel video live broadcast processing method based on dynamic resource allocation, characterized in that: The method is used to control a multi-channel input director station, and the method includes: Get live broadcast signal; In response to the live broadcast signal, obtaining a picture splicing mode, wherein the picture splicing mode is data input by a user through a client or the multi-channel input director station; Based on the dynamic resource allocation strategy, collection pool, processing pool and encoding pool for the multi-input director station, data collection, data splicing and encoding output are performed according to the picture splicing mode.
2. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 1 is characterized in that: The steps of performing data acquisition, data splicing and encoding output according to the picture splicing mode based on the dynamic resource allocation strategy, acquisition pool, processing pool and encoding pool for the multi-channel input director station include: Acquiring target control data and resource allocation data, wherein the target control data includes: acquisition sub-data, picture processing sub-data, and encoding output sub-data; Performing data acquisition, data splicing, and encoding output according to the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the picture splicing mode; The target control data is determined by the following steps: obtaining an analysis signal, obtaining acquisition performance data, output performance data, output mode data, and delay configuration data in response to the analysis signal, and determining the target control data based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data, wherein the delay configuration data is data input by a user through a client or the multi-input director console; The resource allocation data is determined by the following steps: responding to the analysis signal, determining the resource allocation data corresponding to the acquisition pool, the processing pool, and the encoding pool respectively according to the dynamic resource allocation strategy based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data.
3. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 2 is characterized in that: The step of performing data acquisition, data splicing, and encoding output according to the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the picture splicing mode includes: Based on the resource allocation data corresponding to the collection pool and according to the collection sub-data of the target control data, segmenting each input data of the multi-channel input director station using the same segmentation timestamp, generating data segments that conform to the storage structure of the first storage area of the collection pool as first data segments, and storing each first data segment in the first storage area; Among them, the dynamic determination of the acquisition sub-data includes: calculating the segmentation time interval based on the resolution, frame rate and bit rate of the input data of video type as the initial interval, adjusting the initial interval based on the maximum allowable delay threshold corresponding to the delay configuration data as the target interval, determining the segmentation granularity according to the acquisition performance data and the delay configuration data, and using the target interval and the segmentation granularity as the acquisition sub-data.
4. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 3 is characterized in that: The step of performing data acquisition, data splicing and encoding output according to the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool and the picture splicing mode further includes: splicing picture data with the same timestamp for each of the first data segments in the first storage area based on the resource allocation data corresponding to the processing pool and according to the picture splicing mode and the splicing configuration of the splicing sub-data of the target control data to obtain a first picture; Based on the resource allocation data corresponding to the processing pool and according to the quality processing data in the spliced sub-data, performing image quality normalization processing on the first image to obtain a target image; Based on the resource allocation data corresponding to the processing pool, according to each of the target pictures and each of the first data segments of type audio in the first storage area, the picture and audio are matched according to the audio-visual combination data in the splicing sub-data to obtain a second data segment, and the second data segment is stored in the second storage area of the processing pool.
5. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 4 is characterized in that: The step of performing image quality normalization processing on the first picture according to the quality processing data in the splicing sub-data to obtain a target picture includes: performing dynamic correction of a color correction matrix, brightness cross-region interpolation compensation, and resolution adaptive mapping on the first picture according to the quality processing data in the spliced sub-data to obtain a second picture; The second picture is used as the target picture, or a specific object area is identified on the second picture. If a specific object area is identified, dynamic interpolation compensation is applied to the specific object area according to the quality processing data to obtain the target picture.
6. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 5 is characterized in that: The acquiring of the live broadcast signal comprises: Obtaining a switching signal, and obtaining the live broadcast signal in response to the switching signal, wherein the switching signal is a signal input by a user through a client or the multi-channel input director station; The step of using the second picture as the target picture, or identifying a specific object area on the second picture within a preset time period after the switching signal is generated, and applying dynamic interpolation compensation to the specific object area based on the quality processing data to obtain the target picture if the specific object area is identified, includes: performing a smooth transition process on the second picture according to the designated picture to obtain the target picture, or performing a specific object region identification on the second picture; if the specific object region is identified, applying dynamic interpolation compensation and a smooth transition process to the specific object region according to the quality processing data and the designated picture to obtain the target picture; otherwise, performing a smooth transition process on the second picture according to the designated picture to obtain the target picture; The designated picture is the second picture that is closest to the generation time of the switching signal in history.
7. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 4 is characterized in that: The step of performing splicing of the picture data with the same timestamp on each of the first data segments in the first storage area according to the picture splicing mode and the splicing configuration of the splicing sub-data of the target control data to obtain the first picture includes: performing, according to the picture splicing mode and the splicing configuration, a missing type detection of picture data with the same timestamp on each of the first data segments of the video type in the first storage area to obtain a first result; If the first result is incomplete, splicing the first data segments in the first storage area with the same timestamp and the incomplete first result according to the splicing configuration of the splicing sub-data of the target control data based on an exception handling strategy to obtain the first picture, wherein the exception handling strategy is data determined based on the first result and the picture splicing mode; If the first result is complete, the picture data with the same timestamp are spliced for each of the first data segments in the first storage area according to the picture splicing mode and the splicing configuration to obtain the first picture.
8. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 3 is characterized in that: The method further comprises: Based on the resource allocation data corresponding to the collection pool, the timestamp difference and audio delay parameters of the input data of the multi-channel input director station are monitored in real time to generate a synchronization offset set; Based on the resource allocation data corresponding to the acquisition pool, based on the adaptive clock synchronization technology and the dynamic interpolation technology, and according to the synchronization offset set, the first data segment of the video type in the first storage area is dynamically aligned.
9. The multi-channel video live broadcast processing method based on dynamic resource allocation according to claim 8, characterized in that: The step of dynamically aligning the first data segment of the video type in the first storage area according to the synchronization offset set based on the adaptive clock synchronization technology and the dynamic frame insertion technology includes: Performing resolution detection on each of the first data segments in the first storage area that corresponds to the same timestamp and is of video type, using the resolution as input resolution; When the input resolution is lower than a first threshold, a nearest neighbor interpolation mode is used to determine an offset of each data segment on the time axis according to the synchronization offset set, and first data segments of the video type in the first storage area are preliminarily aligned by inserting or deleting frames at corresponding positions according to the offset; When the input resolution is between a first threshold and a second threshold, a bilinear interpolation mode and an adaptive clock synchronization technology are used to dynamically align the first data segment of the video type in the first storage area according to the synchronization offset set; When the input resolution is higher than a second threshold, a bicubic interpolation mode and an adaptive clock synchronization technology are used to dynamically align the first data segment of the video type in the first storage area according to the synchronization offset set.
10. A multi-channel video live broadcast processing system based on dynamic resource allocation, characterized in that: The system includes: a multi-channel input director station and a client, the client is communicatively connected to the multi-channel input director station, and the multi-channel input director station is configured to implement the multi-channel video live broadcast processing method based on dynamic resource allocation as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method for realizing new live broadcasting system
CN107888953A
Live broadcast control method and device
CN112866725A
Video processing method and device, electronic equipment and storage medium
CN114071116A
Multimodal-based director method and system and computer program product
CN116152711A
Method for providing at least one image dataset, storage medium, computer program product, data server, imaging device and telemedicine system
EP3799061A1