Method and system for multi-path video live processing based on dynamic resource allocation

CN120640029BActive Publication Date: 2026-09-22GUANGZHOU MEILU ELECTRONICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510738921.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2026-09-22
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

[0003]基于此,有必要针对现有技术的多路视频直播方案,存在操作繁琐、硬件依赖性强与移动性受限的技术问题,提出了一种基于动态资源分配的多路视频直播处理方法

Benefits of technology

[0006]本申请的基于动态资源分配的多路视频直播处理方法及系统,用户通过客户端或者所述多路输入导播台输入画面拼接模式,不再需要用户通过软件界面手动切换多路输入源的画面布局或执行简单拼接,从而解决了操作繁琐的问题,大大提高了操作效率且降低了出错概率;本申请的技术方案用于控制多路输入导播台,不再依赖高性能计算机端运行导播软件,摆脱了设备体积庞大、功耗高的困扰,而且多路输入导播台便于携带,从而能够很好地适配移动拍摄场景;通过面向多路输入导播台的动态资源分配策略与模块化资源池架构(也就是采集池、处理池和编码池形成的三级资源池),实现了多路输入数据的智能化导播处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640029B_ABST
    Figure CN120640029B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video director technology, and discloses a multi-path video live broadcast processing method and system based on dynamic resource allocation. The method is used for controlling a multi-path input director station. The method comprises the following steps: acquiring a live broadcast signal; in response to the live broadcast signal, acquiring a picture splicing mode, wherein the picture splicing mode is data input by a user through a client or the multi-path input director station; based on a dynamic resource allocation strategy, a collection pool, a processing pool and an encoding pool of the multi-path input director station, data collection, data splicing and encoding output are performed according to the picture splicing mode. Thus, the problem of complicated operation is solved, the operation efficiency is greatly improved, and the error probability is reduced. The method no longer depends on a high-performance computer terminal to run director software, and is free from the troubles of large equipment volume and high power consumption. Moreover, the multi-path input director station is convenient to carry, so that the multi-path input director station can be well adapted to mobile shooting scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video broadcasting technology, and in particular to a method and system for processing multi-channel live video broadcasts based on dynamic resource allocation. Background Technology

[0002] As online live streaming develops towards specialization and mobility, multi-channel video live streaming has become a core requirement for live streaming scenarios such as large-scale events, sports events, and outdoor live streaming. In existing technologies, multi-channel video live streaming mainly relies on computer-based broadcasting software. Typical solutions include manually switching the screen layout of multiple input sources or performing simple splicing through the software interface. However, such traditional solutions have the following systemic defects: (1) Cumbersome operation: Users need to manually adjust the input priority and screen layout of multiple videos, especially in complex splicing modes (such as multi-screen splitting), which is inefficient and prone to errors; (2) Strong hardware dependence and limited mobility: Existing solutions rely on high-performance computers to run broadcasting software, resulting in bulky equipment, high power consumption, and inability to adapt to mobile shooting scenarios. Summary of the Invention

[0003] Based on this, it is necessary to address the technical problems of existing multi-channel video live streaming solutions, such as cumbersome operation, strong hardware dependence, and limited mobility, and propose a multi-channel video live streaming processing method based on dynamic resource allocation.

[0004] Firstly, a method for processing multi-channel live video streaming based on dynamic resource allocation is provided, the method comprising: the method for controlling a multi-channel input broadcast control station, the method comprising: Obtain the live broadcast signal; In response to the live broadcast signal, the video splicing mode is obtained, wherein the video splicing mode is data input by the user through the client or the multi-input control panel; Based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool of the multi-input broadcast control station, data acquisition, data splicing, and encoding output are performed according to the aforementioned screen splicing mode.

[0005] Secondly, a multi-channel video live streaming processing system based on dynamic resource allocation is provided. The system includes a multi-channel input broadcast control station and a client. The client is communicatively connected to the multi-channel input broadcast control station, and the multi-channel input broadcast control station is configured to implement the multi-channel video live streaming processing method based on dynamic resource allocation described in the first aspect.

[0006] This application presents a multi-channel video live streaming processing method and system based on dynamic resource allocation. Users can input the video splicing mode through a client or the multi-channel input control console, eliminating the need for manual switching of the multi-channel input source layout or simple splicing via software interface. This solves the problem of cumbersome operation, greatly improves operational efficiency, and reduces the probability of errors. The technical solution of this application controls the multi-channel input control console, eliminating reliance on high-performance computer-based control software, thus avoiding the problems of bulky equipment and high power consumption. Furthermore, the multi-channel input control console is portable, making it well-suited for mobile shooting scenarios. Through a dynamic resource allocation strategy and modular resource pool architecture (i.e., a three-level resource pool consisting of an acquisition pool, a processing pool, and an encoding pool) for the multi-channel input control console, intelligent control processing of multi-channel input data is achieved. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] in: Figure 1 This is an application environment diagram of a multi-channel video live streaming processing method based on dynamic resource allocation in one embodiment. Figure 2 This is a flowchart of a multi-channel live video processing method based on dynamic resource allocation in one embodiment; Figure 3 This is a schematic diagram of the structure of the multi-input broadcast console of this application; Figure 4 This is a block diagram of a multi-channel video live streaming processing system based on dynamic resource allocation in one embodiment. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0010] The multi-channel video live streaming processing method based on dynamic resource allocation provided in this invention can be applied to applications such as... Figure 1The application environment includes at least two acquisition devices 3, a multi-input broadcast control station 1, and a client 2. Client 2 communicates with multi-input broadcast control station 1 via wired and / or wireless communication technologies. Acquisition devices 3 communicate with multi-input broadcast control station 1 via wired and / or wireless communication technologies.

[0011] Optionally, the multi-input broadcast control station 1 is configured to implement the multi-channel video live streaming processing method based on dynamic resource allocation of this application, specifically including: acquiring a live signal; responding to the live signal to acquire a picture splicing mode, wherein the picture splicing mode is data input by the user through the client 2 or the multi-input broadcast control station 1; and based on the dynamic resource allocation strategy, acquisition pool, processing pool and encoding pool of the multi-input broadcast control station 1, performing data acquisition, data splicing and encoding output according to the picture splicing mode.

[0012] Optionally, client 2 is configured to implement the multi-channel video live streaming processing method based on dynamic resource allocation of this application, specifically including: acquiring a live signal; responding to the live signal to acquire a screen splicing mode, wherein the screen splicing mode is data input by the user through client 2 or the multi-channel input control console 1; and controlling the multi-channel input control console 1 to perform data acquisition, data splicing, and encoding output according to the screen splicing mode based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool for the multi-channel input control console 1.

[0013] Users can input the image splicing mode through client 2 or the multi-input broadcast console 1, eliminating the need for manual switching of the image layout of multiple input sources or simple splicing through the software interface. This solves the problem of cumbersome operation, greatly improves operational efficiency, and reduces the probability of errors. The technical solution of this application is used to control the multi-input broadcast console 1, eliminating the need to rely on a high-performance computer to run the broadcast software, thus avoiding the problems of large device size and high power consumption. Moreover, the multi-input broadcast console 1 is easy to carry, making it well-suited for mobile shooting scenarios. Through the dynamic resource allocation strategy and modular resource pool architecture (i.e., a three-level resource pool consisting of an acquisition pool, a processing pool, and an encoding pool) for the multi-input broadcast console 1, intelligent broadcast processing of multi-input data is achieved.

[0014] Client 2 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.

[0015] The multi-input broadcast control station 1 includes: a communication component, a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is electrically connected to the communication component, and when the processor executes the computer program, it implements the steps of the multi-channel video live streaming processing method based on dynamic resource allocation.

[0016] The communication components include: a wireless communication module and / or a wired communication module. The wireless communication module communicates wirelessly with devices other than the multi-input broadcast control station 1 via wireless communication technology. The wired communication module communicates wiredly with devices other than the multi-input broadcast control station 1 via wired communication technology. The wireless communication module can be selected from existing technologies, and will not be elaborated upon here. The wired communication module can also be selected from existing technologies, and will not be elaborated upon here.

[0017] Please see Figure 3 Optionally, the first side 19 of the multi-input broadcast control station 1 is provided with a first high-definition input interface 11, a second high-definition input interface 12, a high-definition output interface 13, and a TYPE-C (USB interface standard) output interface 14, and the second side 18 of the multi-input broadcast control station 1 is provided with a microphone input interface 15. The multi-input broadcast control station 1 acquires high-definition or standard-definition video data input from the camera 31 (i.e., acquisition device 3) through the first high-definition input interface 11 via a wired communication module. The multi-input broadcast control station 1 acquires high-definition or standard-definition video data input from the camera 32 (i.e., acquisition device 3) through the second high-definition input interface 12 via a wired communication module. The multi-input broadcast control station 1 outputs video data to the display device 41 (i.e., receiving device), such as a monitor or laptop computer, through the high-definition output interface via the wired communication module. The multi-input broadcast control station 1 outputs video data to the desktop computer 42 (i.e., receiving device) through the TYPE-C output interface 14 via the wired communication module. The multi-input broadcast console 1 acquires audio data input from the microphone (i.e., the acquisition device 3) through the microphone interface 15.

[0018] Optionally, the multi-input broadcast control station 1 supports YUY2 and MJPEG formats for video capture. The maximum resolution of the video captured by the multi-input broadcast control station 1 is 1080P, and the refresh rate is 60Hz. The multi-input broadcast control station 1 supports high-definition and 1080P loop-out (i.e., encoded video output). The audio sampling rate of the multi-input broadcast control station 1 is 48kHz. YUY2 is a common video data format, belonging to a sampling format of the YUV (luminance-chrominance) color space. MJPEG (Motion-JPEG) is a video encoding format.

[0019] It is understood that the multi-input broadcast console 1 is equipped with one or more buttons, and users input signals (such as live signals, switching signals, etc.) to the processor by pressing the buttons.

[0020] Optionally, the multi-input broadcast console 1 has a separate button for each video splicing mode, eliminating the need for multiple modes to share a single button, thus simplifying user operation.

[0021] Optionally, each button on the multi-input broadcast control unit 1 corresponds to at least one signal. By multiplexing the buttons, the number of buttons is reduced, thus lowering the cost of the multi-input broadcast control unit 1.

[0022] Optionally, the multi-input broadcast console 1 is equipped with a touch screen, through which users can input signals to the processor and select the video splicing mode.

[0023] It is understood that the data collected by the multi-input control console 1 includes at least two video data streams and at least one audio data stream.

[0024] The present invention will now be described in detail through specific embodiments.

[0025] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a multi-channel video live streaming processing method based on dynamic resource allocation provided in an embodiment of the present invention is shown. The method is used to control a multi-channel input broadcast control station, and includes the following steps: S1: Obtain the live broadcast signal; The live broadcast signal is the signal to start a live broadcast.

[0026] Specifically, it can be data input by the user through the client or the multi-input switcher, or it can be a live signal that is automatically triggered after the multi-input switcher is powered on and initialized.

[0027] S2: In response to the live broadcast signal, obtain the video splicing mode, wherein the video splicing mode is data input by the user through the client or the multi-input control panel; Specifically, when the live broadcast signal is obtained, the video splicing mode input by the user in real time through the client or the multi-input broadcast control station can be obtained, or the video splicing mode stored by the user in advance can be obtained.

[0028] The video stitching mode includes image frame stitching description data and audio stitching description data. Image frame stitching description data describes the stitching of image frames from multiple video streams. Audio stitching description data describes the combination of multiple audio streams. The image frame stitching description data records the stitching method and the video stitching ratio.

[0029] For example, when the video input to the multi-channel input broadcasting station is two channels, the images of the two video channels (that is, the image frames in the video) are spliced ​​in the following ways: single output, equal left and right scale, unequal left and right scale, equal top and bottom scale, unequal top and bottom scale, picture-in-picture, etc.

[0030] Audio splicing description data: highlight one audio stream and mute the other, or mute all audio streams, or allow both audio streams to coexist.

[0031] S3: Based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool of the multi-input broadcasting station, data acquisition, data splicing, and encoding output are performed according to the aforementioned screen splicing mode.

[0032] Specifically, based on the dynamic resource allocation strategy for multi-input broadcast control stations, resource allocation data (e.g., thread resources) corresponding to the acquisition pool, processing pool, and encoding pool are allocated according to the actual description data of the multi-input broadcast control station (e.g., one or more of acquisition performance data, output performance data, and output mode data). Based on the target control data, the acquisition pool is controlled to acquire data according to the resource allocation data corresponding to the acquisition pool; based on the target control data, the processing pool is controlled to perform data splicing according to the resource allocation data corresponding to the processing pool; and based on the target control data, the encoding pool is controlled to perform encoding output according to the resource allocation data corresponding to the encoding pool.

[0033] Dynamic resource allocation strategy is an algorithm that adjusts the allocation of system resources (CPU, memory, bandwidth) in real time. Based on the current input load (such as the number of video channels and resolution), output requirements (such as encoding complexity) and user configuration (such as latency threshold), it dynamically optimizes the proportion of resources in the three stages of acquisition, processing and encoding to ensure efficient and low-latency broadcast processing.

[0034] Optionally, the dynamic resource allocation strategy employs a weighted priority algorithm. Based on the input load (number of video streams, resolution, frame rate, and audio parameters) and output requirements (encoding complexity, latency threshold), the resource allocation weights for the acquisition pool, processing pool, and encoding pool are calculated in real time. The resource allocation weight formula is as follows: Resource allocation weight = α × Input load + β × Output demand + γ × Delay threshold Here, α, β, and γ are adjustable coefficients. The specific values ​​of α, β, and γ can be determined by fitting multiple sets of test data.

[0035] The acquisition pool is a resource pool responsible for the acquisition and temporary storage of multiple video / audio streams. Its core functions include data segmentation, timestamp alignment, and format conversion (such as YUY2 to RGB), and it uses queue management to achieve data buffering and asynchronous processing.

[0036] The processing pool is a resource pool responsible for image stitching, image quality optimization, and audio-visual synchronization. By calling image processing algorithms (such as interpolation and color correction) and audio processing modules (such as mixing and noise reduction), the processing pool integrates multiple input data into a single output stream (i.e., audio and video data) that conforms to the user's stitching mode.

[0037] The encoding pool compresses and encodes the processed audio and video data according to the target format (such as H.264) and pushes it to the resource pool of the output device. The encoding pool supports dynamic bitrate adjustment, multi-protocol streaming (such as RTMP, SRT), and error recovery mechanisms.

[0038] Data acquisition involves processing the data from multiple input channels to the broadcast control station according to preset processing requirements (such as format conversion and segmentation) and then caching it according to a preset caching method (such as caching data segments corresponding to the same input channel in the same queue).

[0039] Data splicing is the process of combining images and audio with data collected to form new audio segments, and then caching these new audio segments.

[0040] Encoding output involves encoding a new audio segment according to the required encoding format, and then transmitting the encoded data to a specified device (such as a laptop) or a specified live streaming application.

[0041] Users can input the video splicing mode through the client or the multi-input broadcast console, eliminating the need for manual switching of the video layout of multiple input sources or simple splicing through the software interface. This solves the problem of cumbersome operation, greatly improves operational efficiency, and reduces the probability of errors. The technical solution of this application is used to control the multi-input broadcast console, eliminating the need to rely on high-performance computer software to run the broadcast software, thus avoiding the problems of bulky equipment and high power consumption. Moreover, the multi-input broadcast console is portable, making it well-suited for mobile shooting scenarios. Through the dynamic resource allocation strategy and modular resource pool architecture (i.e., a three-level resource pool consisting of an acquisition pool, a processing pool, and an encoding pool) for the multi-input broadcast console, intelligent broadcast processing of multi-input data is achieved.

[0042] In one embodiment, the steps of data acquisition, data splicing, and encoding output based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool of a multi-input broadcast control station, according to the video splicing mode, include: S31: Acquire target control data and resource allocation data, wherein the target control data includes: acquisition sub-data, image processing sub-data and encoding output sub-data; S32: Perform data acquisition, data splicing, and encoding output based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the screen splicing mode; The target control data is determined through the following steps: acquiring an analysis signal, responding to the analysis signal, acquiring acquisition performance data, output performance data, output mode data, and delay configuration data, and determining the target control data based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data, wherein the delay configuration data is data input by the user through a client or the multi-input broadcast control console; The resource allocation data is determined through the following steps: in response to the analysis signal, the resource allocation data corresponding to the acquisition pool, the processing pool, and the encoding pool are determined according to the dynamic resource allocation strategy based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data.

[0043] Specifically, the system monitors user actions or system status changes (such as live streaming startup, mode switching, etc.) in real time via hardware interfaces or software modules, triggering analysis signals. These analysis signals may originate from user key presses, client commands, or timed tasks of the program implementing this application.

[0044] The performance data collected includes the hardware capability parameters of the multi-input broadcast console for collecting input data (such as the maximum number of supported input channels, resolution, frame rate, encoding format, etc.).

[0045] Output performance data describes the performance of external output data from a multi-input broadcast control station to another multi-input broadcast control station, such as parameters like the resolution, bit rate, and network transmission latency of the encoded output.

[0046] Output mode data: describes the configuration requirements of the output data of the multi-input broadcasting station, including: (1) Output protocol and transmission method: the output protocol (such as RTMP, SRT, HLS) and transmission method (real-time stream, timed push) set according to the receiving device or application type; (2) Output quality parameters: user-defined or pre-stored target resolution (such as 1080P), bit rate limit (such as 8Mbps), frame rate (such as 60fps), audio encoding format (such as AAC); (3) Terminal adaptation rules: dynamically adjust the encoding parameters (such as resolution downgrading, bit rate adaptation) for different receiving devices (such as mobile terminal, PC terminal).

[0047] The delay configuration data refers to the minimum and maximum allowable delay thresholds set by the user through the client or the multi-input broadcast control console, used to dynamically adjust the data processing rhythm. Delay refers to the time interval between data acquisition and output encoding.

[0048] Optionally, the target control data can be determined using a lookup table method based on the acquired performance data, the output performance data, the output mode data, and the delay configuration data.

[0049] Optionally, regular expressions are used to calculate the index value of each first index based on the acquired performance data, the output performance data, the output mode data, and the delay configuration data, and the target control data is determined by a lookup table method based on each index value.

[0050] The selection range for the first indicator includes, but is not limited to: input load indicators, processing capacity indicators, output demand indicators, and latency tolerance indicators.

[0051] Input load metrics are calculated based on the acquired performance data. Input load metrics include: (1) Number of input channels: the number of matched input interfaces; (2) Average resolution: the value extracted from the resolution of the video of the input multi-channel input broadcast station; (3) Total frame rate: the sum of the frame rates of the video of all multi-channel input broadcast stations; (4) Peak bitrate: the maximum value extracted from the input bitrate of the data of the input multi-channel input broadcast station.

[0052] Processing capacity metrics are calculated based on collected performance data and output performance data. Processing capacity metrics include: (1) CPU utilization: real-time value obtained through system monitoring interface; (2) memory usage: matching the percentage of memory used; (3) current processing latency: calculated based on timestamp differences.

[0053] Output demand metrics are calculated based on output performance data and output mode data. Output demand metrics include: (1) target bitrate: the upper limit of bitrate extracted from output mode data; (2) encoding complexity: graded according to encoding format (e.g., H.264=1, H.265=2); (3) audio synchronization level: extracted from audio-visual combination data.

[0054] The latency tolerance index is an index calculated based on latency configuration data. The latency tolerance index includes: (1) maximum allowable latency; (2) real-time priority: mapped to a level according to the latency threshold set by the user, for example, latency <100ms = high, latency 100-200ms = medium, latency >200ms = low.

[0055] Optionally, the acquired performance data, the output performance data, the output mode data, and the delay configuration data are input into a pre-trained first model for classification prediction. The vector element with the largest value is selected from the predicted vectors, and the control data corresponding to the classification category of the selected vector element is used as the target control data.

[0056] The first model is a pre-trained multi-class classification model. The model structure and training method for the first model can be selected from existing technologies.

[0057] The collected sub-data includes, but is not limited to, dynamically calculating the segment interval based on the resolution, frame rate, and bit rate of the input video, combined with latency configuration data.

[0058] Image processing sub-data includes, but is not limited to, defining splicing modes (such as left and right split-screen ratios) and image quality processing parameters (such as color correction matrices and interpolation algorithms).

[0059] The encoded output sub-data includes, but is not limited to, setting the encoding format, bitrate, and audio synchronization strategy.

[0060] Based on the dynamic resource allocation strategy, system resources (such as CPU threads, memory, and bandwidth) are allocated to the three-level resource pools according to priority to determine the resource allocation data of each resource pool (i.e., the acquisition pool, processing pool, and encoding pool): (1) Acquisition pool: resources are allocated for real-time acquisition and segmented storage of multiple video streams; (2) Processing pool: resources are allocated for image splicing, image quality optimization, and audio-visual synchronization; (3) Encoding pool: resources are allocated for final encoding and output stream push.

[0061] Specifically, based on the acquisition sub-data, the acquisition pool is controlled to acquire data according to the resource allocation data corresponding to the acquisition pool; based on the image processing sub-data, the processing pool is controlled to stitch data according to the resource allocation data corresponding to the processing pool; and based on the encoding output sub-data, the encoding pool is controlled to encode and output data according to the resource allocation data corresponding to the encoding pool.

[0062] Specifically, the acquisition pool is controlled by the resource allocation data corresponding to the acquisition pool. Each input video is divided into timestamp-aligned data segments (e.g., 50ms segments) according to the segmentation interval of the acquisition sub-data. The segmentation granularity is dynamically adjusted according to the latency configuration data (e.g., increasing the segment length and reducing the processing frequency when high latency tolerance is required). The segmented data is stored in the first storage area of ​​the acquisition pool and managed by an independent queue according to the number of input channels.

[0063] By using the resource allocation data corresponding to the processing pool, the processing pool is controlled to extract video segments with the same timestamp from the first storage area, perform pixel-level splicing according to the screen splicing mode (such as left and right split screen), then perform image quality normalization processing, and finally perform audio and video synchronization processing.

[0064] The image quality normalization process specifically includes: (1) Color correction: dynamically matching the color space of each video segment (e.g., sRGB to BT.709); (2) Brightness compensation: interpolating the edge area of ​​the image to eliminate the brightness difference at the splicing point; (3) Resolution mapping: scaling the input of different resolutions to the target resolution (e.g., 1080P); (4) Audio-visual synchronization processing: selecting the main audio stream or mixed output according to the audio-visual combination data, and strictly aligning it with the image timestamp.

[0065] Using the resource allocation data corresponding to the encoding pool, the encoding pool is controlled to compress the spliced ​​video and audio into encoded output sub-data (such as H.264 encoding, with a bitrate of 8Mbps). Adaptive bitrate control technology is used to dynamically adjust the output bitrate according to the network bandwidth and push it to a designated device (such as a live streaming server or local display screen) or store it as a file.

[0066] This embodiment determines target control data and resource allocation data based on acquisition performance data, output performance data, output mode data, and delay configuration data. Then, based on this data, the acquisition pool, processing pool, encoding pool, and image stitching mode, data acquisition, stitching, and encoding output are performed. This enables efficient and flexible processing of multiple data streams from a multi-input broadcast control station. This method can adaptively adjust to various factors such as different equipment performance, output requirements, and user-set delays, ensuring the accuracy of data acquisition, the rationality of data stitching, and the effectiveness of encoded output. This improves the adaptability and operational efficiency of the entire broadcast control process in different application scenarios, and reduces performance degradation or data processing errors caused by unreasonable resource allocation or mismatched settings.

[0067] In one embodiment, the steps of data acquisition, data stitching, and encoding output based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the screen stitching mode include: S3211: Based on the resource allocation data corresponding to the acquisition pool, according to the acquisition sub-data of the target control data, each input data of the multi-channel input broadcasting station is segmented using the same segmentation timestamp to generate a data segment that conforms to the storage structure of the first storage area of ​​the acquisition pool, which is used as the first data segment, and each first data segment is stored in the first storage area. The dynamic determination of the acquired sub-data includes: calculating the segmented time interval based on the resolution, frame rate, and bitrate of the input data (which is video) as the initial interval; adjusting the initial interval based on the maximum allowable delay threshold corresponding to the delay configuration data as the target interval; determining the segmentation granularity based on the acquisition performance data and the delay configuration data; and using the target interval and the segmentation granularity as the acquired sub-data. By calculating the target interval and segmentation granularity in real time, a balance between latency and computational overhead is achieved.

[0068] The first storage area is divided into independent queues according to the number of input paths. Each queue stores the data segments obtained from the segmented processing of the corresponding input data.

[0069] Segmentation processing, based on a segmentation method that strictly aligns the segment timestamps of multiple input data, divides the continuously input video / audio stream into discrete data segments at fixed time intervals, facilitating subsequent splicing and synchronization.

[0070] The granularity of the segments, the smallest time unit for each segment (e.g., 10ms), determines the frequency and precision of data processing. Low-latency scenarios require small segment granularity to achieve high-frequency processing, while high-throughput scenarios can increase the segment granularity.

[0071] Wherein, the initial interval = frame rate × resolution weight / (bit rate × complexity coefficient), the complexity coefficient is determined by a lookup table method according to the encoding format, and the resolution weight is the data determined by a lookup table method based on the average resolution of the input data of type video.

[0072] The initial interval is limited by the maximum delay threshold. For example, if the initial interval is 9ms but the maximum allowable delay is 200ms, the target interval is adjusted to min(9ms, 200ms / N), where N is the number of input channels of the multi-input broadcast console.

[0073] Based on the acquired performance data and the latency configuration data, the segment granularity is determined using a lookup table method.

[0074] It is understood that the collected sub-data includes: target interval and segmentation granularity.

[0075] Optionally, the collected sub-data may further include: queue sharding. Queue sharding is pre-defined data, with each input data stream corresponding to a fixed queue shard.

[0076] Specifically, the timestamps of the multiple input data of the multi-channel input broadcast console are synchronized using a hardware clock or the NTP (Network Time Protocol) to ensure the alignment of all input data. At the end of each interval of the target interval for collecting sub-data, each input data of the multi-channel input broadcast console is segmented according to the segmentation granularity of the collected sub-data using the same segmentation timestamp, resulting in data segments with a length equal to the segmentation granularity (or smaller segments if the input data is insufficient). If the format of the data segment conforms to the storage structure requirements of the first storage area of ​​the acquisition pool, the data segment is used as the first data segment. If the format of the data segment does not conform to the storage structure requirements of the first storage area of ​​the acquisition pool, the data segment is converted into a format and then used as the first data segment. The first data segment is stored in the queue fragment corresponding to the input data of the first data segment.

[0077] It is understandable that the first storage area stores the first data segment in timestamp order.

[0078] Queue sharding achieves resource isolation, and the independent queue design avoids data contention between multiple inputs, thus improving stability.

[0079] This embodiment calculates the initial interval based on the resolution, frame rate, and bitrate of the input data (which is video), then adjusts it to the target interval using latency configuration data. Simultaneously, it considers acquisition performance data to determine the segmentation granularity, thereby generating acquired sub-data. These sub-data are then segmented using the same timestamp. This approach ensures that the generated data segments conform to the storage structure of the first storage area of ​​the acquisition pool, facilitating efficient storage of the first data segment and optimizing data acquisition and storage. This refined processing based on multiple data types better adapts to different input data characteristics and system requirements, improving the rationality and effectiveness of data acquisition and storage, and reducing confusion and errors during data processing.

[0080] In one embodiment, the steps of data acquisition, data stitching, and encoding output based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the screen stitching mode further include: S3221: Based on the resource allocation data corresponding to the processing pool, according to the screen splicing mode and the splicing configuration of the splicing sub-data of the target control data, the screen data with the same timestamp of each first data segment in the first storage area are spliced ​​to obtain the first screen; The splicing configuration describes the geometric alignment rules (such as coordinate offset and scaling ratio) of multiple video streams.

[0081] Specifically, all first data segments with the same timestamp (e.g., 0-50ms) are extracted from the first storage area. The position and proportion of each video are determined according to the image stitching mode. The resolution of the image frames of each input video is adjusted according to the stitching configuration. According to the stitching configuration, the multiple image frames (i.e., image data) with adjusted resolutions are stitched to the same canvas according to the coordinate offset. The image on this canvas is used as the first image. The first image is stored in the temporary buffer corresponding to the processing pool for subsequent image quality normalization processing.

[0082] S3222: Based on the resource allocation data corresponding to the processing pool, and according to the quality processing data in the splicing sub-data, the first image is subjected to image quality normalization processing to obtain the target image; Quality processing data describes a set of parameters used to improve the visual quality of stitched images, including color correction matrices, brightness compensation algorithms, resolution mapping rules, etc.

[0083] Image quality normalization is a process that adjusts the color, brightness, and resolution of the first image frame to a consistent standard, eliminating visual inconsistencies caused by differences in devices.

[0084] The first frame may have issues such as inconsistent colors, brightness differences, and resolution mismatch. To address this, firstly, based on the input video's color space (e.g., sRGB, BT.709), a color correction matrix from the quality processing data is applied to unify each image frame in the first frame to the target color standard (e.g., BT.709) for color correction. Then, based on the brightness compensation algorithm in the quality processing data, the stitching area of ​​the first frame after color correction (e.g., the seam between left and right split screens) is detected, and the brightness differences between the left and right frames are interpolated and smoothed to achieve brightness compensation. Finally, according to the resolution mapping rules in the quality processing data, low-resolution input frames are interpolated and enlarged to the target resolution, or high-resolution frames are downsampled to the target resolution, thus unifying to the target resolution (i.e., scaling the frame). An adaptive sharpening filter is applied to the scaled frame to improve detail clarity, and the frame processed by the adaptive sharpening filter is used as the target frame.

[0085] S3223: Based on the resource allocation data corresponding to the processing pool, according to each target screen and each first data segment of type audio in the first storage area, the screen and audio are matched according to the audio-visual combination data in the splicing sub-data to obtain the second data segment, and the second data segment is stored in the second storage area of ​​the processing pool.

[0086] Audio-visual composite data describes the synchronization rules between audio and video, such as main audio stream selection, audio delay compensation, and mixing strategies.

[0087] The second data segment is a complete data unit after audio and video synchronization, containing the target image and the corresponding audio stream.

[0088] The second storage area is a buffer used to temporarily store the spliced ​​and optimized second data segments. The data segments in the second storage area are arranged in chronological order and can be retrieved by the encoding pool as needed.

[0089] Specifically, firstly, audio is extracted and processed based on the audio-visual composite data. This involves extracting the audio segment corresponding to the timestamp of the target image (i.e., the first data segment of type audio) from the first storage area. If the audio-visual composite data is the main audio stream (e.g., input channel 1), other audio streams are masked; if the audio-visual composite data describes a mix, multiple audio streams are mixed according to the proportions in the audio-visual composite data. Then, each target image and the processed audio stream are encapsulated into a composite data packet (i.e., the second data segment) according to its timestamp. For example, this encapsulation might use MPEG-TS (a video encapsulation format based on the MPEG-2 standard), with the video track in H.264 (one of the ITU-T video codec standards named after the H.26x series) and the audio track in AAC (Advanced Audio Coding). Finally, the second data segment is pushed into the second storage area in chronological order.

[0090] This embodiment first stitches the first data segment with the same timestamp according to the image stitching mode and the stitching configuration of the stitching sub-data in the target control data to obtain the first image, achieving preliminary image integration. Next, the first image is normalized according to the quality processing data to obtain the target image, which helps improve the overall image quality and make it more compliant with requirements. Finally, the target image and audio are mapped according to the audio-visual combination data to obtain the second data segment and stored in the second storage area of ​​the processing pool. This mapping storage of image and audio not only makes the data more organized in subsequent processing but also optimizes the overall data flow from acquisition to preliminary integration processing, improving the accuracy and orderliness of data processing. This better meets the needs of image and audio processing in various application scenarios and improves the overall data processing efficiency of the system.

[0091] In one embodiment, the step of performing image quality normalization processing on the first image based on the quality processing data in the stitched sub-data to obtain the target image includes: S32221: Based on the quality processing data in the splicing sub-data, perform dynamic correction of the color correction matrix, brightness cross-region interpolation compensation, and resolution adaptive mapping on the first image to obtain the second image; Specifically, firstly, based on the color space of the input video (such as sRGB, BT.709), the color correction matrix in the quality processing data is applied to unify each image frame in the first frame to the target color standard (such as BT.709) to achieve color correction; then, based on the brightness compensation algorithm in the quality processing data, the splicing area of ​​the first frame after color correction (such as the middle seam between the left and right split screens) is detected, and the brightness difference between the left and right frames is interpolated and smoothed to achieve brightness compensation and eliminate the visual disjointness; finally, based on the resolution mapping rules in the quality processing data, the low-resolution input frame is interpolated and enlarged, or the high-resolution frame is downsampled and unified to the target resolution (that is, the scaled frame), and the first frame at this time is used as the second frame.

[0092] S32222: Use the second image as the target image, or perform specific object region recognition on the second image. If a specific object region is recognized, apply dynamic interpolation compensation to the specific object region according to the quality processing data to obtain the target image.

[0093] The specific object can be a human body or face, an animal, or a certain type of item.

[0094] The intermediate image (the second image), after color correction, brightness compensation, and resolution mapping, has eliminated basic image quality issues. However, it hasn't yet been optimized for specific object-specific image regions. To further improve image quality, firstly, computer vision algorithms (such as YOLO and OpenCV Haar cascade classifiers) are used to detect image regions corresponding to specific objects (referred to as object-specific regions). For the identified object-specific regions, based on the quality processing data, a higher-precision interpolation algorithm (such as bicubic interpolation) is applied to improve detail clarity and dynamic range. The final output is an optimized image (the target image). The optimized image has advantages such as consistent color, uniform brightness, and resolution adaptation, and the details of the image regions corresponding to specific objects are enhanced. High-computation processing is applied only to the image regions corresponding to specific objects, balancing performance and quality.

[0095] This embodiment obtains a second image by dynamically correcting the color correction matrix of the first image, performing cross-regional brightness interpolation compensation, and adaptive resolution mapping. This series of operations effectively improves the color accuracy, brightness uniformity, and resolution adaptability of the image, optimizing it in multiple dimensions. Furthermore, using the second image as the target image, or applying dynamic interpolation compensation to a specific object region based on quality processing data when that region is identified, allows for targeted processing of specific object regions while maintaining overall image optimization. This enhances the visual effect of those regions, providing better visual presentation when focusing on key areas. Both in terms of visual quality and the processing of specific regions, the image better meets processing requirements, improving overall image quality.

[0096] In one embodiment, acquiring the live stream signal includes: S11: Obtain a switching signal, and in response to the switching signal, obtain the live broadcast signal, wherein the switching signal is a signal input by the user through the client or the multi-input control panel; Switching signals are commands triggered by users through the client or multi-input control panel, used to start live broadcasts, switch screen layouts, etc.

[0097] The live broadcast signal is a control command generated after receiving the switching signal to start the live broadcast process, triggering the entire process of data acquisition, splicing, and encoding output.

[0098] Specifically, the multi-input switcher listens to the physical button signals (such as the "Start Live" button) of the multi-input switcher through the GPIO (General-purpose input / output) interface, and / or the client sends a switching signal to the multi-input switcher.

[0099] The live broadcast signal is obtained in response to the switching signal, and the live broadcast process is initialized.

[0100] Within a preset time period after the start of the generation time of the switching signal, the steps of using the second frame as the target frame, or performing specific object region identification on the second frame, and if a specific object region is identified, applying dynamic interpolation compensation to the specific object region based on the quality processing data to obtain the target frame, include: S322221: Based on the specified image, perform smooth transition processing on the second image to obtain the target image; or, perform specific object region recognition on the second image. If a specific object region is recognized, apply dynamic interpolation compensation and smooth transition processing to the specific object region based on the quality processing data and the specified image to obtain the target image; otherwise, perform smooth transition processing on the second image based on the specified image to obtain the target image. The designated screen is the second screen that is historically closest to the time when the switching signal was generated.

[0101] In a multi-input broadcast control system, a circular buffer is maintained to store the video data from the most recent N seconds (e.g., 5 seconds). The video can be retrieved quickly by timestamp, meaning that a specific video can be obtained from the circular buffer.

[0102] Understandably, if the circular buffer is empty, a black screen or static background will be used as the specified image by default.

[0103] Specifically, the method for smooth transition processing involves alpha blending two images. Alpha blending is an image processing technique used to achieve transparency and blending effects in images.

[0104] Alpha blending of the two images (the specified image and the second image) is performed within a time window (a preset duration after the generation time of the switching signal) to eliminate image jumps and improve viewing smoothness.

[0105] This embodiment first acquires the live broadcast signal by responding to a switching signal. This mechanism allows for flexible acquisition of live content based on signals input by the user through a client or multi-input control console, increasing operational flexibility and user controllability. Then, during image processing within a preset time after the switching signal generation time, whether it's smoothing the second image to obtain the target image based on a specified image (i.e., the second image historically closest to the generation time of the switching signal), or applying dynamic interpolation compensation and smooth transition processing to a specific object area based on quality processing data and a specified image when a specific object area is identified, these operations ensure the continuity and stability of the image during signal switching. Smooth transition processing avoids abruptness that may occur during image switching, while special processing for specific object areas enhances the visual effect of those areas while ensuring a natural overall transition, providing better visual presentation when focusing on key areas and improving the user's live broadcast viewing experience.

[0106] In one embodiment, the step of stitching together the image data with the same timestamp from each of the first data segments in the first storage area according to the image stitching mode and the stitching configuration of the target control data includes: S32211: Based on the image stitching mode and the stitching configuration, perform missing data type detection on each of the first data segments of type video in the first storage area to obtain a first result; The first result is the output of the missing type detection, which marks the data integrity status (such as "complete" or "missing 0-50ms data from input path 2").

[0107] Specifically, firstly, extract all first data segments with the same timestamp from the first storage area, and generate a list to be tested according to the number of input paths; traverse the list to be tested and check whether each input path contains a data segment with the current timestamp; if all input paths contain data, output {"status":"complete"}; if there are missing data, output {"status":"incomplete","missing path":missing_routes,"missing timestamp":current_ts}.

[0108] S32212: If the first result is incomplete, then based on the exception handling strategy, according to the splicing configuration of the splicing sub-data of the target control data, each of the first data segments in the first storage area is spliced ​​with the same timestamp and the first result is incomplete to obtain the first image. The exception handling strategy is data determined based on the first result and the image splicing mode. Anomaly handling strategies are predefined rules based on the type of missing data and the image stitching mode, used to repair or compensate for missing data. Anomaly handling strategies include, but are not limited to, frame interpolation and degradation processing. Frame interpolation: Fills the missing segment with data from the previous frame or a solid color. Degradation processing: Ignores missing paths and only stitches together usable input.

[0109] Specifically, if the first result is incomplete, for the timestamp corresponding to the first result, the first data segment of the next missing timestamp is extracted, and a motion compensation algorithm (such as optical flow) is applied to generate an approximate current frame, or a preset solid color is used as the approximate current frame; the approximate current frame is then added to the first data segment. Based on the splicing configuration of the splicing sub-data of the target control data, each first data segment (compensated data segment and data segment not requiring compensation) is geometrically spliced ​​according to the splicing configuration to generate the first image. Through missing data detection and dynamic compensation, it is ensured that the live stream can continue to output even when some input is abnormal.

[0110] S32213: If the first result is complete, then according to the image splicing mode and the splicing configuration, the image data with the same timestamp of each of the first data segments in the first storage area are spliced ​​together to obtain the first image.

[0111] Specifically, if the first result is complete, for the timestamp corresponding to the first result, there is no need to generate an approximate current frame. Instead, based on the image stitching mode and the stitching configuration, the image data with the same timestamp of each of the first data segments in the first storage area are stitched together to obtain the first image.

[0112] This may result in gaps or misalignments in the image due to missing data, requiring further optimization. To address this issue, this embodiment first performs a missing data type detection for image data with the same timestamp, accurately assessing the integrity of the image data in the first data segment of the video type. When the detection result is incomplete, image stitching is performed based on an anomaly handling strategy determined according to the detection result and the image stitching mode. This approach flexibly handles incomplete data situations, ensuring that a complete image is still obtained as the first image even with missing data, avoiding gaps or misalignments in the entire image. Conversely, when the detection result is complete, image stitching is performed according to the image stitching mode and configuration. This efficiently and accurately stitches image data with the same timestamp when the data is complete, obtaining the first image, thus improving the overall accuracy, stability, and adaptability to different data states of image stitching.

[0113] In one embodiment, the method further includes: S41: Based on the resource allocation data corresponding to the acquisition pool, monitor the timestamp difference and audio delay parameters of the input data of the multi-channel input broadcasting station in real time, and generate a set of synchronization offsets; The synchronization offset set is a dataset that records the timestamp differences (e.g., stream A is 50ms slower than stream B) and audio-visual synchronization errors between multiple input video / audio streams. The synchronization offset set includes: video timestamp differences and audio delay parameters.

[0114] Video timestamp difference: The relative delay (in milliseconds) between multiple input video streams.

[0115] Audio delay parameter: the time difference between the audio and the corresponding video (e.g., audio lags behind video by 20ms).

[0116] Optionally, the system time of all acquisition devices can be synchronized via the NTP protocol or a hardware clock (such as PTP) to reduce initial clock skew.

[0117] Specifically, timestamps are extracted from each input data stream of type video, and the relative delay between the multiple input streams is calculated as the video timestamp difference. The timestamps of each input data stream of type audio (such as audio PTS) are extracted and compared with the timestamps of the corresponding input data of type video to obtain the audio delay parameter.

[0118] S42: Based on the resource allocation data corresponding to the acquisition pool, and based on adaptive clock synchronization technology and dynamic frame interpolation technology, dynamically align the first data segment of type video in the first storage area according to the synchronization offset set.

[0119] Adaptive clock synchronization technology is an algorithm that dynamically adjusts the clock frequency or timestamp of the input stream to eliminate time differences between multiple inputs. By analyzing the synchronization offset, it controls the data acquisition rate or frame processing (inserting / deleting redundant frames) to achieve timeline alignment.

[0120] Specifically, based on the resource allocation data corresponding to the acquisition pool, and using adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set. The timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with large delays (i.e., the analyzed offset is greater than the acquisition offset threshold, such as input path 2 being 50ms slower), their acquisition rate is increased (e.g., from 30fps to 31fps), gradually reducing the time difference. For the first data segment of type video in the first storage area, generated frames are inserted into the positively offset first data segment to catch up with other streams, and redundant frames are deleted from the negatively offset first data segment to slow down the progress.

[0121] Positive offset refers to a delay (lag) in the timestamp of a video relative to a reference timeline.

[0122] Negative offset refers to a video stream's timestamp being ahead of a reference timeline.

[0123] The reference timeline is based on the timestamp of the earliest arriving input data.

[0124] This embodiment generates a synchronization offset set by real-time monitoring of timestamp differences and audio delay parameters of input data from multiple input broadcast consoles based on resource allocation data corresponding to the acquisition pool. This accurately records timestamp differences and other information between multiple input video / audio streams. This synchronization offset set provides crucial information for subsequent processing, enabling precise synchronization of multiple data streams using frame interpolation or frame deletion techniques. Next, based on the resource allocation data corresponding to the acquisition pool, adaptive clock synchronization and dynamic frame interpolation techniques are employed to dynamically align the first data segment of type video in the first storage area according to the synchronization offset set. This effectively solves the problem of time asynchrony among multiple input data streams, improves the synchronization and accuracy of video data, and thus enhances the coordination and stability of the entire process when processing multiple input data, optimizing the final output effect.

[0125] In one embodiment, the step of dynamically aligning a first data segment of type video in the first storage area according to the synchronization offset set based on adaptive clock synchronization technology and dynamic frame interpolation technology includes: S421: Perform resolution detection on each of the first data segments in the first storage area that correspond to the same timestamp and are of type video, and use the resolution as the input resolution; Specifically, all video data segments with the same timestamp (e.g., 0-50ms) are extracted from the first storage area (that is, the first data segment of type video), and the resolution parameters in the video segment header information are parsed to obtain the input resolution.

[0126] S422: When the input resolution is lower than the first threshold, the nearest neighbor interpolation mode is adopted, and the offset of each data segment on the time axis is determined according to the synchronization offset set. Based on the offset, the first data segment of type video in the first storage area is initially aligned by inserting or deleting frames at the corresponding positions. Nearest neighbor interpolation is a simple interpolation algorithm that directly fills new pixels with the value of the nearest pixel. It has low computational cost but may produce jagged edges.

[0127] The offset on the timeline is the amount of delay or advance of a certain input video relative to the reference timeline.

[0128] Initial alignment involves quickly correcting large-scale time shifts by interpolating or deleting frames, but this may sacrifice some image quality.

[0129] Specifically, the offsets of each input path at the same timestamp are read from the synchronization offset set. For the first data segment of type video in the first storage area, if the offset is positive, the number of frames to be inserted is calculated, and the average of the previous frame and the next frame is used as the inserted frame; if the offset is negative, redundant frames are deleted.

[0130] S423: When the input resolution is between the first threshold and the second threshold, a bilinear interpolation mode and an adaptive clock synchronization technique are used to dynamically align the first data segment of type video in the first storage area according to the synchronization offset set. Bilinear interpolation mode is a linear interpolation algorithm based on four adjacent pixels. It has a smoother effect than nearest neighbor and is suitable for medium resolution.

[0131] Dynamic alignment processing combines clock synchronization and high-precision interpolation to achieve a comprehensive alignment scheme that balances timing accuracy and image quality.

[0132] Specifically, based on the resource allocation data corresponding to the acquisition pool, and using adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set. The timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with large delays (i.e., the analyzed offset is greater than the acquisition offset threshold, such as input path 2 being 50ms slower), the acquisition rate is increased (e.g., from 30fps to 31fps) to gradually reduce the time difference. For the first data segment of type video in the first storage area, if there is a positive offset, bilinear interpolation is used to generate intermediate frames for the frames to be inserted, thereby achieving motion-compensated frame interpolation. S424: When the input resolution is higher than the second threshold, a bicubic interpolation mode and adaptive clock synchronization technology are used to dynamically align the first data segment of type video in the first storage area according to the synchronization offset set.

[0133] Bicubic interpolation mode is based on a cubic polynomial interpolation algorithm using 16 adjacent pixels, which preserves more details and is suitable for high-resolution video.

[0134] Specifically, based on the resource allocation data corresponding to the acquisition pool, and using adaptive clock synchronization technology, the offset is analyzed according to the synchronization offset set. The timestamp of the first data segment is dynamically modified according to the offset. Furthermore, for input streams with large delays (i.e., the analyzed offset is greater than the acquisition offset threshold, such as input path 2 being 50ms slower), the acquisition rate is increased (e.g., from 30fps to 31fps) to gradually reduce the time difference. For the first data segment of type video in the first storage area, if there is a positive offset, bicubic interpolation is used to generate intermediate frames for the frames to be inserted, thereby achieving motion-compensated frame interpolation.

[0135] In other words, this embodiment uses nearest neighbor interpolation, which requires less computation, to reduce latency for low-resolution videos; and uses bicubic interpolation to preserve more details for high-resolution videos.

[0136] Optionally, bicubic interpolation is enabled only when the resolution is high and system resources are sufficient; otherwise, it is downgraded to bilinear interpolation.

[0137] System resource adequacy is determined by comprehensively monitoring the following parameters in real time: (1) CPU utilization: The CPU utilization of the current processing pool is lower than the preset threshold (e.g., ≤70%). (2) Available memory: The percentage of remaining memory in the processing pool is higher than the preset threshold (e.g., ≥30%). (3) Processing delay: The processing delay of the current frame is less than 50% of the maximum allowable delay threshold; (4) Thread idle rate: The percentage of idle threads in the processing pool is ≥20%.

[0138] Optionally, when the following conditions are met simultaneously, the system is deemed to have sufficient resources and bicubic interpolation is enabled. The specific conditions include: CPU utilization ≤ 70%, available memory ≥ 30%, processing latency < maximum allowable latency × 0.5, and number of idle threads ≥ total number of threads × 20%. If any judgment condition is not met, it will automatically be downgraded to bilinear interpolation.

[0139] This embodiment first performs resolution detection to obtain the input resolution, providing a basis for judging different processing modes in the future. When the input resolution is lower than a first threshold, the nearest neighbor interpolation mode combined with adaptive clock synchronization technology is used to determine the offset based on the synchronization offset set. The first data segment of the video is initially aligned by frame interpolation or frame deletion. This method can initially solve the alignment problem of data segments on the time axis in a relatively simple and efficient way at low resolutions. When the input resolution is between the first and second thresholds, the bilinear interpolation mode and adaptive clock synchronization technology are used to perform dynamic alignment processing based on the synchronization offset set. This can achieve more accurate alignment of video data at medium resolutions, improving the accuracy of alignment. When the input resolution is higher than the second threshold, the bicubic interpolation mode and adaptive clock synchronization technology are used to perform dynamic alignment processing based on the synchronization offset set. At high resolutions, this can achieve more refined and high-quality dynamic alignment of video data segments, adapting to the alignment needs of video data of different resolutions. Overall, this improves the flexibility, accuracy, and effectiveness of dynamic alignment of video data of different resolutions, optimizing the processing effect of video data.

[0140] Please see Figure 4 As shown, in one embodiment, a multi-channel video live streaming processing system based on dynamic resource allocation is provided. The system includes a multi-channel input broadcast control station 1 and a client 2. The client 2 is communicatively connected to the multi-channel input broadcast control station 1, and the multi-channel input broadcast control station 1 is configured to implement the multi-channel video live streaming processing method based on dynamic resource allocation described above.

[0141] Users can input the video splicing mode through the client or the multi-input broadcast console, eliminating the need for manual switching of the video layout of multiple input sources or simple splicing through the software interface. This solves the problem of cumbersome operation, greatly improves operational efficiency, and reduces the probability of errors. The technical solution of this application is used to control the multi-input broadcast console, eliminating the need to rely on high-performance computer software to run the broadcast software, thus avoiding the problems of bulky equipment and high power consumption. Moreover, the multi-input broadcast console is portable, making it well-suited for mobile shooting scenarios. Through the dynamic resource allocation strategy and modular resource pool architecture (i.e., a three-level resource pool consisting of an acquisition pool, a processing pool, and an encoding pool) for the multi-input broadcast console, intelligent broadcast processing of multi-input data is achieved.

[0142] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0143] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program using signal-related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0145] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for processing multi-channel live video streaming based on dynamic resource allocation, characterized in that, The method is used to control a multi-input broadcast control station, and the method includes: Obtain the live broadcast signal; In response to the live broadcast signal, the video splicing mode is obtained, wherein the video splicing mode is data input by the user through the client or the multi-input control panel; Based on a dynamic resource allocation strategy for a multi-input broadcast control station, comprising an acquisition pool, a processing pool, and an encoding pool, data acquisition, data splicing, and encoding output are performed according to the aforementioned video splicing mode. The acquisition pool is a resource pool responsible for the acquisition and temporary storage of multiple video / audio streams. The processing pool is a resource pool responsible for video splicing, image quality optimization, and audio-visual synchronization. The encoding pool is a resource pool that compresses and encodes the processed audio and video data according to the target format and pushes it to the output device. The dynamic resource allocation strategy is an algorithm that adjusts system resource allocation in real time, dynamically optimizing the resource ratio in the acquisition, processing, and encoding stages based on current input load, output requirements, and latency configuration data. The steps of data acquisition, data splicing, and encoding output based on the dynamic resource allocation strategy, acquisition pool, processing pool, and encoding pool of the multi-input broadcast control station, according to the video splicing mode, include: Acquire target control data and resource allocation data. The target control data includes: acquisition sub-data, image processing sub-data and encoding output sub-data. The acquisition sub-data includes: target interval and segmentation granularity. The image processing sub-data includes defining splicing mode and image quality processing parameters. The encoding output sub-data includes setting encoding format, bitrate and audio synchronization strategy. Data acquisition, data stitching, and encoding output are performed based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the screen stitching mode. Specifically, based on the acquisition sub-data, the acquisition pool is controlled to acquire data according to the resource allocation data corresponding to the acquisition pool; based on the screen processing sub-data and the screen stitching mode, the processing pool is controlled to stitch data according to the resource allocation data corresponding to the processing pool; and based on the encoding output sub-data, the encoding pool is controlled to encode and output data according to the resource allocation data corresponding to the encoding pool. The target control data is determined through the following steps: acquiring an analysis signal, responding to the analysis signal, acquiring acquisition performance data, output performance data, output mode data, and delay configuration data, and determining the target control data based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data, wherein the delay configuration data is the minimum allowable delay threshold and the maximum allowable delay threshold input by the user through the client or the multi-channel input broadcast console; The analysis signal is triggered in real time by monitoring user operations or system status changes through a hardware interface or software module; the output performance data describes the performance of the external output data from the multi-input broadcast console to the multi-input broadcast console; the resource allocation data includes: thread resources; The resource allocation data is determined through the following steps: in response to the analysis signal, the resource allocation data corresponding to the acquisition pool, the processing pool, and the encoding pool are determined according to the dynamic resource allocation strategy based on the acquisition performance data, the output performance data, the output mode data, and the delay configuration data.

2. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 1, characterized in that, The steps of data acquisition, data stitching, and encoding output based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the screen stitching mode include: Based on the resource allocation data corresponding to the acquisition pool, and according to the acquisition sub-data of the target control data, each input data of the multi-channel input broadcasting station is segmented using the same segmentation timestamp to generate a data segment that conforms to the storage structure of the first storage area of ​​the acquisition pool, which is then used as the first data segment. Each of the first data segments is stored in the first storage area. The dynamic determination of the acquired sub-data includes: calculating the segmented time interval based on the resolution, frame rate, and bit rate of the input data of type video, as the initial interval; adjusting the initial interval based on the maximum allowable delay threshold corresponding to the delay configuration data, as the target interval; determining the segmented granularity based on the acquisition performance data and the delay configuration data; and using the target interval and the segmented granularity as the acquired sub-data.

3. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 2, characterized in that, The steps of data acquisition, data stitching, and encoding output based on the target control data, the resource allocation data, the acquisition pool, the processing pool, the encoding pool, and the image stitching mode further include: Based on the resource allocation data corresponding to the processing pool, and according to the splicing configuration of the splicing sub-data of the target control data, the first data segments in the first storage area are spliced ​​together with the same timestamp to obtain the first image. Based on the resource allocation data corresponding to the processing pool, and according to the quality processing data in the splicing sub-data, the first image is subjected to image quality normalization processing to obtain the target image; Based on the resource allocation data corresponding to the processing pool, according to each target screen and each first data segment of type audio in the first storage area, the screen and audio are matched according to the audio-visual combination data in the splicing sub-data to obtain the second data segment, and the second data segment is stored in the second storage area of ​​the processing pool.

4. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 3, characterized in that, The step of performing image quality normalization processing on the first image based on the quality processing data in the stitched sub-data to obtain the target image includes: Based on the quality processing data in the splicing sub-data, the first image is subjected to dynamic correction of the color correction matrix, brightness cross-region interpolation compensation, and resolution adaptive mapping to obtain the second image. Alternatively, the second image can be used as the target image, or a specific object region can be identified in the second image. If a specific object region is identified, dynamic interpolation compensation can be applied to the specific object region based on the quality processing data to obtain the target image.

5. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 4, characterized in that, The acquisition of the live broadcast signal includes: Acquire a switching signal, and in response to the switching signal acquire the live broadcast signal, wherein the switching signal is a signal input by the user through a client or the multi-input control panel; Within a preset time period after the start of the generation time of the switching signal, the steps of using the second frame as the target frame, or performing specific object region identification on the second frame, and if a specific object region is identified, applying dynamic interpolation compensation to the specific object region based on the quality processing data to obtain the target frame, include: Based on the specified image, the second image is subjected to a smooth transition process to obtain the target image; or, a specific object region is identified on the second image. If a specific object region is identified, dynamic interpolation compensation and smooth transition processing are applied to the specific object region based on the quality processing data and the specified image to obtain the target image. Otherwise, based on the specified image, the second image is subjected to a smooth transition process to obtain the target image. The designated screen is the second screen that is historically closest to the time when the switching signal was generated.

6. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 3, characterized in that, The step of stitching together the first data segments in the first storage area with the same timestamp according to the image stitching mode and the stitching configuration of the target control data to obtain the first image includes: Based on the image stitching mode and the stitching configuration, a missing image data type detection with the same timestamp is performed on each of the first data segments of type video in the first storage area to obtain a first result; If the first result is incomplete, then based on the exception handling strategy, according to the splicing configuration of the splicing sub-data of the target control data, each of the first data segments in the first storage area is spliced ​​with the same timestamp and the first result is incomplete to obtain the first image. The exception handling strategy is data determined based on the first result and the image splicing mode. If the first result is complete, then according to the image stitching mode and the stitching configuration, the image data with the same timestamp of each of the first data segments in the first storage area are stitched together to obtain the first image.

7. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 2, characterized in that, The method further includes: Based on the resource allocation data corresponding to the acquisition pool, the timestamp difference and audio delay parameters of the input data of the multi-channel input broadcasting station are monitored in real time to generate a set of synchronization offsets. Based on the resource allocation data corresponding to the acquisition pool, and based on adaptive clock synchronization technology and dynamic frame interpolation technology, the first data segment of type video in the first storage area is dynamically aligned according to the synchronization offset set.

8. The multi-channel video live streaming processing method based on dynamic resource allocation according to claim 7, characterized in that, The step of dynamically aligning the first data segment of type video in the first storage area according to the synchronization offset set, based on adaptive clock synchronization technology and dynamic frame interpolation technology, includes: The resolution of each of the first data segments in the first storage area that corresponds to the same timestamp and is of type video is detected and used as the input resolution. When the input resolution is lower than the first threshold, the nearest neighbor interpolation mode is adopted. The offset of each data segment on the time axis is determined according to the synchronization offset set. Based on the offset, the first data segment of type video in the first storage area is initially aligned by inserting or deleting frames at the corresponding positions. When the input resolution is between the first threshold and the second threshold, a bilinear interpolation mode and an adaptive clock synchronization technique are used to dynamically align the first data segment of type video in the first storage area according to the synchronization offset set. When the input resolution is higher than the second threshold, bicubic interpolation mode and adaptive clock synchronization technology are used to dynamically align the first data segment of type video in the first storage area according to the synchronization offset set.

9. A multi-channel video live streaming processing system based on dynamic resource allocation, characterized in that, The system includes a multi-input broadcast control station and a client, wherein the client is communicatively connected to the multi-input broadcast control station, and the multi-input broadcast control station is configured to implement the multi-channel video live streaming processing method based on dynamic resource allocation as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Live broadcast control method and device

    CN112866725A

  • Video processing method and device, electronic equipment and storage medium

    CN114071116A