An intelligent video source switching system and method for a program director switcher

By performing N-gram feature vector calculation and inter-frame difference analysis on the video source of the director switcher, combining the regular expression generation algorithm of the dictionary tree and the spatiotemporal feature extraction technology, the video switching behavior prediction and decision-making is used using LSTM and dual-deep Q learning models, which solves the problems of inaccurate and insufficient real-time video switching in the existing technology, and realizes efficient and automated video source switching.

CN119316539BActive Publication Date: 2025-06-20ZHONGYI INSTECH TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411417598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-06-20
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

The existing intelligent video switching methods have limitations in understanding video content and scene analysis. It is difficult to accurately capture key information and changes in complex scenarios, and it is difficult to achieve a good balance between real-time and decision-making accuracy, which cannot meet the needs of high-quality live broadcasts.

Method used

By obtaining the input video stream of the video source, N-gram feature vector calculation and inter-frame difference analysis are performed to obtain video frame segmentation data; scene analysis is performed based on the regular expression generation algorithm of the dictionary tree to obtain scene feature signatures; using spatiotemporal feature extraction and fusion technology, combining the lightweight timing convolution network model based on LSTM and the dual-deep Q learning model based on Markov's decision process, video switching behavior prediction and decision-making are performed.

Benefits of technology

It improves the accuracy and coherence of video source switching, enhances the system's adaptability and robustness to different types of program content, realizes adaptive learning to complex environments, and can make the optimal switching strategies based on different scenarios, reduces the workload of manual operations, and improves the working efficiency and automation of the director switcher.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316539B_ABST
    Figure CN119316539B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video source switching, and discloses an intelligent video source switching system and method for a production switcher. The method includes: obtaining an input video stream of a current video source through the production switcher, calculating N-gram feature vectors and performing inter-frame difference analysis to obtain video frame segmentation data; performing scene analysis on the video frame segmentation data to obtain a scene feature signature; extracting and fusing spatio-temporal features for each video frame in the input video stream according to the scene feature signature to obtain a scene spatio-temporal fusion feature set; inputting the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM for video switching behavior prediction to obtain video switching behavior prediction data; inputting the video switching behavior prediction data into a double deep Q learning model based on a Markov decision process for video source switching decision-making to obtain a video source switching decision result. The present application improves the accuracy of intelligent video source switching of the production switcher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video source switching, and particularly to an intelligent video source switching system and method for a production switcher. Background Art

[0002] With the continuous development of television program production and live broadcast technology, production switchers play an increasingly important role in video source switching and video synthesis. Traditional production switcher operations mainly rely on manual experience. Not only do operators need to have rich professional knowledge and quick reaction capabilities, but they are also easily affected by subjective factors, resulting in unstable switching effects. At the same time, with the increase in multi-camera shooting and complex scenes, the pressure and error risk of manual operation also increase.

[0003] In recent years, significant progress has been made in the field of artificial intelligence technology for video processing, providing new ideas for solving the above problems. However, existing intelligent video switching methods still have some limitations. The understanding of video content and scene analysis are not deep enough, making it difficult to accurately capture key information and changes in complex scenes. Secondly, it is difficult to achieve a good balance between real-time performance and decision-making accuracy, and it cannot meet the requirements of high-quality live broadcasts. Summary of the Invention

[0004] This application provides an intelligent video source switching system and method for a production switcher, which is used to improve the accuracy of intelligent video source switching of the production switcher.

[0005] In a first aspect, this application provides an intelligent video source switching method for a production switcher. The intelligent video source switching method for the production switcher includes:

[0006] Obtain the input video stream of the current video source through the production switcher, and perform N-gram feature vector calculation and inter-frame difference analysis on the input video stream to obtain video frame segmentation data;

[0007] Perform scene analysis on the video frame segmentation data based on the regular expression generation algorithm of the trie tree to obtain a scene feature signature;

[0008] Extract and fuse spatio-temporal features for each video frame in the input video stream according to the scene feature signature to obtain a scene spatio-temporal fusion feature set;

[0009] Input the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM for video switching behavior prediction to obtain video switching behavior prediction data;

[0010] Input the video switching behavior prediction data into a double deep Q-learning model based on the Markov decision process for video source switching decision-making to obtain a video source switching decision result.

[0011] In a second aspect, the present application provides an intelligent video source switching system for a production switcher. The intelligent video source switching system for a production switcher includes:

[0012] An acquisition module, configured to acquire an input video stream of a current video source through a production switcher, and perform N-gram feature vector calculation and inter-frame difference analysis on the input video stream to obtain video frame segmentation data;

[0013] An analysis module, configured to perform scene analysis on the video frame segmentation data based on a regular expression generation algorithm of a trie to obtain a scene feature signature;

[0014] A fusion module, configured to perform spatio-temporal feature extraction and fusion on each video frame in the input video stream according to the scene feature signature to obtain a scene spatio-temporal fusion feature set;

[0015] A prediction module, configured to input the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM for video switching behavior prediction to obtain video switching behavior prediction data;

[0016] A decision module, configured to input the video switching behavior prediction data into a double deep Q-learning model based on a Markov decision process for video source switching decision to obtain a video source switching decision result.

[0017] In a third aspect of the present application, an intelligent video source switching device for a production switcher is provided, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the intelligent video source switching device for a production switcher executes the above-mentioned intelligent video source switching method for a production switcher.

[0018] In a fourth aspect of the present application, a computer-readable storage medium is provided, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is made to execute the above-mentioned intelligent video source switching method for a production switcher.

[0019] In the technical solution provided by this application, through N-gram feature vector calculation and inter-frame difference analysis, fine-grained segmentation of video content is achieved, which helps to improve the accuracy and coherence of video source switching. The regular expression generation algorithm based on the trie tree can adaptively generate scene matching rules, greatly improving the flexibility and robustness of scene recognition, enabling the system to better adapt to different types of program content. The spatio-temporal feature extraction and fusion technology fully considers the temporal information and spatial features of the video, enabling the system to more comprehensively understand the video content and thus make more reasonable switching decisions. The lightweight temporal convolutional network model based on LSTM effectively combines long-term and short-term memory capabilities and local feature extraction capabilities, can accurately predict video switching behaviors, and improves the real-time performance and prediction accuracy of the system. The double deep Q-learning model based on the Markov decision process is used for video source switching decision-making, realizing adaptive learning in complex environments and being able to make optimal switching strategies according to different scenarios. The multi-level feature extraction and decision-making process makes the system highly interpretable. While ensuring the switching quality, it greatly reduces the workload of manual operations, improving the working efficiency and automation level of the director switcher. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic diagram of an embodiment of the intelligent video source switching method for a director switcher in an embodiment of this application;

[0022] Figure 2 It is a schematic diagram of an embodiment of the intelligent video source switching system for a director switcher in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of the present application provide an intelligent video source switching system and method for a program director switcher. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 , an embodiment of the intelligent video source switching method for a program director switcher in the embodiments of the present application includes:

[0025] Step S101: Obtain the input video stream of the current video source through the program director switcher, and perform N-gram feature vector calculation and inter-frame difference analysis on the input video stream to obtain video frame segmentation data;

[0026] It can be understood that the execution subject of the present application can be an intelligent video source switching system for a program director switcher, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution subject as an example.

[0027] Specifically, through the input interface of the program switcher, the input video stream of the current video source is obtained. The program switcher can receive multiple video sources in real time and select one of them as the current input video stream. Frame extraction processing is performed on the input video stream, and the video stream is extracted into several video frames at fixed time intervals to form a continuous frame sequence. The color feature vector of each video frame is calculated through the color histogram. The color histogram is statistically obtained based on the color distribution of pixels, and it can effectively reflect the color features of the image. The color feature vector is converted into an N-gram feature vector. The N-gram feature vector is a feature expression method based on sequence information and can capture the local continuity features in the video frame sequence. The cosine similarity between adjacent video frames is calculated to quantify the inter-frame difference. The cosine similarity is a commonly used similarity measurement method. The closer its value is to 1, the more similar the two vectors are, and vice versa, the greater the difference. After obtaining the inter-frame difference value by calculating the cosine similarity, an adaptive threshold segmentation method is used to process the difference value to determine the initial segmentation point set. Adaptive threshold segmentation can dynamically adjust the threshold according to the statistical characteristics of the inter-frame difference value to achieve more accurate segmentation point detection. For the initial segmentation point set, audio feature extraction is performed on the video frame sequence to obtain the audio feature vector. Audio feature extraction is based on the audio signal in the video stream, and common methods include MFCC (Mel Frequency Cepstral Coefficients) and STFT (Short-Time Fourier Transform), etc. The audio feature vector can reflect the audio features corresponding to the video frame. It is weighted and fused with the inter-frame difference value to form a fusion feature sequence. Dynamic programming is performed on the fusion feature sequence to obtain the target segmentation point set. Dynamic programming is an optimization algorithm that can find the optimal segmentation scheme globally. According to the target segmentation point set, the video frame sequence is segmented and verified to ensure that the video frames within each segment have high similarity, while the difference between segments is large, and accurate video frame segmentation data is obtained.

[0028] Step S102: Perform scene analysis on the video frame segmentation data using a regular expression generation algorithm based on a trie tree to obtain a scene feature signature;

[0029] Specifically, multi-scale feature extraction is performed on the video frame segmentation data to obtain a hierarchical feature descriptor, and a multi-level scene representation is constructed based on the hierarchical feature descriptor. The multi-scale feature extraction method can capture feature information at different scales in the video frame, making the scene representation more rich and accurate. By constructing a multi-level scene representation, these feature information can be better organized and managed. The multi-level scene representation is segmented to obtain an initial scene boundary, and the correlation between adjacent scenes is analyzed based on the initial scene boundary to obtain a scene correlation matrix. The determination of the initial scene boundary is based on the change of the feature descriptor. By analyzing the correlation between adjacent scenes, a scene correlation matrix is constructed to reflect the similarity and connection between different scenes. The initial scene boundary is optimized for scene segmentation according to the scene correlation matrix to obtain an optimized scene segmentation result. By analyzing the information in the scene correlation matrix, the initial scene boundary is adjusted and optimized. The key frame sequence of each scene is extracted according to the optimized scene segmentation result to obtain a set of scene representative frames, and a dynamically updated scene feature trie of the set of scene representative frames is constructed based on the regular expression generation algorithm of the trie. The extraction of the key frame sequence is based on the scene segmentation result, and a set of scene representative frames is constructed by selecting representative frames. Based on the regular expression generation algorithm of the trie, the feature information in the set of scene representative frames can be effectively organized and managed, and a dynamically updated scene feature trie is constructed. This trie can be dynamically adjusted as the video content changes to keep the scene features updated and accurate. An adaptive regular expression template is generated according to the dynamically updated scene feature trie to obtain a set of scene matching rules, and a scene recognition strategy is generated according to the set of scene matching rules. The adaptive regular expression template can be adjusted according to the change of the scene features to generate an accurate set of scene matching rules. According to the set of matching rules, a scene recognition strategy is formulated to ensure the accurate recognition and classification of different scenes. Feature matching and encoding are performed on the video frame segmentation data according to the scene recognition strategy, and the feature information in the video frame segmentation data is converted into a scene feature signature to form a complete scene recognition and classification system.

[0030] Step S103: Perform spatio-temporal feature extraction and fusion on each video frame in the input video stream according to the scene feature signature to obtain a set of scene spatio-temporal fusion features;

[0031] Specifically, for each video frame in the input video stream, multi-level feature construction and adaptive feature screening are performed based on the scene feature signature to obtain a scene-related feature set. The multi-level feature construction method can capture feature information of various scales and levels in the video frame, and the adaptive feature screening selects features related to the scene by evaluating the feature importance. Dynamic convolution operations and feature extraction are performed on the scene-related feature set to obtain a scene-adaptive feature map. The dynamic convolution operation captures the changes in the video frame in the time dimension by applying convolutional kernels at different time periods, and the feature extraction further processes these convolution results to generate an adaptive feature map reflecting the scene features. Temporal difference calculation is performed on the scene-adaptive feature map to obtain inter-frame change features. The temporal difference calculation reveals the dynamic features of the video frame changing over time by comparing the feature changes between consecutive frames. A dynamic weight matrix is calculated based on the inter-frame change features, and weighted fusion of the scene-adaptive feature map and the dynamic weight matrix is performed to obtain enhanced spatio-temporal features. The dynamic weight matrix is calculated based on the inter-frame change features and is used to reflect the importance and correlation of the video frame at different time points. By performing weighted fusion of the scene-adaptive feature map and the dynamic weight matrix, the expression ability of the spatio-temporal features is enhanced. A temporal feature sequence is generated based on the enhanced spatio-temporal features, and multi-scale spatio-temporal feature aggregation is performed on the temporal feature sequence to obtain a scene spatio-temporal fusion feature set. The multi-scale spatio-temporal feature aggregation enhances the expression ability of the feature set by integrating the temporal features at different scales, ensuring its good adaptability and expressiveness at different time and space scales.

[0032] Step S104: Input the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM for video switching behavior prediction to obtain video switching behavior prediction data;

[0033] Specifically, the scene spatio-temporal fusion feature set is input into the lightweight temporal convolutional network model based on LSTM, and the input tensor processing is performed on the scene spatio-temporal fusion feature set to obtain an input tensor of T×D dimensions, where T is the time step and D is the feature dimension. The input tensor of T×D dimensions is input into the multi-channel one-dimensional convolutional layer, and convolution operations are performed through F convolutional kernels of different sizes to extract the local patterns and change trends of the input feature set in the time dimension, obtaining a local temporal feature representation of T×F dimensions. The maximum pooling operation is performed on the local temporal feature representation of T×F dimensions to obtain a reduced-dimensional feature map of (T / s)×F dimensions, where the pooling window size is k and the stride is s. The maximum pooling operation can reduce the feature dimension while retaining important temporal information and reducing the computational complexity. The reduced-dimensional feature map of (T / s)×F dimensions is input into the residual block, and a multi-scale temporal feature of (T / s)×F dimensions is output. The residual block contains two layers of one-dimensional convolution and a skip connection, which can effectively capture multi-scale temporal features and solve the gradient vanishing problem through the skip connection to ensure the training stability of the deep network. Channel attention calculation is performed on the multi-scale temporal feature of (T / s)×F dimensions to generate an F-dimensional weight vector and perform element-wise multiplication to obtain a weighted feature map of (T / s)×F dimensions. The channel attention mechanism can adaptively adjust the weights according to the feature importance, highlight the key features, and improve the feature expression ability. The weighted feature map of (T / s)×F dimensions is reshaped into a sequence of (T / s) F-dimensional vectors and input into the LSTM unit. The hidden state of H dimensions is output at each time step, obtaining a long-range dependence feature of (T / s)×H dimensions. The LSTM unit can effectively capture long-term dependence relationships through its long and short-term memory mechanism and has excellent performance in modeling time series data. Self-attention mechanism analysis is performed on the long-range dependence feature of (T / s)×H dimensions, and the attention matrix of (T / s)×(T / s) is calculated and multiplied to obtain a context-enhanced feature of (T / s)×H dimensions. The self-attention mechanism can capture the dependence relationships between positions in the sequence, provide global context information, and enhance the feature representation ability. The context-enhanced feature of (T / s)×H dimensions is input into the gated linear unit, and the information flow is controlled through the sigmoid gating mechanism to output a candidate switching behavior feature of (T / s)×H dimensions. The gated linear unit can screen according to the importance of the input information and improve the accuracy of the feature expression. Multi-task learning head analysis is performed on the candidate switching behavior feature of (T / s)×H dimensions, including N parallel fully connected layers. Each fully connected layer outputs a prediction result of Mi dimensions, obtaining N prediction vectors of different dimensions, where i = 1, 2,..., N. The multi-task learning head can process multiple related tasks simultaneously, improving the generalization ability and prediction accuracy of the model. The N prediction vectors of different dimensions are weighted and summed according to the predefined task weights and normalized through the softmax function to obtain the video switching behavior prediction data, realizing the accurate prediction of the video switching behavior and providing intelligent decision support for the director switcher.

[0034] Step S105: Input the video switching behavior prediction data into the double deep Q - learning model based on the Markov decision process for video source switching decision - making to obtain the video source switching decision result.

[0035] Specifically, obtain the state information of the current video source, and splice the video switching behavior prediction data and the state information of the current video source to obtain an S - dimensional state vector. Input the S - dimensional state vector into the state encoder network. Through two fully - connected layers and the ReLU activation function, obtain an E - dimensional encoded state representation to ensure that the high - dimensional features of the input state vector can be effectively compressed and expressed. Perform an attention mechanism process on the E - dimensional encoded state representation, calculate the similarity with the historical state, and obtain an A - dimensional attention weight vector. The attention mechanism generates an attention weight vector by calculating the similarity between the current state and the historical state, enabling the model to focus on the historical information most relevant to the current state. Perform a weighted sum of the historical state according to the A - dimensional attention weight vector to obtain an H - dimensional historical context vector. Splice the E - dimensional encoded state representation and the H - dimensional historical context vector, and input them into the value network in the double deep Q - learning model based on the Markov decision process. Through three fully - connected layers and the LeakyReLU activation function, obtain a scalar state value estimate. The value network estimates the value of the current state by comprehensively analyzing the state and the historical context, providing a basis for action selection. At the same time, perform an action generation network process on the E - dimensional encoded state representation. Through two fully - connected layers and the Tanh activation function, obtain a K - dimensional candidate action vector. The action generation network generates possible candidate actions by processing the encoded state. Splice the K - dimensional candidate action vector and the E - dimensional encoded state representation, and input them into the advantage function network in the double deep Q - learning model based on the Markov decision process. Through three fully - connected layers and the ELU activation function, obtain a K - dimensional action advantage value. The advantage function network is used to evaluate the advantage degree of each candidate action relative to the current state, so as to select the optimal action among the candidate actions. According to the scalar state value estimate and the K - dimensional action advantage value, use the double Q - learning formula to calculate and obtain a K - dimensional Q - value vector. The double Q - learning formula provides a more stable and accurate action value evaluation by combining the state value and the action advantage. Perform the Softmax function operation with a temperature parameter of τ on the K - dimensional Q - value vector to obtain a K - dimensional action probability distribution. The Softmax function operation can transform the Q - value vector into a probability distribution, ensuring that the action selection has a certain degree of exploration and randomness. Perform importance sampling according to the K - dimensional action probability distribution, select the optimal action index, and look up the table to obtain the corresponding video source switching decision result. In this way, the model can intelligently switch between different video sources, optimizing the operation efficiency of the production switcher and the quality of video switching.

[0036] Perform a concatenation operation on the K-dimensional candidate action vector and the E-dimensional encoded state representation to obtain a (K + E)-dimensional input feature vector, ensuring that action and state information are represented in the same vector and providing a unified input. Perform a matrix multiplication operation on the first fully connected layer of the advantage function network based on the (K + E)-dimensional input feature vector to obtain an M-dimensional intermediate feature vector. The fully connected layer projects the input feature vector into a high-dimensional space through matrix multiplication to extract the deep features of the input features. Perform batch normalization on the M-dimensional intermediate feature vector to obtain a normalized M-dimensional feature vector. Batch normalization reduces the deviation of the data distribution by normalizing the mean and variance of each feature, improving the training stability and convergence speed of the model. Perform an ELU activation function calculation on the normalized M-dimensional feature vector according to batch normalization to obtain an M-dimensional non-linear feature vector. The ELU activation function can introduce non-linearity, enabling the network to have stronger expressive power, and its smoothness in the negative value region helps to alleviate the vanishing gradient problem. Perform a dropout operation on the M-dimensional non-linear feature vector to obtain an M-dimensional sparse feature vector. The dropout operation randomly discards some neurons to prevent the model from overfitting and improve its generalization ability. Perform a matrix multiplication operation on the second fully connected layer of the advantage function network based on the M-dimensional sparse feature vector to obtain an N-dimensional intermediate feature vector. The second fully connected layer further extracts features and performs a deep feature transformation on the input. Perform layer normalization on the N-dimensional intermediate feature vector to obtain a normalized N-dimensional feature vector. Layer normalization is similar to batch normalization, but it normalizes each layer of each feature vector, thus more stably controlling the distribution of feature values. Perform an ELU activation function calculation on the normalized N-dimensional feature vector to obtain an N-dimensional non-linear feature vector. By repeatedly using the ELU activation function, it is ensured that the feature vector can capture rich non-linear information in each layer. Perform a residual connection on the N-dimensional non-linear feature vector and perform an element-wise addition with the (K + E)-dimensional input feature vector to obtain an N-dimensional residual feature vector. The residual connection reduces information loss by directly adding the input information to the output, improving the training effect of the network. Especially in deep networks, the residual connection can effectively alleviate the vanishing gradient problem. Perform a matrix multiplication operation on the third fully connected layer of the advantage function network based on the N-dimensional residual feature vector and perform an ELU activation function calculation to obtain a K-dimensional action advantage value. In this way, the advantage function network can accurately evaluate the advantage of the candidate action in the current state, providing a basis for the final action selection.

[0037] In the embodiments of the present application, through N-gram feature vector calculation and inter-frame difference analysis, fine-grained segmentation of video content is achieved, which helps to improve the accuracy and coherence of video source switching. The regular expression generation algorithm based on the trie tree can adaptively generate scene matching rules, greatly improving the flexibility and robustness of scene recognition, enabling the system to better adapt to different types of program content. The spatio-temporal feature extraction and fusion technology fully considers the temporal information and spatial features of the video, enabling the system to more comprehensively understand the video content and thus make more reasonable switching decisions. The lightweight temporal convolutional network model based on LSTM effectively combines long-term and short-term memory capabilities and local feature extraction capabilities, can accurately predict video switching behavior, and improves the real-time performance and prediction accuracy of the system. The dual deep Q-learning model based on the Markov decision process is used for video source switching decision-making, realizing adaptive learning in complex environments and being able to make optimal switching strategies according to different scenarios. The multi-level feature extraction and decision-making process make the system highly interpretable. While ensuring the switching quality, it greatly reduces the workload of manual operations, improving the working efficiency and automation level of the director switcher.

[0038] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0039] (1) Obtain the input video stream of the current video source through the director switcher, and perform frame extraction processing on the input video stream to obtain a video frame sequence;

[0040] (2) Calculate the color histogram of each video frame according to the video frame sequence to obtain a color feature vector, and transform the color feature vector to obtain an N-gram feature vector;

[0041] (3) Calculate the cosine similarity between adjacent video frames according to the N-gram feature vector to obtain an inter-frame difference value, and perform adaptive threshold segmentation on the inter-frame difference value to obtain an initial segmentation point set;

[0042] (4) Extract the audio feature vector from the video frame sequence according to the initial segmentation point set, and perform weighted fusion of the audio feature vector and the inter-frame difference value to obtain a fusion feature sequence;

[0043] (5) Perform dynamic programming on the fusion feature sequence to obtain a target segmentation point set, and segment and verify the video frame sequence according to the target segmentation point set to obtain video frame segmentation data.

[0044] Specifically, the input video stream of the current video source is obtained through the director switcher, and the video stream is received and processed in real time through the hardware interface and software interface of the director switcher. Frame extraction processing is performed on the video stream, and the extraction interval is usually a fixed time step, such as 30 frames per second, to obtain a video frame sequence. Suppose there is a video stream, and a series of frames are obtained after frame extraction, denoted as {F1, F2, F3, …, F n}, where F i represents the i-th frame. Calculate the color histogram for each video frame according to the video frame sequence. The color histogram is a commonly used method for representing image features, which describes the color distribution of an image by counting the number of pixels of different colors in the image. For each video frame F i , calculate its color histogram to obtain the color feature vector H i . The color histogram is usually divided into several color channels, and each channel is further divided into several color levels. For example, for an RGB image, calculate the histograms for the R, G, and B channels respectively. Transform the color feature vector to obtain the N-gram feature vector. The N-gram feature vector is a feature representation method based on sequence information that can capture the relationship between consecutive elements in a sequence. For the color feature vector H i of each frame, construct the N-gram feature vector by concatenating the color feature vectors of consecutive N frames. For example, for N = 3, the feature vector G i = [H i , H i+1 , H i+2 . Calculate the cosine similarity between adjacent video frames according to the N-gram feature vector to obtain the inter-frame difference value. The cosine similarity is a commonly used similarity measurement method for measuring the similarity between two vectors. For two N-gram feature vectors G i and G i11 , their cosine similarity calculation formula is:

[0045]

[0046] where, G i ·G i+1 represents the dot product of vectors, and ∥G i ∥ and ∥G i + 1∥ represent the L2 norm of the vectors. The inter-frame difference value can then be expressed as:

[0047] diff(G i , G i+1 ) = 1 - cos_sin(G i , G i+1 );

[0048] The difference values between each pair of adjacent video frames are calculated through the above formula. Adaptive threshold segmentation is performed on the inter-frame difference values to obtain an initial segmentation point set. The adaptive threshold segmentation method dynamically adjusts the segmentation threshold by analyzing the statistical characteristics of the inter-frame difference values, and identifies the frames with obvious changes in video content. The initial segmentation point set represents possible scene transition points. Audio feature extraction is performed on the video frame sequence according to the initial segmentation point set to obtain audio feature vectors. Common methods such as MFCC (Mel Frequency Cepstral Coefficients) or STFT (Short-Time Fourier Transform) can be used for audio feature extraction, so as to effectively extract the frequency domain features of the audio signal. For each initial segmentation point, the corresponding audio feature vector is extracted and denoted as A i 。The audio feature vectors and the inter-frame difference values are weighted and fused to obtain a fused feature sequence. By setting weight parameters, the importance of audio features and video features is balanced. Assume that the weight of the audio feature is α and the weight of the inter-frame difference value is β, and the fused feature is expressed as:

[0049] F i = α·A i + β·diff(G i ,G i+1 );

[0050] Dynamic programming is performed on the fused feature sequence to obtain a target segmentation point set, and the video frame sequence is segmented and verified according to the target segmentation point set to obtain video frame segmentation data. Dynamic programming is an optimization algorithm that can find the optimal segmentation scheme globally. By comparing the cost functions of different segmentation schemes, the scheme with the minimum cost is selected. The target segmentation point set represents the finally determined scene transition points. According to the target segmentation point set, the video frame sequence is segmented and verified to ensure that the video frames within each segment have high similarity, while the difference between segments is large, so as to obtain accurate video frame segmentation data.

[0051] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0052] (1) Perform multi-scale feature extraction on the video frame segmentation data to obtain hierarchical feature descriptors, and construct a multi-level scene representation according to the hierarchical feature descriptors;

[0053] (2) Perform segmentation processing on the multi-level scene representation to obtain initial scene boundaries, and perform correlation analysis on adjacent scenes according to the initial scene boundaries to obtain a scene correlation matrix;

[0054] (3) Optimize the scene segmentation of the initial scene boundaries according to the scene correlation matrix to obtain an optimized scene segmentation result;

[0055] (4) Extract the key frame sequence of each scene according to the optimized scene segmentation result to obtain the scene representative frame set, and construct a dynamically updated scene feature trie for the scene representative frame set based on the regular expression generation algorithm of the trie;

[0056] (5) Generate an adaptive regular expression template according to the dynamically updated scene feature trie to obtain the scene matching rule set, and generate a scene recognition strategy according to the scene matching rule set;

[0057] (6) Perform feature matching and encoding on the video frame segmentation data according to the scene recognition strategy to obtain the scene feature signature.

[0058] Specifically, perform multi-scale feature extraction on the video frame segmentation data to obtain a hierarchical feature descriptor. Feature extraction is performed on the video frames at different scales to capture information at different levels. Use a convolutional neural network to obtain multi-scale features in different pooling layers, and form a hierarchical feature descriptor by combining these features. For example, assume there is a video frame sequence, and feature maps of different scales are obtained after feature extraction by a convolutional neural network, denoted as Among them, represents the i-th frame feature map at scale s. Combine these feature maps to obtain the hierarchical feature descriptor D i . Construct a multi-level scene representation according to the hierarchical feature descriptor. The multi-level scene representation forms a multi-level feature space by combining features of different scales, and is used to describe the scene information in the video frame. Combine the hierarchical feature descriptors of all frames to form an overall multi-level scene representation matrix R. Perform segmentation processing on the multi-level scene representation to obtain the initial scene boundary. The segmentation processing can be implemented by a clustering algorithm, such as K-means or DBSCAN. By clustering the feature space, similar frames are segmented into the same scene. Assume that K means clustering is used, and the initial scene boundary can be expressed as {B1, B2, …, B )}}, where B i represents the boundary of the i-th scene. Perform correlation analysis on adjacent scenes according to the initial scene boundary to obtain the scene correlation matrix. The scene correlation matrix is used to describe the similarity between adjacent scenes, and is realized by calculating the cosine similarity between the scene boundary features. Assume that S i and S + are the feature vectors of two adjacent scenes, then their similarity can be expressed as:

[0059]

[0060] By calculating the similarity of all adjacent scene pairs, a scene correlation matrix M is obtained. According to the scene correlation matrix, the initial scene boundary is optimized for scene segmentation to obtain the optimized scene segmentation result. The optimization can be achieved through dynamic programming or graph cut algorithms to minimize the cost function in the similarity matrix, ensuring the maximization of internal similarity within each scene and the maximization of differences between different scenes.

[0061] ″″

[0062] The optimized scene segmentation result can be expressed as {B1, B2, …, B1} ′ where B i represents the boundary of the i-th scene after optimization. According to the optimized scene segmentation result, the key frame sequence of each scene is extracted to obtain the scene representative frame set. A key frame refers to a frame that can represent the characteristics of the entire scene and can be determined by selecting the frame with the largest feature change or the most representative frame in the scene. Assume that the key frames of scene S i are {K1, K2, …, K3}. Based on the trie-based regular expression generation algorithm, a dynamically updated scene feature trie of the scene representative frame set is constructed. A trie is an efficient structured data storage method used to store and match feature sequences. By inserting the feature vectors in the scene representative frame set into the trie and dynamically updating when new scene representative frames appear, a dynamically updated scene feature trie T is formed. According to the dynamically updated scene feature trie, an adaptive regular expression template is generated to obtain the scene matching rule set. The adaptive regular expression template automatically generates a regular expression that can match the scene features by analyzing the feature patterns in the trie. The scene matching rule set can be expressed as {R1, R2, …, R5}, where R i is a regular expression that matches specific scene features. According to the scene matching rule set, a scene recognition strategy is generated. The scene recognition strategy matches and encodes the input video frame segmentation data by applying the regular expression template to identify the scene category to which each video frame belongs. According to the scene recognition strategy, the video frame segmentation data is feature-matched and encoded to obtain the scene feature signature. The scene feature signature is a high-level feature representation that can effectively describe the scene information of the video content.

[0063] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0064] (1) Perform multi-level feature construction and adaptive feature screening on each video frame in the input video stream according to the scene feature signature to obtain the scene-related feature set;

[0065] (2) Perform dynamic convolution operations and feature extraction on the scene-related feature set to obtain the scene-adaptive feature map, and perform temporal difference calculation on the scene-adaptive feature map to obtain the inter-frame change feature;

[0066] (3) Calculate the dynamic weight matrix according to the inter-frame change characteristics, and perform weighted fusion on the scene adaptability feature map and the dynamic weight matrix to obtain enhanced spatio-temporal features;

[0067] (4) Generate a temporal feature sequence based on the enhanced spatio-temporal features, and perform multi-scale spatio-temporal feature aggregation on the temporal feature sequence to obtain a scene spatio-temporal fusion feature set.

[0068] Specifically, perform multi-level feature construction and adaptive feature screening on each video frame in the input video stream according to the scene feature signature to obtain a scene-related feature set. The multi-level feature construction captures various feature information in the video frame through feature extraction networks at different levels, such as convolutional neural networks. Assume that the input video stream is {V1, V2, V3, …, V n}, where V i represents the i-th frame. Through the multi-layer convolution operation of the convolutional neural network, a set of feature maps F i of each video frame is obtained, and these feature maps capture different visual features at different levels. Through the adaptive feature screening method, these features are screened, and those features highly related to the scene feature signature are selected to form a scene-related feature set S i . Perform dynamic convolution operation and feature extraction on the scene-related feature set to obtain a scene adaptability feature map. The dynamic convolution operation refers to applying a convolution kernel at different time periods to capture the temporal changes of the video frame. Assume that the convolution kernel W is used to perform a convolution operation on the feature set S i to obtain a scene adaptability feature map C i , and its formula is expressed as:

[0069] C i = W * S i ;

[0070] where, * represents the convolution operation. Through the dynamic convolution operation, the local patterns and change trends of the video frame in the time dimension are effectively extracted. Perform temporal difference calculation on the scene adaptability feature map to obtain the inter-frame change characteristics. The temporal difference calculation reveals the dynamic characteristics of the video frame changing over time by comparing the feature changes between consecutive frames. Assume that C i and C i+1 are the scene adaptability feature maps of adjacent frames respectively, then the formula for their temporal difference is:

[0071] ΔC i = C i+1 - C i ;

[0072] where, ΔC iRepresents the difference features between the i-th frame and the (i + 1)-th frame. Through calculation, the inter-frame change features between each pair of adjacent frames are obtained. Based on the inter-frame change features, a dynamic weight matrix is calculated, and the scene adaptability feature map and the dynamic weight matrix are weighted and fused to obtain enhanced spatio-temporal features. The dynamic weight matrix is used to measure the importance of each feature in the time dimension. Assume the dynamic weight matrix is W : , then the formula for weighted fusion is:

[0073] E i = W : ·C i ;

[0074] where, E i represents the enhanced spatio-temporal features. Through weighted fusion, the features that change significantly in the time dimension are highlighted to obtain more representative spatio-temporal features. Based on the enhanced spatio-temporal features, a temporal feature sequence is generated, and multi-scale spatio-temporal feature aggregation is performed on the temporal feature sequence to obtain the scene spatio-temporal fusion feature set. By arranging the enhanced spatio-temporal features in chronological order to form a temporal feature sequence, denoted as {E1, E2, E3, …, E n}}. Multi-scale spatio-temporal feature aggregation captures the feature changes in the long term and short term by aggregating features at different time scales. Assume using the multi-scale aggregation function P, then the formula for the scene spatio-temporal fusion feature set is:

[0075] F %= = P(E1, E2, …, E n );

[0076] where, F %= represents the scene spatio-temporal fusion feature set. The multi-scale aggregation function can be a pooling operation based on a time window, or an aggregation method based on a long short-term memory network (LSTM) or an attention mechanism.

[0077] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0078] (1) Input the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM, and perform input tensor processing on the scene spatio-temporal fusion feature set to obtain an input tensor of T×D dimensions, where T is the time step and D is the feature dimension;

[0079] (2) Input the input tensor of T×D dimensions into a multi-channel one-dimensional convolutional layer, and perform convolution operations through F convolutional kernels of different sizes to obtain a local temporal feature representation of T×F dimensions;

[0080] (3) Perform a max pooling operation on the local temporal feature representation of T×F dimensions to obtain a reduced-dimensional feature map of (T / s)×F dimensions, where the pooling window size is k and the stride is s;

[0081] (4) Input the (T / s)×F - dimensional dimensionality - reduced feature map into the residual block, and output the (T / s)×F - dimensional multi - scale temporal features. The residual block contains two layers of one - dimensional convolution and a skip connection;

[0082] (5) Perform channel attention calculation on the (T / s)×F - dimensional multi - scale temporal features, generate an F - dimensional weight vector and perform element - wise multiplication to obtain the (T / s)×F - dimensional weighted feature map;

[0083] (6) Reshape the (T / s)×F - dimensional weighted feature map into (T / s) sequences of F - dimensional vectors, input them into the LSTM unit, and output the H - dimensional hidden state at each time step to obtain the (T / s)×H - dimensional long - range dependence features;

[0084] (7) Perform self - attention mechanism analysis on the (T / s)×H - dimensional long - range dependence features, calculate the (T / s)×(T / s) attention matrix and perform multiplication to obtain the (T / s)×H - dimensional context - enhanced features;

[0085] (8) Input the (T / s)×H - dimensional context - enhanced features into the gated linear unit, control the information flow through the sigmoid gating mechanism, and output the (T / s)×H - dimensional candidate switching behavior features;

[0086] (9) Perform multi - task learning head analysis on the (T / s)×H - dimensional candidate switching behavior features, including N parallel fully - connected layers. Each fully - connected layer outputs a prediction result of Mi dimensions to obtain N prediction vectors of different dimensions, where i = 1, 2,..., N;

[0087] (10) Perform weighted summation on the N prediction vectors of different dimensions according to the predefined task weights and normalize through the softmax function to obtain the video switching behavior prediction data.

[0088] Specifically, input the scene spatio - temporal fusion feature set into the lightweight temporal convolutional network model based on LSTM, and perform input tensor processing on the scene spatio - temporal fusion feature set to obtain a T×D - dimensional input tensor, where T is the time step and D is the feature dimension. Suppose there is a scene spatio - temporal fusion feature set {F1, F2,…, F T} of a video sequence, and the dimension of each feature vector F i is D, then the input tensor X can be expressed as:

[0089] X = [F1, F2,…, F T ;

[0090] The dimension of X is T×D. The T×D-dimensional input tensor is input into a multi-channel one-dimensional convolutional layer, and convolution operations are performed through F convolutional kernels of different sizes to obtain a local temporal feature representation of T×F dimensions. The multi-channel convolutional layer extracts local temporal features of different scales through convolutional kernels of different sizes. For example, using k1, k2, …, k B These convolutional kernels of different sizes perform convolution operations on the input tensor X to obtain a local temporal feature representation C, whose dimension is T×F. A max-pooling operation is performed on the T×F-dimensional local temporal feature representation to obtain a downsampled feature map of (T / s)×F dimensions, where the pooling window size is k and the stride is s. The max-pooling operation reduces the feature dimension by taking the maximum value within the pooling window while retaining important feature information. Assuming the pooling window size is k and the stride is s, the dimension of the downsampled feature map P is (T / s)×F. The (T / s)×F-dimensional downsampled feature map is input into a residual block, and a multi-scale temporal feature of (T / s)×F dimensions is output. The residual block contains two one-dimensional convolutional layers and a skip connection, which solves the problem of gradient disappearance in deep networks through residual connections and effectively extracts multi-scale temporal features. Assuming the input to the residual block is P and the output is R, its dimension is (T / s)×F. Channel attention calculation is performed on the (T / s)×F-dimensional multi-scale temporal feature to generate an F-dimensional weight vector and perform element-wise multiplication to obtain a (T / s)×F-dimensional weighted feature map. The channel attention mechanism highlights key features by assigning different weights to each feature channel. Assuming the weight vector is W D , and the weighted feature map is A, its calculation formula is:

[0091] A = R⊙W D ;

[0092] where ⊙ represents element-wise multiplication. The (T / s)×F-dimensional weighted feature map is reshaped into a sequence of (T / s) F-dimensional vectors and input into an LSTM unit. The hidden state of H dimensions is output at each time step to obtain a long-range dependence feature of (T / s)×H dimensions. The LSTM captures long-range dependence relationships in time series data through its memory unit. Assuming the output of the LSTM unit is L, its dimension is (T / s)×H. Self-attention mechanism analysis is performed on the (T / s)×H-dimensional long-range dependence feature, and an attention matrix of (T / s)×(T / s) is calculated and multiplied to obtain a context-enhanced feature of (T / s)×H dimensions. The self-attention mechanism enhances the feature representation by calculating the correlations between positions in the sequence. Assuming the attention matrix is A = , and the context-enhanced feature is E D , its calculation formula is:

[0093] E D = A = ·L;

[0094] where, A= The dimension is (T / s) × (T / s). The (T / s) × H-dimensional context-enhanced features are input into the gated linear unit, and the information flow is controlled through the sigmoid gating mechanism, and the (T / s) × H-dimensional candidate switching behavior features are output. The gated linear unit controls the information flow through the sigmoid function, retains important information, and suppresses irrelevant information. Assume that the output of the gated linear unit is G, and its dimension is (T / s) × H. The (T / s) × H-dimensional candidate switching behavior features are analyzed by a multi-task learning head, including N parallel fully connected layers, and each fully connected layer outputs M i dimensional prediction results, obtaining N prediction vectors with different dimensions, where i = 1, 2, …, N. Assume that the output of each fully connected layer is O i , and its dimension is M i . According to the predefined task weights, the N prediction vectors with different dimensions are weighted and summed, and normalized through the softmax function to obtain the video switching behavior prediction data. Assume that the task weight is w i , and the weighted sum result is O, and its calculation formula is:

[0095]

[0096] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0097] (1) Obtain the status information of the current video source, and splice the video switching behavior prediction data and the status information of the current video source to obtain an S-dimensional status vector;

[0098] (2) Input the S-dimensional status vector into the status encoder network, and through two layers of fully connected layers and the ReLU activation function, obtain an E-dimensional encoded status representation;

[0099] (3) Perform an attention mechanism process on the E-dimensional encoded status representation, calculate the similarity with the historical status, and obtain an A-dimensional attention weight vector;

[0100] (4) According to the A-dimensional attention weight vector, perform weighted summation on the historical status to obtain an H-dimensional historical context vector;

[0101] (5) Splice the E-dimensional encoded status representation and the H-dimensional historical context vector, input them into the value network in the double deep Q-learning model based on the Markov decision process, and through three layers of fully connected layers and the LeakyReLU activation function, obtain a scalar status value estimate;

[0102] (6) Perform an action generation network process on the E-dimensional encoded status representation, and through two layers of fully connected layers and the Tanh activation function, obtain a K-dimensional candidate action vector;

[0103] (7) Concatenate the K-dimensional candidate action vector and the E-dimensional encoded state representation, and input them into the advantage function network in the double deep Q-learning model based on the Markov decision process. Through three fully connected layers and the ELU activation function, obtain the K-dimensional action advantage value;

[0104] (8) According to the scalar state value estimation and the K-dimensional action advantage value, use the double Q-learning formula to calculate the K-dimensional Q-value vector;

[0105] (9) Perform the Softmax function operation with the temperature parameter τ on the K-dimensional Q-value vector to obtain the K-dimensional action probability distribution;

[0106] (10) Perform importance sampling according to the K-dimensional action probability distribution, select the optimal action index, and look up the table to obtain the corresponding video source switching decision result.

[0107] Specifically, obtain the state information of the current video source, including the frame rate, resolution, and current scene features of the video stream. Concatenate the state information with the video switching behavior prediction data to obtain an S-dimensional state vector. Assume that the state information of the current video source is V-dimensional and the video switching behavior prediction data is P-dimensional, then S = V + P, and the concatenated state vector is denoted as s. Input the S-dimensional state vector s into the state encoder network. Through two fully connected layers and the ReLU activation function, obtain the E-dimensional encoded state representation. The output of the first fully connected layer is:

[0108] h1 = ReLU(W1s + b1);

[0109] where W1 and b1 are the weight matrix and bias vector of the first layer respectively. The output of the second fully connected layer is the encoded state representation:

[0110] h2 = ReLU(W2h1 + b2);

[0111] where W2 and b2 are the weight matrix and bias vector of the second layer respectively. Perform attention mechanism processing on the E-dimensional encoded state representation h2, calculate the similarity with the historical state, and obtain the A-dimensional attention weight vector. Assume that the historical state representation is {h 2,1 , h 2,2 , …, h 2,T}, where T is the number of historical time steps, then the calculation formula for the attention weight vector a is:

[0112]

[0113] where · represents the dot product of vectors. Perform weighted summation on the historical state according to the A-dimensional attention weight vector a to obtain the H-dimensional historical context vector h D . Its calculation formula is:

[0114]

[0115] Concatenate the E - dimensional encoded state representation h2 and the H - dimensional historical context vector h D and input it into the value network in the double - deep Q - learning model based on the Markov decision process. Through three fully - connected layers and the LeakyReLU activation function, obtain the scalar state value estimate. The concatenated vector is denoted as h D\nD]= . The output of the first fully - connected layer is:

[0116] h3 = LeakyReLU(W3h D\nD]= + b3);

[0117] The output of the second fully - connected layer is:

[0118] h _ = LeakyReLU(W _ h3 + b _ );

[0119] The output of the third fully - connected layer is the scalar state value estimate V:

[0120] V = W a h _ + b a ;

[0121] Meanwhile, process the E - dimensional encoded state representation h2 through the action - generation network. Through two fully - connected layers and the Tanh activation function, obtain the K - dimensional candidate action vector. The output of the first fully - connected layer is:

[0122] h a = Tanh(W b h2 + b b );

[0123] The output of the second fully - connected layer is the candidate action vector A:

[0124] A = Tanh(W c h a + b c );

[0125] Concatenate the K - dimensional candidate action vector A and the E - dimensional encoded state representation h2, and input it into the advantage - function network in the double - deep Q - learning model based on the Markov decision process. Through three fully - connected layers and the ELU activation function, obtain the K - dimensional action advantage value. The concatenated vector is denoted as h D\nD]=2 . The output of the first fully - connected layer is:

[0126] h b = ELU(W d h D\nD]=2 + b d );

[0127] The output of the second fully connected layer is:

[0128] h c = ELU(W e h b + b e );

[0129] The output of the third fully connected layer is the action advantage value A f :

[0130] A f = W 1g h c + b 1g ;

[0131] According to the scalar state value estimate V and the K-dimensional action advantage value A f , using the double Q-learning formula, the K-dimensional Q-value vector is calculated. The double Q-learning formula is:

[0132]

[0133] where Q(s,a) represents the Q-value of action a in state s, and A f [k] represents the advantage value of the k-th action. Perform the Softmax function operation with temperature parameter τ on the K-dimensional Q-value vector to obtain the K-dimensional action probability distribution. The Softmax function formula is:

[0134]

[0135] where P(a i ) represents the probability of action a i , and τ is the temperature parameter. Perform importance sampling according to the K-dimensional action probability distribution, select the optimal action index, and look up the table to obtain the corresponding video source switching decision result. For example, assume that the K-dimensional action probability distribution is {P(a1), P(a2), …, P(a j )}, select the action index a * with the largest probability through importance sampling, and the corresponding video source switching decision result is D(a * ).

[0136] In a specific embodiment, the process of performing the step of concatenating the K-dimensional candidate action vector and the E-dimensional encoded state representation, inputting it into the advantage function network in the double deep Q-learning model based on the Markov decision process, and obtaining the K-dimensional action advantage value through three fully connected layers and the ELU activation function may specifically include the following steps:

[0137] (1) Perform a concatenation operation on the K-dimensional candidate action vector and the E-dimensional encoded state representation to obtain a (K + E)-dimensional input feature vector;

[0138] (2) Perform matrix multiplication on the first fully connected layer of the advantage function network using the (K + E)-dimensional input feature vector to obtain an M-dimensional intermediate feature vector;

[0139] (3) Perform batch normalization on the M-dimensional intermediate feature vector to obtain a normalized M-dimensional feature vector;

[0140] (4) Calculate using the ELU activation function based on the normalized M-dimensional feature vector to obtain an M-dimensional non-linear feature vector;

[0141] (5) Perform dropout operation on the M-dimensional non-linear feature vector to obtain an M-dimensional sparse feature vector;

[0142] (6) Perform matrix multiplication on the second fully connected layer of the advantage function network using the M-dimensional sparse feature vector to obtain an N-dimensional intermediate feature vector;

[0143] (7) Perform layer normalization on the N-dimensional intermediate feature vector to obtain a normalized N-dimensional feature vector;

[0144] (8) Calculate using the ELU activation function based on the normalized N-dimensional feature vector to obtain an N-dimensional non-linear feature vector;

[0145] (9) Perform residual connection on the N-dimensional non-linear feature vector and perform element-wise addition with the (K + E)-dimensional input feature vector to obtain an N-dimensional residual feature vector;

[0146] (10) Perform matrix multiplication on the third fully connected layer of the advantage function network using the N-dimensional residual feature vector and perform ELU activation function calculation to obtain a K-dimensional action advantage value.

[0147] Specifically, perform a concatenation operation on the K-dimensional candidate action vector and the E-dimensional encoded state representation to obtain a (K + E)-dimensional input feature vector. Assume the candidate action vector is A, the encoded state representation is H, and the concatenated input feature vector is I, then:

[0148] I = [A; H];

[0149] Among them, the dimension of I is (K + E). Perform matrix multiplication on the first fully connected layer of the advantage function network using the (K + E)-dimensional input feature vector I to obtain an M-dimensional intermediate feature vector. Assume the weight matrix of the first fully connected layer is W1 and the bias vector is b1, then the calculation formula for the intermediate feature vector M1 is:

[0150] M1 = W1I + b1;

[0151] Perform batch normalization on the M-dimensional intermediate feature vector M1 to obtain a normalized M-dimensional feature vector. Batch normalization reduces the bias of the data distribution by normalizing the mean and variance of each feature. Perform the ELU activation function calculation based on the normalized M-dimensional feature vector to obtain an M-dimensional non-linear feature vector. The ELU activation function enables the network to better represent complex feature relationships by introducing non-linearity. Perform a dropout operation on the M-dimensional non-linear feature vector to obtain an M-dimensional sparse feature vector. Dropout prevents model overfitting and improves its generalization ability by randomly discarding some neurons. Perform a matrix multiplication operation on the second fully connected layer of the advantage function network according to the M-dimensional sparse feature vector to obtain an N-dimensional intermediate feature vector. Assume that the weight matrix of the second fully connected layer is W2 and the bias vector is b2. Then the calculation formula for the intermediate feature vector z2 is:

[0152] z2 = W2·d1 + b2;

[0153] Perform layer normalization on the N-dimensional intermediate feature vector to obtain a normalized N-dimensional feature vector. Layer normalization further improves the training effect of the model by normalizing the feature vector so that its mean is 0 and variance is 1. Perform the ELU activation function calculation based on the normalized N-dimensional feature vector to obtain an N-dimensional non-linear feature vector. Perform a residual connection on the N-dimensional non-linear feature vector and perform an element-wise addition with the (K + E)-dimensional input feature vector to obtain an N-dimensional residual feature vector. The residual connection reduces the vanishing gradient problem by directly adding the input to the output. Perform a matrix multiplication operation on the third fully connected layer of the advantage function network according to the N-dimensional residual feature vector and perform the ELU activation function calculation to obtain the K-dimensional action advantage value.

[0154] The above describes the intelligent video source switching method for the production switcher in the embodiments of the present application. Next, the intelligent video source switching system for the production switcher in the embodiments of the present application will be described. Please refer to Figure 2 One embodiment of the intelligent video source switching system for the production switcher in the embodiments of the present application includes:

[0155] An acquisition module 201, configured to obtain an input video stream of the current video source through the production switcher, and perform N-gram feature vector calculation and inter-frame difference analysis on the input video stream to obtain video frame segmentation data;

[0156] An analysis module 202, configured to perform scene analysis on the video frame segmentation data based on the regular expression generation algorithm of the trie tree to obtain a scene feature signature;

[0157] A fusion module 203, configured to perform spatio-temporal feature extraction and fusion on each video frame in the input video stream according to the scene feature signature to obtain a scene spatio-temporal fusion feature set;

[0158] A prediction module 204 for inputting the scene spatio-temporal fusion feature set into a lightweight temporal convolutional network model based on LSTM to predict video switching behaviors and obtain video switching behavior prediction data;

[0159] A decision-making module 205 for inputting the video switching behavior prediction data into a double deep Q-learning model based on the Markov decision process to make a video source switching decision and obtain a video source switching decision result.

[0160] Through the collaborative cooperation of the above-mentioned various components, through N-gram feature vector calculation and inter-frame difference analysis, the refined segmentation of video content is achieved, which helps to improve the accuracy and coherence of video source switching. The regular expression generation algorithm based on the trie can adaptively generate scene matching rules, greatly improving the flexibility and robustness of scene recognition, enabling the system to better adapt to different types of program content. The spatio-temporal feature extraction and fusion technology fully considers the temporal information and spatial features of the video, enabling the system to more comprehensively understand the video content and thus make more reasonable switching decisions. The lightweight temporal convolutional network model based on LSTM effectively combines long-term and short-term memory capabilities and local feature extraction capabilities, can accurately predict video switching behaviors, and improves the real-time performance and prediction accuracy of the system. Using a double deep Q-learning model based on the Markov decision process for video source switching decisions realizes adaptive learning in complex environments and can make optimal switching strategies according to different scenarios. The multi-level feature extraction and decision-making process makes the system highly interpretable. While ensuring the switching quality, it greatly reduces the workload of manual operations, improving the working efficiency and automation degree of the production switcher.

[0161] This application also provides an intelligent video source switching device for a production switcher. The intelligent video source switching device for a production switcher includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes the steps of the intelligent video source switching method for a production switcher in the above-mentioned various embodiments.

[0162] This application also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer executes the steps of the intelligent video source switching method for a production switcher.

[0163] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0164] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0165] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. An intelligent video source switching method for a director switcher, characterized in that: The method comprises: The input video stream of the current video source is obtained through the director switcher, and the N-gram feature vector calculation and inter-frame difference analysis are performed on the input video stream to obtain video frame segmentation data; Performing scene analysis on the video frame segmentation data based on a regular expression generation algorithm of a dictionary tree to obtain a scene feature signature; Extract and fuse spatiotemporal features of each video frame in the input video stream according to the scene feature signature to obtain a scene spatiotemporal fusion feature set; Inputting the scene spatiotemporal fusion feature set into a LSTM-based lightweight temporal convolutional network model to predict video switching behavior, thereby obtaining video switching behavior prediction data; The video switching behavior prediction data is input into a dual-depth Q learning model based on a Markov decision process to make a video source switching decision, and a video source switching decision result is obtained.

2. The intelligent video source switching method for a director switcher according to claim 1, characterized in that: The step of obtaining the input video stream of the current video source through the director switcher, and performing N-gram feature vector calculation and frame difference analysis on the input video stream to obtain video frame segmentation data includes: The input video stream of the current video source is obtained through the director switcher, and frame extraction processing is performed on the input video stream to obtain a video frame sequence; Calculating a color histogram for each video frame according to the video frame sequence to obtain a color feature vector, and converting the color feature vector to obtain an N-gram feature vector; Calculating the cosine similarity between adjacent video frames according to the N-gram feature vector to obtain an inter-frame difference value, and performing adaptive threshold segmentation on the inter-frame difference value to obtain an initial segmentation point set; Extract audio features from the video frame sequence according to the initial segmentation point set to obtain an audio feature vector, and perform weighted fusion of the audio feature vector and the inter-frame difference value to obtain a fused feature sequence; Dynamic programming is performed on the fused feature sequence to obtain a target segmentation point set, and the video frame sequence is segmented and tested according to the target segmentation point set to obtain video frame segmentation data.

3. The intelligent video source switching method for a director switcher according to claim 1, characterized in that: The regular expression generation algorithm based on the dictionary tree performs scene analysis on the video frame segmentation data to obtain a scene feature signature, including: Performing multi-scale feature extraction on the video frame segmentation data to obtain a hierarchical feature descriptor, and constructing a multi-level scene representation according to the hierarchical feature descriptor; Segmenting the multi-level scene representation to obtain an initial scene boundary, and performing correlation analysis on adjacent scenes according to the initial scene boundary to obtain a scene correlation matrix; Performing scene segmentation optimization on the initial scene boundary according to the scene correlation matrix to obtain an optimized scene segmentation result; Extracting a key frame sequence of each scene according to the optimized scene segmentation result to obtain a scene representative frame set, and constructing a dynamically updated scene feature dictionary tree of the scene representative frame set based on a regular expression generation algorithm of the dictionary tree; Generate an adaptive regular expression template according to the dynamically updated scene feature dictionary tree to obtain a scene matching rule set, and generate a scene recognition strategy according to the scene matching rule set; The video frame segmentation data is feature matched and encoded according to the scene recognition strategy to obtain a scene feature signature.

4. The intelligent video source switching method for a director switcher according to claim 1, characterized in that: The step of extracting and fusing spatiotemporal features of each video frame in the input video stream according to the scene feature signature to obtain a scene spatiotemporal fusion feature set includes: Performing multi-level feature construction and adaptive feature screening on each video frame in the input video stream according to the scene feature signature to obtain a scene-related feature set; Performing dynamic convolution operation and feature extraction on the scene-related feature set to obtain a scene-adaptive feature map, and performing temporal difference calculation on the scene-adaptive feature map to obtain inter-frame change features; Calculating a dynamic weight matrix according to the inter-frame variation feature, and performing weighted fusion on the scene adaptability feature map and the dynamic weight matrix to obtain an enhanced spatiotemporal feature; A temporal feature sequence is generated according to the enhanced spatiotemporal feature, and multi-scale spatiotemporal feature aggregation is performed on the temporal feature sequence to obtain a scene spatiotemporal fusion feature set.

5. The intelligent video source switching method for a director switcher according to claim 1, characterized in that: The step of inputting the scene spatiotemporal fusion feature set into a lightweight temporal convolutional network model based on LSTM to predict video switching behavior and obtain video switching behavior prediction data includes: Inputting the scene spatiotemporal fusion feature set into a lightweight temporal convolutional network model based on LSTM, and performing input tensor processing on the scene spatiotemporal fusion feature set to obtain a T×D-dimensional input tensor, where T is the time step and D is the feature dimension; Inputting the T×D-dimensional input tensor into a multi-channel one-dimensional convolutional layer, performing a convolution operation through F convolution kernels of different sizes, and obtaining a T×F-dimensional local temporal feature representation; Performing a maximum pooling operation on the T×F dimensional local temporal feature representation to obtain a (T / s)×F dimensional reduced dimensionality feature map, where the pooling window size is k and the step size is s; Input the (T / s)×F dimensional reduced dimension feature map into a residual block, and output a (T / s)×F dimensional multi-scale temporal feature, wherein the residual block includes two layers of one-dimensional convolution and one skip connection; Perform channel attention calculation on the (T / s)×F dimensional multi-scale time series features, generate an F dimensional weight vector and perform element-by-element multiplication to obtain a (T / s)×F dimensional weighted feature map; Reshape the (T / s)×F-dimensional weighted feature map into a (T / s) F-dimensional vector sequence, input it into the LSTM unit, output an H-dimensional hidden state at each time step, and obtain a (T / s)×H-dimensional long-range dependency feature; Perform a self-attention mechanism analysis on the (T / s)×H dimensional long-range dependency features, calculate the (T / s)×(T / s) attention matrix and multiply them to obtain the (T / s)×H dimensional context enhancement features; The (T / s)×H dimensional context enhancement features are input into a gated linear unit, the information flow is controlled by a sigmoid gating mechanism, and a (T / s)×H dimensional candidate switching behavior feature is output; The (T / s)×H dimensional candidate switching behavior features are subjected to a multi-task learning head analysis, including N parallel fully connected layers, each fully connected layer outputs M i dimensional prediction results, and obtain prediction vectors of N different dimensions, where i = 1, 2, ..., N; The prediction vectors of the N different dimensions are weighted and summed according to the predefined task weights, and normalized by a softmax function to obtain the video switching behavior prediction data.

6. The intelligent video source switching method for a director switcher according to claim 1, characterized in that: The step of inputting the video switching behavior prediction data into a dual-depth Q learning model based on a Markov decision process to make a video source switching decision, and obtaining a video source switching decision result, includes: Acquire the state information of the current video source, and concatenate the video switching behavior prediction data and the state information of the current video source to obtain an S-dimensional state vector; Input the S-dimensional state vector into the state encoder network, and obtain the E-dimensional encoded state representation through two fully connected layers and ReLU activation function; Performing attention mechanism processing on the E-dimensional encoded state representation, calculating the similarity with the historical state, and obtaining an A-dimensional attention weight vector; Performing weighted summation of historical states according to the A-dimensional attention weight vector to obtain an H-dimensional historical context vector; The E-dimensional encoded state representation and the H-dimensional historical context vector are concatenated and input into the value network in the dual-depth Q-learning model based on the Markov decision process, and a scalar state value estimation is obtained through three layers of fully connected layers and LeakyReLU activation function; Performing action generation network processing on the E-dimensional encoded state representation, obtaining a K-dimensional candidate action vector through two fully connected layers and a Tanh activation function; The K-dimensional candidate action vector and the E-dimensional encoding state representation are concatenated and input into the advantage function network in the double-depth Q learning model based on the Markov decision process, and the K-dimensional action advantage value is obtained through three layers of fully connected layers and ELU activation function; According to the scalar state value estimate and the K-dimensional action advantage value, a K-dimensional Q value vector is calculated using a double Q learning formula; Performing a Softmax function operation with a temperature parameter of τ on the K-dimensional Q-value vector to obtain a K-dimensional action probability distribution; Importance sampling is performed according to the K-dimensional action probability distribution, an optimal action index is selected, and a table is looked up to obtain a corresponding video source switching decision result.

7. The intelligent video source switching method for a director switcher according to claim 6, characterized in that: The K-dimensional candidate action vector and the E-dimensional encoding state representation are concatenated and input into the advantage function network in the double-depth Q learning model based on the Markov decision process, and the K-dimensional action advantage value is obtained through three layers of fully connected layers and ELU activation function, including: Performing a concatenation operation on the K-dimensional candidate action vector and the E-dimensional encoding state representation to obtain a (K+E)-dimensional input feature vector; Performing a matrix multiplication operation on the first fully connected layer of the advantage function network according to the (K+E)-dimensional input feature vector to obtain an M-dimensional intermediate feature vector; Performing batch normalization processing on the M-dimensional intermediate feature vector to obtain a standardized M-dimensional feature vector; Performing ELU activation function calculation according to the standardized M-dimensional feature vector to obtain an M-dimensional nonlinear feature vector; Performing a dropout operation on the M-dimensional nonlinear feature vector to obtain an M-dimensional sparse feature vector; Performing a matrix multiplication operation on the second fully connected layer of the advantage function network according to the M-dimensional sparse feature vector to obtain an N-dimensional intermediate feature vector; Performing a layer normalization operation on the N-dimensional intermediate feature vector to obtain a normalized N-dimensional feature vector; Performing ELU activation function calculation according to the normalized N-dimensional feature vector to obtain an N-dimensional nonlinear feature vector; Performing residual connection on the N-dimensional nonlinear feature vector and performing element-wise addition on the (K+E)-dimensional input feature vector to obtain an N-dimensional residual feature vector; A matrix multiplication operation is performed on the third fully connected layer of the advantage function network according to the N-dimensional residual feature vector, and an ELU activation function calculation is performed to obtain a K-dimensional action advantage value.

8. An intelligent video source switching system for a director switcher, characterized in that: The system is used to execute the intelligent video source switching method for a director switcher according to any one of claims 1 to 7, and comprises: An acquisition module is used to acquire the input video stream of the current video source through the director switcher, and perform N-gram feature vector calculation and inter-frame difference analysis on the input video stream to obtain video frame segmentation data; An analysis module, used for performing scene analysis on the video frame segmentation data based on a regular expression generation algorithm of a dictionary tree to obtain a scene feature signature; A fusion module, used for extracting and fusing spatiotemporal features of each video frame in the input video stream according to the scene feature signature to obtain a scene spatiotemporal fusion feature set; A prediction module, used for inputting the scene spatiotemporal fusion feature set into a lightweight temporal convolutional network model based on LSTM to predict video switching behavior and obtain video switching behavior prediction data; The decision module is used to input the video switching behavior prediction data into a dual-depth Q learning model based on a Markov decision process to make a video source switching decision and obtain a video source switching decision result.

9. An intelligent video source switching device for a director switcher, characterized in that: The intelligent video source switching device for a director switch station comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the intelligent video source switching device for a director switcher executes the intelligent video source switching method for a director switcher as described in any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, an intelligent video source switching method for a director switch station as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Live video processing method, apparatus and device, and storage medium

    CN108712661A

  • Studio video feature image enhancement processing method

    CN118075552A