Logic evaluation method and system based on news shot switching

By combining the methods of convolutional neural network, graph neural network and deep Q network, the flexibility and real-time optimization of the lens switching method in news videos is solved, and the intelligence and fluency of lens switching is improved to adapt to the needs of different news types.

CN120378689AActive Publication Date: 2025-07-25CHANGJIANG DRAGON NEW MEDIA CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510863989.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing lens switching methods lack flexibility when processing complex news videos, are difficult to adapt to changes in different types of news, and lack real-time feedback optimization mechanisms, resulting in the lens switching rules being not intelligent and smooth enough.

Method used

The combination of convolutional neural network and graph neural network is adopted to detect lens switching points, extract image features and model timing dependencies, and use the fuzzy C-means clustering algorithm to evaluate the lens switching logic, and use the deep Q network to dynamically optimize the lens switching strategy, and make real-time adjustments based on user feedback.

Benefits of technology

It realizes the flexibility and intelligence of lens switching, improves the fluency and content coherence of lens switching, and can dynamically optimize according to different news types, improving the viewing experience of the audience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378689A_ABST
    Figure CN120378689A_ABST
Patent Text Reader

Abstract

The invention discloses a logic evaluation method and system based on news shot switching, and relates to the technical field of video intelligent processing, and the method comprises the steps: extracting video data from an original news video stream; detecting a lens switching point in the video, and analyzing to obtain a lens switching rule; a convolutional neural network is used to extract image features of lens switching, a time sequence dependency relationship is obtained through graph neural network modeling, and switching features are obtained through fusion; evaluating the lens switching logic by using a fuzzy C-means clustering algorithm, and calculating a lens switching logic score; and dynamically optimizing a lens switching strategy by using a deep Q network according to the lens switching logic score and the switching feature. According to the method provided by the invention, a method of combining deep learning and reinforcement learning is introduced, so that a lens switching strategy can be dynamically adjusted according to real-time feedback of a user, and the smoothness and content coherence of lens switching are improved. And flexibility, intellectualization and real-time optimization on lens switching are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video intelligent processing, and particularly to a logic evaluation method and system based on news shot switching. Background Art

[0002] In recent years, with the development of news video production technology, shot switching, as a crucial part of video editing, has gradually received extensive attention from the academic and industrial communities. Most traditional shot switching methods rely on simple rules or manual marking, and achieve smooth video conversion by manually setting parameters such as shot duration, switching frequency, and interval. With the continuous progress of machine learning and deep learning technologies, more and more automated methods have been applied to news video shot switching. These technologies can help the system automatically learn the dependencies between shots, and then perform more intelligent shot switching to ensure the automation and efficiency of video editing.

[0003] However, existing shot switching methods still face challenges, especially when dealing with shot switching rules in complex news videos. Traditional rule-based shot switching methods often rely on fixed duration and frequency settings, which are not very adaptable to different types of content in news reports (such as breaking news, live reports, and in-depth analysis news). For example, live news reports may require more frequent shot switching and shorter durations, while news analysis programs require smoother shot switching and longer durations for each shot. In addition, most existing technologies cannot accurately understand the temporal dependencies between shots, resulting in the adjustment of shot switching frequency and duration not meeting the actual needs of different news contents. More importantly, these methods lack a real-time dynamic optimization mechanism and cannot automatically adjust the shot switching strategy according to audience feedback or changes in news content. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is that the existing shot switching methods of technologies have problems such as inflexible switching rules, difficulty in adapting to changes in different news types, and lack of real-time feedback optimization.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A logic evaluation method based on news shot switching, including:

[0007] Extract video data from the original news video stream;

[0008] Detect shot switching points in the video and analyze to obtain shot switching rules;

[0009] Use a convolutional neural network to extract the image features of shot transitions, model the temporal dependencies through a graph neural network, and fuse them to obtain transition features;

[0010] Combine the transition features with the shot transition rules, use the fuzzy C-means clustering algorithm to evaluate the shot transition logic, and calculate the shot transition logic score;

[0011] Provide real-time feedback on the shot transition logic score, and use the deep Q-network to dynamically optimize the shot transition strategy based on the shot transition logic score and the transition features.

[0012] As a preferred embodiment of the logic evaluation method for news shot transitions according to the present invention, wherein: the shot transition point is the switching moment of consecutive shots;

[0013] The shot duration is the time difference between the start time of the current shot and the start time of the next shot; the shot transition frequency is the number of shot transitions per unit time, and the shot transition frequency is calculated by each shot transition point in the video and the total duration;

[0014] By calculating the similarity between the current frame and the previous frame , determine whether a shot transition has occurred;

[0015] If the similarity metric is lower than the set threshold, it is determined that a shot transition has occurred, and the current time point is marked as a shot transition point.

[0016] As a preferred embodiment of the logic evaluation method for news shot transitions according to the present invention, wherein: the shot transition rules include analyzing the shot transition duration, frequency, and order in the video based on the shot transition points as the characteristic values of news shot transitions;

[0017] According to the news type, adjust the three characteristic values respectively, and use the adjusted three characteristic values to represent the shot transition rules: for real-time reports, increase the shot transition frequency and shorten the shot duration; for news analysis, reduce the shot transition frequency and extend the shot duration; when switching from a news explanation shot to a live scene shot, extend the shot transition duration.

[0018] As a preferred embodiment of the logic evaluation method for news shot transitions according to the present invention, wherein: the extraction of the image features of the shot includes using a convolutional neural network CNN to process each frame of the image, extracting image features, texture, edges, and color distribution to describe the visual content of the shot; inputting each frame of the image, extracting local features through the convolutional layer, and learning the edge, color, and texture information in the image;

[0019] Perform downsampling operations through the pooling layer and enhance non-linear features through the ReLU activation function; finally, output the feature map , including the image features of each frame of the shot.

[0020] As a preferred solution of the logical evaluation method based on news shot switching according to the present invention, wherein: obtaining the temporal dependence relationship includes using a graph neural network GNN to model the dependence relationship between shots; the shot image features extracted by the CNN , are input into the GNN as graph node features;

[0021] Define that the nodes of the graph represent shots, the temporal dependence relationship between each shot switching point is used as the edge in the graph, and the order, duration, and frequency of shot switching are obtained through weighted summation as the weight of the edge; construct the adjacency matrix of the graph to represent the dependence relationship between shots; the elements of the adjacency matrix represent the shot node and the shot node the strength of the dependence relationship between them;

[0022] The graph neural network propagates information through the edges of the graph, captures the temporal relationship of shot switching, and generates a new temporal feature representation for each shot , for each shot node , the new temporal feature representation is updated by the weighted sum of the temporal features of the neighbor nodes and the adjacency matrix ; is the feature representation of the th shot at the th layer; is the feature representation of the th shot at the th layer;

[0023] Combining the outputs of the CNN and the GNN, fuse the image features and temporal features of each shot to obtain the shot switching feature representation as .

[0024] As a preferred solution of the logical evaluation method based on news shot switching according to the present invention, wherein: evaluating the shot switching logic includes constructing the objective function of fuzzy C-means clustering according to the shot switching duration, frequency, and sequence constraints and the clustering parameters; the clustering parameters include the number of clusters , the fuzzification factor and the class center ; evaluate the smoothness and coherence of shot switching through the membership degree of clustering;

[0025] Optimize the clustering parameters using a genetic algorithm, and use the objective function as the fitness function to generate multiple individuals, where each individual represents a set of clustering parameters; optimize the clustering parameter combination through the selection, crossover, and mutation operations of the genetic algorithm;

[0026] Perform clustering using the parameters optimized by the genetic algorithm to obtain the membership degree and class center of each shot transition; analyze each shot transition feature through the clustering center and membership degree , and assign smoothness and coherence scores to the shot transitions according to the clustering results, and input the scores into the reward function database.

[0027] As a preferred solution of the logic evaluation method for news shot transitions according to the present invention, wherein: the dynamic optimization of the shot transition strategy includes optimizing the shot transition strategy through a deep Q-network (DQN), and dynamically adjusting the duration, frequency, and content coherence characteristics of the shot transitions; defining the state space representation, where the state is composed of the image features and temporal features of each shot, , representing the state at the th moment;

[0028] Among them, represents the image features extracted by a CNN;

[0029] Define the action space, where the action is the decision of the shot transition, representing the action taken at the th moment, including adjusting the shot duration and switching frequency; perform Q-value update, and use the deep Q-network to calculate the Q-value according to the state and the action , and update the policy through a reinforcement learning algorithm, representing the discount factor;

[0030] Define the reward function to feedback the shot transition strategy, representing the reward at the th moment; is obtained by the weighted sum of the smoothness reward and the content coherence reward; through training iterations, the deep Q-network will learn the optimal shot transition strategy and automatically adjust the smoothness and content coherence of the shot transitions.

[0031] As a preferred solution of the logic evaluation system for news shot transitions according to the present invention, wherein: it includes a video data extraction module, an analysis module, a graph vision fusion module, and a policy adjustment module;

[0032] The video data extraction module is used to extract images from the original news video stream, perform preprocessing, and input the processed images into the analysis module;

[0033] The analysis module is used to detect the shot transition points and perform rule analysis based on the extracted features;

[0034] The graph-vision fusion module is used to fuse the image features extracted by CNN and the temporal features extracted by GNN to obtain the comprehensive feature representation of each shot;

[0035] The strategy adjustment module is used to dynamically optimize the shot transition strategy using the Deep Q-Network (DQN) based on the visual and temporal features of the shots.

[0036] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the logic evaluation method based on news shot transition are implemented.

[0037] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the logic evaluation method based on news shot transition are implemented.

[0038] Advantages of the present invention: The logic evaluation method based on news shot transition provided by the present invention solves these technical problems by introducing a method combining deep learning and reinforcement learning. In particular, the combination of convolutional neural network and graph neural network enables the system to consider both the visual content and temporal dependencies of the shots, and optimize the shot transition strategy. In addition, by introducing the Deep Q-Network, the system can dynamically adjust the shot transition strategy according to the user's real-time feedback, improving the smoothness and content coherence of shot transitions. This enables flexibility, intelligence, and real-time optimization during shot transitions. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0040] Figure 1 It is the overall flowchart of a logic evaluation method based on news shot transition provided by the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0042] Example 1, referring to Figure 1 , an embodiment of the present invention provides a logical evaluation method based on news shot switching, including:

[0043] S1: Extract video data from the original news video stream; detect shot switching points in the video and analyze to obtain shot switching rules.

[0044] Furthermore, the detection of shot switching depends on the similarity metric between adjacent frames. Calculate the similarity between the current frame and the previous frame and determine whether a shot switch has occurred. The similarity calculation formula uses the Structural Similarity Index (SSIM) to measure the similarity between two frames. The specific formula is expressed as:

[0045]

[0046] Where: and represent two frames of images respectively; and are the average brightness of the two frames of images; and are the brightness variances of the two frames of images; is the covariance of the two frames of images; and are small constants added for stability. Usually choose:

[0047]

[0048]

[0049] Where and are constants, is the dynamic range of image pixels (for example, for 8-bit images, ). When the similarity is lower than the set threshold, it is determined that a shot switch has occurred, and the current time point is marked as a shot switching point.

[0050] Precisely identify the occurrence of shot transitions based on the similarity between adjacent frames, and further analyze the duration of shots, transition frequency, and other relevant features. The duration of each shot is determined by the time difference between the transition time point of the current shot and the transition time point of the next shot.

[0051] Shot transition rules include analyzing the duration, frequency, and sequence of shot transitions in the video based on the shot transition points, which are used as characteristic values for news shot transitions.

[0052] Indicates the start time of the th shot. Indicates the start time of the th shot. Indicates the start time of the th shot. The formula for calculating the shot duration is:

[0053]

[0054] Where, Indicates the duration of shot , and are the start times of the current shot and the next shot respectively.

[0055] The shot transition frequency calculation reflects the number of shot transitions per unit time. The frequency is determined by calculating the ratio of each shot transition point in the video to the total duration. Indicates the total number of shot transitions. Indicates the total duration of the video. Indicates the shot transition frequency. The formula for calculating the shot transition frequency is:

[0056]

[0057] Where, is the total number of shot transitions, is the total duration of the video, is the shot transition frequency.

[0058] The shot transition frequency reflects the number of shot transitions per unit time. A higher transition frequency may lead to a too fast pace, while a lower frequency may make the content seem dull. Adjust parameters such as the smoothness and duration of shot transitions, select shot types, and regulate the transition rhythm.

[0059] Furthermore, according to the news type, the three characteristic values are adjusted respectively, and the adjusted three characteristic values are used to represent the shot switching rules. For the real-time reporting type, the shot switching frequency is increased and the shot duration is shortened; for the news analysis type, the shot switching frequency is reduced and the shot duration is extended; when switching from the news commentary shot to the live scene shot, the shot switching duration is extended.

[0060] Define shot switching rules based on feature values. These rules are based on the duration, frequency, and content type of shot switching. The specific rules are as follows:

[0061] Rule 1: Minimum interval between lens switching When switching from a "long shot" to a "close shot", the time interval Should not be less than seconds. When switching from "close-up" to "close-up", the time interval Should not be less than Second.

[0062] Rule 2: Lens switching frequency. If the lens switching frequency per unit time Greater than a threshold , then the shot duration It should be extended to prevent the camera from switching too frequently.

[0063] Rule 3: The association between shot switching and content logic.

[0064] In news reporting, the camera switching should take into account the continuity of the content. For example, when switching from the reporter's explanation to the live scene, too fast switching should be avoided.

[0065] Use graph neural networks to model temporal dependencies and combine image features to judge the rationality of shot switching. For example, the switch from "news figures" to "background" shots should be based on a specific temporal structure. Use similar judgment criteria to analyze the coherence of shots and the naturalness of switching.

[0066] The rules are adjusted based on the characteristics of news reports for quantitative representation, and the rules are further adjusted dynamically according to the type of news.

[0067] Rule adjustment (camera switching frequency): For sports news, the camera switching frequency Increase, threshold Can be set to a higher value. For in-depth news analysis, the frequency of camera switching should be reduced, and the threshold Can be set to a lower value (e.g., 1 time / second). Quantization expression:

[0068]

[0069] Rule Adjustment (Shot Length): In sports coverage, shot length can be shortened. In in-depth analysis news, the shot duration should be extended (e.g., set to 5 seconds). Quantitative adjustment:

[0070]

[0071] Among them, Scale is the scaling factor corresponding to the news type (for example, 0.8 for sports news and 1.2 for in-depth analysis news).

[0072] Among them, represents the news shot switching interval threshold, represents the fine switching interval threshold, represents the maximum allowable shot switching frequency per unit time; the fluency is evaluated through the shot switching duration and frequency, and the coherence is evaluated through the shot switching order and similarity index.

[0073] It should be noted that according to the type of news report (such as sports report, breaking news or in-depth news analysis), the shot switching rules are dynamically adjusted. For sports reports, increase the shot switching frequency and appropriately shorten the shot duration. In-depth news analysis reduces the shot switching frequency and extends the shot duration to make the content more immersive and coherent. This process is completed by a deep learning model by adjusting the feature values, combined with a feedback mechanism, enabling the system to automatically optimize according to different news scenarios.

[0074] S2: Use a convolutional neural network to extract the image features of shot switching, model the temporal dependence relationship through a graph neural network, and fuse to obtain the switching features.

[0075] Furthermore, in order to improve the understanding and optimization of shot switching, a method of graph vision fusion is introduced, combining the convolutional neural network CNN to extract image features and the graph neural network GNN to model the temporal dependence between shots. CNN extracts image features and inputs each frame of image. Output the deep features of each frame of image . CNN extracts the low-level and high-level features of each frame of image through multiple convolutional and pooling operations, including texture, edge and color information. Input the image features of each shot and the switching temporal relationship. Output the temporal dependence representation between shots .

[0076] Extracting the image features of the shot includes using the convolutional neural network CNN to process each frame of image, extracting the image features, texture, edge, color distribution, to describe the visual content of the shot; inputting each frame of image, extracting local features through the convolutional layer, and learning the edge, color, texture information in the image.

[0077] Perform downsampling operations through the pooling layer and enhance non-linear features through the ReLU activation function; finally, output the feature map , including the image features of each frame of the shot.

[0078] Combining the graph neural network GNN and the convolutional neural network CNN, a method of graph-visual fusion is introduced, which combines the convolutional neural network CNN to extract the visual features of the shots and the graph neural network GNN to model the temporal dependencies and graph structure information between the shots. The goal is to comprehensively improve the understanding ability of shot transitions, not only relying on single image features, but also considering the temporal relationship and overall structure of shot transitions.

[0079] Use the convolutional neural network CNN to process each frame of the image and extract the low-level and high-level features of the image. These features include texture, edges, color distribution, etc., and are mainly used to describe the visual content of the shot.

[0080] Input each frame of the image. The convolutional layer extracts local features through the convolutional layer and learns information such as edges, colors, and textures in the image. The pooling layer performs downsampling operations to reduce the size of the feature map while retaining the most important features. Enhance non-linear features through the ReLU activation function.

[0081] CNN output: The finally output feature map contains the visual features of each frame of the shot.

[0082] Obtain the temporal dependency relationships, including using the graph neural network GNN to model the dependencies between the shots; the shot image features extracted by the CNN , are input into the GNN as graph node features.

[0083] Define that the nodes of the graph represent the shots, the temporal dependency relationships between each shot transition point are used as the edges in the graph, and the order, duration, and frequency of the shot transitions are obtained through weighted summation as the weights of the edges; construct the adjacency matrix of the graph to represent the dependency relationships between the shots; the elements of the adjacency matrix represent the shot nodes and the shot nodes the strength of the dependency relationship between them.

[0084] The graph neural network propagates information through the edges of the graph, captures the temporal relationships of shot transitions, and generates new temporal feature representations for each shot , for each shot node , the new temporal feature representation is updated through the weighted sum of the temporal features of the neighbor nodes and the adjacency matrix . represents the th shot at the Feature representation of the layer is the temporal feature extracted by the GNN, representing at shots at the layer's feature representation.

[0085] Input of the GNN: The image features of each shot (extracted by CNN) are input into the GNN as node features. The temporal dependencies between each shot transition point are used as the edges in the graph. The order, duration, and frequency of shot transitions will be used as the weights of the edges. Adjacency matrix: Construct the adjacency matrix of the graph to represent the dependencies between shots. Information between shots is propagated through the edges of the graph to obtain the temporal relationship between shots. The calculation formula of the graph neural network is:

[0086]

[0087] where: is the feature representation of the th shot at the layer. represents the adjacent nodes of shot (i.e., the shots related to it). is the weight matrix of the layer, is the bias term. is the activation function (such as ReLU) used to enhance the non-linear modeling ability of the network.

[0088] Combining the outputs of CNN and GNN, the image features and temporal features of each shot are fused together to obtain the shot transition feature representation as .

[0089] Through information propagation, the graph neural network can capture the temporal relationship of shot transitions and generate a new temporal feature representation for each shot, which contains the context information of shot transitions.

[0090] It should be noted that CNN is used to extract the visual features of shots, such as texture, color, edges, etc., to provide rich visual information for each shot. The graph neural network GNN models the temporal relationship: By modeling the temporal dependencies between shots through GNN, it can capture the context relationship and temporal structure of shot transitions, ensuring the rationality of shot transitions. Combining the image features extracted by CNN with the temporal features learned by GNN generates an all-round feature representation for each shot. It can not only understand the visual features of shots but also understand the temporal dependencies of shot transitions, making the switching decision more intelligent. The graph-vision fusion technology improves the perception ability of the system, can accurately grasp the temporal and content relationships of shot transitions, and avoid inappropriate switches.

[0091] S3: Combine the switching features with the shot transition rules, and use the fuzzy C-means clustering algorithm to evaluate the shot transition logic and calculate the shot transition logic score.

[0092] Furthermore, define the fuzzy C-means clustering model. The fuzzy C-means clustering FCM algorithm performs clustering by minimizing the objective function. In this step, the core objective function of FCM will be defined, and the parameters to be optimized will be determined. The objective function of the FCM algorithm is expressed as:

[0093]

[0094] where: is the membership degree of the data point in the category , indicating the degree to which the data point belongs to this category. is the center of the category . is the fuzzification factor, which controls the fuzziness of the membership degree. is the number of clustering categories, indicating the number of categories to be set in FCM. is the total number of data points, is the data point, and the fuzzification factor , which controls the fuzziness of the membership degree, usually takes the value .

[0095] Construct the objective function of fuzzy C-means clustering according to the shot transition duration, frequency, and sequence constraints and the clustering parameters; the clustering parameters include the number of clusters , the fuzzification factor and the category center ; evaluate the smoothness through the shot transition duration and frequency, and evaluate the coherence through the shot transition sequence and similarity index; evaluate the smoothness and coherence of the shot transition through the membership degree of clustering.

[0096] Genetic algorithm (GA) optimizes fuzzy C-means clustering. The genetic algorithm can be used to optimize the parameters in FCM, such as the number of clusters C, the fuzzification factor m, and the initialized category center .

[0097] Define the individuals and populations of the genetic algorithm. Among them, each individual represents a set of FCM parameters. Specifically, an individual includes: the number of clusters C, the fuzzification factor m, and the initial value of the cluster center (the center point of each category) .

[0098] Perform population initialization, randomly generate a population of a certain number of individuals, and each individual contains a set of possible parameter settings. The fitness function design goal of the genetic algorithm is to optimize the objective function of FCM through selection, crossover, and mutation operations to make the clustering results more accurate. Therefore, the fitness function will be designed based on the clustering results and objective function of FCM. Use the objective function of FCM as the fitness function: The lower the fitness value, the better the clustering effect of the current individual. Fitness function:

[0099]

[0100] where, is the objective function value of FCM. The goal is to minimize , so take its reciprocal as the fitness value.

[0101] Perform fuzzy C - means clustering using the parameters optimized by the genetic algorithm. The steps of the clustering process are as follows:

[0102] Initialize the membership matrix U and the cluster centers Initialization stage: The membership matrix U is an N × C matrix, where each element represents the membership degree of the data point belonging to the category j. The initial membership values are usually random.

[0103] Update the membership matrix Update the membership matrix according to the current cluster centers. The membership is calculated by the following formula:

[0104]

[0105] This formula updates the membership based on the distance between each data point and each cluster center. The smaller the distance, the larger the membership. Update the cluster centers According to the updated membership matrix U, update the center of each category :

[0106]

[0107] The cluster center is the weighted average of all data points in each category. Repeat updating the membership matrix and the cluster centers until the objective function converges, or reaches the maximum number of iterations.

[0108] Perform clustering using the parameters optimized by the genetic algorithm to obtain the membership degree and cluster centers of each shot transition; analyze each shot transition feature through the cluster centers and membership degrees , assign smoothness and coherence scores to the shot transitions according to the clustering results, and input the scores into the reward function database.

[0109] It should be noted that fuzzy clustering provides "soft" scores, which are more in line with the subjective experience of the audience than simple right or wrong judgments; DQN online learning can dynamically adjust according to the actual effects and automatically adapt to different news types and emergency scenarios. Overall, this method improves the viewing smoothness of the audience and the information reception efficiency, and realizes the intelligent and self-adaptive optimization of news shot transitions.

[0110] S4: Real-time feedback the shot transition logic score, and dynamically optimize the shot transition strategy using a deep Q-network according to the shot transition logic score and the transition features.

[0111] Furthermore, combining the outputs of CNN and GNN, fuse the image features and temporal features of each shot together to obtain the final shot feature representation , which not only contains the visual content of the image, but also captures the temporal dependencies and overall structural relationships between the shots.

[0112] Dynamically adjust the shot transition strategy according to the real-time feedback of the user to ensure that the shot transitions meet the expectations of the audience and enhance the viewing pleasure of the program.

[0113] Real-time user feedback During the video playback, the user, who can be a news editor or an audience, can provide feedback in real time through a graphical interface: the user can adjust the switching speed or display time of the shots. The user can select the shot type, adjust the rhythm of the shot transitions, and provide scores on the smoothness and frequency of the shot transitions.

[0114] Reinforcement learning to optimize the shot transition strategy Optimize the shot transition strategy through a deep Q-network (DQN), and dynamically adjust features such as the duration, frequency, and content coherence of the shot transitions. State representation: The state is composed of the image features and temporal features of each shot, specifically:

[0115]

[0116] Among them, are the image features extracted by CNN, are the temporal features extracted by GNN. Action representation: The action represents the decision of the shot transition, including adjusting the shot duration, switching frequency, etc. Q-value update: Use the deep Q-network to calculate the Q-value according to the state and the action , and update the strategy through the reinforcement learning algorithm:

[0117]

[0118] Among them, is the reward calculated based on user feedback, is the discount factor.

[0119] The reward function designs the reward function to feedback the quality of shot transitions and help the model learn the optimal strategy. Smoothness reward: A high reward is given when the shot transition is natural, and a low reward otherwise. Content coherence reward: A high reward is given when the shot transition conforms to the logic of the news content, and a low reward if the transition is inappropriate. Reward function design:

[0120]

[0121] Among them, represents the weight coefficient of the smoothness reward, represents the weight coefficient of the content coherence reward, which is used to balance the impact of content coherence on the total reward. FlowReward represents the smoothness reward function, which reflects the smoothness of the current shot transition. During shot transitions, smoother and more natural transitions will receive higher rewards, while abrupt or overly fast transitions will receive lower rewards. CReward represents the content coherence reward function, which evaluates whether the shot transition conforms to the content logic and news structure.

[0122] If the shot transition conforms to the logical order of the content (such as a natural transition from a news reporter to a background shot), a higher reward will be received. If the transition does not conform to the logic (such as a direct jump from a background shot to an unrelated shot), a lower reward will be received. represents the action taken at time , that is, the current shot transition decision. Actions may include adjusting the shot duration, switching frequency, and selecting a new shot type.

[0123] Through continuous training, the deep Q-network will learn the optimal shot transition strategy, automatically adjust the shot transition duration, frequency, and content coherence, making the shot transition smoother, more natural, and meeting the audience's needs.

[0124] It should be noted that it effectively improves the logical evaluation and optimization ability of news shot transitions. The improvement enables the system to not only automatically extract and analyze video data but also intelligently adjust the shot transition rules according to different news types, optimize the extraction of shot features and temporal modeling by combining deep learning techniques CNN and GNN, and optimize the shot transition strategy through a real-time feedback mechanism and reinforcement learning DQN. Ultimately, it can provide smoother, more coherent, and rhythm-compliant shot transitions, improving the audience's viewing experience.

[0125] Example 2, an embodiment of the present invention, provides a logical evaluation method based on news shot switching. To verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0126] First of all, the goal of this embodiment is to verify the effectiveness of the logical evaluation method based on news shot switching, especially how to extract the visual features of shots through a convolutional neural network (CNN), model the temporal dependence relationship between shots by combining with a graph neural network (GNN), and how to optimize the shot switching strategy using a deep Q-network (DQN). The experiment will cover the following aspects: extracting data from the original news video stream, detecting shot switching points, analyzing and obtaining shot switching rules, making rule-based judgments, and how to adjust the rules according to the characteristics of news reports.

[0127] At the beginning stage of the experiment, first extract each frame of image data from the news video and perform preprocessing. Each frame of image is fed into a convolutional neural network (CNN) to extract the visual features of the shot, such as edges, textures, color distributions, etc. Through the pooling layer and activation function, the CNN extracts the high-level features of each frame of image as the input for subsequent temporal modeling.

[0128] Next, use a graph neural network (GNN) to model the temporal dependence relationship between shots. The image features of each shot are used as the node features of the GNN, and the temporal dependence relationship between shot switching points is represented by constructing the adjacency matrix of the graph. The graph neural network will learn the context dependence between shots through the way of information propagation.

[0129] The generation of shot switching rules is based on the detected shot switching points. The duration of each shot is calculated by the time difference between the switching time point of the current shot and the switching time point of the next shot. The shot switching frequency is evaluated by calculating the number of shot switches per unit time.

[0130] For different news types (such as sports news, in-depth news analysis, real-time news), adjust the frequency and duration rules of shot switching. For example, real-time news reports may require increasing the frequency of shot switching and shortening the shot duration, while in-depth news analysis requires reducing the shot switching frequency and lengthening the shot duration. Finally, use a deep Q-network (DQN) to optimize the shot switching strategy and dynamically adjust parameters such as the duration and switching frequency of the shot according to the user's real-time feedback to improve the quality of shot switching.

[0131] A series of news video clips will be used for verification in the experiment. The goal is to test whether the proposed method can effectively optimize shot switching and improve the fluency and logic of news programs under different news types.

[0132] Through the proposed logical evaluation method based on news shot transitions, the shot transition strategies of the experimental videos have been significantly optimized under different types of news reports. In sports news, due to its relatively dynamic nature, the shot transition frequency is relatively high, and the duration is short (the average transition time is 2.5 seconds). For example, in experimental video 1, the "long shot to close-up" transition time is 2.5 seconds, the transition frequency is 3.5 times per second, the user fluency score is 9, and the content coherence score is 8, indicating that the shot transitions of this video perform well in terms of rhythm and fluency.

[0133] In contrast, in-depth news analysis programs require fewer shot transitions, and each shot has a longer duration (for example, in experimental video 2, the "close-up to extreme close-up" transition time is 3.5 seconds). This enables the audience to better focus on the content without being interrupted by frequent shot transitions. The transition frequency of experimental video 2 is relatively low (1.2 times per second), and the user fluency score is relatively high (score of 9), which proves that the system can dynamically adjust the shot transition rules according to the characteristics of news reports to adapt to different rhythm requirements.

[0134] Live news reports, such as experimental videos 3 and 6, usually require a relatively high shot transition frequency to highlight the dynamic changes of news events. In these experiments, the transition frequency is generally higher than 3 times per second, and the transition time is short (for example, in experimental video 3, the "long shot to background" transition time is 2 seconds). These settings can ensure the compactness and timeliness of the news, and at the same time obtain good scores in user feedback.

[0135] Example 3, an embodiment of the present invention, provides a logical evaluation system based on news shot transitions, including a video data extraction module, an analysis module, a graph-vision fusion module, and a strategy adjustment module.

[0136] The video data extraction module is used to extract images from the original news video stream, perform preprocessing, and input the processed images into the analysis module.

[0137] The analysis module is used to detect shot transition points and perform rule analysis based on the extracted features;

[0138] The graph-vision fusion module is used to fuse the image features extracted by CNN and the temporal features extracted by GNN to obtain a comprehensive feature representation of each shot.

[0139] The strategy adjustment module is used to dynamically optimize the shot transition strategy using the deep Q-network DQN based on the visual and temporal features of the shots.

[0140] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0141] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0142] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0143] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by a combination of any of the following techniques well known in the art: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A logical evaluation method based on news shot switching, characterized in that Including: Extracting video data from the original news video stream; Detecting the shot transition points in the video and analyzing to obtain the shot transition rules; Using a convolutional neural network to extract the image features of shot transitions, modeling the temporal dependencies through a graph neural network, and fusing to obtain the transition features; Combining the transition features with the shot transition rules, using the fuzzy C-means clustering algorithm to evaluate the shot transition logic, and calculating the shot transition logic score; Real-time feedback of the shot transition logic score, and dynamically optimizing the shot transition strategy according to the shot transition logic score and the transition features by using a deep Q-network.

2. The logical evaluation method based on news shot switching according to claim 1, wherein: The shot transition point is the transition moment between consecutive shots; The shot duration is the time difference between the start time of the current shot and the start time of the next shot; The shot transition frequency is the number of shot transitions per unit time, and the shot transition frequency is calculated by each shot transition point and the total duration in the video; By calculating the current frame and the previous frame for similarity , determine whether a shot transition has occurred; If the similarity metric is lower than a set threshold, it is determined that a shot transition has occurred, and the current time point is marked as a shot transition point.

3. The logical evaluation method based on news shot switching according to claim 2, characterized in that: The shot transition rules include analyzing the shot transition duration, frequency, and order in the video according to the shot transition points as the eigenvalue of news shot transitions; According to the news type, respectively adjusting the three eigenvalues, and using the adjusted three eigenvalues to represent the shot transition rules: for real-time reports, increasing the shot transition frequency and shortening the shot duration; for news analysis, reducing the shot transition frequency and lengthening the shot duration; when switching from a news explanation shot to a live scene shot, lengthening the shot transition duration.

4. The logic evaluation method based on news shot transitions according to claim 3, wherein: The extraction of the image features of the shot includes using a convolutional neural network CNN to process each frame of the image, extracting the image features, texture, edges, and color distribution to describe the visual content of the shot; inputting each frame of the image, extracting local features through the convolutional layer, and learning the edge, color, and texture information in the image; Perform downsampling operations through the pooling layer and enhance non-linear features through the ReLU activation function; finally, output the feature map , which contains the image features of each frame of the shot.

5. The logic evaluation method based on news shot switching according to claim 4, characterized in that: The obtained temporal dependency relationship includes using a graph neural network (GNN) to model the dependencies between shots; the shot image features extracted by the CNN , are input into the GNN as graph node features; Define that the nodes of the graph represent shots, and the temporal dependency relationships between each pair of shot transition points are used as the edges in the graph. The order, duration, and frequency of shot transitions are weighted and summed to obtain the result as the weight of the edge; construct the adjacency matrix of the graph represent the dependency relationships between shots; the elements of the adjacency matrix represent the shot node and the shot node the strength of the dependency relationship between them; The graph neural network propagates information through the edges of the graph, captures the temporal relationship of shot transitions, and generates a new temporal feature representation for each shot. For each shot node The new temporal feature representation is updated by the weighted sum of the temporal features of neighbor nodes and the adjacency matrix ; is the feature representation of the th shot at the th layer; is the feature representation of the th shot at the th layer; Combining the outputs of CNN and GNN, the image features and temporal features of each shot are fused together to obtain the shot transition feature representation as .

6. The logic evaluation method based on news shot transitions according to claim 5, wherein: The evaluation of the shot transition logic includes constructing an objective function for fuzzy C-means clustering based on shot transition duration, frequency, and order. And clustering parameters; the clustering parameters include the number of clusters , the fuzzification factor And the class center ; evaluate the smoothness through shot transition duration and frequency, and evaluate the coherence through shot transition order and similarity index; evaluate the smoothness and coherence of shot transitions through the membership degree of clustering. Optimize the clustering parameters using a genetic algorithm, and use the objective function as the fitness function to generate multiple individuals, where each individual represents a set of clustering parameters; optimize the combination of clustering parameters through the selection, crossover, and mutation operations of the genetic algorithm; Execute clustering using the parameters optimized by the genetic algorithm to obtain the membership degree and class center of each shot transition; analyze each shot transition feature through the cluster center and membership degree According to the clustering results, assign fluency and coherence scores to the shot transitions, and input the scores into the reward function database.

7. The logical evaluation method based on news shot switching according to claim 6, wherein: The dynamic optimization of the shot transition strategy includes optimizing the shot transition strategy through a deep Q-network DQN and dynamically adjusting the duration, frequency, and content coherence features of the shot transition; Define the state space representation, where the state is composed of the image features and temporal features of each shot, , representing the state at the th moment; Among them, represents the image features extracted by CNN; Define the action space, and the action is the decision for shot transition, indicating the action taken at the moment, including adjusting the shot duration and switching frequency; perform Q-value update, use the deep Q-network to calculate the Q-value according to the state and the action and update the policy through the reinforcement learning algorithm, represents the discount factor; Define the reward function Feedback the shot transition strategy, denote the reward at the moment; obtained by the weighted sum of the fluency reward and the content coherence reward; through training iterations, the deep Q-network will learn the optimal shot transition strategy and automatically adjust the fluency and content coherence of shot transitions.

8. A system adopting the logic evaluation method based on news shot switching as described in any one of claims 1 to 7, characterized in that: Including a video data extraction module, an analysis module, a graph vision fusion module, and a strategy adjustment module; The video data extraction module is used to extract images from the original news video stream, perform preprocessing, and input the processed images into the analysis module; The analysis module is used to detect the shot transition points and perform rule analysis according to the extracted features; The graph vision fusion module is used to fuse the image features extracted by the CNN and the temporal features extracted by the GNN to obtain the comprehensive feature representation of each shot; The strategy adjustment module is used to dynamically optimize the shot transition strategy according to the visual features and temporal features of the shot by using a deep Q-network DQN.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the logic evaluation method based on news shot transitions according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the logic evaluation method based on news shot transitions according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • News division method and apparatus

    CN108551584A

  • Lens switching method and device, equipment, medium and program product

    CN116943176A

  • Switching method and device of dual-lens motion camera and computer equipment

    CN118574002A

  • Lens automatic switching method and system based on intelligent light sensing test

    CN118972696A

  • Intelligent MV generation method, system and device based on AIGC and medium

    CN119788886A