A news intelligent broadcast system and method based on deep learning
Through deep learning technology and dynamic event tree construction, semantic alignment and multi-scale modeling of multi-modal news data are achieved, solving the shortcomings of existing news broadcast systems in multi-modal data processing and user personalized needs, and improving the real-time and accuracy of news broadcasts.
Patent Information
- Application Number
- CN202510339248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing news broadcasting system has significant limitations in handling multimodal data, complex event modeling, user personalized needs and content generation hierarchy, and it is difficult to achieve semantic alignment and dynamic tracking of multimodal data, resulting in a lack of real-time, accuracy and personalization of broadcasting content.
Deep learning technology is adopted, combining dynamic time convolution networks and hierarchical Transformer architecture, and through cross-modal comparison learning and multi-head self-attention mechanism, a dynamic event tree is built to realize semantic alignment and multi-scale modeling of multi-modal data, and to generate multi-modal broadcast content through three-stage generation networks, and optimize the generation rules based on user feedback.
It improves the real-time, content accuracy and personalized adaptability of news broadcasts, and can quickly generate breaking news summary, detailed content and predictive broadcasts, meets users' needs for rapid response to news and future trend insights, and optimizes user experience.
Smart Images

Figure CN119854545B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and news broadcasting, and particularly relates to a news intelligent broadcasting system and method based on deep learning. Background Art
[0002] With the rapid development of Internet and artificial intelligence technologies, the way of news dissemination has undergone multiple revolutions from traditional media to digital media and then to intelligent media. At present, news broadcasting has gradually evolved from a single text or voice form to a comprehensive dissemination form integrating multi-modal information, including the combination of various forms such as text, images, and videos, to meet the user's requirements for real-time, diverse, and accurate information acquisition. However, existing news broadcasting systems still have obvious technical limitations when dealing with multi-modal data processing, complex event modeling, and user personalized needs.
[0003] Currently, many news broadcasting systems rely on template or rule-driven methods for content generation. Although this method can meet basic automation requirements, it shows significant limitations when dealing with complex news events or multi-modal data. Specifically, the templated method usually cannot flexibly adapt to the dynamic development of events, and the generated broadcast content is often rigid and lacks the ability to be adjusted in real time. In addition, when facing multi-modal data such as images and videos, existing technologies are difficult to achieve semantic alignment and deep integration of data, resulting in a low semantic correlation between the generated content and the actual news event, and it is difficult to meet the user's requirements for the authenticity and diversity of news broadcasting.
[0004] Complex news events usually have the characteristics of multi-stage dynamic evolution, including multiple links such as event occurrence, situation development, public reaction, and subsequent processing. This dynamic change in the time dimension poses higher modeling requirements for news broadcasting. However, most existing technologies adopt static modeling methods and cannot comprehensively capture the time dynamic characteristics of news events. For example, for the broadcast of breaking news, existing methods are often limited to quickly extracting the basic information of the event, but lack the ability to dynamically track the development of the event and update semantics in subsequent reports. This defect makes the broadcast content unable to reflect the latest progress of the event, reducing the timeliness and coherence of the broadcast.
[0005] The processing ability of multi-modal data is also an important technical bottleneck faced by existing news broadcast systems. News content is usually presented in the forms of text, images, and videos, etc. Each type of modal data contains specific semantic information. However, when dealing with multi-modal data, existing technologies often adopt independent analysis methods and lack a unified semantic modeling and alignment mechanism for different modal data. Due to the different characteristics of each modal data, there may be differences or conflicts in semantic information between modalities. Existing systems are difficult to solve these problems through effective fusion methods, resulting in the generated news content lacking relevance and consistency between multi-modalities, further affecting the accuracy and expressiveness of the broadcast.
[0006] User interaction feedback is an important part of the news intelligent broadcast system and can directly reflect users' attention and preferences for news content. However, existing technologies are insufficient in using user feedback to optimize the generated content. Although some systems can sort popular content through simple user behavior analysis (such as the number of clicks and browsing duration), they cannot effectively integrate users' real-time interaction data into the generation model for dynamically adjusting the broadcast content. The lack of in-depth utilization of user feedback makes it difficult for news broadcasts to meet users' personalized needs and also unable to improve the accuracy of content generation.
[0007] In addition, existing news broadcast systems also have deficiencies in the hierarchical and predictive capabilities of content generation. News events usually include core information (such as the subject of the event, time clues) and background information (such as relevant historical events or peripheral details). However, existing systems lack effective discrimination of semantic levels when generating broadcast content, resulting in the generated content being lengthy and lacking focus. At the same time, the predictive broadcast ability for the future development trend of events is weak. Existing technologies often only provide descriptions of the current state of events and lack reasonable speculation on possible future evolution directions, which further limits the practicality of the broadcast system.
[0008] Therefore, how to provide a news intelligent broadcast system and method based on deep learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0009] An object of the present invention is to propose a news intelligent broadcast system and method based on deep learning. The present invention combines deep learning technology, a multi-modal data alignment mechanism, and a dynamic event modeling method to comprehensively achieve the intelligence and high efficiency of news broadcasts. By introducing a dynamic time convolutional network and a hierarchical Transformer architecture, the time features and semantic features of news content are accurately extracted. Cross-modal contrast learning is used to achieve semantic alignment of multi-modal data, and the modeling ability of multi-stage evolution of news events is improved by constructing a dynamic event tree. The system has strong real-time performance, accurate content, rich levels, and personalized adaptation ability, and can be widely applied in the fields of intelligent media and online news services, greatly optimizing the user experience.
[0010] A news intelligent broadcast method based on deep learning according to an embodiment of the present invention includes the following steps:
[0011] S1. Collect text data, image data, and video data of news events, construct a multi-modal news dataset and perform semantic encoding to generate preliminary semantic features;
[0012] S2. Based on a dynamic time convolutional network, perform feature extraction on the preliminary semantic features in the time dimension, and perform semantic grading on the news content through a hierarchical Transformer architecture to generate time features and hierarchical semantic features;
[0013] S3. Through event-characteristic-driven cross-modal contrast learning, align the time features and hierarchical semantic features with the multi-modal news data to generate cross-modal fusion features;
[0014] S4. Use a multi-head self-attention mechanism to perform multi-scale semantic modeling on the cross-modal fusion features, and construct a dynamic event tree based on time nodes and semantic relevance to generate a multi-scale dynamic semantic structure;
[0015] S5. Input the multi-scale dynamic semantic structure into an improved three-stage generation network to sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution, and push them to the user terminal in real time;
[0016] S6. Based on user interaction feedback information, use an event feedback optimization mechanism to dynamically adjust the node weights and semantic generation rules of the dynamic event tree to generate optimized broadcast content;
[0017] S7. Synchronously output the optimized broadcast content in multi-modal forms of text, voice, and video.
[0018] Optionally, the S2 specifically includes:
[0019] S21. Perform embedding encoding on the preliminary semantic features, and embed feature vectors through an improved position encoding method:
[0020] ;
[0021] Wherein, represents the partial code value of the pos-th feature vector on the even-dimensional index 2q, represents the partial code value of the pos-th feature vector on the odd-dimensional index 2q + 1, pos represents the position index of the feature vector, exp represents the exponential function, t represents the time node, d represents the dimension of the embedded feature vector, represents the time decay factor;
[0022] S22. Use a dynamic time convolutional network to extract features in the time dimension from the embedded encoded features, and capture the time features in the irregular time series through dynamically adjusted convolutional kernels:
[0023] ;
[0024] Among them, represents the output feature value at time step t, k represents the size of the convolutional kernel, represents the weight parameter of the convolutional kernel, represents the input feature value at time step t - r, represents the time node difference, and b represents the bias term;
[0025] S23. Construct the first layer of the hierarchical Transformer, and use the multi-head self-attention mechanism to perform semantic encoding on the time features extracted by the dynamic time convolutional network, and extract the core semantic features in the news content. The core semantics include the subject, keywords, and time node features of the news event:
[0026] ;
[0027] Among them, represents the core semantic feature, softmax represents the normalization function, represents the query vector of the core semantic feature, represents the key vector of the core semantic feature, represents the transpose of the key vector of the core semantic feature, represents the value vector of the core semantic feature, represents the weight matrix of the time feature, represents the normalization factor;
[0028] S24. Construct the second layer of the hierarchical Transformer, perform background semantic decoding on the core semantic features, and extract secondary semantic features:
[0029] ;
[0030] Among them, represents the secondary semantic feature, represents the query vector of the secondary semantic feature, represents the key vector of the secondary semantic feature, represents the transpose of the key vector of the secondary semantic feature, represents the value vector of the secondary semantic feature, and represent the weight parameters, represents the positional encoding of the background semantics;
[0031] S25. Generate temporal features and hierarchical semantic features through a multi - dynamic weighting mechanism of the temporal dimension and semantic levels based on core semantic features and secondary semantic features:
[0032] ;
[0033] Among them, represents the temporal features and hierarchical semantic features, represents the weight matrix of the core semantic features, represents the weight matrix of the secondary semantic features, represents the temporal features.
[0034] Optionally, the event - characteristic - driven cross - modal contrast learning in S3 specifically includes: obtaining a set of event - characteristic information, including time clues, semantic entities, and background features, as guiding signals for feature alignment; based on cross - modal contrast learning, aligning the temporal features and hierarchical semantic features with multi - modal news data.
[0035] Time clues: refer to the time nodes or time periods when news events occur, such as timestamps, time sequences, event development stages, etc. These information are used to mark and organize the temporal dimension of news data, providing a temporal reference for subsequent alignment. Time clues help the feature alignment model gather data with similar temporal characteristics (such as images and texts at different time points of the same event) into a unified temporal semantic space.
[0036] Semantic entities: refer to the core objects or entities of news events, such as relevant people, places, or themes of a certain news event. These semantic entities are the core information focused on during the alignment process. By extracting semantic entities, the focus on the core content during the alignment process can be enhanced, and irrelevant or noisy information can be ignored.
[0037] Background features: refer to the context information surrounding news events, such as the background environment mentioned in the news, relevant historical events, or supplementary descriptive information. The introduction of background features can provide broader semantic support for alignment. Especially when dealing with multi - modal data (such as video or picture backgrounds), it helps to improve the accuracy and integrity of alignment.
[0038] As guiding signals for feature alignment: After time clues, semantic entities, and background features are extracted, they form a set of multi - dimensional guiding signals (usually called event - characteristic vectors). These vectors are used to guide cross - modal feature alignment, enabling data from different modalities (text, image, and video) to find common reference points in the semantic space.
[0039] Cross - modal contrast learning is a machine - learning method aimed at comparing and matching features of different modalities (such as text, image, video, etc.).
[0040] Core idea: By calculating the similarity of feature vectors from different modalities, data with high semantic correlation is screened out, and low-correlated data is ignored.
[0041] In this solution, the time features and hierarchical semantic features come from the dynamic event tree respectively, while the multi-modal news data comes from different modalities (such as videos, images, and texts).
[0042] Comparison method: By comparing the time features with the hierarchical semantic features, and comparing these features with the data of text, image, and video modalities, the model can learn the semantic alignment relationships between modalities.
[0043] Alignment of time features and hierarchical semantic features:
[0044] Time features: Derived from the time nodes of the dynamic event tree, such as the time sequence of news and the development stages of events. Through alignment, it is ensured that cross-modal data is consistent in the time dimension. For example, image, text, and video content in the same time period can be matched to the corresponding time features.
[0045] Hierarchical semantic features: Derived from the semantic levels (core layer, sub-core layer, and background layer) of the dynamic event tree. Through alignment, the model can associate data of different modalities to the corresponding semantic levels. For example, key semantic information corresponds to the main scene in the image, the core description in the text, etc.
[0046] Alignment of multi-modal news data:
[0047] The alignment process includes: projecting the time features, hierarchical semantic features, and multi-modal news data into a unified semantic space; calculating the semantic similarity between the time features, hierarchical semantic features, and multi-modal data; screening out the most relevant feature combinations to form cross-modal fusion features.
[0048] Optionally, the specific steps of S4 include:
[0049] S41. Divide the cross-modal fusion features according to time nodes to form a feature set of several time segments , where represents the time node corresponding feature subset;
[0050] S42. Use the multi-head self-attention mechanism to perform semantic modeling on the feature subsets within each time node, extract local context associations, and obtain preliminary time node features;
[0051] S43. Compare the different preliminary time node features pairwise, define the semantic correlation degree between time nodes, and the semantic correlation degree is calculated by the weighted combination of feature similarity and time difference:
[0052] ;
[0053] Among them, represents the semantic correlation degree between time nodes, A represents the dimension of the feature vector, and time node The represents the time node The eigenvalue in the k-th dimension, represents the time node The eigenvalue in the k-th dimension, exp represents the exponential function, represents the time node and time node The time difference between them, represents the time decay parameter;
[0054] S44. Construct a dynamic event tree T=(N, E) based on each time node and semantic correlation degree, where represents the set of time nodes, the edge set , represents the semantic correlation degree threshold;
[0055] S45. Hierarchically divide the time nodes in the dynamic event tree:
[0056] ;
[0057] Among them, represents the level to which the time node belongs, represents the semantic contribution degree of the time node , represents the semantic importance threshold, represents the correlation degree stratification threshold, represents the time interval threshold, n represents the total number of time nodes;
[0058] S46. Iteratively update the dynamic event tree, calculate and correct the edge weights and hierarchical divisions, and finally generate a multi-scale dynamic semantic structure.
[0059] Optionally, the specific content of S5 includes:
[0060] S51. Initialize an improved three-stage generation network, including a summary generation module, a detail generation module, and a prediction generation module, and load the core layer features, sub-core layer features, and background layer features in the multi-scale dynamic semantic structure;
[0061] S52. In the summary generation module, extract elements from the core layer features to generate a summary broadcast content of urgent and breaking news. The summary generation module includes a sequence encoder and an attention decoder, and outputs a summary broadcast sequence in the first stage;
[0062] S53. In the detail generation module, receive the summary broadcast sequence output by the summary generation module, the sub-core layer features and background layer features in the multi-scale dynamic semantic structure, and integrate them into the multi-modal decoder. Supplement the details of the event subject, background, and context through the fine-grained attention mechanism, and output the detail broadcast sequence in the second stage;
[0063] S54. In the prediction generation module, combine the summary broadcast sequence, the detail broadcast sequence, and the temporal feature information of the multi-scale dynamic semantic structure. Capture the future trend of event evolution through the recursive temporal network, and output the predictive broadcast content based on event evolution;
[0064] S55. Integrate the summary broadcast content, the detail broadcast content, and the predictive broadcast content based on event evolution into a unified multi-modal output buffer, and perform cross-validation to filter out redundant and contradictory content;
[0065] S56. Output the cross-validated broadcast content to the user terminal through the real-time push channel.
[0066] Optionally, the specific steps of S6 are as follows:
[0067] S61. Collect user interaction feedback information, including the number of user clicks, the duration of stay, the result of comment sentiment analysis, and sharing behavior, to form a user feedback matrix , where represents the comprehensive feedback score of user a at time node :
[0068] ;
[0069] Among them, represents the number of clicks of user a at time node , MaxClicks represents the maximum number of clicks of all users, represents the time that user a stays at time node , MaxTime represents the maximum stay time of all users, represents the comment sentiment score of user a at time node , , and represent the weight parameters of the feedback indicators;
[0070] S62. Dynamically update the weights of each time node of the dynamic event tree based on the user feedback matrix:
[0071] ;
[0072] Among them, represents the time node The updated weight, represents the balance parameter, represents the time node The original weight, represents the time node The average feedback value;
[0073] S63. Adjust the semantic generation rules of the dynamic event tree according to the updated weight, and recalculate the semantic contribution degree of each node:
[0074] ;
[0075] Among them, represents the time node The semantic contribution degree, represents the time node The updated weight, n represents the total number of time nodes, tanh represents the hyperbolic tangent function, ρ represents the amplification coefficient of feedback change, represents the time node The reference feedback value;
[0076] S64. Dynamically optimize the hierarchical structure of the dynamic event tree based on the updated weight and semantic contribution degree, perform pruning operations, and re-divide the core layer and the background layer;
[0077] S65. Use the updated time node weight and semantic generation rules to regenerate the summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution, so that the semantics of the generated content are consistent with the user feedback;
[0078] S66. Push the optimized broadcast content to the user terminal, and record the new round of user feedback information in real time to form a closed-loop optimization mechanism for user feedback and dynamic event tree adjustment.
[0079] A news intelligent broadcast system based on deep learning according to an embodiment of the present invention includes the following modules:
[0080] The data acquisition and encoding module is used to collect the text data, image data, and video data of news events, construct a multi-modal news data set, and perform semantic encoding on the multi-modal news data to generate preliminary semantic features;
[0081] The time and semantic feature extraction module is used to extract the time-dimensional features of the preliminary semantic features based on the dynamic time convolution network, and perform semantic grading on the news content through the hierarchical Transformer architecture to generate time features and hierarchical semantic features;
[0082] A cross-modal alignment module, which is used to align temporal features and hierarchical semantic features with multi-modal news data through event-feature-driven cross-modal contrast learning, and generate cross-modal fusion features;
[0083] A dynamic event tree generation module, which is used to perform multi-scale semantic modeling on the cross-modal fusion features by using the multi-head self-attention mechanism, construct a dynamic event tree based on time nodes and semantic relevance, and generate a multi-scale dynamic semantic structure;
[0084] A three-stage content generation module, which is used to input the multi-scale dynamic semantic structure into an improved three-stage generation network, and sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution;
[0085] An event feedback optimization module, which is used to dynamically adjust the node weights and semantic generation rules of the dynamic event tree based on user interaction feedback information by using an event feedback optimization mechanism, and generate optimized broadcast content;
[0086] A multi-modal output module, which is used to synchronously output the optimized broadcast content in multi-modal forms of text, speech, and video, and transmit it to the user terminal through a real-time push channel.
[0087] The beneficial effects of the present invention are as follows:
[0088] First of all, by introducing a dynamic time convolutional network and a hierarchical Transformer architecture, the present invention improves the system's ability to extract temporal dimension features of news content. The dynamic time convolutional network can accurately capture temporal features in irregular time series and extract multi-stage dynamic changes of news events. The hierarchical Transformer architecture performs semantic grading on news content, distinguishing core semantic features (such as news subjects, keywords, and time clues) from secondary semantic features (such as backgrounds and auxiliary information), and realizing hierarchical expression of news content. This hierarchical processing method makes the generated broadcast content not only highlight the key points but also take into account background details, and can meet the needs of users to quickly obtain core information and deeply understand the news background.
[0089] Secondly, the system realizes semantic alignment and multi-scale modeling of multi-modal news data through event feature-driven cross-modal contrast learning and multi-head self-attention mechanism. Cross-modal contrast learning uses time clues, semantic subjects, and background features as guiding signals to project the features of multi-modal data such as text, images, and videos into a unified semantic space, significantly improving the semantic consistency between different modal data. The multi-head self-attention mechanism further performs multi-scale semantic modeling on the cross-modal fusion features, effectively capturing the dynamic changes of news content at time nodes and the relevance between semantic levels. The dynamic event tree constructed based on time nodes and semantic relevance can not only clearly present the evolution path of news events but also provide high-quality semantic input for subsequent generation tasks.
[0090] In addition, the improved three-stage generation network designed in the present invention significantly enhances the hierarchical nature and diversity of news broadcast content. Through the progressive processing of the summary generation module, detail generation module, and prediction generation module, the system can generate rapid summary broadcasts of emergency events, complete news details, and predictive broadcasts based on the event evolution trend, respectively. This generation mode ensures the comprehensiveness and forward-looking nature of the broadcast content, while meeting the user's rapid response needs for breaking news and the need to gain insights into future trends. Especially in the prediction generation stage, the system combines the time node information of the dynamic event tree and the recursive time network to achieve reasonable speculation on the future development of events, filling the gap in the generation of predictive content in existing news broadcast technologies.
[0091] The introduction of user interaction feedback further enhances the dynamic adaptability and personalization ability of the system. The present invention constructs a user feedback matrix by collecting interaction data such as the number of user clicks, dwell time, emotional comments, and sharing behaviors, and uses the event feedback optimization mechanism to dynamically adjust the node weights and semantic generation rules of the dynamic event tree. This closed-loop optimization mechanism enables the generated broadcast content to adapt to the changing interests and personalized needs of users in real time, significantly improving the accuracy of the content and user satisfaction.
[0092] Finally, the present invention realizes the synchronous output of multi-modal content, supporting broadcasts in various forms such as text, speech, and video. This multi-modal output mechanism can not only enhance the expressiveness of news broadcasts but also provide users with diverse choices, further optimizing the user experience. At the same time, the real-time push function ensures that users can obtain news content in the first place, meeting the high requirements for the timeliness of news broadcasts. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0094] Figure 1 The overall flowchart of a news intelligent broadcast method based on deep learning proposed by the present invention;
[0095] Figure 2 The structural schematic diagram of a news intelligent broadcast system based on deep learning proposed by the present invention. Detailed implementation manners
[0096] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0097] Refer to Figure 1 , a news intelligent broadcast method based on deep learning, comprising the following steps:
[0098] S1. Collect the text data, image data and video data of news events, construct a multi-modal news data set and perform semantic encoding to generate preliminary semantic features;
[0099] S2. Based on a dynamic time convolutional network, perform feature extraction on the preliminary semantic features in the time dimension, and perform semantic grading on the news content through a hierarchical Transformer architecture to generate time features and hierarchical semantic features;
[0100] S3. Through event characteristic-driven cross-modal contrast learning, align the time features and hierarchical semantic features with the multi-modal news data to generate cross-modal fusion features;
[0101] S4. Use a multi-head self-attention mechanism to perform multi-scale semantic modeling on the cross-modal fusion features, and construct a dynamic event tree based on time nodes and semantic relevance to generate a multi-scale dynamic semantic structure;
[0102] S5. Input the multi-scale dynamic semantic structure into an improved three-stage generation network to sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution, and push them to the user terminal in real time;
[0103] S6. Based on the user interaction feedback information, use an event feedback optimization mechanism to dynamically adjust the node weights and semantic generation rules of the dynamic event tree to generate optimized broadcast content;
[0104] S7. Synchronously output the optimized broadcast content in multi-modal forms of text, voice and video.
[0105] In this embodiment, the S2 specifically includes:
[0106] S21. Perform embedding encoding on the preliminary semantic features, and embed the feature vectors through an improved position encoding method:
[0107] ;
[0108] Among them, represents the partial code value of the pos-th feature vector on the even-dimensional index 2q, represents the partial code value of the pos-th feature vector on the odd-dimensional index 2q + 1, pos represents the position index of the feature vector, exp represents the exponential function, t represents the time node, and d represents the dimension of the embedded feature vector, represents the time decay factor;
[0109] S22. Use a dynamic time convolutional network to extract features in the time dimension from the embedded encoded features, and capture the time features in the irregular time series through dynamically adjusted convolutional kernels:
[0110] ;
[0111] Among them, represents the output feature value at time step t, k represents the size of the convolutional kernel, represents the weight parameter of the convolutional kernel, represents the input feature value at time step t - r, represents the time node difference, and b represents the bias term;
[0112] S23. Construct the first layer of the hierarchical Transformer, and use the multi-head self-attention mechanism to perform semantic encoding on the time features extracted by the dynamic time convolutional network, and extract the core semantic features in the news content. The core semantics include the subject, keywords, and time node features of the news event:
[0113] ;
[0114] Among them, represents the core semantic feature, softmax represents the normalization function, represents the query vector of the core semantic feature, represents the key vector of the core semantic feature, represents the transpose of the key vector of the core semantic feature, represents the value vector of the core semantic feature, represents the weight matrix of the time feature, represents the normalization factor;
[0115] S24. Construct the second layer of the hierarchical Transformer, and perform background semantic decoding on the core semantic features to extract secondary semantic features:
[0116] ;
[0117] Among them, represents the secondary semantic feature, the query vector representing the secondary semantic feature, the key vector representing the secondary semantic feature, the transpose of the key vector representing the secondary semantic feature, the value vector representing the secondary semantic feature, and represents the weight parameter, the positional encoding representing the background semantics;
[0118] S25. Generate temporal features and hierarchical semantic features through a multi - dynamic weighting mechanism of the time dimension and semantic hierarchy based on the core semantic feature and the secondary semantic feature:
[0119] ;
[0120] Among them, represents the temporal features and hierarchical semantic features, the weight matrix representing the core semantic feature, the weight matrix representing the secondary semantic feature, represents the temporal feature.
[0121] In this embodiment, the event - characteristic - driven cross - modal contrast learning in S3 specifically includes: obtaining an event - characteristic information set, including time clues, semantic subjects, and background features, as a guiding signal for feature alignment; based on cross - modal contrast learning, aligning the temporal features and hierarchical semantic features with multi - modal news data.
[0122] In this embodiment, S4 specifically includes:
[0123] S41. Divide the cross - modal fusion features according to time nodes to form a feature set of several time segments , where represents the time node the corresponding feature subset;
[0124] S42. Use the multi - head self - attention mechanism to perform semantic modeling on the feature subsets within each time node, extract local context associations, and obtain preliminary time - node features;
[0125] S43. Compare the different preliminary time - node features pairwise, define the semantic association degree between time nodes, and the semantic association degree is calculated by the weighted combination of feature similarity and time difference:
[0126] ;
[0127] Among them, represents the time node and the semantic correlation degree with the time node where A represents the dimension of the feature vector represents the time node the eigenvalue at the k-th dimension represents the time node the eigenvalue at the k-th dimension, exp represents the exponential function represents the time node and the time node the time difference between them represents the time decay parameter;
[0128] S44. Construct a dynamic event tree T=(N, E) based on each time node and the semantic correlation degree, where represents the set of time nodes, the edge set , represents the semantic correlation degree threshold;
[0129] S45. Perform hierarchical partitioning on the time nodes in the dynamic event tree:
[0130] ;
[0131] where represents the level to which the time node belongs, represents the semantic contribution degree of the time node , represents the semantic importance threshold, represents the correlation degree hierarchical threshold, represents the time interval threshold, and n represents the total number of time nodes;
[0132] S46. Iteratively update the dynamic event tree, calculate and correct the edge weights and hierarchical partitioning, and finally generate a multi-scale dynamic semantic structure.
[0133] In this embodiment, the S5 specifically includes:
[0134] S51. Initialize the improved three-stage generation network, including a summary generation module, a detail generation module, and a prediction generation module, and load the core layer features, sub-core layer features, and background layer features in the multi-scale dynamic semantic structure;
[0135] S52. In the summary generation module, extract elements from the core layer features to generate a summary broadcast content of urgent and breaking news. The summary generation module includes a sequence encoder and an attention decoder, and outputs the summary broadcast sequence of the first stage;
[0136] S53. In the detail generation module, receive the summary broadcast sequence output by the summary generation module, the sub-core layer features and background layer features in the multi-scale dynamic semantic structure, and integrate them into the multi-modal decoder. Supplement the details of the event subject, background, and context through the fine-grained attention mechanism, and output the detail broadcast sequence in the second stage;
[0137] S54. In the prediction generation module, combine the summary broadcast sequence, the detail broadcast sequence, and the temporal feature information of the multi-scale dynamic semantic structure. Capture the future trend of event evolution through a recursive temporal network, and output the predictive broadcast content based on event evolution;
[0138] S55. Integrate the summary broadcast content, the detail broadcast content, and the predictive broadcast content based on event evolution into a unified multi-modal output buffer, and perform cross-validation to screen out redundant and contradictory content;
[0139] S56. Output the cross-validated broadcast content to the user terminal through the real-time push channel.
[0140] In this embodiment, the specific steps of S6 are as follows:
[0141] S61. Collect user interaction feedback information, including the number of user clicks, the duration of stay, the result of comment sentiment analysis, and sharing behavior, to form a user feedback matrix , where represents the comprehensive feedback score of user a at time node :
[0142] ;
[0143] Among them, represents the number of clicks of user a at time node , MaxClicks represents the maximum number of clicks of all users, represents the time that user a stays at time node , MaxTime represents the maximum stay time of all users, represents the comment sentiment score of user a at time node , , and represent the weight parameters of the feedback indicators;
[0144] S62. Dynamically update the weight of each time node of the dynamic event tree based on the user feedback matrix:
[0145] ;
[0146] Among them, represents time node The updated weight, represents the balance parameter, represents the time node The original weight, represents the time node The average feedback value;
[0147] S63. Adjust the semantic generation rule of the dynamic event tree according to the updated weight, and recalculate the semantic contribution degree of each node:
[0148] ;
[0149] Among them, represents the time node The semantic contribution degree, represents the time node The updated weight, n represents the total number of time nodes, tanh represents the hyperbolic tangent function, ρ represents the amplification coefficient of feedback change, represents the time node The reference feedback value;
[0150] S64. Dynamically optimize the hierarchical structure of the dynamic event tree based on the updated weight and semantic contribution degree, perform pruning operations, and re-divide the core layer and the background layer;
[0151] S65. Use the updated time node weight and semantic generation rule to regenerate the summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution, so that the semantics of the generated content are consistent with the user feedback;
[0152] S66. Push the optimized broadcast content to the user terminal, and record the new round of user feedback information in real time to form a closed-loop optimization mechanism for user feedback and dynamic event tree adjustment.
[0153] Reference Figure 2 , A news intelligent broadcast system based on deep learning, including the following modules:
[0154] The data acquisition and encoding module is used to collect the text data, image data, and video data of news events, construct a multi-modal news data set, and perform semantic encoding on the multi-modal news data to generate preliminary semantic features;
[0155] The time and semantic feature extraction module is used to extract the time-dimensional features of the preliminary semantic features based on the dynamic time convolutional network, and perform semantic grading on the news content through the hierarchical Transformer architecture to generate time features and hierarchical semantic features;
[0156] A cross-modal alignment module, which is used to align temporal features and hierarchical semantic features with multi-modal news data through event feature-driven cross-modal contrast learning, and generate cross-modal fusion features;
[0157] A dynamic event tree generation module, which is used to perform multi-scale semantic modeling on the cross-modal fusion features by using the multi-head self-attention mechanism, construct a dynamic event tree based on time nodes and semantic relevance, and generate a multi-scale dynamic semantic structure;
[0158] A three-stage content generation module, which is used to input the multi-scale dynamic semantic structure into an improved three-stage generation network to sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution;
[0159] An event feedback optimization module, which is used to dynamically adjust the node weights and semantic generation rules of the dynamic event tree based on user interaction feedback information by using an event feedback optimization mechanism, and generate optimized broadcast content;
[0160] A multi-modal output module, which is used to synchronously output the optimized broadcast content in multi-modal forms of text, speech, and video, and transmit it to the user terminal through a real-time push channel.
[0161] Example 1:
[0162] To verify the feasibility of the present invention in implementation, the present invention is applied to the intelligent news broadcast system of an online news media platform. The platform has approximately 5 million daily active users, covering news events in multiple fields such as politics, technology, sports, and entertainment. The daily updated news volume exceeds 3,000 articles, and users have extremely high requirements for the real-time, diversity, and personalization of news acquisition. However, the traditional news broadcast system has obvious deficiencies in data processing efficiency, broadcast content accuracy, and user satisfaction. This example illustrates the actual effect of the present invention through specific scenario descriptions and data verification.
[0163] During a certain technology news conference, the platform needs to perform real-time broadcasts on the events related to the conference, including multi-modal content such as text summaries, image displays, and video interpretations. The traditional system often can only edit news titles and short articles manually, with low efficiency, and cannot effectively extract key content from the conference videos. The present invention performs semantic encoding on the conference videos, pictures, and press releases through a data collection and encoding module to generate preliminary semantic features, uses a dynamic time convolutional network to extract temporal dimension features, and differentiates core semantics and background semantics through a hierarchical Transformer architecture, thereby quickly constructing a dynamic event tree to completely cover the multi-stage evolution process of the event.
[0164] During the application process, the three-stage generation network of the present invention first generates a summary broadcast content based on the core semantic features. For example, when the press conference just starts, the system generates a breaking news summary within 30 seconds, such as "XX brand releases a new AI chip, featuring high performance and low power consumption". As the press conference content progresses, the system further extracts secondary semantic features to generate detailed broadcast content, including the technical parameters, application scenarios, market forecasts, etc. of the chip. In addition, through the time node modeling of the dynamic event tree, the system predicts the media and market reactions after the press conference, such as "It is expected that this chip will lead a new trend in the industry within the next six months", and pushes it to the user terminal in real time.
[0165] To verify the beneficial effects of the present invention, a detailed analysis was conducted on the operation data of the platform during the press conference. In the comparative experiment, the traditional news broadcast system and the method of the present invention were respectively used to process the same press conference news event, and key indicators such as news generation efficiency, user click-through rate, content matching degree, and user satisfaction were recorded. Table 1 below is the statistical result of the comparative experiment data.
[0166] Table 1 Comparative Effect Analysis Table of Traditional System and the System of the Present Invention
[0167] ;
[0168] As can be seen from the data in Table 1 above, the present invention shows significant advantages in terms of news generation time, user click-through rate, and breaking news response time. In terms of news generation time, the traditional system on average requires 180 seconds, while the present invention only requires 30 seconds, with a 500% improvement. In terms of user click-through rate, the click-through rate of the content generated by the present invention reaches 25.8%, which is a 106.4% increase compared to 12.5% of the traditional system. In addition, in terms of content matching degree and user satisfaction score, the present invention reaches 9.3 and 9.5 respectively, with a 50% and 33.8% increase compared to the traditional system. Especially in terms of breaking news response time, the performance of the present invention is particularly outstanding, and it can complete the generation and push of breaking news within 15 seconds, while the traditional system on average requires 300 seconds.
[0169] The present invention dynamically adjusts the news generation rules through the event feedback optimization module, enabling the system to adapt to the feedback needs of users in real time. For example, when users show a higher interest in technical details, the system strengthens the generation of technical parameter content through the weight adjustment mechanism and reduces the proportion of background information, thereby further improving the user's reading experience and satisfaction.
[0170] In summary, the present invention significantly improves the efficiency, accuracy, and user satisfaction of news broadcasting in practical applications, effectively solves the deficiencies of traditional systems in multi-modal data processing, event dynamic modeling, and personalized content generation, and provides an efficient and intelligent solution for the news media industry.
[0171] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A news intelligent broadcast method based on deep learning, characterized in that, It includes the following steps: S1. Collect the text data, image data, and video data of news events, construct a multi-modal news dataset and perform semantic encoding to generate preliminary semantic features; S2. Based on a dynamic time convolutional network, extract features in the time dimension from the preliminary semantic features, and perform semantic grading on the news content through a hierarchical Transformer architecture to generate time features and hierarchical semantic features; S3. Through event-characteristic-driven cross-modal contrast learning, align the time features and hierarchical semantic features with the multi-modal news data to generate cross-modal fusion features; S4. Use the multi-head self-attention mechanism to perform multi-scale semantic modeling on the cross-modal fusion features, and construct a dynamic event tree based on time nodes and semantic relevance to generate a multi-scale dynamic semantic structure; S5. Input the multi-scale dynamic semantic structure into an improved three-stage generation network to sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution, and push it to the user terminal in real time; S6. Based on the user interaction feedback information, use the event feedback optimization mechanism to dynamically adjust the node weights and semantic generation rules of the dynamic event tree to generate optimized broadcast content; S7. Synchronously output the optimized broadcast content in multi-modal forms of text, voice, and video; The specific content of S2 includes: S21. Perform embedding encoding on the preliminary semantic features, and embed feature vectors through an improved position encoding method: ; Among them, represents the partial code value of the pos-th feature vector on the even-dimensional index 2q, represents the partial code value of the pos-th feature vector on the odd-dimensional index 2q + 1, pos represents the position index of the feature vector, exp represents the exponential function, t represents the time node, d represents the dimension of the embedded feature vector, represents the time decay factor; S22. Use a dynamic time convolutional network to extract features in the time dimension from the embedded encoding features, and capture time features in irregular time series through dynamically adjusted convolutional kernels: ; Among them, represents the output eigenvalue at time step t, k represents the size of the convolution kernel, represents the weight parameter of the convolution kernel, represents the input eigenvalue at time step t - r, represents the time node difference, and b represents the bias term; S23. Construct the first layer of the hierarchical Transformer, use the multi-head self-attention mechanism to perform semantic encoding on the time features extracted by the dynamic time convolutional network, and extract the core semantic features in the news content. The core semantics include the features of the main body, keywords, and time nodes of the news event: ; Among them, represents the core semantic feature, and softmax represents the normalization function, represents the query vector of the core semantic feature, represents the key vector of the core semantic feature, represents the transpose of the key vector of the core semantic feature, represents the value vector of the core semantic feature, represents the weight matrix of the time feature, represents the normalization factor; S24. Construct the second layer of the hierarchical Transformer, perform background semantic decoding on the core semantic features, and extract secondary semantic features: ; Among them, represents a secondary semantic feature, represents a query vector of the secondary semantic feature, represents a key vector of the secondary semantic feature, represents the transpose of the key vector of the secondary semantic feature, represents a value vector of the secondary semantic feature, and represents a weight parameter, represents a positional encoding of the background semantics; S25. Based on the core semantic features and secondary semantic features, generate time features and hierarchical semantic features through a multiple dynamic weighting mechanism in the time dimension and semantic level: ; Among them, represents the time feature and the hierarchical semantic feature, represents the weight matrix of the core semantic feature, represents the weight matrix of the secondary semantic feature, represents the time feature; The specific content of S5 includes: S51. Initialize the improved three-stage generation network, including a summary generation module, a detailed generation module, and a prediction generation module, and load the core layer features, sub-core layer features, and background layer features in the multi-scale dynamic semantic structure; S52. In the summary generation module, extract elements from the core layer features to generate summary broadcast content for urgent and breaking news. The summary generation module includes a sequence encoder and an attention decoder, and outputs the summary broadcast sequence in the first stage; S53. In the detailed generation module, receive the summary broadcast sequence output by the summary generation module and the sub-core layer features and background layer features in the multi-scale dynamic semantic structure, and fuse them into the multi-modal decoder. Supplement the details of the event main body, background, and context through a fine-grained attention mechanism, and output the detailed broadcast sequence in the second stage; S54. In the prediction generation module, by combining the time feature information of the summary broadcast sequence, the detailed broadcast sequence, and the multi-scale dynamic semantic structure, the future trend of event evolution is captured through a recursive time network, and predictive broadcast content based on event evolution is output; S55. Integrate the summary broadcast content, the detailed broadcast content, and the predictive broadcast content based on event evolution into a unified multi-modal output module, and perform cross-verification to screen out redundant and contradictory content; S56. Output the cross-verified broadcast content to the user terminal through a real-time push channel.
2. The news intelligent broadcast method based on deep learning according to claim 1, characterized in that The cross-modal contrast learning driven by event characteristics in S3 specifically includes: obtaining a set of event characteristic information, including time clues, semantic subjects, and background features, as a guiding signal for feature alignment; based on cross-modal contrast learning, aligning the time features and hierarchical semantic features with multi-modal news data.
3. A news intelligent broadcast method based on deep learning according to claim 1, characterized in that, S4 specifically includes: S41. Divide the cross-modal fusion features according to time nodes to form a feature set of several time segments , where represents the time node and the corresponding feature subset; S42. Use the multi-head self-attention mechanism to perform semantic modeling on the feature subsets within each time node, extract local context associations, and obtain preliminary time node features; S43. Compare the different preliminary time node features pairwise, define the semantic association degree between time nodes, and the semantic association degree is calculated by the weighted combination of feature similarity and time difference: ; Among them, represents the semantic correlation degree between and the time node A represents the dimension of the feature vector, represents the time node in the k-th dimension of the eigenvalue, represents the time node in the k-th dimension of the eigenvalue, exp represents the exponential function, represents the time node and the time node the time difference between them, represents the time decay parameter; S44. Construct a dynamic event tree T = (N, E) based on each time node and semantic correlation degree, where represents the set of time nodes, and the edge set , represents the semantic correlation degree threshold; S45. Perform hierarchical division on the time nodes in the dynamic event tree: ; Among them, represents the level to which the time node belongs, represents the semantic contribution degree of the time node , represents the semantic importance threshold, represents the correlation degree stratification threshold, represents the time interval threshold, and n represents the total number of time nodes; S46. Iteratively update the dynamic event tree, calculate and correct the edge weights and hierarchical division, and finally generate a multi-scale dynamic semantic structure.
4. A news intelligent broadcast method based on deep learning according to claim 1, characterized in that S6 specifically includes: S61. Collect user interaction feedback information, including the number of user clicks, dwell time, comment sentiment analysis results, and sharing behavior, to form a user feedback matrix , where represents the comprehensive feedback score of user a at time node : ; Among them, represents the number of clicks of user a on the time node , MaxClicks represents the maximum number of clicks of all users, represents the time that user a stays on the time node , MaxTime represents the maximum stay time of all users, represents the comment sentiment score of user a on the time node , , and represent the weight parameters of the feedback indicators; S62. Dynamically update the weights of each time node in the dynamic event tree based on the user feedback matrix: ; Among them, represents the update weight of the time node , represents the balance parameter represents the original weight of the time node , represents the average feedback value of the time node . S63. Adjust the semantic generation rules of the dynamic event tree according to the updated weights, and recalculate the semantic contribution degree of each node: ; Among them, represents the semantic contribution degree of the time node, represents the update weight of the time node, n represents the total number of time nodes, tanh represents the hyperbolic tangent function, and ρ represents the amplification coefficient of feedback change. represents the benchmark feedback value of the time node; S64. Dynamically optimize the hierarchical structure of the dynamic event tree based on the updated weights and semantic contribution degrees, perform pruning operations, and re-divide the core layer and the background layer; S65. Use the updated time node weights and semantic generation rules to regenerate the summary broadcast content, the detailed broadcast content, and the predictive broadcast content based on event evolution, so that the semantics of the generated content are consistent with the user feedback; S66. Push the optimized broadcast content to the user terminal, and record the new round of user feedback information in real time to form a closed-loop optimization mechanism for user feedback and dynamic event tree adjustment.
5. A news intelligent broadcast system based on deep learning, which executes the news intelligent broadcast method based on deep learning according to any one of claims 1 to 4, characterized in that, It includes the following modules: Data collection and encoding module, which is used to collect text data, image data, and video data of news events, construct a multi-modal news data set, and perform semantic encoding on the multi-modal news data to generate preliminary semantic features; Time and semantic feature extraction module, which is used to extract features in the time dimension from the preliminary semantic features based on a dynamic time convolutional network, and perform semantic grading on the news content through a hierarchical Transformer architecture to generate time features and hierarchical semantic features; Cross-modal alignment module, which is used to align the time features and hierarchical semantic features with multi-modal news data through cross-modal contrast learning driven by event characteristics to generate cross-modal fusion features; A dynamic event tree generation module, which is used to perform multi-scale semantic modeling on cross-modal fusion features by using the multi-head self-attention mechanism, construct a dynamic event tree based on time nodes and semantic relevance, and generate a multi-scale dynamic semantic structure; A three-stage content generation module, which is used to input the multi-scale dynamic semantic structure into an improved three-stage generation network to sequentially generate summary broadcast content, detailed broadcast content, and predictive broadcast content based on event evolution; An event feedback optimization module, which is used to dynamically adjust the node weights and semantic generation rules of the dynamic event tree based on user interaction feedback information by using an event feedback optimization mechanism to generate optimized broadcast content; A multi-modal output module, which is used to synchronously output the optimized broadcast content in the forms of text, speech, and video, and transmit it to the user terminal through a real-time push channel.
Citation Information
Patent Citations
Intelligent news broadcasting method and system based on artificial intelligence
CN119441573A
System and method for conditional marginal distributions at flexible evaluation horizons
US20220383075A1