A news analysis method and system based on multimodal large model

Through the news analysis method based on multimodal large model, the problem of uncontrollable dissemination of hot topics on online news is solved, the semantic analysis and traceability of online news is realized, and the cohesion of news topics is improved.

CN118535978BActive Publication Date: 2025-06-06CHINA ECONOMIC INFORMATION SERVICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410535593.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-06-06
Estimated Expiration
2044-04-30

AI Technical Summary

Technical Problem

It is difficult for existing technology to effectively trace and analyze the source and dissemination of online news hotspots, resulting in uncontrollable after the fermentation and upgrading of online news hotspots, which has adverse effects on social networks.

Method used

The news analysis method based on multimodal large model is adopted to collect news data from multiple modal forms, perform preprocessing, modal transformation, feature extraction and similarity calculation, and generate news relationship networks and topics to correlate and trace online news.

Benefits of technology

It realizes semantic analysis and feature extraction of online news, can relate to split news fragments, quickly trace the news dissemination path, accurately find the source of news, and improves the cohesion of news topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118535978B_ABST
    Figure CN118535978B_ABST
Patent Text Reader

Abstract

The present invention provides a news analysis method and system based on a multimodal big model, which relates to the technical field of news analysis. The method includes: preprocessing the collected multimodal data; converting the preprocessed multimodal data into text data through the multimodal big model and performing feature extraction to obtain multiple semantic feature vectors of multiple news; respectively calculating the similarity values ​​of multiple semantic feature vectors of multiple news to obtain multiple similar feature vectors of multiple news; respectively performing weight calculation on multiple similar feature vectors of multiple news to obtain the best similar news of multiple news; respectively generating corresponding news relationship networks according to the best similar news of multiple news, respectively analyzing multiple news relationship networks to obtain multiple news topics. The present invention associates news with similar information through the generated news topics and news relationship networks, obtains the dissemination path of the news, and realizes rapid and accurate tracing of the news.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of news analysis, and in particular to a news analysis method and system based on a multimodal large model. Background Art

[0002] With the advent of the information age, mobile Internet media has developed rapidly, and the massive amount of online news information generated is prone to breed online news hotspots, which will ferment and escalate in a short period of time. In order to quickly grasp the dissemination dynamics, dissemination paths and development context of online news hotspots, it is necessary to timely trace the source of online news hotspots and analyze and judge them in order to prepare for the response.

[0003] Nowadays, online news presents diversified characteristics when it is spread in mobile Internet media, such as: through text, voice, pictures and videos. However, when forwarding online news, the media will extract news fragments from the whole news to spread, which breeds more news hotspots. However, the dissemination method of excerpted news fragments is often difficult to trace, and it is impossible to understand the source and dissemination direction of online news hotspots, resulting in uncontrollable fermentation and escalation of online news hotspots, which has a negative impact on social networks. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a news analysis method and system based on a multimodal large model in view of the deficiencies in the prior art.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] A news analysis method based on a multimodal large model comprises the following steps:

[0007] Collecting news-related data in multiple modalities from a designated news media platform to obtain multiple multimodal data corresponding to multiple news items, and preprocessing the multiple multimodal data;

[0008] Performing modal conversion on the plurality of pre-processed multimodal data respectively through the multimodal big model to obtain a plurality of semantic text data corresponding to the plurality of news, and performing feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database;

[0009] Calculating the similarities between the multiple semantic feature vectors of the multiple news and the multiple semantic feature vectors in the vector database respectively, obtaining multiple similar feature vectors corresponding to the multiple news, and performing weight calculation on the multiple semantic feature vectors of the multiple news and the corresponding multiple similar feature vectors respectively, to obtain the optimal similar news corresponding to the multiple news;

[0010] A plurality of corresponding news relationship networks are generated respectively according to the optimal similar news corresponding to the plurality of news, and all the semantic feature vectors of the plurality of news in the plurality of news relationship networks are analyzed respectively to obtain news topics corresponding to the plurality of news relationship networks.

[0011] Another technical solution of the present invention to solve the above technical problems is as follows:

[0012] A news analysis system based on a multimodal large model, comprising:

[0013] A data preprocessing module is used to collect news-related data in multiple modes from a designated news media platform, obtain multimodal data corresponding to multiple news, and preprocess the multiple multimodal data;

[0014] A feature generation module, used to perform modal conversion on the plurality of pre-processed multimodal data respectively through a multimodal large model to obtain a plurality of semantic text data corresponding to a plurality of news, and to perform feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database;

[0015] A feature calculation module, used to respectively calculate the similarity between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database, to obtain a plurality of similar feature vectors corresponding to the plurality of news, and to respectively perform weight calculation on the plurality of semantic feature vectors of the plurality of news and the corresponding plurality of similar feature vectors, to obtain the best similar news corresponding to the plurality of news;

[0016] The topic generation module is used to generate multiple corresponding news relationship networks according to the optimal similar news corresponding to multiple news, and analyze all the semantic feature vectors of multiple news in the multiple news relationship networks to obtain news topics corresponding to the multiple news relationship networks.

[0017] The beneficial effects of the present invention are: using multimodal large model technology to realize semantic analysis of information in various network news media, and extracting semantic features from text data after semantic analysis, and performing multiple computational processing on semantic features, it is possible to associate the split phenomenon that occurs when news is forwarded and propagated. Multiple related news are generated into the same news relationship network through feature vectors in related news, and the news relationship network is analyzed through semantic features to generate corresponding news topics, thereby improving the cohesion of news topics. That is, news with similar information are associated through the generated news topics and news relationship network, and the dissemination path of the news is obtained, so that the news can be quickly traced and the source of the news can be accurately found. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1A flowchart of a news analysis method based on a multimodal large model provided in an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of data collection and preprocessing provided by an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of extracting semantic features based on a multimodal large model provided by an embodiment of the present invention;

[0021] Figure 4 A schematic diagram of news relationship analysis provided by an embodiment of the present invention;

[0022] Figure 5 A module block diagram of a news analysis system based on a multimodal large model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0024] With the advent of the information age, mobile Internet media has developed rapidly in online news platforms and public accounts, generating a massive amount of online news information, which can easily breed online news hotspots and ferment and escalate in a short period of time. In order to quickly grasp the dissemination dynamics, dissemination paths and development context of online news hotspots and do a good job in response, it is necessary to timely conduct online news hotspot source tracing analysis and judgment.

[0025] Nowadays, online news information on the Internet presents diversified characteristics, such as: quoting a paragraph of online news in text, in pictures, in short videos, or cutting out a short video clip from a video for dissemination. Therefore, news analysis needs to consider the recognition and processing of multimodal news information, and accurately identify the semantic features between the content, and unify the analysis and processing of news content in various news media. News analysis technology can realize real-time monitoring and analysis of multiple channels at the same time, and can automatically mine information related to online news hotspots in multiple channels (i.e. mobile Internet media), so as to trace the source of news and divide the level of news dissemination.

[0026] The technical means of traditional news analysis mainly include:

[0027] (1) Hot topic and sensitive topic identification technology: mainly using keywords and semantic analysis to identify the sensitivity of topics;

[0028] (2) Topic tracking and detection technology: including report segmentation, topic association identification, new topic discovery and topic tracking;

[0029] (3) Automatic analysis and summarization of online text: by using computer automation and artificial intelligence technology to collect news texts on designated websites, automatically extract the topics based on the content of the news texts, and process the news texts through natural language processing technology to generate text summaries;

[0030] (4) News dissemination path analysis: Construct a dissemination network through the dissemination relationships in the network, discover key disseminators and dissemination paths, and then understand the source and flow of online news information, providing data support for refined news dissemination planning and intervention.

[0031] Therefore, traditional news analysis technology extracts news topics from news text content, only considering the accuracy of implicit semantic topics, and cannot associate news in the form of short videos, voices, pictures and texts. It does not consider the information in the entire news text, which will lead to low cohesion of topics. In addition, since many mobile Internet media forward online news based on news extraction fragments, and even partially modify the news fragments before forwarding, the modified news fragments will lose news semantics, making it difficult to trace the source. Traditional news analysis technology only relies on the propagation of explicit relationships during forwarding to analyze the news propagation path, and cannot efficiently handle the split behavior in the dissemination of network information, that is, it cannot associate multiple news fragments after segmentation, resulting in a disconnection in the analysis of the news propagation path, so the traceability is incomplete, and the source and propagation path of the news cannot be accurately found.

[0032] like Figure 1-Figure 4 As shown, a news analysis method based on a multimodal large model provided by an embodiment of the present invention includes the following steps:

[0033] Collecting news-related data in multiple modalities from a designated news media platform to obtain multiple multimodal data corresponding to multiple news items, and preprocessing the multiple multimodal data;

[0034] Performing modal conversion on the plurality of pre-processed multimodal data respectively through the multimodal big model to obtain a plurality of semantic text data corresponding to the plurality of news, and performing feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database;

[0035] Calculating the similarities between the multiple semantic feature vectors of the multiple news and the multiple semantic feature vectors in the vector database respectively, obtaining multiple similar feature vectors corresponding to the multiple news, and performing weight calculation on the multiple semantic feature vectors of the multiple news and the corresponding multiple similar feature vectors respectively, to obtain the optimal similar news corresponding to the multiple news;

[0036] A plurality of corresponding news relationship networks are generated respectively according to the optimal similar news corresponding to the plurality of news, and all the semantic feature vectors of the plurality of news in the plurality of news relationship networks are analyzed respectively to obtain news topics corresponding to the plurality of news relationship networks.

[0037] It should be understood that online news analysis refers to the process of collecting, analyzing, mining and predicting information such as speeches, comments, topics, etc. in social media using technologies such as artificial intelligence and natural language processing. Online news analysis based on multimodal data fusion technology processes and integrates various forms of social media data to accurately analyze and predict online news. A multimodal large model refers to a machine learning model that can process multiple information from different modalities (such as images, voice, text, etc.). A multimodal large model is usually composed of multiple sub-models, each of which is responsible for processing data of a specific modal type. These sub-models will interact and fuse information, thereby realizing the joint processing of multimodal data, better understanding the associations and contexts between different types of data, and thus improving the performance and generalization ability of the model to obtain semantic feature vectors.

[0038] In the embodiment of the present invention, the multimodal large model technology is used to implement semantic analysis of information in various network news media, and semantic features are extracted from the text data after semantic analysis. The semantic features are subjected to multiple computational processing, and the split phenomenon that occurs when news is forwarded and propagated can be associated. Multiple related news are generated into the same news relationship network through the feature vectors in the related news, and the news relationship network is analyzed through semantic features to generate corresponding news topics, thereby improving the cohesion of news topics. That is, news with similar information are associated through the generated news topics and news relationship network, and the dissemination path of the news is obtained, so that the news can be quickly traced and the source of the news can be accurately found.

[0039] Preferably, data related to news is collected from a designated news media platform in the form of multiple modes to obtain multimodal data corresponding to multiple news, specifically:

[0040] Based on the data provider pushing news data (i.e. external data source) to the receiving center of the pre-built big data processing platform;

[0041] Use crawler technology to collect news data from the Internet into the database of the pre-built big data processing platform.

[0042] It should be understood that data providers can be third-party news platforms and official news platforms (such as news or forums) purchased from outside. Multiple modal data corresponding to multiple news are collected through a pre-built big data processing platform. The big data processing platform uses distributed storage and computing technology, which can effectively process massive data, support multiple data types and data sources, and can be used in multiple fields such as data mining, machine learning, and log analysis; it can also perform data collection, data preprocessing, data storage, data cleaning, data query analysis, and data visualization. The big data processing platform also supports the processing and analysis of real-time collected data to obtain the latest news information in real time.

[0043] Specifically, offline news data (i.e., external data source) is obtained based on the news data pushed by the data provider. The crawler technology is used to collect news information on the Internet according to the preset time parameters to achieve the purpose of real-time news data collection and obtain real-time news data. When collecting offline news data and real-time news data respectively, it includes collecting a variety of media information (such as website type, media name and release channel, etc.), news time, and assigning a corresponding collection ID number to each piece of news data; among them, if the release time of the news can be collected during the collection, the release time is set to the news time, and if the release time cannot be collected, the collection time is set to the news time.

[0044] In the embodiment of the present invention, offline collection and / or real-time collection of already public news data is adopted to obtain multimodal data of multiple news, thereby combining news transmitted by multiple news carriers, enriching the news relationship network, and making news tracing more accurate.

[0045] Preferably, the preprocessing of the plurality of multimodal data is specifically:

[0046] The multimodal data is sharded through a pre-built big data processing platform to obtain a plurality of sharded data groups, wherein the sharded data groups include a plurality of the multimodal data; hash calculations are performed on the plurality of the multimodal data of each sharded data group to obtain a plurality of hashed multimodal data corresponding to each sharded data group; deduplication processing is performed on the plurality of hashed multimodal data corresponding to each sharded data group to obtain a plurality of deduplicated multimodal data corresponding to each sharded data group; data optimization is performed on the deduplicated multimodal data corresponding to each sharded data group to obtain a plurality of optimized multimodal data corresponding to each sharded data group; the plurality of optimized multimodal data corresponding to each sharded data group are merged to obtain a plurality of pre-processed multimodal data.

[0047] Specifically, the 10 collected news data (i.e., offline news data and / or real-time news data) are processed in slices according to the first preset time range (i.e., an interval range of the set news time) or the first preset ID range (i.e., the interval range of the collected ID number) to generate three slice data groups. The three slice data groups are cleaned and processed respectively using a multi-threaded, multi-node deployment method, including generating a unique news ID for each news, Hash calculation (i.e., hash calculation), deduplication and data optimization processing. Data optimization processing includes supplementing information to news data (such as supplementing media information such as website type, media name, and publishing channel) and invalidating news titles and text content (such as removing extra spaces, special characters, and other semantically meaningless characters according to a preset character format). Multiple news data corresponding to the three slice data groups (i.e., optimized multimodal data) are merged to obtain 10 pre-processed news data (i.e., multiple pre-processed multimodal data).

[0048] It should be understood that data sharding is to group all the news data collected at that time to improve the computational efficiency of preprocessing.

[0049] In the embodiment of the present invention, the data is processed in slices according to the time sequence of the development of network news information or the sequence of ID numbers through the big data processing platform, so that a large amount of news data can be processed with high concurrency and high efficiency, thereby improving the preprocessing speed. The integrity of the data is verified by performing hash calculation on the sliced ​​data. The collected data is deduplicated to avoid repeated collection of news and increase unnecessary calculation amount. The news data is supplemented with information so that more relevant information can be obtained when tracing the source of the news, and irrelevant characters are invalidated to reduce the amount of calculation for platform processing.

[0050] Preferably, the multimodal big model is used to perform modal conversion on the plurality of pre-processed multimodal data to obtain a plurality of semantic text data corresponding to the plurality of news, specifically:

[0051] Through the multimodal big model, any pre-processed multimodal data is modally converted to obtain multiple semantic text data of the corresponding news, including:

[0052] If the preprocessed multimodal data is initial text data, semantic segmentation processing is performed on the initial text data according to the specified character width to obtain a plurality of first short text data and a plurality of first long text data;

[0053] If the pre-processed multimodal data is initial audio data, convert the initial audio data into text to obtain audio text data, and perform semantic segmentation processing on the audio text data according to the specified character width to obtain a plurality of second short text data and a plurality of second long text data;

[0054] If the preprocessed multimodal data is initial image data, extract text from the initial image data to obtain first image text data, and perform semantic segmentation processing on the first image text data according to the specified character width to obtain a plurality of third short text data and a plurality of third long text data;

[0055] If the pre-processed multimodal data is initial video data, background audio data is divided from the initial video data, the background audio data is converted into text to obtain background audio text data, a plurality of key frame data are extracted from the initial video data, text is extracted from the plurality of key frame data respectively to obtain a plurality of second image text data, the background audio text data and the plurality of second image text data are aggregated to obtain video text data, and the video text data is semantically segmented according to the specified character width to obtain a plurality of fourth short text data and a plurality of fourth long text data;

[0056] Combining a plurality of the first short text data and / or a plurality of the second short text data and / or a plurality of the third short text data and / or a plurality of the fourth short text data into a plurality of short text data, combining a plurality of the first long text data and / or a plurality of the second long text data and / or a plurality of the third long text data and / or a plurality of the fourth long text data into a plurality of long text data, and using the plurality of the short text data and the plurality of the long text data as a plurality of semantic text data;

[0057] In this way, all pre-processed multimodal data are modally converted through the multimodal large model to obtain multiple semantic text data corresponding to multiple news.

[0058] Specifically, the news source, text content, file suffix and other news types are identified, and the multimodal technology (i.e., multimodal large model) is applied to process the news, including: segmenting the initial long text of the news according to the specified character width of 512 characters to obtain short text and long text, and retaining semantic features during segmentation; extracting text from image news to obtain text; using speech-to-text technology to convert audio news into text; using speech-to-text technology to convert the audio in video news into text, extracting key frames from video news and then extracting text to obtain text, and summarizing video text to obtain text data of various types of news. Among them, the character width of short text is less than or equal to 300 characters, and the character width of long text is greater than 300 characters and less than or equal to 512 characters or 1024 characters.

[0059] It should be understood that if the character length of the text data of each modal news exceeds the specified length (such as 512 characters or 1024 characters), the feature vector will be inaccurate during the subsequent feature extraction, so the text is segmented according to semantics (and the character length of the segmented text data is less than or equal to 512 characters). Since each frame of the video news has images with text and images without text, the images with text in the video news are used as key frames.

[0060] In the embodiment of the present invention, news data of multiple modes can be converted simultaneously through multimodal technology, text can be generated at a high speed and with high precision, and text exceeding a preset length can be segmented according to semantics, so that similar features can be matched more accurately when performing feature matching on each piece of news, making tracing more accurate.

[0061] Preferably, the multimodal large model includes a text processing unit, an audio processing unit, an image processing unit and a video processing unit;

[0062] In the text processing unit, the initial text data is semantically segmented according to the specified character width using the BERT semantic segmentation model to obtain a plurality of first short text data and a plurality of first long text data;

[0063] In the audio processing unit, the initial audio data is converted into text by a speech-to-text model to obtain audio text data, and the audio text data is semantically segmented according to the specified character width by a BERT semantic segmentation model to obtain a plurality of second short text data and a plurality of second long text data;

[0064] In the image processing unit, text is extracted from the initial image data by using an image-text conversion model to obtain first image text data, and semantic segmentation is performed on the first image text data according to the specified character width by using a BERT semantic segmentation model to obtain a plurality of third short text data and a plurality of third long text data;

[0065] In the video processing unit, background audio data is separated from the initial video data, and the background audio data is converted into text through a speech-to-text model to obtain background audio text data. A plurality of key frame data are extracted from the initial video data through a key frame extraction model, and text is extracted from the plurality of key frame data respectively through a picture-to-text conversion model to obtain a plurality of second image text data. The background audio text data and the plurality of second image text data are aggregated to obtain video text data, and the video text data is semantically segmented according to the specified character width through a BERT semantic segmentation model to obtain a plurality of fourth short text data and a plurality of fourth long text data.

[0066] Since multimodal data may include data in four modalities, namely, initial text data, initial audio data, initial image data, and initial video data, it may also include only three of the data, only two of the data, or only one of the data.

[0067] If the traffic multimodal data of a certain traffic news only includes initial text data and initial video data, the multimodal data is modally converted through the multimodal big model to obtain multiple semantic text data of the traffic news, specifically:

[0068] Performing semantic segmentation processing on the initial text data according to the specified character width by using the BERT semantic segmentation model to obtain a plurality of first short text data and a plurality of first long text data;

[0069] The initial video data is divided into background audio data, the background audio data is converted into text by a speech-to-text model to obtain background audio text data, a plurality of key frame data are extracted from the initial video data by a key frame extraction model, text is extracted from the plurality of key frame data respectively by a picture-to-text conversion model to obtain a plurality of second image text data, the background audio text data and the plurality of second image text data are aggregated to obtain video text data, and the video text data is semantically segmented according to the specified character width by a BERT semantic segmentation model to obtain a plurality of fourth short text data and a plurality of fourth long text data;

[0070] Combine multiple first short text data and multiple fourth short text data into multiple short text data, combine multiple first long text data and multiple fourth long text data into multiple long text data, and use multiple short text data and multiple long text data as semantic text data of traffic news.

[0071] It should be understood that the key frame extraction model automatically extracts representative key frames (i.e., images containing text) by learning the spatial and temporal features in the video data.

[0072] Preferably, an image-text conversion model is constructed by an image feature extractor, a global average pooling layer and a text generator, including:

[0073] Connect the input port to the input of the image feature extractor, connect the output of the image feature extractor to the input of the global average pooling layer, connect the output of the global average pooling layer to the input of the text generator, and connect the output of the text generator to the output port to complete the construction of the image-text conversion model.

[0074] The text extraction from the initial image data is performed by the image-text conversion model to obtain the first image text data, specifically:

[0075] The image feature extractor performs feature extraction on the initial image data to obtain a first text feature map, the global average pooling layer performs global average pooling on the first text feature map to obtain a first encoding feature map, and the text generator performs text extraction on the first encoding feature map to obtain first image text data.

[0076] The text extraction is performed on the plurality of key frame data respectively through the image-text conversion model to obtain a plurality of second image text data, specifically:

[0077] The image feature extractor performs feature extraction on the multiple key frame data respectively to obtain multiple second text feature maps, the global average pooling layer performs global average pooling on the multiple second text feature maps respectively to obtain multiple second encoding feature maps, and the text generator performs text extraction on the multiple second encoding feature maps respectively to obtain multiple second image text data.

[0078] Specifically, the image feature extractor of the image-to-text conversion model is a CNN neural network model, and the text generator is built based on the Transformer framework, and the text generator includes a convolutional layer and a self-attention mechanism module; the high-level features in the image (i.e., text feature map) are extracted through the CNN neural network model, the extracted high-level features are feature encoded through the global average pooling layer (i.e., global average pooling), the text representation in the high-level features after feature encoding is extracted through the convolutional layer of the text generator, the dependency in the text representation is obtained through the self-attention mechanism module of the text generator, and a coherent text sequence (i.e., image text data) is generated.

[0079] It also includes, after generating the text sequence, post-processing operations, such as removing redundant words, adjusting sentence structure, etc., to improve the readability and accuracy of the generated text. The text sequence is also stored in the database through embedding to achieve mapping of different data sources to the same dimensional space.

[0080] It should be understood that the text in multiple images containing text is extracted through OCR technology, and the text of the corresponding images is used as a training set to train the image-text conversion model, and the trained image-text conversion model is subjected to model distillation to reduce the model size and improve the inference speed.

[0081] Preferably, the feature extraction is performed on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news, specifically:

[0082] Performing full feature extraction on the multiple short text data of the multiple news items respectively to obtain title feature vectors of the multiple news items and multiple first short text feature vectors;

[0083] Semantic feature extraction is performed on the plurality of long text data of the plurality of news respectively to obtain a plurality of second short text feature vectors of the plurality of news;

[0084] The title feature vectors of multiple news items, multiple first short text feature vectors and multiple second short text feature vectors are used as multiple semantic feature vectors corresponding to the multiple news items.

[0085] Specifically, we vectorize network news based on its characteristics, and use the multimodal large model to vectorize network news based on natural language processing technology. We fully consider the characteristics of various network news when vectorizing, and perform targeted semantic feature vectorization by category. For example, we extract all semantic features from titles and short texts, extract features from long texts according to the specified character width, and retain the semantic features of the text when extracting features, generating feature vectors containing contextual semantic information. We store all extracted feature vectors in the vector database.

[0086] It should be understood that natural language processing (NLP) technology including word segmentation, part-of-speech tagging, named entity recognition, dependency syntax analysis, long and short sentence determination, etc. is used to perform semantic analysis to extract the most representative semantic features of the news content.

[0087] In the embodiment of the present invention, network news data is vectorized and the characteristics of various network news are fully considered during the vectorization, so as to achieve targeted semantic feature vectorization processing.

[0088] Preferably, before the step of respectively calculating the similarities between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database, the method further comprises:

[0089] The plurality of semantic feature vectors of the plurality of news are respectively processed in fragmentation to obtain a plurality of fragmented news groups, wherein the fragmented news groups include a plurality of news respectively composed of a plurality of semantic feature vectors, and the plurality of news in the plurality of fragmented news groups are sorted one by one according to preset conditions to obtain a plurality of news sequences.

[0090] Specifically, 10 news in the second preset time range (i.e., another interval range of the set news time) or the first preset ID range (i.e., the interval range of the news ID) are processed in segments to generate two segmented news groups. The news in the two segmented news groups are sorted one by one in time order or in order of ID numbers from large to small through the sorting algorithm of big data. The news in the first segmented news group is sorted first to generate the first group of incremental data with time order (i.e., news sequence), and then the news in the second segmented news group is sorted to generate the second group of incremental data with time order.

[0091] It should be understood that, since each piece of news is composed of multiple semantic feature vectors, when sorting the news, the multiple semantic feature vectors are used as the news to describe the operation steps as a whole.

[0092] In the embodiment of the present invention, multiple news data are sorted so that the news data are arranged in order. When similar features are matched, multiple features in the news can be matched with feature vectors in the database one by one according to the sorting order of the news, avoiding repeated comparison or missing comparison.

[0093] Preferably, the similarities between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database are calculated respectively to obtain a plurality of similar feature vectors corresponding to the plurality of news, specifically:

[0094] The similarities between the multiple semantic feature vectors of the multiple news in the multiple news sequences and the multiple semantic feature vectors in the vector database are calculated one by one through a similarity matching algorithm, and multiple similarity values ​​corresponding to the multiple semantic feature vectors in each of the news sequences are obtained. The semantic feature vectors corresponding to the similarity values ​​greater than or equal to the preset matching threshold are taken as similarity feature vectors, and similar news corresponding to the multiple news in each of the news sequences are obtained, and each of the similar news includes multiple semantic feature vectors.

[0095] Specifically, multiple semantic feature vectors of multiple news are similarly matched respectively through a similarity matching algorithm, and feature vectors that reach (i.e., are greater than or equal to) a preset value are taken as similar feature vectors by calculating cosine similarity values; news corresponding to semantic feature vectors corresponding to similarity values ​​less than a preset matching threshold are taken as root nodes, and a news relationship network corresponding to the news is constructed; and similar news corresponding to each of the first group of time-ordered incremental data is obtained, and each of the similar news corresponds to multiple semantic feature vectors.

[0096] It should be understood that similar matching algorithms include cosine similarity, Euclidean distance, Manhattan distance, Jaccard similarity, etc. The preset matching threshold is set to 0.85 or 0.9, etc.

[0097] In the embodiment of the present invention, similar feature vectors are obtained by calculating similarity values ​​between vectors, so that news with a relatively high correlation can be screened out from the database.

[0098] Preferably, since each news corresponds to multiple semantic feature vectors, each similar news corresponds to multiple semantic feature vectors, and each similar news has a corresponding title feature vector, multiple first short text feature vectors and multiple second short text feature vectors;

[0099] The weight calculation is performed on the plurality of semantic feature vectors of the plurality of news and the corresponding plurality of similar feature vectors to obtain the optimal similar news corresponding to the plurality of news, specifically:

[0100] Calculating a plurality of first short text feature vectors and a plurality of second short text feature vectors of a plurality of news in each of the news sequences one by one according to preset paragraph parameters, to obtain paragraph vectors corresponding to the plurality of news;

[0101] Calculating a plurality of first short text feature vectors and a plurality of second short text feature vectors of a plurality of similar news one by one according to preset paragraph parameters, to obtain similar paragraph vectors corresponding to the plurality of similar news;

[0102] Calculate the similarity between the paragraph vectors corresponding to the multiple news in each of the news sequences and the similar paragraph vectors of the corresponding similar news one by one through a similar matching algorithm, and obtain the paragraph scores corresponding to the multiple news in each of the news sequences;

[0103] Calculating the similarity between the title feature vectors corresponding to the multiple news in each of the news sequences and the title feature vectors of the corresponding similar news one by one through a similar matching algorithm, and obtaining the title scores corresponding to the multiple news in each of the news sequences;

[0104] Calculate the similarities between the first short text feature vectors corresponding to the news in each of the news sequences and the first short text feature vectors and the second short text feature vectors of the corresponding similar news one by one through a similar matching algorithm, and obtain the text scores corresponding to the news in each of the news sequences;

[0105] The paragraph scores, title scores and text scores corresponding to the multiple news in each of the news sequences are weighted by a weighted decision algorithm to obtain weight values ​​corresponding to the multiple news in each of the news sequences. The news corresponding to the weight values ​​are sorted according to priority in each of the news sequences according to the corresponding weight values ​​to obtain the optimal similar news corresponding to the multiple news in each of the news sequences.

[0106] Specifically, for the first group of incremental data with time sequence, the similarity of the longest paragraph vectors in the two similar news is calculated according to the preset paragraph parameters by the similarity matching algorithm to obtain the paragraph score; the similarity of the title vectors of the two similar news is calculated according to the similarity matching algorithm to obtain the title score; the similarity of the short text vector and all the vectors of the body of the similar news (i.e., the feature vector corresponding to the short text and the feature vector of the short text after the semantic segmentation of the long text) is calculated according to the similarity matching algorithm to obtain the body score. The title score, paragraph score and body score are calculated according to the preset weights by the weighted decision algorithm to obtain the total score, and the news corresponding to the weight score is sorted according to the priority of the total score to obtain the best similar news (i.e., the first news after sorting is taken as the most similar news); wherein, if the title score, paragraph score and body score are all full marks, i.e., the similarity is 100%, it is determined that the two similar news are forwarded from the original text and are directly determined to be the most similar, and there is no need to sort the news corresponding to the weight score according to the priority of the total score.

[0107] The similarity between the multiple first short text feature vectors corresponding to any news in the news sequence and the multiple first short text feature vectors and multiple second short text feature vectors of the corresponding similar news is calculated by a similar matching algorithm to obtain the text score of any news in the news sequence, which is specifically:

[0108] The similarities between the multiple first short text feature vectors corresponding to any news and the multiple first short text feature vectors and multiple second short text feature vectors of the corresponding similar news are calculated through a similarity matching algorithm to obtain multiple short text similarities. The multiple short text similarities of each news are weighted to obtain the text score of the corresponding news.

[0109] It should be understood that the score is calculated by comprehensively considering the news length, semantics, title and other information through a weighted decision algorithm. For example, if the length of the news exceeds a certain number of characters, has certain actual semantics, and has the cosine similarity with a paragraph vector of a long news article, the score is calculated.

[0110] Since each group of incremental data has a sequence, each group of data is processed according to the priority of each group sorting order, and the news in each group is processed according to the priority of the news sorting order. In addition, after the first group of incremental data with time sequence is matched with similar news, weights are calculated for similar feature vectors, and a news relationship network is generated, the second group of incremental data with time sequence is matched with similar news, weights are calculated for similar feature vectors, and a news relationship network is generated.

[0111] In the embodiment of the present invention, the news with the strongest correlation among the news is calculated through a weighted decision algorithm, so as to construct a news relationship network according to the correlation.

[0112] Preferably, the optimal similar news has a corresponding initial news relationship network;

[0113] The generating of a plurality of corresponding news relationship networks according to the optimal similar news corresponding to the plurality of news is specifically as follows:

[0114] The multiple news are respectively used as child nodes of the initial news relationship network where the corresponding optimal similar news is located, so as to generate multiple corresponding news relationship networks.

[0115] Specifically, if the best similar news is the root node in the corresponding initial news relationship network, the news is used as the parent node under the root node in the initial news relationship network to generate a news relationship network; if the best similar news is the parent node in the corresponding initial news relationship network, the news is used as the child node in the initial news relationship network to generate a news relationship network. Among them, if the similarity between a certain news and the best similar news is 100%, if the best similar news is the root node in the corresponding initial news relationship network, the news is used as the parent node under the root node in the initial news relationship network to generate a news relationship network.

[0116] In the embodiment of the present invention, multiple news are associated according to similar features between the news, that is, the multiple news are clustered and a news relationship network is constructed to facilitate subsequent news tracing, find the path of news dissemination, and find the source of news dissemination.

[0117] Preferably, the steps of analyzing all the semantic feature vectors of the multiple news in the multiple news relationship networks to obtain the news topics corresponding to the multiple news relationship networks are as follows:

[0118] The title generation model is constructed through the embedding layer, encoding layer, decoding layer and output layer, including:

[0119] Connecting the output of the embedding layer to the input of the encoding layer, connecting the output of the encoding layer to the input of the decoder, and connecting the output of the decoder to the input of the output layer to complete the construction of the title generation model;

[0120] All the semantic feature vectors of multiple news in the corresponding news relationship network are obtained through the embedding layer, the dependency feature representations between the multiple semantic feature vectors are extracted through the encoding layer, the dependency feature representations are constructed into a topic text sequence through the decoder, and the topic text sequence is converted into a news topic through the output layer, and so on, until the topics of multiple news relationship networks are generated, and the news topics corresponding to the multiple news relationship networks are obtained.

[0121] The method also includes storing the plurality of news relationship networks and corresponding news topics in a relationship database.

[0122] Specifically, the title generation model is built based on the transformer framework. The embedding layer obtains all semantic feature vectors of multiple news in the news relationship network from the vector database, that is, the semantic feature vectors of multiple news in the vector database are extracted as input. The encoder layer is composed of multiple encoder stacks, each of which contains a self-attention mechanism and a feedforward neural network; the self-attention mechanism allows the model to focus on different parts of the input sequence, thereby capturing the dependencies between words, and the feedforward neural network further processes the weighted representation of the attention mechanism. The decoder layer is composed of multiple stacked decoder layers, each of which contains a self-attention mechanism, an encoder-decoder attention mechanism, and a feedforward neural network; the encoder-decoder attention mechanism allows the decoder to focus on specific parts of the encoder output, thereby generating relevant titles based on the content of the article. The output layer converts the output of the decoder into actual title text; it is implemented through a linear layer and a softmax function, outputs the probability distribution of words at each position, and then uses a beam search algorithm to generate multiple candidate titles with priorities, and selects the optimal news title according to the priority. Among them, multiple pre-collected existing news are constructed into a training set to train the title generation model.

[0123] The semantic feature vectors of all news in a news relationship network are learned regularly according to the learning time parameter, and topics are generated through semantic feature analysis. This makes the generated topics more interpretable and pays more attention to contextual semantic similarity, solving the problem that the traditional text representation model only considers the accuracy of implicit semantic topics, but does not consider the entire text information of the news, resulting in low topic cohesion. Among them, the news topic can be a sentence in the form of a title, a paragraph of topic text introduction, or a group of representative keyword words.

[0124] It should be understood that the method of generating news topics also includes: ignoring the contextual semantic similarity, and replacing the learned topics with network media weights, semantic feature intensive areas, and time intensive areas of network news. The topic can also be replaced by the news corresponding to the root node. The traditional text representation model can also be used to replace the generated topic learning. At the same time, some special rules can also be applied to generate topics, such as using the title of multiple news names with core media attributes as the news topic.

[0125] In the embodiment of the present invention, the clustering relationship (i.e., the multiple news relationship networks and the corresponding news topics) is saved in the relationship database, and the relationship between the news can be displayed later using content visualization technologies such as hierarchical trees and family pedigrees. The splitting behavior in the dissemination of network information and the modified parts of the news during the dissemination process (such as paragraph removal or addition) can be analyzed to achieve the identification and traceability analysis of news topics.

[0126] like Figure 5 As shown, a news analysis system based on a multimodal large model provided by an embodiment of the present invention includes:

[0127] A data preprocessing module is used to collect news-related data in multiple modes from a designated news media platform, obtain multimodal data corresponding to multiple news, and preprocess the multiple multimodal data;

[0128] A feature generation module, used to perform modal conversion on the plurality of pre-processed multimodal data respectively through a multimodal large model to obtain a plurality of semantic text data corresponding to a plurality of news, and to perform feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database;

[0129] A feature calculation module, used to respectively calculate the similarity between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database, to obtain a plurality of similar feature vectors corresponding to the plurality of news, and to respectively perform weight calculation on the plurality of semantic feature vectors of the plurality of news and the corresponding plurality of similar feature vectors, to obtain the best similar news corresponding to the plurality of news;

[0130] The topic generation module is used to generate multiple corresponding news relationship networks according to the optimal similar news corresponding to multiple news, and analyze all the semantic feature vectors of multiple news in the multiple news relationship networks to obtain news topics corresponding to the multiple news relationship networks.

[0131] The above-mentioned news analysis system based on a multimodal big model can refer to the implementation content and beneficial effects of the news analysis method based on a multimodal big model specifically described above, which will not be repeated here.

[0132] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A news analysis method based on a multimodal large model, characterized in that: The steps include: Collecting news-related data in multiple modalities from a designated news media platform to obtain multiple multimodal data corresponding to multiple news, and preprocessing the multiple multimodal data, specifically: Slice the plurality of multimodal data through a pre-built big data processing platform to obtain a plurality of slicing data groups, wherein the slicing data groups include a plurality of the multimodal data; hash calculations are performed on the plurality of multimodal data of each slicing data group to obtain a plurality of hashed multimodal data corresponding to each slicing data group; de-duplicate the plurality of hashed multimodal data corresponding to each slicing data group to obtain a plurality of de-duplicated multimodal data corresponding to each slicing data group; optimize the de-duplicated multimodal data corresponding to each slicing data group to obtain a plurality of optimized multimodal data corresponding to each slicing data group; merge the plurality of optimized multimodal data corresponding to each slicing data group to obtain a plurality of pre-processed multimodal data; Performing modal conversion on the plurality of pre-processed multimodal data respectively through the multimodal big model to obtain a plurality of semantic text data corresponding to the plurality of news, and performing feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database; Calculating the similarities between the multiple semantic feature vectors of the multiple news and the multiple semantic feature vectors in the vector database respectively, obtaining multiple similar feature vectors corresponding to the multiple news, and performing weight calculation on the multiple semantic feature vectors of the multiple news and the corresponding multiple similar feature vectors respectively, to obtain the optimal similar news corresponding to the multiple news; Generating a plurality of corresponding news relationship networks according to the optimal similar news corresponding to the plurality of news, respectively, analyzing all the semantic feature vectors of the plurality of news in the plurality of news relationship networks, and obtaining news topics corresponding to the plurality of news relationship networks; The analysis of all the semantic feature vectors of the multiple news in the multiple news relationship networks is performed to obtain the news topics corresponding to the multiple news relationship networks, specifically: The title generation model is constructed through the embedding layer, encoding layer, decoding layer and output layer, including: Connecting the output of the embedding layer to the input of the encoding layer, connecting the output of the encoding layer to the input of the decoding layer, and connecting the output of the decoding layer to the input of the output layer to complete the construction of the title generation model; All the semantic feature vectors of multiple news in the corresponding news relationship network are obtained through the embedding layer, the dependency feature representations between the multiple semantic feature vectors are extracted through the encoding layer, the dependency feature representations are constructed into a topic text sequence through the decoding layer, and the topic text sequence is converted into a news topic through the output layer, and so on, until the topics of multiple news relationship networks are generated, and the news topics corresponding to the multiple news relationship networks are obtained.

2. The news analysis method according to claim 1, characterized in that: The multimodal model is used to perform modal conversion on the multiple pre-processed multimodal data to obtain multiple semantic text data corresponding to the multiple news, specifically: Through the multimodal big model, any pre-processed multimodal data is modally converted to obtain multiple semantic text data of the corresponding news, including: If the preprocessed multimodal data is initial text data, semantic segmentation processing is performed on the initial text data according to the specified character width to obtain a plurality of first short text data and a plurality of first long text data; If the pre-processed multimodal data is initial audio data, convert the initial audio data into text to obtain audio text data, and perform semantic segmentation processing on the audio text data according to the specified character width to obtain a plurality of second short text data and a plurality of second long text data; If the preprocessed multimodal data is initial image data, extract text from the initial image data to obtain first image text data, and perform semantic segmentation processing on the first image text data according to the specified character width to obtain a plurality of third short text data and a plurality of third long text data; If the pre-processed multimodal data is initial video data, background audio data is divided from the initial video data, the background audio data is converted into text to obtain background audio text data, a plurality of key frame data are extracted from the initial video data, text is extracted from the plurality of key frame data respectively to obtain a plurality of second image text data, the background audio text data and the plurality of second image text data are aggregated to obtain video text data, and the video text data is semantically segmented according to the specified character width to obtain a plurality of fourth short text data and a plurality of fourth long text data; Combining a plurality of the first short text data and / or a plurality of the second short text data and / or a plurality of the third short text data and / or a plurality of the fourth short text data into a plurality of short text data, combining a plurality of the first long text data and / or a plurality of the second long text data and / or a plurality of the third long text data and / or a plurality of the fourth long text data into a plurality of long text data, and using the plurality of the short text data and the plurality of the long text data as a plurality of semantic text data; In this way, all pre-processed multimodal data are modally converted through the multimodal large model to obtain multiple semantic text data corresponding to multiple news.

3. The news analysis method according to claim 2, characterized in that: The feature extraction of the plurality of semantic text data of the plurality of news items is performed respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news items, specifically: Performing full feature extraction on the multiple short text data of the multiple news items respectively to obtain title feature vectors of the multiple news items and multiple first short text feature vectors; Semantic feature extraction is performed on the plurality of long text data of the plurality of news respectively to obtain a plurality of second short text feature vectors of the plurality of news; The title feature vectors of multiple news items, multiple first short text feature vectors and multiple second short text feature vectors are used as multiple semantic feature vectors corresponding to the multiple news items.

4. The news analysis method according to claim 3, characterized in that: Before the step of respectively calculating the similarities between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database, the method further includes: The plurality of semantic feature vectors of the plurality of news are respectively processed in fragmentation to obtain a plurality of fragmented news groups, wherein the fragmented news groups include a plurality of news respectively composed of a plurality of semantic feature vectors, and the plurality of news in the plurality of fragmented news groups are sorted one by one according to preset conditions to obtain a plurality of news sequences.

5. The news analysis method according to claim 4, characterized in that: The similarity between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database is calculated respectively to obtain a plurality of similar feature vectors corresponding to the plurality of news, specifically: The similarities between the multiple semantic feature vectors of the multiple news in the multiple news sequences and the multiple semantic feature vectors in the vector database are calculated one by one through a similarity matching algorithm, and multiple similarity values ​​corresponding to the multiple semantic feature vectors in each of the news sequences are obtained. The semantic feature vectors corresponding to the similarity values ​​greater than or equal to the preset matching threshold are taken as similarity feature vectors, and similar news corresponding to the multiple news in each of the news sequences are obtained, and each of the similar news includes multiple semantic feature vectors.

6. The news analysis method according to claim 5, characterized in that: The weight calculation is performed on the plurality of semantic feature vectors of the plurality of news and the corresponding plurality of similar feature vectors to obtain the optimal similar news corresponding to the plurality of news, specifically: Calculating a plurality of first short text feature vectors and a plurality of second short text feature vectors of a plurality of news in each of the news sequences one by one according to preset paragraph parameters, to obtain paragraph vectors corresponding to the plurality of news; Calculating a plurality of first short text feature vectors and a plurality of second short text feature vectors of a plurality of similar news one by one according to preset paragraph parameters, to obtain similar paragraph vectors corresponding to the plurality of similar news; Calculate the similarity between the paragraph vectors corresponding to the multiple news in each of the news sequences and the similar paragraph vectors of the corresponding similar news one by one through a similar matching algorithm, and obtain the paragraph scores corresponding to the multiple news in each of the news sequences; Calculating the similarity between the title feature vectors corresponding to the multiple news in each of the news sequences and the title feature vectors of the corresponding similar news one by one through a similar matching algorithm, and obtaining the title scores corresponding to the multiple news in each of the news sequences; Calculate the similarities between the first short text feature vectors corresponding to the news in each of the news sequences and the first short text feature vectors and the second short text feature vectors of the corresponding similar news one by one through a similar matching algorithm, and obtain the text scores corresponding to the news in each of the news sequences; The paragraph scores, title scores and text scores corresponding to the multiple news in each of the news sequences are weighted by a weighted decision algorithm to obtain weight values ​​corresponding to the multiple news in each of the news sequences. The news corresponding to the weight values ​​are sorted according to priority in each of the news sequences according to the corresponding weight values ​​to obtain the optimal similar news corresponding to the multiple news in each of the news sequences.

7. The news analysis method according to claim 1, characterized in that: The optimal similar news has a corresponding initial news relationship network; The generating of a plurality of corresponding news relationship networks according to the optimal similar news corresponding to the plurality of news is specifically as follows: The multiple news are respectively used as child nodes of the initial news relationship network where the corresponding optimal similar news is located, so as to generate multiple corresponding news relationship networks.

8. A news analysis system based on a multimodal large model, characterized in that: include: The data preprocessing module is used to collect news-related data in multiple modes from the specified news media platform, obtain multimodal data corresponding to multiple news, and preprocess the multiple multimodal data, specifically: Slice the plurality of multimodal data through a pre-built big data processing platform to obtain a plurality of slicing data groups, wherein the slicing data groups include a plurality of the multimodal data; hash calculations are performed on the plurality of multimodal data of each slicing data group to obtain a plurality of hashed multimodal data corresponding to each slicing data group; de-duplicate the plurality of hashed multimodal data corresponding to each slicing data group to obtain a plurality of de-duplicated multimodal data corresponding to each slicing data group; optimize the de-duplicated multimodal data corresponding to each slicing data group to obtain a plurality of optimized multimodal data corresponding to each slicing data group; merge the plurality of optimized multimodal data corresponding to each slicing data group to obtain a plurality of pre-processed multimodal data; A feature generation module, used to perform modal conversion on the plurality of pre-processed multimodal data respectively through a multimodal large model to obtain a plurality of semantic text data corresponding to a plurality of news, and to perform feature extraction on the plurality of semantic text data of the plurality of news respectively to obtain a plurality of semantic feature vectors corresponding to the plurality of news and store them in a vector database; A feature calculation module, used to respectively calculate the similarity between the plurality of semantic feature vectors of the plurality of news and the plurality of semantic feature vectors in the vector database, to obtain a plurality of similar feature vectors corresponding to the plurality of news, and to respectively perform weight calculation on the plurality of semantic feature vectors of the plurality of news and the corresponding plurality of similar feature vectors, to obtain the best similar news corresponding to the plurality of news; A topic generation module, used to generate a plurality of corresponding news relationship networks according to the optimal similar news corresponding to the plurality of news, and to analyze all the semantic feature vectors of the plurality of news in the plurality of news relationship networks to obtain news topics corresponding to the plurality of news relationship networks; In the topic generation module, all the semantic feature vectors of multiple news in the multiple news relationship networks are analyzed respectively to obtain news topics corresponding to the multiple news relationship networks, specifically: The title generation model is constructed through the embedding layer, encoding layer, decoding layer and output layer, including: Connecting the output of the embedding layer to the input of the encoding layer, connecting the output of the encoding layer to the input of the decoding layer, and connecting the output of the decoding layer to the input of the output layer to complete the construction of the title generation model; All the semantic feature vectors of multiple news in the corresponding news relationship network are obtained through the embedding layer, the dependency feature representations between the multiple semantic feature vectors are extracted through the encoding layer, the dependency feature representations are constructed into a topic text sequence through the decoding layer, and the topic text sequence is converted into a news topic through the output layer, and so on, until the topics of multiple news relationship networks are generated, and the news topics corresponding to the multiple news relationship networks are obtained.

Citation Information

Patent Citations

  • Network news hotspot mining method and device

    CN111198946A

  • All-media news intelligent cataloguing method based on multi-modal information fusion understanding

    CN112818906A