A method for arranging and generating information data analysis reports based on large models
By analyzing online news and social media data through large models, accurate news comprehensive reports are generated, which solves the problem of trend shifts in news hotspots and improves the efficiency of news editing and the quality of content.
Patent Information
- Application Number
- CN202411789442.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-12-06
AI Technical Summary
With the rapid development of the Internet and the growth of mobile users, news hot information is no longer limited to journalists' field investigations, but is casually recorded and uploaded by individual users, resulting in a trend shift in hot news events, making it difficult to effectively assist journalists in making decisions.
Through a large-scale model-based information data analysis method, we use the backend API interface of the Internet and social media to obtain real-time news data, perform semantic recognition and vector analysis, combine historical hot search data to evaluate public opinion trends, and generate accurate comprehensive reports on online news.
It has improved the efficiency of news editors in producing manuscripts and improved the quality of news content.
Smart Images

Figure CN119669571B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information retrieval technology, and in particular to a method for collating and generating information data analysis reports based on a large model. Background Art
[0002] Big model information data analysis report compilation refers to the systematic sorting, analysis and summary of massive information data processed by big models, extracting key data points, trends, correlations and insights, and forming a clearly structured and easy-to-understand report so that decision makers can quickly grasp the information behind the data and provide a basis for planning and business decisions.
[0003] Based on the rapid development of the Internet and the growth of domestic mobile users, news hot information is no longer limited to journalists' on-site investigations, but appears randomly as individual users upload their records to the Internet. However, hot news events usually have a trend shift, that is, content similar to hot news is more likely to attract the attention of individual users. Therefore, a decision-making plan to assist journalists is proposed. Summary of the Invention
[0004] In order to solve the above technical problems, a method for compiling and generating information data analysis reports based on a large model is provided. This technical solution solves the above-mentioned rapid development of the Internet and the growth of domestic mobile users. News hot information is no longer limited to journalists' on-site investigations, but appears randomly as personal users upload their records to the Internet. However, hot news events usually have a trend transfer, that is, content similar to hot news is more likely to attract the attention of personal users. Therefore, a set of decision-making solutions to assist journalists is proposed.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A method for collating and generating information data analysis reports based on a large model, comprising:
[0007] Based on the backend API interface of the Internet and social media, obtain real-time news comprehensive sample data and build a comprehensive sample database of online news;
[0008] Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples;
[0009] Based on the topic information parameters of the comprehensive sample of online news, scene vector analysis is performed on the topic information parameters to obtain the theme scene vector of the comprehensive sample of online news;
[0010] Obtain comprehensive sample data of historical trending news, determine the information parameters of trending topics, analyze the keyword change trends in the information parameters of trending topics, and evaluate the public opinion trend vector of the trending topic scene;
[0011] We conducted a fitting analysis on the theme scene vectors of the comprehensive online news sample and the public opinion trend vectors of the hot search theme scenes to evaluate the preference index between the comprehensive online news sample and the hot search topics.
[0012] Generate a comprehensive online news report corresponding to the preference index between the comprehensive online news sample and the hot search topics;
[0013]
[0014] Where G k is the preference index between the comprehensive sample of online news and the kth hot search topic, P(z k |d i ) is the kth topic scene vector of the i-th comprehensive sample of online news, is the mean value of the kth topic scene vector of the i-th online news comprehensive sample, Y k (t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the mean of the public opinion trend vector of the kth hot search theme scene in t unit time, and L is the total number of theme scene vectors.
[0015] Preferably, based on the online news comprehensive sample database, semantic recognition and screening of the online news comprehensive sample is performed to obtain topic information parameters of the online news comprehensive sample, specifically including:
[0016] Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples;
[0017] Preprocess the noise data in the comprehensive sample of online news and split it into image data and text data to obtain news image data and news text data;
[0018] Substitute the news image data into the CNN convolutional neural network model, use the convolution layer to perform convolution and pooling operations on the news impact data to obtain the news image feature map, and input it into the feature extraction layer for global average. The output of the residual block of the pooling layer is substituted into the full connection layer and the softmax layer to generate the feature map of the feature extraction layer. The feature map is flattened into a one-dimensional array and normalized to obtain the news image feature vector.
[0019] Substitute news text data into the BERT model, use the input embedding layer to map word vectors, paragraph vectors, and position vectors for the news text data, and obtain the news text embedding vector. This is then input into the Transformer encoder for multi-head self-attention mechanism and feedforward neural network training, and the news text feature vector is output.
[0020] Based on the news image feature vector and the news text feature vector, feature splicing is performed in NumPy to obtain the topic information parameters of the comprehensive sample of online news.
[0021] Preferably, performing scene vector analysis on the topic information parameters of the comprehensive network news sample to obtain the subject scene vector of the comprehensive network news sample specifically includes:
[0022] Based on the topic information parameters of the comprehensive sample of online news, a topic information parameter matrix of the comprehensive sample of online news is constructed;
[0023] Perform topic posterior probability analysis on each element in the topic information parameter matrix of the comprehensive sample of online news to obtain the topic scenario corresponding to the comprehensive sample of online news;
[0024] Verify the likelihood function value of each element in the topic information parameter matrix for the topic scene, and evaluate the classification probability of the topic scene of the comprehensive sample of online news;
[0025] The classification probability of the theme scene of the comprehensive sample of online news is iterated several times until the likelihood function value reaches convergence, and the theme scene vector of the comprehensive sample of online news is obtained;
[0026] The subject scene vector of the comprehensive online news sample is specifically:
[0027]
[0028] Where, P(z k |d i ) is the kth topic scene vector of the i-th network news comprehensive sample, P(w j |z k ) is the probability distribution of the jth topic information parameter in the kth topic scenario, P(z k |w ij ,d i ) is the probability distribution of the kth topic scenario given the jth topic information parameter of the i-th network news comprehensive sample, δ(w,w ij ) is the indicator function, when w is equal to w ij When , its value is 1, otherwise it is 0, n is the total number of comprehensive samples of online news, and m is the total number of topic information.
[0029] Preferably, obtaining comprehensive sample data of historical hot search news, determining hot search topic information parameters, analyzing keyword change trends in the hot search topic information parameters, and evaluating the public opinion trend vector of the hot search theme scene specifically include:
[0030] Based on the comprehensive sample data of historical hot search news, the keywords in the hot search topic information parameters are marked according to the unit time, and the time series parameters of the topic information of historical hot search news are constructed;
[0031] Based on the time series parameters of the topic information of historical hot search news, the TF-IDF value of each keyword is calculated to determine the topic information word heat distribution value of each historical hot search news;
[0032] Based on text sentiment analysis, the sentiment tendency of the topic information word heat distribution value of each historical hot search news is evaluated, and the sentiment tendency distribution value of the topic information word heat distribution value of each historical hot search news is obtained;
[0033] Based on ARIMA time series autoregression, a public opinion trend analysis model for historical hot search news is constructed;
[0034] The topic information vocabulary heat distribution value of each historical hot search news and the sentiment tendency distribution value of the topic information vocabulary of each historical hot search news are substituted into the public opinion trend analysis model of historical hot search news. The objective function is constructed based on the difference between the predicted value and the actual value. The prediction error function of the training result is used as the end target to generate the public opinion trend vector of the hot search topic scene.
[0035] The public opinion trend analysis model of the historical hot search news is specifically as follows:
[0036]
[0037] Where Y k (t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the predicted value of the zth word heat distribution of the topic information of the hot search news in the historical t unit time, is the predicted value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, is the actual value of the zth word heat distribution of the hot search news topic information in the historical t unit time, is the actual value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, u is the total number of topic information words in the hot search news, and T is the total time.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] This paper proposes a large-scale model-based information data analysis and report generation solution. Based on multi-source data analysis, this solution leverages real-time data from online news and social media, combined with semantic recognition and vectorized analysis algorithms, to automatically identify and extract topic information. It then uses historical trending search data to assess public opinion trends and conducts matching and fitting analysis between topics and public opinion trends, thereby generating accurate online news reports. This approach has the beneficial effects of improving editorial efficiency and enhancing the quality of news content. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flow chart of a method for arranging and generating information data analysis reports based on a large model;
[0041] Figure 2 A flow chart of the method for obtaining topic information parameters of a comprehensive sample of online news;
[0042] Figure 3 Flowchart of the method for obtaining topic scene vectors of comprehensive samples of online news;
[0043] Figure 4 Flowchart of the opinion trend vector method for evaluating hot search topic scenarios. DETAILED DESCRIPTION
[0044] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0045] Reference Figure 1 As shown, a method for collating and generating information data analysis reports based on a large model includes:
[0046] Based on the backend API interface of the Internet and social media, obtain real-time news comprehensive sample data and build a comprehensive sample database of online news;
[0047] Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples;
[0048] Based on the topic information parameters of the comprehensive sample of online news, scene vector analysis is performed on the topic information parameters to obtain the theme scene vector of the comprehensive sample of online news;
[0049] Obtain comprehensive sample data of historical trending news, determine the information parameters of trending topics, analyze the keyword change trends in the information parameters of trending topics, and evaluate the public opinion trend vector of the trending topic scene;
[0050] We conducted a fitting analysis on the theme scene vectors of the comprehensive online news sample and the public opinion trend vectors of the hot search theme scenes to evaluate the preference index between the comprehensive online news sample and the hot search topics.
[0051] Generate a comprehensive online news report corresponding to the preference index between the comprehensive online news sample and the hot search topics;
[0052]
[0053] Where G k is the preference index between the comprehensive sample of online news and the kth hot search topic, P(z k |d i ) is the kth topic scene vector of the i-th comprehensive sample of online news, is the mean value of the kth topic scene vector of the i-th online news comprehensive sample, Y k (t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the mean of the public opinion trend vector of the kth hot search theme scene in t unit time, and L is the total number of theme scene vectors.
[0054] This solution, based on multi-source data analysis, leverages real-time data from online news and social media, combined with semantic recognition and vectorized analysis algorithms, to automatically identify and extract topic information. It then uses historical trending search data to assess public opinion trends and analyze the matching and fitting of topics and public opinion trends, generating accurate online news reports. This has the potential to improve editorial efficiency and enhance the quality of news content.
[0055] Reference Figure 2 As shown, based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples, including:
[0056] Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples;
[0057] Preprocess the noise data in the comprehensive sample of online news and split it into image data and text data to obtain news image data and news text data;
[0058] Substitute the news image data into the CNN convolutional neural network model, use the convolution layer to perform convolution and pooling operations on the news impact data to obtain the news image feature map, and input it into the feature extraction layer for global average. The output of the residual block of the pooling layer is substituted into the full connection layer and the softmax layer to generate the feature map of the feature extraction layer. The feature map is flattened into a one-dimensional array and normalized to obtain the news image feature vector.
[0059] Substitute news text data into the BERT model, use the input embedding layer to map word vectors, paragraph vectors, and position vectors for the news text data, and obtain the news text embedding vector. This is then input into the Transformer encoder for multi-head self-attention mechanism and feedforward neural network training, and the news text feature vector is output.
[0060] Based on the news image feature vector and the news text feature vector, feature splicing is performed in NumPy to obtain the topic information parameters of the comprehensive sample of online news.
[0061] It is understandable that since the comprehensive online news sample contains independent text data or image data, when processing the topic information vector of the comprehensive online news sample, it is necessary to create two independent private variables to encapsulate the image data and text data in the comprehensive online news sample to avoid mutual contamination between the data.
[0062] Reference Figure 3 As shown, based on the topic information parameters of the comprehensive network news sample, scene vector analysis is performed on the topic information parameters to obtain the theme scene vector of the comprehensive network news sample, specifically including:
[0063] Based on the topic information parameters of the comprehensive sample of online news, a topic information parameter matrix of the comprehensive sample of online news is constructed;
[0064] Perform topic posterior probability analysis on each element in the topic information parameter matrix of the comprehensive sample of online news to obtain the topic scenario corresponding to the comprehensive sample of online news;
[0065] Verify the likelihood function value of each element in the topic information parameter matrix for the topic scene, and evaluate the classification probability of the topic scene of the comprehensive sample of online news;
[0066] The classification probability of the theme scene of the comprehensive sample of online news is iterated several times until the likelihood function value reaches convergence, and the theme scene vector of the comprehensive sample of online news is obtained;
[0067] The subject scene vector of the comprehensive online news sample is specifically:
[0068]
[0069] Where, P(z k |d i ) is the kth topic scene vector of the i-th network news comprehensive sample, P(w j |z k ) is the probability distribution of the jth topic information parameter in the kth topic scenario, P(z k |wij ,d i ) is the probability distribution of the kth topic scenario given the jth topic information parameter of the i-th network news comprehensive sample, δ(w,w ij ) is the indicator function, when w is equal to w ij When , its value is 1, otherwise it is 0, n is the total number of comprehensive samples of online news, and m is the total number of topic information.
[0070] This solution uses topic information from online news to construct a parameter matrix. This matrix analysis determines the news's thematic context, and the likelihood function is calculated to verify its accuracy. After multiple iterations of optimization, the thematic context vector for the news is ultimately obtained, enabling accurate and efficient extraction of news thematic information.
[0071] Reference Figure 4 As shown, obtaining comprehensive sample data of historical hot search news, determining hot search topic information parameters, analyzing keyword change trends in hot search topic information parameters, and evaluating the public opinion trend vector of the hot search theme scene specifically include:
[0072] Based on the comprehensive sample data of historical hot search news, the keywords in the hot search topic information parameters are marked according to the unit time, and the time series parameters of the topic information of historical hot search news are constructed;
[0073] Based on the time series parameters of the topic information of historical hot search news, the TF-IDF value of each keyword is calculated to determine the topic information word heat distribution value of each historical hot search news;
[0074] Based on text sentiment analysis, the sentiment tendency of the topic information word heat distribution value of each historical hot search news is evaluated, and the sentiment tendency distribution value of the topic information word heat distribution value of each historical hot search news is obtained;
[0075] Based on ARIMA time series autoregression, a public opinion trend analysis model for historical hot search news is constructed;
[0076] Substitute the topic information word popularity distribution value and the sentiment tendency distribution value of each historical hot search news into the historical hot search news public opinion trend analysis model. Use the difference between the predicted value and the actual value to construct the objective function. Use the prediction error function of the training result as the end target to generate the public opinion trend vector of the hot search topic scene.
[0077] The public opinion trend analysis model of the historical hot search news is specifically as follows:
[0078]
[0079] Where Y k(t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the predicted value of the zth word heat distribution of the topic information of the hot search news in the historical t unit time, is the predicted value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, is the actual value of the zth word heat distribution of the hot search news topic information in the historical t unit time, is the actual value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, u is the total number of topic information words in the hot search news, and T is the total time.
[0080] This solution collects comprehensive sample data from historical trending news, time-series tags the keywords in the trending topic information parameters, calculates the TF-IDF value of each keyword to measure its popularity distribution, and further combines it with text sentiment analysis to obtain the sentiment distribution of each keyword. Subsequently, an ARIMA time-series autoregressive model is used to construct a public opinion trend analysis model. The keyword popularity distribution and sentiment distribution values are substituted into the model for training, with the goal of minimizing prediction error. Ultimately, a public opinion trend vector for the trending topic scenario is generated, providing decision support for journalists.
[0081] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for arranging and generating information data analysis reports based on a large model, characterized in that: include: Based on the backend API interface of the Internet and social media, obtain real-time news comprehensive sample data and build a comprehensive sample database of online news; Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples; Based on the topic information parameters of the comprehensive sample of online news, scene vector analysis is performed on the topic information parameters to obtain the theme scene vector of the comprehensive sample of online news; Obtain comprehensive sample data of historical trending news, determine the information parameters of trending topics, analyze the keyword change trends in the information parameters of trending topics, and evaluate the public opinion trend vector of the trending topic scene; We conducted a fitting analysis on the theme scene vectors of the comprehensive online news sample and the public opinion trend vectors of the hot search theme scenes to evaluate the preference index between the comprehensive online news sample and the hot search topics. According to the preference index between the comprehensive sample of online news and the hot search topics, a comprehensive report on online news corresponding to the preference is generated, including: Where G k is the preference index between the comprehensive sample of online news and the kth hot search topic, P(z k |d i ) is the kth topic scene vector of the i-th comprehensive sample of online news, is the mean value of the kth topic scene vector of the i-th online news comprehensive sample, Y k (t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the mean of the public opinion trend vector of the kth hot search theme scene in t unit time, and L is the total number of theme scene vectors.
2. The method for arranging and generating information data analysis reports based on a large model according to claim 1, characterized in that: Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples, including: Based on the online news comprehensive sample database, semantic recognition and screening of online news comprehensive samples are performed to obtain topic information parameters of the online news comprehensive samples; Preprocess the noise data in the comprehensive sample of online news and split it into image data and text data to obtain news image data and news text data; Substitute the news image data into the CNN convolutional neural network model, use the convolution layer to perform convolution and pooling operations on the news impact data to obtain the news image feature map, and input it into the feature extraction layer for global average. The output of the residual block of the pooling layer is substituted into the full connection layer and the softmax layer to generate the feature map of the feature extraction layer. The feature map is flattened into a one-dimensional array and normalized to obtain the news image feature vector. Substitute news text data into the BERT model, use the input embedding layer to map word vectors, paragraph vectors, and position vectors for the news text data, and obtain the news text embedding vector. This is then input into the Transformer encoder for multi-head self-attention mechanism and feedforward neural network training, and the news text feature vector is output. Based on the news image feature vector and the news text feature vector, feature splicing is performed in NumPy to obtain the topic information parameters of the comprehensive sample of online news.
3. The method for arranging and generating information data analysis reports based on a large model according to claim 2, characterized in that: Based on the topic information parameters of the comprehensive sample of online news, scene vector analysis is performed on the topic information parameters to obtain the theme scene vector of the comprehensive sample of online news. Specifically, the following are included: Based on the topic information parameters of the comprehensive sample of online news, a topic information parameter matrix of the comprehensive sample of online news is constructed; Perform topic posterior probability analysis on each element in the topic information parameter matrix of the comprehensive sample of online news to obtain the topic scenario corresponding to the comprehensive sample of online news; Verify the likelihood function value of each element in the topic information parameter matrix for the topic scene, and evaluate the classification probability of the topic scene of the comprehensive sample of online news; The classification probability of the theme scene of the comprehensive sample of online news is iterated several times until the likelihood function value converges, and the theme scene vector of the comprehensive sample of online news is obtained.
4. The method for arranging and generating information data analysis reports based on a large model according to claim 3, characterized in that: The theme scene vector of the comprehensive online news sample is specifically: Where, P(z k |d i ) is the kth topic scene vector of the i-th network news comprehensive sample, P(w j |z k ) is the probability distribution of the jth topic information parameter in the kth topic scenario, P(z k |w ij ,d i ) is the probability distribution of the kth topic scenario given the jth topic information parameter of the i-th network news comprehensive sample, δ(w,w ij ) is the indicator function, when w is equal to w ij When , its value is 1, otherwise it is 0, n is the total number of comprehensive samples of online news, and m is the total number of topic information.
5. The method for arranging and generating information data analysis reports based on a large model according to claim 4, characterized in that: Obtain comprehensive sample data of historical hot search news, determine the information parameters of hot search topics, analyze the keyword change trends in the information parameters of hot search topics, and evaluate the public opinion trend vector of the hot search theme scene. Specifically, the following steps are involved: Based on the comprehensive sample data of historical hot search news, the keywords in the hot search topic information parameters are marked according to the unit time, and the time series parameters of the topic information of historical hot search news are constructed; Based on the time series parameters of the topic information of historical hot search news, the TF-IDF value of each keyword is calculated to determine the topic information word heat distribution value of each historical hot search news; Based on text sentiment analysis, the sentiment tendency of the topic information word heat distribution value of each historical hot search news is evaluated, and the sentiment tendency distribution value of the topic information word heat distribution value of each historical hot search news is obtained; Based on ARIMA time series autoregression, a public opinion trend analysis model for historical hot search news is constructed; The topic information vocabulary heat distribution value of each historical hot search news and the sentiment tendency distribution value of the topic information vocabulary of each historical hot search news are substituted into the public opinion trend analysis model of historical hot search news. The objective function is constructed based on the difference between the predicted value and the actual value. The prediction error function of the training result is used as the end target to generate the public opinion trend vector of the hot search topic scene.
6. The method for arranging and generating information data analysis reports based on a large model according to claim 5, characterized in that: The public opinion trend analysis model of the historical hot search news is specifically as follows: Where Y k (t) is the public opinion trend vector of the kth hot search theme scene in t unit time, is the predicted value of the zth word heat distribution of the topic information of the hot search news in the historical t unit time, is the predicted value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, is the actual value of the zth word heat distribution of the hot search news topic information in the historical t unit time, is the actual value of the sentiment tendency distribution of the zth word in the topic information of the hot search news in the historical t unit time, u is the total number of topic information words in the hot search news, and T is the total time.
Citation Information
Patent Citations
News recommendation method based on mass news data event popularity
CN112199601A
News analysis method and system based on multi-modal large model
CN118535978A