Common text analysis method and system based on LDA model and LSTM algorithm
By applying the LDA model and LSTM algorithm in public text analysis, the problems of insufficient topic division, single attention analysis and simple sentiment analysis are solved, and higher analysis accuracy and intelligent analysis effects are achieved.
Patent Information
- Application Number
- CN202510206190.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
In the prior art, the topic classification is not accurate enough, the focus analysis dimension is single, and the sentiment analysis is too simple, resulting in low accuracy of text analysis.
The public text analysis method based on the LDA model and LSTM algorithm is adopted to obtain target text data, extract feature words to construct feature word vector matrix, build confusion curves and theme models, determine the topic attention and attention trend characteristics, and use pre-established sentiment analysis models for sentiment analysis.
It has achieved improvement in the accuracy of topic division, multi-dimensional attention analysis, comprehensive grasp of text evolution trends, and improved the accuracy of emotional classification, realizing intelligent analysis of public texts.
Smart Images

Figure CN119721052B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text analysis, and in particular, to a public text analysis method and system based on the LDA model and the LSTM algorithm. Background Art
[0002] With the continuous development of public management, the number of public texts has shown an explosive growth. How to quickly and accurately extract valuable information from a large amount of texts has become an important challenge faced by various departments and enterprises.
[0003] Currently, common public text analysis methods mainly include traditional methods such as manual reading and analysis and keyword matching. These methods often rely on manual experience, have low analysis efficiency, and are difficult to meet the processing requirements of large-scale text data.
[0004] In recent years, with the development of machine learning technology, some studies have begun to attempt to apply topic models to public text analysis. These methods achieve automatic identification and classification of text topics by calculating text features. However, these methods still have some limitations when dealing with public texts, such as the difficulty in determining the number of topics and the single dimension of attention analysis.
[0005] Therefore, the existing technology has the following technical problems: First, there is a lack of an effective topic optimization method, resulting in inaccurate topic division; second, there is a lack of a multi-dimensional attention analysis framework, making it difficult to comprehensively grasp the evolution trend of public texts; third, the sentiment analysis is too simple to accurately identify the specific orientation of the text. Therefore, the text analysis accuracy of the existing technology is relatively low. Summary of the Invention
[0006] In view of this, the present application provides a public text analysis method and system based on the LDA model and the LSTM algorithm, which solves the problems of inaccurate topic division, single dimension of attention analysis, and overly simple sentiment analysis in the existing technology.
[0007] An embodiment of the present application provides a public text analysis method based on the LDA model and the LSTM algorithm, including:
[0008] Obtain the target text data of a scientific and technological project;
[0009] Extract the feature words in the target text data, and construct a feature word vector matrix based on the feature words. The target text data includes word frequency data and inverse document frequency data;
[0010] Based on the feature word vector matrix, construct a perplexity curve, and construct a topic model based on the perplexity curve. The perplexity curve is used to indicate the relationship between the number of topics in the target text data and the perplexity, and the topic model is used to indicate multiple target topics obtained after performing topic clustering on the target text data;
[0011] Based on the topic model, determine the topic attention of the target topic, and determine the attention trend feature based on the topic attention. The topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend feature is used to indicate the change trend of the target topic in the time dimension, space dimension, and topic dimension;
[0012] Use a pre-established sentiment analysis model to analyze the target text data to obtain a sentiment analysis result;
[0013] Based on the target topic, the attention trend feature, and the sentiment analysis result, generate a text analysis result of the scientific and technological project.
[0014] Optionally, the sentiment analysis model is constructed through the following steps, including:
[0015] Perform sentence-level segmentation on the sample text data to obtain multiple text statements;
[0016] Identify the encouraging sentences in the text statements as the first sentiment identification value, identify the restrictive sentences in the text statements as the second sentiment identification value, and identify the ordinary sentences in the text statements as the third sentiment identification value;
[0017] Divide the labeled text statements into a training set and a test set according to a preset ratio;
[0018] Based on the training set, determine the training parameters in a preset LSTM model, where the training parameters include batch size, dropout rate parameter, number of training epochs, sequence length, vocabulary size, and number of output classes;
[0019] Based on the training parameters, use the training set to train the LSTM model;
[0020] Use the test set to test the trained LSTM model to obtain performance evaluation metrics, and combine the performance evaluation metrics to obtain the sentiment analysis model, where the performance evaluation metrics include accuracy, precision, recall, and F1 value.
[0021] Optionally, the determining the topic attention of the target topic based on the topic model and determining the attention trend feature based on the topic attention includes:
[0022] Based on the topic model and the target text data, determine the document topic attention of each record document in the target text data, where the document topic attention is used to indicate the probability distribution of each target topic in the record document within any time period;
[0023] Based on the document topic attention and the total number of topics within any time period, determine the topic attention of the target topic within the time period;
[0024] Based on the topic attention, perform time - dimension analysis to obtain time - distribution characteristics, where the time - distribution characteristics include annual trend analysis results, seasonal analysis results, and stage analysis results. The annual trend analysis results are used to indicate the attention trend and change characteristics of the target topic, the seasonal analysis results are used to indicate the release - time characteristics and periodic laws of the target topic, and the stage analysis results are used to indicate the evolution characteristics and key - transfer paths of the target topic under different project development stages;
[0025] Based on the topic attention, perform space - dimension analysis to obtain space - distribution characteristics, where the space - distribution characteristics include regional - distribution characteristics, provincial - distribution characteristics, and spatial - clustering distribution characteristics. The regional - distribution characteristics are used to indicate the topic - attention distribution in each region and the project differences between regions, the provincial - distribution characteristics are used to indicate the topic - attention conditions in each province and the project characteristics of a specified province, and the spatial - clustering distribution characteristics are used to indicate the geographical - distribution characteristics oriented by the scientific and technological projects and the regional collaborative - development trend;
[0026] Based on the topic attention, perform topic - dimension analysis to obtain topic - distribution characteristics, where the topic - distribution characteristics are used to indicate the importance ranking, topic correlation, and topic - evolution trend among the target topics;
[0027] Based on the time - distribution characteristics, the space - distribution characteristics, and the topic - distribution characteristics, perform multi - dimension cross - analysis to obtain attention - trend characteristics.
[0028] Optionally, the constructing a perplexity curve based on the feature - word vector matrix and constructing a topic model based on the perplexity curve includes:
[0029] Use the feature - word vector matrix to initialize the model parameters of a preset initial topic model, where the model parameters include the number of topics, hyperparameters, and the number of iterations;
[0030] Based on the model parameters and the feature - word vector matrix, perform iterative training on the initial topic model and record the perplexity corresponding to each number of topics to construct a perplexity curve;
[0031] Detect the change rate of the perplexity curve, and obtain the number of topics corresponding to the case where the change rate is lower than a preset change rate threshold, and construct a topic model based on the obtained number of topics.
[0032] Optionally, the extracting feature words from the target text data to construct a feature word vector matrix according to the feature words includes:
[0033] Based on the target text data, calculate the word frequency data of each word in the target text data, and the word frequency data is used to indicate the occurrence frequency of the word in the recorded document;
[0034] Based on the word frequency data, calculate the inverse document frequency data of the word, and the inverse document frequency data is used to indicate the importance of the word;
[0035] Based on the word frequency data and the inverse document frequency data, determine the word feature data of the word;
[0036] Based on the word feature data, screen out feature words.
[0037] Optionally, the training of the LSTM model using the training parameters includes:
[0038] Construct a feature enhancement model, where the feature enhancement model includes an encoder and a decoder, the encoder includes a Mamba block for extracting sequence features and a skip connection structure composed of a UNet structure, the skip connection structure is used to indicate the fusion of different scale features, and the decoder includes an LSTM layer for indicating the processing of temporal dependencies;
[0039] Map the training set to a word vector sequence;
[0040] Construct an input structure with sentence-level features and document-level features, and form an enhanced training sample set from the input structure;
[0041] Based on the enhanced training sample set, perform self-supervised pre-training on the feature enhancement model;
[0042] Introduce labeled data, and fine-tune the pre-trained feature enhancement model based on the enhanced training sample set to generate a feature processing model;
[0043] Use the feature processing model to perform feature enhancement on the word vector sequence to obtain an optimized training set;
[0044] Use the optimized training set to train the LSTM model.
[0045] Optionally, extracting feature words from the target text data and constructing a feature word vector matrix based on the feature words further includes:
[0046] Performing multi-level semantic analysis on the target text data and recursively subdividing the target text data into information granules of different granularities, where the information granules are used to indicate the hierarchical themes of the corresponding granularities;
[0047] Determining the granularity density of each text segment in the information granules at each granularity, where the granularity density is used to indicate the similarity between the text segment and the hierarchical theme of the corresponding granularity;
[0048] When it is detected that the granularity density is lower than a preset density threshold, marking the corresponding text segment as potential abnormal data;
[0049] Optimizing the information granules according to the potential abnormal data to obtain a text feature sequence;
[0050] Combining the text feature sequence, the granularity density, and the target text data to calculate the word frequency data.
[0051] Optionally, obtaining the target text data of the scientific and technological project includes:
[0052] Obtaining the text data to be processed of the scientific and technological project;
[0053] Constructing a data distribution optimization network, where the data distribution optimization network includes a sequence of normalization flow transformation layers, and an attention mechanism is arranged in each normalization flow transformation layer in the sequence of normalization flow transformation layers;
[0054] Mapping the text data to be processed to a preset normal distribution space to generate a preliminary mapping result;
[0055] Calculating the local density of data points in the preliminary mapping result based on a preset kernel density estimation model;
[0056] Constructing data weights based on the local density;
[0057] Performing density optimization on the preliminary mapping result based on the data weights to generate an optimized data distribution of density;
[0058] Constructing and executing a contrastive learning task, where the contrastive learning task is used to indicate extracting the structural features in the optimized data distribution of density;
[0059] Constructing a domain discriminator, where the domain discriminator is used to indicate the data distribution differences of different data sources;
[0060] Configuring a dynamic weight mechanism for indicating weight adjustment;
[0061] Process the density-optimized data distribution based on the contrastive learning task, the domain discriminator, and the dynamic weight mechanism to obtain target text data.
[0062] Optionally, the method further includes:
[0063] Generate a data histogram based on the target text data;
[0064] Generate a corresponding word cloud map based on the word features of the target text data, where the word features include the word frequency data and inverse document frequency data of the words;
[0065] Generate a corresponding attention time-series line chart based on the attention trend features;
[0066] Generate a corresponding sentiment pie chart based on the sentiment analysis results, where the sentiment analysis results include the proportion of sentiment categories;
[0067] Generate a corresponding regional feature distribution map based on the text analysis results of each region regarding the scientific and technological project;
[0068] Visually display the data histogram, the word cloud map, the attention time-series line chart, the sentiment pie chart, and the regional feature distribution map.
[0069] An embodiment of this application also provides a public text analysis system based on the LDA model and the LSTM algorithm, including:
[0070] A data acquisition module, configured to acquire target text data of a scientific and technological project;
[0071] A feature extraction module, configured to extract feature words from the target text data to construct a feature word vector matrix based on the feature words, where the target text data includes word frequency data and inverse document frequency data;
[0072] A model construction module, configured to construct a perplexity curve based on the feature word vector matrix, and construct a topic model based on the perplexity curve. The perplexity curve is used to indicate the relationship between the number of topics in the target text data and the perplexity, and the topic model is used to indicate multiple target topics obtained after topic clustering of the target text data;
[0073] An attention analysis module, configured to determine the topic attention of the target topic based on the topic model, and determine attention trend features based on the topic attention. The topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend features are used to indicate the change trend of the target topic in the time dimension, space dimension, and topic dimension;
[0074] An emotion analysis module for analyzing the target text data by using a pre-established emotion analysis model to obtain an emotion analysis result;
[0075] A text analysis module for generating a text analysis result of the scientific and technological project based on the target theme, the attention trend feature, and the emotion analysis result.
[0076] This application has the following technical effects:
[0077] By constructing a perplexity curve and a topic model through a feature word vector matrix, the number of topics is optimized, the accuracy of topic division is improved, and then the attention trend feature is determined through the topic model to achieve multi-dimensional attention analysis, and the overall trend of text evolution is comprehensively grasped. Through the emotion analysis model constructed by the LSTM deep learning model for emotion analysis, the accuracy of emotion classification is improved. Thus, through the attention trend feature and the emotion analysis result, the intelligent analysis of public text is realized, providing data support for relevant decisions. Description of the Drawings
[0078] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below:
[0079] Figure 1 It is a schematic flowchart of a public text analysis method based on the LDA model and the LSTM algorithm provided by an embodiment of this application;
[0080] Figure 2 It is a schematic flowchart of feature word extraction provided by an embodiment of this application;
[0081] Figure 3 It is a schematic flowchart of topic model construction provided by an embodiment of this application;
[0082] Figure 4 It is a schematic structural diagram of a public text analysis system based on the LDA model and the LSTM algorithm in an embodiment of this application. Detailed Embodiments
[0083] As Figure 1 It is a schematic flowchart of a public text analysis method based on the LDA model and the LSTM algorithm provided by an embodiment of this application. An embodiment of this application provides a public text analysis method based on the LDA model and the LSTM algorithm, including the following steps S1 to S6.
[0084] S1: Obtain the target text data of the scientific and technological project.
[0085] In a specific implementation, step S1 first needs to establish a data distribution optimization network. This network adopts a Normalizing Flow architecture, which includes multiple sequences of transformation layers, and each transformation layer is equipped with an attention mechanism to adaptively adjust the importance of different feature dimensions. For example, when processing technology project texts, the network can pay more attention to feature dimensions related to "technological innovation", "R & D investment", etc.
[0086] The data acquisition process uses distributed crawler technology to collect raw text data from the official websites of various technology projects. The collected content includes fields such as project titles, release times, release departments, and body contents. To ensure data quality, the system will perform preliminary data cleaning to remove duplicate and incorrect entries as well as irrelevant chart data. Approximately 4000 effective target text data can be obtained at this stage.
[0087] Next, the text data to be processed is mapped to a preset normal distribution space. Optionally, this step is achieved through a reversible neural network transformation to ensure the continuity and interpretability of the data distribution. During the mapping process, the system uses the Kernel Density Estimation (KDE) method to calculate the local density of data points, selects the Radial Basis Function (RBF) as the kernel function, and determines the bandwidth parameter through cross - validation.
[0088] Subsequently, based on the calculated local density values, data weights are constructed. The weight calculation comprehensively considers factors such as the coverage and innovation degree of the text, and adopts a weighted combination method of density values and expert scores. These weights are then used to guide the sampling process to ensure that more important samples have a higher probability of being selected.
[0089] To improve the model's adaptability to data from different sources, a contrastive learning task and a domain discriminator are also introduced. The contrastive learning task generates text representations from multiple perspectives (such as synonym replacement, sentence rearrangement, etc.) to learn the internal structure of the data. The domain discriminator adopts the Gradient Reversal Layer technology to prompt the feature extractor to generate feature representations independent of the data source.
[0090] Finally, through a dynamic weight mechanism, the weights of different tasks in the total loss function are dynamically adjusted according to the training progress. For example, more attention is paid to feature extraction in the initial stage of training, and more attention is paid to domain adaptation in the later stage. After these processes, high - quality target text data is finally output, providing a reliable data basis for subsequent topic analysis and sentiment analysis.
[0091] Exemplarily, the experimental results show that this data acquisition method can significantly improve the data quality. The utilization rate of effective samples has increased by 25%, the deviation of data distribution has decreased by 30%, and the detection accuracy of abnormal samples has reached 95%. At the same time, the cross-domain generalization ability has increased by 20%, the precision of feature extraction has increased by 15%, and the model training speed has also accelerated by 35%.
[0092] In some embodiments, step S1 includes:
[0093] Obtain the text data to be processed of the scientific and technological project;
[0094] Construct a data distribution optimization network, where the data distribution optimization network includes a sequence of normalization flow transformation layers, and an attention mechanism is arranged in each normalization flow transformation layer in the sequence of normalization flow transformation layers;
[0095] Map the text data to be processed to a preset normal distribution space to generate a preliminary mapping result;
[0096] Based on a preset kernel density estimation model, calculate the local density of data points in the preliminary mapping result;
[0097] Construct data weights according to the local density;
[0098] Perform density optimization on the preliminary mapping result based on the data weights to generate a density-optimized data distribution;
[0099] Construct and execute a contrastive learning task, and the contrastive learning task is used to indicate extracting structural features in the density-optimized data distribution;
[0100] Construct a domain discriminator, and the domain discriminator is used to indicate the data distribution differences of different data sources;
[0101] Configure a dynamic weight mechanism for indicating weight adjustment;
[0102] Based on the contrastive learning task, the domain discriminator, and the dynamic weight mechanism, process the density-optimized data distribution to obtain the target text data.
[0103] In this embodiment, a distributed crawler framework is constructed in the stage of obtaining text data to be processed. This framework adopts a master-slave architecture, where the master node is responsible for task distribution and scheduling, and the slave nodes are responsible for specific data collection tasks. The crawler system supports access to multiple data sources, including scientific and technological project management platforms, scientific and technological project declaration systems, etc. Optionally, it is implemented using the Scrapy framework, configured with mechanisms such as automatic retry, proxy switching, and request frequency control to ensure the stability and reliability of data collection. The collected raw data includes fields such as project name, release time, release department, project type, project level, funding amount, and project summary.
[0104] Furthermore, for the normalizing flow transformation layer sequence, each normalizing flow transformation layer consists of three sub-modules, including an affine transformation module, a non-linear activation module, and an attention module (i.e., attention mechanism). Optionally, the affine transformation uses an invertible linear transformation to ensure no information loss. The non-linear activation module selects the LeakyReLU function to avoid gradient vanishing. The attention module adopts a multi-head self-attention mechanism with the number of heads set to 8 and the hidden layer dimension to 512. Therefore, in this embodiment, by constructing a data distribution optimization network containing a normalizing flow transformation layer sequence, it can adaptively focus on different feature dimensions and improve the expressive ability of the model.
[0105] Subsequently, first, the text data to be processed is preprocessed, including word segmentation, stop word removal, part-of-speech tagging, etc. Then, a pre-trained Word2Vec model is used to convert the text into a vector representation with the vector dimension set to 300. Next, these vectors are mapped to the standard normal distribution space through the constructed data distribution optimization network to generate a preliminary mapping result. Among them, the mapping process is optimized by maximum likelihood estimation, using the Adam optimizer with a learning rate set to 0.001 and a batch size of 64. To ensure the stability of the mapping, a batch normalization layer and a residual connection are also introduced.
[0106] Furthermore, an improved kernel density estimation algorithm is adopted to determine the kernel density estimation model, and thus determine the local density of the data points in the preliminary mapping result. Specifically, this kernel density estimation model uses the Gaussian kernel function as the basic kernel function, and the kernel bandwidth is adaptively determined by the Silverman criterion. To improve the computational efficiency, the Ball-Tree data structure is used for nearest neighbor search, reducing the computational complexity from O(n²) to O(nlogn). At the same time, an adaptive bandwidth mechanism is introduced, using a larger bandwidth in sparse data regions and a smaller bandwidth in dense regions to improve the estimation accuracy.
[0107] Further, a multi-level weight calculation scheme is determined based on the calculated local density. First, the local density is scaled to the [0,1] interval through min-max normalization. Then, a soft threshold mechanism is introduced, and the basic weight is set as:
[0108] w = 1 / (1+exp(-(d-μ) / σ))
[0109] Among them, w represents the weight of the basic data, d is the normalized local density, and μ and σ are the density mean and standard deviation respectively. In addition, a time decay factor is considered, and recent data is given a higher weight. The final data weight is a weighted combination of the basic data weight and the time weight.
[0110] After the above processing, this embodiment can effectively handle the problem of unbalanced data distribution. Experiments show that the uniformity of the processed data distribution has increased by 40%, the detection accuracy of abnormal samples has reached 92%, and the data quality has been significantly improved. The rationality of the weight calculation is verified through cross-validation. In the subsequent topic analysis task, the model performance of the weighted samples has increased by 15% compared to the unweighted samples.
[0111] Furthermore, in the density optimization stage, a density optimization algorithm based on importance sampling can be adopted. Specifically, when implementing, first construct a priority queue data structure and sort the data samples in descending order of weight. Then, use a stratified sampling strategy to divide the samples into three density levels: high, medium, and low, and set the sampling probabilities to 0.5, 0.3, and 0.2 respectively. To maintain the diversity of the data, a systematic sampling method is adopted within each level, that is, samples are selected at fixed intervals. At the same time, a local density constraint is introduced to ensure that the density difference between adjacent samples does not exceed a preset threshold (the empirical value is set to 0.3). Therefore, through density optimization, this embodiment not only ensures the retention of high-quality samples but also avoids over-concentration of the data distribution.
[0112] Optionally, this embodiment also configures four data augmentation strategies: (1) Synonym replacement, using a pre-trained word vector model to select words with a similarity greater than 0.8 for replacement; (2) Back-translation augmentation, generating semantically similar expressions through Chinese-English back-translation; (3) Sentence rearrangement, adjusting the sentence order while maintaining semantic coherence; (4) Random cropping, retaining 60%-90% of the original text content. Generate two augmented views for each sample, and optimize the feature extractor through a contrastive loss function, with the temperature parameter set to 0.07. The feature extractor adopts a two-layer Transformer structure, with a hidden layer dimension of 768 and 12 attention heads.
[0113] In this embodiment, a domain discriminator is constructed. The domain discriminator is used to indicate the data distribution differences of different data sources. The construction of the domain discriminator adopts the idea of adversarial learning. Among them, the discriminator network consists of three fully connected layers, and the dimensions of the middle layers are 512 and 256 respectively. The LeakyReLU activation function and Dropout regularization (ratio 0.3) are used. To enhance the discrimination ability, a feature normalization layer and a residual connection are added. During the training process, the gradient reversal layer technique is used to make the feature extractor generate domain-independent feature representations through adversarial training. The loss function uses cross-entropy loss, and an L2 regularization term is introduced to prevent overfitting.
[0114] S1.9: Configure a dynamic weight mechanism for indicating adjusted weights;
[0115] Exemplarily, the dynamic weight mechanism is configured with an adaptive weight adjustment strategy. Three weight coefficients are maintained: the contrast learning weight λ1, the domain adversarial weight λ2, and the data reconstruction weight λ3, and their initial values are set to 0.5, 0.3, and 0.2 respectively. During the training process, these weights are dynamically adjusted based on the performance of the validation set. Specifically, after each training epoch, the loss change rate of each task on the validation set is calculated, the weights of the tasks with poorer performance are increased, and the weights of the tasks with better performance are decreased. The adjustment step size adopts a cosine annealing strategy to ensure the stability of training.
[0116] Thus, based on the contrast learning task, the domain discriminator, and the dynamic weight mechanism, the processed data distribution after density optimization is processed to obtain the target text data. Specifically, the final data processing stage adopts a multi-task joint optimization method. The total loss function L is defined as:
[0117] L = λ1Lcon + λ2Ladv + λ3Lrec
[0118] where Lcon is the contrast learning loss, Ladv is the domain adversarial loss, and Lrec is the reconstruction loss. The optimizer is selected as AdamW, the learning rate is set to 2e-5, and the weight decay is 0.01. To improve the training efficiency, a gradient accumulation strategy is adopted, and the accumulation step is 4. An early stopping strategy is used during the training process, and the training is stopped when the performance of the validation set has not improved for 5 consecutive epochs. Finally, an ensemble learning method is used to fuse the prediction results of different checkpoints to obtain the final target text data.
[0119] Experimental results show that this method performs excellently in processing scientific and technological project text data. Compared with traditional methods, the feature extraction accuracy has increased by 23%, the domain adaptation ability has increased by 31%, and the data quality has increased by 28%. In downstream tasks such as topic classification and sentiment analysis, significant performance improvements have also been achieved, and the F1 score has increased by 18% on average.
[0120] S2: Extract the feature words from the target text data, and construct a feature word vector matrix based on the feature words. The target text data includes word frequency data and inverse document frequency data.
[0121] In this application, the extraction of feature words and the construction of the feature word vector matrix are key steps in text analysis. First, preprocess the target text data, including text tokenization, stop word removal, and word stemming. Among them, for tokenization, the custom mode of the jieba tokenizer is used and optimized in combination with a professional dictionary in the field of scientific and technological projects. The professional dictionary contains approximately 20,000 common terms in the scientific and technological fields, covering key fields such as artificial intelligence, biomedicine, and new materials. When removing stop words, in addition to using a general stop word list, specific stop words are supplemented according to the characteristics of scientific and technological project texts, such as common words with little contribution to topic analysis like "project" and "declaration". In the word stemming stage, a method combining rules and statistics is adopted. For English terms, the Porter stemming algorithm is used for processing; for Chinese words, normalization is carried out through a thesaurus of synonyms. For example, expressions such as "artificial intelligence", "AI", and "artificial wisdom" will be unified and normalized to "artificial intelligence". Therefore, this application can significantly reduce the dimension of the feature space and improve the efficiency of subsequent analysis.
[0122] Next, this application calculates the term frequency (TF, Term Frequency) and the inverse document frequency (IDF, Inverse Document Frequency). Optionally, when constructing the feature word vector matrix, an improved TF-IDF weight calculation method is adopted. First, multiply the TF value of the term frequency and the IDF value of the inverse document frequency to obtain the basic TF-IDF score. Then, introduce a position weight factor and a topic relevance factor for adjustment. The position weight considers the position distribution of words in the document, and words in the title and at the beginning and end of paragraphs obtain higher weights; the topic relevance is obtained by calculating the cosine similarity with a predefined topic word list, and the weights of words with high relevance will be appropriately increased.
[0123] To improve the expression ability of features, this application also integrates word vector technology. Use a pre-trained Word2Vec model (the training corpus contains 5 million scientific and technological literatures) to map each feature word into a 300-dimensional dense vector. These vectors are reduced in dimension through principal component analysis (PCA), and the feature dimensions required to explain 95% of the variance are retained (usually between 100 and 150 dimensions). The final feature word vector matrix is a combination of TF-IDF weights and word vector representations, and weighted fusion is carried out through an attention mechanism.
[0124] The experimental results show that compared with the traditional TF-IDF method, the expression ability of features has been improved by 35%, the clustering effect (measured by the silhouette coefficient) in the subsequent topic clustering task has been improved by 28%, and the computing efficiency has been improved by 40%. At the same time, by introducing the position weight and the topic relevance factor, the discriminative ability of features has been significantly enhanced, and the accuracy rate in the keyword extraction task has reached more than 85%.
[0125] Therefore, this application can understand the decision-making basis of the model by analyzing the weight distribution of feature words, which has important value for scientific and technological project managers to understand and apply the analysis results.
[0126] In some embodiments, step S2 includes:
[0127] Based on the target text data, calculate the word frequency data of each word in the target text data, and the word frequency data is used to indicate the occurrence frequency of the word in the recorded document;
[0128] Based on the word frequency data, calculate the inverse document frequency data of the word, and the inverse document frequency data is used to indicate the importance degree of the word;
[0129] Based on the word frequency data and the inverse document frequency data, determine the word feature data of the word;
[0130] Based on the word feature data, screen out the feature words.
[0131] In this embodiment, in the word frequency data calculation stage, the text is first standardized. Exemplarily, the standardization process includes: removing special characters and numbers (retaining numbers helpful for understanding the content, such as "5G", "IPv6", etc.), unifying the character encoding to UTF-8, and converting full-width characters to half-width characters. Then, an improved word frequency calculation method is adopted, which not only considers the original occurrence times of words but also introduces a position weighting mechanism. Specifically, the word frequency calculation formula is:
[0132] TF(t,d) = α * f_title(t,d) + β * f_body(t,d) + γ * f_key(t,d)
[0133] Among them, TF(t,d) represents the word frequency, f_title, f_body, and f_key respectively represent the occurrence frequencies of the word in the title, body, and keyword parts, and the weight coefficients α, β, and γ are respectively set to 0.4, 0.3, and 0.3. This weighting method can better reflect the importance degree of words in different text positions.
[0134] Furthermore, when calculating the inverse document frequency data, a two-layer IDF calculation strategy is adopted. First, the conventional IDF value is calculated, and then the field weight adjustment is introduced, that is, the IDF value is corrected according to the importance of the word in the science and technology project field dictionary. For example, hot words in the fields of "artificial intelligence" and "blockchain" will obtain higher field weights. In addition, this embodiment also establishes a dynamic update mechanism to regularly update the IDF value based on newly added documents to ensure that it reflects the latest distribution of word importance.
[0135] Furthermore, in the process of determining word feature data, word frequency data and inverse document frequency data are combined to construct a multi-dimensional feature vector. In addition to the basic TF-IDF value, the feature vector also contains the following dimensions: 1) syntactic features of words, such as part of speech, syntactic dependencies, etc.; 2) semantic features, using the contextual semantic representation extracted by the BERT model; 3) statistical features, such as word distribution skewness, kurtosis and other statistics. The calculation of feature vectors uses a distributed processing framework, which divides large-scale text data into multiple batches for parallel processing, significantly improving processing efficiency.
[0136] Thus, based on the word feature data, feature words are screened. Among them, in the feature word screening stage, a multi-layer screening mechanism is configured, including: threshold screening, that is, setting the minimum threshold of the TF-IDF value (the empirical value is 0.01) to filter out low-importance words; entropy screening, that is, calculating the information entropy of the words in the document collection, retaining words with large information content (the entropy threshold is set to 1.5); co-occurrence analysis, that is, building a word co-occurrence network, calculating the centrality index of the words, and retaining the words corresponding to important nodes in the network; expert rules, that is, integrating the screening rules defined by domain experts, such as the retention strategy of words in a specific field; part-of-speech screening, that is, giving priority to retaining nouns, verbs and other content words, and appropriately retaining adverbs and adjectives that can reflect the characteristics of the document. Subsequently, a voting mechanism can be used for screening, that is, words need to meet multiple screening conditions at the same time to be finally retained. At the same time, this application regularly evaluates the quality of the screening results (i.e., feature words), and asks domain experts to evaluate and feedback the screening results through random sampling, and dynamically adjusts the screening parameters according to the feedback results.
[0137] Based on the above embodiments, in some embodiments, such as Figure 2 As shown, step S2 also includes:
[0138] S2.1: performing a multi-level semantic analysis on the target text data, and recursively subdividing the target text data into information particles of different granularities, wherein the information particles are used to indicate a hierarchical theme of the granularity;
[0139] S2.2: Determine the particle density of each text segment in the information particle at each granularity, where the particle density is used to indicate the similarity between the text segment and the hierarchical theme to which the granularity belongs;
[0140] S2.3: When it is detected that the particle density is lower than a preset density threshold, mark the corresponding text segment as potentially abnormal data;
[0141] S2.4: Optimize the information particle according to the potentially abnormal data to obtain a text feature sequence;
[0142] S2.5: Combine the text feature sequence, the particle density, and the target text data to calculate the word frequency data.
[0143] In this embodiment, in the multi-level semantic analysis and information particle division stage, a hierarchical topic modeling method can be adopted. First, construct four semantic levels: document level, paragraph level, sentence level, and phrase level. At each level, use an improved Hierarchical Latent Dirichlet Allocation (LDA) model for topic extraction. The hyperparameters of the model are optimized through grid search. The number of topics is set to 20 at the document level, 50 at the paragraph level, 100 at the sentence level, and 200 at the phrase level. To improve the accuracy of topic modeling, this embodiment integrates the semantic representation of the pre-trained language model BERT, and fuses the context representation of BERT with the topic representation of LDA through an attention mechanism. The recursive subdivision process adopts a top-down strategy to ensure semantic consistency between the upper-level topics and the lower-level topics.
[0144] Furthermore, in the particle density calculation stage, a multi-dimensional density calculation method can be adopted. For each text segment, calculate its similarity with the corresponding hierarchical theme, specifically including the following dimensions: semantic similarity, that is, use Word2Vec to calculate the cosine similarity between the text segment and the topic words; topic probability, that is, calculate the probability distribution of the text segment belonging to each topic based on the LDA model; structural similarity, that is, consider the position and context relationship of the text segment in the document structure; lexical overlap, that is, calculate the lexical overlap degree between the text segment and the topic feature words. The final particle density is obtained through the weighted combination of these dimensions, as shown in the following formula:
[0145] Density = Σ(wi * si)
[0146] where Density represents the particle density, wi is the dimension weight of each dimension, and si is the corresponding similarity score. The weights are determined through cross-validation and are 0.4, 0.3, 0.2, and 0.1 respectively.
[0147] In this regard, an adaptive threshold mechanism can be adopted in the abnormal data detection and marking stage. It should be noted that the density threshold is not a fixed value and can be dynamically adjusted according to the density distribution of each level. Exemplarily, on each level, the kernel density estimation (KDE) method is used to fit the distribution of particle density, and the samples below the mean of the distribution minus 1.5 times the standard deviation are marked as potential anomalies. At the same time, this embodiment also considers the correlation between levels. If a text segment is marked as abnormal at multiple levels, its abnormal probability will increase accordingly.
[0148] Furthermore, according to the potential abnormal data, the information granules are optimized to obtain a text feature sequence. Among them, in the information granule optimization stage, this embodiment processes the marked abnormal data. First, the abnormal segment is corrected by context reconstruction, that is, the semantic information of adjacent text segments is used to supplement and correct the abnormal segment. If the density after reconstruction is still low, then consider merging this segment with other segments with similar semantics or demoting it to a lower level for processing. The optimization process adopts an iterative strategy, and the density is recalculated after each optimization until the threshold requirement is met or the maximum number of iterations is reached, obtaining a text feature sequence.
[0149] Thus, combining the text feature sequence, particle density, and target text data, the word frequency data is calculated. That is, in the final word frequency calculation stage, this application considers the text feature sequence, particle density, and original text data. The word frequency calculation formula is:
[0150] TF(t) = base_freq(t) * (1 + α * density_score(t)) * (1 + β * seq_weight(t))
[0151] Among them, TF(t) represents the word frequency data calculated in this embodiment, base_freq(t) is the basic word frequency, density_score(t) is the adjustment factor based on particle density, seq_weight(t) is the weight adjustment based on the feature sequence, and α and β are adjustment parameters, which are set to 0.3 and 0.2 through experiments. Therefore, this embodiment can better reflect the importance of words in different semantic levels and topic structures.
[0152] S3: Based on the feature word vector matrix, a perplexity curve is constructed, and a topic model is constructed based on the perplexity curve. The perplexity curve is used to indicate the relationship between the number of topics in the target text data and the perplexity, and the topic model is used to indicate that after the target text data is subject to topic clustering, multiple target topics are obtained.
[0153] In this application, for constructing a topic model, the appropriate number of topics K is a key parameter determining the model's performance. Too few topics will result in overly coarse topic partitioning and fail to reflect the detailed features of the text. Too many topics will lead to topic overlap and increase the difficulty of interpretation. Therefore, this application determines the optimal number of topics through Perplexity analysis.
[0154] First, this application determines a reasonable search range for the number of topics. Based on experience and literature research, the number of topics for science and technology policy texts is generally more appropriate between 5 and 15. Too few topics are difficult to cover multiple fields of the policy, while too many topics will reduce the distinguishability between topics. Therefore, this application sets the search range from 5 to 15 with a step size of 1, that is, it evaluates the model performance when the number of topics is 5, 6, 7... 15 respectively.
[0155] For each candidate number of topics K, this application needs to fully train an LDA model (i.e., the topic model) and calculate its perplexity. The perplexity calculation formula is: perplexity(D) = exp{-∑log p(wd) / N}, where perplexity(D) represents the perplexity, D is the test set document collection (i.e., the target text data), wd is the word in the target text data, and N is the total number of words. A lower perplexity indicates a stronger prediction ability of the model for the text. In the Python implementation, this application uses the LdaModel class of the gensim library, sets fixed hyperparameters (α = 50 / k, β = 0.01), trains the model through multiple iterations (such as 1000 times), and records the perplexity corresponding to each K value. Subsequently, the perplexities corresponding to different numbers of topics are plotted into a curve graph to form a perplexity curve. Among them, the horizontal axis of the perplexity curve is the number of topics K, and the vertical axis of the perplexity curve is the perplexity value. Usually, as the number of topics increases, the perplexity will first decrease rapidly and then level off. This "inflection point" phenomenon indicates that continuing to increase the number of topics will no longer significantly improve the model performance. Exemplarily, by observing the perplexity curve, it is found that when the number of topics is 9, the perplexity begins to level off. Therefore, K = 9 is selected as the final number of topics.
[0156] To verify the rationality of this selection, this application also needs to conduct a qualitative analysis in combination with the interpretability of the topic. By examining the topic-word distribution when K = 9, it is found that each topic has clear semantic features and can better reflect different aspects of science and technology policies, such as scientific research funds, talent cultivation, market supervision, etc., which further confirms the rationality of the optimal number of topics obtained through perplexity analysis. In practical applications, this application recommends using the cross-validation method to calculate the perplexity multiple times to reduce the influence of randomness. At the same time, other evaluation indicators such as Topic Coherence can be combined to comprehensively determine the final number of topics. In addition, the model needs to be updated regularly according to the newly added policy texts and the optimal number of topics needs to be re-evaluated to adapt to the changes in the policy environment.
[0157] It should be noted that according to the topic model, topic clustering and analysis are carried out to obtain the target topics. Exemplarily, the topics are classified into 4 major categories: First, related to funds, including requirements for the use of scientific research funds, application for project funds, venture capital investment and loans, and evaluation and rewards for scientific research achievements; Second, talent education, including student education, cooperation between university laboratories and talents, and talent cultivation; Third, agricultural science and technology, including popularization and application of agricultural science and technology; Fourth, market supervision, including production quality and market supervision and management.
[0158] In some embodiments, as Figure 3 shown, this step S3 includes:
[0159] S3.1: Use the feature word vector matrix to initialize the model parameters of the preset initial topic model, where the model parameters include the number of topics, hyperparameters, and the number of iterations;
[0160] S3.2: Based on the model parameters and the feature word vector matrix, perform iterative training on the initial topic model and record the perplexity corresponding to each number of topics to construct a perplexity curve;
[0161] S3.3: Detect the change rate of the perplexity curve and obtain the number of topics corresponding to when the change rate is lower than the preset change rate threshold, and construct a topic model based on the obtained number of topics.
[0162] In this embodiment, an adaptive parameter configuration strategy is adopted in the stage of initializing the topic model parameters. First, the search range of the number of topics K is preliminarily set. According to the scale of the dataset and the characteristics of the domain, the search range is set to [5, 100]. For hyperparameters (such as the Dirichlet prior parameter of the document-topic distribution and the Dirichlet prior parameter of the topic-word distribution), a grid search method is used for optimization. The search range of the Dirichlet prior parameter of the document-topic distribution is [0.01, 1.0], and the search range of the Dirichlet prior parameter of the topic-word distribution is [0.01, 0.5], with a step size of 0.01 for both. The number of iterations is set using a dynamic adjustment mechanism. The basic number of iterations is set to 1000, and at the same time, a convergence condition is set: stop iterating when the change in log-likelihood for 10 consecutive iterations is less than 1e-6.
[0163] Optionally, to improve the quality of model initialization, this embodiment also introduces prior information based on expert knowledge. Specifically, according to the domain classification system of scientific and technological projects, a seed word list is constructed, and 10 - 20 high-confidence seed words are configured for each potential topic. These seed words are used to guide the generation of the initial topic-word distribution, improving the model convergence speed and topic quality. At the same time, to handle new words and cross-domain terms, the system reserves a certain proportion (about 20%) of free topic space, allowing the model to discover new topic structures during training.
[0164] Furthermore, based on the model parameters and the feature word vector matrix, the initial topic model is iteratively trained, and the perplexity corresponding to each number of topics is recorded to construct a perplexity curve. Among them, in the model training and perplexity calculation stage, this embodiment uses an improved Gibbs sampling algorithm for model training. To improve the training efficiency, a distributed training architecture is introduced, and the dataset is divided into multiple subsets for parallel processing. Each training round includes the following steps: reassigning the topics of words using Collapsed Gibbs Sampling; updating the topic-word distribution using the variational Bayesian method; calculating the perplexity and evaluating the model performance using the cross-validation method.
[0165] Optionally, during the calculation of the perplexity, penalty terms for topic coherence and topic diversity are added. To ensure the stability of the evaluation, each number of topics is configured to run 5 times, and the average perplexity value is taken. At the same time, this embodiment also records auxiliary indicators such as Topic Coherence and Topic Distinctiveness.
[0166] Furthermore, the change rate of the detection perplexity curve is detected, and the number of corresponding topics when the change rate is lower than the preset change rate threshold is obtained, and a topic model is constructed based on the obtained number of topics. Among them, in the stage of determining the number of topics, this embodiment configures a multi-index comprehensive evaluation mechanism. In addition to the traditional perplexity change rate, the following factors are also considered:
[0167] Topic consistency change: Use PMI (Point Mutual Information) and NPMI (Normalized Point Mutual Information) to evaluate the semantic consistency of topics;
[0168] Topic coverage: Evaluate the coverage of topics for the document set;
[0169] Topic distinctiveness: Calculate the JS divergence between topics to ensure sufficient distinctiveness between topics;
[0170] Computational complexity: Consider the requirements of different numbers of topics for computing resources.
[0171] Optionally, the setting of the change rate threshold adopts an adaptive method. The basic threshold is set to 0.01, but it will be dynamically adjusted according to the dataset size and topic quality metrics. This embodiment uses a sliding window (window size is 5) to calculate the change rate of perplexity. When the change rates of three consecutive windows are all lower than the threshold, it is considered that the appropriate number of topics has been found.
[0172] Optionally, in order to verify the rationality of the topic number selection, this embodiment conducts a stability analysis, including the following steps:
[0173] Perform multiple random samplings on the data to verify the stability of the topic number selection;
[0174] Use different initialization parameters to test the robustness of the results;
[0175] By manually evaluating the sampled topic content, ensure the interpretability of the topics.
[0176] In this regard, the experimental results show that this method for determining the number of topics performs excellently in the analysis of scientific and technological project texts. Compared with the fixed number of topics or simple heuristic methods, the number of topics determined in this embodiment can better balance the model complexity and topic quality. Specifically, the topic consistency has increased by 35%, the topic coverage has increased by 28%, while maintaining a relatively low computational overhead. In practical applications, this embodiment can accurately capture the topic structure of scientific and technological project texts, providing a reliable basis for subsequent project classification and topic analysis.
[0177] S4: Based on the topic model, determine the topic attention of the target topic, and determine the attention trend feature based on the topic attention. The topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend feature is used to indicate the change trend of the target topic in the time dimension, space dimension, and topic dimension.
[0178] In this application, the calculation of topic attention can adopt a multi-dimensional analysis method. Exemplarily, first construct a time window sequence, set the basic time unit as month, and at the same time support aggregation analysis of multiple time granularities such as quarter, half-year, and year. For each time window, calculate the weight distribution of each topic. The formula for calculating topic attention is:
[0179] Attention(t,i) = λ1Doc_Ratio(t,i) + λ2Citation_Score(t,i) + λ3*Impact_Factor(t,i)
[0180] Among them, Attention(t,i) represents topic attention, t represents the time window, i represents the topic, and λ1, λ2, and λ3 are weight coefficients (determined to be 0.4, 0.3, and 0.3 through cross-validation). Doc_Ratio(t,i) represents the proportion of the topic in the document set, Citation_Score(t,i) reflects the citation influence of relevant documents, and Impact_Factor(t,i) represents external factors such as project funding amount and level of participating units.
[0181] Furthermore, in the trend analysis in the time dimension, this application adopts time series decomposition technology to decompose the topic attention sequence into a trend term, a seasonal term, and a random term. The trend term is extracted and smoothed using LOESS (Locally Weighted Scatterplot Smoothing), and the seasonal term is identified for periodic patterns through the STL (Seasonal and Trend decomposition using Loess) method. Optionally, in order to capture sudden changes, this application also introduces a change point detection algorithm and uses the PELT (Pruned Exact Linear Time) method to identify significant change points in attention. These change points usually correspond to key events such as major policy adjustments, technological breakthroughs, or market reforms.
[0182] In the analysis of the spatial dimension, this application constructs a geographical information association network. First, the geographical locations of project execution units are mapped to provincial and municipal administrative divisions, and then the contribution intensities of different regions to different themes are calculated. The spatial distribution characteristics are quantified through spatial autocorrelation analysis, and the Moran's I index is used to evaluate the spatial aggregation degree of theme attention. At the same time, by constructing a spatial weight matrix, the geographical paths of theme propagation and the regional innovation network structure are identified. This application also introduces a regional innovation ability index to adjust the theme weights of different regions to ensure that the analysis results can reflect the differences in innovation strength among regions.
[0183] In the analysis of the theme dimension, this application adopts a dynamic theme evolution model. First, a theme similarity matrix is constructed, and the Jensen-Shannon divergence is used to measure the semantic distance between themes. Based on this matrix, theme groups are identified through hierarchical clustering methods, and the evolution process of these groups over time is tracked. The theme evolution patterns include:
[0184] Theme continuation: The core semantics remain stable, but the expressions may change over time;
[0185] Theme splitting: One theme differentiates into multiple sub-themes;
[0186] Theme merging: Multiple related themes merge into a larger theme category;
[0187] Theme extinction: The attention continuously decreases until it disappears;
[0188] Theme emergence: The emergence of new technologies or new concepts leads to the formation of new themes.
[0189] To improve the accuracy of trend prediction, the system integrates multiple prediction models:
[0190] ARIMA model: Deals with linear time series characteristics;
[0191] Prophet model: Captures trends with seasonal and holiday effects;
[0192] LSTM neural network: Learns complex non-linear patterns;
[0193] XGBoost model: Integrates external features for prediction.
[0194] These models are combined through an ensemble learning method, and the weights are dynamically adjusted according to the historical prediction performance of each model. The prediction results also consider the confidence interval, and Monte Carlo simulation is used to generate possible development paths.
[0195] Optionally, the present application also establishes a trend warning mechanism, which triggers a warning when the following situations are detected: sudden change in attention: the change in attention within a short period exceeds 2 times the historical standard deviation; trend reversal: a significant change occurs in the long-term upward / downward trend; spatial anomaly: an abnormal fluctuation occurs in the attention distribution in a specific area; theme anomaly: a sudden change occurs in the theme structure.
[0196] In some embodiments, step S4 includes:
[0197] Based on the theme model and the target text data, determine the document theme attention of each record document in the target text data, where the document theme attention is used to indicate the probability distribution of each target theme in the record document within any time period;
[0198] Based on the document theme attention and the total number of themes within any time period, determine the theme attention of the target theme within the time period;
[0199] Based on the theme attention, perform time dimension analysis to obtain time distribution characteristics, where the time distribution characteristics include annual trend analysis results, seasonal analysis results, and stage analysis results. The annual trend analysis results are used to indicate the attention trend and change characteristics of the target theme, the seasonal analysis results are used to indicate the release time characteristics and periodic laws of the target theme, and the stage analysis results are used to indicate the evolution characteristics and key transfer paths of the target theme under different project development stages;
[0200] Based on the theme attention, perform spatial dimension analysis to obtain spatial distribution characteristics, where the spatial distribution characteristics include regional distribution characteristics, provincial distribution characteristics, and spatial clustering distribution characteristics. The regional distribution characteristics are used to indicate the theme attention distribution in each region and the project differences between regions, the provincial distribution characteristics are used to indicate the theme attention situation in each province and the project characteristics of a specified province, and the spatial clustering distribution characteristics are used to indicate the geographical distribution characteristics oriented by the scientific and technological projects and the regional coordinated development trend;
[0201] Based on the theme attention, perform theme dimension analysis to obtain theme distribution characteristics, where the theme distribution characteristics are used to indicate the importance ranking, theme relevance, and theme evolution trend among the target themes;
[0202] Based on the time distribution characteristics, the spatial distribution characteristics, and the theme distribution characteristics, perform multi-dimensional cross-analysis to obtain attention trend characteristics.
[0203] In this embodiment, the document theme attention is determined by the probability distribution of theme k in document d at time t. Subsequently, based on the total number of themes within any time period and the document theme attention, the theme attention is determined.
[0204] Furthermore, for the analysis of the time dimension, it includes: annual trend analysis, such as aggregating theme distribution data by year, calculating the annual average attention of each theme, plotting the attention trend graph from 2014 to 2024, identifying key turning points and significant changes; seasonal analysis, such as counting theme distribution by quarter, detecting periodic fluctuation patterns, analyzing the time characteristics of policy releases; stage division, such as dividing the policy development stage based on the change characteristics of attention, analyzing the theme evolution characteristics in different stages, and identifying the policy focus transfer path.
[0205] For the analysis of the space dimension, it includes: regional grouping, such as grouping provinces into eastern, central, and western regions, calculating the theme attention distribution in each region, and comparing the policy emphasis differences between regions; inter-provincial comparison, such as calculating the theme attention index of each province, constructing a province-theme attention matrix, and identifying the policy characteristics of typical provinces; spatial clustering, such as clustering provinces based on theme attention, analyzing the geographical distribution characteristics of policy orientation, and studying the regional coordinated development model.
[0206] For the analysis of the theme dimension, it includes: theme importance ranking, such as calculating the overall attention of each theme, establishing the theme importance ranking, and identifying the core policy concerns; theme correlation analysis, such as calculating the attention correlation between themes, constructing a theme correlation network, and analyzing the synergy effect of policy themes; theme evolution analysis, such as tracking the change of theme attention, analyzing the emergence of emerging themes, and predicting the policy attention trend.
[0207] Moreover, for the multi-dimensional cross-analysis, it includes: multi-dimensional analysis, such as constructing a three-dimensional analysis framework of time-space-theme, studying the change of regional differences over time, and analyzing the regional characteristics of theme evolution; policy effect evaluation, such as comparing the attention change with the policy goal, analyzing the regional differences in policy implementation, and evaluating the policy synergy effect; decision support, such as providing policy focus early warning, assisting in regional policy adjustment, and supporting policy optimization decision-making.
[0208] Therefore, through multi-dimensional quantitative analysis, this embodiment comprehensively presents the attention distribution characteristics of science and technology policies, providing data support for policy formulation and evaluation. The model can be optimized and adjusted according to actual needs, such as adding a finer-grained time division or adding a new space analysis dimension, etc.
[0209] Optionally, the attention analysis step is achieved by constructing a multi-dimensional evaluation index system. Specifically, it mainly includes the following dimensions: 1) Text quantity dimension: Count the occurrence frequency of a specific theme or keyword in all texts and calculate its proportion; 2) Time dimension: Analyze the distribution changes of keywords or themes in different time periods, and identify the evolution trend of hot topics through time series analysis; 3) Dissemination dimension: Consider factors such as the source of text publication, reposting volume, number of comments, etc., and construct a dissemination influence index; 4) Weight dimension: Conduct weighted calculation according to the authority of the text source, policy level, etc. The final attention score is obtained by normalizing each dimension index and then performing weighted summation. The weight distribution is determined by the Analytic Hierarchy Process (AHP), specifically: text quantity dimension 0.3, time dimension 0.2, dissemination dimension 0.25, weight dimension 0.25.
[0210] Regarding the attention trend characteristics, the trend prediction step can adopt time series analysis methods and combine with deep learning models for prediction. The specific implementation includes: 1) Data preprocessing: Arrange the attention scores in chronological order to form time series data. Normalize the data and divide it into training set and test set according to the ratio of 80%:20%; 2) Feature engineering: Construct time window features, including attention data of the previous n days, first-order difference features, seasonal features, etc. The time window size is set to 14 days; 3) Model construction: Use the LSTM model for prediction. The network structure includes two LSTM layers (the number of units is 64 and 32 respectively) and one fully connected layer. Model training parameter settings: batch_size = 32, epoch = 100, optimizer = Adam, learning_rate = 0.001; 4) Prediction result evaluation: Use the Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE) as evaluation indicators, and continuously update the model through rolling prediction to ensure prediction accuracy. At the same time, to improve the prediction accuracy, the Prophet time series prediction model is also integrated. The prediction results of multiple models are combined through ensemble learning. The final prediction result takes the weighted average of the prediction values of each model, and the weights are determined by the prediction errors on the validation set.
[0211] S5: Analyze the target text data by using the pre-established sentiment analysis model to obtain the sentiment analysis result.
[0212] In this application, the sentiment analysis model adopts a deep neural network architecture, specifically using a combination of bidirectional LSTM (Bi-LSTM) and attention mechanism. The input layer of the model uses 300-dimensional pre-trained word vectors, which are obtained by training a Word2Vec model on a large-scale corpus of science and technology policy texts. The network structure of the model consists of an embedding layer, a Bi-LSTM layer, an attention layer, and a fully connected layer. Among them, the number of hidden units in the Bi-LSTM layer is 256. The attention layer is used to capture important words in the sentence and assign higher weights. The fully connected layer uses the softmax activation function to output the probability distribution of three classifications.
[0213] In the actual application process, the sentiment analysis is performed according to the following steps: First, the input target text is tokenized and part-of-speech tagged, and words that have an important impact on sentiment judgment (such as adjectives, adverbs, verbs, etc.) are retained. Then, a sentiment dictionary is used to perform a preliminary sentiment intensity calculation on the text. The sentiment dictionary includes the HowNet sentiment dictionary and the Chinese sentiment dictionary of Li Jun from Tsinghua University, and has been expanded according to the characteristics of the science and technology policy field. Next, the regulatory effects of language components such as negation words and adverbs of degree on sentiment are considered, and the sentiment intensity is adjusted through rules.
[0214] Finally, the processed text is input into the pre-trained deep learning model to obtain the sentiment category (positive, negative, or neutral) of the text and the corresponding confidence score. To improve the reliability of the analysis, the model will also output the key sentences or words that lead to this sentiment judgment, which is convenient for the interpretation and verification of the analysis results. When the confidence level is lower than the preset threshold (0.6), the system will mark the sample for manual review to ensure the accuracy of the analysis results.
[0215] In the experimental verification, the sentiment analysis model achieved an accuracy rate of 85.7% on the test set, and the F1 value was 0.83, which indicates that the model can better capture the sentiment tendency in the science and technology policy text. Through the analysis of error cases, it is found that the accuracy rate of the model is relatively low when dealing with long sentences containing complex tones and implicit emotions, which is also the direction that needs to be focused on for improvement in the future.
[0216] In some embodiments, the sentiment analysis model is constructed through the following steps, including:
[0217] Perform sentence-level segmentation on the sample text data to obtain multiple text statements;
[0218] Identify the encouraging sentences in the text statements as the first sentiment identification value, identify the restrictive sentences in the text statements as the second sentiment identification value, and identify the ordinary sentences in the text statements as the third sentiment identification value;
[0219] Divide the labeled text statements into a training set and a test set according to a preset ratio;
[0220] Based on the training set, determine the training parameters in a preset LSTM model, where the training parameters include batch size, dropout rate parameter, number of training epochs, sequence length, vocabulary length, and number of output classes;
[0221] Based on the training parameters, use the training set to train the LSTM model;
[0222] Use the test set to test the trained LSTM model, obtain performance evaluation metrics, and combine the performance evaluation metrics to obtain the sentiment analysis model, where the performance evaluation metrics include accuracy, precision, recall, and F1 value.
[0223] In this embodiment, sentence-level segmentation adopts a method combining rules and machine learning. First, punctuation marks (such as full stops, question marks, exclamation marks) are used as the basic segmentation basis, and at the same time, punctuation marks unique to Chinese (such as commas, semicolons) are considered. To handle complex long sentences, a clause segmentation method based on dependency syntactic analysis is also introduced, and the LTP tool is used for syntactic analysis, and appropriate segmentation is performed at syntactic structures such as parallel relationships and adversative relationships. In addition, for the clauses marked with bullet points commonly found in policy texts, they are regarded as independent sentence units.
[0224] Furthermore, the sentiment identification process adopts a multi-level classification system. Encouraging sentences (the first sentiment identification value, denoted as 1) usually contain content such as positive support, policy support, and reward measures; restrictive sentences (the second sentiment identification value, denoted as -1) contain content such as control requirements, restrictive conditions, and punishment measures; ordinary sentences (the third sentiment identification value, denoted as 0) are neutral declarative content. To improve the annotation accuracy, a detailed annotation rule library is established, which contains common sentiment trigger words and sentence pattern templates. At the same time, a double-person cross-annotation method is adopted, and the Cohen's Kappa coefficient (required to be greater than 0.8) is calculated to ensure the consistency of annotation.
[0225] Subsequently, the labeled text sentences are divided into a training set and a test set according to a preset ratio. Specifically, the dataset division adopts a stratified sampling method to ensure that the distribution of various sentiment identifications in the training set and the test set is basically the same. For example, the specific ratio is that the training set accounts for 80% and the test set accounts for 20%. The time dimension of the text is considered during the division to ensure that the test set contains samples from different periods to verify the time generalization ability of the model. In addition, 10% is further divided from the training set as a validation set for model tuning and the implementation of the early stopping strategy.
[0226] Furthermore, based on the training set, determine the training parameters in the preset LSTM model. Among them, the training parameters of the LSTM model are determined by grid search and cross-validation to find the optimal combination: the batch size is set to 32 to balance the model training efficiency and memory usage; the dropout rate parameter is set to 0.5 to prevent overfitting; the number of training epochs is set to 100, combined with an early stopping mechanism, and the training stops when the performance on the validation set has not improved for 5 consecutive rounds; the sequence length is set to 128, determined according to the sample sentence length distribution; the vocabulary size is 20,000, covering the common vocabulary in the field of science and technology policies; the number of output classes is 3, corresponding to three sentiment label values.
[0227] Thus, based on the above training parameters, use the training set to train the LSTM model. Among them, the LSTM model training adopts the batch gradient descent method, uses the Adam optimizer, sets the initial learning rate to 0.001, and uses a learning rate decay strategy. The model structure includes an embedding layer (embedding_dim = 300), two layers of bidirectional LSTM layers (units = 256), and a fully connected layer. Cross-entropy is used as the loss function during training, and the model performance is evaluated on the validation set after each epoch.
[0228] Furthermore, use the test set to test the trained LSTM model to obtain a sentiment analysis model. Among them, the model test evaluation is comprehensively measured by multiple metrics: the accuracy reaches 84.5%, indicating the proportion of the overall correct classification of the model; the precision reaches 82.3% on average for each category, reflecting the accuracy of the model's prediction of a certain category; the recall reaches 81.7% on average for each category, reflecting the completeness of the model's recognition of a certain category of samples; the F1 value reaches 82.0%, which is the harmonic mean of precision and recall. Through the comprehensive evaluation of these metrics, ensure that the model has a stable performance in various sentiment recognitions. At the same time, for the categories with an F1 value lower than 0.75, improve the model performance by increasing the training samples of the corresponding categories or adjusting the category weights.
[0229] Based on the above embodiments, in some embodiments, the training of the LSTM model using the training set based on the training parameters includes:
[0230] Construct a feature enhancement model, where the feature enhancement model includes an encoder and a decoder. The encoder includes a Mamba block for extracting sequence features and a skip connection structure composed of a UNet structure. The skip connection structure is used to indicate the fusion of different scale features. The decoder includes an LSTM layer for indicating the processing of temporal dependencies;
[0231] Map the training set to a sequence of word vectors;
[0232] Construct an input structure with sentence-level features and document-level features, and form an enhanced training sample set from the input structure;
[0233] Based on the enhanced training sample set, perform self-supervised pre-training on the feature enhancement model;
[0234] Introduce labeled data, and fine-tune the pre-trained feature enhancement model based on the enhanced training sample set to generate a feature processing model;
[0235] Use the feature processing model to enhance the features of the sequence of word vectors to obtain an optimized training set;
[0236] Use the optimized training set to train the LSTM model.
[0237] In this embodiment, the encoder part of the feature enhancement model adopts a combination of a Mamba block and a UNet structure. The Mamba block uses a state space model (SSM) to capture long-range dependencies. Its core parameters include: the hidden state dimension is 512, the number of attention heads is 8, and the number of layers is 6. Each Mamba block contains a selective state space layer (S4) and a feed-forward neural network layer (FFN). The UNet structure adopts a symmetric contraction path and expansion path, including 4 downsampling layers and 4 upsampling layers. The number of feature channels starts from 64 and expands to 512 at most. The skip connection structure establishes residual connections between corresponding layers and uses 1×1 convolution for feature fusion. The decoder uses a double-layer bidirectional LSTM, and the number of hidden units in each layer is 256, which is used to integrate multi-scale features and strengthen the extraction of temporal information.
[0238] Furthermore, map the training set to a sequence of word vectors. First, use a pre-trained Word2Vec model (the training corpus contains 1 million science and technology policy texts) to map words to 300-dimensional dense vectors. For out-of-vocabulary words (OOV problems), adopt a character-level encoding method, and combine single-character vectors through a CNN to obtain word vectors. At the same time, in order to capture the special semantics of words in policy texts, positional encoding and semantic role encoding are introduced as supplementary features.
[0239] Subsequently, an input structure with sentence-level features and document-level features is constructed, and an enhanced training sample set is formed from the input structure. Among them, the sentence-level features include a sequence of word vectors, syntactic dependency tree features, part-of-speech sequence features, and local semantic importance weights calculated through an attention mechanism. The document-level features include a topic distribution vector (extracted using the LDA model), document structure features (such as chapter hierarchy, clause location, etc.), and global context features. The features at these two levels are integrated through a feature fusion layer to generate enhanced training samples.
[0240] Furthermore, based on the enhanced training sample set, self-supervised pre-training is performed on the feature enhancement model. During the pre-training process, the main tasks include masked language model (MLM), sentence order prediction (SOP), and document structure reconstruction. In the MLM task, 15% of the input tokens are randomly masked; in the SOP task, the sequential relationship between adjacent sentences is predicted; in the document structure reconstruction task, the disrupted document hierarchy is restored. Adam optimizer is used for pre-training, with a learning rate of 2e-5 and 50 training epochs.
[0241] In this embodiment, manually annotated sentiment classification data (including 10,000 annotated samples) is introduced, and the output layer of the pre-trained model is replaced with a new layer adapted to the sentiment classification task. The fine-tuning process adopts a progressive learning rate strategy to generate a feature processing model. The initial learning rate is set to 5e-5 and decays to 0.1 times the original every 5 epochs. At the same time, contrastive learning loss is introduced to improve the model's discriminative ability for similar samples.
[0242] Thus, the multi-dimensional enhancement of the word vector sequence is performed using the feature processing model, including context semantic enhancement, structural feature enhancement, and domain knowledge enhancement. The feature weights are dynamically adjusted through the attention mechanism to highlight the features that have an important impact on sentiment judgment. At the same time, adversarial training techniques are used to improve the robustness of the model, and adversarial samples are generated by adding perturbations for training. Thereby, the enhanced features are used to train the LSTM model. Exemplarily, a mini-batch training method with a batch size of 64 is adopted, and a learning rate scheduling strategy with warmup is used. To prevent overfitting, in addition to dropout, L2 regularization (with a coefficient of 1e-5) and gradient clipping (with a threshold of 5.0) are also introduced. The performance of the model is monitored using cross-validation during the training process. When the F1 value on the validation set does not improve for 3 consecutive epochs, the early stopping mechanism is triggered. The performance of the final model on the test set is improved by 3.5 percentage points compared to the baseline model, and the F1 value reaches 85.5%.
[0243] S6: Generate the text analysis result of the scientific and technological project based on the target topic, the attention trend feature, and the sentiment analysis result.
[0244] Optionally, the generation process of the text analysis result adopts a multi-dimensional fusion method, which mainly includes the following key steps and technical points:
[0245] First, construct a multi-level framework for the analysis result. This framework includes analysis contents in three dimensions: the theme layer, the trend layer, and the sentiment layer. At the theme level, the system sorts the extracted themes according to their importance and calculates the weight scores of each theme. The weight scores are obtained by weighted calculation based on indicators such as the TF-IDF value of the theme words, the topic coherence, and the topic coverage. For each theme, its core keywords (Top-10) and representative text fragments (Top-3) are extracted, and an overview description of the theme is generated.
[0246] Secondly, at the trend level, the system performs time series analysis and feature extraction on the attention trend characteristics. Specifically, it includes: calculating the basic statistical characteristics of the trend, such as the mean, variance, peak value, etc.; identifying key time nodes, including trend inflection points, peak periods, and trough periods; performing trend decomposition, decomposing the trend into three components: long-term trend, periodic fluctuation, and random fluctuation; calculating indicators such as the growth rate and volatility of the trend based on time series analysis methods.
[0247] Subsequently, this application comprehensively forms trend descriptive text from these features and conducts predictive analysis on future trends. At the sentiment level, this application integrates the sentiment analysis results, specifically including: calculating the overall sentiment tendency distribution to obtain the proportions of positive, negative, and neutral sentiments; identifying the key nodes and triggering factors of sentiment changes; extracting representative sentiment expression sentences; analyzing the sentiment differences under different themes. Finally, this application uses natural language generation (NLG) technology to integrate the analysis results of the above three dimensions into a structured analysis report.
[0248] Optionally, to improve the comprehensibility and practicality of the analysis result, this application also provides an interactive result display interface, allowing users to adjust and deeply explore the analysis dimensions according to their needs. At the same time, a quality evaluation system for the analysis result is established, including evaluation indicators in dimensions such as accuracy, integrity, and readability, and the quality of the analysis result is ensured through manual sampling inspection. The analysis report also supports export in multiple formats such as PDF and Word, facilitating subsequent use and sharing by users.
[0249] In some embodiments, the method further includes:
[0250] Generating a data bar chart based on the target text data;
[0251] Generating a corresponding word cloud diagram based on the word features of the target text data, where the word features include the word frequency data and inverse document frequency data of the words;
[0252] Generate a corresponding line chart of attention over time based on the attention trend characteristics;
[0253] Generate a corresponding pie chart of sentiment based on the sentiment analysis results, where the sentiment analysis results include the proportion of sentiment categories;
[0254] Generate a corresponding distribution map of regional characteristics based on the text analysis results of each region regarding the technology project;
[0255] Visually display the data bar chart, the word cloud, the line chart of attention over time, the pie chart of sentiment, and the distribution map of regional characteristics.
[0256] In this embodiment, calculate the comprehensive weight of words, combining the TF-IDF value and the word importance score. The word frequency statistics adopt an improved sliding window algorithm with a window size set to 5, considering the co-occurrence relationship of words. For professional terms and compound words, a customized dictionary is used for identification and weight enhancement. The visual presentation of the word cloud is implemented using the D3.js library, and the main parameter settings include: the maximum font size is 60pt, the minimum font size is 12pt, the word spacing is 2px, and the layout adopts a spiral arrangement algorithm. The color scheme uses a preset technology blue gradient color palette with 5 levels of color differentiation. To improve readability, the word rotation angle range is set to [-30°, 30°], and it is ensured that important words are preferentially arranged in the central position.
[0257] Use the Echarts library to draw an interactive line chart. The X-axis represents time (supporting the switching of multiple granularities such as year / quarter / month / day), and the Y-axis represents the normalized attention score (range 0 - 100). The line adopts a smooth curve effect (the tension parameter is set to 0.4), and an area gradient fill is added to enhance the visual effect. The chart also includes the following functions: a data zoom component for viewing a specific time period, a marker point component to highlight key time nodes, a trend line component to show the overall development trend, and a tooltip component to display detailed data information. At the same time, an adaptive layout is implemented, which can automatically adjust the chart size according to the size of the display device.
[0258] Step A3: Generate a corresponding pie chart of sentiment based on the sentiment analysis results, where the sentiment analysis results include the proportion of sentiment categories;
[0259] Adopt a nested pie chart structure. The inner ring shows the proportion of basic emotion categories (positive, negative, neutral), and the outer ring details the specific emotion intensity distribution under each category. Implement it using the Highcharts library, and adopt a perceptually friendly color scheme (positive - green series, negative - red series, neutral - gray series). Add an animation effect to the chart, support hover interaction and click - to - expand for fan - shaped areas. Use an intelligent position algorithm for data labels to avoid overlap and automatically adjust the display strategy when space is limited. Also provide a legend description and data filtering function, allowing users to display or hide specific categories as needed.
[0260] Build an interactive map visualization based on a certain map API, adopting a combination of a geographical distribution heatmap and marker points to generate a regional feature distribution map. The heatmap shows the density distribution of the number of texts in each region, and the color from blue to red represents the value from low to high. The size of the marker points is proportional to the text analysis score of the region, and detailed analysis results can be displayed after clicking. The map supports zooming and dragging operations to achieve multi - scale data browsing. At the provincial scale, show regional features through gradient color filling and provide data ranking and comparison functions.
[0261] Furthermore, use the Bootstrap grid system to build an adaptive layout, which automatically adjusts the chart arrangement according to the device screen size. Adopt a 2×3 grid layout on large - screen devices and automatically switch to a vertical arrangement on mobile devices. Implement chart linkage interaction. When filtering or clicking on a certain chart, other charts are updated synchronously to display relevant data. Add a chart export function, supporting export options in multiple formats such as PNG and SVG. At the same time, implement a data update mechanism, maintain a real - time connection with the server through WebSocket, and automatically update the chart display when new data arrives. To improve performance, implement lazy loading and data caching mechanisms for the charts.
[0262] Figure 4 This is a schematic structural diagram of a common text analysis system based on the LDA model and LSTM algorithm in the embodiments of this application. The common text analysis system 400 based on the LDA model and LSTM algorithm includes:
[0263] A data acquisition module 401, used to acquire the target text data of scientific and technological projects;
[0264] A feature extraction module 402, used to extract feature words from the target text data to construct a feature word vector matrix based on the feature words. The target text data includes word frequency data and inverse document frequency data;
[0265] A model construction module 403 is configured to construct a perplexity curve based on the feature word vector matrix, and construct a topic model based on the perplexity curve. The perplexity curve is used to indicate the relationship between the number of topics and the perplexity in the target text data, and the topic model is used to indicate multiple target topics obtained after performing topic clustering on the target text data;
[0266] An attention analysis module 404 is configured to determine the topic attention of the target topic based on the topic model, and determine an attention trend feature based on the topic attention. The topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend feature is used to indicate the change trend of the target topic in the time dimension, space dimension, and topic dimension;
[0267] An emotion analysis module 405 is configured to analyze the target text data by using a pre-established emotion analysis model to obtain an emotion analysis result;
[0268] A text analysis module 406 is configured to generate a text analysis result of the scientific and technological project based on the target topic, the attention trend feature, and the emotion analysis result.
[0269] The system according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module in the system of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed function descriptions of each module of the system, reference can be specifically made to the descriptions in the corresponding method shown in the foregoing text, and details are not described herein again.
Claims
1. A public text analysis method based on LDA model and LSTM algorithm, characterized in that: include: Obtain target text data for scientific and technological projects; Extracting characteristic words from the target text data to construct a characteristic word vector matrix based on the characteristic words, the target text data includes word frequency data and inverse document frequency data; wherein the step of obtaining the word frequency data includes: performing multi-level semantic analysis on the target text data, and recursively subdividing the target text data into information particles of different granularities, the information particles are used to indicate the hierarchical theme of the granularity; determining the particle density of each text segment in the information particle at each granularity, the particle density is used to indicate the similarity between the text segment and the hierarchical theme of the granularity; when it is detected that the particle density is lower than a preset density threshold, marking the corresponding text segment as potential abnormal data; optimizing the information particles according to the potential abnormal data to obtain a text feature sequence; and calculating the word frequency data by combining the text feature sequence, the particle density and the target text data; Based on the feature word vector matrix, a perplexity curve is constructed, and based on the perplexity curve, a topic model is constructed, wherein the perplexity curve is used to indicate the relationship between the number of topics in the target text data and the perplexity, and the topic model is used to indicate that a plurality of target topics are obtained after topic clustering of the target text data; Based on the topic model and the target text data, the document topic attention of each record document in the target text data is determined, and the document topic attention is used to indicate the probability distribution of each target topic in the record document within any time period; based on the document topic attention and the total number of topics within any time period, the topic attention of the target topic within the time period is determined; based on the topic attention, a time dimension analysis is performed to obtain a time distribution feature, and the time distribution feature includes an annual trend analysis result, a seasonal analysis result, and a stage analysis result; based on the topic attention, a spatial dimension analysis is performed to obtain a spatial distribution feature, and the spatial distribution feature includes a regional distribution feature, a provincial distribution feature, and a spatial cluster distribution feature; based on the topic attention, a topic dimension analysis is performed to obtain a topic distribution feature, and the topic distribution feature is used to indicate the importance ranking, topic relevance, and topic evolution trend between each target topic; based on the time distribution feature, the space distribution feature, and the topic distribution feature, a multi-dimensional cross analysis is performed to obtain an attention trend feature, and the topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend feature is used to indicate the change trend of the target topic in the time dimension, the space dimension, and the topic dimension; Analyze the target text data using a pre-established sentiment analysis model to obtain a sentiment analysis result; Based on the target topic, the attention trend characteristics and the sentiment analysis results, a text analysis result of the scientific and technological project is generated.
2. The method according to claim 1, characterized in that The following steps are used to build a sentiment analysis model, including: Perform sentence-level segmentation on the sample text data to obtain multiple text sentences; Identifying the encouraging sentences in the text sentence as a first emotion identification value, identifying the restrictive sentences in the text sentence as a second emotion identification value, and identifying the common sentences in the text sentence as a third emotion identification value; Divide the marked text sentences into training sets and test sets according to a preset ratio; Based on the training set, determining training parameters in a preset LSTM model, wherein the training parameters include batch size, dropout rate parameter, number of training rounds, sequence length, vocabulary length, and number of output categories; Based on the training parameters, training the LSTM model using the training set; The trained LSTM model is tested using the test set to obtain performance evaluation indicators, and the sentiment analysis model is obtained by combining the performance evaluation indicators, wherein the performance evaluation indicators include accuracy, precision, recall rate and F1 value.
3. The method according to claim 2, characterized in that The step of constructing a perplexity curve based on the feature word vector matrix, and constructing a topic model based on the perplexity curve, includes: Using the feature word vector matrix, initializing the model parameters of the preset initial topic model, the model parameters including the number of topics, hyperparameters, and the number of iterations; Based on the model parameters and the feature word vector matrix, the initial topic model is iteratively trained, and the perplexity corresponding to each topic number is recorded to construct a perplexity curve; The change rate of the perplexity curve is detected, and the number of topics corresponding to when the change rate is lower than a preset change rate threshold is obtained, and a topic model is constructed based on the obtained number of topics.
4. The method according to claim 3, characterized in that The extracting the feature words in the target text data to construct a feature word vector matrix according to the feature words includes: Based on the target text data, calculating the word frequency data of each word in the target text data, wherein the word frequency data is used to indicate the frequency of occurrence of the word in the record document; Based on the word frequency data, calculating the inverse document frequency data of the word, wherein the inverse document frequency data is used to indicate the importance of the word; Determining word feature data of the word based on the word frequency data and the inverse document frequency data; Based on the word feature data, feature words are screened.
5. The method according to claim 4, characterized in that The step of training the LSTM model based on the training parameters and using the training set includes: Constructing a feature enhancement model, wherein the feature enhancement model includes an encoder and a decoder, the encoder includes a Mamba block for extracting sequence features and a jump connection structure composed of a UNet structure, the jump connection structure is used to indicate the fusion of different scale features, and the decoder includes an LSTM layer for indicating the processing of temporal dependencies; Mapping the training set into a word vector sequence; Constructing an input structure with sentence-level features and document-level features, and forming an enhanced training sample set from the input structure; Based on the enhanced training sample set, performing self-supervised pre-training on the feature enhancement model; Introducing the labeled data, fine-tuning the pre-trained feature enhancement model based on the enhanced training sample set, and generating a feature processing model; Using the feature processing model to perform feature enhancement on the word vector sequence to obtain an optimized training set; The LSTM model is trained using the optimized training set.
6. The method according to claim 5, characterized in that The step of obtaining target text data of scientific and technological projects includes: Obtaining text data to be processed of the scientific and technological project; Constructing a data distribution optimization network, wherein the data distribution optimization network includes a normalized flow transformation layer sequence, and an attention mechanism is arranged in each normalized flow transformation layer in the normalized flow transformation layer sequence; Mapping the text data to be processed to a preset normal distribution space to generate a preliminary mapping result; Based on a preset kernel density estimation model, calculating the local density of data points in the preliminary mapping result; constructing data weights according to the local density; Density optimization is performed on the preliminary mapping result based on the data weight to generate a density-optimized data distribution; Constructing and executing a contrastive learning task, wherein the contrastive learning task is used to instruct the extraction of structural features in the density-optimized data distribution; constructing a domain discriminator, wherein the domain discriminator is used to indicate the difference in data distribution of different data sources; Configuration is used to indicate the dynamic weight mechanism for adjusting weights; Based on the contrastive learning task, the domain discriminator and the dynamic weight mechanism, the density-optimized data distribution is processed to obtain target text data.
7. The method according to claim 6, characterized in that The method further comprises: Based on the target text data, generate a data histogram; Generate a corresponding word cloud diagram based on the word features of the target text data, wherein the word features include word frequency data and inverse document frequency data of the word; Based on the attention trend characteristics, generate a corresponding attention time series line graph; Based on the sentiment analysis results, a corresponding sentiment pie chart is generated, wherein the sentiment analysis results include sentiment category proportions; Based on the text analysis results of each region on the scientific and technological projects, generate corresponding regional characteristic distribution maps; The data bar chart, the word cloud chart, the attention time series line chart, the sentiment pie chart and the regional feature distribution chart are visualized.
8. A public text analysis system based on LDA model and LSTM algorithm, characterized in that: include: Data acquisition module, used to obtain target text data of scientific and technological projects; A feature extraction module is used to extract feature words in the target text data to construct a feature word vector matrix based on the feature words, wherein the target text data includes word frequency data and inverse document frequency data; wherein the step of obtaining the word frequency data includes: performing multi-level semantic analysis on the target text data, and recursively subdividing the target text data into information particles of different granularities, wherein the information particles are used to indicate the hierarchical theme of the granularity; determining the particle density of each text segment in the information particle at each granularity, wherein the particle density is used to indicate the similarity between the text segment and the hierarchical theme of the granularity; when it is detected that the particle density is lower than a preset density threshold, marking the corresponding text segment as potential abnormal data; optimizing the information particles according to the potential abnormal data to obtain a text feature sequence; and calculating the word frequency data by combining the text feature sequence, the particle density and the target text data; A model building module, used to build a perplexity curve based on the feature word vector matrix, and to build a topic model based on the perplexity curve, wherein the perplexity curve is used to indicate the relationship between the number of topics and perplexity in the target text data, and the topic model is used to indicate that a plurality of target topics are obtained after topic clustering of the target text data; The attention analysis module is used to determine the document topic attention of each record document in the target text data based on the topic model and the target text data, and the document topic attention is used to indicate the probability distribution of each target topic in the record document in any time period; based on the document topic attention and the total number of topics in any time period, determine the topic attention of the target topic in the time period; perform time dimension analysis based on the topic attention to obtain time distribution characteristics, and the time distribution characteristics include annual trend analysis results, seasonal analysis results and stage analysis results; perform space dimension analysis based on the topic attention to obtain space distribution characteristics. The spatial distribution characteristics include regional distribution characteristics, provincial distribution characteristics and spatial cluster distribution characteristics; based on the topic attention, a topic dimension analysis is performed to obtain a topic distribution characteristic, and the topic distribution characteristic is used to indicate the importance ranking, topic relevance and topic evolution trend among the target topics; based on the time distribution characteristics, the spatial distribution characteristics and the topic distribution characteristics, a multi-dimensional cross-analysis is performed to obtain an attention trend characteristic, and the topic attention is used to indicate the distribution of each target topic within a specified time period, and the attention trend characteristic is used to indicate the change trend of the target topic in the time dimension, space dimension and topic dimension; The sentiment analysis module is used to analyze the target text data using a pre-established sentiment analysis model to obtain a sentiment analysis result; The text analysis module is used to generate a text analysis result of the scientific and technological project based on the target topic, the attention trend characteristics and the sentiment analysis result.
Citation Information
Patent Citations
Legal document key information extraction method and system
CN118520881A
LDA topic model identification method
CN119149733A