User tendency prediction system based on social media text emotion

By using a user preference prediction system based on social media text sentiment, the problems of real-time response and accurate retrieval of dynamic data in information retrieval are solved. It realizes dynamic adaptation of multi-dimensional index structure and real-time capture of user preferences, thereby improving the accuracy and reliability of sentiment change trend judgment.

CN121233744AInactive Publication Date: 2025-12-30QUZHOU COLLEGE OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511240393.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing information retrieval technologies struggle to achieve accurate retrieval and real-time response when dealing with information that changes frequently, in multiple dimensions, and at multiple levels. They also lack the ability to dynamically track data changes, leading to decreased relevance of search results and difficulty in capturing changes in group preferences.

Method used

The user preference prediction system based on social media text sentiment extracts semantic vectors through a multi-dimensional credit module of literacy text and maps them to topic codes to establish a multi-dimensional index structure. It uses micro-cluster statistical summary and macro-cluster evolution trend sequence to calculate literacy preference deterioration index, achieving dynamic adaptation and real-time aggregation.

Benefits of technology

It improves the fineness and dynamic adaptability of the index, enabling timely capture of the real-time evolution trend of user sentiment, improving the accuracy and reliability of sentiment change trend judgment, and ensuring more timely and accurate risk identification of changes in the sentiment of campus user groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233744A_ABST
    Figure CN121233744A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information retrieval, in particular to a social media text emotion-based user tendency prediction system, which comprises a literacy text multi-dimensional credit investigation module, a literacy text multi-dimensional credit investigation module, a literacy text multi-dimensional credit investigation module, a literacy text multi-dimensional credit investigation module, a literacy text multi-dimensional credit investigation module, a literacy text multi-dimensional credit investigation module and a literacy text multi-dimensional credit investigation module, and establishing a literacy tendency basic index unit. According to the method, the semantic vector is extracted based on the input social media text and mapped to the topic ontology library to obtain the topic code, a semantic and topic dual index structure is constructed, accurate matching of the retrieval view and the user request in the multi-dimensional feature space is achieved, and the fine granularity and the dynamic adaptive capacity of the index are improved; by means of incremental updating of a micro-cluster statistical summary, a macro-cluster is formed based on center point distance real-time aggregation, the real-time evolution trend of emotional tendency of a user is captured, and the problem that an existing retrieval technology lacks timely tracking of dynamic data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, and in particular to a user tendency prediction system based on social media text sentiment. Background Technology

[0002] The field of information retrieval technology primarily studies how to quickly and accurately find information that meets user needs from large-scale, structured or unstructured datasets. This field encompasses techniques such as text parsing, feature extraction, index building, similarity calculation, search optimization, result ranking, and multimodal fusion.

[0003] Current information retrieval technologies generally rely on single feature dimensions or static index structures, neglecting the dynamic changes in semantics and topical relationships. This makes it difficult to respond promptly and accurately to information that changes frequently, in multiple dimensions, and at multiple levels. Index spaces often use fixed thresholds or static partitioning methods, lacking sensitivity to data changes, resulting in inaccurate search scope and decreased relevance of search results. In clustering, they only focus on the static clustering results of data items, lacking the ability to perceive real-time changes in the internal distribution characteristics of data during the clustering process. This makes it difficult to reflect changes in the micro-relationships between data points, and consequently, to capture the true evolutionary trajectory of gradual deterioration or improvement in group tendencies. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a user tendency prediction system based on social media text sentiment.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a user tendency prediction system based on social media text sentiment includes:

[0006] The literacy text multidimensional credit module extracts semantic vectors based on the input social media post text, and maps the text to the campus literacy theme ontology to obtain theme codes and establish basic index units for literacy orientation.

[0007] The user preference vector index module, based on the literacy preference basic index unit, places the vector and encoded data into a multi-dimensional data structure, divides the index space, and establishes a multi-dimensional index structure. When a retrieval request is received, it locates the region within the multi-dimensional index structure that is similar to the requested vector in terms of semantics, topic, and temporal weight, and obtains the user preference retrieval view.

[0008] The literacy topic incremental clustering module updates the statistical summary of the micro-cluster to which the data item belongs in each view based on the user preference retrieval view, obtains the topic micro-cluster statistical summary, collects all the topic micro-cluster statistical summaries at a preset time point, and aggregates them according to the distance of the center point of each micro-cluster in the literacy topic dimension to establish a literacy topic macro-cluster set.

[0009] The user tendency development prediction module, based on the set of literacy theme macroclusters, arranges them in chronological order to obtain a macrocluster evolution trend sequence. Based on the macrocluster evolution trend sequence, it calculates the literacy tendency deterioration index and determines and generates a campus user tendency early warning signal.

[0010] Preferably, the steps for obtaining the basic index unit of literacy orientation are as follows:

[0011] Based on the input social media post text, the post text content is decomposed into lexical units. Taking the semantic position of the lexical units in the corpus as the benchmark, the semantic position distribution of all lexical units in the text in the high-dimensional space is calculated, and the distribution result is represented in vector form to obtain the semantic vector.

[0012] Based on the semantic vector, the post text is compared with the campus literacy theme ontology one by one. The standard code of the corresponding theme is extracted according to the literacy theme with the shortest distance to the semantic vector of the post text in the comparison results, and the theme code is obtained.

[0013] Based on the semantic vector and the topic encoding, the sentiment polarity of all sentences in the post text is analyzed one by one, the average sentiment polarity score of each sentence is calculated, and the difference between the current time and the post posting timestamp is calculated in combination with the post text's posting timestamp. The difference is used to determine the time decay factor, and the semantic vector, topic encoding and time decay factor are concatenated item by item to form a unified high-dimensional space numerical vector to obtain a composite feature vector.

[0014] Based on the composite feature vector, the unique identification code in the post source identifier is extracted, and the identification code is mapped and matched one by one with the composite feature vector. The identification code is the primary key, and the composite feature vector is the attribute. A structured association relationship is established one by one to obtain the basic index unit of literacy tendency.

[0015] Preferably, the steps for obtaining the multidimensional index structure are as follows:

[0016] Based on the literacy orientation basic index unit, the composite feature vectors and identification codes included in the literacy orientation basic index unit are extracted one by one. The values ​​of each dimension in the composite feature vector are divided in the order of semantic dimension, topic dimension and time sequence dimension. The values ​​of semantic dimension are assigned to semantic feature array, the values ​​of topic dimension are assigned to topic feature array, and the values ​​of time sequence dimension are assigned to time sequence feature array. Then, the semantic feature array, topic feature array and time sequence feature array are mapped to the corresponding unique identification code to generate a multi-dimensional index structure.

[0017] Preferably, the step of obtaining the user's preferred search view is as follows:

[0018] Based on the aforementioned multidimensional index structure, the topic-sensitive adaptive search radius is calculated;

[0019] Based on the topic-sensitive adaptive search radius, the weighted Euclidean distance between each composite feature vector and the request vector in the multidimensional index structure is calculated, and all composite feature vectors with a weighted Euclidean distance not greater than the topic-sensitive adaptive search radius and their corresponding unique identification codes are selected. All the selected composite feature vectors and their corresponding unique identification codes are combined into a continuous storage structure to generate a user-preferred search view.

[0020] Preferably, the steps for obtaining the topic micro-cluster statistical summary are as follows:

[0021] Based on the user preference retrieval view, each data item is extracted one by one from the user preference retrieval view. The composite feature vector corresponding to the data item is called. The total number of data points in the micro-cluster to which each data item belongs, the linear sum of the values ​​of each dimension, and the square sum of the values ​​of each dimension are updated according to the values ​​of the composite feature vector in the semantic dimension, topic dimension, and time sequence dimension, respectively, to generate a topic micro-cluster statistical summary.

[0022] Preferably, the steps for obtaining the set of literacy topic macroclusters are as follows:

[0023] Based on the topic micro-cluster statistical summary, all micro-cluster statistical summaries are extracted uniformly at a preset time point. The linear summation of the topic dimension and the total number of data points in each micro-cluster statistical summary are called to calculate the center point position of each micro-cluster in the topic dimension, and generate a set of center points of the literacy topic dimension of all micro-clusters.

[0024] Based on the set of central points of the literacy theme dimension, the numerical difference between any two micro-cluster central points in the set in the theme dimension is calculated one by one. A threshold for the numerical difference of micro-cluster central points is set. If the numerical difference between two micro-cluster central points is less than or equal to the threshold, the corresponding two micro-clusters are merged into one micro-cluster. This process continues until the numerical difference between the central points exceeds the threshold, thus generating a set of literacy theme macro-clusters.

[0025] Preferably, the steps for obtaining the macrocluster evolution trend sequence are as follows:

[0026] Based on the aforementioned set of literacy-themed macroclusters, each macrocluster in the set is extracted one by one. All composite feature vectors stored within the macrocluster are analyzed, and the value of the mental health dimension is extracted for each composite feature vector within the macrocluster. The corresponding timestamp is recorded for each composite feature vector. At the same time, the values ​​of all composite feature vectors within the macrocluster on the mental health dimension are summed sequentially, and the total number of composite feature vectors within the macrocluster is counted. All timestamps are arranged in ascending order to generate a macrocluster evolution trend sequence.

[0027] Preferably, the step of obtaining the campus user tendency warning signal is as follows:

[0028] Based on the macro-cluster evolution trend sequence, calculate the literacy tendency deterioration index;

[0029] Based on the aforementioned literacy tendency deterioration index, the literacy tendency deterioration index is compared with the literacy tendency deterioration index threshold. If the literacy tendency deterioration index is greater than the literacy tendency deterioration index threshold, it is determined that the macrocluster has a tendency to evolve in a negative direction in the mental health dimension, and at the same time, it is determined that the macrocluster has deterioration characteristics in both data scale and internal opinion consistency, generating a campus user tendency warning signal.

[0030] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0031] In this invention, semantic vectors are extracted from input social media text and mapped to a topic ontology library to obtain topic codes. A dual index structure of semantics and topic is constructed to achieve accurate matching between the search view and user requests in a multi-dimensional feature space, improving the fine-grainedness and dynamic adaptability of the index. Incremental updates of micro-cluster statistical summaries are used to form macro-clusters based on the distance between centroids in real time, capturing the real-time evolution trend of user sentiment tendencies and solving the shortcoming of existing retrieval technologies in lacking timely tracking of dynamic data. Furthermore, by establishing a macro-cluster evolution trend sequence, a comprehensive tendency deterioration index is calculated based on the centroid displacement rate, data scale growth rate, and group opinion consistency. This enables collaborative analysis and unified early warning judgment of multi-dimensional sentiment trends, improving the accuracy and reliability of sentiment change trend judgment and ensuring more timely and accurate risk identification of changes in the sentiment of campus user groups. Attached Figure Description

[0032] Figure 1 This is a system flowchart of the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0034] Please see Figure 1 The present invention provides a technical solution: a user tendency prediction system based on social media text sentiment, comprising:

[0035] The literacy text multidimensional credit module extracts semantic vectors based on the input social media post text, and maps the text to the campus literacy theme ontology to obtain theme codes and establish basic index units for literacy orientation.

[0036] The user preference vector index module, based on the basic index unit of literacy preference, places vectors and encoded data into a multi-dimensional data structure, divides the index space, and establishes a multi-dimensional index structure. When a retrieval request is received, it locates the region within the multi-dimensional index structure that is similar to the requested vector in terms of semantics, topic, and temporal weight, and obtains the user preference retrieval view.

[0037] The literacy topic incremental clustering module updates the statistical summary of the micro-cluster to which the data item belongs in each view based on the user's preference search view, obtains the statistical summary of the topic micro-cluster, collects the statistical summary of all topic micro-clusters at a preset time point, and aggregates them according to the distance of the center point of each micro-cluster in the literacy topic dimension to establish a set of literacy topic macro-clusters.

[0038] The user tendency development prediction module, based on a set of literacy-themed macroclusters arranged in chronological order, obtains a macrocluster evolution trend sequence. Based on the macrocluster evolution trend sequence, it calculates the literacy tendency deterioration index and determines and generates a campus user tendency early warning signal.

[0039] The steps for obtaining the basic index unit of literacy orientation are as follows:

[0040] Based on the input social media post text, the post text content is decomposed into lexical units. Taking the semantic position of the lexical units in the corpus as the benchmark, the semantic position distribution of all lexical units in the text in the high-dimensional space is calculated, and the distribution result is represented in vector form to obtain the semantic vector.

[0041] Based on semantic vectors, the semantic similarity of each post text with the campus literacy theme ontology is compared. The standard code of the corresponding theme is extracted according to the literacy theme with the shortest semantic vector distance to the post text in the comparison results, and the theme code is obtained.

[0042] Based on semantic vectors and topic encodings, analyze the sentiment polarity of all sentences in the post text one by one, calculate the average value of the sentiment polarity scores of each sentence, calculate the difference between the current time and the post timestamp based on the post timestamp of the post text, determine the temporal decay factor with the difference, and splice the semantic vector, topic encoding, and temporal decay factor item by item to form a unified high-dimensional space numerical vector, obtaining a composite feature vector;

[0043] Based on the composite feature vector, extract the unique identification code in the post source identifier, map and match the identification code with the composite feature vector one by one, use the identification code as the primary key and the composite feature vector as the attribute, and establish a structured association relationship item by item to obtain the basic index unit of literacy tendency.

[0044] Specifically, based on the input social media post text, first use a Chinese word segmenter based on the maximum entropy model to preprocess the post text content. This word segmenter uses the word frequency and word formation probability learned from a campus literacy education corpus with a scale of more than 50 million words to split the continuous text string into independent lexical units. During the word segmenting process, synchronously remove the words in the predefined stop word list in the text, such as words without actual semantic contributions like "de", "le", "zai", etc., and filter out all non-text emoticons and website links. Then, call the pre-trained BERT model to convert each lexical unit into its context-related word embedding vector. This BERT model is based on a structure containing 12 Transformer encoder layers, 768 hidden units, and 12 self-attention heads, and is secondarily pre-trained using a campus literacy education corpus built specifically for this task. This corpus汇集了近五年内国内五十所重点高校官方新闻网站、学生论坛、理论课程教材以及相关社交媒体公开群组的文本数据,在计算中,将一条帖子文本的所有词汇单元的向量输入到BERT模型中,并提取模型输出层中对应于输入序列起始符[CLS]的768维输出向量,该向量融合了整个帖子的全局语义信息,作为该帖子的初步语义向量,为了进一步优化向量表示,引入词频-逆文档频率(TF-IDF)算法计算帖子中每个词汇单元的权重,其中词频(TF)为该词在当前帖子中出现的次数除以帖子总词数,逆文档频率(IDF)则通过公式 Calculate, where is the total number of documents in the corpus, is the number of documents containing the word, multiply the BERT vector of each lexical unit by its corresponding TF-IDF weight, and then sum and average all the weighted lexical unit vectors in the post to finally obtain a 768-dimensional semantic vector that can accurately reflect the core semantics of the post.

[0045] It should be noted that there is an unclear part in the original text where "汇集了近五年内国内五十所重点高校官方新闻网站、学生论坛、理论课程教材以及相关社交媒体公开群组的文本数据" is not properly translated. You may need to provide the correct English expression for this part for a more accurate translation.Based on semantic vectors, a pre-built campus literacy theme ontology is invoked for theme matching. This ontology was jointly constructed by ten senior experts in the field of literacy education based on the guidelines for mental health education. Its structure is a tree-like hierarchical structure, containing five primary themes, each with 3 to 7 secondary themes, totaling 28 secondary themes. Each secondary theme is associated with 50 to 100 core keywords and three standard explanatory phrases, and is assigned a unique standard code consisting of pure numbers. Before performing the comparison, the system has pre-traversed all secondary themes in the ontology, processed all core keywords and standard explanatory phrases under each theme, and calculated their respective semantic vectors using the same BERT model as described above. These vectors are then averaged to generate a single 768-dimensional theme center vector representing the core semantics of each secondary theme. During the comparison, the system calculates the cosine similarity between the semantic vector of the post and the theme center vectors of all 28 secondary themes in the ontology, using the following formula: ,in The semantic vector of the post. Given a topic center vector, the system will traverse all topics and find the topic with the highest cosine similarity value to the post's semantic vector. For example, if the cosine similarity calculated between the semantic vector of a post and the center vector of the topic "Emotion Management and Stress Coping" (coded 10502) is 0.92, which is higher than all other topics, the system will determine that the post belongs to this topic and extract its corresponding standard code 10502 to obtain the topic code.

[0046] Based on semantic vectors and topic encoding, the system first invokes a sentiment analysis model based on a bidirectional long short-term memory network (Bi-LSTM) and an attention mechanism to perform sentiment polarity analysis on each sentence in the post text. This model has been trained on a corpus of over 200,000 labeled sentiment-related texts (positive, negative, neutral) related to campus life. Its network structure includes an embedding layer, two Bi-LSTM layers, and an attention layer, effectively capturing sentiment keywords and their contextual dependencies within sentences. The model outputs a continuous sentiment polarity score between -1.0 and +1.0 for each sentence, where -1.0 represents extreme negativity, +1.0 represents extreme positivity, and 0.0 represents complete neutrality. After parsing all sentences in the post, the arithmetic mean of these sentiment polarity scores is calculated as the overall sentiment tendency score of the post, i.e., the psychological health dimension value. Simultaneously, the system reads the post's publication timestamp from its metadata and obtains the current system timestamp, calculating the time difference between the two in days. Based on this difference, the time-series decay factor is calculated using the exponential decay function, and its calculation formula is as follows: ,in This is the decay factor, set to reduce the post's influence to half after 30 days, calculated by... ,get Its value is approximately 0.0231. For example, a post published 3 days ago has a time decay factor of [value missing]. Finally, the four parts—the semantic vector with a dimension of 768, the topic code in numerical form (e.g., 10502), the calculated mental health dimension value, and the time decay factor—are concatenated in a predetermined order to form a unified high-dimensional numerical vector. Specifically, the four values ​​or vectors are joined end to end. If the semantic vector is 768-dimensional, the final vector dimension is 768+1+1+1=771-dimensional, resulting in a composite feature vector.

[0047] Based on composite feature vectors, the system locates and extracts the source identifier associated with the post text from the original social media data records. This identifier is a code that ensures the post is uniquely identified throughout the system. Its specific format depends on the data source. For example, for posts from campus BBS, the unique identification code is composed of the BBS's unique identifier, the forum ID, the topic ID, and the floor ID concatenated, such as 'PKUBBS-CS-10358-25'. For posts from Weibo, a globally unique 64-bit integer ID is directly used. After extracting this unique identification code, the system pairs it with the 771-dimensional composite feature vector generated in the previous stage to create a data structure entity. The entity uses a unique identification code as the primary key and a complete composite feature vector as its associated attribute value. This process is achieved by establishing a key-value mapping, where the key is a unique identification code of string or integer, and the value is an array of composite feature vectors containing 771 floating-point numbers. For each processed social media post, the system repeats this operation, generating such structured association records one by one. Finally, these records are aggregated into a set, and each element in the set, that is, each pair of "identification code-composite feature vector", is defined as a literacy orientation basic index unit. This unit encapsulates all the quantitative features of a single post, providing standardized and atomic data input for subsequent multidimensional index construction and retrieval analysis.

[0048] The steps to obtain the multidimensional index structure are as follows:

[0049] Based on the basic index unit of literacy orientation, the composite feature vectors and identification codes included in the basic index unit of literacy orientation are extracted one by one. The values ​​of each dimension in the composite feature vector are divided in the order of semantic dimension, topic dimension and time sequence dimension. The values ​​of semantic dimension are assigned to semantic feature array, the values ​​of topic dimension are assigned to topic feature array, and the values ​​of time sequence dimension are assigned to time sequence feature array. Then, the semantic feature array, topic feature array and time sequence feature array are mapped to the corresponding unique identification code to generate a multi-dimensional index structure.

[0050] Specifically, based on the literacy-oriented basic index unit, the system reads records one by one from this index unit set. For each record, it extracts two core components: a unique identification code and a composite feature vector. The composite feature vector is a 771-dimensional numerical vector containing semantic, topic, and temporal information, while the unique identification code is the unique identifier of the post, such as 'PKUBBS-CS-10358-25'. After extraction, the system performs precise dimensional segmentation on each 771-dimensional composite feature vector. This segmentation strictly follows a preset dimensional division rule: the first 768 dimensions of the vector (index 0 to 767) are defined as semantic dimensions. These values ​​together constitute the semantic representation of the post and are therefore segmented and stored in a newly created semantic feature array. The 769th dimension of the vector (index 768) represents the topic classification code of the post, for example, 10502. The topic feature array is segmented separately, forming a single-element topic feature array. The last two dimensions of the vector (indexes 769 and 770) represent the psychological health dimension value and the time decay factor of the post, respectively. These two values ​​reflect the timeliness and emotional state of the post, and are therefore segmented and included in the time-series feature array. After the construction of the three feature arrays is completed, the system associates and binds these three arrays (semantic feature array, topic feature array, and time-series feature array) with the unique identification code corresponding to the composite feature vector. In specific implementation, the system adopts a hash table-based mapping structure, using the unique identification code as the key of the hash table, and storing a structure or object containing the above three feature arrays as the value. By repeatedly performing this extraction, segmentation, and mapping process on all basic index units of literacy tendencies, a complete multi-dimensional index structure is finally built in memory.

[0051] The steps to obtain the user's preferred search view are as follows:

[0052] Based on the multidimensional index structure, the topic-sensitive adaptive search radius is calculated using the following formula:

[0053] ;

[0054] in, To provide a topic-sensitive adaptive search radius, The request vector is a composite feature vector consisting of numerical values ​​representing semantic, topic, and temporal dimensions. For the first in the multidimensional index structure Composite feature vectors, The number of nearest neighbors associated with a topic, determined by the topic dimension value in the request vector. For the request vector and the first The weighted Euclidean distance between composite feature vectors For request vector and The variance of the weighted Euclidean distance between the composite eigenvectors. A variance sensitivity parameter used to control the degree of influence of distance variance;

[0055] Based on the topic-sensitive adaptive search radius, the weighted Euclidean distance between each composite feature vector and the request vector in the multidimensional index structure is calculated. All composite feature vectors with a weighted Euclidean distance not greater than the topic-sensitive adaptive search radius and their corresponding unique identification codes are selected. All the selected composite feature vectors and their corresponding unique identification codes are combined into a continuous storage structure to generate a user-preferred search view.

[0056] Specifically, the formula: The advantage of this formula lies in its ability to achieve adaptability and intelligence in the retrieval process through dynamic calculation of the search radius. The first part of the formula, averaging the weighted Euclidean distances of topic-related nearest neighbors, sets a base value for the search radius that is closely related to the requested topic, ensuring the relevance of the retrieval. The second part of the formula, incorporating a multiplication factor with a hyperbolic tangent function, finely adjusts the base radius based on the dispersion of the nearest neighbor data points. This adjustment is beneficial when data points similar to the requested vector topic are widely distributed semantically and temporally, and opinions are dispersed (i.e., the distance variance is low). When the data points are highly clustered and opinions converge (i.e., the variance is large), this factor will increase significantly, thereby expanding the search radius to capture a more comprehensive range of potential topics. Conversely, when data points are highly clustered and opinions converge (i.e., the variance is large), this factor will increase significantly, thereby expanding the search radius to capture a more comprehensive range of potential topics. When the factor is small, it is close to 1, keeping the search radius within a small range.

[0057] The steps for obtaining the request vector are as follows: The request vector is a numerical representation obtained by the system after real-time analysis and quantification of the request content when it receives an external retrieval request. Its generation process is completely consistent with the construction of the basic index unit of literacy orientation. For example, when a retrieval request with the content "I feel so much pressure from the final exams, I feel like I'm going to get depressed, is there any way to relieve it?" is received, the system first segments the text and removes stop words, then uses a pre-trained BERT model to extract its 768-dimensional semantic vector. Next, the semantic vector is compared with the cosine similarity of all topic center vectors in the campus literacy topic ontology. It is found that it is closest to the topic "emotion management and stress coping" coded as 10502, so 10502 is used as its topic code. Subsequently, the psychological health dimension value of the text is calculated to be -0.75 by the sentiment analysis model, indicating strong negative emotions. Since it is an immediate request, the time difference between its publication time and the current time is 0, and the time decay factor is calculated to be 1.0. Finally, these four parts are concatenated to obtain a 771-dimensional request vector. This serves as the starting point for subsequent searches.

[0058] (in the multidimensional index structure) The steps for obtaining a composite feature vector are as follows: this vector is a data point that has been stored in the multidimensional index structure in the previous steps, representing a specific social media post in the database. During the retrieval, the system traverses these structured composite feature vectors from the multidimensional index structure, and each... Each corresponds to a unique post record within the system, and the values ​​for its 771 dimensions (semantics, topic, mental health, and time series) are the result of standardized processing of historical data. For example, the system retrieves a vector from the index. Its topic dimension is encoded as 10502, its mental health dimension value is -0.6, its time decay factor is 0.88, and the remaining 768 dimensions are the semantic representation of the post. This vector will be combined with the request vector. Perform distance calculation.

[0059] The steps to obtain the (number of nearest neighbors associated with a topic) are as follows: the value of this parameter is determined by the request vector. The topic dimension value is dynamically determined to adjust the initial number of neighbors in the search based on the breadth of the topic. The system has a built-in mapping table between topic levels and the number of neighbors. This table is formulated by domain experts based on the characteristics of the literacy topic system. First-level topics, due to their broad coverage, require the examination of more related posts, and therefore have a larger value assigned to them. The value is set lower for level 2 or 3 topics, which are more specific. Values ​​are used to achieve precise focus, and the specific mapping relationship is as follows: primary topic Secondary theme Level 3 theme For example, when requesting vectors When the topic code is 10502 ("Emotion Management and Stress Coping", a secondary topic), the system queries the mapping table and... The value is set to 50.

[0060] The steps to obtain the (weighted Euclidean distance) are as follows: this distance is used to measure the request vector. With any vector in the index To account for the comprehensive differences in multidimensional space, each dimension of the vector needs to be normalized before calculation to eliminate the influence of different dimensions having different units. The min-max normalization method is used to transform the values ​​of each dimension to the [0, 1] interval. The calculation formula is as follows: ,in It is the original value. and These are the minimum and maximum values ​​of this dimension in the entire dataset. After normalization, the formula for calculating the weighted Euclidean distance is: Among them, weight , , The weights are determined based on the importance of each dimension in judging user preferences. Through grid search optimization on a validation set containing 5000 expert-annotated data points, the final weights are: semantic weights. Topic weight Time-series weights .

[0061] The steps to obtain the variance of the weighted Euclidean distance are as follows: this parameter is used to measure the variance of the weighted Euclidean distance. To determine the degree of dispersion in the distance between the nearest neighbor data points and the request vector, the system first needs to consider the request vector. The topic code is used to initially filter out post vectors with the same topic code in the multidimensional index structure. Then, from these topic-related vectors, their correlation with the target topic is calculated. Calculate the unweighted Euclidean distance between them and select the closest one. Each vector is used to associate the topic with its nearest neighbors. Then, the system calculates... With this The weighted Euclidean distance between the nearest neighbor vectors , to obtain a containing A set of distance values, then, calculate this The average of the distance values Finally, the standard deviation is calculated using the formula. This value reflects the degree of consistency in content and timeliness of posts with the same topic as the query request.

[0062] The steps to obtain the (variance sensitivity parameter) are as follows: this parameter is used to adjust the search radius for the distance variance. The sensitivity of F1 score is a hyperparameter that needs to be calibrated according to system performance requirements. Its value is set based on a large number of retrieval experiments on historical datasets. The experimental goal is to achieve a balance between precision and recall, that is, to find a value that maximizes the F1 score. Maximize The specific setting process is as follows: Select a test set containing 10,000 posts that have been labeled as relevant, and set... The candidate value range, for example, from 0.01 to 10.0, with a step size of 0.01, is used for each candidate. The value is used to run the retrieval process across the entire test set, calculate its average F1 score, and finally, the value with the highest average F1 score is obtained. The value is determined as the optimal configuration of the system through the above experimental process. The optimal value is 0.2.

[0063] Calculation process:

[0064] Received request vector Its topic code is 10502 (second-level topic), then The system first finds... Find the 50 nearest neighbor vectors with the same topic and calculate... The weighted Euclidean distances to these 50 vectors yield a distance set.

[0065] First, calculate the average of these 50 distances:

[0066] ;

[0067] The average distance was calculated. .

[0068] Next, calculate the variance of these distances:

[0069] ;

[0070] The variance was calculated. .

[0071] Set the variance sensitivity parameter .

[0072] Substituting the above calculation results into the topic-sensitive adaptive search radius formula:

[0073] ;

[0074] ;

[0075] ;

[0076] ;

[0077] ;

[0078] The results show that, for the current request, the system calculates a topic-sensitive adaptive search radius of 0.2288. This value will serve as the core threshold for the next step of filtering relevant posts. Any post whose weighted Euclidean distance to the request vector is less than or equal to 0.2288 will be included in the user's preferred search view. This radius value is dynamic; if the variance of the nearest neighbor distance is larger, it indicates that the opinions are more dispersed, and the calculated radius will be larger, and vice versa. This achieves intelligent contraction and expansion of the search scope.

[0079] Based on a topic-sensitive adaptive search radius, the system initiates a global search operation. This operation traverses all composite feature vectors stored in the multidimensional index structure, and for each composite feature vector... The system calls a predefined weighted Euclidean distance calculation function and compares it with the current request vector. The comparison process considers the differences across three dimensions: semantics, topic, and temporal sequence. Specifically, the system calculates the square of the difference between two vectors across 771 normalized dimensions. Then, it multiplies the squared difference of each dimension by its corresponding weight (0.6 for semantic dimension, 0.3 for topic dimension, and 0.1 for temporal dimension). The system then sums all the weighted squared differences and takes the square root to obtain the final distance value. After calculating the distance value, the system immediately compares it with the topic-sensitive adaptive search radius of 0.2288 obtained in the previous step. The composite feature vector is considered valid only if the calculated weighted Euclidean distance is less than or equal to 0.2288. Only when the data item and its corresponding unique identifier are determined to be highly relevant to the request will the system retrieve this "composite feature vector-unique identifier" data item from the multidimensional index structure and temporarily store it in a temporary result list. This process will continue until all data items in the multidimensional index structure have been compared. Finally, the system will integrate all the "composite feature vector-unique identifier" data items that meet the conditions collected in the temporary result list and store them in a contiguous array structure in memory. This final array structure containing all the filtered results is the user's preferred search view.

[0080] The steps to obtain the topic micro-cluster statistics summary are as follows:

[0081] Based on the user preference retrieval view, each data item is extracted one by one from the user preference retrieval view. The composite feature vector corresponding to the data item is called. The total number of data points in the micro-cluster to which each data item belongs, the linear sum of the values ​​of each dimension, and the square sum of the values ​​of each dimension are updated according to the values ​​of the composite feature vector in the semantic dimension, topic dimension, and time sequence dimension, respectively, to generate a topic micro-cluster statistical summary.

[0082] Specifically, based on the user-preferred search view, the system uses an incremental clustering algorithm to process the data items in the view in real time. First, the system traverses each data item in the user-preferred search view, i.e., the "composite feature vector-unique identification code" pair. For the first data item, the system creates a new micro-cluster and uses the composite feature vector of this data item as the initial center of this micro-cluster. At the same time, it initializes the statistical summary of this micro-cluster, including the total number of data points as 1, the linear sum of values ​​of each dimension (LS) which is the composite feature vector itself, and the sum of squares of values ​​of each dimension (SS) which is the sum of squares of the values ​​of each dimension of the composite feature vector. For each subsequent data item, the system calculates the weighted Euclidean distance between its composite feature vector and all existing micro-cluster centers. The distance calculation method is the same as in the aforementioned search stage, and the system finds the nearest micro-cluster. Then, the system determines whether the radius of the nearest micro-cluster after absorbing this data item will exceed a preset micro-cluster radius threshold. This threshold is determined by repeatedly performing clustering on historical datasets. The K-Means clustering experiment analyzes the distribution of the average radius within a cluster under different numbers of clusters. The 75th quantile of this distribution is used as the initial value, and then fine-tuned based on expert experience, for example, set to 0.15. If the radius after absorption does not exceed the threshold, the data item is assigned to this micro-cluster, and the statistical summary of the micro-cluster is updated immediately: the total number of data points is incremented by 1, the composite feature vector of the data item is accumulated dimension by dimension to the linear summation (LS) of the micro-cluster, and the squared values ​​of each dimension of the composite feature vector are accumulated dimension by dimension to the square summation (SS). If the radius after absorption exceeds the threshold, or the distance between the data item and all existing micro-clusters is greater than twice the micro-cluster radius threshold, the system considers the data item to represent a new viewpoint or topic, creates a brand new micro-cluster for it, and initializes its statistical summary in the above manner. By performing this process on all data items in the user's preferred search view, the system dynamically maintains a series of micro-clusters and generates and updates the corresponding topic micro-cluster statistical summary for each micro-cluster in real time.

[0083] The steps to obtain the set of macroclusters for literacy topics are as follows:

[0084] Based on the topic micro-cluster statistical summary, all micro-cluster statistical summaries are extracted uniformly at a preset time point. The linear summation of the topic dimension and the total number of data points in each micro-cluster statistical summary are called to calculate the center point position of each micro-cluster in the topic dimension, and generate a set of center points of the literacy topic dimension of all micro-clusters.

[0085] Based on the set of central points of literacy theme dimensions, the numerical difference between any two micro-cluster central points in the set in terms of theme dimension is calculated one by one. A threshold for the numerical difference of micro-cluster central points is set. If the numerical difference between two micro-cluster central points is less than or equal to the threshold, the corresponding two micro-clusters are merged into one micro-cluster. This process continues until the numerical difference between the central points exceeds the threshold, thus generating a set of literacy theme macro-clusters.

[0086] Specifically, based on the topic micro-cluster statistical summary, the system initiates a periodic macro-aggregation task at a preset time, such as 2 AM daily. This time point is chosen based on the analysis of campus social media user activity; during this period, the system load is lowest, the amount of newly generated data is minimal, and it is suitable for computationally intensive offline analysis. At the start of the task, the system locks all current topic micro-cluster statistical summaries, forming a static snapshot. Then, it extracts the statistical summary information of each micro-cluster from this snapshot. For each micro-cluster, the system pays particular attention to its statistical data in the topic dimension, namely the linear summation of the topic dimension (LS_theme) and the total number of data points (N) of that micro-cluster. This is achieved by dividing the linear summation of the topic dimension by the total number of data points. The system can accurately calculate the center point of a micro-cluster on the literacy theme dimension. This center point is a floating-point number representing the average tendency of all posts within the micro-cluster on the literacy theme. For example, if the linear sum of the theme dimensions of a micro-cluster is 52510 and the total number of data points is 500, then its theme center point is... This indicates that the micro-cluster mainly focuses on the vicinity of the theme "emotional management and stress coping" coded as 10502. The system repeats this calculation process for all micro-clusters in the snapshot, collects the center point positions of all calculated theme dimensions, and finally generates a set containing the coordinates of all active topic micro-clusters in the ideological and political theme space at the current moment, that is, the set of center points of the literacy theme dimensions of all micro-clusters.

[0087] Based on the set of centroids for the campus literacy themes, the system executes a distance-based hierarchical clustering algorithm to aggregate these micro-clusters at a macro level. First, the system arbitrarily selects the centroids of two micro-clusters from the set and calculates the absolute value of their numerical difference along the theme dimension. For example, if the centroid of micro-cluster A is 105.02 and the centroid of micro-cluster B is 105.08, the difference between them is 0.06. The system compares this difference with a preset threshold for the numerical difference between micro-cluster centroids. This threshold is set based on an understanding of the theme coding system in the campus literacy theme ontology, aiming to merge different subtle discussion points belonging to the same secondary theme. According to the ontology design, the difference between different tertiary theme codes under the same secondary theme is typically between 0.1 and 1.0. Therefore, the threshold is set to 0.5. This value is determined by analyzing the minimum difference between all adjacent secondary theme codes in the ontology and taking half of it. The purpose is to ensure that microclusters belonging to different secondary themes are not mistakenly merged. If the difference between the calculated centroids of two microclusters, such as 0.06, is less than or equal to the threshold of 0.5, the system determines that the two microclusters are highly correlated in terms of literacy themes and should be merged. The merging operation is as follows: the topic microcluster statistical summaries of the two microclusters are integrated, that is, their total number of data points, linear summation of each dimension, and square summation of each dimension are added to form a new, larger statistical summary of the microcluster. The new centroid position is calculated using the new statistical summary. Then, the original two microclusters are replaced with this new merged microcluster. The system repeats this process iteratively, continuously calculating and comparing the difference between any two microcluster centroids in the set, until no pair of microclusters can be found whose centroid difference is less than or equal to the threshold of 0.5. At this point, the clustering process converges, and the final set of microclusters is the set of literacy theme macroclusters.

[0088] The steps for obtaining the macrocluster evolution trend sequence are as follows:

[0089] Based on the set of literacy-themed macroclusters, each macrocluster in the set is extracted one by one. All composite feature vectors stored in the macrocluster are analyzed. The value of the mental health dimension is extracted for each composite feature vector in the macrocluster. The corresponding timestamp is recorded for each composite feature vector. At the same time, the values ​​of all composite feature vectors in the macrocluster on the mental health dimension are summed in turn. The total number of composite feature vectors in the macrocluster is counted. All timestamps are arranged in ascending order to generate a macrocluster evolution trend sequence.

[0090] Specifically, based on a set of macroclusters for literacy themes, the system performs independent time-series analysis on each macrocluster. First, the system extracts a macrocluster from the set and accesses all the original data items contained within it, i.e., the "composite feature vector-unique identification code" pair. For each data item within the macrocluster, the system parses its composite feature vector, specifically locating and extracting the psychological health dimension value representing emotional tendency. This value is a floating-point number between -1.0 and +1.0. Simultaneously, the system extracts the publication timestamp accurate to the second from the metadata of this data item, recording this psychological health dimension value and its corresponding timestamp as a data pair. After traversing all data items within the macrocluster, the system obtains the macro... The system first generates a list of mental health dimension values ​​for all posts within a cluster, along with their respective posting timestamps. Then, it sums the mental health dimension values ​​in the list to obtain a macro-cluster mental health value sum, and counts the total number of data items in the list, i.e., the total number of composite feature vectors. Finally, the system sorts all extracted timestamps in chronological order from earliest to latest, and integrates the sorted timestamp sequence, the corresponding list of mental health dimension values, the macro-cluster mental health value sum, and the total number of composite feature vectors into a structured data record. This record fully describes all the key information about the evolution of a specific literacy-themed macro-cluster over time, generating a macro-cluster evolution trend sequence for that macro-cluster.

[0091] The steps for obtaining early warning signals of campus user preferences are as follows:

[0092] Based on the macro-cluster evolution trend sequence, the literacy tendency deterioration index is calculated using the following formula:

[0093] ;

[0094] in, As a literacy tendency deterioration index, Start timestamp The average value of the composite eigenvectors within the macro-cluster on the mental health dimension. Start timestamp The sum of the macro-cluster mental health values ​​corresponding to the macro-clusters. Start timestamp The total number of macro-cluster data points corresponding to the macro-cluster. End timestamp The average value of the composite eigenvectors within the macro-cluster on the mental health dimension. End timestamp The sum of the macro-cluster mental health values ​​corresponding to the macro-cluster. End timestamp The total number of macro-cluster data points corresponding to the macro-cluster. Start timestamp The variance of the composite eigenvectors within the corresponding macro-cluster on the mental health dimension, End timestamp The variance of the composite eigenvectors within the corresponding macro-cluster on the mental health dimension, This is the end timestamp. The starting timestamp, To prevent extremely small positive numbers with a denominator of zero;

[0095] Based on the literacy tendency deterioration index, the literacy tendency deterioration index is compared with the literacy tendency deterioration index threshold. If the literacy tendency deterioration index is greater than the literacy tendency deterioration index threshold, it is determined that the macrocluster has a trend of evolving in a negative direction in the mental health dimension. At the same time, it is determined that the macrocluster has deterioration characteristics in both data scale and internal opinion consistency, and a campus user tendency warning signal is generated.

[0096] Specifically, the formula: The advantage of this formula lies in its ability to comprehensively assess the potential risks of a macro-cluster of literacy topics from three key perspectives: the rate of change in sentiment, the growth in topic size, and the internal consensus of opinions. The first term (rate of change in mean sentiment) directly measures the speed at which the sentiment of the topic deteriorates; a higher value indicates that negative emotions spread faster. The second term (growing size factor) captures the growth trend of the number of participants in the topic through a logarithmic function, effectively identifying topics that are rapidly escalating and expanding in influence. The third term (variance change factor) uses an exponential function to measure the convergence of opinions within the macro-cluster; a decrease in variance indicates convergence and extremism of viewpoints. The product of these three terms constitutes a sensitive and comprehensive deterioration index. Compared to a single indicator, this index can identify potentially harmful campus topics earlier and more accurately, providing a basis for decision-making for literacy educators.

[0097] (Start timestamp) and The steps for obtaining the (end timestamp) are as follows: These two timestamps define the time window for analyzing the macro-cluster evolution trend. Their selection is based on a preliminary analysis of the macro-cluster evolution trend sequence. The system first checks the timestamps of all posts within the macro-cluster to determine the earliest and latest publication times, forming a total time span. To capture significant trends and avoid the impact of short-term fluctuations, the length of the analysis window is set to a fixed value, such as 7 days. It is set to midnight of the calendar day containing the earliest timestamp in the sequence, and It is then set to be from The starting timestamp is 23:59:59 at the end of the 7th day of the calculation. For example, if the earliest post of a macrocluster was published on October 1, 2023 at 14:00, then the starting timestamp is... The Unix timestamp set to October 1, 2023, 00:00:00, end timestamp. This is the Unix timestamp for October 8, 2023, at 23:59:59.

[0098] , and The steps to obtain these parameters are as follows: these parameters describe the starting timestamp. The system measures the state of macroclusters over time, but since macroclusters accumulate over time, it's impossible to directly obtain the state at a precise point in time. Therefore, the system uses a time slice approach to approximate this state. The system filters out all publication timestamps from the macrocluster evolution trend sequence that fall within the start date of the time window (i.e., from...). arrive Posts within the last 24 hours are considered representative samples of the initial state. The system then calculates the sum of the mental health dimension values ​​for all posts within this sample set to obtain the starting timestamp. The sum of the macro-cluster mental health values ​​corresponding to the macro-cluster Count the number of posts in this sample set to obtain the start timestamp. Total number of macrocluster data points corresponding to the macrocluster Finally, based on the mental health dimension values ​​of all posts in this sample set, the sample variance is calculated to obtain the start timestamp. The variance of the composite eigenvector within the corresponding macro-cluster on the mental health dimension For example, if 120 posts are selected within the initial day and their total mental health score is -36, then... , By calculating the variance of these 120 values, we can obtain... .

[0099] , and The steps to obtain these parameters are as follows: these parameters describe the end timestamp. The state of the macrocluster at any given time is obtained in a similar way to the initial state parameters. The system filters out all publication timestamps that fall within the end of the time window (i.e., from the macrocluster evolution trend sequence) from the macrocluster evolution trend sequence. Arriving in the first 24 hours The system selects posts from this sample set as representative samples of the end state, then calculates the sum of the mental health dimension values ​​for all posts within this sample set to obtain the end timestamp. The sum of the macro-cluster mental health values ​​corresponding to the macro-cluster Count the number of posts in this sample set to obtain the end timestamp. Total number of macrocluster data points corresponding to the macrocluster Finally, based on the mental health dimension values ​​of all posts in this sample set, the sample variance is calculated to obtain the final timestamp. The variance of the composite eigenvector within the corresponding macro-cluster on the mental health dimension For example, if 300 posts are selected by the end of the day and their total mental health score is -180, then... , By calculating the variance of these 300 values, we can obtain... .

[0100] and The steps to obtain these parameters are as follows: The two parameters are the average mental health values ​​of posts within the macro-cluster under the initial and final states, respectively. Their calculation directly depends on the parameters obtained in the aforementioned steps, including the start timestamp. The average value of the composite eigenvector within the corresponding macro-cluster on the mental health dimension Through formula The calculated end timestamp The average value of the composite eigenvector within the corresponding macro-cluster on the mental health dimension Through formula The calculations show that these two average values ​​intuitively reflect the changes in the overall level of macro-cluster sentiment at both ends of the time window.

[0101] (To prevent extremely small positive numbers with a denominator of zero) is set as .

[0102] Calculation process:

[0103] Based on the example obtained from the aforementioned parameters, substitute the specific values ​​into the formula for calculation:

[0104] Start timestamp (Unix timestamp): 1696118400 (2023-10-01 00:00:00);

[0105] End timestamp (Unix timestamp): 1696809599 (2023-10-08 23:59:59);

[0106] Time difference A second is approximately equal to 7 days.

[0107] Initial state parameters: , , ;

[0108] End state parameters: , , ;

[0109] Minimal positive number ;

[0110] First calculate and :

[0111] ;

[0112] ;

[0113] Substitute all parameters into the literacy tendency deterioration index The calculation formula is as follows:

[0114] ;

[0115] ;

[0116] ;

[0117] ;

[0118] ;

[0119] ;

[0120] This result indicates that the literacy tendency deterioration index of this macrocluster is This value itself has no absolute meaning. Its core value lies in comparing it with the indices of other macroclusters or with its own historical indices, as well as with preset thresholds. The positive or negative value and the magnitude of the index reflect the strength of the deterioration trend. A positive value indicates that there is a deterioration trend, the larger the value, the higher the risk, and a negative value indicates that the situation is improving.

[0121] Based on the literacy tendency deterioration index, the system compares it with a pre-set literacy tendency deterioration index threshold. Setting this threshold is a comprehensive decision-making process. First, the system calculates the literacy tendency deterioration index for all occurrences of literacy topic macroclusters on a three-month historical dataset, forming an index distribution. Then, ten experienced literacy education experts are invited to conduct blind reviews and score the actual harm of these historical macroclusters, with scores ranging from 1 to 10, representing no risk to extremely high risk. The system performs regression analysis on the index distribution and expert scores to find the index critical point that best distinguishes expert scores as "high risk" (e.g., 7 points and above). This critical point serves as the initial threshold. For example, the calculated threshold is... Subsequently, this threshold will be periodically reviewed and dynamically adjusted by humans based on the actual accuracy and recall rate of the warnings. If the literacy tendency deteriorates, for example, the calculated threshold will be adjusted accordingly. Greater than the threshold of the literacy tendency deterioration index If the system determines that the macrocluster exhibits a significant negative evolutionary trend in the mental health dimension, that is, students' negative emotions are intensifying. At the same time, since the calculation of this index has incorporated considerations of scale growth and opinion convergence, an index exceeding the threshold also means that the macrocluster is showing deteriorating characteristics in both data scale (number of participants) and internal opinion consistency (convergence of viewpoints). Based on this, the system ultimately generates a campus user tendency warning signal.

Claims

1. A system for predicting user inclination based on sentiment of social media text, characterized in that, The system comprises: The literacy text multi-dimensional credit module extracts semantic vectors based on the input social media post text, and simultaneously maps the text to a campus literacy theme ontology library to obtain theme coding and establish a literacy tendency basic index unit; The user tendency vector index module places vectors and coding data into a multi-dimensional data structure based on the literacy tendency basic index unit, divides the index space, establishes a multi-dimensional index structure, locates areas similar to the request vector in semantics, theme and time sequence weight within the multi-dimensional index structure when a search request is received, and obtains a user tendency search view; The literacy topic incremental clustering module updates the statistical summary of the micro-cluster to which each view data item belongs based on the user tendency search view, obtains a topic micro-cluster statistical summary, collects all the topic micro-cluster statistical summaries at a preset time point, aggregates the micro-clusters according to the distance of the center points in the literacy theme dimension, and establishes a literacy theme macro-cluster set; The user tendency development prediction module arranges the macro-cluster evolution trend sequence in chronological order based on the literacy theme macro-cluster set, calculates a literacy tendency deterioration index based on the macro-cluster evolution trend sequence, and judges to generate a campus user tendency early warning signal. 2.The system for predicting user inclination based on sentiment of social media text according to claim 1, wherein, The obtaining step of the literacy tendency basic index unit is: Based on the input social media post text, the post text content is decomposed into lexical units, the semantic position distribution of all lexical units in the text in a high-dimensional space is calculated based on the semantic position of lexical units in the corpus, the distribution result is represented in the form of a vector, and a semantic vector is obtained; Based on the semantic vector, the post text is compared with the campus literacy theme ontology library for semantic similarity, and the standard code of the corresponding theme is extracted according to the literacy theme with the shortest semantic vector distance in the comparison result, and theme coding is obtained; Based on the semantic vector and the theme coding, the sentiment polarity of all sentences in the post text is analyzed one by one, the average value of the sentiment polarity score of each sentence is calculated, the difference between the current time and the post publication timestamp is calculated combined with the post publication timestamp, the time sequence attenuation factor is determined by the difference, the semantic vector, the theme coding and the time sequence attenuation factor are spliced one by one to form a unified high-dimensional space numerical vector, and a composite feature vector is obtained; Based on the composite feature vector, the unique identification code in the post source identifier is extracted, the identification code and the composite feature vector are matched one by one, the identification code is the primary key, the composite feature vector is the attribute, the structured association relationship is established one by one, and the literacy tendency basic index unit is obtained. 3.The user tendency prediction system based on social media text sentiment according to claim 1, wherein, The obtaining step of the multi-dimensional index structure is: Based on the literacy tendency basis index unit, the composite feature vector included in the literacy tendency basis index unit is extracted with the identification code piece by piece, the values of each dimension in the composite feature vector are sequentially segmented according to the semantic dimension, the theme dimension and the time sequence dimension, the semantic dimension values are classified into a semantic feature array, the theme dimension values are classified into a theme feature array, the time sequence dimension values are classified into a time sequence feature array, and then the semantic feature array, the theme feature array and the time sequence feature array are respectively mapped with the corresponding unique identification code to generate a multi-dimensional index structure. 4.The system for predicting user inclination based on sentiment of social media text according to claim 1, wherein, The acquisition step of the user tendency retrieval view is: Based on the multi-dimensional index structure, a theme-sensitive adaptive search radius is calculated; Based on the theme-sensitive adaptive search radius, the weighted Euclidean distance of each composite feature vector in the multi-dimensional index structure and the request vector is calculated, and the composite feature vectors and the corresponding unique identification codes with the weighted Euclidean distance not greater than the theme-sensitive adaptive search radius are screened out, and the screened composite feature vectors and the corresponding unique identification codes are combined into a continuous storage structure to generate a user tendency retrieval view. 5.The user tendency prediction system based on social media text sentiment according to claim 1, wherein, The acquisition step of the topic micro-cluster statistical summary is: Based on the user tendency retrieval view, each data item in the user tendency retrieval view is extracted one by one, the corresponding composite feature vector of the data item is called, and the total number of data points, the linear cumulative sum of each dimension value and the square cumulative sum of each dimension value in each micro-cluster to which each data item belongs are updated according to the values of the composite feature vector in the semantic dimension, the theme dimension and the time sequence dimension to generate a topic micro-cluster statistical summary. 6.The user tendency prediction system based on social media text sentiment according to claim 1, wherein, The acquisition step of the literacy theme macro-cluster set is: Based on the topic micro-cluster statistical summary, all micro-cluster statistical summaries are extracted at a preset time point, the linear cumulative sum of the theme dimension and the total number of data points in each micro-cluster statistical summary are called to calculate the center point position of each micro-cluster in the theme dimension to generate a set of literacy theme dimension center points of all micro-clusters; Based on the set of literacy theme dimension center points, the value difference between any two micro-cluster center points in the set in the theme dimension is calculated one by one, a micro-cluster center point value difference threshold is set, if the value difference between two micro-cluster center points is less than or equal to the value difference threshold, the two micro-clusters are merged into one micro-cluster, and the process continues until the value difference between the center points exceeds the value difference threshold, and a literacy theme macro-cluster set is generated. 7.The user tendency prediction system based on social media text sentiment according to claim 1, wherein, The acquisition step of the macro-cluster evolution trend sequence is: Based on the literacy theme macro-cluster set, each macro-cluster in the literacy theme macro-cluster set is extracted piece by piece, all composite feature vectors stored in the macro-cluster are parsed, the values of the psychological health dimension of each composite feature vector in the macro-cluster are extracted, and the corresponding time stamp of each composite feature vector is recorded, the values of all composite feature vectors in the psychological health dimension in the macro-cluster are sequentially summed, and the total number of composite feature vectors in the macro-cluster is counted, all time stamps are arranged in ascending order, and a macro-cluster evolution trend sequence is generated. 8.The user tendency prediction system based on social media text sentiment according to claim 1, wherein, The acquisition step of the campus user tendency early warning signal is: Based on the macro-cluster evolution trend sequence, a literacy tendency deterioration index is calculated; Based on the literacy tendency deterioration index, the literacy tendency deterioration index is compared with a literacy tendency deterioration index threshold value. If the literacy tendency deterioration index is greater than the literacy tendency deterioration index threshold value, it is determined that the macro cluster has a tendency to evolve in a negative direction in the psychological health dimension, and at the same time, it is determined that the macro cluster has deterioration characteristics in the data size and internal opinion consistency, and a campus user tendency early warning signal is generated.

Citation Information

Cited By

  • Coding and application method and system constructed based on rule file and corpus

    CN121960503A