Chinese language and literature text automatic grading and recommending method based on big data analysis
The language characteristics of Chinese language and literature texts are extracted and automatically graded through big data analysis technology, and personalized recommendations are made based on user interest preferences, solving the problems of insufficient recommendation relevance and accuracy in the existing technology, and achieving more accurate text grading and recommendation effects.
Patent Information
- Application Number
- CN202510006457.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately capture user interest dynamics and text complex features, resulting in insufficient relevance and accuracy of recommendations, and traditional text grading methods fail to fully analyze the deep language features of text.
Through big data analysis technology, language features are extracted from Chinese language and literary texts, text feature vectors are formed, and the text is automatically classified using the decision tree algorithm. Combining user historical reading data and interest preferences, collaborative filtering and content-based recommendation algorithms are used for personalized recommendations.
Accurate grading and personalized recommendations of texts are achieved, which improves the relevance and user experience of recommendations, and avoids the limitations of traditional methods.
Smart Images

Figure CN120067438A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and in particular to an automatic classification and recommendation method for Chinese language and literature texts based on big data analysis. Background Art
[0002] With the rapid development of information technology, especially the widespread application of big data and artificial intelligence, the way users obtain content on the Internet has undergone profound changes; especially in the field of literature, especially in the texts of Chinese language and literature, users' reading needs and interests have gradually become diversified and personalized; as users' reading behaviors and interests accumulate, how to provide users with accurate personalized recommendations from massive text data has become a difficult problem that needs to be solved urgently in the academic and technical circles; traditional recommendation methods are often limited to simple recommendations based on user historical data, and fail to make full use of the language characteristics and deep semantic analysis of the text itself.
[0003] However, there are still some urgent problems to be solved in the existing text recommendation technology. First, most of the existing recommendation systems adopt the traditional single recommendation algorithm based on content or collaborative filtering, which is difficult to accurately capture the user's interest dynamics and the complex characteristics of the text, resulting in insufficient relevance and accuracy of the recommendation. Secondly, traditional text grading methods often rely only on simple keyword matching or manual annotation, lack of sufficient analysis of the deep language characteristics of the text, and it is difficult to effectively grade the text according to multiple dimensions such as difficulty, theme, and emotion. In addition, user interests change rapidly, and how to obtain and update user interest preferences in real time and dynamically based on big data analysis is still a technical problem. Summary of the invention
[0004] Based on the above objectives, the present invention provides a method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis.
[0005] A method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis, comprising the following steps:
[0006] S1: Acquire Chinese language and literature text data from multiple sources through sharing protocols, imports or interfaces;
[0007] S2: preprocess the text data obtained by S1, including word segmentation, removal of stop words, word frequency statistics and text format standardization;
[0008] S3: Apply big data analysis technology to extract language features from preprocessed text data, including sentence structure, grammatical features, emotional tendencies and rhetoric, and form feature vectors of the text;
[0009] S4: Based on the feature vectors of the texts formed in S3, use the decision tree algorithm to classify the texts, and divide them into different levels according to the difficulty, including elementary, intermediate, and advanced levels;
[0010] S5: According to the user's historical reading data and interest preferences, use the recommendation algorithm to make personalized recommendations for the classified texts.
[0011] Optionally, the specific steps of S1 include:
[0012] S11: By establishing a data sharing agreement with content providers, use the data interface to obtain Chinese language and literature text data from cooperative digital libraries, publishing houses, and academic databases;
[0013] S12: Through the file import module, receive and import Chinese language and literature text files stored locally, supporting multiple file formats, including TXT, PDF, or EPUB;
[0014] S13: Connect to third-party data providers through the API interface, use the standardized data transmission protocol to obtain the latest released Chinese language and literature works data from the partner platform;
[0015] S14: Perform format conversion and encoding unification on the obtained text data to ensure that all text data adopts a unified character encoding standard.
[0016] Optionally, the specific steps of S2 include:
[0017] S21: Use a tokenization tool based on regular expressions to perform tokenization on the Chinese language and literature text data obtained in S1, splitting the continuous text string into individual words or phrases to ensure accurate identification of word boundaries;
[0018] S22: Use a predefined stop word list to automatically filter and remove high-frequency meaningless words in the tokenized text, including common conjunctions, prepositions, and auxiliary words, to reduce text noise;
[0019] S23: Apply the word frequency statistics algorithm to perform frequency analysis on the words in the preprocessed text data, calculate the importance weight of each word in the text, and generate a word frequency matrix;
[0020] S24: Perform text format standardization operations, including unifying the character encoding format, eliminating redundant spaces and special characters in the text, and unifying the representation formats of dates and numbers, to ensure the consistency and standardization of text data.
[0021] Optionally, the specific steps of S23 include:
[0022] S31: First, calculate the term frequency TF of each word in the text. The calculation formula is: Among them, w represents a word;
[0023] S32: Calculate the inverse document frequency IDF, and the calculation formula is: Among them, N is the total number of documents in the text collection, and df(w) is the number of documents containing the word w;
[0024] S33: According to the calculated TF and IDF values, calculate the TF-IDF weight of each word. The calculation formula of the TF-IDF is: TF-IDF(w) = TF(w) × IDF(w), where w represents a word, TF(w) is the word frequency of the word in the text, and IDF(w) is the inverse document frequency of the word;
[0025] S34: According to the calculated TF-IDF values, generate a word frequency matrix for each text. Each row in the word frequency matrix represents a text, each column represents a word, and the value of the matrix is the TF-IDF weight of the word in the text.
[0026] Optionally, the specific steps of S3 are as follows:
[0027] S31: Analyze the grammatical structure of each sentence in the preprocessed text data, and extract the structural features such as the sentence length, the number of clauses, and the depth of the syntactic tree;
[0028] S32: Identify and count the application of grammatical rules in the text data, including verb tenses, voices, subject-verb agreements, and sentence type distribution features;
[0029] S33: Classify the sentiment of the text data to determine the sentiment polarity of the text, including positive, negative, and neutral;
[0030] S34: Detect and count the types of rhetorical devices in the text and their occurrence frequencies, including metaphors, personifications, and parallelisms;
[0031] S35: Standardize all the language features extracted in S31 to S34, and use the Z-score standardization method to convert each feature value into a standard score to eliminate the dimensional differences between different features;
[0032] S36: Based on the standardized language features, combined with the word frequency matrix generated in S23, construct the feature vector of the text.
[0033] Optionally, the specific steps of S36 are as follows:
[0034] S361: Create a feature vector space, where each dimension represents a language feature of the text or the word frequency of a word;
[0035] S362: Perform weighted processing on each language feature in S361. Using the feature weight algorithm, assign a weight w to each language feature f , where w f represents the importance of language feature f in text analysis;
[0036] S363: Multiply each normalized language feature value F f by the corresponding feature weight w f to obtain the weighted language feature value W f , and merge it with the term frequency matrix TF m generated in S23 to form the final text feature vector V.
[0037] Optionally, the specific steps of S4 include:
[0038] S41: Train a decision tree classification model. Using the supervised learning method, utilize the labeled training data set D, where each training sample contains a text feature vector v i and the corresponding difficulty level label L i , where L i ∈{beginner, intermediate, advanced}, and optimize the classification rules of the split nodes and leaf nodes of the decision tree by minimizing the classification error;
[0039] S42: Apply information gain as the feature selection criterion. Calculate the split information amount of each feature at the current node, and select the feature with the largest information gain as the split basis for the current node;
[0040] S43: Recursively apply step S43 until the stop condition is met, including the depth of the tree reaching the preset value, the purity of the leaf nodes meeting the requirements, or the number of node samples being lower than the threshold, to generate a complete decision tree classification model;
[0041] S44: Use the trained decision tree classification model to perform classification prediction on the input text feature vector V, and determine the difficulty level L of the text through the split path of the decision tree.
[0042] Optionally, the specific steps of S5 include:
[0043] S51: Collect and organize the user's historical reading data, including the titles, authors, reading times of the texts read by the user, and the user's ratings or feedback information on the texts, to build the user's reading history profile;
[0044] S52: Analyze the user's interest preferences. Use data mining techniques to perform pattern recognition on the user's reading behavior, extract the features of the user's preferred topics, styles, and difficulty levels, and form the user's interest vector;
[0045] S53: Match the user's interest vector with the text feature vector after grading, and use the collaborative filtering algorithm to identify other users or similar text features similar to the user's interest vector to predict the text that the user may be interested in;
[0046] S54: Apply the content-based recommendation algorithm, combine the language features of the text and the user's interest vector, calculate the similarity score between the text and the user's interest, generate a personalized recommendation list; and provide the user with Chinese language and literature texts that meet their interests and needs according to the personalized recommendation list.
[0047] Optionally, the S53 specifically includes:
[0048] S531: Calculate the similarity between the interest vector U of the target user and the interest vectors U j of all other users;
[0049] S532: According to the similarity calculated in S531, select the k users with the highest similarity to the target user's interest vector U to form a set S of similar users, where k is a preset user quantity threshold;
[0050] S533: Calculate the average score of the texts in the user set S, and screen out the texts with scores exceeding the preset threshold as recommended candidate texts;
[0051] S534: Calculate the similarity between the target user's interest vector U and the text feature vector T i after grading;
[0052] S535: According to the similarity calculated in S534, select the text with the highest similarity from the recommended candidate texts as the text that the predicted user may be interested in.
[0053] Optionally, the S54 specifically includes:
[0054] S541: Normalize the text feature vector V constructed in S36 and the user interest vector U formed in S52 to ensure that they are within the same scale range;
[0055] S542: Calculate the similarity score S i between the normalized text feature vector and the user interest vector, and its calculation formula is: S i =V'·U', where S i represents the similarity score between text i and the user's interest, and V' and U' are the normalized text feature vector and user interest vector respectively;
[0056] S543: According to the calculated similarity score S i, sort all the classified texts in descending order of similarity to generate a preliminary personalized recommendation list R;
[0057] S544: Filter the texts in the recommendation list R to exclude the texts that the user has already read to ensure the novelty of the recommended content.
[0058] Advantages of the present invention:
[0059] In the present invention, by deeply mining the language features of texts and combining the user's historical reading data and interest preferences, the texts can be accurately classified and personalized recommendations can be provided; compared with the prior art, the present invention not only automatically classifies the texts through a decision tree algorithm, making the text classification process more objective and automated, but also effectively avoids the limitations of the traditional methods based on keywords or manual annotation.
[0060] In the present invention, by adopting a combination of collaborative filtering algorithm and content-based recommendation algorithm, the dynamic capture of user interests and accurate recommendations can be realized in a big data environment; by calculating the similarity between the texts and the user interests, not only the relevance of the recommendations is improved, but also the changes in user interests can be adapted in real time to provide more personalized recommendations that meet the user's needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0062] Figure 1 Schematic diagram of the method for automatically classifying and recommending Chinese language and literature texts according to an embodiment of the present invention;
[0063] Figure 2 Schematic diagram of the process for making personalized recommendations according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The present invention will be described in detail below with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0065] It should be noted that in the specification, references to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when describing a specific feature, structure, or characteristic in connection with an embodiment, implementing such feature, structure, or characteristic in connection with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0066] Generally, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, depending at least in part on the context, can alternatively allow for the existence of other factors that may not be explicitly described.
[0067] As Figure 1 - Figure 2 shown, a method for automatically grading and recommending Chinese language and literature texts based on big data analysis includes the following steps:
[0068] S1: Obtain Chinese language and literature text data from multiple sources through a sharing protocol, import, or interface;
[0069] S2: Preprocess the text data obtained in S1, including word segmentation, stop word removal, word frequency statistics, and text format standardization;
[0070] S3: Apply big data analysis techniques to extract language features from the preprocessed text data, including sentence structure, grammar features, sentiment tendency, and rhetorical devices, and form a feature vector of the text;
[0071] S4: Based on the feature vector of the text formed in S3, use a decision tree algorithm to grade the text and divide it into different levels according to difficulty, including elementary, intermediate, and advanced;
[0072] S5: According to the user's historical reading data and interest preferences, use a recommendation algorithm to perform personalized recommendation on the graded text.
[0073] S1 specifically includes:
[0074] S11: Establish a data sharing protocol with content providers and use a data interface to obtain Chinese language and literature text data from cooperative digital libraries, publishing houses, and academic databases to ensure legal acquisition and high quality of the data;
[0075] S12: Receive and import the Chinese language and literature text files stored locally through the file import module, supporting multiple file formats including TXT, PDF, or EPUB, to ensure the comprehensiveness and diversity of text data;
[0076] S13: Connect to third-party data providers through the API interface, using a standardized data transfer protocol (such as RESTful API), to obtain the latest released Chinese language and literature work data from the partner platform, ensuring the real-time update and richness of text data;
[0077] S14: Perform format conversion and encoding unification on the obtained text data, ensuring that all text data adopts a unified character encoding standard (such as UTF-8) for subsequent preprocessing and analysis steps; Through the above steps, Chinese language and literature text data can be systematically and legally obtained from multiple authoritative sources, ensuring the comprehensiveness, real-time nature, and consistency of the data, thus providing a solid data foundation for subsequent text preprocessing, feature extraction, and personalized recommendation.
[0078] S2 specifically includes:
[0079] S21: Use a tokenization tool based on regular expressions to perform tokenization on the Chinese language and literature text data obtained in S1, splitting continuous text strings into individual words or phrases to ensure accurate identification of word boundaries;
[0080] S22: Use a predefined stop word list to automatically filter and remove high-frequency meaningless words in the tokenized text, including common conjunctions, prepositions, and auxiliary words, to reduce text noise and improve the accuracy of subsequent analysis;
[0081] S23: Apply a word frequency statistical algorithm to perform frequency analysis on the words in the preprocessed text data, calculate the importance weight of each word in the text, and generate a word frequency matrix for the construction of feature vectors;
[0082] S24: Perform text format standardization operations, including unifying the character encoding format (such as using UTF-8 encoding), eliminating redundant spaces and special characters in the text, and unifying the representation formats of dates and numbers, to ensure the consistency and standardization of text data for subsequent data processing and analysis.
[0083] S23 specifically includes:
[0084] S31: First, calculate the term frequency TF of each word in the text, and the calculation formula is: where w represents the word; This step is used to measure the frequency of the word in the text;
[0085] S32: Calculate the inverse document frequency IDF, and the calculation formula is: where N is the total number of documents in the text collection, and df(w) is the number of documents containing the word w; this step is used to measure the scarcity of the word in the entire text collection;
[0086] S33: According to the calculated TF and IDF values, calculate the TF-IDF weight of each word. The calculation formula of TF-IDF is: TF-IDF(w) = TF(w) × IDF(w), where w represents the word, TF(w) is the word frequency of the word in the text, and IDF(w) is the inverse document frequency of the word; this step is used to combine the occurrence frequency of the word and its scarcity in the text collection to calculate the weight of each word;
[0087] S34: According to the calculated TF-IDF values, generate a word frequency matrix for each text. Each row in the word frequency matrix represents a text, each column represents a word, and the value of the matrix is the TF-IDF weight of the word in the text; through the above steps, the importance weight of each word in the text can be efficiently calculated, and a word frequency matrix can be generated according to the TF-IDF algorithm; this matrix can not only reflect the importance of the word in a single text, but also reveal the scarcity of the word in the entire text collection, thus providing more accurate data support for subsequent text analysis and feature vector construction.
[0088] S3 specifically includes:
[0089] S31: Analyze the syntactic structure of each sentence in the preprocessed text data, and extract the structural features such as the sentence length, the number of subordinate clauses, and the depth of the syntactic tree;
[0090] S32: Identify and count the application of grammar rules in the text data, including verb tenses, voices, subject-verb agreements, and sentence type distribution features;
[0091] S33: Conduct sentiment classification on the text data to determine the sentiment polarity of the text, including positive, negative, and neutral;
[0092] S34: Detect and count the types of rhetorical devices in the text and their occurrence frequencies, including metaphors, personifications, and parallelisms;
[0093] S35: Standardize all the language features extracted in S31 to S34, and use the Z-score standardization method to convert each feature value into a standard score to eliminate the dimensional difference between different features;
[0094] S36: Based on the standardized language features and combined with the word frequency matrix generated in S23, construct the feature vector of the text to facilitate subsequent analysis and processing by machine learning algorithms; through the above steps, the multi-dimensional language features of Chinese language and literature texts can be comprehensively and systematically extracted, including sentence structure, grammatical features, sentiment tendency, rhetorical devices, etc., and these features are effectively integrated with the word frequency matrix generated in step S23 to form a unified feature vector; this process not only improves the accuracy and meticulousness of text feature extraction, but also ensures the comprehensiveness and comparability of the feature vector, providing a solid data foundation for subsequent text grading and personalized recommendation.
[0095] S36 specifically includes:
[0096] S361: Combine the standardized language features (including the grammatical, sentiment, and rhetorical features extracted in steps S31 to S34) with the word frequency matrix generated in S23; first, create a feature vector space, where each dimension represents a language feature of the text or the word frequency of a word;
[0097] S362: Perform weighted processing on each language feature in S361, using a feature weight algorithm (such as a weighted algorithm based on information gain) to assign a weight w f , where w f represents the importance of language feature f in text analysis; the weight calculation formula is: where, TF f represents the word frequency of language feature f, IDF f represents the inverse document frequency of language feature f, and N represents the total number of features; this formula is used to determine the relative importance of each language feature in the feature vector;
[0098] S363: Multiply each standardized language feature value F f by the corresponding feature weight w f to obtain the weighted language feature value W f , and merge it with the word frequency matrix TF m generated in S23 to form the final text feature vector V, whose expression is: V = [W 1 , W 2 ,..., W f , TF 1 , TF 2 ,..., TF m , where, W f = w f × F f represents the weighted language feature value, and TF mrepresents the m-th word frequency value in the word frequency matrix, where m represents the dimension of the word frequency matrix; this step integrates multi-dimensional language features and word frequency information to generate a comprehensive feature vector; through the above steps, the language features and word frequency information of the text can be efficiently and accurately combined to construct a multi-dimensional text feature vector; such a feature vector can not only comprehensively reflect the language structure, sentiment tendency and rhetorical features of the text, but also effectively integrate the lexical information in the word frequency matrix to ensure the comprehensiveness and comparability of the feature vector.
[0099] S4 specifically includes:
[0100] S41: Train a decision tree classification model using the supervised learning method with the labeled training data set D, where each training sample contains the text feature vector V i and the corresponding difficulty level label L i , where L i ∈{beginner, intermediate, advanced}, and optimize the classification rules of the splitting nodes and leaf nodes of the decision tree by minimizing the classification error;
[0101] S42: Apply information gain as the feature selection criterion, calculate the splitting information amount of each feature at the current node, and select the feature with the largest information gain as the splitting basis for the current node. The calculation formula of information gain is: where IG(T, a) represents the information gain of attribute a for the data set T, Entropy(T) is the entropy of the data set T, Values(a) are all possible values of attribute a, and T v is the data subset when the value of attribute a is v;
[0102] S43: Recursively apply step S43 until the stopping conditions are met, including the depth of the tree reaching the preset value, the purity of the leaf nodes meeting the requirements, or the number of node samples being lower than the threshold, to generate a complete decision tree classification model;
[0103] S44: Use the trained decision tree classification model to perform classification prediction on the input text feature vector V, determine the difficulty level L (beginner, intermediate or advanced) of the text through the splitting path of the decision tree, and assign the prediction result to the text; through the detailed description of the above steps, the method of the present invention can efficiently and accurately classify Chinese language and literature texts using the decision tree algorithm; specifically, using information gain as the feature selection criterion ensures that the decision tree model selects the most discriminative features during the classification process, improving the classification accuracy; recursively splitting and optimizing the decision tree structure enables the model to effectively process text data of different complexities and accurately divide them into three difficulty levels: beginner, intermediate and advanced; this not only improves the automation and intelligence level of text classification, but also provides a precise classification basis for subsequent personalized recommendations, enhancing the overall performance of the system and the reading experience of users.
[0104] S5 specifically includes:
[0105] S51: Collect and organize the user's historical reading data, including the titles, authors, reading times of the texts the user has read, and the user's ratings or feedback information on the texts, and construct the user's reading history profile;
[0106] S52: Analyze the user's interest preferences, use data mining techniques to perform pattern recognition on the user's reading behavior, extract the characteristics of the themes, styles, and difficulty levels preferred by the user, and form the user's interest vector;
[0107] S53: Match the user's interest vector with the classified text feature vectors, and use the collaborative filtering algorithm to identify other users or similar text features similar to the user's interest vector to predict the texts that the user may be interested in;
[0108] S54: Apply the content-based recommendation algorithm, combine the language features of the text and the user's interest vector, calculate the similarity score between the text and the user's interest, and generate a personalized recommendation list; and provide Chinese language and literature texts that meet the user's interests and needs to the user according to the personalized recommendation list; Through the above steps, accurate personalized text recommendations can be achieved based on the user's historical reading data and interest preferences, combined with the collaborative filtering and content-based recommendation algorithms; This not only improves the accuracy of the recommendation system, but also better meets the personalized needs of users, and enhances the user's reading experience and satisfaction.
[0109] S53 specifically includes:
[0110] S531: Calculate the similarity between the interest vector U of the target user and the interest vectors U j of all other users. The formula is: where Sim(U, U j ) represents the similarity between the user's interest vector U and the user's interest vector U j , U represents the interest vector of the target user, and U j represents the interest vector of the jth other user, · represents the dot product of vectors, and ||U|| and ||U j || represent the magnitudes of vectors U and U j respectively;
[0111] S532: According to the similarity calculated in S531, select the k users with the highest similarity to the target user's interest vector U to form a set S of similar users, where k is a preset user quantity threshold;
[0112] S533: Calculate the average rating of the texts in the user set S, and screen out the texts with ratings exceeding the preset threshold as recommended candidate texts;
[0113] S534: Calculate the similarity between the target user interest vector U and the text feature vector T after classification, using the cosine similarity calculation formula: i Among them, T represents the feature vector of the i-th text after classification, and ||T i || is the norm of the vector T i ; i
[0114] S535: According to the similarity calculated in S534, select the text with the highest similarity from the recommended candidate texts as the text that the predicted user may be interested in; through the above steps, the user interest vector and the text feature vector can be accurately matched, and the collaborative filtering algorithm is used to effectively identify other users or similar text features similar to the user's interest, so as to accurately predict the text that the user may be interested in.
[0115] S54 specifically includes:
[0116] S541: Normalize the text feature vector V constructed in S36 and the user interest vector U formed in S52 to ensure that they are within the same scale range to eliminate the influence of dimensional differences on the similarity calculation;
[0117] S542: Calculate the similarity score S i between the normalized text feature vector and the user interest vector, and its calculation formula is: S i = V'·U', where S i represents the similarity score between text i and the user interest, and v' and U' are the normalized text feature vector and user interest vector respectively;
[0118] S543: According to the calculated similarity score S i , sort all the texts after classification, arrange them in descending order of similarity, and generate a preliminary personalized recommendation list R;
[0119] S544: Screen the texts in the recommendation list R to exclude the texts that the user has already read to ensure the novelty of the recommended content; through the above steps, the similarity score between the text and the user interest can be accurately calculated, and a highly personalized recommendation list can be generated; specifically, the cosine similarity algorithm is used to effectively measure the matching degree between the text feature and the user interest, ensuring the relevance and user satisfaction of the recommended content; the normalization process eliminates the influence of dimensional differences and improves the accuracy of the similarity calculation; the filtering mechanism ensures the novelty of the recommended content and the satisfaction of the user's needs by excluding the texts that the user has already read; generally speaking, this method significantly improves the accuracy of personalized recommendation and the user experience, and enhances the intelligent level and practical value of the system.
[0120] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions that are made within the spirit and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis, characterized in that: The following steps are involved: S1: Acquire Chinese language and literature text data from multiple sources through sharing protocols, imports or interfaces; S2: preprocess the text data obtained by S1, including word segmentation, removal of stop words, word frequency statistics and text format standardization; S3: Apply big data analysis technology to extract language features from preprocessed text data, including sentence structure, grammatical features, emotional tendencies and rhetoric, and form feature vectors of the text; S4: Based on the feature vectors of the text formed in S3, the decision tree algorithm is used to grade the text and divide it into different levels according to difficulty, including elementary, intermediate and advanced levels; S5: Based on the user's historical reading data and interest preferences, a recommendation algorithm is used to make personalized recommendations for the graded texts.
2. According to the method of automatic classification and recommendation of Chinese language and literature texts based on big data analysis in claim 1, it is characterized by: The S1 specifically includes: S11: Establish data sharing agreements with content providers and use data interfaces to obtain Chinese language and literature text data from cooperating digital libraries, publishing institutions and academic databases; S12: Receive and import Chinese language and literature text files stored locally through the file import module, supporting multiple file formats, including TXT, PDF or EPUB; S13: Connect with third-party data providers through API interfaces and use standardized data transmission protocols to obtain the latest data on Chinese language and literature works from partner platforms; S14: Convert the format and unify the encoding of the acquired text data to ensure that all text data adopt a unified character encoding standard.
3. According to the method of automatic classification and recommendation of Chinese language and literature texts based on big data analysis in claim 1, it is characterized in that: The S2 specifically includes: S21: using a word segmentation tool based on regular expressions to perform word segmentation processing on the Chinese language and literature text data obtained in S1, dividing continuous text strings into separate words or phrases to ensure accurate identification of word boundaries; S22: Use the predefined stop word list to automatically filter and remove high-frequency meaningless words in the text after word segmentation, including common conjunctions, prepositions and auxiliary words, to reduce text noise; S23: Apply a word frequency statistical algorithm to perform frequency analysis on the words in the preprocessed text data, calculate the importance weight of each word in the text, and generate a word frequency matrix; S24: Perform text format standardization operations, including unifying character encoding formats, eliminating redundant spaces and special characters in text, and unifying the representation formats of dates and numbers.
4. According to claim 3, a method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis is characterized in that: The S23 specifically includes: S31: First, calculate the word frequency TF of each word in the text. The calculation formula is: Among them, w represents a word; S32: Calculate the inverse document frequency IDF, the calculation formula is: Where N is the total number of documents in the text collection, and df(w) is the number of documents containing word w; S33: Calculate the TF-IDF weight of each word according to the calculated TF and IDF values, and the calculation formula of TF-IDF is: TF-IDF(w)=TF(w)×IDF(w), where w represents a word, TF(w) is the word frequency of the word in the text, and IDF(w) is the inverse document frequency of the word; S34: Generate a word frequency matrix for each text according to the calculated TF-IDF value, where each row in the word frequency matrix represents a text, each column represents a word, and the value of the matrix is the TF-IDF weight of the word in the text.
5. According to the method of automatic classification and recommendation of Chinese language and literature texts based on big data analysis in claim 1, it is characterized in that: The S3 specifically includes: S31: parsing the grammatical structure of each sentence in the preprocessed text data, extracting structural features of sentence length, number of clauses, and depth of syntactic tree; S32: Identify and count the application of grammatical rules in text data, including verb tense, voice, subject-verb agreement and sentence type distribution characteristics; S33: Perform sentiment classification on text data to determine the sentiment polarity of the text, including positive, negative and neutral; S34: Detect and count the types of rhetorical devices in the text and their frequencies, including metaphor, personification and parallelism; S35: All the language features extracted from S31 to S34 are standardized, and the Z-score standardization method is used to convert each feature value into a standard score to eliminate the dimensional differences between different features; S36: Based on the standardized language features and the word frequency matrix generated in S23, a feature vector of the text is constructed.
6. The method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis according to claim 5 is characterized in that: The S36 specifically includes: S361: Create a feature vector space, where each dimension represents a language feature of the text or the frequency of a word; S362: Perform weighted processing on each language feature in S361, using a feature weight algorithm to assign a weight w to each language feature f , where w f represents the importance of language feature f in text analysis; S363: Each standardized language feature value F f And the corresponding feature weight w f Multiply them together to get the weighted language feature value W f , and the word frequency matrix TF generated in S23 m Merge to form the final text feature vector V.
7. The method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis according to claim 1 is characterized in that: The S4 specifically includes: S41: Train the decision tree classification model using supervised learning methods and using a labeled training dataset D, where each training sample contains a text feature vector V i and the corresponding difficulty level label L i , where L i ∈{elementary, intermediate, advanced}, optimize the classification rules of the split nodes and leaf nodes of the decision tree by minimizing the classification error; S42: Apply information gain as a feature selection criterion, calculate the splitting information amount of each feature at the current node, and select the feature with the largest information gain as the basis for splitting the current node; S43: recursively apply step S43 until a stop condition is met, including the depth of the tree reaches a preset value, the purity of the leaf node reaches a requirement, or the number of node samples is lower than a threshold, to generate a complete decision tree classification model; S44: Using the trained decision tree classification model, classify and predict the input text feature vector V, and determine the difficulty level L of the text through the splitting path of the decision tree.
8. The method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis according to claim 1 is characterized in that: The S5 specifically includes: S51: Collect and organize the user's historical reading data, including the title, author, reading time of the text that the user has read, and the user's rating or feedback information on the text, to build the user's reading history archive; S52: Analyze the user's interest preferences, use data mining technology to perform pattern recognition on the user's reading behavior, extract the characteristics of the user's preferred topics, styles, and difficulty levels, and form the user's interest vector; S53: Matching the user's interest vector with the classified text feature vector, using a collaborative filtering algorithm to identify other users or similar text features similar to the user's interest vector, so as to predict the text that the user is potentially interested in; S54: Apply a content-based recommendation algorithm, combine the language features of the text and the user's interest vector, calculate the similarity score between the text and the user's interests, and generate a personalized recommendation list; and according to the personalized recommendation list, provide the user with Chinese language and literature texts that meet their interests and needs.
9. The method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis according to claim 8 is characterized in that: The S53 specifically includes: S531: Calculate the interest vector U of the target user and the interest vectors U of all other users j The similarity between S532: According to the similarity calculated in S531, select k users with the highest similarity to the target user's interest vector U to form a similar user set S, where k is a preset user quantity threshold; S533: Calculate the average rating of the texts in the user set S, and select texts with ratings exceeding a preset threshold as candidate recommendation texts; S534: Calculate the target user interest vector U and the classified text feature vector T i The similarity between S535: According to the similarity calculated in S534, the text with the highest similarity is selected from the recommended candidate texts as the text that is predicted to be of potential interest to the user.
10. The method for automatic classification and recommendation of Chinese language and literature texts based on big data analysis according to claim 9 is characterized in that: The S54 specifically includes: S541: normalizing the text feature vector V constructed in S36 and the user interest vector U formed in S52 to ensure that the two are within the same scale range; S542: Calculate the similarity score S between the normalized text feature vector and the user interest vector i , the calculation formula is: S i =V′·U′, where S i represents the similarity score between text i and user interest, V′ and U′ are the normalized text feature vector and user interest vector respectively; S543: Based on the calculated similarity score S i , sort all the graded texts, arrange them from high to low according to similarity, and generate a preliminary personalized recommendation list R; S544: Filter the texts in the recommendation list R and exclude the texts that the user has read to ensure the novelty of the recommended content.