English literature sentiment analysis method and system based on XGBoost algorithm
Through the English literature emotion analysis system based on the XGBoost algorithm, the ratio of literary types and emotional vocabulary is automatically recognized, the granularity is refined and the emotional intensity is marked, which solves the problem of inaccurate traditional emotion analysis methods and achieves efficient and accurate emotional characteristics analysis.
Patent Information
- Application Number
- CN202510085054.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional sentiment analysis methods rely on manually labeled data sets, resulting in inaccurate analysis results and are susceptible to subjective factors of the annotator.
An English literature sentiment analysis system based on the XGBoost algorithm is adopted. The system automatically recognizes the literature type, calculates the emotional vocabulary ratio, divides the units to be analyzed, labels the emotional intensity, and adjusts the initial emotional characteristic value based on these data.
The accurate analysis of the emotional characteristics of English literary works is achieved, the efficiency and accuracy of the analysis is improved, and the reliability of the analysis results are ensured and close to actual emotional expression.
Smart Images

Figure CN120067328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sentiment analysis, and more specifically, to an English literature sentiment analysis method and system based on the XGBoost algorithm. Background Art
[0002] The XGBoost algorithm is a machine learning algorithm based on gradient boosting, which performs excellently in classification and regression tasks.
[0003] With the continuous progress of information technology, sentiment analysis has become a key research direction in the field of natural language processing. The purpose of sentiment analysis is to identify and extract the sentiment tendency in the text, covering aspects such as positive, negative, and neutral. In English literary works, emotional expression is one of the core elements of the works. Through sentiment analysis, the theme of the works, the characters' personalities, and the author's creative intentions can be explored in depth. However, traditional sentiment analysis methods usually rely on manually annotated datasets. This method is not only time-consuming and laborious but also easily interfered by the subjective factors of the annotators, resulting in inaccurate analysis results.
[0004] Therefore, it is necessary to design an English literature sentiment analysis method and system based on the XGBoost algorithm to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes an English literature sentiment analysis method and system based on the XGBoost algorithm, aiming to solve the problem that the analysis results of traditional sentiment analysis methods in the current technology are inaccurate.
[0006] On the one hand, the present invention proposes an English literature sentiment analysis system based on the XGBoost algorithm, including:
[0007] A determination module, configured to determine the literary type of the English literary work to be analyzed, compare the literary type with the historical literary type database, and determine the initial sentiment feature value of the English literary work to be analyzed according to the comparison result; wherein, the value range of the initial sentiment feature value is (-1, 1);
[0008] A judgment module, configured to collect the number of sentiment words in the English literary work to be analyzed, and calculate the first quantity ratio of the number of negative words to the number of sentiment words and the second quantity ratio of the number of positive words to the number of sentiment words respectively; when the initial sentiment feature value is greater than 1, judge whether to adjust the initial sentiment feature value according to the first quantity ratio; when the initial sentiment feature value is less than or equal to 1, judge whether to adjust the initial sentiment feature value according to the second quantity ratio;
[0009] A processing module, configured to divide the English literary work to be analyzed into several units to be analyzed when it is determined that the initial emotional feature value needs to be adjusted, label the emotional intensity of each unit to be analyzed using an emotional dictionary, and construct an emotional intensity data set; determine an adjustment coefficient for the initial emotional feature value based on the emotional intensity data set, and adjust the initial emotional feature value according to the adjustment coefficient to obtain a final emotional feature value;
[0010] A storage module, configured to store the English literary work to be analyzed and the adjustment coefficient.
[0011] Further, the literary types include novels, poems, dramas, and essays.
[0012] Further, when determining the initial emotional feature value of the English literary work to be analyzed according to the comparison result, it includes:
[0013] Calculate the average value of the emotional feature values corresponding to each literary type in the historical literary type database, and denote it as the emotional average value;
[0014] When the literary type matches the historical literary type database, use the emotional average value as the initial emotional feature value;
[0015] When the literary type does not match the historical literary type database, calculate the similarity between the literary type and the historical literary type database, and select the emotional average value corresponding to the literary type with the maximum similarity as the initial emotional feature value.
[0016] Further, when the initial emotional feature value is greater than 1, when determining whether to adjust the initial emotional feature value according to the first quantity ratio, it includes:
[0017] Compare the first quantity ratio with a first quantity ratio threshold, and determine whether to adjust the initial emotional feature value according to the comparison result;
[0018] When the first quantity ratio is within the first quantity ratio threshold, it is determined not to adjust the initial emotional feature value;
[0019] When the first quantity ratio is outside the first quantity ratio threshold, it is determined to adjust the initial emotional feature value.
[0020] Further, when the initial emotional feature value is less than or equal to 1, when determining whether to adjust the initial emotional feature value according to the second quantity ratio, it includes:
[0021] Compare the second quantity ratio with the second quantity ratio threshold, and determine whether to adjust the initial emotional feature value according to the comparison result;
[0022] When the second quantity ratio is within the second quantity ratio threshold, it is determined that the initial emotional feature value is not adjusted;
[0023] When the second quantity ratio is outside the second quantity ratio threshold, it is determined that the initial emotional feature value is adjusted.
[0024] Further, when using an emotion dictionary to label the emotional intensity of each of the units to be analyzed and constructing an emotional intensity data set, it includes:
[0025] Perform word segmentation on each of the units to be analyzed to obtain a word segmentation result;
[0026] Compare the word segmentation result with the emotion dictionary to determine the emotional intensity value corresponding to each word segmentation result;
[0027] According to the emotional intensity value corresponding to each word segmentation result, calculate the average emotional intensity value of the unit to be analyzed, and use the average emotional intensity value as the emotional intensity value of the unit to be analyzed;
[0028] Collect the emotional intensity values of each of the units to be analyzed to construct the emotional intensity data set.
[0029] Further, when determining the adjustment coefficient of the initial emotional feature value based on the emotional intensity data set, it includes:
[0030] Set an emotional intensity standard value, extract all emotional intensity values greater than the emotional intensity value corresponding to the emotional intensity standard value, and calculate the emotional intensity deviation value corresponding to each emotional intensity value and the emotional intensity standard value;
[0031] Determine all emotional intensity deviation values, and generate a first deviation value set according to the emotional intensity deviation values that are less than or equal to the preset deviation value;
[0032] Generate a second deviation value set according to the emotional intensity deviation values that are greater than the preset deviation value;
[0033] Calculate the first average deviation value of the first deviation value set, and calculate the second average deviation value of the second deviation value set;
[0034] Calculate an emotion deviation index based on the first average deviation value and the second average deviation value;
[0035] Determine the adjustment coefficient of the initial emotional feature value according to the emotion deviation index.
[0036] Further, when calculating the sentiment deviation index based on the first average deviation value and the second average deviation value, it includes:
[0037] The sentiment deviation index is obtained by the following formula:
[0038]
[0039] where Z represents the sentiment deviation index; represents the first average deviation value; represents the second average deviation value; Dmax represents the maximum value among the sentiment intensity deviation values.
[0040] Further, when determining the adjustment coefficient of the initial sentiment feature value according to the sentiment deviation index, it includes:
[0041] A sentiment deviation index interval is preset, where the sentiment deviation index interval includes a first preset sentiment deviation index and a second preset sentiment deviation index;
[0042] An adjustment coefficient interval is preset, where the adjustment coefficient interval includes a first preset adjustment coefficient, a second preset adjustment coefficient, and a third preset adjustment coefficient;
[0043] When the first preset condition is recognized, the first preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial sentiment feature value, and the product value of the first preset adjustment coefficient and the initial sentiment feature value is used as the final sentiment feature value;
[0044] When the second preset condition is recognized, the second preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial sentiment feature value, and the product value of the second preset adjustment coefficient and the initial sentiment feature value is used as the final sentiment feature value;
[0045] When the third preset condition is recognized, the third preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial sentiment feature value, and the product value of the third preset adjustment coefficient and the initial sentiment feature value is used as the final sentiment feature value;
[0046] where the first preset condition includes that the sentiment deviation index is less than the first preset sentiment deviation index, the second preset condition includes that the sentiment deviation index is greater than or equal to the first preset sentiment deviation index and less than the second preset sentiment deviation index, and the third preset condition includes that the sentiment deviation index is greater than or equal to the second preset sentiment deviation index.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: The English literature sentiment analysis system based on the XGBoost algorithm of the present invention realizes the accurate analysis of the sentiment characteristics of English literary works through the mutual cooperation of the determination module, the judgment module, the processing module and the storage module. In the determination module, the system can automatically identify the literary type of the work to be analyzed and compare it with the historical literary type database, so as to initially determine the sentiment characteristic value of the work. This process not only improves the analysis efficiency but also ensures the accuracy of the analysis; the judgment module further collects the number of sentiment words in the work and provides a basis for adjusting the initial sentiment characteristic value by calculating the ratio of the number of negative words to the number of positive words. This fully considers the distribution of sentiment words in the work and makes the analysis result closer to the actual sentiment expression of the work; when it is determined that the initial sentiment characteristic value needs to be adjusted, the processing module will play a key role. It first divides the work into multiple units to be analyzed and annotates the sentiment intensity of each unit using the sentiment dictionary. This process not only refines the analysis granularity but also improves the analysis accuracy. Subsequently, the system determines the adjustment coefficient based on the sentiment intensity data set and adjusts the initial sentiment characteristic value to finally obtain the final sentiment characteristic value of the work; the storage module is responsible for storing the work to be analyzed and the adjustment coefficient for subsequent analysis and comparison. This not only ensures the traceability of the analysis result but also provides valuable data support for subsequent research.
[0048] On the other hand, the present invention also proposes an English literature sentiment analysis method based on the XGBoost algorithm, including the following steps:
[0049] S100: Determine the literary type of the English literary work to be analyzed, compare the literary type with the historical literary type database, and determine the initial sentiment characteristic value of the English literary work to be analyzed according to the comparison result; wherein, the value range of the initial sentiment characteristic value is (-1, 1);
[0050] S200: Collect the number of sentiment words in the English literary work to be analyzed, and calculate the first quantity ratio of the number of negative words to the number of sentiment words and the second quantity ratio of the number of positive words to the number of sentiment words respectively; when the initial sentiment characteristic value is greater than 1, determine whether to adjust the initial sentiment characteristic value according to the first quantity ratio; when the initial sentiment characteristic value is less than or equal to 1, determine whether to adjust the initial sentiment characteristic value according to the second quantity ratio;
[0051] S300: When it is determined that the initial emotional feature value needs to be adjusted, divide the English literary work to be analyzed into several units to be analyzed, annotate the emotional intensity of each unit to be analyzed using an emotion dictionary, and construct an emotional intensity data set; determine the adjustment coefficient of the initial emotional feature value based on the emotional intensity data set, and adjust the initial emotional feature value according to the adjustment coefficient to obtain the final emotional feature value;
[0052] S400: Store the English literary work to be analyzed and the adjustment coefficient.
[0053] It can be understood that the above English literary emotion analysis method and system based on the XGBoost algorithm have the same beneficial effects, which will not be elaborated here. Description of the Drawings
[0054] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0055] Figure 1 is a structural block diagram of an English literary emotion analysis system based on the XGBoost algorithm provided by an embodiment of the present invention;
[0056] Figure 2 is a flowchart of an English literary emotion analysis method based on the XGBoost algorithm provided by an embodiment of the present invention. Detailed Embodiments
[0057] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in conjunction with the embodiments.
[0058] Refer to Figure 1 As shown, in some embodiments of the present application, this embodiment provides an English literary emotion analysis system based on the XGBoost algorithm, including:
[0059] A determination module, configured to determine the literary type of the English literary work to be analyzed, compare the literary type with a historical literary type database, and determine an initial emotional feature value of the English literary work to be analyzed according to the comparison result; wherein, the value range of the initial emotional feature value is (-1, 1);
[0060] A judgment module, configured to collect the number of emotional words in the English literary work to be analyzed, and calculate a first quantity ratio of the number of negative words to the number of emotional words and a second quantity ratio of the number of positive words to the number of emotional words respectively; when the initial emotional feature value is greater than 1, judge whether to adjust the initial emotional feature value according to the first quantity ratio; when the initial emotional feature value is less than or equal to 1, judge whether to adjust the initial emotional feature value according to the second quantity ratio;
[0061] A processing module, configured to, when it is determined to adjust the initial emotional feature value, divide the English literary work to be analyzed into several units to be analyzed, label the emotional intensity of each unit to be analyzed by using an emotional dictionary, and construct an emotional intensity data set; determine an adjustment coefficient of the initial emotional feature value based on the emotional intensity data set, and adjust the initial emotional feature value according to the adjustment coefficient to obtain a final emotional feature value;
[0062] A storage module, configured to store the English literary work to be analyzed and the adjustment coefficient.
[0063] It can be seen that the English literature sentiment analysis system based on the XGBoost algorithm in this embodiment realizes the precise analysis of the sentiment characteristics of English literature works through the mutual cooperation of the determination module, the judgment module, the processing module, and the storage module. In the determination module, the system can automatically identify the literary type of the work to be analyzed and compare it with the historical literary type database, thereby initially determining the sentiment characteristic value of the work. This process not only improves the analysis efficiency but also ensures the accuracy of the analysis; the judgment module further collects the number of sentiment words in the work and provides a basis for adjusting the initial sentiment characteristic value by calculating the ratio of the number of negative words to positive words. This fully considers the distribution of sentiment words in the work, making the analysis result closer to the actual sentiment expression of the work; when it is determined that the initial sentiment characteristic value needs to be adjusted, the processing module will play a key role. It first divides the work into multiple units to be analyzed and annotates the sentiment intensity of each unit using a sentiment dictionary. This process not only refines the analysis granularity but also improves the analysis accuracy. Subsequently, the system determines the adjustment coefficient based on the sentiment intensity dataset and adjusts the initial sentiment characteristic value to finally obtain the final sentiment characteristic value of the work; the storage module is responsible for storing the work to be analyzed and the adjustment coefficient for subsequent analysis and comparison. This not only ensures the traceability of the analysis result but also provides valuable data support for subsequent research.
[0064] It can be understood that the English literature sentiment analysis system based on the XGBoost algorithm of the present invention improves the analysis efficiency and accuracy through the mutual cooperation of each module.
[0065] Specifically, the literary types include novels, poems, dramas, and prose.
[0066] In this embodiment, the literary types also include autobiographical literature, epistolary literature, and critical literature.
[0067] Specifically, when determining the initial sentiment characteristic value of the English literature work to be analyzed according to the comparison result, it includes:
[0068] Calculate the average value of the sentiment characteristic values corresponding to each of the literary types in the historical literary type database and denote it as the sentiment average value;
[0069] When the literary type matches the historical literary type database, use the sentiment average value as the initial sentiment characteristic value;
[0070] When the literary type does not match the historical literary type database, calculate the similarity between the literary type and the historical literary type database, and select the sentiment average value corresponding to the literary type with the maximum similarity as the initial sentiment characteristic value.
[0071] It is understandable that matching means that there is a complete or approximate corresponding item in the historical literary genre database for the literary genre of the literary work to be analyzed. For example, if the work to be analyzed is a "novel", and the "novel" genre already exists in the historical literary genre database, and the corresponding emotional characteristic value (average emotion value) also exists, then this is a match. In this case, directly use the average emotion value of this literary genre as the initial emotional characteristic value of this work.
[0072] It is understandable that non - matching means that there is no directly corresponding genre in the historical literary genre database for the literary genre of the literary work to be analyzed, or the difference between the two is so large that it is impossible to find a directly matching emotional characteristic value. For example, if the work to be analyzed is a new or unique literary genre (such as a certain modern cross - literary genre or hybrid literary work), and there is no such genre in the historical database, then this is non - matching. In this case, it is necessary to calculate the similarity between the literary genre to be analyzed and all genres in the historical literary genre database. The calculation steps of the similarity are as follows: First, obtain the feature vector of the literary genre to be analyzed, which can include elements such as keywords, themes, styles of the literary genre; then, obtain the feature vector of each literary genre in the historical literary genre database; next, use similarity calculation methods such as cosine similarity and Euclidean distance to calculate the similarity between the literary genre to be analyzed and each literary genre in the historical literary genre database; finally, select the average emotion value corresponding to the literary genre with the maximum similarity as the initial emotional characteristic value. This process fully considers the similarities and differences between literary genres, enabling the initial emotional characteristic value of the work to be analyzed to be determined more accurately in the case of non - matching, further improving the accuracy and applicability of the analysis.
[0073] Specifically, when the initial emotional characteristic value is greater than 1, when judging whether to adjust the initial emotional characteristic value according to the first quantity ratio, it includes:
[0074] Compare the first quantity ratio with the first quantity ratio threshold, and judge whether to adjust the initial emotional characteristic value according to the comparison result;
[0075] When the first quantity ratio is within the first quantity ratio threshold, it is determined not to adjust the initial emotional characteristic value;
[0076] When the first quantity ratio is outside the first quantity ratio threshold, it is determined to adjust the initial emotional characteristic value.
[0077] Specifically, when the initial emotional characteristic value is less than or equal to 1, when judging whether to adjust the initial emotional characteristic value according to the second quantity ratio, it includes:
[0078] Compare the second quantity ratio with the second quantity ratio threshold, and determine whether to adjust the initial sentiment feature value according to the comparison result;
[0079] When the second quantity ratio is within the second quantity ratio threshold, it is determined not to adjust the initial sentiment feature value;
[0080] When the second quantity ratio is outside the second quantity ratio threshold, it is determined to adjust the initial sentiment feature value.
[0081] It can be understood that when determining whether to adjust the initial sentiment feature value, the system uses the method of comparing the quantity ratio with the threshold. This method is not only simple and intuitive, but also can effectively avoid the analysis deviation caused by the uneven distribution of the quantity of sentiment words. When the quantity ratio of negative words or positive words is within the preset threshold range, the system believes that the initial sentiment feature value is already relatively accurate, so no adjustment is made; while when the quantity ratio exceeds the threshold range, the system determines that the initial sentiment feature value needs to be further adjusted to ensure the accuracy of the analysis result. This fully considers the distribution of sentiment words in literary works and the differences in sentiment expression between different literary works, making the analysis result closer to the actual sentiment expression of the works.
[0082] Specifically, when using the sentiment dictionary to label the sentiment intensity of each of the units to be analyzed and constructing the sentiment intensity data set, it includes:
[0083] Perform word segmentation processing on each of the units to be analyzed to obtain the word segmentation result;
[0084] Compare the word segmentation result with the sentiment dictionary to determine the sentiment intensity value corresponding to each word segmentation result;
[0085] Calculate the average sentiment intensity value of the unit to be analyzed according to the sentiment intensity value corresponding to each word segmentation result, and use the average sentiment intensity value as the sentiment intensity value of the unit to be analyzed;
[0086] Collect the sentiment intensity values of each of the units to be analyzed to construct the sentiment intensity data set.
[0087] It can be understood that the unit to be analyzed refers to the smallest unit into which the English literary work to be analyzed is subdivided during the sentiment analysis process. These units to be analyzed can be sentences, paragraphs, or smaller text fragments, whose sentiment intensity is annotated through a sentiment dictionary and used to construct a sentiment intensity dataset. For example, if the English literary work to be analyzed is a poem about love that contains a large number of words expressing passionate love, then after word segmentation, these words will be compared with the positive words in the sentiment dictionary and corresponding sentiment intensity values will be obtained. Suppose the poem is divided into 10 units to be analyzed, and the average sentiment intensity values of each unit are 0.8, 0.9, 0.75, 0.85, 0.95, 0.7, 0.82, 0.78, 0.92, and 0.88 respectively. Then, these sentiment intensity values will be aggregated to construct the sentiment intensity dataset of the poem.
[0088] Specifically, when determining the adjustment coefficient of the initial sentiment feature value based on the sentiment intensity dataset, it includes:
[0089] Set a sentiment intensity standard value, extract all sentiment intensity values greater than the sentiment intensity value corresponding to the sentiment intensity standard value, and calculate the sentiment intensity deviation value corresponding to each sentiment intensity value and the sentiment intensity standard value;
[0090] Determine all sentiment intensity deviation values, and generate a first deviation value set based on the sentiment intensity deviation values that are all less than or equal to the preset deviation value;
[0091] Generate a second deviation value set based on the sentiment intensity deviation values that are all greater than the preset deviation value;
[0092] Calculate the first average deviation value of the first deviation value set, and calculate the second average deviation value of the second deviation value set;
[0093] Calculate the sentiment deviation index based on the first average deviation value and the second average deviation value;
[0094] Determine the adjustment coefficient of the initial sentiment feature value according to the sentiment deviation index.
[0095] Specifically, when calculating the sentiment deviation index based on the first average deviation value and the second average deviation value, it includes:
[0096] The sentiment deviation index is obtained through the following formula:
[0097]
[0098] Where Z represents the sentiment deviation index; represents the first average deviation value; represents the second average deviation value; Dmax represents the maximum value among the sentiment intensity deviation values.
[0099] It can be understood that when determining the adjustment coefficient, the system first sets a standard value of emotional intensity. This standard value is a benchmark for distinguishing the strength of emotional intensity in the unit to be analyzed.
[0100] Specifically, when determining the adjustment coefficient of the initial emotional feature value according to the emotional deviation index, it includes:
[0101] Preset an emotional deviation index range, where the emotional deviation index range includes a first preset emotional deviation index and a second preset emotional deviation index;
[0102] Preset an adjustment coefficient range, where the adjustment coefficient range includes a first preset adjustment coefficient, a second preset adjustment coefficient, and a third preset adjustment coefficient;
[0103] When the first preset condition is recognized, the first preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotional feature value, and the product value of the first preset adjustment coefficient and the initial emotional feature value is used as the final emotional feature value;
[0104] When the second preset condition is recognized, the second preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotional feature value, and the product value of the second preset adjustment coefficient and the initial emotional feature value is used as the final emotional feature value;
[0105] When the third preset condition is recognized, the third preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotional feature value, and the product value of the third preset adjustment coefficient and the initial emotional feature value is used as the final emotional feature value;
[0106] Among them, the first preset condition includes that the emotional deviation index is less than the first preset emotional deviation index, the second preset condition includes that the emotional deviation index is greater than or equal to the first preset emotional deviation index and less than the second preset emotional deviation index, and the third preset condition includes that the emotional deviation index is greater than or equal to the second preset emotional deviation index.
[0107] It can be understood that when determining the adjustment coefficient, the system uses the emotional deviation index as the judgment basis. The emotional deviation index is an indicator that comprehensively reflects the degree of difference between the emotional intensity distribution in the work to be analyzed and the standard value. By presetting the emotional deviation index range and the adjustment coefficient range, the system can automatically select an appropriate adjustment coefficient according to the size of the emotional deviation index. This design not only improves the automation of the analysis but also ensures that the selection of the adjustment coefficient is more scientific and reasonable. Specifically, when the emotional deviation index is small, it indicates that the emotional intensity in the work to be analyzed is relatively close to the standard value. At this time, the system selects a smaller adjustment coefficient to avoid having too much impact on the initial emotional characteristic value. When the emotional deviation index is large, it indicates that there is a large difference between the emotional intensity in the work to be analyzed and the standard value. At this time, the system selects a larger adjustment coefficient to make a greater adjustment to the initial emotional characteristic value. This strategy fully considers the differences in emotional expressions of different works, making the analysis results more accurate and reliable.
[0108] Refer to Figure 2 As shown, in some embodiments of the present application, this embodiment provides an English literature emotional analysis method based on the XGBoost algorithm, including the following steps:
[0109] S100: Determine the literary type of the English literature work to be analyzed, compare the literary type with the historical literary type database, and determine the initial emotional characteristic value of the English literature work to be analyzed according to the comparison result; wherein, the value range of the initial emotional characteristic value is (-1, 1);
[0110] S200: Collect the number of emotional words in the English literature work to be analyzed, and calculate the first quantity ratio of the number of negative words to the number of emotional words and the second quantity ratio of the number of positive words to the number of emotional words respectively; when the initial emotional characteristic value is greater than 1, judge whether to adjust the initial emotional characteristic value according to the first quantity ratio; when the initial emotional characteristic value is less than or equal to 1, judge whether to adjust the initial emotional characteristic value according to the second quantity ratio;
[0111] S300: When it is determined to adjust the initial emotional characteristic value, divide the English literature work to be analyzed into several units to be analyzed, label the emotional intensity of each unit to be analyzed using an emotional dictionary, and construct an emotional intensity data set; determine the adjustment coefficient of the initial emotional characteristic value based on the emotional intensity data set, and adjust the initial emotional characteristic value according to the adjustment coefficient to obtain the final emotional characteristic value;
[0112] S400: Store the English literature work to be analyzed and the adjustment coefficient.
[0113] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0114] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows Figure 1 or blocks or a combination of multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in Figure 1 one or more flows Figure 1 or blocks or a combination of multiple blocks.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows Figure 1 or blocks or a combination of multiple blocks.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. An English literature sentiment analysis system based on XGBoost algorithm, characterized in that: include: A determination module is configured to determine the literary genre of the English literary work to be analyzed, compare the literary genre with a historical literary genre database, and determine an initial sentiment feature value of the English literary work to be analyzed according to the comparison result; wherein the value range of the initial sentiment feature value is (-1, 1); A judgment module is configured to collect the number of emotional words in the English literary work to be analyzed, and respectively calculate a first quantity ratio of the number of negative words to the number of emotional words, and a second quantity ratio of the number of positive words to the number of emotional words; when the initial emotional feature value is greater than 1, determine whether to adjust the initial emotional feature value according to the first quantity ratio; when the initial emotional feature value is less than or equal to 1, determine whether to adjust the initial emotional feature value according to the second quantity ratio; The processing module is configured to, when it is determined that the initial emotion feature value is to be adjusted, divide the English literary work to be analyzed into a plurality of units to be analyzed, annotate the emotion intensity of each unit to be analyzed using an emotion dictionary, and construct an emotion intensity data set; determine an adjustment coefficient of the initial emotion feature value based on the emotion intensity data set, and adjust the initial emotion feature value according to the adjustment coefficient to obtain a final emotion feature value; The storage module is configured to store the English literary work to be analyzed and the adjustment coefficient.
2. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 1, characterized in that: The literary genres described include fiction, poetry, drama, and prose.
3. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 2, characterized in that: When determining the initial emotional feature value of the English literary work to be analyzed according to the comparison result, it includes: Calculate the average value of the sentiment feature value corresponding to each of the literature types in the historical literature type database, and record it as the sentiment average value; When the literary genre matches the historical literary genre database, taking the sentiment average as the initial sentiment feature value; When the literary genre does not match the historical literary genre database, the similarity between the literary genre and the historical literary genre database is calculated, and the average value of the emotions corresponding to the literary genre corresponding to the maximum similarity value is selected as the initial emotional feature value.
4. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 1, characterized in that: When the initial emotion feature value is greater than 1, judging whether to adjust the initial emotion feature value according to the first quantity ratio includes: Comparing the first quantity ratio with a first quantity ratio threshold, and determining whether to adjust the initial emotion feature value according to the comparison result; When the first quantity ratio is within the first quantity ratio threshold, determining not to adjust the initial emotion feature value; When the first quantity ratio is outside the first quantity ratio threshold, it is determined to adjust the initial emotion feature value.
5. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 4, characterized in that: When the initial emotion feature value is less than or equal to 1, judging whether to adjust the initial emotion feature value according to the second quantity ratio includes: Comparing the second quantity ratio with a second quantity ratio threshold, and determining whether to adjust the initial emotion feature value according to the comparison result; When the second quantity ratio is within the second quantity ratio threshold, determining not to adjust the initial emotion feature value; When the second quantity ratio is outside the second quantity ratio threshold, it is determined to adjust the initial emotion feature value.
6. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 1, characterized in that: When the sentiment intensity of each unit to be analyzed is annotated using a sentiment dictionary and a sentiment intensity data set is constructed, the following steps are included: Performing word segmentation processing on each of the units to be analyzed to obtain a word segmentation result; Compare the word segmentation results with the sentiment dictionary to determine the sentiment intensity value corresponding to each word segmentation result; Calculate the average sentiment intensity value of the unit to be analyzed according to the sentiment intensity value corresponding to each of the word segmentation results, and use the average sentiment intensity value as the sentiment intensity value of the unit to be analyzed; The emotion intensity values of each of the units to be analyzed are collected to construct the emotion intensity data set.
7. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 6, characterized in that: When determining the adjustment coefficient of the initial emotion feature value based on the emotion intensity data set, it includes: Set a standard value for emotional intensity, extract all emotional intensity values that are greater than the standard value for emotional intensity, and calculate an emotional intensity deviation value corresponding to each of the emotional intensity values and the standard value for emotional intensity; Determine all emotion intensity deviation values, and generate a first deviation value set according to all emotion intensity deviation values that are less than or equal to a preset deviation value; Generate a second deviation value set according to all emotion intensity deviation values greater than a preset deviation value; Calculating a first average deviation value of the first deviation value set, and calculating a second average deviation value of the second deviation value set; Calculating a sentiment deviation index based on the first average deviation value and the second average deviation value; An adjustment coefficient of the initial emotional feature value is determined according to the emotional deviation index.
8. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 7, characterized in that: When calculating the sentiment deviation index based on the first average deviation value and the second average deviation value, the method includes: The sentiment bias index is obtained by the following formula: Among them, Z represents the sentiment bias index; represents the first mean deviation value; represents the second average deviation value; Dmax represents the maximum value among the emotion intensity deviation values.
9. The English literature sentiment analysis system based on the XGBoost algorithm according to claim 7, characterized in that: When determining the adjustment coefficient of the initial emotional feature value according to the emotional deviation index, it includes: Presetting an emotional deviation index interval, wherein the emotional deviation index interval includes a first preset emotional deviation index and a second preset emotional deviation index; Presetting an adjustment coefficient interval, wherein the adjustment coefficient interval includes a first preset adjustment coefficient, a second preset adjustment coefficient, and a third preset adjustment coefficient; When the first preset condition is identified, the first preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotion feature value, and the product value of the first preset adjustment coefficient and the initial emotion feature value is used as the final emotion feature value; When the second preset condition is identified, the second preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotion feature value, and the product value of the second preset adjustment coefficient and the initial emotion feature value is used as the final emotion feature value; When the third preset condition is identified, the third preset adjustment coefficient is selected as the adjustment coefficient corresponding to the initial emotion feature value, and the product value of the third preset adjustment coefficient and the initial emotion feature value is used as the final emotion feature value; Among them, the first preset condition includes that the emotional deviation index is less than the first preset emotional deviation index, the second preset condition includes that the emotional deviation index is greater than or equal to the first preset emotional deviation index and less than the second preset emotional deviation index, and the third preset condition includes that the emotional deviation index is greater than or equal to the second preset emotional deviation index.
10. A method for sentiment analysis of English literature based on XGBoost algorithm, applied to the English literature sentiment analysis system based on XGBoost algorithm as claimed in any one of claims 1 to 9, characterized in that: include: Determine the literary genre of the English literary work to be analyzed, compare the literary genre with a historical literary genre database, and determine the initial sentiment feature value of the English literary work to be analyzed according to the comparison result; wherein the value range of the initial sentiment feature value is (-1, 1); Collect the number of emotional words in the English literary work to be analyzed, and calculate a first quantity ratio of the number of negative words to the number of emotional words, and a second quantity ratio of the number of positive words to the number of emotional words; when the initial emotional feature value is greater than 1, determine whether to adjust the initial emotional feature value according to the first quantity ratio; when the initial emotional feature value is less than or equal to 1, determine whether to adjust the initial emotional feature value according to the second quantity ratio; When it is determined that the initial emotion feature value is to be adjusted, the English literary work to be analyzed is divided into a plurality of units to be analyzed, the emotion intensity of each unit to be analyzed is annotated using an emotion dictionary, and an emotion intensity data set is constructed; an adjustment coefficient of the initial emotion feature value is determined based on the emotion intensity data set, and the initial emotion feature value is adjusted according to the adjustment coefficient to obtain a final emotion feature value; The English literary work to be analyzed and the adjustment coefficient are stored.