Word processing system and method based on language literature

By designing a word processing system based on language and literature, using technologies such as word segmentation, word vector generation and real-time data collection, it solves the problem that traditional readers find it difficult to effectively segment and extract the main content of literary works, and achieves the effect of improving reading efficiency and depth of understanding.

CN120197618AInactive Publication Date: 2025-06-24烟台理工学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510253280.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When traditional readers read literary works, it is difficult for them to effectively segment the textual context of the work, and they cannot accurately extract the main content of each paragraph, resulting in a decrease in reading efficiency.

Method used

Design a word processing system based on language and literature, including a word input unit, a preprocessing unit, a data acquisition unit, a data transmission unit, a data storage unit, a data processing unit, an intelligent generation unit and a remote monitoring unit. The system realizes intelligent segmentation and keyword extraction of literary works through word segmentation, word vector generation, stop word filtering, real-time data collection and analysis.

Benefits of technology

The system can help readers quickly grasp the structure and level of the work, improve reading efficiency and depth of understanding, and provide real-time display through remote monitoring units, increasing processing transparency and user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197618A_ABST
    Figure CN120197618A_ABST
Patent Text Reader

Abstract

The invention discloses a word processing system based on language literature, which comprises a word input unit, a preprocessing unit, a data acquisition unit, a data transmission unit, a data storage unit, a data processing unit, an intelligent generation unit and a remote monitoring unit, and belongs to the technical field of word processing. Word vector generation and stop word and low-frequency word filtering are carried out on literature works through the preprocessing unit, and a high-quality data basis is provided for subsequent processing. A data acquisition unit acquires key information of word weights, paragraph lengths, word quantity, part of speech and positions in real time, comprehensive and accurate data support is provided for deep analysis of works, and then an intelligent generation unit can intelligently segment literature works according to an analysis result of a data processing unit, so that the segmentation efficiency of the literature works is improved. And keywords are extracted for language organization to generate a summary of each segment, so that a reader is helped to quickly grasp the structure and the hierarchy of the works, and the reading efficiency and the understanding depth are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of word processing, and more specifically discloses a word processing system and method based on language and literature. Background Art

[0002] With the rapid development of information technology, word processing has been widely applied in various fields. However, most of the existing word processing systems focus on the basic editing and format adjustment of text, and lack effective processing means for text types such as literary works with unique writing contexts and rich connotations. Traditional readers often rely on personal understanding and experience to grasp the structure and levels of literary works when reading them. However, due to the usually complex writing context and rich symbolic meanings of literary works, it is difficult for readers to accurately divide paragraphs according to the author's intention. The structures of many literary works are not simple linear narratives, but show progressive plots or emotional fluctuations through delicate descriptions, dialogues, inner monologues, etc. This makes it often difficult for readers to clearly grasp the core theme or main content of each paragraph during the reading process, thus affecting the understanding and overall grasp of the deep meaning of the work. Such a vague understanding often leads to a decline in reading efficiency. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to solve the problem that traditional readers cannot effectively divide paragraphs according to the writing context of literary works when reading them, cannot accurately extract the main content of each paragraph, and will reduce the reading efficiency of readers for literary works.

[0004] To solve the above technical problems, according to one aspect of the present invention, more specifically, a word processing system based on language and literature includes: a text input unit, a preprocessing unit, a data acquisition unit, a data transmission unit, a data storage unit, a data processing unit, an intelligent generation unit, and a remote monitoring unit; Text Input Unit: Used to receive the text data of the literary work to be processed; Preprocessing Unit: Used to generate word vectors after word segmentation of the input text, and filter out stop words and low-frequency words; Data Acquisition Unit: Used to collect information data generated in real time during the processing of the text of the literary work; Data Transmission Unit: Used to transmit the information data collected by the data acquisition unit in real time; Data Storage Unit: Used to store the information data transmitted in real time by the data transmission unit; Data processing unit: It is used to process and analyze the information data stored in the data storage unit, obtain the processing and analysis results, generate a text processing report according to the processing and analysis results, and then forward the obtained processing and analysis results to the remote monitoring unit and the intelligent generation unit; Intelligent generation unit: It is used to segment a literary work and extract keywords according to the analysis results of the data processing unit, and organize the extracted keywords to achieve a summary of each segment; Remote monitoring unit: It is used to display the information data stored in the data storage unit and the analysis results of the data processing unit in real time on the remote platform terminal.

[0005] Furthermore, the data acquisition unit includes: a weight acquisition module, a word length acquisition module, a word count acquisition module, a part-of-speech acquisition module, and a position acquisition module; Weight acquisition module: It is used to collect the TF-IDF weight information data of words in a literary work in real time; Word length acquisition module: It is used to collect the paragraph length information data in a literary work in real time; Word count acquisition module: It is used to collect the word quantity information data in a literary work in real time; Part-of-speech acquisition module: It is used to collect the part-of-speech information data of words in a literary work in real time; Position acquisition module: It is used to collect the position information data of words in a literary work in real time.

[0006] Furthermore, the data processing unit includes: a data acquisition module, an analysis and judgment module, a result forwarding module, and a report generation module; Data acquisition module: It is used to obtain the information data stored in the data storage unit in real time; Analysis and judgment module: It is used to process and analyze the information data obtained by the data acquisition module and obtain the analysis and judgment results; Result forwarding module: It is used to forward the results processed and analyzed by the analysis and judgment module to the remote monitoring unit; Report generation module: It is used to generate a text processing report according to the analysis and judgment results.

[0007] Furthermore, the intelligent generation unit includes: an intelligent segmentation module, a key extraction module, a language organization module, and a summary output module; Intelligent segmentation module: It is used to segment a literary work intelligently according to the analysis results of the data processing unit; Key extraction module: It extracts keywords from a literary work according to the analysis results of the data processing unit; Language organization module: It is used to organize the extracted keywords; Summary output module: used to output the results by organizing the language organized by the language organization module.

[0008] Furthermore, the remote monitoring unit includes: a data receiving module and a data display module; Data receiving module: used to receive in real time the information data stored in the data storage unit and the analysis result information of the data processing unit; Data display module: used to display the information data received by the data receiving module on the display screen of the remote platform terminal.

[0009] Furthermore, the analysis and judgment module can obtain the TF-IDF weight information data of the words in the literary work from the data acquisition module, and obtain the paragraph similarity index through the analysis of the data obtained above: Wherein, is the paragraph similarity index, is the paragraph in the th word's TF-IDF weight, is the paragraph in the th word's TF-IDF weight, is the total number of common words between adjacent paragraphs and , is the part-of-speech weight coefficient of the th word, and the in the denominator is the modulus length for normalization, that is, calculating the length of each paragraph vector.

[0010] Furthermore, the analysis and judgment module can obtain the paragraph length information data in the literary work from the data acquisition module, and obtain the optimized segmentation index through the analysis of the data obtained above: Wherein, is the optimized segmentation index, is the th paragraph of text, is the total number of segments after segmentation, is the th paragraph of text, is the length of the paragraph , is the preset average paragraph length, is the similarity between adjacent paragraphs and .

[0011] Furthermore, the analysis and judgment module can obtain the number of words, the part-of-speech of words, and the position information data of words in a literary work from the data acquisition module, and analyze the obtained data to obtain a keyword index: Among them, is the keyword index, is the number of times the keyword appears in the literary work, is the total number of words in the literary work, is the total number of segments in the literary work, is the number of segments containing the keyword in the literary work, is the part-of-speech weight of the keyword, is the total number of part-of-speech types contained in the literary work, is the position weight of the keyword, is the preset conversion coefficient of the keyword.

[0012] According to another aspect of the present invention, a text processing method based on language and literature is provided. This method is implemented based on the above text processing system for language and literature, and specifically includes the following steps: S1. Receive the text data of the literary work to be processed, and then generate word vectors by segmenting the input text through the preprocessing unit, and filter out stop words and low-frequency words; S2. During the process of processing the text of the literary work, through the data acquisition unit, the word weight, paragraph length, number of words in the literary work, part-of-speech and position information data of words are collected in real time; S3. Transmit and store the collected information data, and perform real-time analysis on the information data to obtain an analysis result; S4. According to the analysis result, through the intelligent generation unit, segment the literary work and extract keywords, and organize the language of the keywords to summarize each segment.

[0013] The beneficial effects of a text processing system and method based on language and literature according to the present invention are as follows: Through the preprocessing unit, word vectors are generated for literary works, and stop words and low-frequency words are filtered, providing a high-quality data basis for subsequent processing. The data acquisition unit collects key information such as word weights, paragraph lengths, word counts, part-of-speech, and positions in real time, providing comprehensive and accurate data support for in-depth analysis of the works. Subsequently, based on the analysis results of the data processing unit, the intelligent generation unit can intelligently segment literary works, extract keywords for language organization, and generate summaries for each paragraph. This function not only helps readers quickly grasp the structure and hierarchy of the works but also effectively improves reading efficiency and understanding depth. The remote monitoring unit can display information data and analysis results in real time, enabling users to understand the processing progress and results of literary works at any time, increasing the transparency of the processing and user participation. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present invention will be further described in detail below with reference to the drawings and specific implementation methods.

[0015] Figure 1 is a schematic diagram of the system principle; Figure 2 is a schematic diagram of the method flow. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The present invention will be described in detail below with reference to the drawings and embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0017] According to one aspect of the present invention, as Figure 1-2 shown, a text processing system and method based on language and literature are provided. First, the text input unit receives the text data of the literary work to be processed, and then the preprocessing unit generates word vectors after word segmentation of the input text, filtering stop words and low-frequency words.

[0018] Then, during the processing of the text of the literary work, through the weight acquisition module, word length acquisition module, word count acquisition module, part-of-speech acquisition module, and position acquisition module in the data acquisition unit, the TF-IDF weights of the words in the literary work, the paragraph lengths in the literary work, the word counts in the literary work, the part-of-speech of the words, and the position information data of the words are collected in real time.

[0019] Then, the collected information data is transmitted and stored, and the information data is analyzed in real time to obtain analysis results, and a text processing report can be generated according to the processing analysis results, and then the obtained processing analysis results are forwarded to the remote monitoring unit and the intelligent generation unit.

[0020] After that, through the intelligent segmentation module, according to the analysis results of the data processing unit, the literary work is intelligently segmented. Then, through the key extraction module, keywords are extracted from the literary work, and through the language organization module, the extracted keywords are organized. Finally, through the summary output module, the language organized by the language organization module is output as a result.

[0021] The data receiving module in the remote monitoring unit can receive in real time the information data stored in the data storage unit and the analysis results of the data processing unit, and display the received information data on the background terminal display screen through the data display module.

[0022] The analysis and judgment module can obtain the TF-IDF weight information data of the words in the literary work from the data acquisition module, and analyze the obtained data to obtain the paragraph similarity index: Among them, is the paragraph similarity index, is the th TF-IDF weight of the word in the paragraph, is the th TF-IDF weight of the word in the paragraph, is the total number of common words in adjacent paragraphs and , is the part-of-speech weight coefficient of the th word (for example, for a noun = 1.2, for a verb = 1.0). The in the denominator is the modulus length used for normalization, that is, calculating the length of each paragraph vector. It ensures that the calculated similarity is a normalized value and eliminates the influence of paragraph length on similarity.

[0023] The analysis and judgment module can obtain the paragraph length information data in the literary work from the data acquisition module, and analyze the obtained data to obtain the optimized segmentation index: If the optimized segmentation index between adjacent paragraphs is low, it indicates that there are significant differences in their content, and the segmentation result is more in line with the logical flow of the text. Among them, is the optimized segmentation index, is the th paragraph text (the result after segmentation), is the total number of paragraphs after segmentation, is the th paragraph text (the result after segmentation), is the length (number of words or sentences) of paragraph , is the preset average paragraph length, For adjacent segments and the similarity therebetween (which can be calculated by the formula of the paragraph similarity index). Suppose there is a text, and the possible segmentation results after preliminary analysis are as follows: Segment 1: Introduction of background (Paragraphs A, B, C); Segment 2: Development of events (Paragraphs D, E); Segment 3: Climax part (Paragraphs F, G); Segment 4: Ending (Paragraphs H, I). Then, calculate the optimized segmentation index of adjacent paragraphs: (Segment 1, Segment 2); (Segment 2, Segment 3); (Segment 3, Segment 4), and obtain the objective function value. By adjusting the segmentation points (such as classifying Paragraph C into Segment 2), find the segmentation scheme that minimizes the objective function value.

[0024] The analysis and judgment module can obtain the data of the number of words, the part-of-speech of words, and the position information of words in the literary work from the data acquisition module, and analyze and obtain the keyword index through the above-obtained data analysis: Among them, is the keyword index, is the number of times the keyword appears in the literary work, is the total number of words in the literary work, is the total number of segments in the literary work, is the number of segments containing the keyword in the literary work, is the part-of-speech weight of the keyword, (if the part-of-speech of the keyword is a noun or a proper noun, the part-of-speech weight is 1.5; if the part-of-speech of the keyword is a verb, the part-of-speech weight is 1.2; for other parts-of-speech, the part-of-speech weight is 1), is the total number of types of parts-of-speech contained in the literary work, is the position weight of the keyword, (if the keyword appears in the title, at the beginning of a paragraph or at the end of a paragraph, the position weight of the keyword is 1.5; for other positions, the position weight of the keyword is 1), is the conversion coefficient of the preset keyword.

[0025] The above-described embodiments only represent one implementation manner of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. A text processing system based on language and literature, characterized in that: include: Text input unit, preprocessing unit, data acquisition unit, data transmission unit, data storage unit, data processing unit, intelligent generation unit and remote monitoring unit; Text input unit: used for receiving text data of literary works to be processed; Preprocessing unit: used to generate word vectors after word segmentation of input text and filter out stop words and low-frequency words; Data collection unit: used to collect information data generated in real time during the text processing of literary works; Data transmission unit: used for real-time transmission of information data collected by the data collection unit; Data storage unit: used to store the information data transmitted in real time by the data transmission unit; Data processing unit: used to process and analyze the information data stored in the data storage unit, and obtain the processing and analysis results, and can generate a text processing report based on the processing and analysis results, and then forward the obtained processing and analysis results to the remote monitoring unit and the intelligent generation unit; Intelligent generation unit: used to segment and extract keywords from literary works according to the analysis results of the data processing unit, and organize the extracted keywords into language to summarize each segment; Remote monitoring unit: used to display the information data stored in the data storage unit and the analysis results of the data processing unit in real time on the remote platform terminal.

2. A language and literature-based word processing system according to claim 1, characterized in that: The data collection unit includes: a weight collection module, a word length collection module, a word number collection module, a part of speech collection module, and a position collection module; Weight collection module: used to collect TF-IDF weight information data of words in literary works in real time; Word length collection module: used to collect paragraph length information data in literary works in real time; Word count collection module: used to collect word count information data in literary works in real time; Part-of-speech acquisition module: used to collect the part-of-speech information data of words in literary works in real time; Position collection module: used to collect the position information data of words in literary works in real time.

3. A language and literature based word processing system according to claim 2, characterized in that: The data processing unit includes: a data acquisition module, an analysis and judgment module, a result forwarding module and a report generation module; Data acquisition module: used for real-time acquisition of information data stored in the data storage unit; Analysis and judgment module: used to process and analyze the information data obtained by the data acquisition module and obtain the results of analysis and judgment; Result forwarding module: used to forward the results of analysis processed by the analysis and judgment module to the remote monitoring unit; Report generation module: used to generate text processing reports based on the results of analysis and judgment.

4. A language and literature based word processing system according to claim 3, characterized in that: The intelligent generation unit includes: an intelligent segmentation module, a key extraction module, a language organization module and a summary output module; Intelligent segmentation module: used to intelligently segment literary works according to the analysis results of the data processing unit; Key extraction module: extract keywords from literary works based on the analysis results of the data processing unit; Language organization module: used to organize the extracted keywords; Summary output module: used to organize the language from the language organization module and output the results.

5. A language and literature based word processing system according to claim 4, characterized in that: The remote monitoring unit comprises: a data receiving module and a data display module; Data receiving module: used for receiving the information data stored in the data storage unit and the analysis result information of the data processing unit in real time; Data display module: used to display the information data received by the data receiving module on the remote platform terminal display screen.

6. A language and literature based word processing system according to claim 5, characterized in that: The analysis and judgment module can obtain the TF-IDF weight information data of the words in the literary works from the data acquisition module, and obtain the paragraph similarity index through the above acquired data analysis: in, is the paragraph similarity index, For paragraphs Middle The TF-IDF weight of the word, For paragraphs Middle The TF-IDF weight of the word, For adjacent paragraphs and The total number of common words, For the The part-of-speech weight coefficient of the word, , is the modulus length used for standardization, that is, calculating the length of each paragraph vector.

7. A language and literature based word processing system according to claim 6, characterized in that: The analysis and judgment module can obtain the paragraph length information data in the literary work from the data acquisition module, and obtain the optimized segmentation index through the analysis of the above acquired data: in, To optimize the segmentation index, For the paragraph text, is the total number of segments after segmentation, For the paragraph text, For paragraphs Length, is the preset average paragraph length, For adjacent segments and The similarity between .

8. A language and literature based word processing system according to claim 7, characterized in that: The analysis and judgment module can obtain the number of words, word parts and word location information data in the literary works from the data acquisition module, and obtain the keyword index through the above acquired data analysis: in, is the keyword index, is the number of times the keyword appears in the literary work, is the total number of words in the literary work, is the total number of segments in a literary work, is the number of segments containing keywords in the literary works, is the part-of-speech weight of the keyword, It is the total type of parts of speech contained in literary works. is the position weight of the keyword, The conversion coefficient of the preset keyword.

9. A text processing method based on language and literature, characterized in that: The method is implemented based on a language and literature-based text processing system according to any one of claims 1 to 8, and specifically comprises the following steps: S1, receiving the text data of the literary work to be processed, and then generating word vectors for the input text after word segmentation through the preprocessing unit, and filtering stop words and low-frequency words; S2. In the process of processing the text of the literary work, the word weight, paragraph length, number of words in the literary work, word part of speech and position information data are collected in real time through the data collection unit; S3, transmitting and storing the collected information data, and performing real-time analysis on the information data to obtain analysis results; S4. Based on the analysis results, the literary works are segmented and keywords are extracted through intelligent generation units, and the keywords are organized linguistically to summarize each segment.