Literature data visualization method based on AI large model

Through the literary data visualization method based on AI large models, combining character experiences with timelines and dynamically laying out event nodes, the problem of excessive data volume in novels is solved, and a focused display of complex narrative structures and an efficient reading experience are achieved.

CN120670637AInactive Publication Date: 2025-09-19HEFEI JIAYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510575014.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Data visualization of long literary works suffers from the problem of excessive data volume, which leads to information overload and cognitive fatigue for users. Traditional methods are unable to effectively focus on and display complex narrative structures.

Method used

A method based on an AI big model is adopted. By combining the character experience with the timeline, the Euclidean distance of the event nodes is determined based on the relevance degree K, and the weights of sentences and paragraphs are obtained through the AI ​​big model. Event nodes are dynamically arranged, and only sub-sorting related to the user is displayed. Deep learning and back-propagation algorithms are integrated for model training and verification.

Benefits of technology

It achieves a focused presentation of the complex narrative clues of the novel, reduces visual clutter, improves reading efficiency and analytical ability, can more accurately portray the relationship levels between characters and events, and provide a more efficient reading and research experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670637A_ABST
    Figure CN120670637A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI (artificial intelligence), and particularly discloses a literature data visualization method based on an AI large model, which comprises the following steps of: sorting all things experienced by a single person according to a time axis sequence to obtain a first sequence; determining a main sequence, and obtaining a sub-sequence of a single event in the main sequence; and determining a correlation degree K between the target person and the event Si based on an AI large model, determining an Euclidean distance between the sub-order and a node based on the correlation degree K, and connecting a starting point of the sub-order and the node through a straight line to complete visualization of a single node. According to the method, the complexity of the visualized literature data can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI technology, and in particular to a literary data visualization method based on an AI large model. Background Art

[0002] Literary data visualization refers to a research method that converts textual information in literary works into visual charts or graphs in order to more intuitively analyze and interpret literary content. It combines literary research and data science, and helps discover patterns that are difficult to detect in traditional reading by quantifying text features (such as character relationships, emotional changes, topic distribution, etc.) and presenting them in charts.

[0003] For long literary works, even after data visualization, the problem of excessive data volume may still exist, and users may still experience information overload and cognitive fatigue when viewing them. Because novels often contain complex character networks and multi-line narrative structures, even if the text is converted into a timeline through visualization, the resulting chart may still show dense data nodes and complex connection lines. Therefore, to avoid this problem, we propose a literary data visualization method based on a large AI model. Summary of the Invention

[0004] The purpose of this invention is to provide a literary data visualization method based on an AI large model to solve the following technical problems:

[0005] Even after data visualization, long-form literary works may still present the problem of excessive data volume, potentially leading users to experience information overload and cognitive fatigue. Because novels often contain complex character networks and multi-line narrative structures, even if the text is converted into a timeline through visualization, the resulting charts may still display dense data nodes and complex connection lines.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] A literary data visualization method based on an AI large model includes the following steps:

[0008] Get the things that happened in the literary work, sort all the things experienced by a single character in chronological order, and get the first order;

[0009] Mark the first ranking corresponding to the protagonist in the literature as the main ranking, obtain the i-th event in the main ranking, record it as event Si, obtain the character that appears in event Si, record it as the target character, and intercept the part from the first event in the first ranking to event Si, record it as the sub-ranking;

[0010] Based on the AI ​​big model, the correlation degree K between the target person and the event Si is determined, and based on the correlation degree K, the Euclidean distance between the sub-sorting and the node is determined. The node is the event Si. A straight line is used to connect the starting point of the sub-sorting and the node to complete the visualization of a single node.

[0011] As a further solution of the present invention: the process of determining the relevance K between the target person and the event Si based on the AI ​​big model includes:

[0012] Obtain the paragraph corresponding to the event Si, record it as the target paragraph, obtain the sentence in which the target person appears in the target paragraph, record it as the target sentence;

[0013] Obtain the weight of the target sentence relative to the target paragraph based on the AI ​​large model, generate a weight set kjhj={kj,1,kj,2,…,kj,n}, where kj,n represents the weight of the j-th target sentence relative to the n-th target paragraph, and obtain the maximum weight kmaxj=max(kjhj);

[0014] The mean of the maximum weights is calculated as the correlation degree K.

[0015] As a further solution of the present invention: obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​large model includes:

[0016] Establishing a database, wherein the database stores target sentences and corresponding target paragraphs with labeled weights;

[0017] An AI big model is established based on deep learning, the AI ​​big model is trained and verified through the database, the target sentence and the target paragraph are input into the verified AI big model, and the weight of the target sentence relative to the target paragraph is obtained.

[0018] As a further solution of the present invention: the weights of the target sentences and the corresponding target paragraphs in the database are based on manual annotation.

[0019] As a further solution of the present invention: when a user views an event Si, only the sub-rankings connected to the event Si are displayed to the user.

[0020] As a further solution of the present invention: determining the Euclidean distance between the sub-sort and the node based on the correlation degree K includes: calculating the Euclidean distance D=η / K, where η is a preset standard Euclidean distance.

[0021] As a further solution of the present invention: the AI ​​large model is trained using a back-propagation algorithm, and the AI ​​large model is verified through seven-fold cross-validation.

[0022] As a further solution of the present invention, the process of obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​large model further includes:

[0023] Randomly select m target sentences and target paragraphs corresponding to weights, and manually obtain the weights of the target sentences relative to the target paragraphs, where m is a preset number;

[0024] When the difference between the manually obtained weights and the weights obtained by the AI ​​large model is greater than the preset difference threshold, the weights obtained by the AI ​​large model are marked as abnormal weights. When the ratio of the number of abnormal weights to m is greater than 0.4, the AI ​​large model is retrained and verified.

[0025] The beneficial effects of the present invention are as follows:

[0026] 1) The present invention combines the character's experience with the timeline and dynamically arranges the event nodes based on the relevance K, thereby achieving a focused presentation of the complex narrative clues in the novel. Compared with traditional visualization methods, the present invention can not only display the main plot of the protagonist, but also appropriately hide or fold secondary characters and side plots, thereby reducing visual clutter. By only displaying the sub-sorts closely related to the events the user is looking for, readers can quickly grasp the key plot trends and, when necessary, deeply expand and examine the implicit information and character interactions. This greatly reduces the cognitive load when faced with huge amounts of data, significantly improves the readability and analysis efficiency of literary content, and provides subsequent researchers with a more convenient macro-micro combined reading mode;

[0027] 2) By obtaining the relevant weights of sentences and paragraphs based on the AI ​​big model, the present invention can more accurately portray the relationship levels between characters and events. Compared with simple extensive statistics based on word frequency or traditional text analysis, the deep learning model of the present invention can automatically identify the emotions, motivations and implications hidden in the semantics, providing richer dimensions for visualization charts. The mean of the maximum weight is used as the correlation degree K, which can not only highlight the characters that have a significant impact on the development of the plot, but also make the interactive relationship of minor characters more appropriately positioned in the visualization, helping users to quickly grasp the character levels and event focus during the reading process, greatly improving the insight into literary structure, and providing more powerful technical support for multi-angle research, especially when interpreting deep-level content such as character personality and emotional transitions;

[0028] 3) The present invention also incorporates a dynamic detection and correction mechanism into the process of visualization generation, which greatly enhances the robustness and scalability of the system. By randomly extracting m groups of data for manual verification, once it is found that the weights output by the AI ​​large model are significantly different from the manual annotations, the system can trigger a new round of model training and verification, thereby continuously improving the model performance. This self-iteration capability can not only iterate more accurate visualization results from the continuously accumulated data, but can also adapt to long literary works of different types and sizes, so that whether it is a popular novel with complex character relationships or an epic masterpiece with multiple plot lines, a clearer, concise and easy-to-interact visualization can be obtained, achieving a more efficient reading and research experience, and further stimulating the potential and value of literary data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention will be further described below with reference to the accompanying drawings.

[0030] Figure 1 This is a flowchart of a literary data visualization method based on an AI large model of the present invention. DETAILED DESCRIPTION

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0032] See also Figure 1 As shown, the present invention is a literary data visualization method based on an AI large model, comprising the following steps:

[0033] Get the things that happened in the literary work, sort all the things experienced by a single character in chronological order, and get the first order;

[0034] It should be noted that through natural language processing technology and manual verification, all key events experienced by a character in a literary work are identified. These events are then arranged in chronological order based on the time clues provided in the text (for example, enrollment in one year and a conflict the following year, or the relationship between events). This constructs the character's complete experience from the beginning to the end of the story, which is recorded as the first order. By arranging these events in sequence, the character's complete activity trajectory can be presented in the time dimension, laying the foundation for subsequent visualization and in-depth analysis.

[0035] Mark the first ranking corresponding to the protagonist in the literature as the main ranking, obtain the i-th event in the main ranking, record it as event Si, obtain the character that appears in event Si, record it as the target character, and intercept the part from the first event in the first ranking to event Si, record it as the sub-ranking;

[0036] It is worth noting that after establishing the first ranking of different characters, the sequence corresponding to the protagonist can be marked as the main ranking to highlight the main thread of the story. Then, an event at a certain moment is selected from this main ranking, recorded as the i-th event Si, and other characters that appear in the event are retrieved, called target characters. Since these target characters may also have their own complete event sequences, we only intercept the part from the first event to the Si event in their respective first rankings to form a sub-ranking to represent the character activity trajectory that was related to the event before the protagonist's i-th event. In this way, no matter how many characters are involved in the protagonist's i-th event, their respective sub-rankings can be used to quickly understand their causes and experiences, which makes it easier to highlight the relationship between the protagonist's event and multiple characters in visualization or subsequent analysis;

[0037] Determine the correlation degree K between the target person and the event Si based on the AI ​​big model, determine the Euclidean distance between the sub-sort and the node based on the correlation degree K, and connect the starting point of the sub-sort and the node with a straight line to complete the visualization of a single node;

[0038] In another preferred embodiment of the present invention, determining the Euclidean distance between the sub-ranking and the node based on the correlation degree K includes: calculating the Euclidean distance D=η / K, where η is a preset standard Euclidean distance;

[0039] By analyzing the strength of the association between the target person and event Si in the text through the AI ​​large model, a numerical correlation degree K can be obtained. Subsequently, the Euclidean distance between the sub-sort and event Si is determined based on this K value: if the K value is larger, it means that the connection between the person and event Si is closer, and the end point of the sub-sort can be visually brought closer to the position of Si; if the K value is smaller, it will be further away. In the final visualization diagram, event Si will be presented as a node, similar to a circular mark or icon that highlights an important event in a relationship diagram, and then connected to the starting point of the sub-sort with a straight line. This method can not only intuitively show the closeness between the character's experience and the event, but also make it easier for readers to quickly identify and understand the depth of interaction between each character and the event in a complex network relationship.

[0040] In a preferred embodiment of the present invention, the process of determining the relevance K between the target person and the event Si based on the AI ​​large model includes:

[0041] Obtain the paragraph corresponding to the event Si, record it as the target paragraph, obtain the sentence in which the target person appears in the target paragraph, record it as the target sentence;

[0042] Obtain the weight of the target sentence relative to the target paragraph based on the AI ​​large model, generate a weight set kjhj={kj,1,kj,2,…,kj,n}, where kj,n represents the weight of the j-th target sentence relative to the n-th target paragraph, and obtain the maximum weight kmaxj=max(kjhj);

[0043] Calculate the mean of the maximum weights as the correlation degree K;

[0044] It is worth noting that obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​big model includes:

[0045] Establishing a database, wherein the database stores target sentences and corresponding target paragraphs with labeled weights;

[0046] An AI big model is established based on deep learning, the AI ​​big model is trained and verified through the database, the target sentence and the target paragraph are input into the verified AI big model, and the weight of the target sentence relative to the target paragraph is obtained.

[0047] It should be noted that the weights of the target sentences and the corresponding target paragraphs in the database are based on manual annotations;

[0048] It is understandable that by defining the sentences in the target paragraph where the target character appears as target sentences and using the AI ​​model to calculate their weight relative to the target paragraph, it is possible to more accurately capture the key information in the text that is closely related to the target character. By extracting the maximum weight of each target sentence and then taking the average, a comprehensive indicator can be formed, thereby more comprehensively measuring the degree of influence of the character on the event Si. This not only avoids a one-sided understanding of the text content, but also provides a more scientific basis for subsequent visual layout, ensuring that when presenting the relationship between characters and events, the truly decisive interactive level can be highlighted, ultimately helping readers or researchers to more quickly understand the story development context and the role of the characters;

[0049] It is understandable that the AI ​​large model is trained using a back-propagation algorithm and validated using seven-fold cross validation;

[0050] It is worth noting that the process of obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​big model also includes:

[0051] Randomly select m target sentences and target paragraphs corresponding to weights, and manually obtain the weights of the target sentences relative to the target paragraphs, where m is a preset number;

[0052] When the difference between the manually obtained weights and the weights obtained by the AI ​​large model is greater than the preset difference threshold, the weights obtained by the AI ​​large model are marked as abnormal weights. When the ratio of the number of abnormal weights to m is greater than 0.4, the AI ​​large model is retrained and verified;

[0053] It should be noted that the AI ​​model is trained through backpropagation and uses 7-fold cross-validation to ensure its accuracy and robustness. When calculating the weight of the target sentence relative to the target paragraph, some sentences and paragraphs are randomly sampled for manual verification. This allows for the timely detection of abnormal weights that deviate significantly from the manual annotation results. Once the number of these abnormal weights exceeds a certain proportion, retraining and verification are triggered, and the model is continuously iterated and optimized. This not only ensures the credibility of the visualization data itself, but also allows the model to gradually improve its perception of the deep semantics of the text as it continuously absorbs new annotation experience, thereby providing a more accurate and valuable reference basis for the final visualization of the relationship between people and events.

[0054] In another preferred embodiment of the present invention, when a user views an event Si, only the sub-rankings connected to the event Si are displayed to the user.

[0055] It's important to note that displaying only the relevant sub-rankings when a user is viewing a specific event Si effectively avoids presenting excessive redundant information in complex literary works, thereby reducing visual overload and cognitive strain. By focusing on the content most relevant to the currently searched event, users no longer need to sift through massive amounts of data for useful information. This allows them to more fully grasp key plot points and character connections, while also creating a more flexible interactive environment for subsequent viewing of other events or characters. Ultimately, this helps readers understand the story structure more efficiently and intuitively, significantly enhancing the visual reading experience.

[0056] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A literary data visualization method based on AI big model, characterized by: The following steps are involved: Get the things that happened in the literary work, sort all the things experienced by a single character in chronological order, and get the first order; Mark the first ranking corresponding to the protagonist in the literature as the main ranking, obtain the i-th event in the main ranking, record it as event Si, obtain the character that appears in event Si, record it as the target character, and intercept the part from the first event in the first ranking to event Si, record it as the sub-ranking; Based on the AI ​​big model, the correlation degree K between the target person and the event Si is determined, and based on the correlation degree K, the Euclidean distance between the sub-sorting and the node is determined. The node is the event Si. A straight line is used to connect the starting point of the sub-sorting and the node to complete the visualization of a single node.

2. The method for visualizing literary data based on an AI large model according to claim 1, characterized in that: The process of determining the relevance K between the target person and the event Si based on the AI ​​big model includes: Obtain the paragraph corresponding to the event Si, record it as the target paragraph, obtain the sentence in which the target person appears in the target paragraph, record it as the target sentence; Based on the AI ​​big model, the weight of the target sentence relative to the target paragraph is obtained to generate the weight set kjh j ={k j,1 , k j,2 ,…,k j,n }, k j,n Represents the weight of the j-th target sentence relative to the n-th target paragraph, and obtains the maximum weight kmax j =max(kjh j ); The mean of the maximum weights is calculated as the correlation degree K.

3. The method for visualizing literary data based on an AI large model according to claim 2, characterized in that: Obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​big model includes: Establishing a database, wherein the database stores target sentences and corresponding target paragraphs with labeled weights; An AI big model is established based on deep learning, the AI ​​big model is trained and verified through the database, the target sentence and the target paragraph are input into the verified AI big model, and the weight of the target sentence relative to the target paragraph is obtained.

4. The method for visualizing literary data based on an AI large model according to claim 3, characterized in that: The weights of the target sentences and the corresponding target paragraphs in the database are based on manual annotations.

5. The method for visualizing literary data based on an AI large model according to claim 1, characterized in that: When a user views an event Si, only the sub-sequences connected to the event Si are displayed to the user.

6. The method for visualizing literary data based on an AI large model according to claim 1, characterized in that: Determining the Euclidean distance between the sub-ranking and the node based on the correlation degree K includes: calculating the Euclidean distance D=η / K, where η is a preset standard Euclidean distance.

7. The method for visualizing literary data based on an AI large model according to claim 3, characterized in that: The AI ​​large model is trained using a back-propagation algorithm and verified using seven-fold cross-validation.

8. The method for visualizing literary data based on an AI large model according to claim 2, characterized in that: The process of obtaining the weight of the target sentence relative to the target paragraph based on the AI ​​big model also includes: Randomly select m target sentences and target paragraphs corresponding to weights, and manually obtain the weights of the target sentences relative to the target paragraphs, where m is a preset number; When the difference between the manually obtained weights and the weights obtained by the AI ​​large model is greater than the preset difference threshold, the weights obtained by the AI ​​large model are marked as abnormal weights. When the ratio of the number of abnormal weights to m is greater than 0.4, the AI ​​large model is retrained and verified.