Text data analysis method and device, computer equipment and storage medium

By extracting entities from text and building a thermal knowledge graph, combined with the deep understanding of a large language model, the problem of low text evaluation efficiency is solved, and fast and accurate text quality assessment is achieved.

CN120849577APending Publication Date: 2025-10-28SHANGHAI IQIYI NEW MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510692533.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing text evaluation scenarios rely on professional experts, which are inefficient and highly subjective, making it difficult to quickly and accurately evaluate the text content quality of massive scripts or long TV drama scripts.

Method used

By extracting entities from the text to be analyzed, constructing a heatmap knowledge graph, and combining it with the deep understanding results of a pre-set large language model, text analysis results are generated.

Benefits of technology

It improves the efficiency and accuracy of text evaluation, and can accurately determine the results of text analysis, making it more accurate and efficient than manual evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849577A_ABST
    Figure CN120849577A_ABST
Patent Text Reader

Abstract

The invention relates to a text data analysis method and device, computer equipment and a storage medium. The method comprises the following steps: visually displaying each entity feature and the importance of an entity relationship through a thermodynamic knowledge graph, and synthesizing a deep understanding result of a to-be-analyzed text by the thermodynamic knowledge graph and a preset large language model so as to accurately determine a text analysis result of the to-be-analyzed text. Compared with manual judgment of the text quality, the text evaluation efficiency and the accuracy of the evaluation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular to a text data analysis method, apparatus, computer device, and storage medium. Background Technology

[0002] Existing text evaluation scenarios include script evaluation, novel evaluation, and other text evaluation scenarios. Traditional text evaluation mainly relies on professional experts, which is inefficient and highly subjective. Faced with a massive amount of scripts or long TV drama scripts, how to quickly and accurately evaluate the quality of text content has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides a text data analysis method, apparatus, computer device, and storage medium to address the problem of how to quickly and accurately assess the quality of text content.

[0004] Firstly, this application provides a text data analysis method, the method comprising:

[0005] Entity extraction is performed on the text to be analyzed to obtain multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0006] A heat map is constructed based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0007] The deep understanding results of the text to be analyzed by the preset large language model are obtained, and the text analysis results are obtained by combining the deep understanding results and the heat map.

[0008] Secondly, this application provides a text data analysis device, the device comprising:

[0009] The extraction module is used to extract entities from the text to be analyzed, and obtain multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0010] A construction module is used to construct a heat map knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0011] The analysis module is used to obtain the deep understanding results of the preset large language model on the text to be analyzed, and to obtain the text analysis results by combining the deep understanding results with the heat map.

[0012] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described text data analysis method.

[0013] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for executing the above-described text data analysis method.

[0014] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application extracts entities from the text to be analyzed, obtaining multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship; constructs a heat map knowledge graph based on the multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship; obtains the deep understanding result of the preset large language model on the text to be analyzed, and combines the deep understanding result and the heat map knowledge graph to obtain the text analysis result.

[0015] Based on the above method, the importance of each entity's features and relationships is visualized through a heat map knowledge graph. By combining the heat map knowledge graph with the pre-set large language model to deeply understand the text to be analyzed, the text analysis results can be accurately determined. Compared with manual evaluation of text quality, this method improves the efficiency of text evaluation and the accuracy of evaluation results. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0019] Figure 1 A flowchart illustrating a text data analysis method provided in an embodiment of this application;

[0020] Figure 2 A schematic diagram illustrating the extraction results of entity features and entity relationships provided in an embodiment of this application;

[0021] Figure 3 A schematic diagram of the structure of a thermal knowledge graph provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram illustrating the effect of a quantitative analysis result provided in an embodiment of this application.

[0023] Figure 5 A structural block diagram of a text data analysis device provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0027] In one embodiment, Figure 1 This is a flowchart illustrating a text data analysis method in one embodiment, with reference to... Figure 1 This paper provides a text data analysis method. This embodiment primarily applies this method to a terminal or server equipped with a pre-set large language model. The pre-set large language model refers to a deep learning model (LLM) trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. The text data analysis method specifically includes the following steps:

[0028] Step S210: Entity extraction is performed on the text to be analyzed to obtain multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0029] Specifically, the text to be analyzed is any narrative text requiring text quality assessment, such as scripts, novels, biographies, stories, fairy tales, myths, etc. Text quality assessment refers to evaluating the narrative structure, character relationships, emotional trajectory, or subplots to determine their plausibility and whether they are padded. Specifically, entity extraction can be performed on the text using dictionary matching, regular expressions, corpus-trained machine learning models, or neural network models to obtain multiple entity features and the relationships between them. Then, the frequency of occurrence of each entity feature or relationship can be statistically analyzed to determine the importance of each feature and relationship.

[0030] The entity features extracted from the text to be analyzed include people, items, organizations, scenes, events, etc. The importance of entity features and entity relationships refers to the degree of influence on the text to be analyzed.

[0031] Step S220: Construct a heat map knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0032] Specifically, the heatmap knowledge graph is used to visualize the features of various entities and the relationships between them. It displays entity features and relationships of varying importance using different visual effects, with higher-importance features and relationships being displayed more prominently. Different display styles are used for entity features and relationships of different importance, including color style, font style, line style, display scale, display position, and display frame. The display style for higher-importance entity features and relationships is more prominent. For example, the display style for a highly important entity feature is: red color, bold font, thick solid line, 10% display scale of the heatmap knowledge graph, central area of ​​the heatmap knowledge graph, and square frame.

[0033] Step S230: Obtain the deep understanding results of the preset large language model on the text to be analyzed, and combine the deep understanding results with the heat map to obtain the text analysis results.

[0034] Specifically, a pre-defined large language model is used to perform deep semantic analysis on the text to be analyzed to obtain deep understanding results. The deep understanding results include at least the hidden logic of the text, the text quality assessment results, and the market analysis results. The deep understanding results are combined with the narrative structure in the heat map to generate text analysis results. The text analysis results include standardized text quality assessment results and market analysis results. The text quality assessment results include assessment results from different dimensions such as plot completeness, character development ability, theme expression completeness, and logical loopholes. The market analysis results include audience acceptance, feasibility assessment results for film and television adaptation, and suggestions for content conversion across multiple platforms.

[0035] Based on the above method, the importance of each entity's features and relationships is visualized through a heat map knowledge graph. By combining the heat map knowledge graph with the pre-set large language model to deeply understand the text to be analyzed, the text analysis results can be accurately determined. Compared with manual evaluation of text quality, this method improves the efficiency of text evaluation and the accuracy of evaluation results.

[0036] In one embodiment, the entity extraction of the text to be analyzed, obtaining multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship, includes:

[0037] Obtain the original text and perform standard preprocessing on the original text to obtain the text to be analyzed;

[0038] The preset large language model is used to extract entities from the text to be analyzed, resulting in multiple entity features and entity relationships between the entity features.

[0039] The importance of multiple entity features and entity relationships in the text to be analyzed is obtained by using the preset large language model, thereby obtaining the importance of each entity feature and the importance of each entity relationship.

[0040] Specifically, the original text is the narrative text that has not undergone preprocessing. Standardized preprocessing of the original text mainly includes format conversion and chapter division. Format conversion refers to converting the initial format of the original text to a unified standard format. Chapter division is to divide the original text according to the scenes, episodes, and paragraphs to divide it into different paragraphs or episodes, thereby obtaining the standard format text to be analyzed with paragraph divisions.

[0041] The system utilizes a pre-defined large language model to extract entity features from the text to be analyzed, thereby extracting multiple different entity features and entity relationships. The system then performs an importance analysis on the extracted entity features and entity relationships, specifically determining the importance of each entity feature and entity relationship based on the frequency of their occurrence and their influence on the narrative direction of the text to be analyzed.

[0042] Reference Figure 2 "I Am a Detective" (script title 1), "The Witness" (script title 2), and "The Hidden Corner" (script title 3) are three different scripts. Qin Chuan (character name 1) in ['Qin Chuan', 375.0] is an entity characteristic corresponding to a character in "I Am a Detective," and 375.0 represents the importance of that entity characteristic. Wu Deying (character name 2)_Qin Chuan in ['Wu Deying_Qin Chuan', 87.0] is an entity relationship in "I Am a Detective," and 87.0 represents the importance of that entity relationship. Based on... Figure 2 It allows for an intuitive understanding of the characteristics of each entity and the importance of entity relationships.

[0043] In one embodiment, constructing a heatmap knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship includes:

[0044] Construct a text knowledge graph based on multiple entity features and the entity relationships between these features;

[0045] Based on the importance of each entity feature and the importance of each entity relationship, determine the display mode of each entity feature and the display mode of each entity relationship;

[0046] In the text knowledge graph, each entity feature and each entity relationship is displayed according to the display mode of each entity feature and the display mode of each entity relationship, respectively, to obtain a heat map knowledge graph.

[0047] Specifically, a text knowledge graph is constructed based on multiple entity features and the entity relationships between them. While the text knowledge graph displays the narrative structure of the text to be analyzed, it cannot visualize the importance of each entity feature and entity relationship. Therefore, based on the text knowledge graph, the information is differentiated and displayed according to the importance of each entity feature and entity relationship, referencing... Figure 3 ( Figure 3 and Figure 4(For illustration purposes only; some ambiguous content does not affect the understanding of this scheme.) Different display styles are used to show entity features and entity relationships of different importance, with emphasis on highlighting entity features and entity relationships of higher importance, ultimately forming a heat map that can intuitively display the core structure and important information of the text to be analyzed.

[0048] In one embodiment, obtaining the deep understanding results of the preset large language model on the text to be analyzed, and combining the deep understanding results with the heatmap knowledge graph to obtain the text analysis results, includes:

[0049] Obtain the deep understanding results of the text to be analyzed by the preset large language model;

[0050] Cluster analysis was performed on the aforementioned thermodynamic knowledge graph to obtain quantitative analysis results;

[0051] The text analysis results are generated based on the deep understanding results and the quantitative analysis results.

[0052] Specifically, a pre-defined large language model is used to perform deep understanding of the text to be analyzed in order to generate deep understanding results, which capture information such as the overall theme, sentiment, and narrative structure of the text to be analyzed. Based on this information, the hidden logic of the text, the text quality assessment results, and the market analysis results are derived.

[0053] Cluster analysis of the heatmap knowledge graph is used to reveal the structural features of the text to be analyzed, such as character relationships, plot distribution, and narrative structure. This yields quantitative analysis results, which clearly demonstrate the overall picture of character relationships, the patterns in plot distribution, and the overall narrative flow, providing strong support for a deeper understanding of the text content and creative logic. The deep understanding results are then combined with the heatmap knowledge graph to automatically generate standardized text analysis results.

[0054] In one embodiment, the clustering analysis of the thermal knowledge graph to obtain quantitative analysis results includes:

[0055] The thermal knowledge graph is subjected to feature clustering to obtain multiple cluster sets;

[0056] The correlation degree is calculated for multiple cluster sets to determine the degree of correlation between each cluster set;

[0057] Determine the edge features of each of the cluster sets;

[0058] The quantitative analysis results are generated based on the edge features of each cluster set and the degree of correlation between each cluster set.

[0059] Specifically, hierarchical clustering is performed on the heatmap knowledge graph to group similar entity features or entity relationships together. This involves clustering entity features or relationships based on the frequency of character interactions, the location of events, etc., according to their importance. For example, frequently appearing characters are grouped together, indicating they may belong to the same social circle or storyline. Multiple cluster sets are obtained through clustering, and multiple entity features or entity relationships within the same cluster set belong to the same character group or the same plot segment.

[0060] Calculating the degree of association between different clusters involves calculating the frequency of interaction between different character groups and the probability of transitions between plot segments. This analysis helps to clarify the connections between different character groups and plot segments, clarify the relationships between various parts of the text being analyzed, and help to discover structural features such as the density of character relationship networks and the smoothness of plot transitions. It also allows for a direct understanding of which elements are closely connected and which are relatively independent.

[0061] Identifying edge features within each cluster involves identifying marginal features, which are either entity features or entity relationships. Edge features may have relatively weak connections to other elements within their own cluster, but still have some connection to other clusters. For example, a supporting character might have limited interaction within their own group but significant interaction with other groups. By identifying these edge points, we can uncover key elements in the script that connect different parts, which is crucial for understanding the narrative flow, as these edge points are often key nodes in plot twists and the expansion of character relationships.

[0062] Quantitative analysis results are generated based on the correlation between clusters and the edge features within the clusters. These results further subdivide the knowledge graph of the clusters based on the thermal knowledge graph, referencing... Figure 4 Classes 1, 2, 3, and 4 in the graph indicate different cluster sets. The quantitative analysis results provide an overall analysis of the structure presented by the heatmap, intuitively reflecting the topological structure of the character relationship network (such as star-shaped, network-shaped, etc.) and the hierarchical structure of plot development (such as linear, non-linear, etc.). The quantitative analysis results are used to comprehensively and systematically reveal the structural characteristics of the script, clearly presenting the full picture of character relationships, the pattern of plot distribution, and the overall narrative trend, providing strong support for a deeper understanding of the script's content and creative logic.

[0063] In one embodiment, the quantitative analysis results are generated based on the edge features of each cluster set and the degree of correlation between the cluster sets, including:

[0064] The edge features of each cluster set, the degree of correlation between each cluster set, and the quantitative analysis conditions are input into the preset large language model, and the output of the preset large language model for the quantitative analysis conditions is used as the quantitative analysis result.

[0065] Specifically, by using a pre-defined large language model, based on the edge features of each cluster set, the degree of correlation between each cluster set, and the quantitative analysis conditions, the quantitative analysis results that meet the quantitative analysis conditions are output. The quantitative analysis conditions can be customized according to actual analysis needs. The quantitative analysis conditions specifically include multiple quantitative analysis dimensions, including the topological structure of the character relationship network and the hierarchical architecture of the plot development, thereby realizing the automatic generation of quantitative analysis results.

[0066] In one embodiment, generating the text analysis result based on the deep understanding result and the quantitative analysis result includes:

[0067] The deep understanding results, the quantitative analysis results, and the text analysis conditions are input into the preset large language model, and the output of the preset large language model for the text analysis conditions is used as the text analysis result.

[0068] Specifically, by utilizing a pre-defined large language model based on deep understanding results, quantitative analysis results, and text analysis conditions, the system outputs text analysis results that meet the text analysis conditions. These conditions can be customized according to actual analysis needs and include multiple text analysis dimensions, such as plot density distribution, rhythm assessment data, emotional trend rationality, behavioral logic score, thematic relevance, logical loophole score, film and television adaptation cost estimation, and market potential assessment, thereby achieving automated generation of text analysis results.

[0069] Figure 1 This is a flowchart illustrating a text data analysis method in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0070] In one embodiment, such as Figure 5 As shown, a text data analysis device is provided, comprising:

[0071] Extraction module 310 is used to extract entities from the text to be analyzed, and obtain multiple entity features, the importance of each entity feature, the entity relationship between each entity feature, and the importance of each entity relationship;

[0072] Construction module 320 is used to construct a heat map knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship.

[0073] The analysis module 330 is used to obtain the deep understanding results of the preset large language model on the text to be analyzed, and combine the deep understanding results with the heat map to obtain the text analysis results.

[0074] In one embodiment, the extraction module 310 is further configured to:

[0075] Obtain the original text and perform standard preprocessing on the original text to obtain the text to be analyzed;

[0076] The preset large language model is used to extract entities from the text to be analyzed, resulting in multiple entity features and entity relationships between the entity features.

[0077] The importance of multiple entity features and entity relationships in the text to be analyzed is obtained by using the preset large language model, thereby obtaining the importance of each entity feature and the importance of each entity relationship.

[0078] In one embodiment, the building module 320 is further configured to:

[0079] Construct a text knowledge graph based on multiple entity features and the entity relationships between these features;

[0080] Based on the importance of each entity feature and the importance of each entity relationship, determine the display mode of each entity feature and the display mode of each entity relationship;

[0081] In the text knowledge graph, each entity feature and each entity relationship is displayed according to the display mode of each entity feature and the display mode of each entity relationship, respectively, to obtain a heat map knowledge graph.

[0082] In one embodiment, the analysis module 330 is further configured to:

[0083] Obtain the deep understanding results of the text to be analyzed by the preset large language model;

[0084] Cluster analysis was performed on the aforementioned thermodynamic knowledge graph to obtain quantitative analysis results;

[0085] The text analysis results are generated based on the deep understanding results and the quantitative analysis results.

[0086] In one embodiment, the analysis module 330 is further configured to:

[0087] The thermal knowledge graph is subjected to feature clustering to obtain multiple cluster sets;

[0088] The correlation degree is calculated for multiple cluster sets to determine the degree of correlation between each cluster set;

[0089] Determine the edge features of each of the cluster sets;

[0090] The quantitative analysis results are generated based on the edge features of each cluster set and the degree of correlation between each cluster set.

[0091] In one embodiment, the analysis module 330 is further configured to:

[0092] The edge features of each cluster set, the degree of correlation between each cluster set, and the quantitative analysis conditions are input into the preset large language model, and the output of the preset large language model for the quantitative analysis conditions is used as the quantitative analysis result.

[0093] In one embodiment, the analysis module 330 is further configured to:

[0094] The deep understanding results, the quantitative analysis results, and the text analysis conditions are input into the preset large language model, and the output of the preset large language model for the text analysis conditions is used as the text analysis result.

[0095] like Figure 6 As shown, this application provides a computer device including a processor 711, a communication interface 712, a memory 713, and a communication bus 714, wherein the processor 711, the communication interface 712, and the memory 713 communicate with each other through the communication bus 714.

[0096] Memory 713 is used to store computer programs;

[0097] When the processor 711 executes the program stored in the memory 713, it implements the text data analysis method provided in any of the foregoing method embodiments.

[0098] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0099] In one embodiment, the text data analysis apparatus provided in this application can be implemented as a computer program, which can be implemented in the form of, for example... Figure 6 It runs on the computer device shown. The computer device's memory can store the various program modules that make up the text data analysis device, for example, Figure 5 The extraction module 310, construction module 320, and analysis module 330 are shown. The computer program comprised of these modules causes the processor to execute the text data analysis methods of the various embodiments of this application described in this specification.

[0100] Figure 6 The computer equipment shown can be used as follows Figure 5 The extraction module in the text data analysis device shown performs entity extraction on the text to be analyzed, obtaining multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. The computer device, through the construction module, constructs a heatmap knowledge graph based on the multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. The computer device, through the analysis module, obtains the deep understanding results of the text to be analyzed from a preset large language model, and combines the deep understanding results with the heatmap knowledge graph to obtain the text analysis results.

[0101] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the text data analysis method provided in any of the foregoing method embodiments.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0104] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also mean including the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that alternatives or substitutions may be used.

[0105] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A text data analysis method, characterized in that, The method includes: Entity extraction is performed on the text to be analyzed to obtain multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. A heat map is constructed based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. The deep understanding results of the text to be analyzed by the preset large language model are obtained, and the text analysis results are obtained by combining the deep understanding results and the heat map.

2. The method according to claim 1, characterized in that, The entity extraction process for the text to be analyzed, which yields multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship, includes: Obtain the original text and perform standard preprocessing on the original text to obtain the text to be analyzed; The preset large language model is used to extract entities from the text to be analyzed, resulting in multiple entity features and entity relationships between the entity features. The importance of multiple entity features and entity relationships in the text to be analyzed is obtained by using the preset large language model, thereby obtaining the importance of each entity feature and the importance of each entity relationship.

3. The method according to claim 1, characterized in that, The step of constructing a heatmap knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship includes: Construct a text knowledge graph based on multiple entity features and the entity relationships between these features; Based on the importance of each entity feature and the importance of each entity relationship, determine the display mode of each entity feature and the display mode of each entity relationship; In the text knowledge graph, each entity feature and each entity relationship is displayed according to the display mode of each entity feature and the display mode of each entity relationship, respectively, to obtain a heat map knowledge graph.

4. The method according to claim 1, characterized in that, The step of obtaining the deep understanding results of the preset large language model on the text to be analyzed, and combining the deep understanding results with the heat map to obtain the text analysis results, includes: Obtain the deep understanding results of the text to be analyzed by the preset large language model; Cluster analysis was performed on the aforementioned thermodynamic knowledge graph to obtain quantitative analysis results; The text analysis results are generated based on the deep understanding results and the quantitative analysis results.

5. The method according to claim 4, characterized in that, The cluster analysis of the thermal knowledge graph yields quantitative analysis results, including: The thermal knowledge graph is subjected to feature clustering to obtain multiple cluster sets; The correlation degree is calculated for multiple cluster sets to determine the degree of correlation between each cluster set; Determine the edge features of each of the cluster sets; The quantitative analysis results are generated based on the edge features of each cluster set and the degree of correlation between each cluster set.

6. The method according to claim 5, characterized in that, The quantitative analysis results are generated based on the edge features of each cluster and the degree of correlation between the clusters, including: The edge features of each cluster set, the degree of correlation between each cluster set, and the quantitative analysis conditions are input into the preset large language model, and the output of the preset large language model for the quantitative analysis conditions is used as the quantitative analysis result.

7. The method according to claim 4, characterized in that, The step of generating the text analysis result based on the deep understanding result and the quantitative analysis result includes: The deep understanding results, the quantitative analysis results, and the text analysis conditions are input into the preset large language model, and the output of the preset large language model for the text analysis conditions is used as the text analysis result.

8. A text data analysis device, characterized in that, The device includes: The extraction module is used to extract entities from the text to be analyzed, and obtain multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. A construction module is used to construct a heat map knowledge graph based on multiple entity features, the importance of each entity feature, the entity relationships between the entity features, and the importance of each entity relationship. The analysis module is used to obtain the deep understanding results of the preset large language model on the text to be analyzed, and to obtain the text analysis results by combining the deep understanding results with the heat map.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.