Target text generation method based on big data

Through the big data-based target text generation method, the graph attention network and graph propagation algorithm are used to solve the problem of flexible organization and automatic update of text generation method in the field of consulting design planning, realizing the adaptive organization and efficient management of documents.

CN120278129AActive Publication Date: 2025-07-08ZHONGTONG INFORMATION SERVICE CO LTD +1

Patent Information

Application Number
CN202510772659.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In the field of professional fields such as consulting design planning, text generation methods lack a deep understanding of content semantics and flexible organization capabilities, resulting in the rigid structure of the generated document, which is difficult to adapt to the specific needs of different types of documents, and requires a lot of manual intervention when updating documents, which is inefficient and prone to errors.

Method used

The target text generation method based on big data is adopted to analyze the document structure through the graph attention network, build document association diagrams and dependency diagrams, use graph propagation algorithm to generate text content, and automatically update the document structure and content when source data or document changes.

Benefits of technology

It realizes adaptive document organization, improves the logic and readability of documents, accurately recognizes complex relationships between documents, reduces manual maintenance costs, and improves the automation level and update efficiency of document management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278129A_ABST
    Figure CN120278129A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and natural language processing, and discloses a big data-based target text generation method, which comprises the following steps of: modeling a document structure into a directed graph, and predicting an optimal document structure by utilizing a graph attention network; constructing a document association graph and analyzing a relationship between documents by adopting a graph neural network; constructing a document dependency graph to form a complete document knowledge network; based on the three graph structures, a graph propagation algorithm is used for text content generation; when source data changes, a graph propagation algorithm is automatically triggered to perform intelligent updating; according to the method, the limitation of a traditional templated structure is broken through, self-adaptive document organization is achieved, the complex incidence relation between documents is accurately recognized, an organic knowledge network is formed, intelligent propagation of document updating is achieved, the dependency relation and version consistency are automatically maintained, and the document generation quality and efficiency of the consultation design planning industry are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and more specifically, it relates to a method for generating target texts based on big data. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, data-driven intelligent text generation technologies have been widely applied in multiple fields. However, in professional fields such as consulting design and planning, the current text generation technologies still face many challenges.

[0003] Traditional text generation methods mainly rely on preset templates and rules, lacking the ability to deeply understand content semantics and flexibly organize them. These methods usually adopt techniques such as keyword matching and sentence pattern replacement, and cannot effectively handle the generation requirements of complex professional documents. Especially when a large amount of data needs to be comprehensively analyzed and a highly logical and professional document needs to be generated, the limitations of traditional methods are particularly obvious.

[0004] In terms of document structure, existing technologies mostly use fixed templates or simple statistical methods to determine the document structure, and cannot be intelligently adjusted according to specific content and usage scenarios. This results in a rigid document structure, making it difficult to adapt to the specific requirements of different types of documents, and affecting the overall quality and practical value of the document.

[0005] In addition, existing technologies also have obvious deficiencies in handling document updates and version management. When the basic data or related documents change, it is difficult for existing systems to automatically identify and update all affected documents, requiring a large amount of manual intervention, with low efficiency and prone to errors.

[0006] Therefore, there is a need for a method for generating target texts that can intelligently understand the document structure, accurately identify the association relationships between documents, and support automatic updates to meet the requirements for generating high-quality documents in professional fields such as consulting design and planning. Summary of the Invention

[0007] The present invention provides a method for generating target texts based on big data, which solves the technical problems in related technologies of lacking the ability to deeply understand content semantics and flexibly organize them, and being unable to effectively handle the generation requirements of complex professional documents.

[0008] The present invention provides a method for generating target texts based on big data, including the following steps: Model the structure of the target document as a directed graph, and use a graph attention network to analyze the document structure graph, calculate the importance weights of each node through an attention algorithm, and predict the optimal document structure organization method; Construct a document association graph based on the predicted document structure, and at the same time use a graph neural network to analyze the semantic similarity and logical association strength between documents to identify the implicit relationships between documents; Construct a document dependency graph based on the document association graph, clarify the dependency relationships and version correspondence relationships between documents, and form a complete document knowledge network structure; Based on the constructed document structure graph, document association graph and dependency graph, use the graph propagation algorithm for text content generation; When the source data or related documents change, automatically trigger the graph propagation algorithm, and intelligently update the structure and content of the affected documents according to the document dependency relationships.

[0009] In a preferred embodiment, the graph attention network adopts a multi-level cascade structure, including 3 to 5 attention layers, with 8 to 16 attention heads used in each layer. The graph attention network includes a multi-head attention layer, a feature transformation layer, and an output layer.

[0010] In a preferred embodiment, the graph neural network includes a graph convolutional layer, a pooling layer, and a classification layer. The graph convolutional layer is responsible for feature propagation and aggregation on the document graph, the pooling layer summarizes the graph-level features, and the classification layer outputs the association strength scores between documents.

[0011] In a preferred embodiment, the document dependency graph includes a dependency relationship recognition module, a version management module, and a consistency check module.

[0012] In a preferred embodiment, the graph propagation algorithm adopts a multi-round iterative propagation method and gradually converges to a stable state through 3 to 10 rounds of iteration. The graph propagation algorithm includes an information aggregation unit, a weight calculation unit, and a content generation unit.

[0013] In a preferred embodiment, the propagation weights are dynamically calculated based on the semantic similarity, citation frequency, and time correlation between documents.

[0014] In a preferred embodiment, the dynamic update adopts an incremental update strategy and only updates the affected parts of the documents.

[0015] In a preferred embodiment, the graph attention network adopts residual connections and layer normalization techniques to improve the training stability and convergence speed of the network.

[0016] In a preferred embodiment, the big data-based target text generation method is applied to the big data text generation scenarios in the consulting design and planning industry, and processes project reports, technical documents, market analyses, and various types of document collections.

[0017] In a preferred embodiment, the big data-based target text generation system, which is used to execute the big data-based target text generation method, includes: A data acquisition module for obtaining industry forefront information from multiple data sources; A structure prediction module for modeling the structure of a target document as a directed graph and analyzing and predicting it using a graph attention network; An association analysis module for constructing a document association graph and a document dependency graph to form a complete document knowledge network structure; A content generation module for generating text content based on graph structure information and a graph propagation algorithm; A dynamic update module for automatically triggering the graph propagation algorithm for intelligent update when the source data or related documents change.

[0018] The beneficial effects of the present invention are as follows: Through the intelligent structure prediction of the graph neural network, the limitations of the traditional templated structure are broken through, and the adaptive document organization based on content features is realized, significantly improving the logic and readability of the generated documents.

[0019] The construction and analysis technology of the multi-document association graph accurately identifies the complex association relationships between documents, forms an organic knowledge network, and improves the overall consistency and knowledge integration effect of the document set.

[0020] The application of the graph propagation algorithm realizes the intelligent propagation of document updates, automatically maintains the dependency relationships and version consistency between documents, greatly reduces the manual maintenance cost, and improves the automated management level of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of the method for generating a target text based on big data according to the present invention; Figure 2 is a bar chart of the comparison of the rationality scores of the document structures according to the present invention; Figure 3 is a line chart of the change of the system processing efficiency over time according to the present invention; Figure 4 is a radar chart of the comparison of the performances of different algorithms according to the present invention; Figure 5 is a pie chart of the distribution of the system resource consumption according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and the functions and arrangements of the elements discussed can be changed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0023] In at least one embodiment of the present invention, a method for generating target text based on big data is disclosed. As Figure 1 shown, it includes the following steps: Step 1: Model the structure of the target document as a directed graph, and use a graph attention network to analyze the document structure graph. Calculate the importance weights of each node through an attention algorithm, and predict the optimal document structure organization method; Specifically, it includes the following steps: Step 1.1: Model the structure of the target document as a directed graph: ; Among them, represents the document structure graph; represents the set of each part node of the document, including hierarchical structures such as chapters, paragraphs, and sub - paragraphs; represents the set of hierarchical organization relationship edges between parts.

[0024] Each node represents a structural unit in the document, and has an attribute set including node type (such as chapter, section, paragraph, etc.), node hierarchical depth, node title, node content, etc.; Each edge represents the hierarchical relationship from node to node , indicating that is the parent unit of , has directionality, and can be assigned a weight value to represent the association strength.

[0025] This directed graph structure can completely express the organizational framework of the document and the subordinate relationship between each part.

[0026] Step 1.2: Use a graph attention network to analyze the document structure graph, calculate the importance weights of each node through an attention algorithm, and predict the optimal document structure organization method; The calculation formula of the graph attention network is: ; Among them, represents the updated feature vector of node in the th layer after being calculated by the graph attention network, that is, the output feature of the current layer; represents the feature vector of the neighbor node in the th layer, which is used as the input feature of the current calculation; represents the attention weight coefficient, which is used to measure the influence of the neighbor node on the current node The degree of importance, where a higher weight indicates that the information provided by the neighboring node is more important; Denotes a node The set of neighboring nodes; Denotes the Trainable weight matrix of the l-th layer, used for linear transformation of node features; Denotes the All neighboring nodes of node For weighted summation operation; Denotes a non-linear activation function, such as ReLU, tanh, etc., used to introduce non-linear transformation to enhance the model's expressive power.

[0027] The graph attention network includes a multi-head attention layer, a feature transformation layer, and an output layer.

[0028] The multi-head attention layer captures different types of document structure relationships by parallel computing multiple attention heads; The feature transformation layer performs non-linear transformation on node features to enhance the expressive power; The output layer generates the final document structure prediction result.

[0029] In the scenario of consulting design planning, such as generating a project feasibility report, the graph attention network can identify the strong correlation between the "Market Analysis" section and the "Risk Assessment" section, thereby adjusting the document structure to make the relevant content more logically compact.

[0030] In some embodiments, the graph attention network adopts a multi-level cascade structure, including 3 to 5 attention layers, and each layer uses 8 to 16 attention heads.

[0031] For example, for a complex urban master plan document, a 5-layer attention network with 16 attention heads per layer can be adopted to fully capture the complex correlation between different planning elements.

[0032] In addition, residual connection and layer normalization techniques are adopted to improve the training stability and convergence speed of the network.

[0033] Step 2: Based on the predicted document structure, construct a document association graph, and at the same time use a graph neural network to analyze the semantic similarity and logical association strength between documents, and identify the implicit relationships between documents; Specifically, it includes the following steps: Step 2.1: Construct a document association graph: ; Among them, Denotes the entire document association graph structure, which is a mathematical representation form used to model the relationships between documents in the document set; Denotes the set of nodes of each document in the document set; The edge set representing the logical association relationships between documents.

[0034] Step 2.2: Use a graph neural network to analyze the semantic similarity and logical association strength between documents, identify the implicit relationships between documents, generate a document association weight matrix, and provide basic data for subsequent structure prediction.

[0035] In addition, the graph neural network includes a graph convolutional layer, a pooling layer, and a classification layer.

[0036] The graph convolutional layer is responsible for feature propagation and aggregation on the document graph to learn the representation vectors of documents; The pooling layer summarizes the graph-level features; The classification layer outputs the association strength scores between documents.

[0037] In a specific application, for example, when processing a document set of a large infrastructure project, the graph neural network can automatically identify a strong dependency relationship between the "Environmental Impact Assessment Report" and the "Construction Plan Document" because the environmental assessment results directly affect the formulation of the construction plan.

[0038] Step 3: Based on the document association graph, construct a document dependency graph, clarify the dependency relationships and version correspondence relationships between documents, and form a complete document knowledge network structure; Construct a document dependency graph: ; Among them, represents the document dependency graph; represents the document node set, and each node represents an independent document in the system, containing basic attribute information such as document ID, document type, creation time, etc.; represents the edge set of the dependency relationships between documents, and each edge defines the dependency type, dependency strength, and dependency direction between two documents; represents the version information of documents, including the version history record, version number, update timestamp, change content summary, etc. of each document, and is used to track the evolution process of documents.

[0039] Therefore, this model clearly represents the dependency relationships and version correspondence relationships between documents through a triple structure, providing a structured basis for the dynamic update of the document set.

[0040] The document dependency graph includes a dependency relationship recognition module, a version management module, and a consistency check module.

[0041] The dependency relationship recognition module determines the dependency strength between documents by analyzing the document content and citation relationships; The version management module maintains the version history and change records of each document; The consistency check module verifies the logical consistency of the document set.

[0042] For example, in an urban planning project, when the "Traffic Planning Scheme" document is updated, the system can automatically identify that relevant documents such as "Land Use Planning" and "Environmental Protection Planning" need to be adjusted accordingly and establish a clear dependency relationship chain.

[0043] Step 4: Based on the constructed document structure diagram, document association diagram, and dependency diagram, use the graph propagation algorithm for text content generation; Specifically, it includes the following steps: Step 4.1: Based on the constructed document association diagram and dependency diagram, use the graph propagation algorithm for text content generation.

[0044] Among them, the graph propagation algorithm realizes information propagation through the following formula: ; Among them, represents the total propagated update information received by node , that is, the update content that the document needs to perform according to the changes in relevant documents; represents the summation operation on all neighbor nodes of node , that is, considering all other documents related to the current document; represents the set of all neighbor nodes adjacent to node , that is, all other documents directly related to the current document i; represents the propagation weight between node and neighbor node . The value range is from 0 to 1. The greater the weight, the greater the impact of the change in document on document . This weight is dynamically calculated based on factors such as the strength of the dependency relationship between documents, content relevance, and time relevance; represents the change information of neighbor node , including change elements such as content updates, structural adjustments, and data corrections of document , which is the input signal of the propagation algorithm.

[0045] Step 4.2: Combine the propagated information with the original big data content to generate the target text content that meets the document structure requirements, thereby ensuring the relevance and consistency of the generated content.

[0046] The graph propagation algorithm includes an information aggregation unit, a weight calculation unit, and a content generation unit.

[0047] The information aggregation unit collects the change information from relevant document nodes; The weight calculation unit dynamically adjusts the propagation weight according to the association strength between documents; The content generation unit generates new text content based on the aggregated information.

[0048] In practical applications, for example, when generating an enterprise strategic planning report, when the data in the "Market Environment Analysis" section is updated, the graph propagation algorithm can spread this changed information to relevant chapters such as "Development Strategy Formulation" and "Risk Response Measures" to ensure the consistency and timeliness of the content of the entire report.

[0049] In some embodiments, the graph propagation algorithm adopts a multi-round iterative propagation method and gradually converges to a stable state through 3 - 10 rounds of iteration.

[0050] Propagation weight It is dynamically calculated according to the semantic similarity, citation frequency, and time correlation between documents.

[0051] For example, when processing a technical document set of a new energy project, when the key technical parameters in the "Technical Feasibility Analysis" document change, the algorithm first spreads the changed information to the directly related "Equipment Selection Plan" document, then affects the "Investment Estimation Report" in the second round of propagation, and finally updates the "Project Risk Assessment" in the third round of propagation, realizing the orderly spread of information and the intelligent update of content.

[0052] Step 5, when the source data or related documents change, automatically trigger the graph propagation algorithm, and intelligently update the structure and content of the affected documents according to the document dependency relationship.

[0053] In addition, through version identification and dependency tracking, maintain the version consistency and logical integrity of the entire document set to achieve the automated management of the document set.

[0054] In some embodiments, the dynamic update adopts an incremental update strategy, only updating the affected parts of the documents instead of regenerating them in full, thereby improving the update efficiency.

[0055] For example, when the population prediction data of a traffic planning project changes, the system only updates the traffic demand analysis chapter related to the population data, while keeping the content of other unrelated chapters unchanged, significantly reducing the consumption of computing resources and the update time.

[0056] Application examples of this embodiment: Application scenarios: A large-scale urban comprehensive development project needs to generate 120 related documents including a master plan report, environmental impact assessment, traffic planning scheme, economic feasibility analysis, etc. There are complex logical associations and data dependencies between these documents, and it is difficult for traditional methods to ensure the consistency and timeliness between documents.

[0057] Implementation process example: First, the system models the master plan report as a document structure diagram containing 15 main chapters, with each chapter containing 3 - 8 sub - paragraph nodes.

[0058] The graph attention network identifies a strong association weight of 0.85 between "land use planning" and "transportation planning" and an association weight of 0.72 between "environmental protection" and "industrial layout" by analyzing the semantic associations between chapters.

[0059] Based on these association analyses, the system automatically adjusts the document structure and tightly organizes the strongly related chapters logically.

[0060] The system constructs a multi - document association graph containing 120 document nodes.

[0061] The graph neural network identifies a two - way strong dependence relationship with an association strength of 0.91 between the "environmental impact assessment report" and the "master plan report"; There is a one - way dependence relationship with an association strength of 0.78 between the "transportation planning scheme" and the "land use planning document".

[0062] Based on this, the system establishes a complete document dependence graph, clarifying 896 dependence relationships among 120 documents.

[0063] When the population prediction data of the project is adjusted from the original 800,000 people to 950,000 people, the graph propagation algorithm is automatically activated.

[0064] In the first round of propagation, the changed information is propagated to the directly related "housing demand analysis" and "public facility allocation" documents; In the second round of propagation, the influence extends to the "transportation demand prediction" and "infrastructure scale" documents; In the third round of propagation, it finally affects the "investment estimate" and "implementation schedule" documents.

[0065] The entire propagation process involves content updates of 32 documents, and the system automatically completes the regeneration of relevant chapters.

[0066] Technical effect verification: Through the application of this implementation method, the document management effect of the urban comprehensive development project is significantly improved.

[0067] The rationality score of the document structure is increased from the original 6.2 points to 8.5 points, with an increase rate of 37%.

[0068] The satisfaction degree of the expert review team on the logical consistency of the document is increased from the original 65% to 92%, with an increase rate of 42%.

[0069] In terms of system efficiency, the document update time has been shortened from the original average of 72 hours to 18 hours, with a 75% improvement in efficiency.

[0070] When the core data changes, the automatic update coverage rate of relevant documents reaches 94%, while the update coverage rate of traditional methods is only 56%.

[0071] The manual maintenance cost is reduced by 68%, and the overall workload of project document management is reduced by approximately 60%.

[0072] Figures 2 to 5 They are respectively the comparison of the rationality score of the document structure; the change of system processing efficiency over time; the comparison of the performance of different algorithms; the distribution of system resource consumption.

[0073] These data fully verify the significant advantages of the technical solution of this application in improving document quality and management efficiency, providing effective technical support for the intelligent management of documents in large and complex projects.

[0074] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A method for generating a target text based on big data, characterized in that, It includes the following steps: Model the structure of the target document as a directed graph, and use the graph attention network to analyze the document structure graph. Calculate the importance weights of each node through the attention algorithm to predict the optimal document structure organization method; Based on the predicted document structure, construct a document association graph. At the same time, use the graph neural network to analyze the semantic similarity and logical association strength between documents to identify the implicit relationships between documents; Based on the document association graph, construct a document dependency graph to clarify the dependency relationships and version correspondence relationships between documents, and form a complete document knowledge network structure; Based on the constructed document structure graph, document association graph, and dependency graph, use the graph propagation algorithm for text content generation; When the source data or related documents change, automatically trigger the graph propagation algorithm to intelligently update the structure and content of the affected documents according to the document dependency relationships.

2. The method for generating a target text based on big data according to claim 1, wherein, The graph attention network adopts a multi-level cascade structure, including 3 to 5 attention layers, with 8 to 16 attention heads used in each layer. The graph attention network includes a multi-head attention layer, a feature transformation layer, and an output layer.

3. The method for generating a target text based on big data according to claim 1, wherein The graph neural network includes a graph convolutional layer, a pooling layer, and a classification layer. The graph convolutional layer is responsible for feature propagation and aggregation on the document graph, the pooling layer summarizes the graph-level features, and the classification layer outputs the association strength scores between documents.

4. The method for generating a target text based on big data according to claim 1, wherein The document dependency graph includes a dependency relationship recognition module, a version management module, and a consistency check module.

5. The method for generating a target text based on big data according to claim 1, wherein, The graph propagation algorithm adopts a multi-round iterative propagation method and gradually converges to a stable state through 3 to 10 rounds of iteration. The graph propagation algorithm includes an information aggregation unit, a weight calculation unit, and a content generation unit.

6. The method for generating a target text based on big data according to claim 5, wherein, The propagation weights are dynamically calculated based on the semantic similarity, citation frequency, and temporal correlation between documents.

7. The method for generating a target text based on big data according to claim 1, wherein The dynamic update adopts an incremental update strategy and only updates the affected parts of the documents.

8. The method for generating a target text based on big data according to claim 2, wherein, The graph attention network adopts residual connection and layer normalization techniques to improve the training stability and convergence speed of the network.

9. The method for generating a target text based on big data according to claim 1, wherein It is applied to the big data text generation scenario in the consulting design and planning industry to process project reports, technical documents, market analysis, and various types of document collections.

10. A target text generation system based on big data, which is used to execute the big-data-based target text generation method described in any one of claims 1-9, and is characterized in that It includes: A data acquisition module for obtaining industry forefront information from multiple data sources; A structure prediction module for modeling the structure of the target document as a directed graph and using the graph attention network for analysis and prediction; An association analysis module for constructing a document association graph and a document dependency graph to form a complete document knowledge network structure; A content generation module for generating text content based on the graph structure information and the graph propagation algorithm; A dynamic update module for automatically triggering the graph propagation algorithm for intelligent update when the source data or related documents change.

Citation Information

Patent Citations

  • Social media rumor detection method based on network information propagation graph modeling

    CN112199608A

  • Data dependency analysis method, device and system and medium

    CN116401307A

  • Multi-document abstract generation method and device, equipment, storage medium and program product

    CN118626636A

  • Version mismatch delay and update for a distributed system

    US20120109914A1

  • System and method for maintaining links and revisions

    US20240028823A1

Cited By

  • System and method for AI to automatically generate technology commercialization feasibility report

    CN121072506A