A document generation system and method based on a big data model

The document generation system, which integrates and optimizes big data models, solves the problems of accuracy, logical coherence, and personalization in existing document generation technologies, and achieves high-quality, personalized document generation.

CN120124596BActive Publication Date: 2025-11-11CCID CONSULTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510082013.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-11-11
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Existing document generation systems are inadequate in terms of content accuracy, logical coherence, content quality, and personalization, making it difficult to generate documents that meet specific user needs.

Method used

A document generation system based on a big data model is adopted. By integrating, verifying, labeling and extracting features from data sources, a document generation model is constructed, and the model is optimized and evaluated. Knowledge graph and reinforcement learning techniques are combined to polish and customize the documents.

Benefits of technology

It improves the accuracy and quality of document generation, enables personalized document content customization to meet specific user needs, and ensures the logical coherence and consistency of documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124596B_ABST
    Figure CN120124596B_ABST
Patent Text Reader

Abstract

This invention provides a document generation system and method based on a big data model, belonging to the field of question-and-answer management technology. The method includes: data source integration; data source verification; data annotation and feature extraction; construction of a document generation model; model optimization and evaluation; generation of a draft document; document polishing and improvement; and in-depth understanding of user needs and customized modification of the document. The document generation system and method based on a big data model provided by this invention improves the quality, accuracy, and logical coherence of document content, and enables personalized customization of document content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of question-and-answer management technology, and in particular to a document generation system and method based on a big data model. Background Technology

[0002] Automatic text generation is one of the core problems in the field of natural language processing. Generating text automatically based on the author's ideas can greatly reduce the author's workload; however, generating text on specific aspects of the author's ideas also presents certain challenges. The following are the problems with existing document generation systems:

[0003] 1. Regarding the accuracy of content, sometimes inaccurate information is generated. For example, when generating documents related to historical events, the events, locations, or relationships between people involved may be confused. This is because the data sources may be inaccurate or outdated, and there is a lack of sufficient verification mechanisms.

[0004] The generated documents may have issues with logical coherence. For example, when arguing a point, there may be a lack of reasonable deductive relationships between the evidence and the argument, or incorrect connections may appear when explaining the causal relationships of events. This is because the document generation system is not perfect in constructing its logical framework and cannot fully simulate human logical thinking abilities.

[0005] 2. In terms of content quality, the generated documents are often rather simplistic in their expression and lack literary flair. The language style may be rather mechanical, mostly using common sentence structures and vocabulary, making it difficult to employ rich rhetorical devices and diverse sentence structures to increase the article's appeal like a human author.

[0006] For some complex topics, document generation systems often only produce superficial overviews, failing to delve into the deeper meaning of the subject. For example, in academic paper generation, while they can cobble together papers that meet format requirements, they are clearly lacking in in-depth analysis of research questions and unique insights.

[0007] 3. Regarding personalization, when users have very specific personalized needs, such as industry-specific terminology or a company's internal style, the document generation system may not be able to accurately meet them. It generates documents based more on general data and is less adaptable to special requirements.

[0008] The system cannot fully and accurately understand the user's intent. The instructions entered by the user may be somewhat ambiguous. The document generation system may not be able to further clarify the user's intent through questioning or other means like a human can, resulting in a significant gap between the generated document and the user's expectations.

[0009] Therefore, there is an urgent need in this field for a technical solution that generates documents with high accuracy and quality, and can achieve personalized customization.

[0010] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0011] The purpose of this invention is to provide a document generation system and method with high content accuracy and quality, and capable of personalized customization.

[0012] To achieve the above objectives, the present invention provides the following solution:

[0013] A document generation system based on a big data model includes:

[0014] The parameter settings module is used to configure the knowledge base, model selection, and role settings;

[0015] The report list view module allows you to view the document title, tags, and creation time; it also allows you to search for documents by their creation time and document tags.

[0016] The feedback module is used to input feedback on the system;

[0017] The system login module is used by users to log in by entering their username and password;

[0018] The model management module is used to enable and disable model deployment;

[0019] The knowledge base management module is used to view, close, and delete knowledge bases. It can also search for documents in the knowledge base by document title, file name, document tags, and creation time.

[0020] The conversation management module includes a user management unit and a sensitive word management unit;

[0021] The user management unit is used to search for users based on user ID, user name, and creation time, and can also view users' historical conversations;

[0022] The sensitive word unit is used to search for sensitive words based on sensitive words, tags, whether it is enabled, and creation time, as well as to modify and delete sensitive words; it can also add new sensitive words.

[0023] The feedback management module is used to manage feedback.

[0024] The system management module is used to manage multiple functional modules.

[0025] Optionally, the system management module includes:

[0026] The tenant management unit includes a tenant list sub-unit and a tenant package sub-unit.

[0027] The tenant list sub-unit contains a tenant list that can be searched by tenant name, contact person, contact phone number, tenant status, and creation time. It also allows modification of tenant information and deletion of tenants.

[0028] The tenant package sub-unit is used to search based on package name, status, and modification time. The tenant package sub-unit contains a package list.

[0029] The user management module is used to manage users. It includes an intelligent customer service portal for managing and allocating users according to the head office, branches, and various departments within the company. The module also features a search function that allows users to search by name, phone number, status, and creation time. Furthermore, it allows modification of user information.

[0030] The role management module is used to manage different roles. It can search for roles by role name, role identifier, status, and creation time. Role names include test accounts, management roles, and super administrators. It can also modify role information, change permissions, and delete roles.

[0031] The menu management module is used to manage the menu list and can search for menus by name and status. The menu names include: Boot Development Documentation, Cloud Development Documentation, Model Management, Knowledge Base Management, Dialogue Management, Feedback Management, System Management, Payment Management, Report Management, Workflow, Product Center, Order Center, Marketing Center, and Official Account Management. The module also allows modification of menus.

[0032] The department management module is used to manage the various departments of the head office and branches; it allows users to modify, add, and delete departments; and it also allows users to search by department name and status.

[0033] The job management module is used to modify, add, and delete jobs. Job information includes: job number, job code, job name, job sorting, and status. It can also search for jobs by job code, job name, and status.

[0034] The dictionary management module allows users to modify, add, and delete dictionaries. Dictionary information includes: dictionary number, dictionary name, dictionary type, status, and creation time. Users can also search for dictionaries by name, dictionary type, status, and creation time.

[0035] The notification and announcement module is used to modify, add, and delete notifications and announcements. Announcement information includes: announcement title, announcement type, status, and creation time. It also allows searching by announcement title and status.

[0036] The operation log module is used to manage operation logs; log information includes log number, operation module, operation name, operation type, operator, operation result, operation date, and execution duration; it can be searched by system module, operator, type, status, and operation time;

[0037] The login log module is used to manage login information. Login information includes log type, username, login address, userAgent, result, and login date. It allows searching by login address, username, status, and login time.

[0038] The application management module is used to manage application clients; client information includes: client ID, client key, application name, application icon, status, access token validity period, refresh token validity period, authorization type, and creation time; it can search by application name and status;

[0039] The token management module is used to manage tokens; token information includes: access token, refresh token, user ID, user type, creation time, and expiration time; it can search by user ID, client ID, and user type;

[0040] The SMS management module is used to manage SMS messages. SMS information includes: SMS signature, channel code, activation status, SMS API account, SMS API key, and creation time. It can also be searched based on SMS signature, activation status, and creation time.

[0041] A document generation method based on a big data model includes:

[0042] Data source integration;

[0043] Data source verification;

[0044] Data annotation and feature extraction;

[0045] Build a document generation model;

[0046] Model optimization and evaluation;

[0047] Generate a first draft of the document;

[0048] Document polishing and improvement;

[0049] Gain a deep understanding of user needs and customize the documentation accordingly.

[0050] Optionally, the data source integration includes:

[0051] Use Apache Spark's distributed file system interface to collect document-related data from multiple data sources;

[0052] Identify the types of multiple data sources, including relational databases, non-relational databases, files in the local file system, and files in network storage;

[0053] For document-related data from different data sources, analyze their format in detail;

[0054] Configure the Apache Spark environment and ensure the Spark cluster is configured correctly; ensure that the necessary dependency libraries are imported into the Spark project, especially the driver libraries related to the data source to be connected.

[0055] Data is collected using Spark's distributed file system interface, specifically...

[0056] In Spark, the first step is to create a SparkSession object, which is the entry point for interacting with Spark.

[0057] To connect to different data sources, use Spark's JDBC interface for relational databases and the Spark-MongoDB connector for non-relational databases.

[0058] Spark's DataFrame operations can be used to integrate data from different data sources.

[0059] The collected data is initially cleaned to remove noisy data and incomplete data records.

[0060] Optionally, the data source verification includes:

[0061] Data sources are categorized into primary and secondary sources;

[0062] For each data point, label its source type;

[0063] For different sources, an initial credibility weight is assigned based on their historical accuracy and authority; this is represented by a matrix C, where C... ij This represents the confidence weight of the i-th data point originating from the j-th type of source;

[0064] When multiple sources provide the same or similar data points, calculate the degree of consistency of their data.

[0065] Suppose that a data point x is provided by n sources, and the values ​​of these sources are x1, x2, ..., xn. n ; Calculate the consistency index in

[0066] If I is greater than a set threshold, the data is considered reliable in terms of consistency.

[0067] Update the source credibility weight based on the data consistency check results;

[0068] If the data consistency is good, the credibility weight of the source providing the data point increases. Where n is the number of sources providing the data point;

[0069] If data consistency is poor, the credibility weight of the source providing that data point will be reduced.

[0070] As new data is continuously input, repeat the above steps to continuously update the credibility matrix and verify the accuracy of the data;

[0071] For sources with a credibility weight below a certain threshold, their data should be marked as requiring further manual review or discarded directly.

[0072] The semantic and logical relationships of the data are analyzed, and the causal and correlation relationships between the data are analyzed by establishing a logical relationship model, so as to further verify the data;

[0073] Let the logical relationship be represented by the function L(x, y), where x and y are the relevant data points. If the value of L(x, y) exceeds a reasonable range, the relevant data will be re-examined or marked as suspicious data.

[0074] Optionally, the data annotation and feature extraction include:

[0075] The cleaned data is labeled. If it is used to generate a specific type of document, it can be labeled manually or automatically according to the document's theme and style.

[0076] Extract features from the data, including word frequency features, part-of-speech features, and semantic features; integrate with Spark using a natural language processing library; use Spark's mapPartitions operation to call NLTK functions to compute features within each data partition, and then merge the results.

[0077] Optionally, the construction of the document generation model includes:

[0078] Using a neural network language model, in the Spark environment, the model is trained using distributed computing. Large-scale text data is divided into multiple small batches, and the model is trained in parallel on multiple Spark nodes.

[0079] We employ transfer learning to fine-tune a pre-trained model on a large dataset to adapt it to a specific document generation task.

[0080] Optionally, the model optimization and evaluation includes:

[0081] Optimize the model on a Spark cluster using stochastic gradient descent and its variants for parameter updates; Spark's distributed computing capabilities can accelerate the optimization process by computing gradients and updating parameters in parallel across multiple nodes.

[0082] Multiple evaluation metrics are used to evaluate the model's performance. A cross-validation evaluation method is implemented on Spark, which divides the data into training, validation, and test sets. The model is evaluated on different subsets and its hyperparameters are tuned.

[0083] Optionally, the generated initial draft document includes:

[0084] Using a trained model, a draft document is generated based on the given input. In the Spark environment, the input data is distributed to multiple nodes, each node generates a portion of the document content based on the model, and then these contents are merged. If a long document is to be generated, the document is divided into multiple paragraphs, and each node is responsible for generating one paragraph.

[0085] For the generated draft, preliminary formatting adjustments can be made by writing custom rules or using natural language processing tools.

[0086] Optionally, the document polishing and improvement includes:

[0087] Additional language processing techniques can be used to polish the initial draft of the document. In Spark, external grammar checking tools can be called or predefined word substitution rules can be used to process the generated document.

[0088] The document was further improved based on specific needs, and the tone of the document was adjusted and citations or case studies were added, taking into account domain knowledge and user feedback.

[0089] The process of gaining a deep understanding of user needs and customizing documents includes:

[0090] First, the user-input commands are analyzed lexically and syntactically; the commands are broken down into basic semantic units, recognizing not only individual words, but also phrases and sentence structures.

[0091] Semantic role labeling technology is used to determine the role of each semantic unit in the entire instruction; a semantic network is constructed to connect the parsed semantic units according to their role relationships;

[0092] Construct a large-scale knowledge graph containing professional knowledge from various industries, internal culture and style of different companies, etc.; the nodes of this knowledge graph represent different concepts, and the edges represent the relationships between concepts.

[0093] The semantic network obtained by the intent parsing module is matched with the knowledge graph; the most matching knowledge fragment is found by calculating the similarity between nodes in the semantic network and nodes in the knowledge graph, as well as the similarity between edges.

[0094] Based on the results of knowledge graph matching, the templates and rules for document generation are dynamically adjusted; if industry-specific professional terms are matched, the vocabulary selection and grammatical structure are adjusted to conform to the expression habits of that industry.

[0095] Reinforcement learning techniques are used to adjust the dynamic adaptation strategy based on how well the generated document matches the user's expectations;

[0096] After the document is generated, a user feedback mechanism is provided; users can rate the generated document or provide specific suggestions for modification.

[0097] The user feedback is used as input to re-enter the intent parsing module and adjust the understanding of the user's intent. Based on the new intent understanding, the document is regenerated through the knowledge graph matching and dynamic adaptation modules until the user is satisfied.

[0098] The formula is expressed as follows:

[0099] Let the user input command be I, and the semantic network obtained by the intent parsing module be SN = f1(I), where f1 is the intent understanding function;

[0100] Let the knowledge graph be KG, and the matching degree function be m(SN, KG) = g(SN, KG), where g is a function for calculating the matching degree between the semantic network SN and the knowledge graph KG;

[0101] Let the generated document template be T, and the dynamically adapted template be T' = h(T, m(SN, KG)), where h is a function that adjusts the template based on the matching degree;

[0102] Let the document generation function be D(T',I)=d(T',I), and the final generated document be Doc=d(T',I);

[0103] In the feedback loop, let the user feedback be F, and the updated intent be I' = k(F, I), where k is a function that updates the intent based on the feedback. Then, the document is generated again according to the above process.

[0104] The algorithm flow is as follows:

[0105] Step 1: Intent Analysis

[0106] Receive user input command I;

[0107] Lexical and syntactic analysis is performed on I to obtain basic semantic units;

[0108] Perform semantic role labeling and construct a semantic network SN = f1(I);

[0109] Step 2: Knowledge graph matching;

[0110] Build or load a knowledge graph (KG);

[0111] Calculate the matching degree m(SN, KG) = g(SN, KG);

[0112] Step 3: Dynamic Adaptation:

[0113] Based on the matching degree m(SN, KG), adjust the document generation template T to obtain T' = h(T, m(SN, KG));

[0114] Step 4: Document Generation

[0115] Based on the dynamically adapted template T' and the user input command I, generate document Doc = d(T', I);

[0116] Step 5: Feedback Loop

[0117] If the user provides feedback F, update the intent I' = k(F, I);

[0118] Repeat step one until the user is satisfied.

[0119] Compared with the prior art, the present invention has the following beneficial effects:

[0120] This invention provides a document generation system and method based on a big data model, capable of generating the following documents:

[0121] Beneficial effects:

[0122] 1. Data source integration:

[0123] Data comprehensiveness:

[0124] Integrating multiple data sources can make the data more comprehensive. For example, when building a document generation model, data from different sources can cover more topics, domains, and information types, thus providing the model with a broader knowledge base and improving the richness of document generation.

[0125] Reduce data bias:

[0126] By integrating multiple data sources, the potential biases of relying on a single source can be avoided. If only one data source is used, the model may produce biased results due to the specific biases of that source (such as the views of a particular region or group). Integration makes the data more representative and improves the accuracy of the model.

[0127] 2. Data source verification:

[0128] Data reliability:

[0129] Verifying data sources ensures data reliability. Reliable data is fundamental when building document generation models. Using unverified data can lead to incorrect or misleading content being generated by the model. Validation eliminates false, inaccurate, or outdated data, thereby increasing the credibility of the model's output.

[0130] Accuracy Guarantee:

[0131] It helps ensure data accuracy. In today's data environment, data is diverse and complex. Verifying data sources can ensure data accuracy and prevent model errors caused by using inaccurate data sources.

[0132] 3. Data labeling and feature extraction

[0133] Feature recognition:

[0134] The ability to accurately identify features in data is crucial for building document generation models. For example, in natural language processing, data annotation and feature extraction can identify the part-of-speech, semantic features, and other characteristics of words. Models can then use these features to better understand the input content and generate reasonable documents.

[0135] Improve model targeting:

[0136] This allows the model to learn and generate more effectively. By extracting key features relevant to document generation, the model can focus on this important information rather than being distracted by irrelevant information, thereby improving the quality and efficiency of document generation.

[0137] 4. Construct a document generation model

[0138] Automated document generation:

[0139] It can automate document generation. For large-scale document needs, such as news reports and business documents, the model can quickly generate drafts based on input requirements, saving manpower and time costs.

[0140] Consistency and Standardization:

[0141] Ensure consistency and standardization in document generation. The model generates documents according to pre-defined rules and algorithms, ensuring that the generated documents maintain consistency in format, style, and content structure, conforming to specific standards and requirements.

[0142] 5. Model Optimization and Evaluation

[0143] Performance improvements:

[0144] Through optimization and evaluation, the performance of a model can be continuously improved. For example, optimization algorithms can increase the model's running speed, and evaluation metrics can identify the model's shortcomings, allowing for targeted improvements that make the model more accurate and efficient in document generation.

[0145] Enhanced adaptability:

[0146] This enables the model to better adapt to different tasks and data environments. As data changes and task requirements evolve, optimization and evaluation can adjust the model's parameters and structure to maintain good adaptability and effectively generate documents under various conditions.

[0147] 6. Generate the first draft of the document

[0148] Quickly start creating:

[0149] It provides a foundation for quickly starting the creative process. Whether writing reports, articles, or other documents, a first draft can serve as a starting point for further refinement, saving time compared to starting from scratch and allowing creators to move more quickly into the revision and improvement stage.

[0150] Content framework construction:

[0151] It helps in building the content framework of the document. The first draft often contains the basic structure and main content points of the document, providing framework support for subsequent polishing, supplementation and customized modifications.

[0152] 7. Document polishing and improvement

[0153] Improve document quality:

[0154] It can significantly improve document quality. Polishing can improve the language expression of a document, making it more fluent, accurate, and professional; refining can supplement the content missing in the first draft, adjust the logical structure, and make the document more complete and rigorous.

[0155] Style matching:

[0156] It allows you to tailor the style of a document to a specific audience or purpose. For example, you can transform a scientific document from an academic style to a popular science style to better meet the needs of different readers.

[0157] 8. Thoroughly understand user needs and customize document modifications.

[0158] Meeting personalized needs

[0159] Ensure that documents meet users' personalized needs. Different users have different requirements for documents. By deeply understanding these needs and customizing them, documents can be tailored to users' specific purposes in terms of content, structure, and style, thereby improving user satisfaction.

[0160] Enhanced targeting:

[0161] Enhance the relevance of documents. Customized documents can better convey information and achieve the desired effect for specific audiences, tasks, or scenarios, such as business proposals for a specific client or research reports for an academic conference. Attached Figure Description

[0162] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0163] Figure 1 This is a schematic diagram of the homepage of the document generation system provided in an embodiment of the present invention.

[0164] Figure 2 This is a schematic diagram of the parameter setting module interface provided in an embodiment of the present invention.

[0165] Figure 3 This is a schematic diagram of the report list viewing module interface provided in an embodiment of the present invention.

[0166] Figure 4 This is a schematic diagram of the system login interface provided in an embodiment of the present invention.

[0167] Figure 5 This is a schematic diagram of the system backend homepage provided in an embodiment of the present invention.

[0168] Figure 6 This is a diagram of the model management module interface provided in an embodiment of the present invention.

[0169] Figure 7 Interface diagram of the knowledge base management module provided in the embodiments of the present invention.

[0170] Figure 8 This is a diagram of the dialogue management module interface provided in an embodiment of the present invention.

[0171] Figure 9 This is a diagram of the feedback module interface provided in an embodiment of the present invention.

[0172] Figure 10 This is a diagram of the system management module interface provided in an embodiment of the present invention. Detailed Implementation

[0173] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0174] The purpose of this invention is to provide a system and method for generating documents with high accuracy and quality, and capable of personalized customization.

[0175] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0176] Example 1:

[0177] This embodiment provides a document generation system based on a big data model, such as... Figure 1-10 As shown, it includes:

[0178] The parameter settings module is used to configure the knowledge base, model selection, and role settings;

[0179] The report list view module allows you to view the document title, tags, and creation time; it also allows you to search for documents by their creation time and document tags.

[0180] The feedback module is used to input feedback on the system;

[0181] The system login module is used by users to log in by entering their username and password;

[0182] The model management module is used to enable and disable model deployment;

[0183] The knowledge base management module is used to view, close, and delete knowledge bases. It can also search for documents in the knowledge base by document title, file name, document tags, and creation time.

[0184] The conversation management module includes a user management unit and a sensitive word management unit;

[0185] The user management unit is used to search for users based on user ID, user name, and creation time, and can also view users' historical conversations;

[0186] The sensitive word unit is used to search for sensitive words based on sensitive words, tags, whether it is enabled, and creation time, as well as to modify and delete sensitive words; it can also add new sensitive words.

[0187] The feedback management module is used to manage feedback.

[0188] The system management module is used to manage multiple functional modules.

[0189] In one embodiment, the system management module includes:

[0190] The tenant management unit includes a tenant list sub-unit and a tenant package sub-unit.

[0191] The tenant list sub-unit contains a tenant list that can be searched by tenant name, contact person, contact phone number, tenant status, and creation time. It also allows modification of tenant information and deletion of tenants.

[0192] The tenant package sub-unit is used to search based on package name, status, and modification time. The tenant package sub-unit contains a package list.

[0193] The user management module is used to manage users. It includes an intelligent customer service portal for managing and allocating users according to the head office, branches, and various departments within the company. The module also features a search function that allows users to search by name, phone number, status, and creation time. Furthermore, it allows modification of user information.

[0194] The role management module is used to manage different roles. It can search for roles by role name, role identifier, status, and creation time. Role names include test accounts, management roles, and super administrators. It can also modify role information, change permissions, and delete roles.

[0195] The menu management module is used to manage the menu list and can search for menus by name and status. The menu names include: Boot Development Documentation, Cloud Development Documentation, Model Management, Knowledge Base Management, Dialogue Management, Feedback Management, System Management, Payment Management, Report Management, Workflow, Product Center, Order Center, Marketing Center, and Official Account Management. The module also allows modification of menus.

[0196] The department management module is used to manage the various departments of the head office and branches; it allows users to modify, add, and delete departments; and it also allows users to search by department name and status.

[0197] The job management module is used to modify, add, and delete jobs. Job information includes: job number, job code, job name, job sorting, and status. It can also search for jobs by job code, job name, and status.

[0198] The dictionary management module allows users to modify, add, and delete dictionaries. Dictionary information includes: dictionary number, dictionary name, dictionary type, status, and creation time. Users can also search for dictionaries by name, dictionary type, status, and creation time.

[0199] The notification and announcement module is used to modify, add, and delete notifications and announcements. Announcement information includes: announcement title, announcement type, status, and creation time. It also allows searching by announcement title and status.

[0200] The operation log module is used to manage operation logs; log information includes log number, operation module, operation name, operation type, operator, operation result, operation date, and execution duration; it can be searched by system module, operator, type, status, and operation time;

[0201] The login log module is used to manage login information. Login information includes log type, username, login address, userAgent, result, and login date. It allows searching by login address, username, status, and login time.

[0202] The application management module is used to manage application clients; client information includes: client ID, client key, application name, application icon, status, access token validity period, refresh token validity period, authorization type, and creation time; it can search by application name and status;

[0203] The token management module is used to manage tokens; token information includes: access token, refresh token, user ID, user type, creation time, and expiration time; it can search by user ID, client ID, and user type;

[0204] The SMS management module is used to manage SMS messages. SMS information includes: SMS signature, channel code, activation status, SMS API account, SMS API key, and creation time. It can also be searched based on SMS signature, activation status, and creation time.

[0205] A document generation method based on a big data model includes:

[0206] Data source integration;

[0207] Data source verification;

[0208] Data annotation and feature extraction;

[0209] Build a document generation model;

[0210] Model optimization and evaluation;

[0211] Generate a first draft of the document;

[0212] Document polishing and improvement;

[0213] Gain a deep understanding of user needs and customize the documentation accordingly.

[0214] In one embodiment, the data source integration includes:

[0215] By leveraging Apache Spark's distributed file system (such as HDFS) interface, document-related data can be collected from multiple data sources (such as internal enterprise databases, data obtained from web crawlers, sensor data, etc.). For example, Spark's SparkContext can be used to read data files of different formats, such as the textFile method for reading text files.

[0216] Identify the types of multiple data sources, including relational databases (such as MySQL and PostgreSQL), non-relational databases (such as MongoDB and Cassandra), files in the local file system (such as CSV, JSON, and XML files), and files in network storage;

[0217] For document-related data from different data sources, analyze their format in detail. For example, the table structure and field types in relational databases; the document structure in non-relational databases (such as the JSON-like document structure in MongoDB); and the data format in files (such as column delimiters and data types in CSV files).

[0218] Configure the Apache Spark environment, ensuring the Spark cluster is correctly configured; this includes setting up the Master and Worker nodes. If using Spark on YARN or Spark Standalone mode, you need to configure it appropriately based on the cluster resources, such as allocating sufficient memory and CPU cores.

[0219] Ensure that the necessary dependency libraries are imported into your Spark project, especially the driver libraries related to the data source you want to connect to; for example, if you want to connect to MySQL, you need to import the MySQL JDBC driver; if you want to process JSON data, you may need to import the relevant JSON parsing library (Spark itself has some support for JSON, but additional configuration may be required).

[0220] Data is collected using Spark's distributed file system interface, specifically...

[0221] In Spark, the first step is to create a SparkSession object, which is the entry point for interacting with Spark.

[0222] To connect to different data sources, use Spark's JDBC interface for relational databases and the Spark-MongoDB connector for non-relational databases.

[0223] Spark's DataFrame operations can be used to combine data from different data sources; for example, the union operation can be used to merge multiple DataFrames together.

[0224] The collected data is initially cleaned to remove noisy data and incomplete data records.

[0225] In one embodiment, the data source verification includes:

[0226] Data sources are categorized into primary sources (such as authoritative academic databases and official government statistics) and secondary sources (such as industry reports and non-governmental research).

[0227] For each data point, label its source type;

[0228] For different sources, an initial credibility weight is assigned based on their historical accuracy and authority; for example, the initial credibility weight for authoritative academic databases is 0.8, for official government statistics it is 0.9, for industry reports it is 0.6, and for independent research it is 0.5. This is represented by a matrix C, where C... ij This represents the confidence weight of the i-th data point originating from the j-th type of source.

[0229] When multiple sources provide the same or similar data points, calculate the degree of consistency of their data.

[0230] Suppose that a data point x is provided by n sources, and the values ​​of these sources are x1, x2, ..., xn. n ; Calculate the consistency index in

[0231] If I is greater than a set threshold, the data is considered reliable in terms of consistency.

[0232] Update the source credibility weight based on the data consistency check results;

[0233] If the data consistency is good, the credibility weight of the source providing the data point increases. Where n is the number of sources providing the data point;

[0234] If data consistency is poor, the credibility weight of the source providing that data point will be reduced.

[0235] As new data is continuously input, repeat the above steps to continuously update the credibility matrix and verify the accuracy of the data;

[0236] For sources with a credibility weight below a certain threshold, their data should be marked as requiring further manual review or discarded directly.

[0237] In addition to numerical verification, the semantics and logical relationships of the data are also analyzed.

[0238] For example, if one set of data indicates population growth in a region, while another set indicates stagnant housing construction and significant population outflow, these two sets of data are logically contradictory. By establishing a logical relationship model, the causal and correlational relationships between the data can be analyzed to further validate the data.

[0239] Let the logical relationship be represented by the function L(x, y), where x and y are the relevant data points. If the value of L(x, y) exceeds a reasonable range, the relevant data will be re-examined or marked as suspicious data.

[0240] In one embodiment, the data annotation and feature extraction include:

[0241] The cleaned data is labeled. If it is used to generate a specific type of document, it can be labeled manually or automatically according to the document's theme and style.

[0242] Extract features from the data, including word frequency features, part-of-speech features, and semantic features; integrate with Spark using a natural language processing library; use Spark's mapPartitions operation to call NLTK functions to compute features within each data partition, and then merge the results.

[0243] In one embodiment, the construction of the document generation model includes:

[0244] Using a neural network language model, in the Spark environment, the model is trained using distributed computing. Large-scale text data is divided into multiple small batches, and the model is trained in parallel on multiple Spark nodes.

[0245] We employ transfer learning to fine-tune a pre-trained model on a large dataset to adapt it to a specific document generation task.

[0246] In one embodiment, the model optimization and evaluation includes:

[0247] Optimize the model on a Spark cluster using stochastic gradient descent and its variants for parameter updates; Spark's distributed computing capabilities can accelerate the optimization process by computing gradients and updating parameters in parallel across multiple nodes.

[0248] Multiple evaluation metrics are used to evaluate the model's performance. A cross-validation evaluation method is implemented on Spark, which divides the data into training, validation, and test sets. The model is evaluated on different subsets and its hyperparameters are tuned.

[0249] In one embodiment, the generated initial document draft includes:

[0250] Using a trained model, a draft document is generated based on the given input. In the Spark environment, the input data is distributed to multiple nodes, each node generates a portion of the document content based on the model, and then these contents are merged. If a long document is to be generated, the document is divided into multiple paragraphs, and each node is responsible for generating one paragraph.

[0251] For the generated draft, preliminary formatting adjustments can be made by writing custom rules or using natural language processing tools.

[0252] In one embodiment, the document polishing and improvement includes:

[0253] Additional language processing techniques can be used to polish the initial draft of the document. In Spark, external grammar checking tools can be called or predefined word substitution rules can be used to process the generated document.

[0254] The document was further improved based on specific needs, and the tone of the document was adjusted and citations or case studies were added, taking into account domain knowledge and user feedback.

[0255] The process of gaining a deep understanding of user needs and customizing documents includes:

[0256] The system performs lexical and syntactic analysis on user-input instructions, breaking them down into basic semantic units, recognizing not only individual words but also phrases and idioms. For example, for instructions containing industry-specific terminology, it identifies the full semantic scope of the terminology rather than simply matching words.

[0257] Semantic role labeling is used to determine the role of each semantic unit in the entire instruction, such as agent, recipient, or tool. This helps to understand the relationships between the elements in the instruction.

[0258] Construct a semantic network that connects the parsed semantic units according to their role relationships; this network can better represent the overall intent of the instruction, rather than looking at each word or phrase in isolation.

[0259] Construct a large-scale knowledge graph containing professional knowledge from various industries, internal culture and style of different companies, etc.; the nodes of this knowledge graph represent different concepts (such as industry terms, company departments, etc.), and the edges represent the relationships between concepts (such as ownership relationships, process relationships, etc.).

[0260] The semantic network obtained from the intent parsing module is matched with the knowledge graph. By calculating the similarity between nodes in the semantic network and nodes in the knowledge graph, as well as the similarity between edges, the most matching knowledge fragment is found. For example, if the user instruction involves the internal style of a specific company, style nodes related to that company and their relationships in the knowledge graph are found.

[0261] Based on the results of knowledge graph matching, the templates and rules for document generation are dynamically adjusted; if industry-specific professional terms are matched, the vocabulary selection and grammatical structure are adjusted to conform to the expression habits of that industry; for the internal style of a specific company, the tone, format and common expressions of the document are adjusted.

[0262] Reinforcement learning techniques are used to adjust the dynamic adaptation strategy based on how well the generated document matches the user's expectations (through user feedback or predefined evaluation metrics). For example, if the generated document does not meet the user's expectations in terms of the company's internal style, the reinforcement learning algorithm will adjust the parameters related to the company style in the dynamic adaptation module so that a more suitable document can be generated next time.

[0263] After the document is generated, a user feedback mechanism is provided; users can rate the generated document or provide specific suggestions for modification.

[0264] The user feedback is used as input to re-enter the intent parsing module and adjust the understanding of the user's intent. Based on the new intent understanding, the document is regenerated through the knowledge graph matching and dynamic adaptation modules until the user is satisfied.

[0265] The formula is expressed as follows:

[0266] Let the user input command be I, and the semantic network obtained by the intent parsing module be SN = f1(I), where f1 is the intent understanding function;

[0267] Let the knowledge graph be KG, and the matching degree function be m(SN, KG) = g(SN, KG), where g is a function for calculating the matching degree between the semantic network SN and the knowledge graph KG;

[0268] Let the generated document template be T, and the dynamically adapted template be T' = h(T, m(SN, KG)), where h is a function that adjusts the template based on the matching degree;

[0269] Let the document generation function be D(T',I)=d(T',I), and the final generated document be Doc=d(T',I);

[0270] In the feedback loop, let the user feedback be F, and the updated intent be I' = k(F, I), where k is a function that updates the intent based on the feedback. Then, the document is generated again according to the above process.

[0271] The algorithm flow is as follows:

[0272] Step 1: Intent Analysis

[0273] Receive user input command I;

[0274] Lexical and syntactic analysis is performed on I to obtain basic semantic units;

[0275] Perform semantic role labeling and construct a semantic network SN = f1(I);

[0276] Step 2: Knowledge graph matching;

[0277] Build or load a knowledge graph (KG);

[0278] Calculate the matching degree m(SN, KG) = g(SN, KG);

[0279] Step 3: Dynamic Adaptation:

[0280] Based on the matching degree m(SN, KG), adjust the document generation template T to obtain T' = h(T, m(SN, KG));

[0281] Step 4: Document Generation

[0282] Based on the dynamically adapted template T' and the user input command I, generate document Doc = d(T', I);

[0283] Step 5: Feedback Loop

[0284] If the user provides feedback F, update the intent I' = k(F, I);

[0285] Repeat step one until the user is satisfied.

[0286] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0287] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A document generation method based on a big data model, characterized in that, include: Data source integration; Data source verification; Data annotation and feature extraction; Build a document generation model; Model optimization and evaluation; Generate a first draft of the document; Document polishing and improvement; Gain a deep understanding of user needs and customize the documentation accordingly; The data source verification includes: Data sources are categorized into primary and secondary sources; For each data point, label its source type; For different sources, an initial credibility weight is assigned based on their historical accuracy and authority; a matrix is ​​used to... C To indicate, among which C ij Indicates the first i The data point comes from the first... j Credibility weights for different source types; When multiple sources provide the same or similar data points, calculate the degree of consistency of their data. Set data points x There are n sources providing the values, and the values ​​from these sources are respectively x 1, x 2,···, x n ; Calculate the consistency index ; if I If the data exceeds a set threshold, it is considered reliable in terms of consistency. Update the source credibility weight based on the data consistency check results; If the data consistency is good, the credibility weight of the source providing the data point increases. ;in n This refers to the number of sources providing that data point; If data consistency is poor, the credibility weight of the source providing that data point will be reduced. ; As new data is continuously input, repeat the above steps to continuously update the credibility matrix and verify the accuracy of the data; For sources with a credibility weight below a certain threshold, their data should be marked as requiring further manual review or discarded directly. The semantic and logical relationships of the data are analyzed, and the causal and correlation relationships between the data are analyzed by establishing a logical relationship model, so as to further verify the data; Let logical relationships be expressed using functions. L ( x , y ) indicates that among them x and y These are relevant data points, if L ( x , y If the value of ) exceeds a reasonable range, the relevant data will be re-examined or marked as suspicious data; The document generation method based on a big data model employs a document generation system based on a big data model, which includes: The parameter settings module is used to configure the knowledge base, model selection, and role settings; The report list view module allows you to view the document title, tags, and creation time; it also allows you to search for documents by their creation time and document tags. The feedback module is used to input feedback on the system; The system login module is used by users to log in by entering their username and password; The model management module is used to enable and disable model deployment; The knowledge base management module is used to view, close, and delete knowledge bases. It can also search for documents in the knowledge base by document title, file name, document tags, and creation time. The conversation management module includes a user management unit and a sensitive word management unit; The user management unit is used to search for users based on user ID, user name, and creation time, and can also view users' historical conversations; The sensitive word unit is used to search for sensitive words based on sensitive words, tags, whether it is enabled, and creation time, as well as to modify and delete sensitive words; it can also add new sensitive words. The feedback management module is used to manage feedback. The system management module is used to manage multiple functional modules; The system management module includes: The tenant management unit includes a tenant list sub-unit and a tenant package sub-unit. The tenant list sub-unit contains a tenant list that can be searched by tenant name, contact person, contact phone number, tenant status, and creation time. It also allows modification of tenant information and deletion of tenants. The tenant package sub-unit is used to search based on package name, status, and modification time. The tenant package sub-unit contains a package list. The user management module is used to manage users. It includes an intelligent customer service portal for managing and allocating users according to the head office, branches, and various departments within the company. The module also features a search function that allows users to search by name, phone number, status, and creation time. Furthermore, it allows modification of user information. The role management module is used to manage different roles. It can search for roles by role name, role identifier, status, and creation time. Role names include test accounts, management roles, and super administrators. It can also modify role information, change permissions, and delete roles. The menu management module is used to manage the menu list and can search for menus by name and status. The menu names include: Boot Development Documentation, Cloud Development Documentation, Model Management, Knowledge Base Management, Dialogue Management, Feedback Management, System Management, Payment Management, Report Management, Workflow, Product Center, Order Center, Marketing Center, and Official Account Management. The module also allows modification of menus. The department management module is used to manage the various departments of the head office and branches; it allows users to modify, add, and delete departments; and it also allows users to search by department name and status. The job management module is used to modify, add, and delete jobs. Job information includes: job number, job code, job name, job sorting, and status. It can also search for jobs by job code, job name, and status. The dictionary management module allows users to modify, add, and delete dictionaries. Dictionary information includes: dictionary number, dictionary name, dictionary type, status, and creation time. Users can also search for dictionaries by name, dictionary type, status, and creation time. The notification and announcement module is used to modify, add, and delete notifications and announcements. Announcement information includes: announcement title, announcement type, status, and creation time. It also allows searching by announcement title and status. The operation log module is used to manage operation logs; log information includes log number, operation module, operation name, operation type, operator, operation result, operation date, and execution duration; it can be searched by system module, operator, type, status, and operation time; The login log module is used to manage login information. Login information includes log type, username, login address, userAgent, result, and login date. It allows searching by login address, username, status, and login time. The application management module is used to manage application clients; client information includes: client ID, client key, application name, application icon, status, access token validity period, refresh token validity period, authorization type, and creation time; it can search by application name and status; The token management module is used to manage tokens; token information includes: access token, refresh token, user ID, user type, creation time, and expiration time; it can search by user ID, client ID, and user type; The SMS management module is used to manage SMS messages. SMS information includes: SMS signature, channel code, activation status, SMS API account, SMS API key, and creation time. It can also be searched based on SMS signature, activation status, and creation time.

2. The document generation method according to claim 1, characterized in that, The data sources integrated include: Use Apache Spark's distributed file system interface to collect document-related data from multiple data sources; Identify the types of multiple data sources, including relational databases, non-relational databases, files in the local file system, and files in network storage; For document-related data from different data sources, analyze their format in detail; Configure the Apache Spark environment and ensure the Spark cluster is configured correctly; ensure that the dependency libraries are imported into the Spark project, including the driver libraries related to the data source to be connected. Data is collected using Spark's distributed file system interface, specifically... In Spark, the first step is to create a SparkSession object, which is the entry point for interacting with Spark. To connect to different data sources, use Spark's JDBC interface for relational databases and the Spark-MongoDB connector for non-relational databases. Spark's DataFrame operations can be used to integrate data from different data sources. The collected data is initially cleaned to remove noisy data and incomplete data records.

3. The document generation method according to claim 1, characterized in that, The data annotation and feature extraction include: The cleaned data is labeled. If it is used to generate a specific type of document, it is labeled manually or automatically according to the document's theme and style. Extract features from the data, including word frequency features, part-of-speech features, and semantic features; integrate with Spark using a natural language processing library; use Spark's mapPartitions operation to call NLTK functions to compute features within each data partition, and then merge the results.

4. The document generation method according to claim 1, characterized in that, The document generation model includes: Using a neural network language model, in the Spark environment, the model is trained using distributed computing. Large-scale text data is divided into multiple small batches, and the model is trained in parallel on multiple Spark nodes. We employ transfer learning to fine-tune a pre-trained model on a large dataset to adapt it to a specific document generation task.

5. The document generation method according to claim 1, characterized in that, The model optimization and evaluation include: Model optimization is performed on a Spark cluster, using stochastic gradient descent and its variants for parameter updates; Spark's distributed computing capabilities accelerate the optimization process by computing gradients and updating parameters in parallel across multiple nodes. Multiple evaluation metrics are used to evaluate the model's performance. A cross-validation evaluation method is implemented on Spark, which divides the data into training, validation, and test sets. The model is evaluated on different subsets and its hyperparameters are tuned.

6. The document generation method according to claim 1, characterized in that, The generated initial draft document includes: Using a trained model, a draft document is generated based on the given input. In the Spark environment, the input data is distributed to multiple nodes, each node generates a portion of the document content based on the model, and then these contents are merged. If a long document is to be generated, the document is divided into multiple paragraphs, and each node is responsible for generating one paragraph. For the generated draft, preliminary formatting adjustments can be made by writing custom rules or using natural language processing tools.

7. The document generation method according to claim 1, characterized in that, The document polishing and improvement includes: Additional language processing techniques can be used to polish the initial draft of the document. In Spark, external grammar checking tools can be called or predefined word substitution rules can be used to process the generated document. The document was further improved based on specific needs, and the tone of the document was adjusted and citations or case studies were added, taking into account domain knowledge and user feedback. The process of gaining a deep understanding of user needs and customizing documents includes: Perform lexical and syntactic analysis on user-input commands; break down commands into basic semantic units, recognizing not only individual words, but also phrases and sentence structures; Semantic role labeling technology is used to determine the role of each semantic unit in the entire instruction; a semantic network is constructed to connect the parsed semantic units according to their role relationships; Construct a large-scale knowledge graph containing professional knowledge from various industries, internal culture and style information from different companies; the nodes of this knowledge graph represent different concepts, and the edges represent the relationships between concepts. The semantic network obtained by the intent parsing module is matched with the knowledge graph; the most matching knowledge fragment is found by calculating the similarity between nodes in the semantic network and nodes in the knowledge graph, as well as the similarity between edges. Based on the results of knowledge graph matching, the templates and rules for document generation are dynamically adjusted; if industry-specific professional terms are matched, the vocabulary selection and grammatical structure are adjusted to conform to the expression habits of that industry. Reinforcement learning techniques are used to adjust the dynamic adaptation strategy based on how well the generated document matches the user's expectations; After the document is generated, a user feedback mechanism is provided; users can rate the generated document or provide specific suggestions for modification. Using user feedback as input, the system re-enters the intent parsing module to adjust its understanding of user intent. Based on the new intent understanding, the document is then regenerated through the knowledge graph matching and dynamic adaptation modules. Until the user is satisfied; The formula is expressed as follows: Let the user input command be I The semantic network obtained by the intent parsing module is SN = f 1( I ),in f 1 is the intention understanding function; Let the knowledge graph be KG The matching degree function is m ( SN , KG )= g ( SN , KG ),in g It is a computational semantic network SN With knowledge graph KG A function for matching degree; Let the document generation template be T The dynamically adapted template is T’ = h ( T , m (S N , KG )),in h It is a function that adjusts the template based on the matching degree; Let the document generation function be... D ( T’ , I )= d ( T ', I The final generated document is Doc = d ( T’ , I ); In the feedback loop, let the user feedback be... F The updated intent is I’ = k ( F , I ),in k It is a function that updates the intent based on feedback, and then regenerates the document according to the above process; The algorithm flow is as follows: Step 1: Intent Analysis Receive user input commands I ; right I Lexical and syntactic analysis is performed to obtain basic semantic units; Perform semantic role labeling and construct a semantic network. SN = f 1( I ); Step 2: Knowledge graph matching; Building or loading knowledge graphs KG ; Calculate matching degree m ( SN , KG )= g ( SN , KG ); Step 3: Dynamic Adaptation: Based on matching degree m ( SN , KG Adjust the document generation template. T ,get T’ = h ( T , m (S N , KG )); Step 4: Document Generation Based on the dynamically adapted template T’ and user input commands I Generate document Doc = d ( T’ , I ); Step 5: Feedback Loop If the user provides feedback F, update the intent. I’ = k ( F , I ); Repeat step one until the user is satisfied.

Citation Information

Patent Citations

  • Document generation system, method and equipment based on knowledge base and large model and medium

    CN117556010A