Document generation system and method based on big data model
Through a document generation system based on the big data model, integrating and verifying data sources, extracting features, building and optimizing document generation models, generating and polishing documents, and deeply understanding user needs, the shortcomings of the existing document generation system in terms of accuracy, quality and personalization are solved, and efficient and personalized document generation is achieved.
Patent Information
- Application Number
- CN202510082013.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing document generation system has shortcomings in content accuracy, logical coherence, content quality, personalized customization and user intention understanding, resulting in the generated documents that may be inaccurate, logically confusing, low quality, inability to meet specific needs and inability to fully understand user intentions.
A document generation system based on a big data model is adopted to improve the accuracy, quality and personalization of document generation through data source integration, verification, annotation and feature extraction, construction of document generation models, model optimization and evaluation, generation of document first drafts, document polishing and improvement, and in-depth understanding of user needs and customizing and modifying documents.
It achieves high content accuracy and high-quality document generation, can meet personalized customization needs, and improves the logical coherence and user satisfaction of the document.
Smart Images

Figure CN120124596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of question and answer management, and particularly to a document generation system and method based on a big data model. Background Art
[0002] Automatic text generation is one of the core issues in the field of natural language processing. Automatically generating text according to the author's ideas can greatly reduce the author's workload; however, at the same time, it is also challenging to generate text in a specified aspect according to the author's ideas. The following are the problems existing in the existing document generation systems:
[0003] 1. In terms of content accuracy, sometimes content that does not conform to the facts is generated. For example, when generating documents related to historical events, the events, locations, or relationships between characters where the events occurred may be confused. This is because the data sources may be inaccurate or outdated, and there is a lack of sufficient verification mechanisms.
[0004] The generated documents may have problems in logical coherence. For example, when discussing viewpoints, there is a lack of a reasonable derivation relationship between the arguments and the viewpoints, or incorrect associations occur when elaborating on the causal relationships of events. This is because the document generation system is not perfect enough when constructing the logical framework and is difficult to fully simulate human logical thinking ability.
[0005] 2. In terms of content quality, the generated documents are often relatively single in expression and lack literary grace. The language style may be relatively mechanical, mostly using some common sentence patterns and vocabulary, and it is difficult to use rich rhetorical devices and diverse sentence structures like human authors to increase the attractiveness of the article.
[0006] For some complex topics, the content generated by the document generation system is often only a superficial overview and it is difficult to deeply explore the connotation of the topic. For example, in the generation of academic papers, although papers that meet the format requirements can be pieced together, there are obvious deficiencies in the in-depth analysis of research questions, unique insights, etc.
[0007] 3. In terms of personalized customization, when users have very specific personalized requirements, such as professional terms for a specific industry, the internal style of a specific company, etc., the document generation system may not be able to accurately meet them. It is more based on general data for generation and has poor adaptability to special requirements.
[0008] The system cannot fully and accurately understand the user's intention. The instructions input by the user may be somewhat ambiguous, and the document generation system may not be able to further clarify the user's intention like a human being through inquiries, etc., resulting in a large gap between the generated document and the user's expectations.
[0009] Therefore, there is an urgent need in the art for a technical solution for document generation that has high content accuracy, high content quality, and can achieve personalized customization.
[0010] The information disclosed in this background art section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention
[0011] The object of the present invention is to provide a document generation system and method that have high content accuracy, high content quality, and can achieve personalized customization.
[0012] To achieve the above object, the present invention provides the following solutions:
[0013] A document generation system based on a big data model, comprising:
[0014] A parameter setting module for setting a knowledge base, model selection, and role setting;
[0015] A report list viewing module for viewing the title, tags, and creation time of a document; and capable of searching for a document by searching for the generation time and document tags of the document;
[0016] An opinion feedback module for inputting feedback opinions on the system;
[0017] A system login module for a user to input an account number and password for login;
[0018] A model management module for enabling and disabling model deployment;
[0019] A knowledge base management module for viewing, closing, and deleting a knowledge base, and capable of searching for documents in the knowledge base according to the document title, file name, document tags, and creation time;
[0020] A dialogue management module, including a user management unit and a sensitive word management unit;
[0021] The user management unit is used to search for users according to user ID, user name, and creation time, and can also view the historical conversations of users;
[0022] The sensitive word unit is used to search for sensitive words according to sensitive words, tags, whether it is enabled, and creation time, and to modify and delete sensitive words; and can also add new sensitive words;
[0023] A feedback management module for managing feedback opinions
[0024] A system management module for managing multiple functional modules.
[0025] Optionally, the system management module includes:
[0026] A tenant management unit, which is provided with a tenant list subunit and a tenant package subunit;
[0027] The tenant list subunit is provided with a tenant list, which can be searched according to tenant name, contact person, contact mobile phone, tenant status, and creation time, and can also modify tenant information and delete tenants;
[0028] The tenant package subunit is used to search according to package name, status, and modification time, and a package list is set in the tenant package subunit;
[0029] A user management module for managing users; an intelligent customer service company port is set in the user management to manage and allocate users according to the head office, branch offices, and each department of the company. The user management module is provided with a search port, which can be searched according to user name, mobile phone number, status, and creation time; it can also modify user information;
[0030] A role management module for managing different roles, which can search for roles according to role name, role identifier, status, and creation time. Role names include test accounts, management roles, and super administrators; it can modify role information, change permissions, and delete roles;
[0031] A menu management module for managing the menu list, which can search for menus according to menu name and status; the menu names include: Boot Development Documentation, Cloud Development Documentation, Model Management, Knowledge Base Management, Conversation Management, Feedback Management, System Management, Payment Management, Report Management, Workflow, Product Center, Order Center, Marketing Center, and Official Account Management; it can also modify the menu;
[0032] A department management module for managing each department of the head office and branch offices; it can modify, add, and delete departments; it can also search according to department name and status;
[0033] A position management module for modifying, adding, and deleting positions; position information includes: position number, position code, position name, position sorting, and status; it can search for positions according to position code, position name, and status;
[0034] A dictionary management module for modifying, adding, and deleting dictionaries; dictionary information includes: dictionary number, dictionary name, dictionary type, status, and creation time; it can search according to dictionary name, dictionary type, status, and creation time;
[0035] The notice and announcement module is used to modify, add, and delete notices and announcements; the announcement information includes: announcement title, announcement type, status, and creation time; it is possible to search according to the announcement title and announcement status;
[0036] The operation log module is used to manage operation logs; the log information includes log number, operation module, operation name, operation type, operator, operation result, operation date, and execution duration; it is possible to search according to the system module, operator, type, status, and operation time;
[0037] The login log module; is used to manage login situations; the login information includes log type, user name, login address, userAgent, result, and login date; it is possible to search according to the login address, user name, status, and login time;
[0038] The application management module is used to manage application clients; the client information includes: client number, client secret, application name, application icon, status, validity period of access token, validity period of refresh token, authorization type, and creation time; it is possible to search according to the application name and status;
[0039] The token management module is used to manage tokens; the token information includes: access token, refresh token, user number, user type, creation time, and expiration time; it is possible to search according to the user number, client number, and user type;
[0040] The SMS management module is used to manage SMS; the SMS information includes: SMS signature, channel code, enabled status, account of SMS API, secret key of SMS API, and creation time; it is possible to search according to the SMS signature, enabled status, and creation time.
[0041] A document generation method based on a big data model, including:
[0042] Integrating data sources;
[0043] Verifying data sources;
[0044] Data annotation and feature extraction;
[0045] Constructing a document generation model;
[0046] Model optimization and evaluation;
[0047] Generating a preliminary draft of the document;
[0048] Polishing and perfecting the document;
[0049] Deeply understanding user needs and making customized modifications to the document.
[0050] Optionally, the data source integration includes:
[0051] Collect document-related data from multiple data sources using the distributed file system interface of Apache Spark;
[0052] Determine the types of multiple data sources, including relational databases, non-relational databases, files in the local file system, and files in network storage;
[0053] For the document-related data in different data sources, analyze its format in detail;
[0054] Configure the Apache Spark environment to ensure that the Spark cluster is correctly configured; ensure that the necessary dependency libraries are imported in the Spark project, especially the driver libraries related to the data sources to be connected;
[0055] Use the distributed file system interface of Spark to collect data. Specifically,
[0056] In Spark, first create a SparkSession object, which is the entry point for interacting with Spark;
[0057] Connect to different data sources. For relational databases, use the JDBC interface of Spark; for non-relational databases, the Spark-MongoDB connector is required;
[0058] Use the DataFrame operations of Spark to integrate the data obtained from different data sources;
[0059] Perform preliminary cleaning on the collected data to remove noisy data and incomplete data records.
[0060] Optionally, the data source verification includes:
[0061] Classify the data sources into primary sources and secondary sources;
[0062] Mark the source type for each data point;
[0063] For different source types, assign an initial credibility weight according to their historical accuracy and authority; represented by a matrix C, where C ij represents the credibility weight that the i-th data point comes from the j-th type of source;
[0064] When multiple sources provide the same or similar data points, calculate the degree of consistency of their data;
[0065] Suppose there are n sources providing the data point x, and the values of these sources are x 1 , x 2 , ···, xn ; Calculate the consistency index where
[0066] If I is greater than a set threshold, the data is considered reliable in terms of consistency;
[0067] Update the credibility weight of the source according to the result of the data consistency check;
[0068] If the data consistency is good, for the source providing this data point, its credibility weight increases where n is the number of sources providing this data point;
[0069] If the data consistency is poor, for the source providing this data point, its credibility weight decreases
[0070] As new data is continuously input, repeat the above steps to continuously update the credibility matrix and verify the accuracy of the data;
[0071] For sources with credibility weights below a lower limit, mark their data as requiring further manual review or directly discard it;
[0072] Analyze the semantic and logical relationships of the data. By establishing a logical relationship model, analyze the causal and correlative relationships between the data to further verify the data;
[0073] Let the logical relationship be represented by the function L(x, y), where x and y are related data points. If the value of L(x, y) exceeds a reasonable range, re-review the relevant data or mark it as suspicious data.
[0074] Optionally, the data annotation and feature extraction include:
[0075] Annotate the cleaned data. If it is used to generate a specific type of document, it can be manually or automatically annotated according to the theme and style of the document;
[0076] Extract the features of the data, including word frequency features, part-of-speech features, and semantic features; Integrate with the natural language processing library using Spark; Through the mapPartitions operation of Spark, call the functions of NLTK within each data partition to calculate the features, and then merge the results.
[0077] Optionally, the construction of the document generation model includes:
[0078] Use a neural network language model. In the Spark environment, utilize distributed computing to train the model. Divide the large-scale text data into multiple small batches and train the model in parallel on multiple nodes of Spark;
[0079] Adopt the method of transfer learning, and use the pre-trained model to fine-tune on a large dataset to adapt to specific document generation tasks.
[0080] Optionally, the model optimization and evaluation include:
[0081] Optimize the model on the Spark cluster, and use stochastic gradient descent and its variants for parameter updates; the distributed computing ability of Spark can accelerate the optimization process by calculating gradients and updating parameters in parallel on multiple nodes;
[0082] Adopt multiple evaluation metrics to evaluate the performance of the model, implement the cross-validation evaluation method on Spark, divide the data into training set, validation set and test set, evaluate the model on different subsets and adjust the hyperparameters of the model.
[0083] Optionally, the generation of the initial draft of the document includes:
[0084] Use the trained model to generate the initial draft of the document according to the given input; in the Spark environment, distribute the input data to multiple nodes, each node generates part of the document content according to the model, and then merge these contents; if generating a long document, divide the document into multiple paragraphs, and each node is responsible for generating one paragraph;
[0085] For the generated initial draft, perform preliminary format adjustment by writing custom rules or using natural language processing tools.
[0086] Optionally, the document polishing and improvement include:
[0087] Use additional language processing techniques to polish the initial draft of the document. In Spark, call external grammar checking tools or use predefined vocabulary replacement rules to process the generated document;
[0088] Further improve the document according to the specific requirements of the document, and adjust the tone of the document and add references or cases in combination with some domain knowledge and user feedback;
[0089] The in-depth understanding of user requirements and customized modification of the document include:
[0090] First, perform lexical and syntactic analysis on the instructions input by the user; decompose the instructions into basic semantic units, not just identifying individual words, but also including phrase and phrasal structures;
[0091] Adopt semantic role labeling technology to determine the role of each semantic unit in the whole instruction; construct a semantic network and connect the parsed semantic units according to their role relationships;
[0092] Construct a large-scale knowledge graph that contains information such as professional knowledge in various industries, the internal cultures and styles of different companies, etc.; the nodes of this knowledge graph represent different concepts, and the edges represent the relationships between concepts.
[0093] Match the semantic network obtained by the intent parsing module with the knowledge graph; by calculating the similarity between the nodes in the semantic network and the nodes in the knowledge graph, as well as the similarity between the edges and edges, find the most matching knowledge fragment.
[0094] According to the results of the knowledge graph matching, dynamically adjust the templates and rules for document generation; if specific industry-specific terms are matched, adjust the word selection and grammar structure to conform to the expression habits of that industry.
[0095] Adopt reinforcement learning technology to adjust the dynamic adaptation strategy according to the matching degree between the generated document and the user's expectations.
[0096] After the document is generated, provide a user feedback mechanism; users can rate the generated document or provide specific modification suggestions.
[0097] Take the user feedback as input, re-enter the intent parsing module, and adjust the understanding of the user's intent; according to the new understanding of the intent, go through the knowledge graph matching and dynamic adaptation modules again to regenerate the document until the user is satisfied.
[0098] The formula is expressed as follows:
[0099] Let the user input instruction be I, and the semantic network obtained by the intent parsing module be SN = f 1 (I), where f 1 is the intent understanding function;
[0100] Let the knowledge graph be KG, and the matching degree function be m(SN, KG) = g(SN, KG), where g is the function for calculating the matching degree between the semantic network SN and the knowledge graph KG.
[0101] Let the document generation template be T, and the dynamically adapted template be T' = h(T, m(SN, KG)), where h is the function for adjusting the template according to the matching degree.
[0102] Let the document generation function be D(T', I) = d(T', I), and the finally generated document be Doc = d(T', I);
[0103] In the feedback loop, let the user feedback be F, and the updated intent be I' = k(F, I), where k is the function for updating the intent according to the feedback, and then regenerate the document according to the above process.
[0104] The algorithm flow is as follows:
[0105] Step 1: Intention Analysis:
[0106] Receive the user input instruction I;
[0107] Perform lexical and syntactic analysis on I to obtain basic semantic units;
[0108] Perform semantic role labeling to construct a semantic network SN = f 1 (I);
[0109] Step 2: Knowledge Graph Matching;
[0110] Construct or load the knowledge graph KG;
[0111] Calculate the matching degree m(SN, KG) = g(SN, KG);
[0112] Step 3: Dynamic Adaptation:
[0113] According to the matching degree m(SN, KG), adjust the document generation template T to obtain T' = h(T, m(SN, KG));
[0114] Step 4: Document Generation:
[0115] Generate a document Doc = d(T', I) according to the dynamically adapted template T' and the user input instruction I;
[0116] Step 5: Feedback Loop:
[0117] If the user provides feedback F, update the intention I' = k(F, I);
[0118] Restart from Step 1 until the user is satisfied.
[0119] Compared with the prior art, the present invention has the following beneficial effects:
[0120] The present invention provides a document generation system and method based on a big data model, which can produce the following
[0121] Beneficial effects:
[0122] 1. Data Source Integration:
[0123] Data Comprehensiveness:
[0124] By integrating multiple data sources, the data can be made more comprehensive. For example, when constructing a document generation model, data from different sources can cover more topics, fields, and information types, thereby providing a broader knowledge base for the model and improving the richness of document generation.
[0125] Reduction of Data Bias:
[0126] By integrating multiple data sources, the bias that may exist in a single source can be avoided. If only relying on one data source, the model may produce one-sided results due to the specific tendencies of that source (such as the views of a specific region or specific population). After integration, the data can be made more representative and the accuracy of the model can be improved.
[0127] 2. Data source verification:
[0128] Data reliability:
[0129] Verifying the data source ensures the reliability of the data. When building a document generation model, reliable data is the foundation. If unverified data is used, it may cause the model to generate incorrect or misleading content. Verification can exclude false, inaccurate, or outdated data, thereby improving the credibility of the model output.
[0130] Accuracy guarantee:
[0131] Helps to guarantee the accuracy of the data. In today's data environment, the types of data are intricate. Verifying the data source can ensure the accuracy of the data and avoid the model generating incorrect documents due to inaccurate data sources.
[0132] 3. Data annotation and feature extraction
[0133] Feature recognition:
[0134] Can accurately identify the features in the data, which is crucial for building a document generation model. For example, in natural language processing, through data annotation and feature extraction, features such as the part of speech and semantics of words can be identified, and the model can better understand the input content based on these features and generate reasonable documents.
[0135] Improve the model's pertinence:
[0136] Enables the model to learn and generate more pertinently. By extracting the key features related to document generation, the model can focus on these important information instead of being interfered by irrelevant information, thereby improving the quality and efficiency of document generation.
[0137] 4. Build a document generation model
[0138] Automated document generation:
[0139] Can achieve the automated generation of documents. For a large number of document requirements, such as news reports, business documents, etc., the model can quickly generate a draft according to the input requirements, saving labor and time costs.
[0140] Consistency and standardization:
[0141] Ensure the consistency and standardization of document generation. The model generates documents according to pre-set rules and algorithms, enabling the generated documents to be consistent in format, style, content structure, etc., and meet specific standards and requirements.
[0142] 5. Model Optimization and Evaluation
[0143] Performance Improvement:
[0144] Through optimization and evaluation, the performance of the model can be continuously improved. For example, optimizing algorithms can increase the running speed of the model, and evaluation metrics can point out the deficiencies of the model, enabling targeted improvements to make the model more accurate and efficient in document generation.
[0145] Enhanced Adaptability:
[0146] Enable the model to better adapt to different tasks and data environments. As data changes and task requirements evolve, optimization and evaluation can adjust the parameters and structure of the model to maintain good adaptability and effectively generate documents in various situations.
[0147] 6. Generate the First Draft of the Document
[0148] Quick Start of Creation:
[0149] Provide a basis for quickly starting creation. Whether writing reports, articles, or other documents, the first draft can serve as a starting point for further improvement, saving the time of starting from scratch and allowing the creator to enter the stage of modification and improvement faster.
[0150] Content Framework Construction:
[0151] Help to construct the content framework of the document. The first draft often contains the basic structure and main content points of the document, providing a framework support for subsequent polishing, supplementation, and customization modification.
[0152] 7. Document Polishing and Improvement
[0153] Improve the Document Quality:
[0154] Can significantly improve the quality of the document. Polishing can improve the language expression of the document, making it smoother, more accurate, and more professional; improvement can supplement the content omitted in the first draft and adjust the logical structure to make the document more complete and rigorous.
[0155] Style Adaptation:
[0156] Enable the style of the document to adapt to specific audiences or purposes. For example, polish a technical document from an academic style to a popular science style to better meet the needs of different readers.
[0157] 8. Deeply understand user needs and customize the document
[0158] Meet personalized needs
[0159] Ensure that the document can meet the personalized needs of users. Different users have different requirements for the document. By deeply understanding the needs and making customized modifications, the document can meet the specific purposes of users in terms of content, structure, style, etc., improving user satisfaction.
[0160] Enhanced targeting:
[0161] Enhance the targeting of the document. For specific audiences, tasks, or scenarios, the customized document can better convey information and achieve the expected effect, such as a business proposal for a specific customer or a research report for a specific academic conference. Brief description of the drawings
[0162] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0163] Figure 1 It is a schematic diagram of the home page of the document generation system provided by the embodiment of the present invention.
[0164] Figure 2 It is a schematic diagram of the interface of the parameter setting module provided by the embodiment of the present invention.
[0165] Figure 3 It is a schematic diagram of the interface of the report list viewing module provided by the embodiment of the present invention.
[0166] Figure 4 It is a schematic diagram of the system login interface provided by the embodiment of the present invention.
[0167] Figure 5 It is a schematic diagram of the home page of the system background provided by the embodiment of the present invention.
[0168] Figure 6 It is a diagram of the interface of the model management module provided by the embodiment of the present invention.
[0169] Figure 7 It is a diagram of the interface of the knowledge base management module provided by the embodiment of the present invention
[0170] Figure 8 It is a diagram of the interface of the dialogue management module provided by the embodiment of the present invention.
[0171] Figure 9 It is a diagram of the interface of the feedback module provided by the embodiment of the present invention.
[0172] Figure 10 This is the interface diagram of the system management module provided by the embodiments of the present invention. Specific embodiments
[0173] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0174] The purpose of the present invention is to provide a system and method for document generation with high content accuracy, high content quality, and capable of realizing personalized customization.
[0175] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0176] Embodiment 1:
[0177] This embodiment provides a document generation system based on a big data model, as Figure 1-10 shown, including:
[0178] A parameter setting module for setting the knowledge base, model selection, and role setting;
[0179] A report list viewing module for viewing the title, tags, and creation time of the document; it can also search for documents by searching for the generation time and document tags of the document;
[0180] An opinion feedback module for inputting feedback opinions on the system;
[0181] A system login module for users to input their accounts and passwords for login;
[0182] A model management module for starting and closing model deployment;
[0183] A knowledge base management module for viewing, closing, and deleting the knowledge base, and it can also search for documents in the knowledge base according to the document title, file name, document tags, and creation time;
[0184] A dialogue management module including a user management unit and a sensitive word management unit;
[0185] The user management unit is used to search for users according to the user ID, user name, and creation time, and it can also view the historical conversations of the users;
[0186] The sensitive word unit is used to search for sensitive words based on sensitive words, tags, whether it is enabled, and creation time, as well as modify and delete sensitive words; it can also add new sensitive words;
[0187] The feedback management module is used to manage feedback opinions
[0188] The system management module is used to manage multiple functional modules.
[0189] In one embodiment, the system management module includes:
[0190] The tenant management unit, in which a tenant list subunit and a tenant package subunit are set;
[0191] In the tenant list subunit, a tenant list is set, which can search according to tenant name, contact person, contact mobile phone, tenant status, and creation time, and can also modify tenant information and delete tenants;
[0192] The tenant package subunit is used to search according to package name, status, and modification time, and a package list is set in the tenant package subunit;
[0193] The user management module is used to manage users; an intelligent customer service company port is set in the user management to manage and allocate users according to the head office, branch offices, and each department of the company. The user management module has a search port, which can search according to user name, mobile phone number, status, and creation time; it can also modify user information;
[0194] The role management module is used to manage different roles, and can search for roles according to role name, role identifier, status, and creation time. The role names include test accounts, management roles, and super administrators; it can modify role information, change permissions, and delete roles;
[0195] The menu management module is used to manage the menu list and can search for menus according to menu name and status; the menu names include: Boot development documents, Cloud development documents, model management, knowledge base management, dialogue management, feedback management, system management, payment management, report management, work flow, product center, order center, marketing center, and official account management; it can also modify menus;
[0196] The department management module is used to manage each department of the head office and branch offices; it can modify, add, and delete departments; it can also search according to department name and status;
[0197] The position management module is used to modify, add, and delete positions; the position information includes: position number, position code, position name, position sorting, and status; it is capable of searching for positions based on the position code, position name, and status;
[0198] The dictionary management module allows users to modify, add, and delete dictionaries; the dictionary information includes: dictionary number, dictionary name, dictionary type, status, and creation time; it is capable of searching based on the dictionary name, dictionary type, status, and creation time;
[0199] The notice and announcement module is used to modify, add, and delete notices and announcements; the announcement information includes: announcement title, announcement type, status, and creation time; it is capable of searching based on the announcement title and announcement status;
[0200] The operation log module is used to manage operation logs; the log information includes log number, operation module, operation name, operation type, operator, operation result, operation date, and execution duration; it is capable of searching based on the system module, operator, type, status, and operation time;
[0201] The login log module; is used to manage login situations; the login information includes log type, user name, login address, userAgent, result, and login date; it is capable of searching based on the login address, user name, status, and login time;
[0202] The application management module is used to manage application clients; the client information includes: client number, client secret key, application name, application icon, status, validity period of access token, validity period of refresh token, authorization type, and creation time; it is capable of searching based on the application name and status;
[0203] The token management module is used to manage tokens; the token information includes: access token, refresh token, user number, user type, creation time, and expiration time; it is capable of searching based on the user number, client number, and user type;
[0204] The SMS management module is used to manage SMS; the SMS information includes: SMS signature, channel code, enabled status, account of SMS API, secret key of SMS API, and creation time; it is capable of searching based on the SMS signature, enabled status, and creation time.
[0205] A document generation method based on a big data model, including:
[0206] Integrating data sources;
[0207] Validating data sources;
[0208] Data annotation and feature extraction;
[0209] Build a document generation model;
[0210] Model optimization and evaluation;
[0211] Generate a preliminary draft of the document;
[0212] Document polishing and improvement;
[0213] Deeply understand the user's needs and make customized modifications to the document.
[0214] In one embodiment, the data source integration includes:
[0215] Use the distributed file system interface of Apache Spark (such as HDFS) to collect document-related data from multiple data sources (such as enterprise internal databases, data obtained from web crawlers, sensor data, etc.). For example, Spark's SparkContext can be used to read data files in different formats, such as the textFile method for reading text files.
[0216] Determine the types of multiple data sources, including relational databases (such as MySQL, PostgreSQL), non-relational databases (such as MongoDB, Cassandra), files in the local file system (such as files in CSV, JSON, XML formats), and files in network storage;
[0217] For the document-related data in different data sources, analyze its format in detail. For example, the table structure and field types in a relational database; the document structure in a non-relational database (such as the JSON-like document structure in MongoDB); the data format in a file (such as the column delimiter and data type in a CSV file).
[0218] Configure the Apache Spark environment to ensure that the Spark cluster is correctly configured; including setting up the Master node and Worker nodes. If using Spark on YARN or Spark Standalone mode, reasonable configuration is required according to the cluster resources, such as allocating sufficient memory, CPU cores, etc.
[0219] Ensure that the necessary dependency libraries are imported in the Spark project, especially the driver libraries related to the data sources to be connected; for example, if connecting to MySQL, the MySQL JDBC driver needs to be imported; if processing JSON data, relevant JSON parsing libraries may need to be imported (Spark itself has some support for JSON, but additional configuration may be required).
[0220] Use the distributed file system interface of Spark to collect data. Specifically,
[0221] In Spark, first create a SparkSession object, which is the entry point for interacting with Spark;
[0222] Connect to different data sources. For relational databases, use Spark's JDBC interface; for non-relational databases, use the Spark-MongoDB connector;
[0223] Use Spark's DataFrame operations to integrate the data obtained from different data sources; for example, the union operation can be used to merge multiple DataFrames together.
[0224] Perform preliminary cleaning on the collected data to remove noisy data and incomplete data records.
[0225] In one embodiment, the data source verification includes:
[0226] Classify the data sources into primary sources (such as authoritative academic databases, government official statistics, etc.) and secondary sources (such as industry reports, civilian research, etc.).
[0227] Mark the source type for each data point;
[0228] For different source types, assign an initial credibility weight according to their historical accuracy and authority; for example, the initial credibility weight of an authoritative academic database is 0.8, government official statistics is 0.9, industry reports is 0.6, and civilian research is 0.5. Represent it with a matrix C, where C ij represents the credibility weight that the i-th data point comes from the j-th type of source.
[0229] When multiple sources provide the same or similar data points, calculate the degree of consistency of their data;
[0230] Suppose there are n sources providing the data point x, and the values of these sources are x 1 , x 2 , ···, x n ; Calculate the consistency index where
[0231] If I is greater than a set threshold, it is considered that the data is reliable in terms of consistency;
[0232] Update the credibility weight of the source according to the result of the data consistency check;
[0233] If the data consistency is good, for the source providing this data point, its credibility weight increases where n is the number of sources providing the data point;
[0234] If the data consistency is poor, the credibility weight of the source providing the data point is reduced
[0235] As new data is continuously input, repeat the above steps to continuously update the credibility matrix and verify the accuracy of the data;
[0236] For sources with a credibility weight lower than a lower limit, mark their data as requiring further manual review or directly discard it;
[0237] In addition to numerical verification, the semantic and logical relationships of the data are also analyzed.
[0238] For example, if one data indicates population growth in a certain region, while another data indicates stagnant housing construction and a large outflow of population in the same region, these two sets of data are logically contradictory. By establishing a logical relationship model and analyzing the causal, correlative, etc. relationships between the data, the data is further verified.
[0239] Let the logical relationship be represented by the function L(x, y), where x and y are related data points. If the value of L(x, y) exceeds a reasonable range, reexamine the relevant data or mark it as suspicious data.
[0240] In one embodiment, the data annotation and feature extraction include:
[0241] Annotate the cleaned data. If it is used to generate a specific type of document, it can be manually or automatically annotated according to the theme and style of the document;
[0242] Extract the features of the data, including word frequency features, part-of-speech features, and semantic features; integrate with the natural language processing library using Spark; through the mapPartitions operation of Spark, call the functions of NLTK within each data partition to calculate the features, and then merge the results.
[0243] In one embodiment, the construction of the document generation model includes:
[0244] Use a neural network language model. In the Spark environment, utilize distributed computing to train the model. Divide the large-scale text data into multiple small batches and train the model in parallel on multiple nodes of Spark;
[0245] Adopt the method of transfer learning and fine-tune the pre-trained model on a large dataset to adapt to a specific document generation task.
[0246] In one embodiment, the model optimization and evaluation include:
[0247] Optimize the model on the Spark cluster and use stochastic gradient descent and its variants for parameter updates; the distributed computing ability of Spark can accelerate the optimization process by computing gradients and updating parameters in parallel on multiple nodes.
[0248] Adopt multiple evaluation metrics to evaluate the performance of the model, implement the cross-validation evaluation method on Spark, divide the data into training set, validation set and test set, evaluate the model on different subsets and adjust the hyperparameters of the model.
[0249] In one embodiment, the generation of the initial draft of the document includes:
[0250] Use the trained model to generate the initial draft of the document according to the given input; in the Spark environment, distribute the input data to multiple nodes, each node generates part of the document content according to the model, and then combine these contents; if generating a long document, divide the document into multiple paragraphs, and each node is responsible for generating one paragraph.
[0251] For the generated initial draft, perform preliminary format adjustment by writing custom rules or using natural language processing tools.
[0252] In one embodiment, the polishing and improvement of the document includes:
[0253] Polish the initial draft of the document using additional language processing techniques. In Spark, call external grammar checking tools or use predefined vocabulary replacement rules to process the generated document.
[0254] Further improve the document according to the specific requirements of the document, and adjust the tone of the document and add references or cases in combination with some domain knowledge and user feedback.
[0255] The in-depth understanding of user needs and customized modification of the document includes:
[0256] Perform lexical and syntactic analysis on the instructions input by the user; decompose the instructions into basic semantic units, not only identifying individual words, but also including word groups and phrase structures; for example, for instructions containing specific industry terms, identify the complete semantic scope of the terms, rather than simple word matching.
[0257] Adopt Semantic Role Labeling technology to determine the role of each semantic unit in the entire instruction; such as agent, patient, tool, etc. This helps to understand the relationship between the elements in the instruction.
[0258] Construct a semantic network and connect the parsed semantic units according to their role relationships; this network can better represent the overall intention of the instruction rather than looking at each word or phrase in isolation.
[0259] Build a large-scale knowledge graph that contains information such as professional knowledge in various industries, the internal cultures and styles of different companies, etc.; the nodes of this knowledge graph represent different concepts (such as industry terms, company departments, etc.), and the edges represent the relationships between concepts (such as belonging relationships, process relationships, etc.);
[0260] Match the semantic network obtained by the intention parsing module with the knowledge graph; by calculating the similarity between the nodes in the semantic network and the nodes in the knowledge graph, as well as the similarity between the edges and the edges, find the most matching knowledge fragment. For example, if the user instruction involves the internal style of a specific company, find the style nodes related to that company and their related relationships in the knowledge graph.
[0261] Dynamically adjust the templates and rules for document generation according to the results of the knowledge graph matching; if specific industry professional terms are matched, adjust the word selection and grammar structure to conform to the expression habits of that industry; for the internal style of a specific company, adjust the tone, format, and common expression methods of the document.
[0262] Adopt reinforcement learning technology to adjust the dynamic adaptation strategy according to the matching degree between the generated document and the user's expectations (through user feedback or predefined evaluation metrics); for example, if it is found that the generated document does not meet the user's expectations in terms of the company's internal style, the reinforcement learning algorithm will adjust the parameters related to the company's style in the dynamic adaptation module so that a more compliant document can be generated next time.
[0263] After the document is generated, provide a user feedback mechanism; users can rate the generated document or provide specific modification opinions;
[0264] Take the user feedback as input, re-enter the intention parsing module, and adjust the understanding of the user's intention; according to the new intention understanding, go through the knowledge graph matching and dynamic adaptation module again to regenerate the document until the user is satisfied.
[0265] The formula is expressed as follows:
[0266] Let the user input instruction be I, and the semantic network obtained by the intention parsing module be SN = f 1 (I), where f 1 is the intention understanding function;
[0267] Let the knowledge graph be KG, and the matching degree function be m(SN, KG) = g(SN, KG), where g is the function for calculating the matching degree between the semantic network SN and the knowledge graph KG;
[0268] Let the document generation template be T, and the dynamically adapted template be T’ = h(T, m(SN, KG)), where h is a function for adjusting the template according to the matching degree;
[0269] Let the document generation function be D(T’, I) = d(T’, I), and the finally generated document be Doc = d(T’, I);
[0270] In the feedback loop, let the user feedback be F, and the updated intention be I’ = k(F, I), where k is a function for updating the intention according to the feedback, and then regenerate the document according to the above process;
[0271] The algorithm process is as follows:
[0272] Step 1: Intention parsing:
[0273] Receive the user input instruction I;
[0274] Perform lexical and syntactic analysis on I to obtain basic semantic units;
[0275] Perform semantic role labeling to construct the semantic network SN = f 1 (I);
[0276] Step 2: Knowledge graph matching;
[0277] Construct or load the knowledge graph KG;
[0278] Calculate the matching degree m(SN, KG) = g(SN, KG);
[0279] Step 3: Dynamic adaptation:
[0280] According to the matching degree m(SN, KG), adjust the document generation template T to obtain T’ = h(T, m(SN, KG));
[0281] Step 4: Document generation:
[0282] According to the dynamically adapted template T’ and the user input instruction I, generate the document Doc = d(T’, I);
[0283] Step 5: Feedback loop:
[0284] If the user provides feedback F, update the intention I’ = k(F, I);
[0285] Restart from Step 1 until the user is satisfied.
[0286] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0287] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A document generation system based on a big data model, characterized in that: include: Parameter setting module, used to set the knowledge base, model selection and role setting; Report list viewing module, used to view the title, tags and creation time of the document; You can also search for documents by searching for the time when the document was created and the document tags; Feedback module, used to input feedback on the system; System login module, used by users to enter their account and password to log in; Model management module, used to enable and disable model deployment; Knowledge base management module, which is used to view, close, and delete the knowledge base, and can also search for documents in the knowledge base based on document title, file name, document tag, and creation time; The dialogue management module includes a user management unit and a sensitive word management unit; The user management unit is used to search for users based on user ID, user name, and creation time, and can also view the user's historical conversations; The sensitive word unit is used to search for sensitive words according to sensitive words, tags, whether they are enabled and creation time, and to modify and delete sensitive words; and can also add new sensitive words; Feedback management module, used to manage feedback The system management module is used to manage multiple functional modules.
2. The document generation system based on the big data model according to claim 1, characterized in that: The system management module includes: A tenant management unit, wherein the tenant management unit is provided with a tenant list subunit and a tenant package subunit; The tenant list subunit is provided with a tenant list, which can be searched according to tenant name, contact person, contact phone number, tenant status and creation time, and can also modify tenant information and delete tenants; The tenant package subunit is used to search according to package name, status and modification time, and a package list is set in the tenant package subunit; A user management module is used to manage users; the user management module is provided with an intelligent customer service company port, which is used to manage and allocate users according to the head office, branches and various departments of the company. The user management module is provided with a search port, which can search according to user name, mobile phone number, status and creation time; and can also modify user information; The role management module is used to manage different roles. It can search for roles based on role name, role ID, status, and creation time. Role names include test account, management role, and super administrator. It can modify role information, change permissions, and delete roles. The menu management module is used to manage the menu list and can search the menu according to the menu name and status; the menu names include: Boot development documents, Cloud development documents, model management, knowledge base management, dialogue management, feedback management, system management, payment management, report management, workflow, product center, order center, marketing center and public account management; the menu can also be modified; The department management module is used to manage the various departments of the head office and branches; it can modify, add and delete departments; it can also search by department name and status; The position management module is used to modify, add and delete positions; position information includes: position number, position code, position name, position ranking and status; positions can be searched based on position code, position name and status; Dictionary management module, users can modify, add and delete dictionaries; dictionary information includes: dictionary number, dictionary name, dictionary type, status and creation time; can search based on dictionary name, dictionary type, status and creation time; The notification module is used to modify, add and delete notifications. The notification information includes: the title, type, status and creation time of the notification. It can search based on the title and status of the notification. The operation log module is used to manage the operation log; the log information includes the log number, operation module, operation name, operation type, operator, operation result, operation date and execution time; it can be searched according to the system module, operator, type, status and operation time; Login log module; used to manage login information; login information includes log type, user name, login address, userAgent, result and login date; can search based on login address, user name, status and login time; Application management module, used to manage application clients; client information includes: client number, client key, application name, application icon, status, validity period of access token, validity period of refresh token, authorization type and creation time; can search by application name and status; The token management module is used to manage tokens; token information includes: access token, refresh token, user number, user type, creation time and expiration time; it can search based on user number, client number and user type; The SMS management module is used to manage SMS messages; SMS information includes: SMS signature, channel code, activation status, SMS API account, SMS API key and creation time; it can search based on SMS signature, activation status and creation time.
3. A document generation method of a document generation system based on a big data model according to any one of claims 1-2, characterized in that: include: Integration of data sources; Data source verification; Data annotation and feature extraction; Build a document generation model; Model optimization and evaluation; Generate a first draft of the document; Document polishing and improvement; Deeply understand user needs and customize documents.
4. The document generation method according to claim 3, characterized in that: The data source integration includes: Use Apache Spark's distributed file system interface to collect document-related data from multiple data sources; Identify the types of multiple data sources, including relational databases, non-relational databases, files in the local file system, and files in network storage; Analyze the formats of document-related data in different data sources in detail; Configure the Apache Spark environment and ensure that the Spark cluster is configured correctly; ensure that the necessary dependent libraries are imported into the Spark project, especially the driver libraries related to the data source to be connected; Use Spark's distributed file system interface to collect data. Specifically, In Spark, first create a SparkSession object, which is the entry point for interacting with Spark; Connect to different data sources. For relational databases, use Spark's JDBC interface; for non-relational databases, use the Spark-MongoDB connector; Use Spark's DataFrame operation to integrate data obtained from different data sources; The collected data are preliminarily cleaned to remove noise data and incomplete data records.
5. The document generation method according to claim 3, characterized in that: The data source verification includes: Categorize data sources into primary and secondary sources; For each data point, label its source type; For different types of sources, an initial credibility weight is assigned according to their historical accuracy and authority; it is represented by a matrix C, where C ij represents the credibility weight of the i-th data point coming from the j-th type of source; When multiple sources provide the same or similar data points, calculate the degree of consistency of their data; Assume that a data point x is provided by n sources, and the values of these sources are x1, x2, ···, x n ; Calculate consistency index in If I is greater than a set threshold, the data is considered reliable in terms of consistency; Update the credibility weight of the source based on the results of the data consistency check; If the data consistency is good, the credibility of the source providing the data point increases. Where n is the number of sources providing that data point; If the data consistency is poor, the credibility of the source providing the data point is reduced. As new data is continuously input, the above steps are repeated to continuously update the credibility matrix and verify the accuracy of the data; For sources whose credibility weight is below a lower limit, their data are marked as requiring further manual review or directly discarded; Analyze the semantics and logical relationships of the data, and further verify the data by establishing a logical relationship model to analyze the causal and correlation relationships between the data; Assume that the logical relationship is represented by the function L(x, y), where x and y are related data points. If the value of L(x, y) exceeds a reasonable range, the related data is reviewed or marked as suspicious data.
6. The document generation method according to claim 3, characterized in that: The data annotation and feature extraction include: Label the cleaned data. If it is used to generate a specific type of document, it can be manually or automatically labeled according to the subject and style of the document. Extract data features, including word frequency features, part-of-speech features, and semantic features; use the natural language processing library to integrate with Spark; use Spark's mapPartitions operation to call NLTK functions in each data partition to calculate features, and then merge the results.
7. The document generation method according to claim 3, characterized in that: The document generation model construction includes: Use a neural network language model to train the model using distributed computing in the Spark environment, divide large-scale text data into multiple small batches, and train the model in parallel on multiple Spark nodes; A transfer learning approach is adopted to fine-tune the pre-trained model on a large dataset to adapt it to the specific document generation task.
8. The document generation method according to claim 3, characterized in that: The model optimization and evaluation includes: Optimize the model on a Spark cluster and use stochastic gradient descent and its variants to update parameters; Spark's distributed computing capabilities can speed up the optimization process by computing gradients and updating parameters in parallel on multiple nodes; A variety of evaluation indicators are used to evaluate the performance of the model. The cross-validation evaluation method is implemented on Spark. The data is divided into training set, validation set and test set. The model is evaluated on different subsets and the hyperparameters of the model are adjusted.
9. The document generation method according to claim 3, characterized in that: The generating of the first draft of the document includes: Use the trained model to generate a draft document based on the given input. In the Spark environment, distribute the input data to multiple nodes, each node generates part of the document content based on the model, and then merges the content. If a long document is generated, divide the document into multiple paragraphs, and each node is responsible for generating one paragraph. For the generated draft, preliminary formatting adjustments can be made by writing custom rules or using natural language processing tools.
10. The document generation method according to claim 3, characterized in that: The document polishing and improvement includes: Use additional language processing techniques to polish the first draft of the document. In Spark, call external grammar checkers or use predefined word replacement rules to process the generated document. Further improve the document according to its specific needs, adjust the tone of the document based on some domain knowledge and user feedback, and add citations or cases; The in-depth understanding of user needs and customized modification of documents include: Perform lexical and syntactic analysis on the instructions input by the user; decompose the instructions into basic semantic units, not only recognizing individual words, but also including word groups and phrase structures; The semantic role labeling technology is used to determine the role of each semantic unit in the entire instruction; a semantic network is constructed to connect the parsed semantic units according to their role relationships; Build a large-scale knowledge graph that contains information such as professional knowledge in various industries, the internal culture and style of different companies, etc. The nodes of this knowledge graph represent different concepts, and the edges represent the relationship between concepts. Match the semantic network obtained by the intent parsing module with the knowledge graph; find the most matching knowledge fragment by calculating the similarity between nodes in the semantic network and nodes in the knowledge graph, as well as the similarity between edges; According to the results of knowledge graph matching, the templates and rules for document generation are dynamically adjusted; if the professional terms of a specific industry are matched, the vocabulary selection and grammatical structure are adjusted to conform to the expression habits of the industry; Reinforcement learning technology is used to adjust the dynamic adaptation strategy according to the matching degree between the generated documents and user expectations; After the document is generated, a user feedback mechanism is provided; users can rate the generated document or provide specific modification suggestions; Take user feedback as input, re-enter the intent parsing module, and adjust the understanding of user intent. Based on the new understanding of intent, re-generate the document through the knowledge graph matching and dynamic adaptation modules until the user is satisfied. The formula is as follows: Assume that the user input command is I, and the semantic network obtained by the intention parsing module is SN=f1(I), where f1 is the intention understanding function; Let the knowledge graph be KG, and the matching function be m(SN, KG) = g(SN, KG), where g is the function for calculating the matching degree between the semantic network SN and the knowledge graph KG; Suppose the document generation template is T, and the template after dynamic adaptation is T'=h(T,m(SN,KG)), where h is a function that adjusts the template according to the matching degree; Assume that the document generation function is D(T',I) = d(T',I), and the final generated document is Doc = d(T',I); In the feedback loop, let the user feedback be F, and the updated intent be I'=k(F,I), where k is the function that updates the intent based on the feedback, and then regenerate the document according to the above process; The algorithm flow is as follows: Step 1: Intent analysis: Receive user input command I; Perform lexical and syntactic analysis on I to obtain basic semantic units; Perform semantic role labeling and construct a semantic network SN=f1(I); Step 2: Knowledge graph matching; Build or load knowledge graph KG; Calculate the matching degree m(SN, KG)=g(SN, KG); Step 3: Dynamic Adaptation: According to the matching degree m(SN, KG), adjust the document generation template T to obtain T'=h(T,m(SN,KG)); Step 4: Document generation: Generate document Doc=d(T',I) according to the dynamically adapted template T' and the user input instruction I; Step 5: Feedback loop: If the user provides feedback F, update the intention I' = k(F, I); Repeat step 1 until the user is satisfied.
Citation Information
Patent Citations
Document generation system, method and equipment based on knowledge base and large model and medium
CN117556010A
Order synchronization method of business management platform
CN117725122A
Financial text credibility scoring method based on privacy protection
CN117932614A
Local knowledge base question answering system, method and equipment based on LangChain4J and medium
CN118503388A
BOM data multi-dimensional analysis system and analysis method based on large language model
CN119004372A