Dynamic interactive document generation method and device based on generative AI, computer equipment and computer readable storage medium

An interactive document generation method using multimodal input and dynamic weight adjustment solves the problems of template limitations and lack of interactivity in existing technologies, and achieves efficient and personalized professional document generation.

CN121072503APending Publication Date: 2025-12-05ADVANCED SYST DEV

Patent Information

Application Number
CN202511184011.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing document generation technologies suffer from limitations such as static templates, inefficient prompts, and a lack of interactivity, making it difficult to generate documents that meet users' personalized needs and professional standards.

Method used

Employing multimodal input data fusion, dynamic weight adjustment, and interactive optimization mechanisms, the system generates composite prompt words through a cross-modal semantic alignment model, and combines user historical interaction data and real-time quality detection to achieve dynamic document generation.

Benefits of technology

It enables personalized document generation, reduces the number of manual corrections, improves document generation efficiency and quality, and meets the standard requirements of professional fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121072503A_ABST
    Figure CN121072503A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic interactive document generation method and device based on generative AI, computer equipment and a storage medium. The method comprises the following steps: receiving multi-modal input data; generating a composite cue word by adopting a cross-modal semantic alignment model of a QWen text encoder and a BriVL image encoder; dynamically adjusting the semantic weight of the prompt word based on a weight formula of a knowledge graph matching degree, a user historical acceptance degree and an AI confidence degree; calling the generative AI model to generate document content; a quality detection module identifies questions and actively inquires a user by adopting a layered interaction strategy; correcting the content through an incremental cue word according to the feedback; and optimizing a cue word generation strategy by adopting a PPO reinforcement learning algorithm. According to the method, the problems of static template limitation, low efficiency of cue words and lack of interactivity in traditional document generation are solved, dynamic and interactive personalized document generation is realized, and the document generation quality and efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, and particularly relates to a dynamic interactive document generation method and device based on generative AI, a computer device and a computer readable storage medium. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, automatic document generation based on generative AI has become an important means to improve work efficiency. Currently, more and more enterprises and individuals begin to rely on AI technology to assist in generating various structured documents, including legal documents, business reports, technical solutions, and other professional documents.

[0003] Existing document generation technology is mainly based on preset templates and static prompt word mechanisms. Users input simple text instructions, and the system generates corresponding document content according to fixed template frameworks. This method has certain practicality in processing standardized and relatively fixed format documents, and can produce a basic document draft that meets the requirements in a short time.

[0004] However, with the continuous deepening of actual application needs, the existing technology exposes many limitations. First, the limitation of static templates is increasingly prominent. Traditional document generation systems rely heavily on pre-set fixed templates. These templates can ensure the basic structural integrity of the document, but cannot dynamically adjust the content structure and organization form of the document according to the specific input content and actual needs of the user. When the user needs to generate a document with special structural requirements or personalized content, the fixed template often cannot meet the needs, resulting in a lack of pertinence and applicability of the generated document.

[0005] Secondly, the inefficiency of the prompt word mechanism seriously affects the quality and efficiency of document generation. Existing systems usually use a single text instruction as a prompt word. This simple instruction form is prone to semantic ambiguity, leading to deviation in AI's understanding of the user's true intention. Due to the lack of effective semantic understanding and context association mechanism, the content generated by AI often deviates from the user's actual needs, requiring the user to make multiple manual corrections and adjustments. This not only increases the user's workload, but also greatly reduces the overall efficiency of document generation.

[0006] In addition, the existing technology generally lacks effective interaction mechanisms. During the document generation process, the system cannot provide real-time feedback and quality evaluation, and the user can only discover problems and make modifications after the document is completely generated. This lagging feedback mechanism leads to error accumulation, and when the generated document has logical contradictions, missing key information, or unreasonable structure, etc., it often needs to be reworked significantly, seriously affecting work efficiency. At the same time, the system cannot learn and improve from the user's modification behavior, lacking the ability to continuously optimize.

[0007] In particular, when dealing with highly specialized documents such as legal contracts, technical patent applications, etc., the limitations of existing technologies become more apparent. Such documents not only require accuracy and completeness of content, but also need to comply with specific industry standards and professional requirements. However, existing document generation systems are difficult to effectively integrate professional domain knowledge, and cannot identify and handle domain-specific terminology, rules and constraints, resulting in the quality of generated professional documents being difficult to meet the actual application requirements. SUMMARY

[0008] The present application aims to provide a dynamic interactive document generation method, device, computer equipment and computer readable storage medium based on generative AI, which can overcome the shortcomings of the prior art and realize truly dynamic and interactive document generation to meet the increasingly complex and diverse document generation needs.

[0009] To achieve the above-mentioned purpose, the present application realizes the following technical solutions: A dynamic interactive document generation method based on generative AI, comprising the following steps: S1: receiving multi-modal input data, the multi-modal input data comprising at least two of text data, voice data, image data and video data; S2: processing the multi-modal input data through a cross-modal semantic alignment model to generate a composite prompt word; S3: dynamically adjusting the semantic weight of the composite prompt word based on user historical interaction data, the semantic weight being calculated through a preset weight formula; S4: calling a generative AI model to generate initial document content according to the adjusted composite prompt word; S5: real-time detecting the initial document content through a quality detection module to identify logical contradictions or data omissions; S6: when detecting problems or AI confidence being lower than a preset threshold, actively initiating multiple rounds of dialogue inquiries to clarify user intent; S7: according to user feedback, guiding the generative AI model to correct the document content through incremental prompt word supplementation; S8: collecting user modification traces of the generated document, and optimizing the prompt word generation strategy through reinforcement learning.

[0010] Further: weight = a x term professionalism + b x historical acceptance + c x real-time confidence, wherein a, b and c are preset coefficients, and a + b + c = 1. α , β , γ , α , β , γ , α , β , γ = 1. The term professional degree a is calculated by knowledge graph matching degree, α = ω1 × entity coverage + ω2 × relationship consistency + ω3 × attribute validity; The history acceptance beta is calculated by user history adoption rate, β = λ × (number of accepted content / total number of content) + (1- λ ) × time decay factor; The real-time confidence gamma is calculated by AI model output probability calibration, γ = temperature scaling probability × uncertainty penalty factor.

[0011] Further: the entity coverage = number of matched entities / total number of entities, the relationship consistency is calculated by TransE model to calculate relationship triplets score, and the time decay factor , wherein Δt is the month difference between the current time and the latest review time.

[0012] Further: the cross-modal semantic alignment model comprises: A text encoder that encodes text data using a QWen base model; An image encoder that encodes image data using a BriVL base model; A modal projection layer that maps embedding vectors of different modalities to a unified semantic space of 512-768 dimensions through linear transformation; A cross-modal interaction module that uses cross-attention mechanism to realize bidirectional alignment of text to image and image to text.

[0013] Further: the linear transformation formula of the modal projection layer is: ; ; , wherein , is the projection weight matrix, , is the bias vector.

[0014] Further: the user history interaction data includes modification records, document annotation preferences and term usage frequency, and the dynamic adjustment includes determining document type weight and professional field weight according to the history interaction data.

[0015] Further, the quality detection module comprises a logic consistency detection unit, an integrity detection unit and a professional term detection unit, the professional term detection unit recognizes patent sensitive words based on a pre-constructed domain knowledge graph and provides replacement suggestions.

[0016] Further, the reinforcement learning adopts a PPO algorithm, comprising: Record low-confidence events to the reinforcement learning sample pool to establish negative feedback training data; Design a reward function:

[0017] wherein is a direct feedback reward, is a continuity reward, is a diversity reward; Train the prompt word generation strategy based on the reward function and update the domain knowledge graph node weight.

[0018] Further, the update of the domain knowledge graph node weight comprises: when a certain technical term is corrected multiple times, the semantic weight of its associated concept is increased by a preset proportion; when a frequently occurring ambiguous term is added to the initial prompt word, the corresponding definition requirement is added.

[0019] The application also provides a dynamic interactive document generation device based on generative AI, comprising: A multi-modal input module for receiving multi-modal input data comprising at least two of text data, voice data, image data and video data; A semantic alignment module for processing the multi-modal input data through a cross-modal semantic alignment model to generate a composite prompt word, the cross-modal semantic alignment model comprising a QWen text encoder, a BriVL image encoder and a modal projection layer; A weight adjustment module for dynamically adjusting the semantic weight of the composite prompt word based on user historical interaction data; A document generation module for calling a generative AI model to generate initial document content according to the adjusted composite prompt word; A quality detection module for real-time detection of the initial document content to identify logical contradictions or data omissions; An interactive clarification module for actively initiating a multi-round dialogue inquiry to clarify user intent when detecting problems or AI confidence below a preset threshold; A content correction module for guiding the generative AI model to correct the document content through incremental prompt word supplementation according to user feedback; A strategy optimization module for collecting user modification traces of the generated document and optimizing the prompt word generation strategy through reinforcement learning.

[0020] Further, the weight adjustment module comprises: a knowledge graph matching unit for calculating the professional degree of the term, and comprehensively evaluating through entity coverage, relationship consistency and attribute effectiveness; a historical adoption rate calculation unit for calculating the historical acceptance of the user, and dynamically adjusting the weight in combination with a time decay factor; a confidence calibration unit for calculating the AI generation confidence, and performing probability calibration through temperature scaling and uncertainty penalty.

[0021] Further, the strategy optimization module adopts a PPO reinforcement learning algorithm, and comprises a sample pool management unit, a reward function calculation unit and a knowledge graph updating unit, for continuously optimizing the prompt word generation strategy based on user feedback.

[0022] The application also provides a computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the above-mentioned dynamic interactive document generation method based on generative AI when executing the computer program.

[0023] The application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the above-mentioned dynamic interactive document generation method based on generative AI.

[0024] Compared with the prior art, the application has the following beneficial effects: I. The problem of static template limitation is solved. The application can automatically adjust the document content structure and organization form according to different inputs such as text, voice and image of the user through the multi-modal input fusion and dynamic weight adjustment mechanism, and thus breaks the dependence of the traditional system on fixed templates and realizes truly personalized document generation.

[0025] II. The problem of low efficiency of prompt words is solved. The application establishes a quantitative weight calculation system based on knowledge graph matching degree, historical acceptance of the user and AI confidence, converts single text instructions into composite prompt words, significantly reduces the situation that the AI generated content deviates from the user's demand, and greatly reduces the number of manual corrections.

[0026] III. The problem of lack of interactivity is solved. The application can discover problems and seek user confirmation in real time during the document generation process through real-time quality detection and active dialogue inquiry mechanism, changes the lag feedback mode of the traditional system that problems can only be discovered after the generation is completed, and realizes real-time optimization and error prevention during the generation process. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1A flowchart schematic diagram of the method for generating a dynamic interactive document based on generative AI according to an embodiment of the present application; Figure 2 A module structure schematic diagram of the method for generating a dynamic interactive document based on generative AI according to an embodiment of the present application; Figure 3 A hardware structure schematic diagram of the computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0029] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0030] The present application provides a method for generating a dynamic interactive document based on generative AI, which realizes high-quality personalized automatic document generation through multi-modal input fusion, dynamic weight adjustment and interactive optimization and other technical means.

[0031] In step S1, the system receives multi-modal input data. The multi-modal input data can include any two or more combinations of text data, voice data, image data and video data. The text data can be a user's direct input of a literal description, an existing document fragment or a structured data table; the voice data can be a user's voice instruction, a conference recording or a business negotiation record; the image data can be a technical sketch, a flowchart, a schematic diagram or a picture containing text information; the video data can be a demonstration video containing voice explanation or an operation guidance video. The system preprocesses these different formats of input data, including format conversion, noise removal and data standardization and other operations.

[0032] In step S2, the system processes the multi-modal input data through a cross-modal semantic alignment model. The model uses QWen as a text encoder, which can effectively understand and encode the semantic information of Chinese text. The image encoder uses the BriVL model, which is optimized for Chinese image-text understanding tasks, and can accurately extract visual features and text information from images. The modal projection layer maps the embedding vectors of different modalities to a unified semantic space of 512-768 dimensions through linear transformation. The specific transformation formula is: ; ; where , is the projection weight matrix, , is the bias vector.

[0033] The cross-modal interaction module uses a cross-attention mechanism to calculate the attention weights of text features on image features and the attention weights of image features on text features, achieving bidirectional semantic alignment from text to image and from image to text. In this way, the system can deeply understand the semantic association between different modalities of information and generate more accurate and rich composite prompt words.

[0034] In step S3, the system dynamically adjusts the semantic weight of the composite prompt word based on user historical interaction data. The semantic weight is calculated through a pre-set weight formula: Weight = α × Term Professionalism + β × Historical Acceptance + γ × Real-time Confidence where Term Professionalism a is calculated by knowledge graph matching degree, specifically: α = ω1 × Entity Coverage + ω2 × Relationship Consistency + ω3 × Attribute Validity.

[0035] Entity coverage is equal to the ratio of the number of matched entities to the total number of entities, measuring the coverage of professional terms in generated content. Relationship consistency calculates the relationship triple score through the TransE model to evaluate the rationality of the relationship between entities. The calculation formula is: , where h, r, t are the embedding vectors of the head entity, relationship, and tail entity, respectively.

[0036] Attribute validity checks whether numeric attributes meet value range constraints and whether enumeration type attributes exist in the pre-set list. Historical acceptance β is calculated by combining user historical adoption rate and time decay factor: β = λ x (number of accepted content / total number of content) + (1- λ ) x time decay factor, where the time decay factor , Δt is the difference in months between the current time and the most recent review time, ensuring that recent user feedback has a higher weight. Real-time confidence γ is calculated by AI model output probability calibration, using temperature scaling method to adjust the confidence of predicted probability, and combining Monte Carlo Dropout sampling to calculate the prediction variance as an uncertainty penalty factor.

[0037] In step S4, the system calls the generative AI model to generate initial document content based on the adjusted composite prompt. The generative AI model can choose GPT, Claude or other large language models, and be selected and configured according to the specific needs of the application scenario. The system inputs the weight-adjusted composite prompt to guide the AI model to generate document content that meets the user's needs.

[0038] In step S5, the system detects the initial document content in real time through the quality detection module. The quality detection module includes a logical consistency detection unit, an integrity detection unit and a professional term detection unit. The logical consistency detection unit identifies logical contradictions in the document, such as time sequence errors, causal relationship conflicts, etc., through rule engines and semantic analysis algorithms. The integrity detection unit checks whether the document contains necessary structural components, such as whether the contract document contains party information, rights and obligations clauses, breach of contract responsibilities and other key elements. The professional term detection unit, based on a pre-built domain knowledge graph, identifies whether the use of patent-sensitive words and industry-specific terms is appropriate, and provides replacement suggestions to expand the scope of protection or improve expression accuracy.

[0039] In step S6, when the quality detection finds problems or the AI confidence is below the preset threshold, the system initiates multiple rounds of dialogue inquiries to clarify the user's intention. The system adopts a hierarchical threshold strategy for interactive decision-making: when the AI confidence is in the range of 0.65-0.75, the system triggers basic interactive inquiry mechanism; when the confidence drops to the range of 0.55-0.65, the system uses a two-choice question to prompt, providing clear selection items for the user to simplify the decision-making process; when the confidence is below 0.55, the system provides multiple options and gives open-ended guidance, allowing users to express and supplement more flexibly. The system selects the appropriate inquiry method from the pre-set question template library according to the detected specific problem type and confidence level, and adjusts the question expression according to the user's historical preferences and current context. The dialogue management module maintains the dialogue state, tracks the asked questions and the user's answers, avoids repeated inquiries and ensures the coherence of the dialogue.

[0040] In step S7, the system corrects the document content by incrementally supplementing the prompt words based on user feedback. When the user provides explicit modification suggestions or selects specific solutions, the system converts these feedback into incremental prompt words, which are appended to the original prompt words to guide the AI model to make targeted content corrections. This incremental prompt word supplementing method maintains the coherence of the original content and accurately corrects specific problems, avoiding changes to other content that may occur when the entire document is regenerated.

[0041] In step S8, the system collects user modification traces of the generated document and optimizes the prompt word generation strategy through reinforcement learning. The system uses the PPO algorithm for reinforcement learning, taking user modification behavior and satisfaction evaluation as reward signals. The reward function is designed as:

[0042] wherein is the direct feedback reward, based on the user's acceptance or modification of the generated content; is the coherence reward, evaluating the logical coherence and structural integrity of the generated content; is the diversity reward, encouraging the system to generate diverse expression methods.

[0043] The system records low-confidence events into the reinforcement learning sample pool, establishes negative feedback training data, and updates the parameters of the prompt word generation strategy through the policy gradient method. At the same time, the system also updates the domain knowledge graph node weights. When a technical term is modified multiple times, the semantic weight of its associated concepts is increased by a preset percentage, usually 20%. When ambiguous terms frequently appear, the corresponding definition requirements are automatically added to the initial prompt words in the subsequent generation.

[0044] Taking the contract generation scenario as an example, the user uploads the business negotiation recording and the technical specification document. The system first performs noise reduction processing and automatic speech recognition on the voice data to extract key business clause information; at the same time, it analyzes the technical specification document to identify structured information such as technical parameters and performance indicators. The cross-modal semantic alignment model correlates and matches the business requirements in the voice with the technical requirements in the document to generate an initial prompt word: “Party A’s technical delivery standard needs to include [performance parameter] [test method], the delivery time is [specific date], and the payment method is [payment condition]”. In the dynamic weight adjustment stage, the system finds that the term “test method” field has a low degree of professionalism, and users often modify such clauses historically, so the weight of this part is reduced. The quality detection module finds that “test method” does not specify specific standards or processes, with a confidence level of 0.65, which is lower than the preset threshold of 0.75, triggering interactive inquiry: “It is detected that the technical delivery standard lacks specific test methods. Do you need to reference ISO9001 standards, or do you want to use other specific test specifications?” After the user selects to reference ISO9001 standards, the system supplements through incremental prompt words: “The test method of the technical delivery standard should meet the requirements of the ISO9001 quality management system, including product inspection, process monitoring and continuous improvement, etc.”, and finally generates a contract draft containing complete clauses.

[0045] In the technical document writing scenario, the user inputs a technical sketch and a technical idea summary. The image encoder identifies the module structure, connection relationship and label information in the sketch, and the text encoder understands the core concept and implementation path of the technical idea. The system generates a technical scheme framework: “Based on the communication interface between [data processing unit] and [control module], through [signal modulation technology], the technical effect of [data transmission optimization] is achieved”. The professional term detection unit finds that “data processing unit” is described relatively broadly in the patent context, and suggests replacing it with “digital signal processor” to clarify the technical implementation method, and the knowledge graph detection shows that “control module” is a potential patent sensitive word, and suggests replacing it with “logic control unit” to expand the protection range. The system finally outputs a technical design document that meets the technical specification and patent application requirements.

[0046] In another embodiment, the present application also provides a dynamic interactive document generation device based on generative AI. The device includes a plurality of functional modules, each module working together to realize intelligent document generation function.

[0047] The multi-modal input module is responsible for receiving and preprocessing various input data of the user. The module is configured with a text input interface, a voice collection interface, an image upload interface, and a video processing interface, and can simultaneously process data in different formats such as text, voice, image, and video. The module has a built-in data format converter and preprocessor, which can reduce noise and standardize the format of voice data, adjust the size and optimize the quality of image data, and extract frames and separate audio and video of video data, to ensure that the subsequent processing modules can effectively process these multi-modal data.

[0048] The semantic alignment module is the core processing unit of the device, which integrates a cross-modal semantic alignment model. The module has a built-in QWen text encoder and BriVL image encoder as basic feature extraction components, which map the feature vectors of different modalities to a 512-768 dimensional semantic space through a modal projection layer. The cross-modal interaction submodule uses a cross-attention mechanism to calculate the mutual attention between text features and image features, achieving deep semantic alignment. The module outputs a composite prompt word after fusion processing, providing high-quality semantic guidance for subsequent document generation.

[0049] The weight adjustment module is responsible for dynamically optimizing the weight of the prompt word according to the user's historical data. The module includes three core components: a knowledge graph matching unit, a historical adoption rate calculation unit, and a confidence calibration unit. The knowledge graph matching unit is connected to a pre-built domain knowledge base, which calculates the term professional degree through entity recognition, relationship verification, and attribute inspection. The historical adoption rate calculation unit maintains a user behavior database, records the user's document use, modification, and acceptance, and dynamically calculates the historical acceptance rate combined with the time decay algorithm. The confidence calibration unit uses temperature scaling and Monte Carlo sampling techniques to calibrate the output probability of the AI model, providing reliable confidence evaluation.

[0050] The document generation module serves as an execution unit, which calls the configured generative AI model to generate document content according to the adjusted composite prompt word. The module supports multiple mainstream large language model interfaces, and can flexibly configure model types and parameters according to application requirements. The module has a content formatter built-in, which can structure the generated content according to the format requirements of different document types.

[0051] The quality detection module performs multi-dimensional quality evaluation on the generated document content. The logical consistency detection unit identifies logical contradictions and timing errors in the content through a rule engine and semantic analysis algorithm. The integrity detection unit checks the integrity of necessary elements based on the document type template. The professional term detection unit connects the domain knowledge graph to verify the accuracy of the use of professional terms and provide optimization suggestions.

[0052] The interactive clarification module implements the intelligent conversation function with the user. According to the quality detection result and the confidence level, the module adopts a hierarchical interaction strategy to actively ask the user. When the confidence is in the range of 0.65-0.75, the basic inquiry is triggered, when the confidence is in the range of 0.55-0.65, the two-choice question is provided, and when the confidence is below 0.55, the multiple-choice and open-ended guidance are provided. The module has a built-in conversation state manager, which maintains the context information of multiple rounds of conversation, and ensures the coherence and effectiveness of the interaction.

[0053] The content revision module receives user feedback and performs targeted content adjustment. The module converts the user's modification suggestions into incremental prompt words to guide the generative AI model to make accurate content revisions. The module adopts a local update strategy to minimize the impact on other parts of the document and maintain the stability of the overall content.

[0054] The strategy optimization module implements the continuous learning and improvement function of the system. The module adopts the PPO reinforcement learning algorithm, including a sample pool management unit, a reward function calculation unit and a knowledge graph update unit. The sample pool management unit collects and manages user interaction data to establish a training sample library. The reward function calculation unit calculates the comprehensive reward value according to the user feedback, including direct feedback, coherence and diversity. The knowledge graph update unit dynamically adjusts the knowledge node weight according to the learning result to realize the continuous optimization of the system knowledge base.

[0055] The device can be deployed on a cloud server or a local workstation and provide document generation services for users through a network interface. The device supports multi-user concurrent access, has good scalability and stability, and can meet the performance requirements of enterprise-level applications.

[0056] In another embodiment, the present application also provides a computer device, comprising a memory 20, a processor 10 and a computer program stored on the memory and executable on the processor, the memory 20 and the processor 10 are connected through a communication interface 30, and the processor implements the above-mentioned generative AI-based dynamic interactive document generation method when executing the computer program.

[0057] In another embodiment, the present application also provides a computer readable storage medium having a computer program stored thereon, the computer program is executed by a processor to implement the above-mentioned generative AI-based dynamic interactive document generation method.

[0058] The technical scheme of the present application fully considers the special needs of different application scenarios, and realizes good scalability and adaptability through modular design. The system can adjust the parameter configuration of each module according to the specific deployment environment and user demand, and realize the optimized application in the generation of different types of documents such as legal documents, business reports and technical schemes.

[0059] The above embodiments are only for illustrating the technical concept and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and implement it, and cannot limit the protection scope of the present application. Any equivalent transformation or modification according to the spirit and essence of the present application should be covered within the protection scope of the present application.

Claims

1. A method for dynamic interactive document generation based on generative AI, characterized in that, The method comprises the following steps: S1: receiving multi-modal input data, the multi-modal input data comprising at least two of text data, voice data, image data and video data; S2: processing the multi-modal input data through a cross-modal semantic alignment model to generate a composite prompt word; S3: dynamically adjusting the semantic weight of the composite prompt word based on user historical interaction data, the semantic weight being calculated through a preset weight formula; S4: calling a generative AI model to generate initial document content according to the adjusted composite prompt word; S5: detecting the initial document content in real time through a quality detection module to identify logical contradictions or data omissions; S6: when a problem is detected or the AI confidence is lower than a preset threshold, actively initiating a multi-round dialogue inquiry to clarify the user's intention; S7: according to user feedback, guiding the generative AI model to correct the document content through incremental prompt word supplementation; S8: collecting user modification traces of the generated document and optimizing the prompt word generation strategy through reinforcement learning.

2. The dynamic interactive document generation method based on generative AI according to claim 1, characterized in that, weight = α × term professionalism + β × historical acceptance + γ × real-time confidence, where α , β , γ is a preset coefficient, and α + β + γ = 1;​​​​​ The term professional degree a is calculated by knowledge graph matching degree, α = ω1 × entity coverage + ω2 × relationship consistency + ω3 × attribute validity; The history acceptance β is calculated by a user history adoption rate, β = λ × (number of accepted contents / total number of contents) + (1- λ ) × time decay factor; The real-time confidence γ is calculated by AI model output probability calibration, γ = temperature scaled probability x uncertainty penalty factor.

3. The method of claim 2, wherein, The entity coverage = the number of matched entities / the total number of entities, the relationship consistency is calculated by the TransE model relationship triple score, the time decay factor wherein Δt is the month difference between the current time and the last review time.

4. The method of claim 1, wherein, The cross-modal semantic alignment model comprises: a text encoder that encodes text data using a QWen base model; an image encoder that encodes image data using a BriVL base model; a modal projection layer that maps embedding vectors of different modalities to a unified semantic space of 512-768 dimensions through linear transformation; a cross-modal interaction module that uses a cross-attention mechanism to achieve bidirectional alignment of text to image and image to text.

5. The method of claim 4, wherein, The linear transformation formula of the modal projection layer is: ; ; wherein , is a projection weight matrix, , is a bias vector.

6. The method of claim 1, wherein, The user historical interaction data includes modification records, document annotation preferences and term usage frequency, and the dynamic adjustment includes determining document type weights and professional field weights based on the historical interaction data.

7. The method of claim 1, wherein, The quality detection module includes a logical consistency detection unit, an integrity detection unit and a professional term detection unit, the professional term detection unit identifies patent sensitive terms based on a pre-built domain knowledge graph and provides replacement suggestions.

8. The method of claim 1, wherein, The reinforcement learning uses the PPO algorithm, which includes: record low confidence events to the reinforcement learning sample pool to establish negative feedback training data; design a reward function: wherein is a direct feedback reward, is a coherence reward, is a diversity reward; train the prompt word generation strategy based on the reward function and update the node weights of the domain knowledge graph.

9. The method of claim 8, wherein, The update of the node weights of the domain knowledge graph includes: when a certain technical term is corrected multiple times, the semantic weight of its associated concept is increased by a preset proportion; when a frequently occurring ambiguous term appears, the corresponding definition requirement is added to the initial prompt word.

10. A dynamic interactive document generation apparatus based on generative AI, characterized by, The method comprises: a multi-modal input module for receiving multi-modal input data comprising at least two of text data, voice data, image data and video data; a semantic alignment module for processing the multi-modal input data through a cross-modal semantic alignment model to generate a composite prompt word, the cross-modal semantic alignment model comprising a QWen text encoder, a BriVL image encoder and a modal projection layer; a weight adjustment module for dynamically adjusting the semantic weight of the composite prompt word based on user historical interaction data. The document generation module is configured to call a generative AI model and generate initial document content according to the adjusted composite prompt word; The quality detection module is configured to detect the initial document content in real time and identify logical contradictions or data omissions; The interactive clarification module is configured to actively initiate a multi-round dialogue inquiry to clarify the user's intention when detecting a problem or an AI confidence level below a preset threshold; The content correction module is configured to correct the document content by incrementally supplementing the prompt word based on user feedback and guiding the generative AI model; The strategy optimization module is configured to collect user modification traces of the generated document and optimize the prompt word generation strategy through reinforcement learning.

11. The apparatus of claim 10, wherein, The weight adjustment module includes: A knowledge graph matching unit is configured to calculate the professional degree of a term by comprehensively evaluating entity coverage, relationship consistency, and attribute effectiveness; A historical adoption rate calculation unit is configured to calculate user historical acceptance and dynamically adjust the weight by combining a time decay factor; A confidence calibration unit is configured to calculate the AI generation confidence and perform probability calibration through temperature scaling and uncertainty penalty.

12. The apparatus of claim 10, wherein, The strategy optimization module adopts a PPO reinforcement learning algorithm, including a sample pool management unit, a reward function calculation unit, and a knowledge graph update unit, for continuously optimizing the prompt word generation strategy based on user feedback.

13. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the generative AI-based dynamic interactive document generation method of any one of claims 1-9.

14. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the generative AI-based dynamic interactive document generation method of any one of claims 1-9.

Citation Information

Patent Citations

  • Figure graph optimization method based on human feedback reinforcement learning

    CN117593628A

  • Multi-round dialogue processing method based on RAG technology dynamic cue word retrieval

    CN119538918A

  • Tibetan news generation method, system and equipment and storage medium

    CN120068811A

Cited By

  • Forging field cue word generation method based on multi-modal perception and template splicing

    CN121562578A

  • Document generation method and system based on self-evolution intelligent agent

    CN121766288A