On-duty document generation system and method based on large reasoning model
By designing a duty document generation system, using intention acquisition, knowledge base and interactive generation modules combined with user feedback mechanism, the problem of existing technology that cannot generate personalized and industry-specific documents is solved, and efficient and personalized document generation is achieved to meet the needs of special industries.
Patent Information
- Application Number
- CN202510785508.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing document generation methods based on large inference models cannot meet the customization needs of individuality and industry characteristics. Especially in special industries such as medical care, finance, and law, general templates are difficult to generate high-quality documents that meet industry standards and dynamic requirements.
A duty document generation system was designed, which includes an intention acquisition module, a knowledge base module, an interactive generation module, and a style conversion module. The intention acquisition module obtains task requirement text, the knowledge base module stores and retrieves business knowledge text, and the interactive generation module generates and infers prompt words. Combined with the user feedback mechanism, personalized duty documents are generated, and the style of the documents is adjusted through the style conversion module.
It improves the richness and completeness of duty documents, can generate personalized documents that conform to industry characteristics, meet users' diverse customization needs, and improve the efficiency and quality of document generation.
Smart Images

Figure CN120654689A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text data processing, and specifically relates to a system and method for generating duty documents based on a large inference model. Background Art
[0002] In recent years, with the rapid development of science and technology, deep learning and natural language processing (NLP) technologies have also made remarkable progress. Against this backdrop, large inference models such as GPT-4o and Qwen 2.5 have emerged in the field of text generation, and their capabilities have been significantly improved. These large inference models, supported by powerful algorithms and massive amounts of data, are able to produce high-quality results in a variety of text generation tasks, including news reports, fiction writing, and advertising copy, thereby greatly improving the efficiency and quality of writing. At the same time, they also have the ability to accurately understand structured instructions and can automatically fill in the corresponding content according to preset templates. This undoubtedly lays a solid technical foundation for the generation of standardized documents, making large inference models a highly promising application direction in the field of document generation.
[0003] In today's society, paperwork is an essential and important part of daily operations for businesses and institutions alike, yet it often consumes a significant amount of time and energy. From writing reports to compiling meeting minutes, from the timely release of press releases to the meticulous crafting of academic papers, each step requires the support of high-quality text output. Traditional manual writing methods are not only relatively inefficient but also susceptible to interference from factors such as the writer's personal emotions and expertise, resulting in inconsistent text quality. Furthermore, the explosive growth of information today poses unprecedented challenges to the accuracy and timeliness of written content.
[0004] Duty daily reports, a type of daily paperwork, are specifically documents completed by on-duty personnel during their shifts to record the day's work progress, incident handling, and important information. These reports are typically completed as soon as possible after their shift ends to ensure the timeliness and accuracy of the information. This provides the latest reference for subsequent work and serves as a crucial vehicle for information transfer, work handover, and accountability.
[0005] However, for daily clerical work, many frontline staff members lack standard handwriting standards, generating a large number of personal daily reports. Compiling these reports is time-consuming and laborious, and failing to compile them can easily miss potential risks. Against this backdrop, the document generation capabilities of the inference big model are timely and timely, providing a new and highly promising solution for improving the efficiency and quality of clerical work. This solution is expected to revolutionize traditional clerical work and significantly boost the efficient operations of businesses and institutions.
[0006] The existing document generation method based on the inference big model first extracts key information from the original data, then uses the inference big model to process the key information to obtain semantically coherent text content, and finally uses the standard document template to accurately fill in or reasonably splice the generated text content to finally obtain the document material. However, as the integration of the inference big model with daily business becomes increasingly close, the demand for business understanding continues to deepen, and the daily scenarios that can be adapted by writing with a limited number of templates are gradually decreasing, and can no longer meet the increasingly diverse customization needs of users. In addition, for some companies with special industry backgrounds or professional field needs, such as medical, financial, and legal, the general template writing method cannot fully take into account the characteristics and needs of the industry. It is relatively rigid and difficult to generate high-quality documents that meet industry standards and dynamic requirements. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a system and method for generating duty documents based on a large inference model, which can generate personalized duty documents that meet industry characteristics.
[0008] In order to solve the above technical problems, the first aspect of the embodiment of the present invention discloses a duty document generation system, which includes an intention acquisition module, a knowledge base module, an interaction generation module and a style conversion module;
[0009] The intention acquisition module is data-connected with the knowledge base module and the interaction generation module, and is used to acquire the task requirement text, the extension flag and the reservation flag;
[0010] The knowledge base module is connected to the interactive generation module for storing preset business knowledge texts and performing retrieval processing based on external input texts to be retrieved to obtain retrieval results; the retrieval results include a reference text set and a reorganized text set;
[0011] The interactive generation module is data-connected to the style conversion module, and is used to store a preset expert template set, and to generate and infer prompt words based on the task requirement text, the extension flag, the retention flag, and the search results to obtain a target duty document; the expert template set includes N expert templates, where N is an integer greater than 1;
[0012] The style conversion module is used to perform style conversion processing on the initial duty document based on a preset style template text to obtain a target duty document.
[0013] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the knowledge base module includes a text segmentation unit, a vector index construction unit, a reference counting unit, a similarity calculation unit, a retrieval unit and a text reorganization unit;
[0014] The text block unit is data-connected to the vector index building unit and the retrieval unit, and is used to block the business knowledge text to obtain a text block set; the text block set includes P text blocks; P is an integer greater than 1;
[0015] The vector index construction unit is data-connected to the similarity calculation unit and the text reorganization unit, and is used to process the task requirement text and the text block set to obtain a task vector and an index vector set, and in response to receiving a text to be retrieved from an external input, process the text to be retrieved to obtain a vector to be retrieved; the index vector set includes an index vector corresponding to each text block;
[0016] The reference counting unit is data-connected to the similarity calculation unit and the text reorganization unit, and is used to store a reference count set and update the reference count set based on the reorganized text set from the text reorganization unit; the reference count set includes the reference count corresponding to each of the text blocks;
[0017] The similarity calculation unit is data-connected to the retrieval unit and is used to process the vector to be retrieved, the index vector set, and the reference count set using a similarity calculation model to obtain a similarity set; the similarity set includes P similarity values;
[0018] The retrieval unit is data-connected to the text reorganization unit and is used to generate a reference text set and a non-reference text set based on the similarity set and the text block set; the reference text set and the non-reference text set include Q1 and Q2 text blocks, respectively; Q1+Q2<=P, and Q1 and Q2 are both integers greater than 1;
[0019] The text reorganization unit is used to process the reference text set, the non-reference text set, the task vector and the vector to be retrieved to obtain the reorganized text set.
[0020] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the text reorganization unit includes a relevance calculation subunit, a screening subunit, and a merging subunit:
[0021] The relevance calculation subunit is data-connected to the screening subunit and is used to process the non-reference text set, the index vector set, the vector to be retrieved, and the task vector using a relevance calculation model to obtain a relevance set; the relevance set includes Q2 relevance values;
[0022] The screening subunit is data-connected to the merging subunit, and is used to sort the non-reference text set based on the relevance set to obtain a sorted text set, and delete the last Q3 text blocks of the sorted text set to obtain a supplemented text set; Q3 is an integer greater than 1 and less than Q2;
[0023] The merging subunit is used to calculate the union of the reference text set and the supplementary text set to obtain a recombined text set.
[0024] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the expression of the correlation calculation model is:
[0025] CD j =cos β (VI j ,VC)(1-cos(VI j ,VA)) β
[0026] In the formula, CD j is the jth said association value in said association set CD; VI j is the j-th text block in the non-reference text set I, and the corresponding index vector in the index vector set; VA and VC are the vector to be retrieved and the task vector respectively; β is the preset correlation coefficient; j is an integer from 1 to Q2.
[0027] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the interactive generation module includes an expert template storage unit, a first prompt word generation unit, a second prompt word generation unit, a third prompt word generation unit, a fourth prompt word generation unit, a fifth prompt word generation unit, and an inference unit;
[0028] The expert template storage unit is used to store a preset expert template set;
[0029] The first prompt word generation unit is data-connected to the intention acquisition module, the knowledge base module, and the reasoning unit, and is used to perform a first combination process on the task requirement text and the reference text set to obtain an initial prompt word;
[0030] The second prompt word generating unit is data-connected to the reasoning unit and is used to perform a second combination process on the externally input current question text, current expert template and current text block set to obtain the question extension prompt word;
[0031] The third prompt word generation unit is data-connected to the intention acquisition module and the reasoning unit, and is used to perform a third combination process on the task requirement text and the externally input current expert template to obtain the associated question prompt word;
[0032] The fourth prompt word generating unit is data-connected to the reasoning unit and is used to perform a fourth combination process on the externally input current question text, the current expert template, and the current text block set to obtain a current prompt word;
[0033] The fifth prompt word generating unit is data-connected to the inference unit and is used to perform a fifth combination process on the externally input current answer text and the current on-duty clerk to obtain the document update prompt word;
[0034] The reasoning unit is used to process the document construction prompt words to obtain the corresponding answer text; the document construction prompt words are the initial prompt words, the question expansion prompt words, the related question prompt words, the current prompt words or the document update prompt words.
[0035] A second aspect of an embodiment of the present invention discloses a method for generating a duty document, the method comprising:
[0036] S1. Use the intent acquisition module to obtain the task requirement text;
[0037] S2. Using the knowledge base module to store preset business knowledge texts;
[0038] S3. Using the interactive generation module and the knowledge base module, construct a current duty document based on the task requirement text;
[0039] S4. Using the intention acquisition module, the interaction generation module, and the knowledge base module, the current duty document is updated to obtain an updated current duty document.
[0040] S5, repeat S4 until receiving the end instruction input by the user;
[0041] S6. Determine the initial duty clerk as the current duty clerk;
[0042] S7. Using a style conversion module, based on a preset style template text, perform style conversion on the initial duty document to obtain a target duty document.
[0043] As an optional implementation, in the second aspect of the embodiment of the present invention, the interactive generation module and the knowledge base module are used to construct the current duty document based on the task requirement text, including:
[0044] S31, using the knowledge base module to perform search processing on the task requirement text to obtain a corresponding reference text set and a reorganized text set;
[0045] S32, using a first prompt word generating unit, performing a first combination process on the task requirement text and the reference text set to obtain an initial prompt word;
[0046] S33: Using the inference unit, the initial prompt word is processed to obtain the current duty clerk.
[0047] As an optional implementation, in the second aspect of the embodiment of the present invention, the updating of the current duty document by using the intention acquisition module, the interaction generation module, and the knowledge base module to obtain the updated current duty document includes:
[0048] S41, initializing the current text block set to empty; initializing the number of loops l to 1; initializing the current question text to the task requirement text;
[0049] S42, setting the current expert template as the lth expert template in the expert template set;
[0050] S43, using the intention acquisition module and the reasoning unit, based on the current expert template, the current text block set and the task requirement text, updating the current question text;
[0051] S44, using the knowledge base module to perform search processing on the current question text to obtain a corresponding reference text set and a reorganized text set; updating the current text block set to the reorganized text set;
[0052] S45, using the reasoning unit, the fourth prompt word generation unit, and the fifth prompt word generation unit to update the current duty clerk based on the current expert template, the current text block set, and the current question text;
[0053] S46, add 1 to the value of l;
[0054] S47. Repeat S42 to S46 until l is greater than N.
[0055] As an optional implementation, in the second aspect of the embodiment of the present invention, the updating of the current question text based on the current expert template, the current text block set, and the task requirement text using the intention acquisition module and the reasoning unit includes:
[0056] S431, using the intention acquisition module to acquire the extension flag;
[0057] S432, judging the extension flag;
[0058] When the extension flag is yes, execute S433;
[0059] When the extension flag is no, executing S435;
[0060] S433: Using a second prompt word generating unit, performing a second combination process on the current question text, the current expert template, and the current text block set to obtain a question extension prompt word;
[0061] S434: Using the reasoning unit, process the question extension prompt word to obtain an extended question text; update the current question text to the extended question text; and execute S44;
[0062] S435: Using a third prompt word generating unit, performing a third combination process on the task requirement text and the current expert template to obtain a related question prompt word;
[0063] S436: Utilize the inference unit to process the associated question prompt words to obtain an associated question text; and update the current question text to the associated question text.
[0064] As an optional implementation, in the second aspect of the embodiments of the present invention, using the reasoning unit, the fourth prompt word generation unit, and the fifth prompt word generation unit to update the current duty clerk based on the current expert template, the current text block set, and the current question text includes:
[0065] S451: Using a fourth prompt word generating unit, performing a fourth combination process on the current question text, the current expert template, and the current text block set to obtain a current prompt word;
[0066] S452: Using the inference unit, process the current prompt word to obtain a current answer text;
[0067] S453, using the intention acquisition module to acquire a reservation flag;
[0068] S454, judging the reserved flag;
[0069] When the reserved flag is yes, execute S455;
[0070] When the retention flag is no, the current duty document remains unchanged; execute S46;
[0071] S455: Using a fifth prompt word generating unit, perform a fifth combination process on the current answer text and the current clerk on duty to obtain a clerk update prompt word;
[0072] S456. Utilize the inference unit to process the document update prompt word to obtain an intermediate text; and replace the content of the current duty document with the intermediate text.
[0073] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0074] By introducing a user feedback mechanism in multiple rounds of conversations and dynamically generating and updating duty documents, the richness and completeness of the generated duty documents can be improved, and the user's deep understanding of the industry can be fully utilized to generate personalized duty documents that conform to industry characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0076] Figure 1 The present invention is a structural diagram of a duty document generation system disclosed in an embodiment of the present invention.
[0077] Figure 2 The present invention is a schematic diagram of the structure of a knowledge base module of a duty document generation system disclosed in an embodiment of the present invention.
[0078] Figure 3 The present invention is a structural diagram of a text reorganization unit of a duty document generation system disclosed in an embodiment of the present invention.
[0079] Figure 4 The present invention is a schematic diagram of the structure of an interactive generation module of a duty document generation system disclosed in an embodiment of the present invention.
[0080] Figure 5 The present invention is a flowchart of a method for generating duty documents disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0081] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0082] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0083] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0084] Example 1
[0085] See also Figure 1-4 . Figure 1 This is a structural diagram of a duty document generation system disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a knowledge base module of a duty document generation system disclosed in an embodiment of the present invention; Figure 3 This is a structural diagram of a text reorganization unit of a duty document generation system disclosed in an embodiment of the present invention; Figure 4 This is a structural diagram of an interactive generation module of a duty document generation system disclosed in an embodiment of the present invention. Figure 1 The duty document generation system described is applied to the field of text data processing, such as duty document generation, and the embodiments of the present invention do not limit this. Figure 1 As shown, the system includes an intention acquisition module, a knowledge base module, an interaction generation module and a style conversion module.
[0086] The intention acquisition module is connected to the knowledge base module and the interaction generation module data to obtain the task requirement text, extension flag and reservation flag.
[0087] It should be noted that the above-mentioned intention acquisition module is used to obtain the task requirement text, expansion flag and retention flag input by the user, among which: (1) The task requirement text is used to describe the user's requirements for the duty document to be generated in text form, such as "Generate a duty document, the content of which is a list of priority targets." (2) The value of the expansion flag is yes or no. When the value is yes, it means that the question asked by the previous expert needs to be expanded to obtain an expanded question text suitable for the current expert's identity, and the current expert needs to be asked based on the expanded question text; when the value is no, it means that the current expert needs to generate a related question text based on the overall task requirement text, and the current expert needs to be asked using the related question text to avoid topic repetition or narrowing of content after multiple rounds of dialogue. (3) The value of the retention flag is yes or no. When the value is yes, it means that the user is satisfied with the content of the current expert's answer and it needs to be merged into the current duty document; when the value is no, it means that the user is not satisfied with the content of the current expert's answer and there is no need to change the current duty document.
[0088] It can be seen that through the expansion flag and the retention flag, the user decides whether to expand the topic and retain the content. This can introduce a user feedback mechanism in the generation process of duty documents, improve user participation, and avoid topic repetition or narrowing after multiple rounds of conversations, thereby improving the richness and completeness of the generated duty documents. It can also make full use of the user's in-depth understanding of the industry to generate personalized duty documents that conform to industry characteristics.
[0089] The above-mentioned knowledge base module is data-connected to the above-mentioned interactive generation module, and is used to store preset business knowledge texts, and perform retrieval processing based on externally input texts to be retrieved to obtain retrieval results; the above-mentioned retrieval results include a reference text set and a reorganized text set.
[0090] It should be noted that the above-mentioned business knowledge text includes business knowledge specific to the industry, such as description text of specific concepts, daily production data and business logic, etc., which is not limited in the embodiment of the present invention.
[0091] The above-mentioned interactive generation module is data-connected to the above-mentioned style conversion module, and is used to store a preset expert template set, and to generate and infer prompt words based on the above-mentioned task requirement text, the above-mentioned extension flag, the above-mentioned retention flag and the above-mentioned search results to obtain the target duty document; the above-mentioned expert template set includes N expert templates; N is an integer greater than 1.
[0092] It should be noted that each expert template is used to describe the expert role in a specific field of the profession in text form. For example, the expert template corresponding to a weather expert can be in the following form:
[0093]
[0094] The style conversion module is used to perform style conversion processing on the initial duty document based on a preset style template text to obtain a target duty document.
[0095] It should be noted that the above-mentioned style template text may be a historical duty document or other reference document, and the embodiment of the present invention does not limit this.
[0096] In an optional embodiment, if Figure 2 As shown, the above-mentioned knowledge base module includes a text segmentation unit, a vector index construction unit, a reference counting unit, a similarity calculation unit, a retrieval unit and a text reorganization unit.
[0097] The above-mentioned text segmentation unit is data-connected with the above-mentioned vector index construction unit and the above-mentioned retrieval unit, and is used to segment the above-mentioned business knowledge text into blocks to obtain a text block set; the above-mentioned text block set includes P text blocks; P is an integer greater than 1.
[0098] It should be noted that the above-mentioned block processing is to divide the business knowledge text into a group of paragraphs containing multiple sentences, where each paragraph constitutes a text block.
[0099] The above-mentioned vector index construction unit is data-connected to the above-mentioned similarity calculation unit and the above-mentioned text reorganization unit, and is used to process the above-mentioned task requirement text and the above-mentioned text block set to obtain a task vector and an index vector set, and in response to receiving the text to be retrieved from external input, process the above-mentioned text to be retrieved to obtain a vector to be retrieved; the above-mentioned index vector set includes an index vector corresponding to each of the above-mentioned text blocks.
[0100] It should be noted that the above-mentioned vector index construction unit can be constructed based on an embedding model such as BERT or OpenAI Embeddings and a vector database such as FAISS or Pinecone, wherein: (1) the embedding model is used to process the task requirement text and the text to be retrieved to obtain the task vector and the vector to be retrieved, respectively, and to process each text block in the text block set to obtain the corresponding index vector; (2) the vector database is used to store the obtained task vector, vector to be retrieved and index vector set.
[0101] The above-mentioned reference counting unit is data-connected to the above-mentioned similarity calculation unit and the above-mentioned text reorganization unit, and is used to store a reference count set and update the above-mentioned reference count set based on the reorganized text set from the above-mentioned text reorganization unit; the above-mentioned reference count set includes the reference count corresponding to each of the above-mentioned text blocks.
[0102] It should be noted that the number of references corresponding to each text block is initialized to 0.
[0103] It should be noted that the above updating of the citation count set based on the reorganized text set is to increase the corresponding citation count of each text block in the reorganized text set by 1 after each text reorganization unit calculates the reorganized text set.
[0104] The similarity calculation unit is data-connected to the retrieval unit and is used to process the vector to be retrieved, the index vector set and the reference count set using a similarity calculation model to obtain a similarity set; the similarity set includes P similarity values.
[0105] The above-mentioned retrieval unit is data-connected to the above-mentioned text reorganization unit, and is used to generate a reference text set and a non-reference text set based on the above-mentioned similarity set and the above-mentioned text block set; the above-mentioned reference text set and the non-reference text set include Q1 and Q2 of the above-mentioned text blocks respectively; Q1+Q2≤P, and Q1 and Q2 are both integers greater than 1.
[0106] It should be noted that the above generation of the reference text set and the non-reference text set based on the similarity set and the text block set is to sort the text block set in descending order based on the similarity value to obtain the sorted text block set, and then take the first Q1 and the first Q2 text blocks in the sorted text block set in turn to obtain the reference text set and the non-reference text set.
[0107] The text reorganization unit is used to process the reference text set, the non-reference text set, the task vector and the vector to be retrieved to obtain the reorganized text set.
[0108] It should be noted that only after receiving the external input text to be retrieved each time, the vector index construction unit, similarity calculation unit, retrieval unit and text reorganization unit will execute the corresponding functions in sequence, and finally drive the reference counting unit to update the reference count set.
[0109] In another optional embodiment, the similarity calculation model is expressed as follows:
[0110]
[0111] Where SM i is the i-th similarity value in the similarity set; i is an integer from 1 to P; λ is a preset weight coefficient; VA is the vector to be searched; VB i is the i-th index vector in the above index vector set; α is the preset adjustment coefficient; u i is the i-th citation number in the above citation number set.
[0112] Optionally, the value range of the above weight coefficient is [0.7, 0.9].
[0113] Preferably, the above weight coefficient is 0.8.
[0114] Optionally, the adjustment coefficient has a value range of [0.1, 1].
[0115] Preferably, the adjustment coefficient is 0.6.
[0116] It should be noted that the expression of the above similarity calculation model includes two parts: cosine similarity and similarity based on the number of citations. The closer the value of the cosine similarity is to 1, the closer VA and VB are. i The closer they are in semantic space, the more similar the searched text and the i-th text block are. Citation-based similarity takes into account the number of citations of the corresponding text block. Multiple citations of a text block increase its similarity value, placing it higher in the ranked text block set and thus gradually increasing its importance. Furthermore, by weighting the two components using a weight coefficient, we can adjust their respective contributions to the similarity score.
[0117] It can be seen that by calculating the similarity value through the similarity calculation model, semantic similarity and the number of citations can be comprehensively considered, so that during retrieval, on the basis of giving priority to recalling text blocks with high semantic similarity, high-quality text blocks that are frequently cited can be further given priority.
[0118] In another optional embodiment, Figure 3 As shown, the above-mentioned text reorganization unit includes a relevance calculation subunit, a screening subunit and a merging subunit.
[0119] The above-mentioned relevance calculation subunit is data-connected to the above-mentioned screening subunit, and is used to use the relevance calculation model to process the above-mentioned non-reference text set, the above-mentioned index vector set, the above-mentioned vector to be retrieved and the above-mentioned task vector to obtain a relevance set; the above-mentioned relevance set includes Q2 relevance values.
[0120] The above-mentioned screening subunit is connected to the data of the above-mentioned merging subunit, and is used to sort the above-mentioned non-reference text set based on the above-mentioned association set to obtain a sorted text set, and delete the last Q3 text blocks of the above-mentioned sorted text set to obtain a supplementary text set; Q3 is an integer greater than 1 and less than Q2.
[0121] It should be noted that the above sorting is to sort the text blocks in the non-reference text set in descending order of the corresponding relevance values.
[0122] The merging subunit is used to calculate the union of the reference text set and the supplementary text set to obtain a recombined text set.
[0123] In another optional embodiment, the above correlation calculation model is expressed as follows:
[0124] CD j =cos β (VI j ,VC)(1-cos(VI j ,VA)) β
[0125] In the formula, CD j is the jth correlation value in the above correlation set CD; VI j is the j-th text block in the non-reference text set I, and the corresponding index vector in the index vector set; VA and VC are the vector to be retrieved and the task vector respectively; β is the preset correlation coefficient; j is an integer from 1 to Q2.
[0126] Optionally, the value range of the above correlation coefficient is [0,1].
[0127] Preferably, the above correlation coefficient is 0.5.
[0128] It should be noted that the above correlation calculation model includes two multiplication factors, where the first factor represents the cosine similarity. The closer the value is to 1, the better the VI j The closer it is to VC in the semantic space, the more similar the j-th text block in the task requirement text and the non-reference text set is; the second factor represents VI j The degree of difference between β and VA. The larger the value, the greater the difference between the searched text and the j-th text block in the non-referenced text set. Furthermore, calculating the β power for each of the above two factors can amplify the impact of the two factors on the association value. The closer β is to 0, the more stable the calculated association value, while the closer β is to 1, the more differentiated the calculated association value is.
[0129] It can be seen that through the above-mentioned correlation calculation model, the calculated correlation value can comprehensively reflect the degree of difference between the text blocks in the non-referenced text set and the text to be retrieved, as well as the degree of similarity with the task requirement text, so that in the sorted text set, the text blocks that are similar to the overall task requirements and have a high degree of difference with the text to be retrieved are placed higher, ensuring that the text blocks in the supplementary text set neither deviate from the overall task requirements nor contain as much related information as possible.
[0130] In another optional embodiment, Figure 4As shown, the above-mentioned interaction generation module includes an expert template storage unit, a first prompt word generation unit, a second prompt word generation unit, a third prompt word generation unit, a fourth prompt word generation unit, a fifth prompt word generation unit and an inference unit.
[0131] The expert template storage unit is used to store a preset expert template set.
[0132] The first prompt word generation unit is data-connected to the intention acquisition module, the knowledge base module and the reasoning unit, and is used to perform a first combination process on the task requirement text and the reference text set to obtain an initial prompt word.
[0133] It should be noted that the first combination process mentioned above is to fill the reference text set and the task requirement text into the preset first prompt word template. The first prompt word template mentioned above can be in the following form, where positions ① and ② are used to fill the reference text set and the task requirement text respectively:
[0134]
[0135] The second prompt word generating unit is data-connected to the reasoning unit and is used to perform a second combination process on the externally input current question text, current expert template and current text block set to obtain question extension prompt words.
[0136] It should be noted that the second combination process mentioned above is to fill the current question text, the current expert template, and the current text block set into the preset second prompt word template. The second prompt word template mentioned above can be in the following form, where positions ③, ④, and ⑤ are used to fill the current text block set, the current expert template, and the current question text, respectively:
[0137]
[0138] The third prompt word generating unit is data-connected to the intention acquisition module and the reasoning unit, and is used to perform a third combination process on the task requirement text and the externally input current expert template to obtain the associated question prompt words.
[0139] It should be noted that the third combination process mentioned above is to fill the task requirement text and the current expert template into the preset third prompt word template. The third prompt word template mentioned above can be in the following form, where positions 6 and 7 are used to fill the current expert template and the task requirement text respectively:
[0140]
[0141] The fourth prompt word generating unit is data-connected to the reasoning unit and is used to perform a fourth combination process on the externally input current question text, the current expert template and the current text block set to obtain the current prompt word.
[0142] It should be noted that the fourth combination process mentioned above is to fill the current question text, the current expert template, and the current text block set into the preset fourth prompt word template. The fourth prompt word template mentioned above can be in the following form, where positions ⑧, ⑨, and ⑩ are used to fill the current text block set, the current expert template, and the current question text, respectively:
[0143]
[0144] The fifth prompt word generating unit is connected to the data of the inference unit and is used to perform a fifth combination processing on the current answer text input from the outside and the current on-duty clerk to obtain the document update prompt word.
[0145] It should be noted that the fifth combination process is to fill the current answer text and the current duty clerk into the preset fifth prompt word template. The fifth prompt word template can be in the following form, where the position and Used to fill in the current answer text and the current on-duty clerk respectively:
[0146]
[0147] The above-mentioned reasoning unit is used to process the document construction prompt words to obtain the corresponding answer text; the above-mentioned document construction prompt words are the above-mentioned initial prompt words, the above-mentioned question expansion prompt words, the above-mentioned related question prompt words, the above-mentioned current prompt words or the above-mentioned document update prompt words.
[0148] It should be noted that the above-mentioned reasoning unit is a large reasoning model such as GPT-4o and Qwen 2.5, which is not limited in the embodiment of the present invention.
[0149] In another optional embodiment, the above-mentioned style conversion module is a pre-trained style conversion model; the above-mentioned style conversion model includes a style embedding vector generation unit, a style encoder unit, a source embedding vector set generation unit, a source encoder unit, a feature vector construction unit, a decoder unit and a splicing unit.
[0150] The style embedding vector generating unit is used to first split the preset style template text into several sentences, calculate the embedding vector of each sentence of the style template text, and then set the style embedding vector to the mean of the embedding vectors of all style template sentences.
[0151] It should be noted that the above-mentioned style template text may be a historical duty text or other specified text, and the embodiment of the present invention does not limit this.
[0152] The style encoder unit is used to perform style encoding processing on the style embedding vector to obtain a style vector.
[0153] Optionally, the style encoder unit is constructed based on MLP (Multi-Layer Perceptron).
[0154] The source embedding vector set generating unit is used to split the initial duty document into several clauses, and calculate the embedding vector of each clause of the initial duty document to obtain the source embedding vector set.
[0155] The source encoder unit is used to process each clause of the initial duty document to obtain a corresponding encoding vector.
[0156] Optionally, the source encoder unit is constructed based on a BiLSTM network.
[0157] The feature vector construction unit is used to sequentially concatenate the style vector with each encoding vector to obtain a feature vector corresponding to each encoding vector.
[0158] The decoder unit is used to decode each feature vector in turn to obtain the corresponding target clause.
[0159] Optionally, the above decoder is built based on an LSTM network.
[0160] The above-mentioned splicing unit is used to splice the target clauses in sequence to obtain the target duty document.
[0161] It can be seen that the on-duty document generation system disclosed in the embodiment of the present invention can improve the richness and completeness of the generated on-duty documents by introducing a user feedback mechanism in multiple rounds of dialogue, and can fully utilize the user's in-depth understanding of the industry to generate personalized on-duty documents that conform to industry characteristics.
[0162] Example 2
[0163] See also Figure 5 , Figure 5 This is a flow chart of a method for generating duty documents disclosed in an embodiment of the present invention. Figure 5 The process structure diagram of the duty document generation method described above is applied to text data processing, such as duty document generation, and is not limited in the embodiment of the present invention. Figure 5 As shown, the duty document generation method includes:
[0164] S1. Use the intent acquisition module to obtain the task requirement text.
[0165] S2. Use the knowledge base module to store preset business knowledge texts.
[0166] S3. Using the interactive generation module and the above-mentioned knowledge base module, the current duty document is constructed based on the above-mentioned task requirement text.
[0167] S4. Utilize the intention acquisition module, the interaction generation module, and the knowledge base module to update the current duty document to obtain an updated current duty document.
[0168] S5. Repeat S4 until an end instruction input by the user is received.
[0169] S6. Determine the initial duty clerk as the current duty clerk.
[0170] S7. Using a style conversion module, based on a preset style template text, perform style conversion on the initial duty document to obtain a target duty document.
[0171] In an optional embodiment, the interactive generation module and the knowledge base module are used to construct the current duty document based on the task requirement text, including:
[0172] S31. Using the above-mentioned knowledge base module, the above-mentioned task requirement text is searched and processed to obtain a corresponding reference text set and a reorganized text set.
[0173] It should be noted that the above-mentioned use of the knowledge base module to retrieve and process the task requirement text is to input the task requirement text as the text to be retrieved into the knowledge base module, so that the vector index construction unit, the similarity calculation unit and the retrieval unit are executed in sequence to obtain the corresponding reference text set and the reorganized text set.
[0174] S32: Using a first prompt word generating unit, perform a first combination process on the task requirement text and the reference text set to obtain an initial prompt word.
[0175] S33. Using the inference unit, the initial prompt word is processed to obtain the current duty clerk.
[0176] It should be noted that the above-mentioned current duty document is the answer text obtained after the reasoning unit processes the initial prompt word.
[0177] It can be seen that since the text blocks in the reference text set retain the information most similar to the task requirement text, constructing the first prompt word on this basis and generating the current duty document can make the initially generated current duty document contain as accurate information as possible.
[0178] In another optional embodiment, the intention acquisition module, the interaction generation module, and the knowledge base module are used to update the current duty document to obtain the updated current duty document, including:
[0179] S41. Initialize the current text block set to empty; initialize the number of loops l to 1; and initialize the current question text to the above-mentioned task requirement text.
[0180] S42: Set the current expert template as the lth expert template in the expert template set.
[0181] S43. Utilize the intention acquisition module and the reasoning unit to update the current question text based on the current expert template, the current text block set, and the task requirement text.
[0182] S44: Utilize the knowledge base module to perform search processing on the current question text to obtain a corresponding reference text set and a reorganized text set; and update the current text block set to the reorganized text set.
[0183] It should be noted that the above-mentioned use of the knowledge base module to perform retrieval processing on the current question text is to input the current question text as the text to be retrieved into the knowledge base module, so that the vector index construction unit, the similarity calculation unit and the retrieval unit are executed in sequence to obtain the corresponding reference text set and the reorganized text set.
[0184] S45. Utilize the reasoning unit, the fourth prompt word generating unit, and the fifth prompt word generating unit to update the current duty clerk based on the current expert template, the current text block set, and the current question text.
[0185] S46. Add 1 to the value of l.
[0186] S47. Repeat S42 to S46 until l is greater than N.
[0187] In yet another optional embodiment, the above-mentioned use of the above-mentioned intention acquisition module and the above-mentioned reasoning unit to update the above-mentioned current question text based on the above-mentioned current expert template, the above-mentioned current text block set and the above-mentioned task requirement text includes:
[0188] S431. Utilize the intention acquisition module to acquire the extension flag.
[0189] S432, judging the above extension flag;
[0190] When the extension flag is yes, execute S433;
[0191] When the extension flag is negative, execute S435.
[0192] S433: Using a second prompt word generating unit, perform a second combination process on the current question text, the current expert template, and the current text block set to obtain question extension prompt words.
[0193] S434. Use the above-mentioned reasoning unit to process the above-mentioned question extension prompt words to obtain the extended question text; update the above-mentioned current question text to the above-mentioned extended question text; and execute S44.
[0194] It should be noted that the above-mentioned extended question text is the answer text obtained after the reasoning unit processes the question extension prompt words.
[0195] It should be noted that the text blocks in the reference text set obtained in step S44 retain the information most similar to the current question text, while the reorganized text set adds richer related information based on the reference text set. Thus, by updating the current text block set to this reorganized text set, further constructing question extension prompts, generating an extended question text, and finally updating the current question text to the extended question text, the topic scope of the current question text can be effectively expanded after multiple rounds of dialogue.
[0196] S435 , using a third prompt word generating unit, performing a third combination process on the task requirement text and the current expert template to obtain a related question prompt word.
[0197] S436: Utilize the above-mentioned reasoning unit to process the above-mentioned related question prompt words to obtain the related question text; and update the above-mentioned current question text to the above-mentioned related question text.
[0198] It should be noted that the above-mentioned related question text is the answer text obtained after the reasoning unit processes the related question prompt words.
[0199] In yet another optional embodiment, the above-mentioned updating of the current duty clerk based on the current expert template, the current text block set, and the current question text by utilizing the above-mentioned reasoning unit, the fourth prompt word generation unit, and the fifth prompt word generation unit includes:
[0200] S451 : Using a fourth prompt word generating unit, perform a fourth combination process on the current question text, the current expert template, and the current text block set to obtain a current prompt word.
[0201] S452: Utilize the above-mentioned reasoning unit to process the above-mentioned current prompt word to obtain the current answer text.
[0202] S453. Utilize the above-mentioned intention acquisition module to obtain a reservation flag.
[0203] S454, judging the above-mentioned reservation flag;
[0204] When the above-mentioned reservation flag is yes, execute S455;
[0205] When the retain flag is no, the current duty document remains unchanged; and S46 is executed.
[0206] S455. Utilize the fifth prompt word generating unit to perform a fifth combination process on the current answer text and the current clerk on duty to obtain a clerk update prompt word.
[0207] S456. Use the above-mentioned reasoning unit to process the above-mentioned document update prompt word to obtain an intermediate text; and replace the content of the above-mentioned current duty document with the above-mentioned intermediate text.
[0208] It should be noted that the above-mentioned intermediate text is the answer text obtained after the reasoning unit processes the document update prompt words.
[0209] It can be seen that the implementation of the duty document generation method disclosed in the embodiment of the present invention can improve the richness and completeness of the generated duty documents by introducing a user feedback mechanism in multiple rounds of conversations, and can fully utilize the user's in-depth understanding of the industry to generate personalized duty documents that conform to industry characteristics.
[0210] The device embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0211] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0212] Finally, it should be noted that the on-duty document generation system and method based on the inference big model disclosed in the embodiment of the present invention is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that it is still possible to modify the technical solutions recorded in the aforementioned embodiments, or to replace some of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A duty document generation system, characterized in that: It includes the intention acquisition module, knowledge base module, interaction generation module and style conversion module: The intention acquisition module is data-connected with the knowledge base module and the interaction generation module, and is used to acquire the task requirement text, the extension flag and the reservation flag; The knowledge base module is connected to the interactive generation module for storing preset business knowledge texts and performing retrieval processing based on external input texts to be retrieved to obtain retrieval results; the retrieval results include a reference text set and a reorganized text set; The interactive generation module is data-connected to the style conversion module, and is used to store a preset expert template set, and to generate and infer prompt words based on the task requirement text, the extension flag, the retention flag, and the search results to obtain the target duty document; The expert template set includes N expert templates; N is an integer greater than 1; The style conversion module is used to perform style conversion processing on the initial duty document based on a preset style template text to obtain a target duty document.
2. The duty document generation system according to claim 1, characterized in that: The knowledge base module includes a text segmentation unit, a vector index construction unit, a reference counting unit, a similarity calculation unit, a retrieval unit and a text reorganization unit; The text block unit is data-connected to the vector index building unit and the retrieval unit, and is used to block the business knowledge text to obtain a text block set; the text block set includes P text blocks; P is an integer greater than 1; The vector index construction unit is data-connected to the similarity calculation unit and the text reorganization unit, and is used to process the task requirement text and the text block set to obtain a task vector and an index vector set, and in response to receiving a text to be retrieved from an external input, process the text to be retrieved to obtain a vector to be retrieved; the index vector set includes an index vector corresponding to each text block; The reference counting unit is data-connected to the similarity calculation unit and the text reorganization unit, and is used to store a reference count set and update the reference count set based on the reorganized text set from the text reorganization unit; the reference count set includes the reference count corresponding to each of the text blocks; The similarity calculation unit is data-connected to the retrieval unit and is used to process the vector to be retrieved, the index vector set, and the reference count set using a similarity calculation model to obtain a similarity set; the similarity set includes P similarity values; The retrieval unit is data-connected to the text reorganization unit and is used to generate a reference text set and a non-reference text set based on the similarity set and the text block set; the reference text set and the non-reference text set include Q1 and Q2 text blocks, respectively; Q1+Q2≤P, and Q1 and Q2 are both integers greater than 1; The text reorganization unit is used to process the reference text set, the non-reference text set, the task vector and the vector to be retrieved to obtain the reorganized text set.
3. The duty document generation system according to claim 2, characterized in that: The text reorganization unit includes a relevance calculation subunit, a screening subunit and a merging subunit: The relevance calculation subunit is data-connected to the screening subunit and is used to process the non-reference text set, the index vector set, the vector to be retrieved, and the task vector using a relevance calculation model to obtain a relevance set; the relevance set includes Q2 relevance values; The screening subunit is data-connected to the merging subunit, and is used to sort the non-reference text set based on the relevance set to obtain a sorted text set, and delete the last Q3 text blocks of the sorted text set to obtain a supplemented text set; Q3 is an integer greater than 1 and less than Q2; The merging subunit is used to calculate the union of the reference text set and the supplementary text set to obtain a recombined text set.
4. The duty document generation system according to claim 3, characterized in that: The expression of the correlation calculation model is: CD j =basket β (VI j ,VC)(1-cos(VI j ,VA)) β In the formula, CD j is the jth said association value in said association set CD; VI j is the j-th text block in the non-reference text set I, and the corresponding index vector in the index vector set; VA and VC are the vector to be retrieved and the task vector respectively; β is the preset correlation coefficient; j is an integer from 1 to Q2.
5. The duty document generation system according to claim 1, characterized in that: The interactive generation module includes an expert template storage unit, a first prompt word generation unit, a second prompt word generation unit, a third prompt word generation unit, a fourth prompt word generation unit, a fifth prompt word generation unit and an inference unit; The expert template storage unit is used to store a preset expert template set; The first prompt word generation unit is data-connected to the intention acquisition module, the knowledge base module, and the reasoning unit, and is used to perform a first combination process on the task requirement text and the reference text set to obtain an initial prompt word; The second prompt word generating unit is data-connected to the reasoning unit and is used to perform a second combination process on the externally input current question text, current expert template and current text block set to obtain the question extension prompt word; The third prompt word generation unit is data-connected to the intention acquisition module and the reasoning unit, and is used to perform a third combination process on the task requirement text and the externally input current expert template to obtain the associated question prompt word; The fourth prompt word generating unit is data-connected to the reasoning unit and is used to perform a fourth combination process on the externally input current question text, the current expert template, and the current text block set to obtain a current prompt word; The fifth prompt word generating unit is data-connected to the inference unit and is used to perform a fifth combination process on the externally input current answer text and the current on-duty clerk to obtain the document update prompt word; The reasoning unit is used to process the document construction prompt words to obtain the corresponding answer text; the document construction prompt words are the initial prompt words, the question expansion prompt words, the related question prompt words, the current prompt words or the document update prompt words.
6. A method for generating duty documents, characterized in that: The method applied to the duty document generation system according to any one of claims 1 to 5 comprises: S1. Use the intent acquisition module to obtain the task requirement text; S2. Using the knowledge base module to store preset business knowledge texts; S3. Using the interactive generation module and the knowledge base module, construct a current duty document based on the task requirement text; S4. Using the intention acquisition module, the interaction generation module, and the knowledge base module, the current duty document is updated to obtain an updated current duty document. S5, repeat S4 until receiving the end instruction input by the user; S6. Determine the initial duty clerk as the current duty clerk; S7. Using a style conversion module, based on a preset style template text, perform style conversion on the initial duty document to obtain a target duty document.
7. The method for generating duty documents according to claim 6, characterized in that: The interactive generation module and the knowledge base module are used to construct a current duty document based on the task requirement text, including: S31, using the knowledge base module to perform search processing on the task requirement text to obtain a corresponding reference text set and a reorganized text set; S32, using a first prompt word generating unit, performing a first combination process on the task requirement text and the reference text set to obtain an initial prompt word; S33: Using the inference unit, the initial prompt word is processed to obtain the current duty clerk.
8. The method for generating duty documents according to claim 6, characterized in that: The updating of the current duty document by using the intention acquisition module, the interaction generation module, and the knowledge base module to obtain the updated current duty document includes: S41, initializing the current text block set to empty; initializing the number of loops l to 1; initializing the current question text to the task requirement text; S42, setting the current expert template as the lth expert template in the expert template set; S43, using the intention acquisition module and the reasoning unit, based on the current expert template, the current text block set and the task requirement text, updating the current question text; S44, using the knowledge base module to perform search processing on the current question text to obtain a corresponding reference text set and a reorganized text set; updating the current text block set to the reorganized text set; S45, using the reasoning unit, the fourth prompt word generation unit, and the fifth prompt word generation unit to update the current duty clerk based on the current expert template, the current text block set, and the current question text; S46, add 1 to the value of l; S47. Repeat S42 to S46 until l is greater than N.
9. The method for generating duty documents according to claim 8, characterized in that: The updating of the current question text based on the current expert template, the current text block set, and the task requirement text by utilizing the intention acquisition module and the reasoning unit includes: S431, using the intention acquisition module to acquire the extension flag; S432, judging the extension flag; When the extension flag is yes, execute S433; When the extension flag is no, executing S435; S433: Using a second prompt word generating unit, performing a second combination process on the current question text, the current expert template, and the current text block set to obtain a question extension prompt word; S434: Using the reasoning unit, process the question extension prompt word to obtain an extended question text; update the current question text to the extended question text; and execute S44; S435: Using a third prompt word generating unit, performing a third combination process on the task requirement text and the current expert template to obtain a related question prompt word; S436: Utilize the inference unit to process the associated question prompt words to obtain an associated question text; and update the current question text to the associated question text.
10. The method for generating duty documents according to claim 8, characterized in that: The updating of the current duty clerk based on the current expert template, the current text block set, and the current question text by using the reasoning unit, the fourth prompt word generating unit, and the fifth prompt word generating unit includes: S451: Using a fourth prompt word generating unit, performing a fourth combination process on the current question text, the current expert template, and the current text block set to obtain a current prompt word; S452: Using the inference unit, process the current prompt word to obtain a current answer text; S453, using the intention acquisition module to acquire a reservation flag; S454, judging the reserved flag; When the reserved flag is yes, execute S455; When the retention flag is no, the current duty document remains unchanged; execute S46; S455: Using a fifth prompt word generating unit, perform a fifth combination process on the current answer text and the current clerk on duty to obtain a clerk update prompt word; S456. Utilize the inference unit to process the document update prompt word to obtain an intermediate text; and replace the content of the current duty document with the intermediate text.
Citation Information
Patent Citations
Book content retrieval method and device based on intelligent word segmentation and computer equipment
CN119513290A
Computing System for Inferring Demographics Using Deep Learning Computations and Social Proximity on a Social Data Network
US20170357890A1
Method and system for generating document, computing device, and medium
WO2025107896A1
Cited By
Document auxiliary generation method and device, equipment and medium
CN120930624A