Large model writing method and system based on structured meta-knowledge base
By adopting a large-scale model writing method based on a structured meta-knowledge base in the field of article writing, the problems of insufficient knowledge timeliness and unstable quality of generated content are solved, and efficient, accurate and customizable professional document automation creation is achieved.
Patent Information
- Application Number
- CN202510527960.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art has problems in the field of article writing, insufficient knowledge timeliness, unstable quality of generated content, and lack of support for complex structured documents.
The large-scale model writing method based on the structured meta-knowledge base is adopted to extract the meta-knowledge of articles and paragraphs through structured learning, dynamically manage the knowledge base, ensure the timeliness and quality of knowledge, and use search enhancement generation technology to generate logically coherent and clear structured content.
It improves the timeliness and accuracy of knowledge, improves the quality and stability of generated content, enhances the customization and efficiency of writing, and realizes closed-loop optimization.
Smart Images

Figure CN120087341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a large model writing method and system based on a structured meta knowledge base. Background Art
[0002] With the rapid development of information technology and artificial intelligence, text generation technology has shown great potential in the field of content creation. Combining natural language processing (NLP) and big data analysis technologies, automated writing tools can play an important role in content generation from general text to specific domain reports. However, existing collaboration tools are mainly limited to shallow generation in general scenarios, lack support for complex structured documents, and are insufficient in specific content requirements in specific domains. These tools usually rely on static knowledge bases and cannot update information in real time, resulting in the generated content being possibly outdated or inaccurate.
[0003] In specific document generation tasks, such as the writing of reports or documents, the following requirements exist: First, it is necessary to extract the latest information from an ever-updating knowledge base; second, the generated content needs to ensure clear logic and no semantic deviation; finally, it can flexibly generate document content with different styles and structures according to user needs to achieve high customization. However, existing technologies have significant deficiencies in these aspects. The knowledge bases in traditional systems are mostly statically stored with a long update cycle and cannot meet the real-time writing needs; in the long text scenario, the generated content lacks consistency and the problem of logical disconnection is more prominent; there is no closed-loop between retrieval and generation, and users need to manually operate or adjust, increasing the usage cost. Therefore, existing technologies still face challenges in the field of article writing. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a large model writing method and system based on a structured meta knowledge base, aiming to systematically solve the problems of insufficient knowledge timeliness and unstable generated content quality through automatic extraction and dynamic management of meta knowledge, a method for generating logically coherent long texts, and a closed-loop optimized knowledge base update mechanism, and finally achieve efficient, accurate, and customizable professional document automation creation.
[0005] To solve the above technical problems, the present application is implemented as follows: In a first aspect, the embodiments of the present application provide a large model writing method based on a structured meta knowledge base, and the method includes: Step 1, perform structured learning on the target article, extract the meta knowledge of the article and paragraphs for storage, so as to construct an article writing meta knowledge base; Step 2, update the quality and usage amount of the meta knowledge items in the article writing meta knowledge base according to the knowledge base management mechanism to obtain a final structured meta knowledge base; Step 3: Based on the structured meta-knowledge base, generate an outline and outline content according to the input article summary and relevant constraints to integrate and obtain the article writing content.
[0006] As an optional implementation manner of the first aspect of the present application, step 1 includes: Step 11: Select a target article from the original knowledge base and perform core analysis and summary and paragraph splitting summary to obtain an article summary and a paragraph summary respectively; Step 12: Based on the article summary and the paragraph summary, extract meta-knowledge information, where the meta-knowledge information includes article meta-knowledge and paragraph meta-knowledge; Step 13: After performing similarity matching between the meta-knowledge information and the meta-knowledge in the original knowledge base, store it in the database; Step 14: Vectorize the meta-knowledge in the database to obtain an article writing meta-knowledge base.
[0007] As an optional implementation manner of the first aspect of the present application, in step 12, the article meta-knowledge includes audience analysis, theme development logic, article structure, and logical coherence, and the paragraph meta-knowledge includes paragraph structure, grammar, vocabulary, citation, transition sentence, rhetoric, sentence pattern, tone, and information supplement.
[0008] As an optional implementation manner of the first aspect of the present application, in step 2, the knowledge base management mechanism includes: Step 21: Update the quality score of the meta-knowledge item, where each meta-knowledge item is a unit representation of the system framework of a kind of meta-knowledge; Step 22: Use retrieval-enhanced generation technology to count the usage frequency of the meta-knowledge item in the target article; Step 23: When the quality score and the usage frequency of the meta-knowledge item are respectively lower than the first threshold and the second threshold, delete the meta-knowledge item and transfer the paragraph text of the meta-knowledge item to a similar high-quality meta-knowledge item.
[0009] As an optional implementation manner of the first aspect of the present application, step 22 includes: Set the initial quality score and usage frequency of the newly added meta-knowledge item to 0; Record the usage frequency of each meta-knowledge item. Each time an article is generated, the usage frequency of the used meta-knowledge item is incremented by 1, and each time the retrieval-enhanced generation technology is used for statistics, the usage frequency of the meta-knowledge item will be increased.
[0010] As an optional implementation manner of the first aspect of the present application, step 3 includes: Step 31: Use a large language model and retrieval-enhanced generation technology to construct an article outline and an article synopsis from the key information in the article writing meta-knowledge base; Step 32: Based on the article outline and the article synopsis, use retrieval-enhanced generation technology to obtain the meta-knowledge and writing methods related to each paragraph to generate article writing content.
[0011] As an alternative implementation of the first aspect of the present application, in step 32, the content of article writing includes short texts and long texts. The short text is a text not exceeding two thousand words, which is generated by retrieving paragraph writing meta-knowledge from the article writing meta-knowledge base and retrieval enhancement generation technology. The long text is a text exceeding two thousand words, which is obtained by integrating and optimizing the specific content of the section titles generated by the retrieval enhancement generation technology for all section titles and summaries of the article outline.
[0012] In a second aspect, an embodiment of the present application provides a large model writing system based on a structured meta-knowledge base. The system includes: A meta-knowledge base construction module, configured to perform structured learning on a target article, extract meta-knowledge of the article and paragraphs for storage, so as to construct an article writing meta-knowledge base; A knowledge base update module, configured to update the quality and usage amount of meta-knowledge items in the article writing meta-knowledge base according to a knowledge base management mechanism, so as to obtain a final structured meta-knowledge base; An article content generation module, configured to generate an outline and outline content based on the structured meta-knowledge base according to an input article summary and related constraints, so as to integrate and obtain article writing content.
[0013] In a third aspect, an embodiment of the present application provides an electronic device. The electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0015] Compared with the prior art, the present invention proposes a large model writing method based on a structured meta-knowledge base, which has the following beneficial effects: (1) Improve the timeliness and accuracy of knowledge: Continuously extract the meta-knowledge of the target article through structured learning, and combine with a dynamic update mechanism to ensure the timeliness and quality of the meta-knowledge base, solving the problem of stale information in traditional static knowledge bases.
[0016] (2) Improve the quality and stability of the generated content: Writing based on the structured meta-knowledge base, using the meta-knowledge at the article and paragraph levels to guide content generation. Especially for long texts, through step-by-step generation and integration optimization, the logical coherence and structural clarity of the content are ensured, and the stability of the generated quality is improved.
[0017] (3) Enhance the customization and efficiency of writing: Users can input a brief description of the article and relevant constraints, and the system generates customized outlines and content based on the meta-knowledge base to meet the needs of documents with different styles and structures; the automated process reduces manual intervention and improves writing efficiency.
[0018] (4) Achieve closed-loop optimization: The construction, use, and update of the knowledge base form a closed loop, and the system can dynamically manage according to the actual usage effects (quality scores, usage frequencies) of the meta-knowledge items, continuously optimize the knowledge base, and improve the overall performance. Description of the Drawings
[0019] Figure 1 is a flowchart of a large model writing method based on a structured meta-knowledge base provided by the first embodiment of the present invention; Figure 2 is a schematic structural diagram of a large model writing system based on a structured meta-knowledge base provided by the second embodiment of the present invention. Specific Embodiments
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0022] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0023] Embodiment 1 Please refer to Figure 1 , which is a flowchart of a large model writing method based on a structured meta-knowledge base proposed in the first embodiment of the present application. The steps of the proposed method are as follows.
[0024] Step 1: Perform structured learning on the target article, extract the meta-knowledge of the article and paragraphs for storage, and build an article writing meta-knowledge base.
[0025] Specifically, Step 1 includes the following sub-steps.
[0026] Step 11: Select the target article from the original knowledge base and perform core analysis and summary and paragraph splitting and summary to obtain the article summary and paragraph summary respectively.
[0027] Specifically, first receive the target article uploaded by the user or from the internal article library (supporting docx, doc, markdown, txt formats) and start preprocessing, including format conversion (converting non-text formats to plain text) and removing irrelevant content (headers, footers, page numbers, Markdown syntax, etc.).
[0028] After preprocessing, input the text into a large language model (such as GPT-4o and DeepSeek, etc.) and perform core analysis and summary step by step through the chain of thought. This chain of thought includes: (1) Overall analysis: Let the model read and understand the whole article, identify the style (research report, work summary, notice, etc.), summarize the core theme of the article, and extract keywords; (2) Audience and theme development analysis: The model will refer to the analysis results of the first step, further understand the article, identify the target audience of the article and their characteristic preferences, and infer their reactions to the article; at the same time, the model will also analyze how each part in the article structure supports and develops the theme. (3) Article structure and coherence analysis: The model will identify the main parts of the article and their functions, and describe the logical relationship between each part; (4) Outline generation: The model generates the outline of the article according to the analysis results of the previous steps.
[0029] The core analysis and summary will be stored in the article table of the article database in the knowledge base together with the basic information of the target article.
[0030] Subsequently, perform paragraph summary, including paragraph splitting and embedding processing. Use regular expressions to identify the paragraph delimiters in the article, split to obtain the text content of each paragraph, and record its position in the article. After splitting, save each paragraph text and its position into the paragraph table of the article database, and set a foreign key relationship with the article it belongs to to represent the affiliated relationship. Input each paragraph text into an embedding model (such as text-embedding-3-large or nomic-text-embed, etc.) to obtain the vectorized representation of the text. Then save the vector into the paragraph vector table in the vector repository and perform a foreign key association with the record of this paragraph in the article database to represent that this vector is the vectorized representation of this paragraph text.
[0031] For easy understanding, the following example is given: 1. Receive the file "Research on the Current Situation and Countermeasures of the Construction of the Science and Technology Management Cadre Team in Zhejiang Province" (format: docx) and perform preprocessing to obtain the article text.
[0032] 2. Core analysis and summary of the article: (1) Overall analysis: The text style is a work summary, and the core theme is "Research on the Current Situation and Countermeasures of the Construction of the Science and Technology Management Cadre Team in Zhejiang Province". The keywords are "science and technology management cadres", "team construction", "countermeasure suggestions", "Zhejiang Province", and "technological innovation". Through the investigation of the science and technology management cadre team in Zhejiang Province, the achievements and deficiencies of the current team construction are analyzed. The potential importance of science and technology management in economic development is pointed out, the composition and quality of the cadre team are investigated, the problems faced in practice and training are put forward, and the necessary countermeasures for strengthening the team construction are proposed.
[0033] (2) Analysis of the audience and theme development: The target audience includes science and technology management cadres, local administrators, enterprise management, educators in scientific research institutions and institutions of higher learning, and policymakers. The logical development of the theme is "Through the investigation of the science and technology management cadre team in Zhejiang Province, analyze its construction achievements and problems faced, put forward targeted management and training countermeasures, emphasize the importance of science and technology management to economic development, and emphasize the indispensability of the cadre team in technological innovation." (3) Analysis of the article structure and coherence: The article is divided into 5 parts, including the introduction (introducing the importance of the construction of the science and technology management cadre team and the research background), the investigation process (elaborating on the process and coverage of the investigation, providing data support), the construction achievements (showing the achievements and successful cases of the construction of the science and technology management cadre team in Zhejiang Province), the problem analysis (analyzing the problems and challenges existing in the construction of the science and technology management cadre team), and the countermeasure suggestions (putting forward countermeasure suggestions for the problems, emphasizing the reform of training and the cadre growth mechanism). The overall structure of the article is clear. The theme is introduced through the introduction, and the subsequent parts gradually unfold the investigation results, achievements, and existing problems, and finally end with specific countermeasures. Each part complements each other and effectively promotes the development of the theme. The part that most effectively promotes the theme is the countermeasure suggestions because it directly targets reform and optimization.
[0034] (4) Generate an outline: Introduction: Background introduction, research purpose and significance.
[0035] Investigation process: Distribution and collection of questionnaires, special symposiums, data analysis methods.
[0036] Construction achievements: Degree of attention to technological innovation and science and technology management, construction of science and technology management institutions and systems, improvement of the quality of science and technology management cadres, improvement of science and technology management level, social recognition and support.
[0037] Problem analysis: Insufficient number of science and technology management personnel and excessive workload, mismatch between quality and modern requirements, insufficient training, lack of growth opportunities for grass-roots cadres.
[0038] Countermeasure suggestions: formulating medium- and long-term development plans, implementing knowledge update plans, promoting the exchange of cadres, cultivating enterprise science and technology management talents, establishing a professional title sequence for scientific research management, and increasing the establishment.
[0039] 3. Save the article summary: Use Structured Query Language (SQL) or Object Relational Mapping (ORM) to save the generated core analysis summary of the article into the article table in the article database.
[0040] 4. Paragraph summary: Use regular expressions to split paragraphs and record their positions, then store the paragraph text in the paragraph table and associate it with the article table. For example, its text summary is "The science and technology management level has been improved, the functions have been expanded, the coordination ability has been enhanced, and the team has received favorable comments", and the vector representation is a high-dimensional array (partial dimensions: [-0.005, -0.017, -0.007, -0.003, -0.005, …]), which is saved in the vector database and associated.
[0041] Step 12: Based on the article summary and paragraph summary, extract meta-knowledge information, which includes the meta-knowledge of the article and the meta-knowledge of the paragraph.
[0042] It should be noted that the meta-knowledge of the article may include audience analysis, the logical development of the theme, the article structure, and logical coherence, and the meta-knowledge of the paragraph includes paragraph structure, grammar, and vocabulary, etc.
[0043] Meta-knowledge consists of two parts: "description" and "suggestion". "Description" is used to elaborate on the understanding of this knowledge item, and "suggestion" provides the usage scenarios and methods of this knowledge item.
[0044] Extract the meta-knowledge of the article in two times through the large language model. The first time is to extract the meta-knowledge of audience analysis and the logical development of the theme to understand the audience characteristics and the unfolding logic of the article theme; the second time is to extract the meta-knowledge of the article structure and logical coherence to analyze the organizational framework of the article and the logical relationship between its parts.
[0045] Subsequently, use dynamic selection prompt words to analyze each paragraph and extract paragraph meta-knowledge. Analyze the paragraph text through the large language model to generate the meta-knowledge categories and their characteristics that appear in this paragraph. Then, generate prompt words for guiding analysis based on these categories and characteristics, and the model uses these prompt words to further extract the meta-knowledge of the paragraph.
[0046] The meta-knowledge categories of the paragraph include structure, grammar, vocabulary, citation, transitional sentence, rhetoric, sentence pattern, tone and intonation, and information supplement. These categories respectively focus on the organizational form of the paragraph, the standardization of language expression, the usage effect of vocabulary, the accuracy of citation, the connection between sentences, the application of rhetorical devices, the diversity of sentence patterns, the expression of tone and intonation, and the integrity of information.
[0047] For the sake of easy understanding, the following examples are given: Extract the article records stored in step 11 from the article database. Combine the information of the article records with the prompt words and input them into the model to obtain the meta-knowledge of all 4 articles in two times. Example of model output: Audience analysis Description: The audience mainly includes science and technology management cadres, managers, and leaders of universities and research institutes. They have a certain management and science and technology background, and are concerned about science and technology policies and management efficiency. They show positive expectations for improving the quality and ability of science and technology management cadres, and hope to obtain practical suggestions and solutions to better cope with management challenges.
[0048] Suggestion: In writing, it is necessary to use concise language to present the research results and actual cases, emphasize the importance and necessity of the construction of the science and technology management cadre team, and at the same time provide practical countermeasures to enhance the sense of participation and trust of the audience.
[0049] Subsequently, extract the meta-knowledge of the paragraphs. Select any paragraph in the article and combine it with the prompt words to input into the large language model to obtain the first model output. Obtain the meta-knowledge of this paragraph containing structure, grammar, vocabulary, transitional sentences, sentence patterns, and information supplementation. Among them, the structure and information citation are used well, and the vocabulary usage is poor.
[0050] The second model outputs the meta-knowledge of structure, information supplementation, and vocabulary. Here, the meta-knowledge of a high-quality and a low-quality paragraph are respectively shown.
[0051] Structure (high quality) Description: The paragraph presents a clear logical structure, from introducing the research background to result analysis, then to deficiencies and countermeasure suggestions, following the general - sub - general pattern.
[0052] Suggestion: It is recommended to add subheadings in the preliminary result analysis to better guide the readers and enhance the logic of the writing.
[0053] Vocabulary (low quality) Description: The paragraph uses relatively professional terms such as "technological innovation" and "resource - dependent", which are suitable for the target readers, but some expressions can be more concise.
[0054] Suggestion: Simplify some complex expressions, such as making a more intuitive brief description of "the contradiction between limited supply of production factors and continuous expansion of the economic aggregate".
[0055] Step 13: After performing similarity matching between the meta - knowledge information and the meta - knowledge in the original knowledge base, store it in the database.
[0056] Specifically, the extracted meta-knowledge is embedded to generate a vector representation. Then, in the knowledge item vector table of the vector database, for each knowledge item with the same knowledge category as this meta-knowledge, the Euclidean distance between its vector and the vector of this meta-knowledge is calculated one by one. If the Euclidean distance is less than the threshold, it is considered that the two meta-knowledges are too similar and will not be stored repeatedly; otherwise, if the Euclidean distance is greater than or equal to the threshold, the new meta-knowledge is saved to the knowledge item table.
[0057] For ease of understanding, the following example is given: Taking the audience analysis meta-knowledge in the article meta-knowledge as an example, the processing methods of the meta-knowledges of other articles and paragraphs are the same as it.
[0058] Convert this meta-knowledge into plain text format and use an embedding model for vectorization. The output result of the model is [-0.009, -0.015, -0.019, 0.036, 0.032, …].
[0059] Traverse all vectors in the knowledge item vector table of the vector database and calculate the Euclidean distance from the vector of this meta-knowledge. If there is no vector with a Euclidean distance less than 0.4 (threshold), save this meta-knowledge into the knowledge item database using Structured Query Language or Object Relational Mapping.
[0060] Step 14: Vectorize the meta-knowledge in the database to obtain the article writing meta-knowledge base.
[0061] Specifically, for the newly stored knowledge item, its vectorized representation is stored in the knowledge item vector table of the vector database, and a foreign key relationship is set with the record of this knowledge item in the knowledge item database, indicating that this vector is the vectorized expression of this knowledge item.
[0062] Then the knowledge item will be linked with the records in the article database through a foreign key. The meta-knowledge of the article is linked with the records in the article table, and the paragraph meta-knowledge is linked with the records in the paragraph table, indicating that this knowledge item is extracted from this article / paragraph.
[0063] For ease of understanding, the following example is given (still taking the audience analysis meta-knowledge in the previous step as an example): If the audience analysis meta-knowledge in the previous step is a newly added meta-knowledge to the knowledge item database, save the corresponding vector obtained in the previous step into the knowledge item vector table of the vector database. And use Structured Query Language or Object Relational Mapping to establish a foreign key link with the corresponding record in the knowledge item database. Finally, establish a foreign key link between the record in the knowledge item database and the corresponding article record in the article database.
[0064] If the audience analysis meta-knowledge in the previous step is replaced by an existing meta-knowledge in the knowledge item database. Then establish a foreign key link between this existing meta-knowledge and the corresponding article record in the article database.
[0065] Step 2: Update the quality and usage amount of meta - knowledge items in the article writing meta - knowledge base according to the knowledge base management mechanism to obtain the final structured meta - knowledge base.
[0066] Specifically, the knowledge base management mechanism in this step is as follows: Step 21: Update the quality score of meta - knowledge items. Each meta - knowledge item is a unit representation of the system framework of a type of meta - knowledge.
[0067] Specifically, update the quality score of meta - knowledge items regularly (e.g., weekly or monthly). During an update cycle, randomly select a batch of meta - knowledge items of different categories and without association. For each selected meta - knowledge item, select another meta - knowledge item of the same category from the library as a comparison. At the same time, obtain the article or paragraph text associated with these two meta - knowledge items from the article data table. By comparing (e.g., manual evaluation or model - automatic evaluation) the description accuracy and context applicability of these two meta - knowledge items to their associated text, judge the superiority and inferiority. Add 1 to the quality score of the winner, and keep the score of the inferior party unchanged.
[0068] Step 22: Use retrieval - augmented generation technology to count the usage frequency of meta - knowledge items in the target article.
[0069] It should be noted that the initial quality score and usage frequency of newly added meta - knowledge items can be set to 0. Record the usage frequency of each meta - knowledge item. In the subsequent article generation step (Step 3), when a certain meta - knowledge item is retrieved and used through retrieval - augmented generation technology to assist in generating article content (e.g., optimizing the article structure or generating paragraphs), the usage frequency of this meta - knowledge item is incremented by 1, and the usage frequency of the meta - knowledge item will increase each time the retrieval - augmented generation technology is used for statistics. Therefore, the process of counting the usage frequency is closely related to the application process of RAG (retrieval - augmented generation).
[0070] Step 23: When the quality score and usage frequency of a meta - knowledge item are respectively lower than the first threshold and the second threshold, delete the meta - knowledge item and transfer the paragraph text of the meta - knowledge item to a similar high - quality meta - knowledge item.
[0071] Specifically, regularly check the quality score and usage frequency of meta-knowledge items. If the quality score of a certain meta-knowledge item is lower than a preset first threshold (for example, the score is less than 3), and its usage frequency is lower than a preset second threshold (for example, the number of uses is less than 5 times), then this meta-knowledge item is marked as a low-quality meta-knowledge item. Then, through a similarity matching algorithm (for example, using a similarity threshold slightly looser than that in step 13), find the high-quality meta-knowledge item (with a higher quality score and usage frequency) that is most similar to this low-quality meta-knowledge item in the same knowledge category. If a suitable high-quality meta-knowledge item is found, transfer the article or paragraph text information associated with the low-quality meta-knowledge item (for example, its association record in the paragraph table) to this high-quality meta-knowledge item, and then delete the low-quality meta-knowledge item and its vector representation.
[0072] Step 3: Based on the structured meta-knowledge base, generate an outline and the content of the outline according to the input article summary and relevant constraints to integrate and obtain the article writing content.
[0073] Specifically, step 3 includes the following sub-steps.
[0074] Step 31: Use the large language model and retrieval-augmented generation technology to construct an article outline and an article plan from the key information in the article writing meta-knowledge base.
[0075] Specifically, before generating the article, receive the initial information input by the user, including but not limited to the style, title, expected word count, article summary, and supplementary information for a specific style (such as the background, purpose, data source, method, sampling, etc. of a research report). The user can also upload relevant files as references. The uploaded files are chunked and embedded, and stored in a temporary database table to provide supplementary information for subsequent content generation.
[0076] Subsequently, generate the article outline and plan through a series of (for example, three) large language model calls: The first call: The large model generates a preliminary audience analysis and keywords for the article based on the initial information.
[0077] The second call: Use the RAG technology to embed the initial information, the audience and keywords generated by the first call, and then retrieve the meta-knowledge items with high similarity and related to "topic development logic" and "audience analysis" in the knowledge item vector table of the meta-knowledge base. Combine these retrieved meta-knowledge items and the initial information, and let the large model generate a more accurate article topic and topic development logic.
[0078] Third call: If there are files uploaded by the user (already stored in the temporary database table), use the RAG technology to retrieve key supplementary information related to the current topic from them. Combine the article information (topic, logic, audience, etc.) generated in the previous two steps and the retrieved supplementary information, and finally let the large model generate a detailed article outline.
[0079] For example, the user inputs the research report requirement "The Application and Future Development of Artificial Intelligence in the Manufacturing Industry", with a word count of 1500, the overview and background as described in the original text, and uploads the file "Examples of AI Applications in the Manufacturing Industry.docx".
[0080] First model call: Generate the audience (manufacturing industry practitioners, policymakers, etc.) and keywords (artificial intelligence, manufacturing industry, etc.).
[0081] Second model call: Embed the input information and the generated audience / keywords, and retrieve meta-knowledge items such as "the development logic of the research report topic" and "audience analysis for professionals" from the meta-knowledge base. Combine this information to generate the topic (The Application and Future Development of Artificial Intelligence in the Manufacturing Industry) and the topic development logic (discuss applications, analyze relationships, challenges and impacts, provide suggestions, etc.).
[0082] Third model call: Embed information such as the topic and logic, and at the same time retrieve the chunked content of "Examples of AI Applications in the Manufacturing Industry.docx" in the temporary database, and find examples related to each part of the outline. Combine all the information to generate the following detailed outline example.
[0083] Introduction: Research background and purpose, overview of the report structure.
[0084] Current situation of the development of artificial intelligence technology: The basic concept of artificial intelligence technology, the main areas of technological progress in recent years, and existing cases of the application of artificial intelligence in the manufacturing industry.
[0085] Current situation and development trend of the manufacturing industry: Challenges currently faced by the manufacturing industry, trend analysis of changes in market demand, and the necessity of the transformation and upgrading of the manufacturing industry.
[0086] Application potential of artificial intelligence in the manufacturing industry: Applications in the production link, role in supply chain management, and impact on quality control and maintenance.
[0087] Market demand and policy environment analysis: Overview of the market demand for artificial intelligence, support of relevant policies and regulatory background, and impact of the policy environment on the development of the manufacturing industry.
[0088] Challenges faced: Obstacles to technology implementation, data privacy and security issues, and employee skills and training needs.
[0089] Strategic recommendations: formulate long-term technology development plans, strengthen cooperation with scientific research institutions, and cultivate talents that meet the needs.
[0090] Conclusion: The profound impact of artificial intelligence on manufacturing, summarizes the main findings and suggestions of the research, and looks forward to future development directions.
[0091] Step 32: Based on the article outline and article summary, use retrieval enhancement generation technology to obtain meta-knowledge and writing methods related to each paragraph to generate article writing content.
[0092] It should be noted that article writing content can be divided into two situations: short text and long text: Short text generation (for example, not exceeding a preset threshold, such as 2,000 words): It can be generated in one go. Using RAG technology, the complete outline is embedded, and meta-knowledge items related to "article structure" and "logical coherence" are retrieved from the meta-knowledge base, as well as paragraph writing meta-knowledge applicable to each outline point (such as "paragraph structure", "quotation", "transition sentence", etc.) and similar paragraph examples. At the same time, relevant reference content is retrieved from the temporary database. The outline, retrieved meta-knowledge items, paragraph examples and reference content are provided to the large language model to generate a complete short text article.
[0093] Long text generation (for example, exceeding a preset threshold, such as 2,000 words): adopt a step-by-step generation and reintegration approach. First, for each section (or subsection) title and its overview / key points in the outline, a separate RAG call is made. The title and key points of the section are embedded, and relevant paragraph-level meta-knowledge (such as specific paragraph structure, tone, vocabulary suggestions) and similar paragraph examples are retrieved in the meta-knowledge base, and relevant content is retrieved from the temporary database. Call the large language model to generate specific content for the section. Repeat this process for all sections in the outline to obtain the content of each part. Finally, call the large model again, use RAG technology to embed the entire outline, retrieve meta-knowledge items related to "overall structure of the article", "logical coherence", and "chapter transition", integrate, optimize, and add transitions to the previously generated content of each part to ensure that the entire long text is logically smooth and structurally complete.
[0094] For example, if the 1,500 words required by the user is less than the threshold, the short text generation method is used. The outline is embedded, and the relevant article structure meta-knowledge, paragraph writing meta-knowledge (such as the introduction paragraph structure, data citation method, conclusion paragraph writing method, etc.), and examples in the temporary library are retrieved to generate a full report of about 1,500 words at one time.
[0095] If the user requests 5,000 words, the long text generation method is adopted. For each section in the outline, such as "Introduction" and "Current Status of Artificial Intelligence Technology Development", RAG calls are made separately: embed the section title and key points, retrieve relevant paragraph meta-knowledge and examples, and generate the content of this section. For example, when generating the section "Current Status of Artificial Intelligence Technology Development", the paragraph structure meta-knowledge about "Technical Concept Explanation" and "List of Cases" may be retrieved, and relevant AI application examples are found from the temporary library. After all section contents are generated, another RAG call is made: embed the overall outline, retrieve the article structure and logical coherence meta-knowledge, and integrate and optimize the contents of each section into a complete 5,000-word report.
[0096] Embodiment 2 Please refer to Figure 2 , which shows a schematic structural diagram of a large model writing system based on a structured meta-knowledge base proposed in the second embodiment of the present application. The system includes: A meta-knowledge base construction module 100, configured to perform structured learning on a target article, extract meta-knowledge of the article and paragraphs for storage, so as to construct an article writing meta-knowledge base; A knowledge base update module 200, configured to update the quality and usage amount of meta-knowledge items in the article writing meta-knowledge base according to a knowledge base management mechanism, so as to obtain a final structured meta-knowledge base; An article content generation module 300, configured to generate an outline and outline content based on the structured meta-knowledge base according to the input article summary and relevant constraints, so as to integrate and obtain article writing content.
[0097] A large model writing system based on a structured meta-knowledge base in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0098] An large model writing system based on a structured meta knowledge base in an embodiment of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0099] An large model writing system based on a structured meta knowledge base provided in an embodiment of the present application can implement Figure 1 each process implemented in a method embodiment of an large model writing method based on a structured meta knowledge base. To avoid repetition, it will not be described in detail here.
[0100] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of an large model writing method based on a structured meta knowledge base and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0101] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above method embodiment of an large model writing method based on a structured meta knowledge base and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0102] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0103] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0104] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0105] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can still make many forms, all of which fall within the protection scope of the present application.
Claims
1. A large model writing method based on a structured meta-knowledge base, characterized in that: include: Step 1: Perform structured learning on the target article, extract meta-knowledge of the article and paragraphs and store them to build a meta-knowledge base for article writing; Step 2: updating the quality and usage of meta-knowledge items in the article writing meta-knowledge base according to the knowledge base management mechanism to obtain a final structured meta-knowledge base; Step 3: Based on the structured meta-knowledge base, an outline and outline content are generated according to the input article summary and related constraints to integrate the article writing content.
2. The large model writing method based on a structured meta-knowledge base according to claim 1 is characterized in that: Step 1 includes: Step 11: Select the target article from the original knowledge base and perform core analysis and summary and paragraph splitting and summary to obtain article summary and paragraph summary respectively; Step 12: extracting meta-knowledge information based on the article summary and the paragraph summary, wherein the meta-knowledge information includes meta-knowledge of the article and meta-knowledge of the paragraph; Step 13: Perform similarity matching between the meta-knowledge information and the meta-knowledge in the original knowledge base and store them in a database; Step 14: vectorize the meta-knowledge of the database to obtain an article writing meta-knowledge base.
3. The large model writing method based on structured meta-knowledge base according to claim 2 is characterized in that: In step 12, the meta-knowledge of the article includes audience analysis, topic development logic, article structure and logical coherence, and the meta-knowledge of the paragraph includes paragraph structure, grammar, vocabulary, quotations, transition sentences, rhetoric, sentence patterns, tone and intonation, and information supplementation.
4. The large model writing method based on a structured meta-knowledge base according to claim 1 is characterized in that: In step 2, the knowledge base management mechanism includes: Step 21, updating the quality score of the meta-knowledge item, each meta-knowledge item being a unit representation of a meta-knowledge system framework; Step 22: Count the usage frequency of the meta-knowledge item in the target article using retrieval enhancement generation technology; Step 23: When the quality score and the usage frequency of the meta-knowledge item are respectively lower than the first threshold and the second threshold, the meta-knowledge item is deleted and the paragraph text of the meta-knowledge item is transferred to a similar high-quality meta-knowledge item.
5. The large model writing method based on structured meta-knowledge base according to claim 4 is characterized in that: Step 22 includes: Set the initial quality score and usage frequency of the newly added meta-knowledge items to 0, record the usage frequency of each meta-knowledge item, and increase the usage frequency of the meta-knowledge item by 1 each time an article is generated. Each use of the retrieval enhancement generation technology statistics will increase the usage frequency of the meta-knowledge item.
6. The large model writing method based on structured meta-knowledge base according to claim 1 is characterized in that: Step 3 includes: Step 31: Use the large language model and retrieval enhancement generation technology to construct the article outline and article summary from the key information of the article writing meta-knowledge base; Step 32: Based on the article outline and the article summary, using retrieval enhancement generation technology, obtain meta-knowledge and writing methods related to each paragraph to generate article writing content.
7. The large model writing method based on structured meta-knowledge base according to claim 6 is characterized in that: In step 32, the article writing content includes short text and long text. The short text is a text of no more than two thousand words, which is generated by performing paragraph writing meta-knowledge retrieval using the article writing meta-knowledge base and retrieval enhancement generation technology. The long text is a text of more than two thousand words, which is obtained by integrating and optimizing the specific contents of all section titles and overviews of the article outline generated by the retrieval enhancement generation technology.
8. A large model writing system based on a structured meta-knowledge base, characterized in that: The system comprises: The meta-knowledge base construction module is used to perform structured learning on the target article, extract meta-knowledge of articles and paragraphs for storage, and build a meta-knowledge base for article writing; A knowledge base updating module, used to update the quality and usage of meta-knowledge items in the article writing meta-knowledge base according to the knowledge base management mechanism to obtain a final structured meta-knowledge base; The article content generation module is used to generate an outline and outline content based on the structured meta-knowledge base and the input article summary and related constraints to integrate and obtain the article writing content.
9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a large model writing method based on a structured meta-knowledge base as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps of a large model writing method based on a structured meta-knowledge base as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Method and device used for generating article
CN106970898A
Intelligent writing method, device, equipment and medium
CN115630640A
Article generation method and apparatus based on large language model, and electronic device
CN117709295A
Article automatic generation method and device, equipment and storage medium
CN117744620A
Long text generation method, device and equipment based on large language model
CN118536502A