Content generation method and device, computing equipment and medium
By integrating large language models and local knowledge bases, a comprehensive knowledge analysis and generation system is built, the integration problem between large models and local knowledge bases is solved, and local data processing is realized without cloud dependencies is achieved, the security and efficiency of content generation is improved, and the efficient content output of enterprises in complex scenarios is met.
Patent Information
- Application Number
- CN202510820415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, it is difficult for large models to deeply integrate with the enterprise's local knowledge base during content generation, resulting in low knowledge management efficiency and inability to meet the needs of diversified scenarios; the high dependence on the cloud leads to data privacy risks, and the intelligent document generation capabilities are insufficient, making it difficult to meet the enterprise's efficient content output needs in complex scenarios.
By deeply integrating the general large language model and local knowledge base, a comprehensive knowledge analysis, management, retrieval and generation system is built, and a fully privatized deployment solution is adopted to deploy AI functions in PC clients or embedded hardware devices, realizing cloud-free local data storage and processing. Combining the user's local knowledge base and large-model generation capabilities, it supports dynamic template adjustment and content customization.
It realizes accurate and efficient customized knowledge services, ensures data privacy and business security, improves document generation efficiency and quality, and meets the efficient content output needs of enterprises in complex scenarios such as bids, reports and solutions.
Smart Images

Figure CN120353901A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a content generation method, apparatus, device, and medium based on a local knowledge base and a large model. Background Art
[0002] With the rapid development of artificial intelligence technology, the demand for content generation in various industries (such as slide PPTs, tender documents, or solution documents, etc.) is increasing day by day. Currently, using large models to process and generate content required in various industries has been widely applied and is still developing at a high speed, and is even gradually applied to fields with higher data privacy requirements.
[0003] Therefore, improving the generation and processing ability of current large models, making the content output by the large model as close as possible to the content actually required, and having high data privacy, is a current research focus. Summary of the Invention
[0004] According to one aspect of the present disclosure, there is provided a content generation method, including: obtaining task requirement information related to a content generation task, and performing feature extraction on the input text information corresponding to the task requirement information to determine problem features; obtaining auxiliary knowledge information from a local knowledge base according to the problem features, where the local knowledge base is constructed based on documents in different formats; and using a locally deployed large language model to generate content based on the problem features and with reference to the auxiliary knowledge information.
[0005] According to an embodiment of the present disclosure, the content generation method may further include: obtaining a target template related to the content generation task; where the large language model further refers to the target template when generating the content.
[0006] According to an embodiment of the present disclosure, where the problem features include vector features generated from the input text information, and obtaining auxiliary knowledge information from the local knowledge base according to the problem features includes: matching the vector features with a vector database in the local knowledge base representing at least part of the local knowledge to obtain a matching text vector; and obtaining the matching text vector from the local knowledge base as at least part of the auxiliary knowledge information.
[0007] According to an embodiment of the present disclosure, wherein the problem features include keyword features in the input text information, and obtaining auxiliary knowledge information from the local knowledge base according to the problem features includes: based on the problem features, performing keyword matching on all the words in the documents representing at least part of the local knowledge in the local knowledge base to obtain matching text fragments; and obtaining the matching text fragments from the local knowledge base as at least part of the auxiliary knowledge information.
[0008] According to an embodiment of the present disclosure, wherein the local knowledge base is constructed through the following steps: converting the obtained documents in different formats into text information, and segmenting the corresponding text information according to a text segmentation strategy to obtain each segmented text slice; and using an embedding model to convert each text slice into a text vector and storing it in a vector database that supports a vector search engine to form the local knowledge base.
[0009] According to an embodiment of the present disclosure, wherein the local knowledge base further includes a knowledge graph, and the knowledge graph is constructed through the following steps: converting the obtained documents in different formats into text information, and extracting entities, entity attributes, and relationship information between entities in the knowledge graph from the text information; and constructing the knowledge graph based on the extracted entities, the entity attributes, and the relationship information.
[0010] According to an embodiment of the present disclosure, wherein extracting entities, entity attributes, and relationship information between entities in the knowledge graph from the text information includes: using a pre-trained model to extract multiple entities and their attributes and relationship information between the multiple entities from the text information, wherein the pre-trained model includes one or more of the large language model, long short-term memory network (LSTM), or BERT model.
[0011] According to an embodiment of the present disclosure, wherein obtaining auxiliary knowledge information from the local knowledge base according to the problem features further includes: based on the knowledge graph, retrieving entities, entity attributes, and relationship information associated with the entities corresponding to the problem features from the local knowledge base as at least part of the auxiliary knowledge information.
[0012] According to an embodiment of the present disclosure, wherein obtaining a target template related to the content generation task includes: sending template indication information associated with the content generation task to a template library, the template indication information at least including the type of the content generation task; and obtaining a target template corresponding to the type of the content generation task from the template library.
[0013] According to an embodiment of the present disclosure, generating content by using a locally deployed large language model based on the problem features and referring to the auxiliary knowledge information includes: forming a context prompt for the locally deployed large language model by using the retrieved auxiliary knowledge information, the target template, and the task requirement information; and inputting the context prompt into the large language model so that the large language model generates and outputs the content with reference to the context prompt.
[0014] According to an embodiment of the present disclosure, generating content by using a locally deployed large language model based on the problem features and referring to the auxiliary knowledge information includes: forming a context prompt for the locally deployed large language model by using the retrieved auxiliary knowledge information and the problem features; and inputting the context prompt into the large language model so that the large language model generates and outputs the content with reference to the context prompt; or inputting the retrieved auxiliary knowledge information into the locally deployed large language model so that the large language model summarizes the auxiliary knowledge information and generates and outputs the content based on the summarized content and the problem features; or determining the content, format, and template of the retrieved auxiliary knowledge information, and generating and outputting the content according to at least one of the content, format, and template of the auxiliary knowledge information and for the problem features.
[0015] According to an embodiment of the present disclosure, the content generation method may further include: after generating the content, obtaining modification indication information associated with the content generation task, converting the modification indication information into text information to obtain additional problem features; and using the large language model to further modify the generated content based on the additional problem features.
[0016] According to an embodiment of the present disclosure, where the local knowledge base includes text information converted from documents in different formats and corresponding storage paths, the content generation method further includes: in response to an operation mode of keyword query, retrieving a text segment matching the keyword included in the input text information from the local knowledge base, and generating a storage path of the document corresponding to the retrieved text segment; or analyzing the problem features corresponding to the input text information to determine the user intention, and retrieving a document that meets the user intention from the local knowledge base according to the user intention.
[0017] According to another aspect of the present disclosure, there is also provided a content generation device, which includes: a feature extraction module configured to obtain task requirement information related to a content generation task and perform feature extraction on the input text information corresponding to the task requirement information to determine problem features; an acquisition module configured to obtain auxiliary knowledge information from a local knowledge base according to the problem features, where the local knowledge base is constructed based on documents in different formats; and a generation module configured to use a large language model deployed locally to generate content based on the problem features and with reference to the auxiliary knowledge information.
[0018] According to another aspect of the present disclosure, there is also provided a computing device, including: a processor; and a memory storing a computer program, a local knowledge base, and a template library thereon, and when the computer program is executed by the processor, the processor and the processor of the embedded computing device jointly execute the content generation method as described above.
[0019] According to another aspect of the present disclosure, there is also provided an embedded computing device, including: a processor; and a memory storing a computer program, a local knowledge base, and a template library thereon, and when the computer program is executed by the processor, the processor executes the content generation method as described above.
[0020] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, a local knowledge base, and a template library thereon, and when the computer program is executed by a processor, the processor executes the content generation method as described above.
[0021] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, where the computer program runs an agent application platform when executed, where one or more agent applications are integrated on the agent application platform, and where the one or more agent applications include a content generation application implemented based on a large language model, and the content generation application is configured to execute the content generation method as described above when executed.
[0022] Through the content generation method of the present disclosure embodiment, by deeply integrating a general large language model (LLM) with a local knowledge base, a comprehensive knowledge parsing, management, retrieval, and generation system is constructed, enabling efficient parsing and structured storage of various formats of documents (such as PPT, PDF, Word, pictures, etc.), and using advanced retrieval technologies (such as the vector retrieval and / or keyword retrieval mentioned above), combined with the generation ability of the large model, it can provide users with accurate and efficient customized knowledge services. In addition, by adopting a completely privatized deployment solution (including local knowledge base management and the offline inference ability of the large language model), by deploying AI functions in PC clients or embedded hardware devices, local data storage and processing without cloud dependence are realized, ensuring data privacy and business security, and by optimizing the lightweight large language model and its deployment method, it can support efficient operation on low-power and low-computing embedded devices, thus meeting the strict requirements for data privacy, computing power limitations, and offline application scenarios. By combining the user's local knowledge base and the large model generation ability, it can support template dynamic adjustment, format optimization, and content customization functions based on industry needs, thus significantly improving the document generation efficiency and quality. This intelligent document generation ability can meet the high-efficiency content output requirements of enterprises in complex scenarios such as tender documents, reports, and proposal writing, while ensuring the high adaptability and personalization of the output results. Additionally, the content generation application can be used as an intelligent agent application on the intelligent agent application platform, so that it can interact or coordinate with other integrated intelligent agent applications, thus enabling better content generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings illustrate various embodiments of aspects of the present disclosure, and they are used together with the description to explain the principles of the present disclosure. Those skilled in the art in this technical field understand that the specific embodiments shown in the drawings are merely exemplary and they are not intended to limit the scope of the present disclosure. In the drawings: Figure 1 Shows an architecture diagram of a system for implementing content generation according to an embodiment of the present disclosure.
[0024] Figure 2 Shows a flowchart of a content generation method according to an embodiment of the present disclosure.
[0025] Figure 3 Shows an example process of generating the complete content of a slide PPT using an offline large model based on a local knowledge base according to an embodiment of the present disclosure.
[0026] Figure 4 Shows a schematic diagram of the deployment architecture of a large language model for implementing content generation according to an embodiment of the present disclosure.
[0027] Figure 5A schematic diagram showing another deployment architecture of a large language model for content generation according to an embodiment of the present disclosure.
[0028] Figure 6 A structural block diagram of a content generation device according to an embodiment of the present disclosure.
[0029] Figure 7 A schematic block diagram of a computing device according to an embodiment of the present disclosure. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0031] As mentioned above, current large models have been widely applied to the content generation field, such as for generating answers to questions, slide PPTs, tender documents or solution documents, etc. At the same time, the demand for intelligent knowledge management and content generation in all walks of life is increasing day by day, especially in the fields of private deployment, intelligent question answering and automatic document generation. However, the existing technology solutions still have the following prominent pain points in meeting the enterprise needs: 1. Insufficient integration of local private knowledge bases When enterprises process and manage a large amount of business, technical and customer knowledge, they often face problems such as "data islands", incoherent information and low query efficiency. Moreover, existing AI large models are difficult to deeply integrate with enterprise local knowledge bases, thus unable to efficiently support real-time question answering, accurate retrieval and personalized knowledge applications. This results in low efficiency of enterprise knowledge management and utilization, and it is difficult to meet the diverse scenario requirements.
[0032] 2. High dependence on the cloud The AI services on which traditional large models are based usually rely on cloud computing and network interactions, which pose significant application barriers to industries with high requirements for data privacy. These industries have strict requirements for data security, business privacy and system independence, and the centralized processing mode of AI services in the cloud easily leads to the risk of privacy leakage. In addition, cloud dependence limits the application potential of AI technology under offline or low computing power conditions, and it is difficult to meet the actual needs of enterprises for autonomous operation and independent control.
[0033] 3. Insufficient intelligent document generation ability Existing document generation tools (such as tender or report generation tools) have obvious deficiencies in terms of flexibility and intelligence, resulting in a cumbersome and inefficient document writing process, and it is difficult to guarantee the quality of the generated results. At the same time, existing tools have limited performance in adapting to the specific needs of enterprises and achieving high customization, and cannot effectively support the document generation needs of enterprises in complex scenarios, such as the rapid generation and optimization of tender bids, reports or technical solutions.
[0034] In summary, enterprises urgently need an innovative solution that integrates private deployment, in-depth integration of knowledge bases, intelligent question answering and document generation to solve the key problems in current technologies and improve the efficiency, security and customization level of knowledge management and content generation.
[0035] Figure 1 The architecture diagram of the system for implementing content generation according to an embodiment of the present disclosure is shown.
[0036] For example, the system 100 may include a content generation device 110 and a user interface 120. The content generation device 110 may have processing and storage functions, for example, including one or more processors and one or more memories.
[0037] The memory may store various computer programs and various data / information, and when the computer program is executed, the process of content generation based on a large language model may be implemented. The user may interact with the content generation device 110 through the user interface to input task requirement information about the content to be generated (generate answers to the questions asked, generate tender bids, generate complete PPT documents or other solution documents, etc.) to the content generation device 110, and the content generation device 110 may execute the content generation process for the task requirement information and output the generated content through the user interface after generating the corresponding content. During the process of content generation by the content generation device 110, the user may interact with the content generation device 110 in one or more rounds, so that the finally generated content better meets the user's needs.
[0038] More implementation details of the operation of the content generation device 110 will be described in detail later.
[0039] Figure 2 The flowchart of the content generation method according to an embodiment of the present disclosure is shown. This content generation method may be executed by Figure 1 the content generation device 110 shown.
[0040] As Figure 2 shown, in step S210, task requirement information related to the content generation task is obtained, and feature extraction is performed on the input text information corresponding to the task requirement information to determine the problem features.
[0041] For example, the content generation task may include generating answers to the questions asked, abstracts, complete PPT documents, tender documents, and / or solution documents, etc. The user can input task requirement information about the content to be generated through the user interface. The user can provide task requirement information through text input or voice input, and when using voice input, the input voice can be first converted into text to obtain the input text information. Then, feature extraction can be performed on the input text information. For example, keyword features or vector features can be extracted as question features. For example, keyword extraction can be implemented using existing algorithms (such as keyword extraction based on Word2Vec word clustering, keyword extraction based on statistical features (TF, TF-IDF), keyword extraction based on word graph models, or keyword extraction based on topic models (LDA), etc.). In addition, the vector features can also be obtained by performing feature extraction on the input text information based on existing feature extraction algorithms (such as the bag-of-words model) and converting the extracted features into numerical vectors, where the text feature representation methods can include sparse vector representation, word embedding, and / or sentence and document embedding (for example, implemented using embedding models such as the BERT model).
[0042] Optionally, for different application scenarios, the task requirement information may include some descriptions of the content to be generated. For example, for the tender scenario, the task requirement information may include information descriptions of the project to be tendered for.
[0043] In step S220, auxiliary knowledge information is obtained from the local knowledge base according to the question features, where the local knowledge base is constructed based on documents in different formats.
[0044] For example, the data sources of the local knowledge base may include network data and / or local document data. For example, the user can upload local documents in different formats (such as PPT, PDF, Word, pictures, etc.) or import online data (for example, by providing a network link, the search engine associated with the local knowledge base will obtain relevant text content from the network). Optionally, when the local document is in PDF or picture format, the document parsing module associated with the local knowledge base can first convert the PDF or picture into text format.
[0045] For example, when constructing the local knowledge base based on documents in different formats, the obtained documents in different formats can be converted into text information, and the corresponding text information can be segmented according to the text segmentation strategy to obtain each segmented text slice. Then, an embedding model is used to convert each text slice into a text vector and store it in a vector database that supports a vector search engine to form the local knowledge base. Optionally, the text vectors are stored in the local knowledge base in association with the corresponding text slices.
[0046] Optionally, when splitting text information, the text splitting strategy may include one or more of paragraph markers, maximum paragraph length (e.g., the number of tokens), paragraph overlap length, and text preprocessing rules.
[0047] As an example, an embedding model of a pre-trained model such as BERT or SentenceTransformers can be used to generate text embeddings (vectors). According to the application scenario, the pre-trained model can also be fine-tuned to improve the accuracy in a specific domain. Then, the generated vectors can be stored in a high-performance vector database such as Pinecone or Weaviate to support fast similarity search for subsequent matching processes. Another example is that the Nomic-Embed-Text vector model can be used to segment, encode, transform, and store each text into a vector database. By converting each piece of knowledge into a point in a high-dimensional vector space, similarity calculation, classification, clustering, retrieval, etc. can be performed.
[0048] Optionally, after constructing the local knowledge base, during subsequent use, in response to the operation information of the knowledge in the local knowledge base, operations such as addition, deletion, or modification can be performed on the local knowledge base. That is, the user can operate the local knowledge base through the user interface. For example, when new knowledge needs to be added, the user can upload a new document, and the associated document parsing module can parse the uploaded document and perform the same splitting process to generate text vectors, and then store the text vectors in the local knowledge base in association with the corresponding text slices.
[0049] Optionally, according to the selection of the operation mode (e.g., content generation mode and content search mode) input by the user, for example, the constructed local knowledge base can be operated according to the selected operation mode, so multi-mode switching can be achieved to meet the usage habits of different users. For example, in the case where the operation mode is the content generation mode, the matching content can be obtained by matching against the local knowledge base as will be described later, and the matching content can be used for the large language model to generate corresponding content. In the case where the operation mode is the content search mode, at this time, keyword-driven search can be realized, and content or files related to the keyword can be searched in the local knowledge base, and / or the storage path (storage location) where the searched content or file is located can be output, so the efficient file and content location function based on keywords can be realized, and the user can quickly find relevant documents and knowledge points. In other embodiments, in addition to keyword extraction or instead of keyword extraction, the large language model can also analyze and understand the input text information to determine the user's intention, so as to find the document that meets the user's needs according to this intention.
[0050] Optionally, the knowledge in the local knowledge base can also be classified and stored in an associated manner according to categories, so that the local knowledge base can include multiple sub-knowledge bases, thereby facilitating the reduction of the amount of calculation in the subsequent matching process.
[0051] For example, when the problem features include vector features generated from the input text information, when obtaining auxiliary knowledge information from the local knowledge base, the vector features can be matched with the vector database representing at least part of the local knowledge in the local knowledge base to obtain a matching text vector; and the matching text vector can be obtained from the local knowledge base as at least part of the auxiliary knowledge information. During the matching process, the matching can be performed based on vector similarity, and the text vector with a vector similarity greater than a preset threshold is used as the matching text vector.
[0052] For example, in response to the user's selection of the type of knowledge in the local knowledge base, the vector features can be matched with the sub-knowledge base (vector database) of the knowledge with the selected category in the local database, thereby obtaining a matching text vector. In this case, there is no need to match the vector features with all the vectors in the local database, so the amount of calculation can be reduced.
[0053] For another example, when the problem features include keyword features in the input text information, when obtaining auxiliary knowledge information from the local knowledge base, keyword matching can be performed on all the words in the document representing at least part of the local knowledge in the local knowledge base based on the problem features to obtain a matching text segment; and the matching text segment can be obtained from the local knowledge base as at least part of the auxiliary knowledge information.
[0054] Similarly, in response to the user's selection of the type of knowledge in the local knowledge base, the keyword features can be matched with the sub-knowledge base (vector database) of the knowledge with the selected category in the local database, thereby obtaining a matching text segment. In this case, the amount of calculation can also be reduced.
[0055] Optionally, in some cases, the local knowledge base can be matched based on vector features and keyword features simultaneously, and the corresponding local knowledge obtained by the matching is used as at least part of the auxiliary knowledge information.
[0056] In addition, to further improve the accuracy of content generation, a knowledge graph can be further introduced. For example, the local knowledge base can include a knowledge graph, and the knowledge graph is constructed through the following steps: converting the obtained documents in different formats into text information, and extracting entities, entity attributes, and relationship information between entities in the knowledge graph from the text information; and constructing the knowledge graph based on the extracted entities, the entity attributes, and the relationship information. Optionally, a pre-trained model can be used to extract multiple entities and their attributes and the relationship information between the multiple entities from this text information. As an example, the pre-trained model can include one or more of a large language model for content generation, a long short-term memory network (LSTM), or a BERT model. In this case, the auxiliary knowledge information obtained from the local knowledge base can also include: based on the constructed knowledge graph, retrieving entities, entity attributes, and relationship information associated with the entities corresponding to the above problem features from the local knowledge base as at least a part of the auxiliary knowledge information. That is to say, by constructing a knowledge graph (GRAPH), in-depth association management of entities, relationships, and contexts in the local knowledge base can be carried out, and the accuracy and relevance of question-and-answer content generation can be significantly improved.
[0057] In step S230, a locally deployed large language model is used to generate content based on the task requirement information and with reference to the auxiliary knowledge information.
[0058] For example, if the task requirement information is to generate an answer to the asked question, the large language model can generate a corresponding answer with reference to the obtained auxiliary knowledge information; if the task requirement information is to generate certain specific documents (such as PPTs, tender documents), the large language model can output documents with higher adaptability and accuracy with reference to the obtained auxiliary knowledge information.
[0059] At this time, during the content generation process, the task requirement information and the auxiliary knowledge information obtained in step S220 can be formed into context prompts for the large language model, and the context prompts are input into the locally deployed large language model, so that the large language model generates and outputs the content with reference to the context prompts.
[0060] Alternatively, during the content generation process, the auxiliary knowledge information retrieved in step S220 can also be input into the locally deployed large language model. Since the large language model also has strong analysis and summarization capabilities, the large language model can summarize the auxiliary knowledge information and generate and output the content based on the summarized content and the problem features. Alternatively, information related to the generated content, such as the content, format, and template of the retrieved auxiliary knowledge information, can be determined, and the content can be generated and output based on at least one of these information (the content, format, and template of the auxiliary knowledge information) and in response to the problem features.
[0061] That is to say, the knowledge stored in the local knowledge base (such as document content) can be utilized. When generating corresponding types of content, content can be generated based on the large model itself, generated based on the summary of the content in the local knowledge base by the large model, or the content, format, template, etc. of the content in the local knowledge base can be directly inherited to ensure the usability of the generated content (document). Such benefits can include, in the case where the content of the user's local knowledge base is relatively large, accurately inheriting the content required by the user to generate corresponding content that better conforms to the user's habits, and eliminating the need for the user to search for it themselves.
[0062] In addition, the large language model in the embodiments of the present disclosure is offline and can be lightweight, so as to be able to work properly under low computing power. That is, different from most current large language models that rely on cloud computing power and are online, through local deployment (offline), it can ensure local data storage and processing without cloud dependence, ensuring data privacy and business security. For example, in some embodiments, the large language model (such as a lightweight one), as well as the above-mentioned local database, template library, and associated text parsing module, indexing and retrieval module, can be integrated in a computing device (i.e., in the form of a server). That is to say, the computing device itself has the ability to generate content, is equipped with artificial intelligence (AI) hardware devices, supports efficient operation under low computing power conditions, and is also called an AI box or an embedded computing device. It can obtain task requirement information through the user interface of other computing devices (such as a personal computer or a user terminal), and perform content generation operations in response to the task requirement information. That is, the AI box runs the large language model, and other computing devices only serve as the operation end. Another example is that in some other embodiments, the large language model is carried on an AI hardware device (also called an AI box, in the form of a pure client), and the above-mentioned local database, template library, and associated text parsing module, indexing and retrieval module can be carried on a computing device (such as a personal computer or a user terminal). That is to say, the computing device and the AI hardware device cooperate to complete the content generation process, and the AI box only provides the computing power of the large language model.
[0063] In addition, when specific complete documents such as PPTs or tender documents need to be generated, considering that these documents generally have template or format requirements, in order to make the generated content more in line with the user's preferences or meet specific format requirements, in addition to the content, format, and template of the auxiliary knowledge information that can be retrieved as described above, the content can also be generated by referring to the target template. For example, the corresponding target template can be obtained from an existing template library.
[0064] Therefore, in this case, the content generation method can also include obtaining a target template related to the content generation task.
[0065] At this time, during the content generation process, the task requirement information, the auxiliary knowledge information obtained in step S220, and the target template obtained in step S230 can be formed into context prompts for the large language model, and the context prompts are input into the locally deployed large language model, so that the large language model generates and outputs the content by referring to the context prompts.
[0066] For example, template indication information associated with the content generation task can be sent to the template library, and the target template corresponding to the type of the content generation task can be obtained from the template library. The template indication information at least includes the type of the content generation task (for example, generating a PPT, generating a tender document, etc.).
[0067] Optionally, the template library can be dynamically adjusted and personalized based on user input. For example, the user can delete, modify, etc. the templates in the template library through the user interface, thus ensuring the high quality and high customization of the generated content.
[0068] Optionally, in some cases, the local knowledge base and the template library can also be interrelated. For example, a part of the documents in different formats in the local knowledge base can be, for example, tender documents, PPTs, etc. Therefore, the templates applied to the tender documents, PPTs, etc. in these local knowledge bases can be stored in the template library for calling during the content generation process.
[0069] Optionally, in some cases, in order to generate the required content more accurately, knowledge can also be obtained from an external knowledge base as part of the auxiliary indication information for reference during the content generation process. For example, a query request can be sent to the external knowledge base (the query request can similarly include problem features), and the associated index and retrieval engine of the external knowledge base can be used to retrieve the matching content (such as text vectors or text fragments) in the external knowledge base, so that the matching content can be obtained from the external knowledge base.
[0070] Optionally, to further improve the formatting accuracy of the generated document, after the large language model generates a document associated with the content generation task, a format checking tool can be called to optimize and check the format of the generated content (e.g., paragraph font typesetting check, automatic revision, or intelligent optimization, etc.). In addition, if the generated content is a tender, a compliance checking tool can be called to check the compliance of the tender, thus ensuring the high quality and compliance of the generated document.
[0071] Optionally, to further improve the accuracy of the generated content, modification instruction information associated with the content generation task can be further obtained (e.g., further input by the user through the user interface), and it can be converted into text information again to obtain additional question features, and the content that has been generated can be further modified and adjusted based on the additional question features. Optionally, additional auxiliary knowledge information can be further obtained from the local knowledge base according to the additional question features and used for content modification.
[0072] Therefore, through the content generation method of the present disclosure embodiment, at least the following key technical problems can be solved. First, the problem of the efficient combination of a private local knowledge base and a general large model can be solved. For example, by deeply integrating a general large language model (LLM) with the local knowledge base, a comprehensive knowledge parsing, management, retrieval, and generation system can be constructed, enabling efficient parsing and structured storage of various formats of documents (such as PPT, PDF, Word, pictures, etc.), and using advanced retrieval technologies (such as the vector retrieval and / or keyword retrieval described above), combined with the generation ability of the large model, precise and efficient customized knowledge services can be provided for users. Second, the problem of offline private deployment and data privacy protection can be solved. For example, by adopting a completely private deployment solution (including local knowledge base management and the offline inference ability of the large language model), by deploying AI functions in PC clients or embedded hardware devices, local data storage and processing without cloud dependence are achieved, ensuring data privacy and business security, and by optimizing the lightweight large language model and its deployment method, efficient operation on low-power and low-computing-power embedded devices can be supported, thus meeting the strict requirements for data privacy, computing power limitations, and offline application scenarios. Finally, the problem of insufficient intelligent document generation ability can also be solved. For example, by combining the user's local knowledge base and the large model generation ability, an intelligent generation process for content can be constructed, and functions such as dynamic adjustment of templates based on industry requirements, format optimization, and content customization can be supported, thus significantly improving the efficiency and quality of document generation. This intelligent document generation ability can meet the high-efficiency content output requirements of enterprises in complex scenarios such as tender writing, report writing, and solution writing, while ensuring the high adaptability and personalization of the output results.
[0073] Figure 3An example process for generating the complete content of a slide PPT using an offline large model based on a local knowledge base according to an embodiment of the present disclosure is shown.
[0074] As Figure 3 shown, regarding the construction of the local knowledge base, the user can upload documents in different formats (such as PPT, PDF, Word, pictures, etc.) or voice materials that can be converted into text (collectively referred to as text information) or download text information from the network through an external link, and perform parsing processing by using a file parsing module. For example, it can include parsing it into a text format and / or identifying the table of contents and headings, etc., and then slicing the text information in text format according to a text segmentation strategy (such as dividing by paragraphs and character lengths, etc.) to obtain text slices. Then, an index can be constructed for the processed text slices. In addition, text vectors corresponding to each text slice can also be generated to form a vector database as at least a part of the local knowledge base. In addition, the local knowledge base generally also has an associated index and retrieval engine (including a vector search engine), so that it can be used to obtain the knowledge required during the content generation process from the local knowledge base.
[0075] Regarding the content generation process, here generating a complete PPT document is taken as an example. According to some embodiments, as Figure 3 shown, relevant matching knowledge can be obtained from the local knowledge base through a simple keyword matching process, and an initial PPT outline can be generated based on the relevant matching knowledge. Considering the accuracy of the keyword matching process, the best-matched knowledge can be obtained through multiple rounds of conversations, and a final PPT outline can be generated. In this process, the PPT outline is generated based on the local knowledge base, and instead of using a large language model, existing intelligent document generation tools can be used. After obtaining the PPT outline, in some aspects, existing intelligent document generation tools can be used to generate the final complete PPT document according to this PPT outline. In other aspects, the PPT outline can be used as a context prompt for the large language model to enable the large language model to generate the final complete PPT document.
[0076] Therefore, according to whether to generate the complete PPT document based on a large language model, Figure 3 the content generation process shown can have two generation modes, which can be selected by the user. The embodiments of the present disclosure mainly focus on the process of generating the required content based on a large language model.
[0077] For example, different from Figure 3The PPT generation process shown, in an embodiment of the present disclosure, replaces the existing intelligent document generation tool and uses a large language model to generate a complete PPT document. For example, it can generate task-related task requirement information based on the input content, and extract the problem features of the input text information in the task requirement information, and match the problem features against the local knowledge base (such as vector matching or keyword matching) to obtain relevant knowledge, where the relevant knowledge (as described above, may further include information obtained from the knowledge graph), the problem features, and the optionally obtained target template are jointly used as context prompts for the large language model to generate a complete PPT document. Optionally, after generating the complete PPT document, it can further obtain modification instruction information from the user, and convert it into text information again to obtain additional problem features, and can further modify and adjust the generated content based on the additional problem features. Optionally, it can further obtain additional auxiliary knowledge information from the local knowledge base according to the additional problem features and use it for content modification and improvement.
[0078] Figure 4 Shows a schematic diagram of the deployment architecture of a large language model for implementing content generation according to an embodiment of the present disclosure.
[0079] As Figure 4 Shown, the AI box (the first computing device) runs the large language model, and other computing devices (the second computing device) only serve as the user operation end. Specifically, the large language model (for example, lightweight), as well as the above-mentioned local database, template library, and associated text parsing module, indexing and retrieval module, can be integrated in one computing device (i.e., in the form of a server), that is, the computing device itself has the ability to generate content, and its built-in artificial intelligence (AI) hardware supports efficient operation under low computing power conditions. Therefore, this computing device is also called an AI box or an embedded computing device, which can obtain task requirement information through the user interface of the second computing device (such as a personal computer or a user terminal) serving as the user operation end, and execute content generation operations in response to the task requirement information. The AI box can perform data / information transfer with the second computing device via the application programming interface API.
[0080] For example, the second computing device ( Figure 4A client user interface (UI) is installed on a second computing device (shown as a PC in the figure). The user can input through this user interface, so that the second computing device can generate task requirement information based on the user input and provide the task to the large language model (LLM) in the AI box via the task scheduling module associated with the AI box. After receiving the task, the large language model (LLM) can start the indexing and retrieval module associated with the local knowledge base to perform matching processing on the local knowledge base (such as the vector matching or keyword matching described above), so that the large language model can obtain matching knowledge from the local knowledge base (and / or further identify relevant entities, entity attributes, and relationships from the knowledge graph). On the other hand, the large language model can also call the template library, for example, send template indication information to the template library to obtain the target template corresponding to the content generation task from the template library. In this way, the large language model can use the task requirement information, matching knowledge, and target template as context prompts to generate the required content. The large language model can return the generated content to the client user interface at the second computing device via the task scheduling module to present the generated content to the user.
[0081] In addition, the AI box (the first computing device) can also include a document parsing module for generating and / or updating the local knowledge base based on the obtained documents in different formats, as described in detail above. In addition, the user can also operate on the local knowledge base in the first computing device through the user interface and task scheduling module of the second computing device, such as adding, deleting, or managing, etc. For example, the document parsing module can obtain the original document from the local knowledge base, perform parsing processing on it, and then return the result of the parsing processing to the local knowledge base.
[0082] Figure 5 The figure shows a schematic diagram of another deployment architecture of a large language model for implementing content generation according to an embodiment of the present disclosure.
[0083] As Figure 5 shown, the AI box only provides the computing power of the large language model. Specifically, the large language model is carried on the AI hardware device (also called the AI box), and the above-mentioned local database, template library, and associated text parsing module and indexing and retrieval module can be carried on the computing device (for example, a personal computer or a user terminal). That is to say, the computing device and the AI hardware device cooperate to complete the content generation process.
[0084] For example, the computing device ( Figure 5A client user interface (UI) (not shown) is installed on a computing device (hereinafter referred to as PC). Users can input through this user interface, so that the computing device can generate task requirement information according to the user input and provide tasks to the large language model (LLM) in the AI box via the task scheduling module associated with the AI box. The AI box can perform data / information transfer with the computing device via the application programming interface (API). After receiving the task, the large language model (LLM) can start the indexing and retrieval module associated with the local knowledge base in the computing device via the task scheduling module to perform matching processing on the local knowledge base (such as the vector matching or keyword matching described above), so that the large language model can obtain matching knowledge from the local knowledge base (and / or further identify relevant entities, entity attributes, and relationships from the knowledge graph). On the other hand, the large language model in the AI box can also call the template library via the task scheduling module, for example, send template indication information to the template library to obtain the target template corresponding to the content generation task from the template library via the task scheduling module. In this way, the large language model can use the task requirement information, matching knowledge, and target template as context prompts to generate the required content. The large language model in the AI box can return the generated content to the client user interface at the computing device via the task scheduling module to present the generated content to the user.
[0085] Similarly, the computing device may further include a document parsing module for generating and / or updating the local knowledge base based on the obtained documents in different formats, as described in detail above. In addition, the user can also operate on the local knowledge base in the computing device through the user interface and task scheduling module of the computing device, such as adding / deleting / managing, etc. For example, the document parsing module can obtain the original document from the local knowledge base, perform parsing processing on it, and then return the result of the parsing processing to the local knowledge base.
[0086] Therefore, in the embodiments of the present disclosure, the AI box has the ability of rapid deployment. It can be used as a completely independent operating platform or cooperate with the computing device client, and provide flexible application support through the local knowledge base and AI function modules. In this way, local data storage and processing without relying on the cloud are realized, which can ensure data privacy and business security. And by optimizing the lightweight large language model and its deployment method, it can support efficient operation on low-power and low-computing-power embedded devices, thus meeting the strict requirements for data privacy, computing power limitation, and offline application scenarios.
[0087] It should be noted that Figure 4 and Figure 5As shown, various modules (e.g., a file parsing module, an indexing and retrieval module, or a task scheduling module) can be one or more computer programs or instructions executed on a hardware processor, so that the hardware processor can implement corresponding functions.
[0088] According to another aspect of the present disclosure, a content generation device is also provided.
[0089] Figure 6 The structural block diagram of the content generation device according to an embodiment of the present disclosure is shown. The content generation device 600 can be like Figure 1 the content generation device 110 shown.
[0090] As Figure 6 shown, the content generation device 600 can include a feature extraction module 610, an acquisition module 620, and a generation module 630.
[0091] The feature extraction module 610 can be used to obtain task requirement information related to the content generation task, and perform feature extraction on the input text information corresponding to the task requirement information to determine problem features.
[0092] For example, as described above, the content generation task can include generating answers, abstracts, complete PPT documents, tender documents, and / or solution documents, etc. for the questions asked. Optionally, when it is necessary to generate a complete PPT document or a tender document, etc., the acquisition module 620 of the content generation device 600 can also obtain a target template from the template library, so that the generated content can be generated with reference to the target template.
[0093] More details about feature extraction have been described above, so they will not be repeated here. The feature extraction module 610 can correspond to Figure 4 or Figure 5 at least a part of the indexing and retrieval module in
[0094] or an additional module (not shown) independent of the indexing and retrieval module, and is used to perform feature extraction on the input text information corresponding to the task requirement information.
[0095] For example, as described above, when constructing a local knowledge base based on documents in different formats, the obtained documents in different formats can be converted into text information, and the corresponding text information can be segmented according to a text segmentation strategy to obtain each text slice after segmentation. Then, an embedding model is used to convert each text slice into a text vector and store it in a vector database that supports a vector search engine to form the local knowledge base. Optionally, the text vectors are stored in the local knowledge base in association with the corresponding text slices.
[0096] In addition, a matching process (such as keyword matching or vector matching) can be performed on the local knowledge base for the question features generated based on the input text information, so that knowledge associated with the content generation task can be obtained for subsequent content generation.
[0097] The specific process of determining the auxiliary knowledge information from the local knowledge base has been described in detail above, so it will not be repeated here. For example, the acquisition module 620 can be Figure 4 and Figure 5 the indexing and retrieval module in or a module (not shown) associated with a large language model, which is used to implement the matching process to obtain the matching knowledge.
[0098] The generation module 630 can be used to generate content by using a locally deployed large language model based on the question features and referring to the auxiliary knowledge information.
[0099] For example, if the task requirement information is to generate an answer to the question asked, the large language model can refer to the obtained auxiliary knowledge information to generate the corresponding answer; if the task requirement information is to generate certain specific documents (such as PPTs, tender documents), the large language model can refer to the obtained auxiliary knowledge information to output documents with higher adaptability and accuracy.
[0100] In addition, the large language model in the embodiments of the present disclosure is offline and can be lightweight so as to work properly under low computing power. That is, different from most current large language models that are based on cloud computing power and are online, through the local deployment (offline) method, it can be ensured that local data storage and processing without cloud dependence are achieved, ensuring data privacy and business security.
[0101] During the content generation process, the task requirement information, the obtained auxiliary knowledge information, and optionally the target template can be formed into context prompts for the large language model, and the context prompts are input into the locally deployed large language model, so that the large language model refers to the context prompts to generate and output the content. Alternatively, it can also be generated at least partially based on the summary of the content in the local knowledge base by the large model as described above, or based on the content inheritance method.
[0102] The generation module 630 may be Figure 4 and Figure 5 a large language model in
[0103] For more details about the operations performed by each module in the content generation device 600, reference may be made to the description made above Figures 2 to 5 and will not be repeated here.
[0104] In addition, although the above-mentioned modules are shown by way of example in Figure 6 it should be understood that according to different functions, the content generation device 600 may also be divided into more or fewer modules, or each module may be further divided into sub-modules. In some exemplary embodiments, each module or further divided sub-module may be implemented by electronic hardware (e.g., a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware component, etc.), computer software, a program, or an instruction set (e.g., which may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable ROM (EPROM), etc.) or a combination of both.
[0105] Therefore, through the content generation device of the present disclosure, at least the following key technical problems can be solved. First, it can solve the problem of the efficient combination of a private local knowledge base and a general large model. For example, by deeply integrating a general large language model (LLM) with the local knowledge base, a comprehensive knowledge parsing, management, retrieval, and generation system can be constructed, enabling the efficient parsing and structured storage of various formats of documents (such as PPT, PDF, Word, pictures, etc.), and using advanced retrieval technologies (such as the vector retrieval and / or keyword retrieval mentioned above), combined with the generation ability of the large model, precise and efficient customized knowledge services can be provided for users. Second, it can solve the problems of offline private deployment and data privacy protection. For example, by adopting a completely private deployment scheme (including local knowledge base management and the offline inference ability of the large language model), by deploying AI functions on a PC client or an embedded hardware device, local data storage and processing without cloud dependence are achieved, ensuring data privacy and business security. And by optimizing the lightweight large language model and its deployment method, it can support efficient operation on low-power and low-computing-power embedded devices, thus meeting the strict requirements for data privacy, computing power limitations, and offline application scenarios. Finally, it can also solve the problem of insufficient intelligent document generation ability. For example, by combining the user's local knowledge base and the generation ability of the large model, an intelligent generation process from document outlines to complete content can be constructed, and it can support template dynamic adjustment, format optimization, and content customization functions based on industry needs, thus significantly improving the efficiency and quality of document generation. This intelligent document generation ability can meet the efficient content output requirements of enterprises in complex scenarios such as tender documents, reports, and proposal writing, while ensuring the high adaptability and personalization of the output results.
[0106] According to another aspect of the present disclosure, a computing device is also provided.
[0107] Figure 7 A schematic block diagram of a computing device according to an embodiment of the present disclosure is shown. For example, the computing device may be Figure 4 the embedded AI hardware device (the first computing device) shown.
[0108] As Figure 7 shown, the computing device 700 includes one or more processors, one or more memories connected by a system bus, and optionally a network interface, an input device, and a display screen. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the terminal stores an operating system and may also store computer-executable programs or computer-readable codes. When the computer-executable programs or computer-readable codes are executed by the processor, the processor can implement as described above with reference to Figures 2 to 5Various operations of the described method. The internal memory may also store a computer-executable program or computer-readable code, which, when executed by the processor, enables the processor to perform the references as described above Figures 2 to 5 Various operations of the described method. For example, a computer-executable program, a local knowledge base, and a template library may be stored on the memory. When the computer-executable program is executed by the processor, it enables the processor to perform the references as described above Figures 2 to 5 Various operations of the content generation method described above.
[0109] For another example, the computing device may be Figure 5 The computing device shown, separated from the embedded AI hardware device (such as a personal computer PC). In this way, a computer program, a local knowledge base, and a template library are stored on the computing device. When the computer program is executed by the processor of the computing device, it enables the processor to jointly perform the references with the processor of the embedded computing device (such as an AI box including a large language model) Figures 2 to 5 Various operations of the content generation method described above.
[0110] Optionally, the content generation method of the present disclosure may correspond to an intelligent agent application of a content generation application (such as for PPT, tender document generation, and inspection, etc.). In this case, for example, the intelligent agent application platform running on the processor may include one or more intelligent agent applications, and the one or more intelligent agent applications may include a content generation application. The intelligent agent application platform is generally software and can be carried on a computer storage medium. One or more intelligent applications can be integrated into the intelligent agent application platform in different ways. For example, a mini-program application can be integrated into the intelligent agent application platform through a URL, an intelligent agent plugin can be integrated into the intelligent agent application platform by integrating an SDK, etc., and an intelligent agent application program can be integrated into the intelligent agent application platform by integrating an SDK, a URL, etc., and so on. Different intelligent agent applications can also interact and coordinate with each other to jointly optimize their respective functions.
[0111] The processor can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., and can be of the X84 architecture or the ARM architecture.
[0112] The non-volatile memory may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. It should be noted that the memory of the method described in the present disclosure is intended to include, but is not limited to, these and any other suitable categories of memory.
[0113] An optional display screen of the computing device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computing device may be a touch layer covered on the display screen, or may be a button, trackball or touchpad provided on the terminal housing, or may also be an external keyboard, touchpad or mouse, etc.
[0114] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to perform various operations of the content generation method as described above with reference to Figures 2 to 5 the content described.
[0115] According to still another aspect of the present disclosure, there is also provided a computer program product including a computer program, which when executed by a processor, implements various operations of the content generation method as described above with reference to Figures 2 to 5 the content described.
[0116] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the methods and apparatuses according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or part of code that contains at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0117] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A content generation method, comprising: Obtaining task requirement information related to a content generation task, and performing feature extraction on input text information corresponding to the task requirement information to determine problem features; Obtaining auxiliary knowledge information from a local knowledge base according to the problem features, wherein the local knowledge base is constructed based on documents in different formats; And Using a locally deployed large language model to generate content based on the problem features and with reference to the auxiliary knowledge information.
2. The content generation method according to claim 1, further comprising: Obtaining a target template related to the content generation task; Wherein, the large language model further refers to the target template when generating the content.
3. The content generation method according to claim 1, wherein Using a locally deployed large language model to generate content based on the problem features and with reference to the auxiliary knowledge information, including: Forming a context prompt for the locally deployed large language model with the retrieved auxiliary knowledge information and the problem features; and inputting the context prompt into the large language model, so that the large language model generates and outputs the content with reference to the context prompt; or Inputting the retrieved auxiliary knowledge information into the locally deployed large language model, so that the large language model summarizes the auxiliary knowledge information and generates and outputs the content based on the summarized content and the problem features; or Determining the content, format, and template of the retrieved auxiliary knowledge information, and generating and outputting the content according to at least one of the content, format, and template of the auxiliary knowledge information and for the problem features.
4. The content generation method according to claim 1, wherein, The problem features include vector features generated from the input text information, Wherein, obtaining auxiliary knowledge information from a local knowledge base according to the problem features includes: Matching the vector features with a vector database in the local knowledge base that represents at least partially local knowledge to obtain matching text vectors; and Obtaining the matching text vectors from the local knowledge base as at least a part of the auxiliary knowledge information.
5. The content generation method according to claim 1, wherein, The problem features include keyword features in the input text information, Wherein, obtaining auxiliary knowledge information from a local knowledge base according to the problem features includes: Performing keyword matching on all the words in the documents in the local knowledge base that represent at least partially local knowledge based on the problem features to obtain matching text segments; and Obtaining the matching text segments from the local knowledge base as at least a part of the auxiliary knowledge information.
6. The content generation method according to claim 4, wherein, The local knowledge base is constructed through the following steps: Converting the obtained documents in different formats into text information, and segmenting the corresponding text information according to a text segmentation strategy to obtain each segmented text slice; And Using an embedding model to convert each text slice into a text vector and storing it in a vector database that supports a vector search engine to form the local knowledge base.
7. The content generation method according to claim 1, wherein, The local knowledge base further includes a knowledge graph, and the knowledge graph is constructed through the following steps: Convert the obtained documents in different formats into text information, and extract entities, entity attributes, and relationship information between entities in the knowledge graph from the text information; and Construct the knowledge graph based on the extracted entities, the entity attributes, and the relationship information.
8. The content generation method according to claim 7, wherein, Extracting entities, entity attributes, and relationship information between entities in the knowledge graph from the text information includes: Using a pre-trained model to extract multiple entities and their attributes and relationship information between the multiple entities from the text information, wherein the pre-trained model includes one or more of the large language model, long short-term memory network (LSTM), or BERT model.
9. The content generation method according to claim 7, wherein, Obtaining auxiliary knowledge information from the local knowledge base according to the problem features further includes: Based on the knowledge graph, retrieving entities, entity attributes, and relationship information associated with the entities corresponding to the problem features from the local knowledge base as at least a part of the auxiliary knowledge information.
10. The content generation method according to claim 2, wherein, Obtaining a target template related to the content generation task includes: Sending template indication information associated with the content generation task to the template library, where the template indication information at least includes the type of the content generation task; and Obtaining a target template corresponding to the type of the content generation task from the template library.
11. The content generation method according to claim 2, wherein, Using a locally deployed large language model to generate content based on the problem features and referring to the auxiliary knowledge information includes: Forming a context prompt for the locally deployed large language model with the retrieved auxiliary knowledge information, the target template, and the task requirement information; and Inputting the context prompt into the large language model so that the large language model generates and outputs the content with reference to the context prompt.
12. The content generation method according to claim 1, further includes: After generating the content, obtaining modification indication information associated with the content generation task, and converting the modification indication information into text information to obtain additional problem features; and Using the large language model to further modify the generated content based on the additional problem features.
13. The content generation method according to claim 1, wherein, The local knowledge base includes the text information converted from the documents in different formats and the corresponding storage paths, The content generation method further includes: in response to the operation mode of keyword query, retrieving a text segment matching the keyword included in the input text information from the local knowledge base, and generating a storage path of the document corresponding to the retrieved text segment; or Analyzing the problem features corresponding to the input text information to determine the user intention, and retrieving a document that meets the user intention from the local knowledge base according to the user intention.
14. A content generation device includes: A feature extraction module, configured to obtain task requirement information related to a content generation task, and perform feature extraction on the input text information corresponding to the task requirement information to determine problem features; An acquisition module, configured to obtain auxiliary knowledge information from a local knowledge base according to the problem features, where the local knowledge base is constructed based on documents in different formats; and A generation module, configured to use a locally deployed large language model to generate content based on the problem features and with reference to the auxiliary knowledge information.
15. A computing device, comprising: A processor; and A memory, on which a computer program, a local knowledge base, and a template library are stored. When the computer program is executed by the processor, the processor and the processor of the embedded computing device jointly execute the content generation method according to any one of claims 1-13.
16. An embedded computing device, comprising: A processor; and A memory, on which a computer program, a local knowledge base, and a template library are stored. When the computer program is executed by the processor, the processor executes the content generation method according to any one of claims 1-13.
17. A computer-readable storage medium, on which a computer program, a local knowledge base, and a template library are stored. When the computer program is executed by a processor, the processor executes the content generation method according to any one of claims 1-13.
18. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it runs an agent application platform, where one or more agent applications are integrated on the agent application platform, and wherein the one or more agent applications include a content generation application implemented based on a large language model, and the content generation application, when executed, is configured to execute the content generation method according to any one of claims 1-13.
Citation Information
Patent Citations
Knowledge base construction method and question and answer dialogue method and system based on generative large language model
CN117056471A
Knowledge base response method and system based on large language model
CN117951249A
Document generation method and system based on large language model and medium
CN117951272A
LLM-based intelligent questioning and answering method, system and equipment in power field and medium
CN118277521A
Knowledge base construction method for large language model, retrieval method and related device
CN118916441A
Cited By
Content generation method and device
CN121412373A
Vectorization construction method and system for manufacturing equipment maintenance knowledge
CN122240677A