Document generation method, device and equipment
By combining multimodal model and large language model, the initial document and review meeting audio are converted into text to generate high-quality target documents, solving the problems of time and efficiency in document writing, and achieving rapid and efficient document generation and retention of review opinions.
Patent Information
- Application Number
- CN202510688243.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The document writing process in a company or school is complicated and time-consuming, which leads to the long document writing cycle and low efficiency. Especially because the writer is busy with work, it is difficult to retain important content in the review opinions.
By obtaining the initial document, using a multimodal model to convert the image into description text, and combining it with the initial document text, input a large language model for preprocessing, and generating the target document. In addition, automatic speech recognition is used to convert review meeting audio into text, preprocessing is performed and the final document is generated in combination with the initial document content.
It realizes the generation of high-quality target documents without a lot of manual time, saves writers' time, improves document writing efficiency, retains review opinions, and improves document integrity and quality.
Smart Images

Figure CN120197597A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, and equipment for document generation. Background Art
[0002] In enterprises or schools, document writing is a complex and time-consuming process that usually goes through multiple stages. First, the writer has a preliminary idea and forms an initial document. Then, multiple parties review the content of the initial document. After the review, the writer needs to supplement and improve the document content according to the review opinions.
[0003] However, due to the busy work of the document writer himself, there is often not enough time to write detailed document content, resulting in a too long document writing cycle and low efficiency. Therefore, a method that can improve the document writing efficiency is needed to assist the writer in completing the document writing. Summary of the Invention
[0004] This application provides a method, device, and equipment for document generation, which can automatically generate a final target document for the initial draft of the document to improve the efficiency of the writer in writing the document. The technical solutions are as follows: In the first aspect, a method for document generation is provided. The method includes: Obtain an initial document. When the initial document includes an image, input the image into a multi-modal model to obtain an image description text, and combine the text in the initial document and the image description text to obtain a first text, where the image description text is the text used to describe the image. Input the first text into a large language model to preprocess the first text to obtain a first preprocessed text. Generate a target document according to the first preprocessed text.
[0005] In the technical solution provided by this application, a final target document can be automatically generated according to the initial document provided by the user and the large language model. In this way, after the writer completes a rough initial document, there is no need to spend a lot of time perfecting the initial document by himself, saving the user's time and improving the efficiency of the writer in writing the document. In addition, when there is an image in the initial document, the technical solution provided by this application can also retain the image description content and generate a target document instead of discarding the image content, which can ensure the integrity of the target document.
[0006] In a possible implementation, the method further includes: Obtain a target audio, convert the target audio into an audio description text, where the audio description text is the text used to describe the content of the target audio. Input the audio description text into the large language model to preprocess the audio description text to obtain a second preprocessed text. Furthermore, generate a target document according to the first preprocessed text and the second preprocessed text.
[0007] In the technical solution provided by this application, after the writer finishes writing the initial document, relevant personnel can review the initial document. In the review meeting, the participants make speeches and put forward opinions on modification and supplementation for the initial document. The speeches of the participants can be recorded to obtain the target audio. Then, the target audio is converted into an audio description text, and the audio description text is preprocessed to obtain the second preprocessed text. The preprocessing can include correcting grammar and logic errors, supplementing and expanding the content, etc. Then, the target document can be jointly generated by combining the first preprocessed text and the second preprocessed text obtained above. In this way, it can avoid the problem that when the writer manually writes the target document, the review opinions in the review meeting are forgotten, resulting in defects in the target document.
[0008] In a possible implementation, obtaining the target audio includes: Receiving the target audio sent by the terminal; or, receiving the target audio sent by the conference system.
[0009] In the technical solution provided by this application, the target audio can be the speeches of each participant in the review meeting for reviewing the initial document. Specifically, the review meeting can be an online meeting held in the conference system. The conference system records the review meeting and uses the real-time recorded audio stream as the target audio, or at the end of the review meeting, the writer downloads the complete recording of the review meeting through the terminal.
[0010] In a possible implementation, generating the target document according to the first preprocessed text includes: Identifying the target technical field to which the content of the first preprocessed text belongs, and obtaining the prompt words corresponding to the target technical field. Inputting the first preprocessed text and the prompt words into the large language model so that the large language model processes the first preprocessed text according to the prompt words to obtain the optimized text. Generating the target document according to the optimized text in a specified format.
[0011] In the technical solution provided by this application, in order to improve the accuracy and integrity of the target document generation, the corresponding prompt words can be matched according to the technical field to which the content of the preprocessed document belongs, and when generating the target document, the matched prompt words and the preprocessed text are jointly input into the large language model, so as to instruct the large language model to generate a target document that meets the requirements of the prompt words for the preprocessed text.
[0012] In a possible implementation, generating the target document according to the first preprocessed text and the second preprocessed text includes: Identify the content of the first preprocessed text and the target technical field to which the second preprocessed text belongs, and obtain the prompt words corresponding to the target technical field. Input the first preprocessed text, the second preprocessed text, and the prompt words into the large language model, so that the large language model processes the first preprocessed text and the second preprocessed text according to the prompt words to obtain the optimized text. Generate the target document according to the optimized text in the specified format.
[0013] In the technical solution provided by this application, in order to improve the accuracy and integrity of the target document generation, the corresponding prompt words can be matched according to the technical field to which the content of the preprocessed document belongs, and when generating the target document, the matched prompt words and the preprocessed text are jointly input into the large language model, so that the large language model is instructed by the prompt words to generate the target document that meets the requirements of the prompt words for the preprocessed text.
[0014] In a second aspect, a document generation device is provided, and the device includes: A data input module, configured to obtain an initial document; An image-to-text conversion module, configured to, when the initial document includes an image, input the image into a multi-modal model to obtain an image description text, and combine the text in the initial document and the image description text to obtain a first text, where the image description text is the text used to describe the image; A generation module, configured to input the first text into the large language model to preprocess the first text to obtain a first preprocessed text; generate a target document according to the first preprocessed text.
[0015] In a possible implementation, the data input module is further configured to: Obtain a target audio; The device further includes a speech-to-text conversion module, configured to: Convert the target audio into an audio description text, where the audio description text is the text used to describe the content of the target audio; The generation module is configured to input the audio description text into the large language model to preprocess the audio description text to obtain a second preprocessed text; generate a target document according to the first preprocessed text and the second preprocessed text.
[0016] In a possible implementation, the data input module is configured to: Receive the target audio sent by the terminal; or, Receive the target audio sent by the conference system.
[0017] In a possible implementation, the generation module is configured to: Identify the target technical field to which the content of the first pre-processed text belongs; Obtain the prompt words corresponding to the target technical field; Input the first pre-processed text and the prompt words into the large language model, so that the large language model processes the first pre-processed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
[0018] In a possible implementation, the generating module is configured to: Identify the target technical fields to which the content of the first pre-processed text and the second pre-processed text belong; Obtain the prompt words corresponding to the target technical field; Input the first pre-processed text, the second pre-processed text and the prompt words into the large language model, so that the large language model processes the first pre-processed text and the second pre-processed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
[0019] In a third aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the document generation method as described in the first aspect and its possible implementations above.
[0020] In a fourth aspect, a computer-readable storage medium is provided, including computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the document generation method as described in the first aspect and its possible implementations above.
[0021] In a fifth aspect, a computer program product including instructions is provided, and when the instructions are run by a computing device cluster, the computing device cluster executes the document generation method as described in the first aspect and its possible implementations above. Description of the Drawings
[0022] Figure 1 is a schematic flowchart of a document generation method provided by an embodiment of the present application; Figure 2 is a schematic system framework diagram of a document generation provided by an embodiment of the present application; Figure 3 is a schematic flowchart of a document generation method provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a system framework for document generation provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of a device for document generation provided by an embodiment of the present application; Figure 6 It is a schematic diagram of a computing device provided by an embodiment of the present application; Figure 7 It is a schematic diagram of a computing device cluster provided by an embodiment of the present application; Figure 8 It is a schematic diagram of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0024] To facilitate the understanding of the embodiments of the present application, the following explains the nouns involved in the embodiments of the present application: Multimodal Large Model (MLM): A large machine learning model capable of processing multiple data types (such as text, images, audio, etc.). For example, an image can be input into the multimodal model, and the multimodal model outputs text describing the image, and the text describing the image is the text used to describe the content of the image.
[0025] Large Language Model (LLM): A deep learning model with a large number of parameters, specifically used for generating and understanding natural language. For example, it can understand the input text content and process it according to requirements to generate text that meets the requirements.
[0026] Automatic Speech Recognition (ASR): A technology that converts audio signals into text, which can be implemented through an ASR model, and the ASR model is a machine learning model.
[0027] Document: In the embodiments of the present application, the document includes, but is not limited to, various technical documents, solution designs, white papers, etc. The document format can be a slide presentation document (PowerPoint, PPT), a word document, etc. The application scenarios include, but are not limited to, writing design solutions, writing user manuals, writing reporting materials, and so on.
[0028] Prompt: Input text for guiding a large language model to generate specific content.
[0029] In enterprises or schools, document writing is a complex and time-consuming process that usually goes through multiple stages. First, the writer has a preliminary idea and forms a first draft. Then, it enters the review stage, where the first draft content is jointly reviewed and expanded from different perspectives. After the review is completed, the writer needs to supplement and modify the first draft content according to the review comments to form the final document. For some technical documents, they usually include a large amount of content such as technical background, existing technologies, technical problems to be solved by the solution, main ideas, solution architecture diagrams, specific scenarios, and detailed implementation processes. However, due to their busy work, enterprise personnel often do not have enough time to supplement and modify the first draft content, resulting in an overly long document writing cycle. In addition, since the time when the writer actually supplements and modifies the first draft is often a long time after the review, the opinions put forward at the review meeting may be overlooked, resulting in a poor quality of the finally completed document. Therefore, a more efficient method for assisting the writer in generating documents is needed.
[0030] An embodiment of this application provides a method for document generation. This method can be implemented by a computing device. In this method, the computing device obtains an initial document. When the initial document includes an image, the image is input into a multimodal model to obtain an image description text, and the text in the initial document and the image description text are combined to obtain a first text. Furthermore, the first text is input into a large language model to preprocess the first text through the large language model. The preprocessing can include the revision of incorrect content and can also include the expansion of the solution content in the first text. Finally, a target document is generated according to the preprocessed text. The whole process is automated and does not require manual participation, saving the time of the writer and improving the document writing efficiency.
[0031] The following describes the method for document generation provided by the embodiment of this application with reference to the accompanying drawings. See Figure 1 , this method may include the following steps: Step 101, obtain an initial document.
[0032] In implementation, when a technical document needs to be written, after having an idea, the writer can first simply write an initial document. The initial document can be a word document, a PPT document, etc. The initial document can include simple descriptions of parts such as technical background, existing technologies, technical problems to be solved by the solution, main ideas, solution architecture diagrams, specific scenarios, and detailed implementation processes.
[0033] Then, the writer can upload the initial document to the computing device.
[0034] Step 102: When the initial document includes an image, input the image into the multimodal model to obtain an image description text, and combine the text in the initial document and the image description text to obtain a first text.
[0035] The image description text is the text used to describe the content of the image.
[0036] In practice, when the writer composes the initial document, for the purpose of saving time or vivid description, the writer may insert an image into the initial document. The image can be a picture or a video. In this case, the computing device can read the initial document. When the initial document contains an image, extract the image in the initial document, and input the image into the multimodal model to obtain an image description text. Then, extract the text in the initial document, and replace the position of the image with the corresponding image description text to obtain a first text.
[0037] Alternatively, when the initial document contains an image, input the initial document into the multimodal model. The multimodal model automatically recognizes the image in the initial document and generates an image description text. Then, output the first text. The first text contains the text in the initial document and the image description text.
[0038] Step 103: Input the first text into a large language model to preprocess the first text and obtain a first preprocessed text.
[0039] In practice, the computing device inputs the first text into the large language model, and the large language model preprocesses the first text and outputs a first preprocessed text. Among them, the preprocessing can include correcting error content, expanding the solution content, etc. The error content can be various errors or defects such as technical description errors, grammar errors, and logical errors existing in the first text.
[0040] In addition, when inputting the first text into the large language model, prompt words can also be input at the same time. For example, the prompt words can be "Please correct the error content in the input text and reasonably expand the solution described in the text".
[0041] Step 104: Generate a target document according to the first preprocessed text.
[0042] In practice, in order to improve the quality of the finally generated target document, some prompt words can be designed and stored in advance for different technical fields. For example, the technical fields can include cloud computing technology, database technology, optical communication technology, wireless communication technology, new energy technology, automotive technology, artificial intelligence technology, etc. The embodiments of the present application do not limit the specific division of technical fields. As shown in Table 1 below: Table 1
[0043] In this step 104, the target technical field to which the content of the first pre-processed text belongs can be identified, and the prompt words corresponding to the target technical field can be obtained from the stored correspondence between technical fields and prompt words. Here, to identify the target technical field, the keywords in the first pre-processed text can be matched with the keywords of the technical field to identify its target technical field.
[0044] Then, the first pre-processed text and the prompt words corresponding to the target technical field are input into the large language model, so that the large language model processes the first pre-processed text according to the prompt words to obtain the optimized text.
[0045] Furthermore, the optimized text is used to generate the target document according to the specified format. For example, a document template can be pre-stored, and the document template includes multiple headings, such as technical background, existing technologies, technical problems to be solved by the solution, main ideas, solution architecture diagrams, specific scenarios, and detailed implementation processes, etc. Correspondingly, the optimized text can also include the above headings, so the content in the optimized text can be filled into the document template according to the corresponding headings to obtain the target document. To further improve the accuracy and completeness of the target document, after obtaining the target document, the target document can be input into the large language model to enable the large language model to optimize the target document to ensure the rigor of logic and the correctness and completeness of the content. When inputting the target document into the large language model, prompt words can be input simultaneously, such as "Please check the logic and content in the input document and correct the incorrect content and unrigorous logic".
[0046] In addition, for parts such as the above solution architecture diagram, existing technologies, and detailed implementation processes, the optimized text output by the large language model may have drawing codes. In this case, if drawing codes are detected in the optimized text, the drawing application corresponding to the drawing codes is called, and based on the drawing codes, corresponding architecture diagrams, flowcharts, etc. are generated, and the drawing codes in the above optimized text are replaced with the architecture diagrams, flowcharts, etc. generated based on the drawing codes. Among them, the drawing codes can be mermaid codes, PlantUML codes, etc., and the embodiments of the present application do not limit this.
[0047] In a possible implementation, the corresponding relationship between the technical field and technical knowledge can also be pre-stored. Correspondingly, in step 104 above, after identifying the target technical field to which the content of the first pre-processed text belongs, the corresponding technical knowledge of the target technical field can be obtained from the corresponding relationship between the technical field and technical knowledge. Furthermore, the first pre-processed text, the prompt words corresponding to the target technical field, and the technical knowledge corresponding to the target technical field can be input into the large language model, so that the large language model processes the first pre-processed text according to the prompt words and technical knowledge to obtain the optimized text. The technical knowledge mentioned here can be existing technical knowledge, and the existing technical knowledge is the technical knowledge that has been made public in the forms of papers, patents, conferences, etc.
[0048] The following explains how to identify the target technical field to which the content of the first pre-processed text belongs: The computing device can be deployed with a pre-trained technical field identification model. After obtaining the first pre-processed text, the computing device inputs the first pre-processed text into the technical field identification model, and the technical field identification model outputs the target technical field to which the content of the first pre-processed text belongs.
[0049] Alternatively, after obtaining the first pre-processed text, the computing device inputs the first pre-processed text into the large language model, and the large language model outputs the target technical field to which the content of the first pre-processed text belongs. In addition, when inputting the first pre-processed text into the large language model, prompt words also need to be input at the same time. For example, the prompt words are "Please analyze the technical field to which the input text belongs", or possible technical fields can be added to the prompt words for the large language model to select, such as "Please analyze which of the following technical fields the input text belongs to: cloud computing technology, database technology, optical communication technology, wireless communication technology, new energy technology, automotive technology, artificial intelligence technology".
[0050] In another possible implementation, after the writer finishes writing the initial document, relevant personnel can review the initial document. In the review meeting, the participants make speeches and put forward opinions on modification and supplementation for the initial document. The speeches of the participants can be recorded to obtain the target audio. For example, the participants can elaborate on the content that needs to be added under a certain title, or the participants only roughly state that a brief introduction of a specified technology needs to be added under a certain title, and the specified technology is existing technology, or the participants propose that the specified delivery content of a specified team needs to be added under a certain title.
[0051] There are various methods for obtaining the target audio. The following exemplarily lists several for explanation: Method 1: After the review meeting, the writer can upload the complete audio of the review meeting, or the audio processed from the complete audio, as the target audio, to the computing device. Among them, the processing of the complete audio can include deleting irrelevant content, optimizing the audio quality, etc. The irrelevant content can be content unrelated to technical communication, such as greetings between participants.
[0052] Method 2: As Figure 2 shown, a document system and a meeting system can be deployed in the computing device. The embodiments of the present application can be implemented by the document system. The document system and the meeting system communicate with each other. The above review meeting can be carried out online through the meeting system. During the review meeting, the meeting system records the speeches of the participants and sends the real-time audio stream as the target audio to the document system. Or, after the review meeting, the meeting system sends the complete audio of the review meeting, or the audio processed from the complete audio, as the target audio, to the document system.
[0053] After obtaining the target audio, the computing device inputs the target audio into the ASR model. The ASR model converts the target audio into an audio description text, where the audio description text is the text used to describe the content of the target audio. Then, the audio description text is input into the large language model to preprocess the audio description text and obtain the second preprocessed text. Among them, the preprocessing can include correcting error content, expanding the solution content, etc.
[0054] In addition, when inputting the audio description text into the large language model, a prompt word can also be input at the same time. For example, the prompt word can be "Please correct the error content in the input text and reasonably expand the solution described in the text".
[0055] Combined with the above several possible implementations, the processing of generating the optimized text in step 104 can be as follows: The computing device identifies the target technical field to which the content of the first preprocessed text and the second preprocessed text belongs, obtains the prompt word corresponding to the target technical field in the stored correspondence between the technical field and the prompt word, and obtains the technical knowledge corresponding to the target technical field in the stored correspondence between the technical field and the technical knowledge. Furthermore, the first preprocessed text, the second preprocessed text, the prompt word corresponding to the target technical field, and the technical knowledge corresponding to the target technical field are input into the large language model, so that the large language model processes the first preprocessed text and the second preprocessed text according to the prompt word and the technical knowledge to obtain the optimized text.
[0056] Of course, in the processing here, it is also possible to input only the first preprocessed text and the prompt words, or the first preprocessed text and technical knowledge, or the first preprocessed text, the second preprocessed text and the prompt words, or the first preprocessed text, the second preprocessed text and technical knowledge. The above contents can be combined for input, and the embodiments of the present application do not limit this.
[0057] The following describes the target technical field to which the contents of the first preprocessed text and the second preprocessed text belong: The computing device can be deployed with a pre-trained technical field recognition model. After obtaining the first preprocessed text and the second preprocessed text, the computing device inputs the first preprocessed text and the second preprocessed text into the technical field recognition model, and the technical field recognition model outputs the target technical field to which the contents of the first preprocessed text and the second preprocessed text belong.
[0058] Alternatively, after obtaining the first preprocessed text and the second preprocessed text, the computing device inputs the first preprocessed text and the second preprocessed text into a large language model, and the large language model outputs the target technical field to which the contents of the first preprocessed text and the second preprocessed text belong. In addition, when inputting the first preprocessed text and the second preprocessed text into the large language model, prompt words also need to be input at the same time. For example, the prompt words are "Please analyze which is the technical field to which the two input texts belong together", or possible technical fields can be added to the prompt words for the large language model to select, such as "Please analyze which of the cloud computing technology, database technology, optical communication technology, wireless communication technology, new energy technology, automotive technology, and artificial intelligence technology is the technical field to which the two input texts belong together".
[0059] The following takes the cloud computing technology as an example of the target technical field and exemplarily gives a prompt word corresponding to the cloud computing technology: "Hello, I have a task that needs your help. I have three texts in the cloud computing technology field here, namely the first draft I wrote, the text corresponding to the review recording, and some technical knowledge related to cloud computing. My goal is to sort out the technical background of the solution, the existing technology, the technical problems to be solved by the solution, the main idea, the solution architecture diagram, the specific scenario, and the detailed implementation process from these texts. If the system architecture is involved, describe the system architecture in detail in several paragraphs and generate an architecture diagram represented by mermaid code. If the method flow chart is involved, describe the method flow in detail in several paragraphs and generate a method flow chart represented by mermaid code. In addition, the content of the text I provided can be expanded through Internet search. It should be noted that the output content should not involve various sensitive words".
[0060] The following are several exemplary processes performed by the large language model based on the second preprocessed text: Process 1: If the second preprocessed text includes the statement "Add the first content under the main idea of the solution" said by the participants, then when generating the optimized content, the large language model can add the above first content under the title "Main idea of the solution".
[0061] Process 2: If the second preprocessed text includes the statement "Add a brief introduction to the specified technology under the existing technology" said by the participants, then when generating the optimized content, the large language model obtains the relevant introduction of the specified technology through online search and can add the obtained relevant introduction of the specified technology under the title "Existing technology".
[0062] Process 3: If the second preprocessed text includes the statement "Add the specified delivery content of the specified team under the detailed implementation process" said by the participants, then when generating the optimized content, the large language model accesses the internal knowledge base to obtain the specified delivery content of the specified team and can add the obtained specified delivery content under the title "Detailed implementation process".
[0063] In a possible implementation, in combination with the above target audio, the above initial document may only include titles such as technical background, existing technology, technical problems to be solved by the solution, main idea, solution architecture diagram, specific scenario, and detailed implementation process.
[0064] The following describes the technical solutions provided in the embodiments of the present application with reference to the accompanying drawings: As Figure 3As shown in the figure, in the technical solution provided by the embodiment of the present application, for the initial document, through a multi-modal model, the images therein are converted into image description texts, and combined with the original texts of the initial document to form a first text, which is input into a large language model for correcting error content and expanding solution content, so as to obtain a first pre-processed text. For the target audio of the review meeting of the initial document, through an ASR model, the target audio is converted into an audio description text, and the audio description text is input into the large language model for correcting error content and expanding solution content, so as to obtain a second pre-processed text. Furthermore, obtain the prompt words and technical knowledge corresponding to the target technical field to which the first pre-processed text and the second pre-processed text belong, and input the first pre-processed text, the second pre-processed text, the prompt words corresponding to the target technical field, and the technical knowledge corresponding to the target technical field into the large language model, so that the large language model processes the first pre-processed text according to the prompt words and technical knowledge to obtain an optimized text. Then, for the drawing code in the optimized text, the corresponding drawing application program can be called to generate corresponding flowcharts, framework diagrams, etc., add the generated flowcharts and framework diagrams to the optimized text, and generate a target document based on the optimized text in a specified format. Finally, the target document can be input into the large language model again, so that the large language model checks and corrects the logic and content of the target document.
[0065] The embodiment of the present application also provides a method for document generation. In this method, there may be no initial document, or the initial document is blank. In this case, a target document can be generated based on the target audio and a large language model. The target audio can be the real-time audio stream of the review meeting sent by the cloud conference system. See Figure 4 , the method may include the following steps: Step 201, obtain the target audio.
[0066] In implementation, relevant personnel can conduct a review meeting through the cloud conference system. In the meeting, the initial document can be shared. The initial document can be blank or only have a title. The cloud conference system records the speeches of the participants and generates an audio stream and sends it to the computing device.
[0067] In addition, the cloud conference system can also identify which page the currently shared initial document is, and send the page indication information and the audio stream generated at the same time to the computing device, where the page indication information is used to indicate which page the currently shared initial document is.
[0068] Step 202, convert the target audio into an audio description text.
[0069] Among them, the audio description text is the text used to describe the content of the target audio.
[0070] In implementation, after the computing device receives the audio stream, it can input the audio stream into the ASR model, and the ASR model converts the audio stream into an audio description text.
[0071] Step 203: Input the audio description text into the large language model, preprocess the audio description text, and obtain the preprocessed text.
[0072] In implementation, the computing device inputs the page indication information and the audio description text corresponding to the audio stream received simultaneously with the page indication information into the large prediction model, so as to preprocess the audio description text through the large language model and obtain the preprocessed text.
[0073] The following are several exemplary ways to generate the preprocessed text: Method 1: The attendee detailed the content that needs to be added to the page indicated by the page indication information. This part of the audio description text can be directly used as the preprocessed text within the page indicated by the page indication information.
[0074] Method 2: The attendee only roughly stated that a brief introduction to the specified technology can be added to the page indicated by the page indication information, and the specified technology is an existing technology. Then the large language model can obtain a simple introduction to the specified technology through online search and use it as the processed text within the page indicated by the page indication information.
[0075] Method 3: The attendee proposed that the specified delivery content of the specified team needs to be added to the page indicated by the page indication information. Then relevant information about the specified delivery content may be searched in the enterprise internal knowledge base based on the team name and department name of the specified team and used as the processed text within the page indicated by the page indication information.
[0076] Method 4: The large language model performs AI-assisted content expansion based on the audio description text and uses the expanded content as the preprocessed text.
[0077] Step 204: Generate the target document according to the preprocessed text.
[0078] In implementation, in order to improve the quality of the finally generated target document, some prompt words can be designed and stored in advance for different technical fields. For example, the technical fields may include cloud computing technology, database technology, optical communication technology, wireless communication technology, new energy technology, automotive technology, artificial intelligence technology, etc. The embodiments of the present application do not limit the specific division of technical fields.
[0079] In step 204, the target technical field to which the content of the first preprocessed text belongs can be identified, and the prompt words corresponding to the target technical field are obtained from the stored correspondence between technical fields and prompt words. Here, to identify the target technical field, the keywords in the first preprocessed text can be matched with the keywords of the technical field to identify its target technical field.
[0080] Then, the first preprocessed text and the prompt words corresponding to the target technical field are input into the large language model, so that the large language model processes the first preprocessed text according to the prompt words to obtain the optimized text.
[0081] Furthermore, the optimized text is used to generate the target document according to the specified format. For example, a document template can be pre-stored. The document template includes multiple headings, such as technical background, existing technologies, technical problems to be solved by the solution, main ideas, solution architecture diagrams, specific scenarios, and detailed implementation processes. Correspondingly, the optimized text can also include the above headings. Then, the content in the optimized text can be filled into the document template according to the corresponding headings to obtain the target document.
[0082] In addition, for parts such as the above solution architecture diagram, existing technologies, and detailed implementation processes, there may be drawing codes in the optimized document output by the large language model. In this case, the drawing programs corresponding to these drawing codes can be called to generate corresponding architecture diagrams, flowcharts, etc., and the above drawing codes are replaced with the generated architecture diagrams, flowcharts, etc.
[0083] Based on the same technical concept, an embodiment of the present application further provides a device for document generation. This device can be applied to a computing device. Refer to Figure 5 , and this device may include: A data input module 510, configured to obtain an initial document; An image-to-text conversion module 520, configured to, when the initial document includes an image, input the image into a multi-modal model to obtain an image description text, and combine the text in the initial document and the image description text to obtain a first text, where the image description text is the text used to describe the image; A generation module 530, configured to input the first text into a large language model to preprocess the first text to obtain a first preprocessed text; and generate a target document according to the first preprocessed text.
[0084] In a possible implementation, the data input module 510 is further configured to: Obtain a target audio; The device further includes a speech-to-text conversion module, configured to: Convert the target audio into an audio description text, where the audio description text is the text for describing the content of the target audio; The generating module 530 is configured to input the audio description text into the large language model to preprocess the audio description text to obtain a second preprocessed text; and generate a target document according to the first preprocessed text and the second preprocessed text.
[0085] In a possible implementation, the data input module 510 is configured to: Receive the target audio sent by the terminal; or, Receive the target audio sent by the conference system.
[0086] In a possible implementation, the generating module 530 is configured to: Identify the target technical field to which the content of the first preprocessed text belongs; Obtain the prompt words corresponding to the target technical field; Input the first preprocessed text and the prompt words into the large language model, so that the large language model processes the first preprocessed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
[0087] In a possible implementation, the generating module 530 is configured to: Identify the target technical field to which the content of the first preprocessed text and the second preprocessed text belong; Obtain the prompt words corresponding to the target technical field; Input the first preprocessed text, the second preprocessed text and the prompt words into the large language model, so that the large language model processes the first preprocessed text and the second preprocessed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
[0088] In the technical solution provided by the present application, a final target document can be automatically generated according to the initial document provided by the user and the large language model. In this way, after the writer completes a rough initial document, there is no need to spend a lot of time to perfect the initial document by himself, saving the user's time and improving the efficiency of the writer in writing the document. In addition, in the case where there are images in the initial document, the technical solution provided by the present application can also retain the image description content and generate a target document instead of discarding the image content, which can ensure the integrity of the target document.
[0089] It should be noted that: When the document generation device provided in the above embodiments performs document generation, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computing device is divided into different functional modules to complete all or part of the functions described above. In addition, the document generation device provided in the above embodiments and the method embodiments of document generation belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0090] Among them, the data input module 510, the image text conversion module 520, and the generation module 530 can all be implemented by software or by hardware. Exemplarily, next, taking the generation module 530 as an example, the implementation manner of the feature content generation module 340 will be introduced. The implementation manners of similar data input modules 510 and image text conversion modules 520 can refer to the implementation manner of the generation module 530.
[0091] As an example of a software functional unit, the generation module 530 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instances may be one or more. For example, the generation module 530 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs. Each AZ includes one data center or multiple geographically close data centers. Among them, usually one region may include multiple AZs.
[0092] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Among them, usually one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0093] As an example of a hardware functional unit, the generating module 530 may include at least one computing device, such as a server or the like. Alternatively, the generating module 530 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0094] The multiple computing devices included in the generating module 530 may be distributed in the same region or in different regions. The multiple computing devices of the generating module 530 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the generating module 530 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0095] An embodiment of the present application also provides a computing device 700, which may be used as the above server or terminal. As Figure 6 shown, the computing device 700 includes: a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other through the bus 702. The computing device 700 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.
[0096] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 702 may include a path for transmitting information between various components of the computing device 700 (for example, the memory 706, the processor 704, and the communication interface 708).
[0097] The processor 704 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0098] The memory 706 may include volatile memory, such as random access memory (RAM). The memory 706 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0099] The memory 706 stores executable program code, and the processor 704 executes the executable program code to respectively implement the functions of the foregoing data input module 510, image text conversion module 520, and generation module 530, thereby implementing the method for allocating storage resources. That is, the memory 706 stores instructions for executing the method for document generation.
[0100] Alternatively, the memory 706 stores executable code, and the processor 704 executes the executable code to respectively implement the functions of the foregoing document generation apparatus, thereby implementing the method for document generation. That is, the memory 706 stores instructions for executing the method for document generation.
[0101] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.
[0102] The embodiment of the present application further provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0103] As Figure 7 shown, the computing device cluster includes at least one computing device 700. The memory 706 in one or more of the computing devices 700 in the computing device cluster may store the same instructions for executing the method for document generation.
[0104] In some possible implementations, parts of the instructions for executing the method for document generation may also be stored respectively in the memories 706 of one or more computing devices 700 in the computing device cluster. In other words, the combination of one or more computing devices 700 can jointly execute the instructions for the method for document generation.
[0105] It should be noted that the memories 706 in different computing devices 700 in the computing device cluster can store different instructions, respectively for storing partial functions of resource allocation. That is, the instructions stored in the memories 706 of different computing devices 700 can implement the functions of one or more of the data input module 510, the image-text conversion module 520, and the generation module 530.
[0106] In some possible implementations, one or more computing devices in the computing device cluster can be connected via a network. Among them, the network can be a wide area network or a local area network, etc. Figure 8 A possible implementation is shown. As Figure 8 shown, two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the data input module 510 and the image-text conversion module 520. At the same time, the memory 706 in the computing device 700B stores instructions for executing the function of the generation module 530.
[0107] Figure 8 The connection method between the computing device clusters shown may be considered because document generation requires a large amount of computing resources. Therefore, it is considered to hand over the function of the generation module 530 to the computing device 700B for execution.
[0108] It should be understood that Figure 8 the functions of the computing device 700A shown in can also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be completed by multiple computing devices 700.
[0109] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another.
[0110] It should also be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the various processes do not imply the sequence of execution, and the execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0111] In the present application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality" refers to two or more. For example, a plurality of second devices refers to two or more second devices. In this article, the terms "system" and "network" are often used interchangeably.
[0112] It should be understood that the terms used in the description of the various examples herein are only for the purpose of describing specific examples and are not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0113] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the preceding and following associated objects.
[0114] It should also be understood that the terms "if" and "when" can be interpreted to mean "when" or "upon determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined..." or "if [stated condition or event] is detected" can be interpreted to mean "when determining..." or "in response to determining..." or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]".
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
[0116] All kinds of information, data, and signals involved in this application are authorized by users or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0117] The embodiments of this application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the method for allocating storage resources.
[0118] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method for allocating storage resources.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
[0120] All kinds of information, data, and signals involved in this application are authorized by users or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
Claims
1. A method for document generation, characterized in that, The method includes: Obtain an initial document; When the initial document includes an image, input the image into a multi-modal model to obtain an image description text, and combine the text in the initial document and the image description text to obtain a first text, where the image description text is the text used to describe the image; Input the first text into a large language model to preprocess the first text and obtain a first preprocessed text; Generate a target document according to the first preprocessed text.
2. The method according to claim 1, wherein The method further includes: Obtain a target audio; Convert the target audio into an audio description text, where the audio description text is the text used to describe the content of the target audio; Input the audio description text into the large language model to preprocess the audio description text and obtain a second preprocessed text; The generating a target document according to the first preprocessed text includes: Generate a target document according to the first preprocessed text and the second preprocessed text.
3. The method according to claim 2, wherein The obtaining a target audio includes: Receive the target audio sent by a terminal; or, Receive the target audio sent by a conference system.
4. The method according to claim 1, wherein The generating a target document according to the first preprocessed text includes: Identify the target technical field to which the content of the first preprocessed text belongs; Obtain the prompt words corresponding to the target technical field; Input the first preprocessed text and the prompt words into the large language model so that the large language model processes the first preprocessed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
5. The method according to claim 2 or 3, characterized in that, The generating a target document according to the first preprocessed text and the second preprocessed text includes: Identify the target technical field to which the content of the first preprocessed text and the second preprocessed text belong; Obtain the prompt words corresponding to the target technical field; Input the first preprocessed text, the second preprocessed text and the prompt words into the large language model so that the large language model processes the first preprocessed text and the second preprocessed text according to the prompt words to obtain an optimized text; Generate a target document according to the optimized text in a specified format.
6. A device for document generation, characterized in that, The device includes: A data input module for obtaining an initial document; An image-text conversion module for, when the initial document includes an image, inputting the image into a multi-modal model to obtain an image description text, and combining the text in the initial document and the image description text to obtain a first text, where the image description text is the text used to describe the image; A generation module for inputting the first text into a large language model to preprocess the first text and obtain a first preprocessed text; generating a target document according to the first preprocessed text.
7. The device according to claim 6, characterized in that The data input module is further used for: Obtaining a target audio; The device further includes a voice-text conversion module for: Convert the target audio into an audio description text, where the audio description text is the text used to describe the content of the target audio; The generating module is configured to input the audio description text into the large language model to preprocess the audio description text to obtain a second preprocessed text; and generate a target document according to the first preprocessed text and the second preprocessed text.
8. A cluster of computing devices, characterized in that, Comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the document generation method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, Comprising computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the document generation method according to any one of claims 1 to 5.
10. A computer program product comprising instructions, characterized in that, When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the document generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Requirement document generating method and related equipment
CN110265024A
Processing method and device based on multi-modal input document, equipment and storage medium
CN118940742A
Technical document generation method and device, equipment, medium and product
CN119003733A
Emergency rescue assisting method based on large language model and computer equipment
CN119597921A
Content generation method, computer device, and storage medium
US20250095252A1