Method, device, equipment and medium for generating document based on generative large model
By acquiring relevant paragraph information of the target document through a generative large model, the target document is generated, which solves the problems of low efficiency and insufficient accuracy in the existing technology, and realizes efficient and accurate document generation and enhances the richness of document content.
Patent Information
- Application Number
- CN202410114873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2044-01-26
AI Technical Summary
Existing technologies require manual marking of the merging order when generating target documents, which is inefficient and inaccurate, and cannot effectively enhance the richness of document content.
A generative large model is adopted, which obtains information from multiple relevant paragraphs related to the target topic based on the target topic and a pre-generated document database. The target document is then generated through the generative large model, avoiding human intervention and improving accuracy.
It improves the efficiency and accuracy of generating target documents, enhances the richness and accuracy of document content, and enables the annotation of source information in documents to improve readability.
Smart Images

Figure CN117992569B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of text processing and artificial intelligence, and in particular to a method, apparatus, device and medium for generating documents based on a generative large model. Background Technology
[0002] In many scenarios using existing technologies, it is necessary to generate a target document based on multiple existing documents.
[0003] In traditional techniques, the process involves first identifying multiple documents with similar titles from a document library based on the target title of the desired document. Then, staff manually mark the merging order of these documents, and finally, based on this order, merge the contents of the multiple documents to generate the target document. Summary of the Invention
[0004] This disclosure provides a method, apparatus, device, and medium for generating documents based on a generative large model.
[0005] According to one aspect of this disclosure, a method for generating documents based on a generative large model is provided, comprising:
[0006] Obtain the target topic of the document to be generated;
[0007] Based on the target topic and a pre-generated document database, information on multiple related paragraphs related to the target topic is obtained; the document database includes information on paragraphs included in each of the multiple original documents;
[0008] Based on the information from the multiple related paragraphs, a pre-trained generative large model is used to generate the target document corresponding to the target topic.
[0009] According to another aspect of this disclosure, an apparatus for generating documents based on a generative large model is provided, comprising:
[0010] The topic acquisition module is used to obtain the target topic of the document to be generated;
[0011] The paragraph acquisition module is used to acquire information about multiple related paragraphs related to the target topic based on the target topic and a pre-generated document database; the document database includes information about the paragraphs included in each of the multiple original documents;
[0012] The document generation module is used to generate target documents corresponding to the target topic based on the information of the multiple related paragraphs and using a pre-trained generative large model.
[0013] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.
[0017] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.
[0018] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.
[0019] According to the technology disclosed herein, the efficiency of generating target documents can be effectively improved, the content of generated target documents can be enhanced, and the accuracy of generated target documents can be improved.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0022] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0023] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0024] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0025] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0026] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0027] Figure 6 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0028] Figure 7This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0030] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0031] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0032] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0033] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown in the figure, this embodiment provides a method for generating documents based on a generative large model, which may specifically include the following steps:
[0034] S101. Obtain the target topic of the document to be generated;
[0035] In this embodiment, the target topic of the document to be generated can be manually entered to limit the topic of the desired document. For example, the target topic in this embodiment can be a sentence, or it can be a combination of multiple words, etc.
[0036] S102. Based on the target topic and a pre-generated document database, obtain information on multiple related paragraphs related to the target topic; the document database includes information on the paragraphs included in each of the multiple original documents;
[0037] S103. Based on information from multiple related paragraphs, a pre-trained generative large model is used to generate the target document corresponding to the target topic.
[0038] In this embodiment, the generative large model used can also be called a generative language model (General Language Model; GLM) or a generative large language model.
[0039] The execution subject of the method for generating documents based on a generative large model in this embodiment can be a device for generating documents based on a generative large model. This device can be an electronic entity or a software-integrated application. In use, a device for generating documents based on a generative large model can be used to generate target documents corresponding to the target topic based on a pre-trained generative large model and a pre-generated document database.
[0040] The method for generating documents based on a generative large model in this embodiment can first obtain information on multiple related paragraphs related to the target topic based on the target topic and a pre-generated document database, which can accurately locate the paragraphs related to the target topic. Then, based on the information of multiple related paragraphs, a pre-trained generative large model is used to generate the target document corresponding to the target topic, which can effectively improve the accuracy of the generated target document.
[0041] Compared to existing methods that merge multiple similar documents according to a manually marked merging order to obtain a target document, the technical solution of this embodiment does not require manual intervention and can effectively improve the efficiency of generating target documents. Moreover, it acquires information from paragraphs related to the target topic, and then, based on the information from multiple related paragraphs, uses a generative large model to generate the target document corresponding to the target topic. This not only effectively improves the accuracy of the generated target document, but also, instead of simply and crudely merging the original document or paragraphs, it effectively enhances the content of the generated target document, thus improving its accuracy.
[0042] Figure 2 This is a schematic diagram according to the second embodiment of this disclosure; the method for generating documents based on a generative large model in this embodiment, in the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 2 As shown, the method for generating documents based on a generative large model in this embodiment may specifically include the following steps:
[0043] S201. Obtain the target topic of the document to be generated;
[0044] S202. Based on the target topic and document database, obtain the vector representation of each relevant paragraph among multiple related paragraphs related to the target topic;
[0045] S203. Based on the vector representation of each relevant paragraph in multiple relevant paragraphs and the document database, obtain the position identifier of each relevant paragraph respectively;
[0046] Specifically, the location identifier of each relevant paragraph can be identified by using both the original document identifier of the relevant paragraph and its location identifier within the original document. For example, a relevant paragraph can be identified by both the original document identifier 1 and the location identifier 1-2 within the original document. Here, 1 can be used to represent the original document with identifier 1, and 1-2 can be used to represent the second paragraph within the original document with identifier 1.
[0047] S204. Based on the location identifiers of each relevant paragraph and the document database, obtain the content of each relevant paragraph among multiple relevant paragraphs;
[0048] In this embodiment, steps S202-S204 are described above. Figure 1 The following is an example of one implementation of step S102 in the illustrated embodiment. Specifically, in this implementation, the information obtained regarding multiple related paragraphs related to the target topic may include the content and position identifier of each related paragraph among the multiple related paragraphs related to the target topic.
[0049] In this embodiment, the pre-generated document database may include multiple source documents. Specifically, it may include the content of each paragraph in each source document, the vector representation of each paragraph, and the position identifier of each paragraph. Each source document may include one, two, or more paragraphs. In practical applications, a pre-trained vector representation model can be used to obtain the vector representation of each paragraph in each source document based on the content of each paragraph. Furthermore, the position identifier of each paragraph is obtained based on the identifier of each source document and the position of each paragraph within its respective source document.
[0050] Based on the above content included in the document database, it can be considered that the data in the document database in this embodiment is stored according to the granularity of paragraphs in the original document. For each paragraph in each original document, information in three fields can be stored, such as the content of the paragraph, the vector representation of the paragraph, and the position identifier of the paragraph.
[0051] Specifically, step S202 may include the following steps:
[0052] (a1) Obtain the vector representation of the target topic;
[0053] Specifically, a pre-trained vector representation model can be used. The target topic is input into the vector representation model, and the vector representation model can output a vector representation of the target topic.
[0054] (b1) Based on the vector representation of the target topic and the vector representation of the paragraphs included in each original document in the document database, retrieve the vector representation of multiple related paragraphs related to the target topic.
[0055] In this embodiment, based on the vector representation of the target topic, the vector representations of multiple related paragraphs can be retrieved from the vector representations of all paragraphs included in the document database. Specifically, the similarity between the vector representation of the target topic and the vector representations of each paragraph in the document database can be calculated, i.e., the cosine similarity between the two vectors can be calculated. Then, paragraphs with a similarity greater than or equal to a preset similarity threshold are selected as multiple related paragraphs related to the target topic, thus obtaining the vector representations of multiple related paragraphs.
[0056] Further optionally, in the specific implementation of step S203, the position identifier of each related paragraph can be obtained based on the vector representation of each related paragraph in multiple related paragraphs, the vector representation of the paragraphs included in each original document in the document database, and the position identifier. The position identifier of each related paragraph can locate the position of each related paragraph in its respective original document.
[0057] Alternatively, in the specific implementation of step S204, the content of each relevant paragraph in the corresponding original document can be obtained based on the position identifier of each relevant paragraph, the position identifier of each paragraph included in each original document in the document database, and the content.
[0058] In other words, for each relevant paragraph's vector representation, the location identifier of that relevant paragraph can be obtained from the document database based on its vector representation. Furthermore, based on the location identifier of that relevant paragraph, the content of that relevant paragraph—that is, the text content of that relevant paragraph in the original document—can be obtained from the document database.
[0059] In this embodiment, by adopting the above steps S202-S204, the content of each relevant paragraph in multiple related paragraphs related to the target topic can be obtained accurately and efficiently, providing effective material support for the subsequent generation of the target document.
[0060] Further, optionally, the following steps may be included before step S202:
[0061] (a2) Collect multiple source documents;
[0062] (b2) Convert the format of multiple original documents to make the format of multiple original documents uniform;
[0063] For example, in this embodiment, the formats of the multiple original documents may include various text formats such as Word, PPT, PDF, WPS, and TXT. To facilitate the establishment of a text database, format conversion can be performed to convert the multiple original documents into text data with a uniform format.
[0064] (c2) Obtain the content of each paragraph in the original document, the vector representation of the paragraph, and the position identifier of the paragraph;
[0065] Referring to the description in the above embodiments, the vector representation of each paragraph can input the content of the corresponding paragraph into a pre-trained vector representation model, and the vector representation model outputs the vector representation of the paragraph.
[0066] (d2) Generate a document database based on the content of paragraphs in each of the multiple original documents, the vector representation of the paragraphs, and the position identifier of the paragraphs.
[0067] The content of each paragraph in each original document, the vector representation of each paragraph, and the position identifier of each paragraph are stored in the document database according to the corresponding relationship, and finally the generated document database is obtained.
[0068] The above methods can accurately and efficiently generate a document database, providing effective support for the subsequent generation of target documents.
[0069] S205. Based on the content of each relevant paragraph in multiple related paragraphs, a generative large model is used to generate the target document corresponding to the target topic.
[0070] Specifically, in this embodiment, step S205 is as described above. Figure 1 One specific implementation of step S103 in the illustrated embodiment.
[0071] In this implementation, the content of each relevant paragraph from multiple related paragraphs related to the target topic can be directly input into the generative big model. The generative big model can generate a target document based on the input content, which is the document corresponding to the target topic.
[0072] The method for generating documents based on a generative large model in this embodiment first obtains the content of each relevant paragraph in multiple relevant paragraphs related to the target topic, and then uses a generative large model to generate the target document corresponding to the target topic based on the content of each relevant paragraph in multiple relevant paragraphs. This can effectively improve the accuracy of the content of the generated target document, and by using a generative large model to generate the target document, the content of the generated target document can be enhanced, thereby improving the generation efficiency of the target document.
[0073] Figure 3This is a schematic diagram according to the third embodiment of this disclosure; the method for generating documents based on a generative large model in this embodiment, in the above... Figure 2 Based on the technical solution of the illustrated embodiment, the specific implementation of step S205 is further described in more detail, which may include the following steps:
[0074] S301. Segment the content of each related paragraph in multiple related paragraphs to obtain multiple segmented words;
[0075] Specifically, each related paragraph in multiple related paragraphs can be segmented according to a preset word segmentation strategy, ultimately resulting in multiple word segments included in the multiple related paragraphs. Word segments can be considered the smallest unit within a related paragraph.
[0076] Alternatively, a pre-trained word segmentation model can be used to segment each relevant paragraph. In practice, each relevant paragraph is input into the word segmentation model, which then segments the input paragraph and outputs the segmentation results.
[0077] S302. Count the number of multiple word segments;
[0078] S303. Detect whether the number of multiple word segments is greater than the maximum number threshold that the generative large model can accept; if the number of multiple word segments is not greater than the maximum number threshold, proceed to step S304; if the number of multiple word segments is greater than the maximum number threshold, proceed to step S305.
[0079] S304. Based on multiple word segmentations, a generative large model is used to generate the target document corresponding to the target topic, and the process ends.
[0080] In this embodiment, the granularity of the input information to the generative large model is not the entirety of the relevant paragraphs, but rather the granularity of the word segments included in the relevant paragraphs. If the number of multiple word segments does not exceed the maximum threshold, then multiple word segments can be directly input into the generative large model, which can then generate the target document corresponding to the target topic based on the input information. Specifically, word segments of each relevant paragraph in multiple relevant paragraphs can be sequentially input into the generative large model until all word segments are input into the generative large model.
[0081] In this embodiment, the multiple word segments input to the generative large model are the tokens input to the model. When the total number of tokens included in multiple related paragraphs does not exceed the maximum data threshold, the tokens are directly input into the generative large model, as described in this embodiment. The generative large model can then generate the target document corresponding to the target topic based on all the input tokens.
[0082] S305. Rewrite the content of each relevant paragraph to obtain the rewritten content of each relevant paragraph, so that the number of words in the rewritten content of each relevant paragraph is less than the content of the relevant paragraph; proceed to step S306.
[0083] In this embodiment, a preset rewriting strategy can be used to rewrite the content of each relevant paragraph. This rewriting strategy can be configured to require that the rewritten content of the relevant paragraphs retains the same semantics as the original content, but has fewer words. For example, based on the requirements of the above rewriting strategy, operations such as removing adjectives from some relevant paragraphs or replacing longer words with shorter words of the same semantic meaning can be performed to rewrite the content of each relevant paragraph, ensuring that the rewritten content of the relevant paragraphs has fewer words than the original content.
[0084] For example, step S305 can also be rewritten using a generative large model. Specifically, it can include the following steps:
[0085] (a3) Obtain rewriting prompt information. The rewriting prompt information requires that the number of words in the rewritten content of each relevant paragraph be less than the content of the relevant paragraph.
[0086] The rewrite prompt can be entered by staff. Alternatively, it can be represented by a rewrite example. This example can include an original statement and a rewritten statement, where the original and rewritten statements have the same semantic meaning, but the rewritten statement is shorter than the original.
[0087] (b3) Based on the content of each relevant paragraph and the rewriting prompt information, a generative large model is used to obtain the rewritten content of each relevant paragraph.
[0088] For each relevant paragraph, the content of the relevant paragraph and the rewriting prompt words can be input into the generative big data model. The generative big data model can then rewrite the content of the relevant paragraph based on the rewriting prompt words.
[0089] Based on the above, it can be seen that when there are a large number of related paragraphs and a large number of words included in multiple related paragraphs, it may exceed the maximum number threshold that the generative large model can accept. At this time, it is necessary to use the generative large model to rewrite the content of each related paragraph. The purpose of rewriting in this embodiment is to reduce the number of words in the rewritten content of each related paragraph while keeping the semantics of each related paragraph unchanged.
[0090] S306. Based on the rewritten content of each relevant paragraph in multiple related paragraphs, a generative large model is used to generate the target document corresponding to the target topic, and the process ends.
[0091] Specifically, when implementing step S306, the methods described in steps S301-S306 above can continue. For example, the rewritten content of each related paragraph in multiple related paragraphs can be further segmented into multiple words. If the number of multiple words is not greater than the maximum number threshold, the target document corresponding to the target topic can be directly generated using the method in step S304. However, if the number of multiple words is still greater than the maximum number threshold, the rewritten content of each related paragraph can be rewritten again according to the method in step S305 of the above embodiment, so that the number of words in the rewritten content of the related paragraphs is less. This process is repeated until the number of multiple words is not greater than the maximum number threshold, at which point the target document corresponding to the target topic is generated using the method in step S304.
[0092] The document generation method based on a generative large model in this embodiment can use the word segmentation of each relevant paragraph as input information for the generative large model, which can effectively refine the input information of the generative large model, thereby effectively improving the accuracy and precision of the target document generated by the generative large model. Moreover, if the number of word segments of multiple relevant paragraphs exceeds the maximum number threshold that the generative large model can accept, the content of each relevant paragraph can be rewritten to reduce the number of words in the rewritten content. Then, the multiple word segments of the rewritten content of each relevant paragraph can be used as input information for the generative large model to generate the target document based on the input information. This allows the content of relevant paragraphs in various situations to generate the target document, effectively improving the efficiency of target document generation.
[0093] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure; this embodiment provides a method for generating documents based on a generative large model, in the above... Figures 1-3 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 4 As shown, the method for generating documents based on a generative large model in this embodiment may specifically include the following steps:
[0094] S401. Obtain the target topic of the document to be generated;
[0095] S402. Based on the target topic and document database, obtain the vector representation of each relevant paragraph among multiple related paragraphs related to the target topic;
[0096] With the above Figure 2 The difference between the illustrated embodiment and the actual embodiment is that in this embodiment, step S402 is the same as described above. Figure 1The following is an example of one implementation of step S102 in the illustrated embodiment. Specifically, in this implementation, the information obtained regarding multiple related paragraphs related to the target topic may include vector representations of each related paragraph among the multiple related paragraphs related to the target topic.
[0097] S403. Based on the vector representation of each relevant paragraph in multiple relevant paragraphs and the document database, obtain the position identifier of each relevant paragraph respectively;
[0098] S404. Based on the location identifiers of each relevant paragraph and the document database, obtain the content of each relevant paragraph among multiple relevant paragraphs;
[0099] The specific implementation methods for steps S403-S404 can be found above. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.
[0100] S405. Obtain the identifier of the original document to which each relevant paragraph belongs;
[0101] S406. Obtain the document prompt information of the generated target document. The document prompt information is used to limit the source information of each sentence or paragraph in the generated target document.
[0102] In this embodiment, the document prompt information can be input by the staff and is used to limit the annotation of source information for each sentence or paragraph in the generated target document. Specifically, the annotation of source information can be set to sentence level or paragraph level as needed.
[0103] When the source information is annotated at the statement level, for example, if statement 1 in the generated target document is based on related paragraphs A, B, and C, where related paragraph A belongs to original document 1, related paragraph B belongs to original document 2, and related paragraph C belongs to original document 3, then [1, 2, 3] can be added after statement 1 in the target document to link the source information of statements in the target document.
[0104] If the source information is labeled at the paragraph level, for example, if paragraph 2 in the generated target document is based on related paragraphs D and E, where related paragraph D belongs to the original document 4 and related paragraph E belongs to the original document 5, then [4, 5] can be added after paragraph 2 in the target document to link the source information of the paragraphs in the target document.
[0105] The document prompt information in this embodiment can also be sample information of a document. This sample information can include the content of multiple related paragraphs, the original document identifiers of each related paragraph, and source information for each sentence or paragraph in the generated target document. This allows the generative large model to learn the source information annotations based on the document sample information in the document prompt information and to annotate the source information of each sentence or paragraph in the generated target document.
[0106] S407. Based on the content of each related paragraph in multiple related paragraphs, the identifier of the original document to which each related paragraph belongs, and the document prompt information, a generative large model is used to generate the target document corresponding to the target topic; the target document is marked with the source information of each sentence or paragraph.
[0107] In this embodiment, steps S403-S407 are as described above. Figure 1 The following is an example of one implementation of step S103 in the illustrated embodiment.
[0108] In this implementation, the input to the generative large model can include three aspects: the content of each relevant paragraph, the identifier of the original document to which each relevant paragraph belongs, and document prompt information. The generative large model can generate a target document based on the content of each relevant paragraph. Simultaneously, based on the document prompt information and the identifier of the original document to which each relevant paragraph belongs, it can embed source information into each sentence or paragraph in the generated target document. This source information can identify which relevant paragraphs in the original documents each sentence or paragraph in the generated target document was based on.
[0109] Alternatively, in the specific implementation of step S407, the above-mentioned methods may be adopted. Figure 3 The implementation principle of steps S301-304 in the illustrated embodiment is used. For example, when generating the target document corresponding to the target topic using a generative large model based on the content of each related paragraph in multiple related paragraphs, the identifier of the original document to which each related paragraph belongs, and the document prompt word information, the content of each related paragraph in multiple related paragraphs is segmented according to the method of step S301. Based on steps S302-S305, when the number of segmented words in multiple related paragraphs is not greater than the maximum number threshold, the multiple segmented words included in the multiple related paragraphs, the identifier of the original document to which each related paragraph belongs, and the document prompt word information are input into the generative large model. At this time, the generative large model can generate the target document based on the multiple segmented words included in the multiple related paragraphs. At the same time, the generative large model also links source information to each sentence or paragraph in the generated target document based on the document prompt word information and the identifier of the original document to which each related paragraph belongs.
[0110] Furthermore, in this embodiment, when the number of words in multiple related paragraphs exceeds the maximum threshold, the content of each related paragraph is rewritten using the method described in step S305 above, resulting in rewritten content for each related paragraph. Then, the rewritten content of multiple related paragraphs is segmented. If the number of words obtained is not greater than the maximum threshold, the target document is generated according to the above method. However, if the number of words obtained after segmentation still exceeds the maximum threshold, the content of each related paragraph is rewritten again using step S305 until the number of words corresponding to the rewritten content of multiple related paragraphs is no greater than the maximum threshold. Then, the target document is generated according to the above method, and source information is linked to the target document. This not only effectively enhances the content of the generated target document but also allows for the annotation of source information, improving the readability of the target document.
[0111] By employing the above method, target documents can be generated based on multiple word segments from multiple related paragraphs. By refining the input information of the generative large-scale model, the accuracy and precision of the generated target documents can be effectively improved. Furthermore, embedding source information into the target document further enhances its content.
[0112] Moreover, in this embodiment, the target document is generated by a generative large model. Compared with the existing method of merging multiple documents to generate a target document, this method can improve the emotionality and fullness of the generated target document and enhance its content.
[0113] In practical applications, the above Figure 2 Step 205 in the illustrated embodiment, and Figure 3 Step 306 in the illustrated embodiment can also be implemented in the manner of steps S405-S407 of this embodiment, and will not be described again here.
[0114] The method for generating documents based on a generative large model in this embodiment can also input document prompt words and the identifiers of the original documents to which each relevant paragraph belongs into the generative large model, so that the generative large model can mark the source information of each sentence or paragraph in the generated target document, thereby further and effectively enhancing the content of the target document.
[0115] Figure 5 This is a schematic diagram according to the fifth embodiment of this disclosure; as shown Figure 5 As shown, this embodiment provides an apparatus 500 for generating documents based on a generative large model, including:
[0116] The topic acquisition module 501 is used to acquire the target topic of the document to be generated.
[0117] The paragraph acquisition module 502 is used to acquire information on multiple related paragraphs related to the target topic based on the target topic and a pre-generated document database; the document database includes information on the paragraphs included in each of the multiple original documents;
[0118] The document generation module 503 is used to generate a target document corresponding to the target topic based on the information of the multiple related paragraphs and using a pre-trained generative large model.
[0119] The apparatus 500 for generating documents based on a generative large model in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0120] Figure 6 This is a schematic diagram according to the sixth embodiment of this disclosure; as shown Figure 6 As shown, this embodiment provides an apparatus 600 for generating documents based on a generative large model, including the above-mentioned... Figure 5 The modules with the same names and functions shown are: Topic Acquisition Module 601, Paragraph Acquisition Module 602, and Document Generation Module 603.
[0121] In this embodiment, the paragraph acquisition module 602 is used for:
[0122] Based on the target topic and the document database, obtain the vector representation of each relevant paragraph among multiple related paragraphs related to the target topic;
[0123] Based on the vector representation of each related paragraph in the plurality of related paragraphs and the document database, the position identifier of each related paragraph is obtained respectively;
[0124] Based on the location identifiers of each related paragraph and the document database, the content of each related paragraph among the plurality of related paragraphs is obtained.
[0125] Further optionally, in one embodiment of this disclosure, the paragraph acquisition module 602 is configured to:
[0126] Obtain the vector representation of the target topic;
[0127] Based on the vector representation of the target topic and the vector representation of the paragraphs included in each original document in the document database, the vector representations of multiple related paragraphs related to the target topic are retrieved.
[0128] Further optionally, in one embodiment of this disclosure, the paragraph acquisition module 602 is configured to:
[0129] Based on the vector representation of each related paragraph in the plurality of related paragraphs, the vector representation of the paragraphs included in each original document in the document database, and the position identifier, the position identifier of each related paragraph is obtained respectively.
[0130] Further optionally, in one embodiment of this disclosure, the paragraph acquisition module 602 is configured to:
[0131] Based on the position identifiers of each relevant paragraph, the position identifiers of the paragraphs included in each original document in the document database, and the content, the content of each relevant paragraph in the corresponding original document is obtained.
[0132] Further optional, such as Figure 6 As shown, the apparatus 600 for generating documents based on a generative large model in this embodiment further includes:
[0133] Acquisition module 604 is used to acquire the multiple original documents;
[0134] The format transcoding module 605 is used to transcode the multiple original documents to make the format of the multiple original documents uniform;
[0135] The information acquisition module 606 is used to acquire the content of each paragraph, the vector representation of the paragraph, and the position identifier of the paragraph in each original document.
[0136] The database generation module 607 is used to generate the document database based on the content of paragraphs, the vector representation of paragraphs, and the position identifier of paragraphs in each of the plurality of original documents.
[0137] Further optionally, in one embodiment of this disclosure, the document generation module 603 is configured to:
[0138] Based on the content of each of the multiple related paragraphs, the generative large model is used to generate the target document corresponding to the target topic.
[0139] Further optionally, in one embodiment of this disclosure, the document generation module 603 is configured to:
[0140] The content of each of the multiple related paragraphs is segmented into words to obtain multiple word segments;
[0141] Count the number of the multiple word segments;
[0142] Detect whether the number of the multiple word segments exceeds the maximum number threshold that the generative large model can accept;
[0143] In response to the condition that the number of the multiple word segments is not greater than the maximum number threshold, the target document corresponding to the target topic is generated based on the multiple word segments and using the generative big model.
[0144] Further, optionally, in one embodiment of this disclosure, the document generation module 603 is also used for:
[0145] In response to the fact that the number of the multiple word segments is greater than the maximum number threshold, the content of each of the relevant paragraphs is rewritten to obtain the rewritten content of each of the relevant paragraphs, such that the number of words in the rewritten content of each of the relevant paragraphs is less than the content of the relevant paragraph.
[0146] Based on the rewritten content of each of the multiple related paragraphs, the generative large model is used to generate the target document corresponding to the target topic.
[0147] Further optionally, in one embodiment of this disclosure, the document generation module 603 is configured to:
[0148] Obtain rewriting prompt information, wherein the rewriting prompt information requires that the number of words in the rewritten content of each relevant paragraph is less than the content of the relevant paragraph;
[0149] Based on the content of each relevant paragraph and the rewriting prompt information, the generative large model is used to obtain the rewritten content of each relevant paragraph.
[0150] Further optionally, in one embodiment of this disclosure, the document generation module 603 is configured to:
[0151] Obtain the identifier of the original document to which each of the relevant paragraphs belongs;
[0152] Obtain document prompt information for generating the target document; the document prompt information is used to limit the source information of each sentence or paragraph in the generated target document.
[0153] Based on the content of each of the multiple related paragraphs, the identifier of the original document to which each related paragraph belongs, and the document prompt information, the generative large model is used to generate the target document corresponding to the target topic; wherein the source information of each sentence or paragraph is marked in the target document.
[0154] The apparatus 600 for generating documents based on a generative large model in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0155] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0156] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0157] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0158] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0159] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0160] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).
[0161] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0165] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0166] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0167] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0168] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating documents based on a generative large model, comprising: Obtain the target topic of the document to be generated; Based on the target topic and a pre-generated document database, the vector representations of each relevant paragraph in a plurality of related paragraphs related to the target topic are obtained; the document database includes the content of each paragraph in each of the original documents, the vector representations of each paragraph, and the position identifiers of each paragraph; Based on the vector representation of each related paragraph in the plurality of related paragraphs and the document database, the position identifier of each related paragraph is obtained respectively; Based on the location identifiers of each related paragraph and the document database, the content of each related paragraph among the plurality of related paragraphs is obtained; The content of each of the multiple related paragraphs is segmented into words to obtain multiple word segments; Count the number of the multiple word segments; Detect whether the number of the multiple word segments exceeds the maximum number threshold that the generative large model can accept; In response to the number of the multiple word segments exceeding the maximum number threshold, the content of each of the relevant paragraphs is rewritten to obtain the rewritten content of each of the relevant paragraphs, such that the number of characters in the rewritten content of each of the relevant paragraphs is less than the number of characters in the content of each of the relevant paragraphs, but the semantics remain unchanged; Based on the rewritten content of each of the multiple related paragraphs, the generative large model is used to generate the target document corresponding to the target topic.
2. The method according to claim 1, wherein, Based on the target topic and the document database, vector representations of multiple related paragraphs related to the target topic are obtained, including: Obtain the vector representation of the target topic; Based on the vector representation of the target topic and the vector representation of the paragraphs included in each original document in the document database, the vector representations of multiple related paragraphs related to the target topic are retrieved.
3. The method according to claim 1, wherein, Based on the vector representation of each relevant paragraph among the multiple relevant paragraphs and the document database, the position identifier of each relevant paragraph is obtained, including: Based on the vector representation of each related paragraph in the plurality of related paragraphs, the vector representation of the paragraphs included in each original document in the document database, and the position identifier, the position identifier of each related paragraph is obtained respectively.
4. The method according to claim 1, wherein, Based on the location identifiers of each of the related paragraphs and the document database, the content of each of the multiple related paragraphs is obtained, including: Based on the position identifiers of each relevant paragraph, the position identifiers of the paragraphs included in each original document in the document database, and the content, the content of each relevant paragraph in the corresponding original document is obtained.
5. The method according to claim 1, wherein, Before obtaining information on multiple related paragraphs related to the target topic based on the target topic and a pre-generated document database, the method further includes: Collect the aforementioned multiple source documents; The multiple original documents are converted to a uniform format. Obtain the content of each paragraph in the original document, the vector representation of the paragraph, and the position identifier of the paragraph; The document database is generated based on the content of paragraphs, the vector representation of paragraphs, and the position identifiers of paragraphs in each of the multiple original documents.
6. The method according to claim 1, wherein, The method further includes: In response to the condition that the number of the multiple word segments is not greater than the maximum number threshold, the target document corresponding to the target topic is generated based on the multiple word segments and using the generative big model.
7. The method according to claim 1, wherein, The content of each relevant paragraph is rewritten to obtain the rewritten content of each relevant paragraph, including: Obtain rewriting prompt information, wherein the rewriting prompt information requires that the number of words in the rewritten content of each relevant paragraph is less than the number of words in the content of each relevant paragraph; Based on the content of each relevant paragraph and the rewriting prompt information, the generative large model is used to obtain the rewritten content of each relevant paragraph.
8. The method according to any one of claims 1-7, wherein, Based on the information from the multiple related paragraphs, a pre-trained generative large model is used to generate target documents corresponding to the target topic, including: Obtain the identifier of the original document to which each of the relevant paragraphs belongs; Obtain document prompt information for generating the target document; the document prompt information is used to limit the source information of each sentence or paragraph in the generated target document. Based on the content of each of the multiple related paragraphs, the identifier of the original document to which each related paragraph belongs, and the document prompt information, the generative large model is used to generate the target document corresponding to the target topic; wherein the source information of each sentence or paragraph is marked in the target document.
9. An apparatus for generating documents based on a generative large model, comprising: The topic acquisition module is used to obtain the target topic of the document to be generated; The paragraph retrieval module is used for: Based on the target topic and a pre-generated document database, the vector representations of each relevant paragraph in a plurality of related paragraphs related to the target topic are obtained; the document database includes the content of each paragraph in each of the original documents, the vector representations of each paragraph, and the position identifiers of each paragraph; Based on the vector representation of each related paragraph in the plurality of related paragraphs and the document database, the position identifier of each related paragraph is obtained respectively; Based on the location identifiers of each related paragraph and the document database, the content of each related paragraph among the plurality of related paragraphs is obtained; The document generation module is used for: The content of each of the multiple related paragraphs is segmented into words to obtain multiple word segments; Count the number of the multiple word segments; Detect whether the number of the multiple word segments exceeds the maximum number threshold that the generative large model can accept; In response to the number of the multiple word segments exceeding the maximum number threshold, the content of each of the relevant paragraphs is rewritten to obtain the rewritten content of each of the relevant paragraphs, such that the number of characters in the rewritten content of each of the relevant paragraphs is less than the number of characters in the content of each of the relevant paragraphs, but the semantics remain unchanged; Based on the rewritten content of each of the multiple related paragraphs, the generative large model is used to generate the target document corresponding to the target topic.
10. The apparatus according to claim 9, wherein, The paragraph acquisition module is used for: Obtain the vector representation of the target topic; Based on the vector representation of the target topic and the vector representation of the paragraphs included in each original document in the document database, the vector representations of multiple related paragraphs related to the target topic are retrieved.
11. The apparatus according to claim 9, wherein, The paragraph acquisition module is used for: Based on the vector representation of each related paragraph in the plurality of related paragraphs, the vector representation of the paragraphs included in each original document in the document database, and the position identifier, the position identifier of each related paragraph is obtained respectively.
12. The apparatus according to claim 9, wherein, The paragraph acquisition module is used for: Based on the position identifiers of each relevant paragraph, the position identifiers of the paragraphs included in each original document in the document database, and the content, the content of each relevant paragraph in the corresponding original document is obtained.
13. The apparatus according to claim 9, wherein, Also includes: The acquisition module is used to acquire the multiple original documents; The format transcoding module is used to transcode the multiple original documents to make the format of the multiple original documents uniform; The information acquisition module is used to acquire the content of each paragraph in the original document, the vector representation of the paragraph, and the position identifier of the paragraph. The database generation module is used to generate the document database based on the content of paragraphs, the vector representation of paragraphs, and the position identifier of paragraphs in each of the multiple original documents.
14. The apparatus according to claim 10, wherein, The document generation module is also used for: In response to the condition that the number of the multiple word segments is not greater than the maximum number threshold, the target document corresponding to the target topic is generated based on the multiple word segments and using the generative big model.
15. The apparatus according to claim 9, wherein, The document generation module is used for: Obtain rewriting prompt information, wherein the rewriting prompt information requires that the number of words in the rewritten content of each relevant paragraph is less than the number of words in the content of each relevant paragraph; Based on the content of each relevant paragraph and the rewriting prompt information, the generative large model is used to obtain the rewritten content of each relevant paragraph.
16. The apparatus according to any one of claims 9-15, wherein, The document generation module is used for: Obtain the identifier of the original document to which each of the relevant paragraphs belongs; Obtain document prompt information for generating the target document; the document prompt information is used to limit the source information of each sentence or paragraph in the generated target document. Based on the content of each of the multiple related paragraphs, the identifier of the original document to which each of the related paragraphs belongs, and the document prompt information, the generative big model is used to generate the target document corresponding to the target topic; The source information for each statement or paragraph is marked in the target document.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Text generation method and device, electronic equipment and storage medium
CN116882372A