Professional document generation method and device, equipment, storage medium and program product
Patent Information
- Application Number
- CN202411419398.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-10-11
AI Technical Summary
[0003]但是,上述专业文书生成过程中,存在专业文书的撰写效率比较低的问题
[0045]上述专业文书生成方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,服务器首先获取初始文本,并根据初始文本和专业文书生成模型,生成初始专业文书,而后,根据文书规范处理模型,对初始专业文书的文书结构进行规范处理,得到规范专业文书,其中,文书规范处理模型为根据多个样本初始专业文书和各样本初始专业文书对应的样本规范专业文书训练得到的,在该专业文书生成方法中,通过专业文书生成模型能够快速地根据初始文本生成初始专业文书,提高得到初始专业文书的效率,进一步的,根据文书规范处理模型对初始专业文书的文书结构进行规范处理,能够提高初始专业文书的规范度和专业度,从而能够得到符合专业文书撰写规范的规范专业文书,避免了专业文书不符合撰写规范造成无效撰写的问题,进而提高专业文书的撰写效率。
Smart Images

Figure CN119416759B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating professional documents. Background Technology
[0002] Typically, in the process of generating professional documents, the completed draft needs to be reviewed before submission or publication. For example, with patent applications, a draft can be written based on the technical disclosure document, and then the patent application can be submitted after the draft has been reviewed and approved.
[0003] However, the process of generating the aforementioned professional documents suffers from relatively low writing efficiency. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the writing efficiency of professional documents, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a method for generating professional documents. The method includes:
[0006] Obtain the initial text, and generate the initial professional document based on the initial text and the professional document generation model;
[0007] Based on the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document;
[0008] The document standardization processing model is trained based on multiple initial sample professional documents and the corresponding sample standard professional documents for each initial sample professional document.
[0009] In one embodiment, the document standardization processing model includes a first slicing layer, a first matching layer, and a paragraph processing layer. The step of standardizing the document structure of the initial professional document according to the document standardization processing model to obtain a standardized professional document includes:
[0010] The initial professional document is segmented into paragraphs using the first layering method to obtain multiple paragraphs to be matched.
[0011] For each paragraph to be matched, the first matching layer is used to perform paragraph matching processing on the paragraph to be matched to obtain at least one matched paragraph;
[0012] Using the paragraph processing layer, the document structure of the initial professional document is standardized according to each matched paragraph to obtain the standardized professional document.
[0013] In one embodiment, the paragraph processing layer includes a second segmentation layer, a second matching layer, and a sentence processing layer. The paragraph processing layer is used to standardize the document structure of the initial professional document based on each matched paragraph to obtain the standardized professional document, including:
[0014] The second segmentation layer is used to perform sentence segmentation on each of the matched paragraphs to obtain multiple sentences to be matched.
[0015] For each of the statements to be matched, the second matching layer is used to perform statement matching processing on the statements to be matched to obtain at least one matching statement;
[0016] Using the statement processing layer, the document structure of the initial professional document is standardized according to each of the matching statements to obtain the standardized professional document.
[0017] In one embodiment, generating the initial professional document based on the initial text and the professional document generation model includes:
[0018] The initial text is evaluated for similarity based on a preset query knowledge base, and the evaluation result is obtained. The query knowledge base is constructed based on multiple standard professional documents.
[0019] If the evaluation result indicates that the initial text meets the conditions for generating professional documents, the initial professional document is generated based on the initial text and the professional document generation model.
[0020] In one embodiment, the query knowledge base includes a first vector corresponding to each of the standard professional documents, and the step of performing a similarity assessment on the initial text based on the preset query knowledge base to obtain an assessment result includes:
[0021] The initial text is vectorized to obtain a second vector corresponding to the initial text.
[0022] Based on the second vector, a target vector similar to the second vector is determined from each of the first vectors;
[0023] The evaluation result is determined based on the second vector and the target vector.
[0024] In one embodiment, determining the evaluation result based on the second vector and the target vector includes:
[0025] Calculate the similarity between the second vector and the target vector;
[0026] If the similarity is less than a preset similarity threshold, then the evaluation result is determined to be that the initial text meets the professional document generation conditions;
[0027] If the similarity is greater than or equal to the similarity threshold, then the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0028] In one embodiment, the query knowledge base further includes query paragraphs corresponding to each of the first vectors, and the step of generating the initial professional document based on the initial text and the professional document generation model includes:
[0029] Obtain the target query paragraph corresponding to the target vector from the query knowledge base;
[0030] The initial professional document is generated based on the target query paragraph, the initial text, and the professional document generation model.
[0031] In one embodiment, generating the initial professional document based on the target query paragraph, the initial text, and the professional document generation model includes:
[0032] Obtain supplementary text related to the initial text;
[0033] The target query paragraph, the initial text, and the supplementary text are input into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0034] In one embodiment, the method further includes:
[0035] Obtain the aforementioned standard professional documents and perform paragraph segmentation on each of the aforementioned standard professional documents to obtain multiple segmented paragraphs;
[0036] Based on the segmented segments, multiple segments to be processed are obtained, and each segment to be processed is subjected to segment enhancement processing to obtain multiple query segments;
[0037] Each query segment is vectorized to obtain the first vector corresponding to each query segment, and the correspondence between each query segment and each first vector is stored in the query knowledge base.
[0038] Secondly, this application also provides a professional document generation device. The device includes:
[0039] The generation module is used to obtain initial text and generate initial professional documents based on the initial text and the professional document generation model;
[0040] The processing module is used to standardize the document structure of the initial professional document according to the document standardization processing model to obtain a standardized professional document.
[0041] The document standardization processing model is trained based on multiple initial sample professional documents and the corresponding sample standard professional documents for each initial sample professional document.
[0042] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in the first aspect above.
[0043] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in the first aspect above.
[0044] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.
[0045] The aforementioned professional document generation method, apparatus, computer equipment, computer-readable storage medium, and computer program product involve a server first acquiring initial text and generating initial professional documents based on the initial text and a professional document generation model. Then, according to a document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document. The document standardization processing model is trained using multiple sample initial professional documents and corresponding sample standardized professional documents. In this professional document generation method, the professional document generation model can quickly generate initial professional documents from the initial text, improving the efficiency of obtaining initial professional documents. Furthermore, standardizing the document structure of the initial professional document according to the document standardization processing model can improve the standardization and professionalism of the initial professional document, thereby obtaining a standardized professional document that conforms to professional document writing standards. This avoids the problem of invalid writing caused by professional documents not conforming to writing standards, thus improving the efficiency of professional document writing. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a diagram illustrating the application environment of a professional document generation method in one embodiment.
[0048] Figure 2 This is a flowchart illustrating a professional document generation method in one embodiment;
[0049] Figure 3 This is a flowchart illustrating step 202 in another embodiment;
[0050] Figure 4 This is a flowchart illustrating step 303 in another embodiment;
[0051] Figure 5 This is a flowchart illustrating step 201 in another embodiment;
[0052] Figure 6 This is a flowchart illustrating step 501 in another embodiment;
[0053] Figure 7 This is a flowchart illustrating step 502 in another embodiment;
[0054] Figure 8 This is a flowchart illustrating a professional document generation method in another embodiment;
[0055] Figure 9 This is a flowchart illustrating a professional document generation method in another embodiment;
[0056] Figure 10 This is a structural block diagram of a professional document generation device in one embodiment;
[0057] Figure 11 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0059] In the application of professional document drafting, such as patent drafting, as companies increasingly value intellectual property protection, they promptly apply for patents to safeguard their rights whenever new technological or product innovations emerge. Typically, drafted patent documents require review, often involving manual review to ensure compliance before revisions according to patent format. This process not only involves searching relevant patent documents but also manually comparing innovative points and evaluating the patent's innovativeness, consuming significant manpower and time. Furthermore, overlapping patent protection points within the same company can lead to duplicate drafts and invalid work. Therefore, traditional patent drafting methods suffer from relatively low efficiency.
[0060] In view of this, this application provides a method for generating professional documents. The server first obtains initial text and generates initial professional documents based on the initial text and a professional document generation model. Then, the document structure of the initial professional documents is standardized according to a document standardization processing model to obtain standardized professional documents. The document standardization processing model is trained based on multiple sample initial professional documents and sample standardized professional documents corresponding to each sample initial professional document. In this professional document generation method, the professional document generation model can quickly generate initial professional documents based on the initial text, improving the efficiency of obtaining initial professional documents. Furthermore, standardizing the document structure of the initial professional documents according to the document standardization processing model can improve the standardization and professionalism of the initial professional documents, thereby obtaining standardized professional documents that conform to the writing standards of professional documents, avoiding the problem of invalid writing caused by professional documents not conforming to the writing standards, and thus improving the writing efficiency of professional documents.
[0061] The professional document generation method provided in this application embodiment can be applied to, for example, Figure 1 The implementation environment shown includes a server, which can be a standalone server or a server cluster consisting of multiple servers. A data storage system stores the data that the server needs to process. The data storage system can be integrated onto the server or located in the cloud or on other network servers. Specifically, the server acquires initial text and generates initial professional documents based on the initial text and a professional document generation model. Then, according to a document standardization processing model, the document structure of the initial professional documents is standardized to obtain standardized professional documents. The document standardization processing model is trained using multiple sample initial professional documents and corresponding sample standardized professional documents for each sample initial professional document.
[0062] In other possible implementations, the professional document generation method provided in this application can also be applied to a terminal. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc.
[0063] In one exemplary embodiment, such as Figure 2 As shown, a professional document generation method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 201 and 202:
[0064] Step 201: Obtain the initial text and generate the initial professional document based on the initial text and the professional document generation model.
[0065] It should be noted that the professional document generation method provided in this application can be applied to various scenarios of generating professional documents, such as patent documents and legal documents. This application takes patent documents as an example of professional documents to describe the process of obtaining standardized professional documents.
[0066] The initial text refers to the text to be written. For example, if the initial text is the initial text corresponding to a patent document, then the initial text can be the text content used to describe a technology or a product. The inventive points and innovative points of the technology or product can be determined through the initial text.
[0067] It is understandable that after obtaining the initial text, it is necessary to generate corresponding professional documents based on the initial text. In order to improve the efficiency and accuracy of professional document generation, in this embodiment, a trained professional document generation model can be used to generate professional documents. Here, the professional document generation model is a neural network model used to generate professional documents. For example, the professional document generation model can be a large language model, or it can be a recurrent neural network model. This embodiment does not limit this.
[0068] In this embodiment, the server can receive data packets containing initial text sent by other terminals, servers, etc. via the network. By parsing the data packets, the initial text can be obtained from the parsing results. Then, the initial text can be input into the professional document generation model, and the output result of the professional document generation model can be used as the initial professional document.
[0069] Step 202: Based on the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document; wherein, the document standardization processing model is trained based on multiple sample initial professional documents and the sample standardized professional documents corresponding to each sample initial professional document.
[0070] The document standardization processing model refers to a neural network model that performs standardization processing on the generated initial professional documents. Optionally, the document standardization processing model can be a convolutional neural network model or a recurrent neural network model. This embodiment does not limit this.
[0071] Understandably, the initial professional document merely transforms the initial text into a document conforming to a professional document template. However, different types of professional documents may correspond to different document structures, such as title, abstract, body, and appendices, as well as the claims, description, and abstract of a patent document. Therefore, it is necessary to standardize the initial professional document according to the corresponding document structure to obtain a standardized professional document that conforms to the professional document structure. The standardized professional document is obtained after standardizing the document structure corresponding to the initial professional document.
[0072] To reduce the workload in the initial professional document standardization process, a document standardization model can be used to standardize the initial professional documents. In this embodiment, during the process of obtaining the document standardization model, multiple sample initial professional documents can be obtained as training samples, and the sample standardized professional documents corresponding to these multiple sample initial professional documents can be used as the gold standard. Then, the initial document standardization model is trained using the training samples and the gold standard until the loss function value corresponding to the initial document standardization model reaches a minimum or a stable value during a certain training process. The initial document standardization model corresponding to this training process is then used as the document standardization model.
[0073] In this embodiment, the server can input the initial professional document into the document standardization processing model, and use the document standardization processing model to adjust the document structure of the initial professional document to obtain the adjusted initial professional document. The adjusted initial professional document output by the document standardization processing model is then used as the standardized professional document.
[0074] In the aforementioned professional document generation method, the server first obtains the initial text and generates an initial professional document based on the initial text and the professional document generation model. Then, according to the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document. The document standardization processing model is trained based on multiple sample initial professional documents and sample standardized professional documents corresponding to each sample initial professional document. In this professional document generation method, the professional document generation model can quickly generate initial professional documents from the initial text, improving the efficiency of obtaining initial professional documents. Furthermore, standardizing the document structure of the initial professional document according to the document standardization processing model can improve the standardization and professionalism of the initial professional document, thereby obtaining a standardized professional document that conforms to the professional document writing standards, avoiding the problem of invalid writing caused by professional documents that do not conform to the writing standards, and thus improving the writing efficiency of professional documents.
[0075] based on Figure 2 The illustrated embodiment can be found in [reference]. Figure 3 The document standardization processing model includes a first slicing layer, a first matching layer, and a paragraph processing layer. This embodiment relates to how the server standardizes the document structure of an initial professional document according to the document standardization processing model to obtain a standardized professional document. Step 202 includes... Figure 3 Steps 301-303 are shown.
[0076] Step 301: Use the first layering method to perform paragraph segmentation on the initial professional document to obtain multiple paragraphs to be matched.
[0077] It should be noted that in the document standardization processing model, to improve the standardization of the output results, the initial professional document can be standardized at both the paragraph and sentence levels. The first segmentation layer refers to the processing layer that segments the initial professional document into paragraphs. In this embodiment, the first initial segmentation layer in the initial document standardization processing model can be trained using the paragraph segmentation results corresponding to the sample professional document and the sample standardized professional document to obtain the first segmentation layer.
[0078] In this embodiment, the server can use the first segmentation layer in the document standardization processing model to segment the initial professional document into multiple paragraphs according to punctuation marks and paragraph marks, and use these multiple paragraphs as multiple paragraphs to be matched.
[0079] Step 302: For each paragraph to be matched, use the first matching layer to perform paragraph matching processing to obtain at least one matched paragraph.
[0080] The first matching layer refers to the processing layer that matches the paragraph to be matched to obtain paragraphs similar to the paragraph to be matched. In this embodiment, the first matching layer can search for and obtain paragraphs similar to the paragraph to be matched from multiple historical standard sample documents, and use them as matching paragraphs.
[0081] Optionally, the matching paragraph may include one paragraph similar to the paragraph to be matched, or the matching paragraph may include multiple paragraphs similar to the paragraph to be matched. The matching paragraph may be similar in content to the paragraph to be matched, or it may have a similar structure to the paragraph to be matched, or it may have a similar writing style to the paragraph to be matched.
[0082] In this embodiment, the server can use the first matching layer in the document standardization processing model to find paragraphs similar to the paragraph to be matched from the paragraphs included in multiple historical standardization sample documents, and use the similar paragraphs as the matching paragraphs.
[0083] As an optional implementation, if there are multiple paragraphs similar to the paragraph to be matched, the first n paragraphs can be selected as matching paragraphs based on the similarity between each paragraph and the paragraph to be matched.
[0084] Step 303: Using the paragraph processing layer, the document structure of the initial professional document is standardized according to each matching paragraph to obtain a standardized professional document.
[0085] The paragraph processing layer refers to the layer that processes the matched paragraphs to transform them into standardized paragraphs. In this embodiment, the server can utilize the paragraph processing layer in the document standardization processing model to standardize the document structure of the initial professional document based on the writing style, sentence description style, etc., corresponding to each matched paragraph, thereby obtaining a standardized professional document.
[0086] In this embodiment, the server uses a first segmentation layer to segment the initial professional document into paragraphs, obtaining multiple paragraphs to be matched. Then, for each paragraph to be matched, the first matching layer performs paragraph matching processing to obtain at least one matching paragraph. Subsequently, the paragraph processing layer is used to standardize the document structure of the initial professional document based on each matching paragraph, resulting in a standardized professional document. Since the initial professional document is segmented into multiple paragraphs to be matched and each paragraph is matched separately, the accuracy of the matching process is improved, thereby increasing the matching degree of the obtained matching paragraphs. Furthermore, by standardizing the document structure of the initial professional document at the paragraph level through each matching paragraph, the standardization and professionalism of the obtained standardized professional document are improved.
[0087] based on Figure 3 The illustrated embodiment can be found in [reference]. Figure 4The paragraph processing layer includes a second segmentation layer, a second matching layer, and a sentence processing layer. This embodiment relates to how the server utilizes the paragraph processing layer to standardize the document structure of the initial professional document based on each matched paragraph, thereby obtaining a standardized professional document. Step 303 includes... Figure 4 Steps 401-403 are shown.
[0088] Step 401: Use the second layering method to perform sentence segmentation on each matching paragraph to obtain multiple sentences to be matched.
[0089] The second segmentation layer refers to the processing layer that segments each matched paragraph into sentences. In this embodiment, the second initial segmentation layer in the initial document standardization processing model can be trained using the sentence segmentation results corresponding to the sample professional documents and the sample standard professional documents to obtain the second segmentation layer.
[0090] In this embodiment, the server can use the second segmentation in the document standardization processing model to segment each matching paragraph into multiple sentences according to punctuation marks, keywords, and stop words, and use these multiple sentences as multiple sentences to be matched.
[0091] Step 402: For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0092] The second matching layer refers to the processing layer that matches the statement to be matched to obtain statements similar to the statement to be matched. In this embodiment, the second matching layer can search for and obtain statements similar to the statement to be matched from multiple historical specification sample documents, and use these as matching statements.
[0093] Optionally, the matching statement may include a statement similar to the statement to be matched, or the matching statement may include multiple statements similar to the statement to be matched. The matching statement may have a similar meaning to the statement to be matched, or the matching statement may have similar wording to the statement to be matched, or the matching statement may have similar writing style to the statement to be matched.
[0094] In this embodiment, the server can use the second matching layer in the document standardization processing model to find statements similar to the statement to be matched from the statements included in multiple historical standardization sample documents, and use the similar statements as matching statements.
[0095] As an optional implementation, if there are multiple statements similar to the statement to be matched, the first n statements can be selected as matching statements based on the degree of similarity between each statement and the statement to be matched.
[0096] Step 403: Using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement to obtain a standardized professional document.
[0097] The statement processing layer refers to the layer that processes the matched statements to transform them into standardized statements. In this embodiment, the server can utilize the statement processing layer in the document standardization processing model to standardize the document structure of the initial professional document based on the writing style and statement description style corresponding to each matched statement, thereby obtaining a standardized professional document.
[0098] In this embodiment, the server uses a second segmentation layer to segment each matching paragraph into multiple sentences to be matched. Then, for each sentence to be matched, the second matching layer performs sentence matching processing to obtain at least one matching sentence. Subsequently, the sentence processing layer uses each matching sentence to standardize the document structure of the initial professional document, resulting in a standardized professional document. By segmenting each paragraph into multiple sentences to be matched and performing matching processing on each sentence separately, the accuracy and precision of the matching process are improved, thereby increasing the matching degree of the obtained matching sentences. Furthermore, by standardizing the document structure of the initial professional document at the sentence level through each matching sentence, the standardization and professionalism of the obtained standardized professional document are improved.
[0099] based on Figure 2 The illustrated embodiment can be found in [reference]. Figure 5 This embodiment relates to the process by which a server generates an initial professional document based on initial text and a professional document generation model. Step 201 includes... Figure 5 Steps 501 and 502 are shown.
[0100] Step 501: Perform a similarity assessment on the initial text based on a preset query knowledge base to obtain the assessment results. The query knowledge base is constructed based on multiple standard professional documents.
[0101] The query knowledge base includes multiple standardized professional documents after processing. For example, if the query knowledge base corresponds to a patent document, it will include multiple submitted or granted patent documents. Different types of professional documents can correspond to different query knowledge bases. Understandably, when drafting an initial text into a professional document, to avoid invalid drafting due to identical initial texts already being drafted, it is necessary to first determine whether the initial text can be drafted.
[0102] In this embodiment, the server can compare the initial text with standard professional documents included in a preset query knowledge base, evaluate the similarity of the initial text based on the comparison results, obtain the evaluation results, and determine whether the initial text can be written based on the evaluation results.
[0103] Step 502: If the evaluation result shows that the initial text meets the conditions for generating professional documents, generate the initial professional document based on the initial text and the professional document generation model.
[0104] The "professional document generation condition" refers to a low similarity between the content of the initial text and the content of the standard professional document. In other words, the professional document written from the initial text differs significantly from the standard professional document, and the content of the initial text cannot be disclosed by the standard professional document. For example, taking patent documents as an example, the professional document generation condition means that the professional document corresponding to the initial text has a low similarity to the submitted patent document, and the submitted patent document cannot disclose the professional document corresponding to the initial text.
[0105] In this embodiment, the server can determine whether the initial text can be written based on the evaluation results. If the initial text can be written, it can be determined that the initial text meets the conditions for generating professional documents. The server can then input the initial text into the professional document generation model and determine the output of the professional document generation model as the initial professional document.
[0106] In this embodiment, the server performs a similarity assessment on the initial text based on a preset query knowledge base and obtains the assessment result. Since the query knowledge base is constructed based on multiple standard professional documents, it can accurately assess whether the initial text can be written. Thus, if the assessment result indicates that the initial text meets the conditions for generating a professional document, the initial professional document is generated based on the initial text and the professional document generation model. This avoids the problem of invalid loading and unloading and wasted writing resources due to the inability to write the initial text, thereby improving the writing efficiency of professional documents.
[0107] based on Figure 5 The illustrated embodiment can be found in [reference]. Figure 6 The query knowledge base includes the first vector corresponding to each standard professional document. This embodiment involves how the server performs a similarity assessment on the initial text based on the preset query knowledge base to obtain the assessment result. Step 501 includes... Figure 6 Steps 601-603 are shown.
[0108] Step 601: Vectorize the initial text to obtain the second vector corresponding to the initial text.
[0109] It should be noted that when evaluating the similarity between texts, since text data is unstructured and difficult to process, the text content can be converted into a numerical form that computer devices can recognize, that is, the text data can be converted into vector data. Optionally, the vectorization algorithm can be the bag-of-words model, or the vectorization algorithm can also be Word2Vec; this embodiment does not limit this. The vector corresponding to each standard professional document in the knowledge base is the first vector, and the vector corresponding to the initial text is the second vector.
[0110] In this embodiment, the server can use a vectorization processing algorithm to vectorize the initial text and determine the result of the vectorization processing as the second vector.
[0111] Step 602: Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors.
[0112] Here, the target vector refers to the vector in the first vector that is similar to the second vector. It can be understood that if there exists a vector in the first vector that is similar to the second vector, then that first vector can be identified as the target vector.
[0113] In this embodiment, the server can calculate the cosine similarity between the second vector and each first vector, and determine whether the first vector and the second vector are similar based on the cosine similarity, thereby identifying the first vector that is similar to the second vector as the target vector.
[0114] Step 603: Determine the evaluation result based on the second vector and the target vector.
[0115] The evaluation result refers to whether the initial text meets the conditions for generating professional documents. Optionally, the evaluation result can be that the initial text meets the conditions for generating professional documents, or the evaluation result can be that the initial text does not meet the conditions for generating professional documents.
[0116] In one possible implementation, step 603 above can be achieved by calculating the similarity between the second vector and the target vector, and determining the evaluation result based on the similarity.
[0117] The similarity calculation methods can include Euclidean distance, Pearson similarity, Jaccard similarity, etc.
[0118] In this embodiment, the server can use a similarity calculation method to calculate the similarity between the second vector and the target vector, and then determine the evaluation result corresponding to the similarity based on the similarity.
[0119] In one optional implementation, if the similarity is less than a preset similarity threshold, the evaluation result is determined to be that the initial text meets the conditions for generating professional documents.
[0120] In another alternative implementation, if the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0121] The similarity threshold can be a similarity threshold determined based on the historical initial text and the corresponding historical standard professional documents. For example, the similarity threshold can be 10%, or it can be 20%. This embodiment does not limit this.
[0122] In this implementation, the server calculates the similarity between the second vector and the target vector, and determines the evaluation result based on the similarity. If the similarity is less than a preset similarity threshold, the server determines that the initial text meets the conditions for generating professional documents. If the similarity is greater than or equal to the similarity threshold, the server determines that the initial text does not meet the conditions for generating professional documents. Since the process of judging the relationship between similarity and similarity threshold is relatively simple, it can quickly and accurately determine whether the initial text meets the conditions for generating professional documents.
[0123] In this embodiment, the server vectorizes the initial text to obtain a second vector corresponding to the initial text. Then, based on the second vector, a target vector similar to the second vector is determined from each of the first vectors. Subsequently, the evaluation result is determined based on the second vector and the target vector. Since vector data can be recognized by computer devices compared to text data, the server can quickly process the vector data. By using the second vector corresponding to the initial text and the first vector corresponding to the standard professional document, it can determine whether the initial text meets the conditions for generating a professional document. The server can quickly and accurately determine the target vector from the first vector, thereby improving the efficiency of determining whether to generate an initial professional document based on the initial text according to the evaluation result.
[0124] based on Figure 5 The illustrated embodiment can be found in [reference]. Figure 7 The knowledge base query also includes query paragraphs corresponding to each first vector. This embodiment relates to the process by which the server generates initial professional documents based on the initial text and the professional document generation model. Step 502 includes... Figure 7 Steps 701 and 702 are shown.
[0125] Step 701: Obtain the target query paragraph corresponding to the target vector from the query knowledge base.
[0126] Here, the target query segment refers to the text data corresponding to the target vector. It can be understood that the query knowledge base can store the correspondence between multiple query segments and their corresponding vectors.
[0127] In this embodiment, the server can obtain the query paragraph corresponding to the target vector from the query knowledge base based on the target vector, and use the query paragraph as the target query paragraph.
[0128] Step 702: Generate the initial professional document based on the target query paragraph, the initial text, and the professional document generation model.
[0129] In this embodiment, the server can input the target query paragraph and the initial text into the professional document generation model. In the professional document generation model, the content of the initial text can be modified and written according to the text structure and language description of the target query paragraph, thereby obtaining the initial professional document.
[0130] In one possible implementation, step 702 above can also be achieved through the following steps:
[0131] Obtain supplementary text related to the initial text; input the target query paragraph, the initial text, and the supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0132] The supplementary text is text related to the content of the initial text. For example, the supplementary text may be background information related to the initial text, or the initial text may be text describing the beneficial effects of the initial text.
[0133] In this embodiment, the server can input the supplementary text, the target query paragraph, and the initial text into the professional document generation model, so that the professional document generation model can use the supplementary text to expand and supplement the content of the initial text, thereby outputting the initial professional document.
[0134] In this implementation, since the server obtains the supplementary text corresponding to the initial text, it can expand the content of the initial text using the supplementary text. Then, the target query paragraph, the initial text, and the supplementary text are input into the professional document generation model, which can improve the richness of the obtained initial professional document.
[0135] In this embodiment, the server obtains the target query paragraph corresponding to the target vector from the query knowledge base. Based on the target query paragraph, the initial text, and the professional document generation model, it can generate an initial professional document. Since the target query paragraph is obtained from the query knowledge base and is similar to the initial text, when the professional document generation model writes the initial text based on the target query paragraph, the resulting initial professional document can conform to the text content of the initial text as well as the writing style of the target query paragraph, thereby improving the accuracy and professionalism of the resulting initial professional document.
[0136] In one embodiment, based on the above embodiments, see [link to embodiment]. Figure 8 This embodiment describes how a server constructs a query knowledge base. The server can perform actions such as... Figure 8 Steps 801-803 shown implement this process.
[0137] Step 801: Obtain each standard professional document and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs.
[0138] In this context, each standard professional document refers to a professional document that has undergone standardized processing based on the initial text within a historical time period. In this embodiment, the server can segment the standard professional document into paragraphs based on paragraph marks, key sentences, or other identifying phrases or words, thereby obtaining multiple segmented paragraphs.
[0139] Step 802: Based on each segmented segment, obtain multiple segments to be processed, and perform segment enhancement processing on each segment to obtain multiple query segments.
[0140] Understandably, in the process of segmenting various standard professional documents into multiple segmented paragraphs, the segmentation based on paragraph markers, key sentences, or other identifying phrases or words may result in some segmented paragraphs having fewer words. This could lead to inaccurate evaluation results when using the segmented paragraphs to assess whether the initial text can be written. Therefore, a paragraph word count threshold can be set. Based on the segmented paragraph number, if the word count of multiple adjacent paragraphs is less than the threshold, the multiple paragraphs can be merged, and the merged paragraph can be used as the paragraph to be processed.
[0141] The "paragraphs to be processed" refer to those requiring paragraph enhancement. Paragraph enhancement involves adding pre-processing to the first sentence of each paragraph to further strengthen the distinction between them. The "query paragraphs" refer to those paragraphs that need to be compared with the initial text when evaluating its suitability for writing using a knowledge base.
[0142] In this embodiment, the server can first merge paragraphs with fewer than the paragraph word count threshold based on the word count of each segmented paragraph and the paragraph word count threshold. Then, the merged paragraphs and paragraphs with word counts equal to or greater than the paragraph word count threshold are taken as paragraphs to be processed. Subsequently, a paragraph enhancement processing algorithm is used to enhance each paragraph to be processed, and the paragraphs with enhanced paragraphs are taken as query paragraphs.
[0143] Step 803: Vectorize each query segment to obtain the first vector corresponding to each query segment, and store the correspondence between each query segment and each first vector in the query knowledge base.
[0144] In this embodiment, the server can vectorize each query segment according to the vector processing algorithm to obtain the first vector corresponding to each query segment. Then, a correspondence between each query segment and the first vector corresponding to each query segment can be established. After that, the established correspondence can be stored in the query knowledge base.
[0145] In this embodiment, the server obtains various standard professional documents and performs paragraph segmentation on each document to obtain multiple segmented paragraphs. Then, based on each segmented paragraph, it obtains multiple paragraphs to be processed and performs paragraph enhancement processing on each paragraph to obtain multiple query paragraphs. Subsequently, by vectorizing each query paragraph, it obtains the first vector corresponding to each query paragraph and stores the correspondence between each query paragraph and each first vector in the query knowledge base. This allows the query knowledge base to be updated, ensuring that it includes the correspondence between multiple standard professional documents and their corresponding first vectors. Furthermore, based on the evaluation results in the query knowledge base that determine whether the initial text can be written, the accuracy of writing the initial text into an initial professional document is improved.
[0146] In one embodiment, a professional document generation method is provided for use on a server, such as... Figure 9 As shown, the method includes the following steps:
[0147] Step 901: Obtain each standard professional document and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs.
[0148] Step 902: Based on each segmented segment, obtain multiple segments to be processed, and perform segment enhancement processing on each segment to obtain multiple query segments.
[0149] Step 903: Vectorize each query segment to obtain the first vector corresponding to each query segment, and store the correspondence between each query segment and each first vector in the query knowledge base.
[0150] Step 904: Obtain the initial text.
[0151] Step 905: Vectorize the initial text to obtain the second vector corresponding to the initial text.
[0152] Step 906: Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors.
[0153] Step 907: Calculate the similarity between the second vector and the target vector.
[0154] Step 908: If the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0155] Step 909: If the similarity is less than the preset similarity threshold, then the evaluation result is determined to be that the initial text meets the conditions for generating professional documents.
[0156] Step 9010: Obtain the target query paragraph corresponding to the target vector from the query knowledge base.
[0157] Step 9011: Obtain supplementary text related to the initial text.
[0158] Step 9012: Input the target query paragraph, initial text, and supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0159] Step 9013: Use the first layering method to perform paragraph segmentation on the initial professional document to obtain multiple paragraphs to be matched.
[0160] Step 9014: For each paragraph to be matched, use the first matching layer to perform paragraph matching processing on the paragraph to be matched, and obtain at least one matched paragraph.
[0161] Step 9015: Use the second layering method to perform sentence segmentation on each matching paragraph to obtain multiple sentences to be matched.
[0162] Step 9016: For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0163] Step 9017: Using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement to obtain a standardized professional document.
[0164] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0165] Based on the same inventive concept, this application also provides a professional document generation apparatus for implementing the professional document generation method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more embodiments of the professional document generation apparatus provided below can be found in the limitations of the professional document generation method described above, and will not be repeated here.
[0166] In one exemplary embodiment, such as Figure 10 As shown, a professional document generation device is provided, comprising:
[0167] The generation module 1001 is used to obtain the initial text and generate the initial professional document based on the initial text and the professional document generation model;
[0168] The processing module 1002 is used to standardize the document structure of the initial professional document according to the document standardization processing model, so as to obtain a standardized professional document.
[0169] The document standardization processing model was trained based on multiple initial professional documents and the corresponding standardized professional documents for each initial professional document.
[0170] In one embodiment, the document standardization processing model includes a first slicing layer, a first matching layer, and a paragraph processing layer. The processing module 1002 includes:
[0171] The first acquisition unit is used to perform paragraph segmentation on the initial professional document using the first cutting layer to obtain multiple paragraphs to be matched.
[0172] The second acquisition unit is used to perform paragraph matching processing on each paragraph to be matched using the first matching layer to obtain at least one matched paragraph.
[0173] The third acquisition unit is used to standardize the document structure of the initial professional document based on each matched paragraph using the paragraph processing layer, so as to obtain a standardized professional document.
[0174] In one embodiment, the paragraph processing layer includes a second segmentation layer, a second matching layer, and a statement processing layer. The third acquisition unit is specifically used to: perform statement segmentation processing on each matched paragraph using the second segmentation layer to obtain multiple statements to be matched.
[0175] For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0176] By using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement, resulting in a standardized professional document.
[0177] In one embodiment, the generation module 1001 includes:
[0178] The fourth acquisition unit is used to perform similarity assessment on the initial text based on a preset query knowledge base and obtain the assessment result. The query knowledge base is constructed based on multiple standard professional documents.
[0179] The generation unit is used to generate an initial professional document based on the initial text and the professional document generation model, provided that the evaluation result indicates that the initial text meets the conditions for generating professional documents.
[0180] In one embodiment, the knowledge base is queried, which includes a first vector corresponding to each standard professional document. The fourth acquisition unit is specifically used to: perform vectorization processing on the initial text to obtain a second vector corresponding to the initial text.
[0181] Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors;
[0182] The evaluation result is determined based on the second vector and the target vector.
[0183] In one embodiment, the fourth acquisition unit is specifically used to: calculate the similarity between the second vector and the target vector;
[0184] If the similarity is less than the preset similarity threshold, the evaluation result is determined to be that the initial text meets the conditions for generating professional documents;
[0185] If the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0186] In one embodiment, the query knowledge base further includes query paragraphs corresponding to each first vector, and a generation unit is specifically used to: obtain the target query paragraphs corresponding to the target vectors from the query knowledge base;
[0187] Based on the target query paragraph, initial text, and professional document generation model, generate initial professional documents.
[0188] In one embodiment, the generating unit is specifically used to: obtain supplementary text related to the initial text;
[0189] Input the target query paragraph, initial text, and supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0190] In one embodiment, the device further includes:
[0191] The first acquisition module is used to acquire various standard professional documents and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs.
[0192] The second acquisition module is used to acquire multiple paragraphs to be processed based on each segmented paragraph, and to perform paragraph enhancement processing on each paragraph to be processed to obtain multiple query paragraphs.
[0193] The third acquisition module is used to vectorize each query paragraph to obtain the first vector corresponding to each query paragraph, and store the correspondence between each query paragraph and each first vector in the query knowledge base.
[0194] Each module in the aforementioned professional document generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the server's processor in hardware form or independent of it, or stored in the server's memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0195] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores professional document data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a professional document generation method.
[0196] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0197] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0198] Obtain the initial text, and generate the initial professional document based on the initial text and the professional document generation model;
[0199] Based on the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document;
[0200] The document standardization processing model was trained based on multiple initial professional documents and the corresponding standardized professional documents for each initial professional document.
[0201] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0202] The initial professional document is segmented into paragraphs using the first layering method, resulting in multiple paragraphs to be matched.
[0203] For each paragraph to be matched, the first matching layer is used to perform paragraph matching processing on the paragraph to be matched, so as to obtain at least one matched paragraph.
[0204] By using the paragraph processing layer, the document structure of the initial professional document is standardized according to each matching paragraph, resulting in a standardized professional document.
[0205] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0206] The second layer of segmentation is used to segment each matching paragraph, resulting in multiple sentences to be matched.
[0207] For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0208] By using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement, resulting in a standardized professional document.
[0209] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0210] The initial text is evaluated for similarity based on a pre-defined query knowledge base, and the evaluation results are obtained. The query knowledge base is constructed based on multiple standard professional documents.
[0211] If the evaluation result indicates that the initial text meets the conditions for generating professional documents, an initial professional document is generated based on the initial text and the professional document generation model.
[0212] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0213] The initial text is vectorized to obtain the second vector corresponding to the initial text;
[0214] Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors;
[0215] The evaluation result is determined based on the second vector and the target vector.
[0216] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0217] Calculate the similarity between the second vector and the target vector;
[0218] If the similarity is less than the preset similarity threshold, the evaluation result is determined to be that the initial text meets the conditions for generating professional documents;
[0219] If the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0220] In one embodiment, the processor specifically implements the following steps when executing a computer program:
[0221] Retrieve the target query paragraph corresponding to the target vector from the knowledge base;
[0222] Based on the target query paragraph, initial text, and professional document generation model, generate initial professional documents.
[0223] In one embodiment, the processor also performs the following steps when executing a computer program:
[0224] Retrieve supplementary text related to the initial text;
[0225] Input the target query paragraph, initial text, and supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0226] In one embodiment, the processor also performs the following steps when executing a computer program:
[0227] Obtain various standard professional documents and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs;
[0228] Based on each segmented paragraph, multiple paragraphs to be processed are obtained, and paragraph enhancement processing is performed on each paragraph to be processed to obtain multiple query paragraphs;
[0229] Each query segment is vectorized to obtain the first vector corresponding to each query segment, and the correspondence between each query segment and each first vector is stored in the query knowledge base.
[0230] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0231] Obtain the initial text, and generate the initial professional document based on the initial text and the professional document generation model;
[0232] Based on the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document;
[0233] The document standardization processing model was trained based on multiple initial professional documents and the corresponding standardized professional documents for each initial professional document.
[0234] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0235] The initial professional document is segmented into paragraphs using the first layering method, resulting in multiple paragraphs to be matched.
[0236] For each paragraph to be matched, the first matching layer is used to perform paragraph matching processing on the paragraph to be matched, so as to obtain at least one matched paragraph.
[0237] By using the paragraph processing layer, the document structure of the initial professional document is standardized according to each matching paragraph, resulting in a standardized professional document.
[0238] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0239] The second layer of segmentation is used to segment each matching paragraph, resulting in multiple sentences to be matched.
[0240] For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0241] By using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement, resulting in a standardized professional document.
[0242] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0243] The initial text is evaluated for similarity based on a pre-defined query knowledge base, and the evaluation results are obtained. The query knowledge base is constructed based on multiple standard professional documents.
[0244] If the evaluation result indicates that the initial text meets the conditions for generating professional documents, an initial professional document is generated based on the initial text and the professional document generation model.
[0245] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0246] The initial text is vectorized to obtain the second vector corresponding to the initial text;
[0247] Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors;
[0248] The evaluation result is determined based on the second vector and the target vector.
[0249] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0250] Calculate the similarity between the second vector and the target vector;
[0251] If the similarity is less than the preset similarity threshold, the evaluation result is determined to be that the initial text meets the conditions for generating professional documents;
[0252] If the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0253] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0254] Retrieve the target query paragraph corresponding to the target vector from the knowledge base;
[0255] Based on the target query paragraph, initial text, and professional document generation model, generate initial professional documents.
[0256] In one embodiment, the computer program, when executed by a processor, further performs the following steps:
[0257] Retrieve supplementary text related to the initial text;
[0258] Input the target query paragraph, initial text, and supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0259] In one embodiment, the computer program, when executed by a processor, further performs the following steps:
[0260] Obtain various standard professional documents and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs;
[0261] Based on each segmented paragraph, multiple paragraphs to be processed are obtained, and paragraph enhancement processing is performed on each paragraph to be processed to obtain multiple query paragraphs;
[0262] Each query segment is vectorized to obtain the first vector corresponding to each query segment, and the correspondence between each query segment and each first vector is stored in the query knowledge base.
[0263] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0264] Obtain the initial text, and generate the initial professional document based on the initial text and the professional document generation model;
[0265] Based on the document standardization processing model, the document structure of the initial professional document is standardized to obtain a standardized professional document;
[0266] The document standardization processing model was trained based on multiple initial professional documents and the corresponding standardized professional documents for each initial professional document.
[0267] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0268] The initial professional document is segmented into paragraphs using the first layering method, resulting in multiple paragraphs to be matched.
[0269] For each paragraph to be matched, the first matching layer is used to perform paragraph matching processing on the paragraph to be matched, so as to obtain at least one matched paragraph.
[0270] By using the paragraph processing layer, the document structure of the initial professional document is standardized according to each matching paragraph, resulting in a standardized professional document.
[0271] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0272] The second layer of segmentation is used to segment each matching paragraph, resulting in multiple sentences to be matched.
[0273] For each statement to be matched, the second matching layer is used to perform statement matching processing to obtain at least one matched statement.
[0274] By using the statement processing layer, the document structure of the initial professional document is standardized according to each matching statement, resulting in a standardized professional document.
[0275] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0276] The initial text is evaluated for similarity based on a pre-defined query knowledge base, and the evaluation results are obtained. The query knowledge base is constructed based on multiple standard professional documents.
[0277] If the evaluation result indicates that the initial text meets the conditions for generating professional documents, an initial professional document is generated based on the initial text and the professional document generation model.
[0278] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0279] The initial text is vectorized to obtain the second vector corresponding to the initial text;
[0280] Based on the second vector, determine the target vector that is similar to the second vector from each of the first vectors;
[0281] The evaluation result is determined based on the second vector and the target vector.
[0282] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0283] Calculate the similarity between the second vector and the target vector;
[0284] If the similarity is less than the preset similarity threshold, the evaluation result is determined to be that the initial text meets the conditions for generating professional documents;
[0285] If the similarity is greater than or equal to the similarity threshold, the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
[0286] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps:
[0287] Retrieve the target query paragraph corresponding to the target vector from the knowledge base;
[0288] Based on the target query paragraph, initial text, and professional document generation model, generate initial professional documents.
[0289] In one embodiment, the computer program, when executed by a processor, further performs the following steps:
[0290] Retrieve supplementary text related to the initial text;
[0291] Input the target query paragraph, initial text, and supplementary text into the professional document generation model to obtain the initial professional document output by the professional document generation model.
[0292] In one embodiment, the computer program, when executed by a processor, further performs the following steps:
[0293] Obtain various standard professional documents and perform paragraph segmentation on each standard professional document to obtain multiple segmented paragraphs;
[0294] Based on each segmented paragraph, multiple paragraphs to be processed are obtained, and paragraph enhancement processing is performed on each paragraph to be processed to obtain multiple query paragraphs;
[0295] Each query segment is vectorized to obtain the first vector corresponding to each query segment, and the correspondence between each query segment and each first vector is stored in the query knowledge base.
[0296] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0297] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0298] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for generating professional documents, characterized in that, The method includes: Obtain the initial text, and generate the initial professional document based on the initial text and the professional document generation model; The initial professional document is segmented into paragraphs using the first layer of the document standardization processing model to obtain multiple paragraphs to be matched. For each paragraph to be matched, the first matching layer in the document standardization processing model is used to perform paragraph matching processing on the paragraph to be matched based on multiple historical standardization sample documents to obtain at least one matching paragraph. Using the paragraph processing layer in the document standardization processing model, the document structure of the initial professional document is standardized according to each matched paragraph to obtain a standardized professional document; The document standardization processing model is trained based on multiple initial sample professional documents and corresponding sample standard professional documents for each initial sample professional document; the document structure includes the paragraph-level content of the initial professional documents, and the standardization processing includes adjusting the paragraph-level content of the initial professional documents based on each matching paragraph.
2. The method according to claim 1, characterized in that, The paragraph processing layer includes a second segmentation layer, a second matching layer, and a sentence processing layer. The paragraph processing layer in the document standardization processing model standardizes the document structure of the initial professional document based on each matched paragraph, resulting in the standardized professional document, including: The second segmentation layer is used to perform sentence segmentation on each of the matched paragraphs to obtain multiple sentences to be matched. For each of the statements to be matched, the second matching layer is used to perform statement matching processing on the statements to be matched to obtain at least one matching statement; Using the statement processing layer, the document structure of the initial professional document is standardized according to each of the matching statements to obtain the standardized professional document.
3. The method according to any one of claims 1-2, characterized in that, The step of generating an initial professional document based on the initial text and the professional document generation model includes: The initial text is evaluated for similarity based on a preset query knowledge base, and the evaluation result is obtained. The query knowledge base is constructed based on multiple standard professional documents. If the evaluation result indicates that the initial text meets the conditions for generating professional documents, the initial professional document is generated based on the initial text and the professional document generation model.
4. The method according to claim 3, characterized in that, The query knowledge base includes a first vector corresponding to each of the standard professional documents. The step of performing a similarity assessment on the initial text based on the preset query knowledge base to obtain the assessment result includes: The initial text is vectorized to obtain a second vector corresponding to the initial text. Based on the second vector, a target vector similar to the second vector is determined from each of the first vectors; The evaluation result is determined based on the second vector and the target vector.
5. The method according to claim 4, characterized in that, Determining the evaluation result based on the second vector and the target vector includes: Calculate the similarity between the second vector and the target vector; If the similarity is less than a preset similarity threshold, then the evaluation result is determined to be that the initial text meets the professional document generation conditions; If the similarity is greater than or equal to the similarity threshold, then the evaluation result is determined to be that the initial text does not meet the conditions for generating professional documents.
6. The method according to claim 4, characterized in that, The query knowledge base also includes query paragraphs corresponding to each of the first vectors, and the generation of the initial professional document based on the initial text and the professional document generation model includes: Obtain the target query paragraph corresponding to the target vector from the query knowledge base; The initial professional document is generated based on the target query paragraph, the initial text, and the professional document generation model.
7. The method according to claim 4, characterized in that, The method further includes: Obtain the aforementioned standard professional documents and perform paragraph segmentation on each of the aforementioned standard professional documents to obtain multiple segmented paragraphs; Based on the segmented segments, multiple segments to be processed are obtained, and each segment to be processed is subjected to segment enhancement processing to obtain multiple query segments; Each query segment is vectorized to obtain the first vector corresponding to each query segment, and the correspondence between each query segment and each first vector is stored in the query knowledge base.
8. A professional document generation device, characterized in that, The device includes: The generation module is used to obtain initial text and generate initial professional documents based on the initial text and the professional document generation model; The processing module is used to perform paragraph segmentation on the initial professional document using the first slicing layer in the document standardization processing model to obtain multiple paragraphs to be matched; for each paragraph to be matched, the first matching layer in the document standardization processing model is used to perform paragraph matching on the paragraph to be matched based on multiple historical standard sample documents to obtain at least one matching paragraph; the paragraph processing layer in the document standardization processing model is used to standardize the document structure of the initial professional document according to each matching paragraph to obtain a standardized professional document. The document standardization processing model is trained based on multiple initial sample professional documents and corresponding sample standard professional documents for each initial sample professional document. The document structure includes the paragraph-level content of the initial professional documents, and the standardization processing includes adjusting the paragraph-level content of the initial professional documents based on each matching paragraph.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Sentence vector generation method and device, computer equipment and storage medium
CN114444471A
Automatic typesetting method and system for documents
CN116738934A
Document adjustment method and device, equipment and storage medium
CN117829116A
Bulletin document generation method, device and equipment
CN118485054A