Training method, device and equipment for generating document recognition model of PPT (Power Point)

By training the document recognition model and using a large model and reward strategy to generate PPTs, the problems of low efficiency and accuracy in generating PPTs in existing technologies are solved, and efficient and accurate PPT generation is achieved.

CN120805876APending Publication Date: 2025-10-17BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510846851.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The process of generating PPTs in existing technologies requires users to invest a lot of time and energy, and AI technology cannot quickly and accurately generate the PPTs that users expect, resulting in low generation efficiency and accuracy.

Method used

By obtaining pre-collected documents to be trained, using the large model and prompt word information to determine the recognition results, combining the preset reward strategy to train the document recognition model, output outline information and content blocks for generating PPT.

Benefits of technology

It achieves targeted and professional training of document recognition models, improves the efficiency and accuracy of PPT generation, and reduces the need for manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805876A_ABST
    Figure CN120805876A_ABST
Patent Text Reader

Abstract

The invention provides a training method, device and equipment for generating a PPT document recognition model, and relates to the field of artificial intelligence, in particular to the field of large models. According to the specific implementation scheme, a to-be-trained document collected in advance is obtained, and an identification result of the to-be-trained document is determined based on a preset large model and prompt word information; the identification result comprises outline information and content blocks, the outline information comprises a plurality of titles, the titles represent the content architecture of the document, and the content blocks represent the document content corresponding to the titles; according to the to-be-trained document and the recognition result of the to-be-trained document, training a to-be-trained model based on a preset reward strategy to obtain a trained document recognition model; the preset reward strategy is used for carrying out preset dimension matching on an identification result of the to-be-trained document and an output result of the to-be-trained document by the to-be-trained model, the document identification model is used for outputting outline information and content blocks corresponding to the document, and the outline information and the content blocks are used for generating the PPT.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of large models in the field of artificial intelligence, in particular to a training method, device and equipment of a document recognition model for generating PPT. BACKGROUND

[0002] PPT (PowerPoint) is an indispensable display tool in people's life and work, covering various application scenarios such as teaching plan design, work summary, and personal resume.

[0003] Currently, people usually need to create PPT according to the established document to meet the specific display requirements. However, this way requires users to invest a lot of time and effort in design and production, increasing the time and labor cost of creation, and the generation efficiency of PPT is low. SUMMARY

[0004] The present disclosure provides a training method, device and equipment of a document recognition model for generating PPT.

[0005] According to a first aspect of the present disclosure, a training method of a document recognition model for generating PPT is provided, comprising:

[0006] obtaining a pre-collected training document, determining the recognition result of the training document based on a pre-set large model and prompt word information, wherein the prompt word information is used to assist the large model in understanding the input document, the recognition result includes outline information and content blocks, the outline information includes at least one title, the title represents the content architecture of the document, and the content block represents the document content corresponding to the title;

[0007] According to the training document and the recognition result of the training document, a pre-set reward strategy is used to train the training model, and a trained document recognition model is obtained; wherein the pre-set reward strategy is used to match the recognition result of the training document and the output result of the training model to the training document in a pre-set dimension, the document recognition model is used to output the outline information and the content blocks corresponding to the document, and the outline information and the content blocks are used to generate PPT.

[0008] According to a second aspect of the present disclosure, a training device of a document recognition model for generating PPT is provided, comprising:

[0009] The document recognition unit is configured to obtain a pre-collected training document, determine a recognition result of the training document based on a preset large model and prompt word information, wherein the prompt word information is used to assist the large model in understanding an input document, and the recognition result includes outline information and content blocks, the outline information includes at least one title, the title represents a content architecture of the document, and the content blocks represent document contents corresponding to the title.

[0010] The model training unit is configured to train a training model based on a preset reward strategy according to the training document and the recognition result of the training document, and obtain a trained document recognition model, wherein the preset reward strategy is used to match the recognition result of the training document and an output result of the training model on the training document in a preset dimension, and the document recognition model is used to output outline information and content blocks of a document, and the outline information and the content blocks are used to generate a PPT.

[0011] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0012] at least one processor; and

[0013] a memory connected with the at least one processor in communication;

[0014] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of the present disclosure.

[0015] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to enable the computer to perform the method of the first aspect of the present disclosure.

[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method of the first aspect of the present disclosure.

[0017] According to the technology of the present disclosure, the document recognition model is trained in a targeted and professional manner, thereby improving the generation efficiency and accuracy of the PPT.

[0018] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0020] Figure 1 FIG. 4 is a flowchart of a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0021] Figure 2 FIG. 4 is a flowchart of a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0022] Figure 3 FIG. 5 is a schematic diagram of a model training process based on a preset dimension according to an embodiment of the disclosure;

[0023] Figure 4 FIG. 4 is a flowchart of a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0024] Figure 5 FIG. 4 is a flowchart of a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0025] Figure 6 FIG. 6 is a structural block diagram of a training device of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0026] Figure 7 FIG. 6 is a structural block diagram of a training device of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0027] Figure 8 FIG. 7 is a block diagram of an electronic device for implementing a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure;

[0028] Figure 9 FIG. 7 is a block diagram of an electronic device for implementing a training method of a document recognition model for generating a PPT according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0029] Exemplary embodiments of the disclosure are described herein with reference to the accompanying drawings to assist in a comprehensive understanding of the disclosure by those of ordinary skill in the art, and the disclosure should be considered as including all such modifications and alterations to the examples described herein. Accordingly, it should be understood that various modifications and changes can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Also, the descriptions of the known functions and configurations are omitted herein for clarity and conciseness.

[0030] The library platform is a diversified document resource supply station deeply rooted in professional fields, and is committed to fully meeting the diverse document needs of users in office collaboration, education and training, and other dimensions. On this platform, users can easily retrieve a variety of document resources covering lesson plan design, work summary, and personal resume, accurately matching the specific needs of individuals or teams. Typically, users will expect to convert the retrieved documents into PPT for easy viewing and presentation of the content.

[0031] Currently, there are mainly three schemes for generating PPT from documents: (1) users create PPT from documents to meet specific display needs, which requires users to invest a lot of time and effort in design and production, increasing the time and labor costs of creation; (2) search for similar PPTs through retrieval, which has the advantage of significantly saving user time and effort, however, search engines mainly provide results based on text matching, which often limits the ability to understand customized needs in depth, and users may need to invest additional time and effort to screen and modify after retrieving PPTs to ensure that these PPTs better meet their needs; (3) through AI technology, users can quickly generate PPTs from uploaded documents with one click, which is efficient, but requires high AI technology, and current AI technology cannot quickly and accurately generate PPTs as users expect.

[0032] The present disclosure provides a training method, device and equipment for a document recognition model for generating PPTs, which is applied to the large model field in the field of artificial intelligence to realize the targeted and professional training of the document recognition model, thereby improving the generation efficiency and accuracy of PPTs.

[0033] It should be noted that the model in this embodiment is not a model for a specific user and cannot reflect the personal information of a specific user. It should be noted that the data in this embodiment comes from a public data set.

[0034] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0035] To enable the reader to more deeply understand the implementation principles of the present disclosure, the following Figures 1-9 The embodiments are further refined.

[0036] Figure 1 A flowchart of a training method for a document recognition model for generating PPTs according to an embodiment of the present disclosure is provided. The method can be executed by a training device for a document recognition model for generating PPTs. As shown in Figure 1 the method includes the following steps:

[0037] In S101, a pre-collected training document is obtained, and a recognition result of the training document is determined based on a pre-set large model and prompt information. The prompt information is used to assist the large model in understanding the input document. The recognition result includes outline information and content blocks. The outline information includes at least one title, and the title represents the content architecture of the document. The content blocks represent the document content corresponding to the title.

[0038] For example, the training document can be a document containing text data. A plurality of documents containing different text content are pre-collected as training documents. For example, a plurality of papers can be collected as training documents.

[0039] The prompt information, i.e., prompt, is pre-written. The prompt information can be text data, which is used to assist the large model in understanding the input document. The large model is a pre-constructed large language model. The input data of the large model is a document, and the large model can recognize the input document. The output data is the recognition result of the document. That is, the training document and the prompt information can be input into the large model to output the recognition result. The recognition result can include outline information and content blocks. The outline information includes one or more titles. The title is the topic of the content introduced in the document. The document can include multiple titles. The content architecture of the document can be represented by the title. For example, it can include a first-level title, a second-level title, and a third-level title. The content blocks represent the document content corresponding to the title. The specific document content can or can not be included under the title. That is, the number of content blocks can be less than or equal to the number of titles. For example, a first-level title in the document is directly followed by a second-level title, and the second-level title is specifically introduced. The first-level title has no corresponding content, and the second-level title has a corresponding content block.

[0040] A specific example of the training document can be:

[0041] “Introduction of Benzodiazepines

[0042] I. Benzodiazepine drugs (BZDs)

[0043] Drug effects:

[0044] Sedation: By activating benzodiazepine receptors, it produces a sedative effect.

[0045] Anxiolytic: It can relieve anxiety symptoms and stabilize the patient's mood.

[0046] Muscle relaxation: It has a certain muscle relaxing effect and can be used to treat muscle tension or spasm.

[0047] Anticonvulsant: It can inhibit excessive excitation of the central nervous system and thus play an anticonvulsant role.

[0048] Drug classification:

[0049] According to the length of half-life, it can be divided into short-acting, medium-acting and long-acting.

[0050] Drug characteristics:

[0051] Oral absorption is good, and it is metabolized by the liver.

[0052] It can shorten the time to fall asleep, reduce the time and number of awakenings, and increase the total sleep time.

[0053] Safety and tolerability are good, but long-term use may produce drug dependence, rebound after stopping medication, and memory loss, etc.

[0054] The specific examples of cue information can be:

[0055] "**Task description**:\nExtract the outline of

text

[0056] \n\n**Task requirements**:\n

[0057] 1. The outline strictly follows the structure of

text

[0058] 2. When there is no clear title in

text

[0059] 3. The extracted second-level title or third-level title must be the same as the original text, and cannot be self-expressed; do not extract;

[0060] 4. When

text

[0061] 5. The first-level outline is "#", and there is only one; the second-level outline is "##"; the third-level outline is "###"

[0062] 6. Prohibit output of "-", fourth-level outline; \n

[0063] 7. Prohibit output of ordered numbers, such as "1, 2, 3", "1.1, 1.2, 1.3", "(1), (2), (3)", "①, ②, ③", "First chapter, second chapter, third chapter", "One, two, three", "(One), (Two), (Three)"\n

[0064] 8. The number of output second and third level outlines does not exceed 30 characters!!! \n\n

[0065] 9. When the end of the

text

[0066] 10. The

text

[0067] 11. When the last paragraph is a summary, you can add a secondary title to cover the content of the last paragraph.

[0068] **Correct output example**: \n\n

[0069] # Quarterly Human Cost Budget Report

[0070] ## Introduction

[0071] ## Last Quarter Human Cost Budget Execution

[0072] ### Overall Execution

[0073] ### Salary Execution

[0074] ### Welfare Execution

[0075]

text

[0076] {}

[0077] Among them, {} can be the specific content of the document.

[0078] Input the above example of the document to be trained and the prompt word information into the pre-set large model, and the output recognition result can be:

[0079] " # Introduction to Benzodiazepines

[0080] ## Benzodiazepine Drugs (BZDs)

[0081] ['Start string: Benzodiazepine drugs', 'End string: memory loss and other side effects.']

[0082] ### Drug Effects

[0083] ['Start string: Drug effects', 'End string: anticonvulsant effects.']

[0084] ### Drug Classification

[0085] ['Start string: Drug classification', 'End string: and long-acting classes.']

[0086] ### Drug Characteristics

[0087] ['Start string: Drug characteristics', 'End string: memory loss and other side effects.']"

[0088] Wherein, "#", "##", "###" are titles, and constitute outline information, and "[]" represents a content block. The content block can include all text content under the corresponding title, or can only contain a start string and an end string in all text content under the corresponding title. The start string can be the first few words of the text content under the title, and the end string can be the last few words. Reducing the data volume of the content block improves the efficiency of subsequent model training.

[0089] S102, according to the to-be-trained document and the recognition result of the to-be-trained document, training the to-be-trained model based on the preset reward strategy, obtaining the trained document recognition model; wherein, the preset reward strategy is used to match the recognition result of the to-be-trained document and the output result of the to-be-trained model on the to-be-trained document in a preset dimension, and the document recognition model is used to output the outline information and the content block corresponding to the document, and the outline information and the content block are used to generate the PPT.

[0090] Exemplarily, the to-be-trained model is constructed in advance, the to-be-trained model is used for document recognition, and the document recognition model is obtained after training, which can output the outline information and the content block corresponding to the document, that is, the to-be-trained model can process the to-be-trained document in the training process, and output the predicted outline information and the content block. In this embodiment, the model architecture of the to-be-trained model is not limited, for example, the QwQ-32B model can be used as the to-be-trained model.

[0091] The recognition result output by the large model is used as the label of model training, the to-be-trained document and the recognition result of the to-be-trained document are input into the to-be-trained model, and model training is performed by means of back propagation and the like, to obtain the document recognition model.

[0092] When the model is trained, the reward strategy is set in advance, and the reward strategy is used to match the recognition result of the to-be-trained document and the output result of the to-be-trained model on the to-be-trained document in a preset dimension. The recognition result is the outline information and the content block output by the large model, and the output result is the outline information and the content block output by the to-be-trained model. The reward strategy can include a plurality of matching rules in a preset dimension, for example, the outline information in the recognition result and the outline information in the output result can be matched in the number of words, the content block in the recognition result and the content block in the output result can be matched in the number, and the like. Through the matching of the reward strategy, the training situation of the to-be-trained model can be obtained, that is, whether the to-be-trained model is trained is determined.

[0093] The loss function of the model training can also be preset, and in this embodiment, the preset loss function is not specifically limited. The loss function can also be used to determine whether the to-be-trained model is trained. By combining the loss function and the reward strategy, the model training is not trained by the loss function alone, but by the RFT (Reinforcement Fine-Tuning) technology, thereby improving the training accuracy of the model.

[0094] After obtaining the document recognition model, the document for which a PPT needs to be generated can be input into the model to obtain the outline information and the content block of the document, i.e., the content architecture of the document and the specific content of each title under the content architecture. The outline information and the content block of the document are input into a preset PPT rendering tool to output the final PPT. In this embodiment, the preset PPT rendering tool is not specifically limited.

[0095] In the embodiment of the present disclosure, a pre-collected to-be-trained document is obtained and prompt word information is written, the to-be-trained document and the prompt word information are input into a preset large model to determine a recognition result of the to-be-trained document. The recognition result includes outline information and a content block, the outline information includes multiple titles in the document, and the content block represents the document content corresponding to the title. The recognition result of the to-be-trained document is taken as a label of model training, and a to-be-trained model is trained based on a preset reward strategy to obtain a trained document recognition model. The reward strategy can be used to match the recognition result of the to-be-trained document and the output result of the to-be-trained model in different dimensions, and the outline information and the content block output by the document recognition model can be used to automatically generate a PPT. By using the large model, training in combination with different models is realized, manual setting of labels is not required, and the efficiency of model training is improved. By setting the reward function, the to-be-trained model can be trained in a targeted manner by RFT, the to-be-trained model is improved in terms of the pertinence and professionalism in generating a PPT, and the generation efficiency and accuracy of the PPT are further improved.

[0096] Figure 2 A flowchart of a training method of a document recognition model for generating a PPT provided in an embodiment of the present disclosure is shown, which is an optional embodiment based on the above-described embodiment.

[0097] In this embodiment, according to the to-be-trained document and the recognition result of the to-be-trained document, the to-be-trained model is trained based on a preset reward strategy to obtain a trained document recognition model, including: inputting the to-be-trained document into the to-be-trained model to obtain an output result; wherein the output result represents the outline information and the content block recognized by the to-be-trained model from the to-be-trained document; according to the recognition result of the to-be-trained document and the output result, a reward score of the to-be-trained document is obtained based on a preset reward strategy; wherein the reward score represents the recognition accuracy of the to-be-trained model on the to-be-trained document; and the to-be-trained model is gradient trained according to the reward score and a preset loss function to obtain the trained document recognition model.

[0098] As shown in Figure 2 the method comprises the following steps:

[0099] S201, a pre-acquired to-be-trained document is obtained, and a recognition result of the to-be-trained document is determined based on a preset large model and prompt word information; wherein the prompt word information is used to assist the large model in understanding the input document, and the recognition result includes outline information and content blocks, the outline information includes at least one title, and the title represents the content architecture of the document, and the content blocks represent the document content corresponding to the title.

[0100] Exemplarily, this step can refer to the above step S101, and will not be described here.

[0101] S202, the to-be-trained document is input into the to-be-trained model to obtain an output result; wherein the output result represents the outline information and the content block recognized by the to-be-trained model from the to-be-trained document.

[0102] Exemplarily, the to-be-trained document is first input into the large model to obtain the recognition result output by the large model. Then the to-be-trained document is input into the to-be-trained model, and the to-be-trained model identifies the title and the content block corresponding to the title contained in the to-be-trained document to obtain the output result of the to-be-trained model. That is, the output result can represent the outline information and the content block recognized by the to-be-trained model from the to-be-trained document. The recognition result of the large model and the output result of the to-be-trained model can be the same or different.

[0103] S203, according to the recognition result of the to-be-trained document and the output result, a reward score of the to-be-trained document is obtained based on a preset reward strategy; wherein the reward score represents the recognition accuracy of the to-be-trained model on the to-be-trained document.

[0104] Exemplarily, based on the preset reward strategy, the recognition result of the to-be-trained document is compared with the output result to obtain a reward score corresponding to the to-be-trained document. Each to-be-trained document corresponds to a reward score, and the reward score represents the recognition accuracy of the to-be-trained model on the to-be-trained document. For example, the higher the consistency between the recognition result and the output result of the to-be-trained document, the higher the reward score.

[0105] The recognition result is the standard outline information and content block corresponding to the to-be-trained document. The recognition result is used as a label to determine whether the output result of the to-be-trained model is correct. The reward strategy can include a preset determination condition and a calculation method of the reward score. For example, if the title in the recognition result is completely consistent with the title in the output result, the reward score is 1 point, otherwise, the reward score is 0 point.

[0106] In this embodiment, different dimensions of reward strategies are preset, the model is trained through the reward strategies, the accuracy of the model training is improved, the model training process based on the RFT technology is realized, the subsequent document recognition accuracy is improved, and the generation efficiency and accuracy of the PPT are further improved.

[0107] In this embodiment, the preset reward strategy includes a matching rule in the preset dimension, and the matching rule is used to match the recognition result and the output result of the to-be-trained document. Based on the recognition result and the output result of the to-be-trained document, a reward score of the to-be-trained document is obtained based on the preset reward strategy, including: based on the recognition result and the output result of the to-be-trained document, a matching score of the to-be-trained document in the preset dimension is determined based on the matching rule in the preset dimension; wherein the matching score represents the consistency of the recognition result and the output result of the to-be-trained document in the preset dimension; and the reward score of the to-be-trained document is determined according to the matching scores of the to-be-trained document in all preset dimensions.

[0108] Specifically, different dimensions are preset, and the preset dimension refers to the angle of comparing and matching the recognition result and the output result. For example, the comparison and matching can be performed from the title structure, or the comparison and matching can be performed from the data amount of the content block. The preset reward strategy can include a matching rule in the preset dimension, and multiple dimensions are preset, each dimension corresponding to its own matching rule, and the matching rule is used to match the recognition result and the output result of the to-be-trained document.

[0109] For each preset dimension, the recognition result and the output result of the to-be-trained document are matched according to the matching rule under the preset dimension, to obtain a matching score of the to-be-trained document under the preset dimension. For example, the higher the consistency of the recognition result and the output result under the preset dimension, the higher the matching score. Each preset dimension corresponds to a matching score, that is, multiple matching scores can be obtained. The multiple matching scores are combined to obtain a reward score of the to-be-trained document. For example, the sum of the matching scores can be determined as the reward score.

[0110] The beneficial effect of such an arrangement is that each preset dimension can correspond to a matching score, comprehensive evaluation of the to-be-trained document is achieved, the multiple matching scores are combined to obtain the final reward score, and the training accuracy of the model is improved.

[0111] In this embodiment, the preset dimension is a structure dimension; the matching score of the to-be-trained document under the preset dimension is determined according to the recognition result and the output result of the to-be-trained document based on the matching rule under the preset dimension, including: the outline information in the recognition result of the to-be-trained document is determined as a first outline, and the outline information in the output result of the to-be-trained document is determined as a second outline; the titles in the first outline and the titles in the second outline are traversed respectively, to determine the number of matched titles in the first outline and the second outline; and the matching score of the to-be-trained document under the structure dimension is determined according to the number of matched titles in the first outline and the second outline.

[0112] Specifically, the preset dimension can include a structure dimension, which refers to the dimension of the content architecture of the document represented by the outline information, that is, the to-be-trained document is evaluated from the perspective of the titles in the outline information.

[0113] For the structural dimension, the preset matching rule can be to judge the title in the recognition result and the title in the output result, and add 1 point to the matching score for each consistent title. That is, the outline information in the recognition result of the to-be-trained document can be extracted first, and the outline information in the recognition result is determined as the first outline, and the outline information in the output result of the to-be-trained document is extracted, and the outline information in the output result is determined as the second outline. The titles in the first outline and the titles in the second outline are traversed respectively, and the title in the first outline traversed at present is compared with the title in the second outline in the same position sequence, that is, the first title in the first outline is compared with the first title in the second outline, the second title in the first outline is compared with the second title in the second outline, and so on until the last title is compared. According to the comparison result, the number of matched titles in the first outline and the second outline is determined, that is, the number of titles in the first outline and the second outline located in the same title sequence and having consistent title text is determined. According to the number of matched titles in the first outline and the second outline, the matching score of the to-be-trained document in the structural dimension is determined. For example, the number of titles can be taken as the matching score in the structural dimension.

[0114] Alternatively, if the title in the first outline is completely the same as the title in the second outline, the matching score is +1; if the title in the first outline is different from the title in the second outline, the matching score is 0. For example, there are five titles in the first outline and five titles in the second outline, and the titles in the first outline completely cover the titles in the second outline, that is, the titles are completely matched, so the matching score is +1.

[0115] The beneficial effect of such setting is that the output result of the to-be-trained model is rewarded in structure from the structural dimension, and the training accuracy and comprehensiveness of the model are improved.

[0116] In this embodiment, the preset dimension is the content dimension; according to the recognition result and the output result of the to-be-trained document, the matching score of the to-be-trained document in the preset dimension is determined based on the matching rule in the preset dimension, which includes: determining the title corresponding to the content block in the recognition result of the to-be-trained document as the first title, and determining the title corresponding to the content block in the output result of the to-be-trained document as the second title; matching the first title and the second title to obtain the matching score of the to-be-trained document in the content dimension.

[0117] Specifically, the preset dimension can include the content dimension, which refers to the dimension of the coverage degree of the content block to the document, that is, the to-be-trained document is evaluated from the perspective of the coverage of the content block to the document.

[0118] For the content dimension, the preset matching rule can be to determine whether the coverage of the content blocks in the recognition result on the document is consistent with the coverage of the content blocks in the output result on the document. That is, the content blocks can be extracted from the recognition result of the to-be-trained document, the titles corresponding to the content blocks are determined, the titles corresponding to the content blocks in the recognition result are determined as first titles, and the content blocks are extracted from the output result of the to-be-trained document, the titles corresponding to the content blocks are determined, and the titles corresponding to the content blocks in the output result are determined as second titles. According to the order of the content blocks in the recognition result, each first title is matched with each second title in turn, that is, it is determined whether there is a second title consistent with the first title in the output result. If there is, it means that the output result contains the content block corresponding to the first title, that is, the content of the first title is covered in the output result. If the first title and the second title are completely matched one by one, it means that the coverage of the output result and the recognition result on the document content is consistent, that is, the number of content blocks in the two results is consistent, and the titles represented by the content blocks are also consistent. According to the matching result of the first title and the second title, the matching score of the to-be-trained document in the content dimension can be determined. For example, if the first title and the second title are completely matched one by one, the matching score is +1; if there is a mismatch, the matching score is 0.

[0119] Alternatively, the number of matched titles between the first title and the second title is determined. Each matched title means that the output result covers the chapter content of the title. According to the number of matched titles, the matching score is determined. For example, there are five first titles and five second titles. Each matched title adds 1 point to the matching score, and the highest score is 5 points. That is, each chapter is 1 point, and if all five chapters are covered, the score is 1x5=5 points.

[0120] The beneficial effect of such setting is that the output result of the to-be-trained model is rewarded in terms of content from the content dimension, improving the training accuracy and comprehensiveness of the model.

[0121] In this embodiment, the preset dimension is the boundary dimension; according to the recognition result and the output result of the to-be-trained document, the matching score of the to-be-trained document in the preset dimension is determined based on the matching rule in the preset dimension, including: if the recognition result and the output result of the to-be-trained document contain the same title, the content block of the title is determined from the recognition result of the to-be-trained document as a first content block, and the content block of the title is determined from the output result of the to-be-trained document as a second content block; according to the characters in the first content block and the characters in the second content block, the matching score of the to-be-trained document in the boundary dimension is determined.

[0122] Specifically, the preset dimension can include a boundary dimension, which refers to a dimension of specific text content represented by the content block, i.e., the to-be-trained document is evaluated from the perspective of the text content in the content block.

[0123] For the boundary dimension, the preset matching rule can be to determine whether the text content in the content block in the recognition result and the text content in the content block in the output result are consistent. For each recognized content block, only the starting text content and the ending text content in the content block can be obtained, i.e., the text content at the boundary is used for determination.

[0124] In the recognition result and the output result, the content to be matched needs to belong to the same title, that is, it can be determined whether the recognition result and the output result of the to-be-trained document have the same title. If not, the matching score of the boundary dimension can be directly assigned as 0 or negative. If there is a same title, the content block of the title in the recognition result of the to-be-trained document can be determined as a first content block, and the content block of the title in the output result of the to-be-trained document can be determined as a second content block.

[0125] According to the characters in the first content block and the characters in the second content block, the matching score of the to-be-trained document in the boundary dimension is determined. The characters in the content block are the text content in the content block. The content block can include all the content under the corresponding title, or only the first few characters and the last few characters in the content under the corresponding title. For example, the first five characters in the content under the corresponding title can be taken as a starting string, and the last five characters in the content under the corresponding title can be taken as an ending string. Each content block only contains the corresponding starting string and ending string.

[0126] When the characters in the content block are matched, the characters in the first content block and the characters in the second content block can be matched in whole or in part to determine whether the first content block and the second content block are consistent. For example, the content block contains all the content in the corresponding title, and the overall similarity of the first content block and the second content block can be calculated. If the similarity exceeds a preset threshold, it is considered that the first content block and the second content block are consistent. For another example, the content block only contains the starting string and the ending string. The starting string in the first content block is matched with the starting string in the second content block, and the ending string in the first content block is matched with the ending string in the second content block. If the matching is consistent, it is considered that the first content block and the second content block are consistent.

[0127] Alternatively, if the starting string in the first content block and the starting string of the second content block are consistent, 0.5 points are added; if the ending string in the first content block and the ending string of the second content block are consistent, 0.5 points are added; if the starting string in the first content block and the starting string of the second content block are inconsistent, 0 points are added; if the ending string in the first content block and the ending string of the second content block are inconsistent, 0 points are added. For example, there are five first content blocks and five second content blocks, and the starting string and the ending string are matched respectively. If the starting strings are consistent, each chapter is added 0.5 points, and a total of five chapters are added 2.5 points; if the ending strings are consistent, each chapter is added 0.5 points, and a total of five chapters are also added 2.5 points.

[0128] The beneficial effect of such a setting is that, starting from the boundary dimension, the output result of the to-be-trained model is rewarded with the text content of the content block, improving the training accuracy and comprehensiveness of the model.

[0129] In this embodiment, the preset dimension is the redundancy dimension; according to the recognition result and the output result of the to-be-trained document, the matching score of the to-be-trained document in the preset dimension is determined based on the matching rule in the preset dimension, including: determining the similarity between the recognition result and the output result of the to-be-trained document; wherein the similarity represents the number of redundant characters between the recognition result and the output result; according to the similarity, the matching score of the to-be-trained document in the redundancy dimension is determined.

[0130] Specifically, the preset dimension can include the redundancy dimension, which refers to the dimension of whether the output result contains redundant information, i.e., from the perspective of the overall content of the output result, the to-be-trained document is evaluated.

[0131] For the redundancy dimension, the preset matching rule can be that, for the recognition result and the output result, the similarity between the recognition result and the output result of the to-be-trained document is calculated, and the similarity can represent the number of redundant characters between the recognition result and the output result, i.e., whether there are redundant characters in the output result is determined, and the more the number of redundant characters, the lower the similarity. In this embodiment, the calculation formula of the similarity is not specifically limited.

[0132] According to the calculated similarity, the matching score of the to-be-trained document in the redundancy dimension is determined. For example, the higher the similarity, the higher the matching score; the lower the similarity, the lower the matching score. If the matching score of the redundancy dimension is set to be the highest 0 points, then in the case of no redundancy, the matching score is 0 points; if there is redundancy, the matching score is negative.

[0133] The beneficial effect of such an arrangement is that, starting from the redundant dimension, the output result of the to-be-trained model is rewarded in the redundant condition, avoiding the model from outputting redundant information, and improving the training accuracy and comprehensiveness of the model.

[0134] In this embodiment, the preset dimensions can include a structure dimension, a content dimension, a boundary dimension, a redundant dimension, and the like. Figure 3 A schematic diagram of a model training process based on preset dimensions is shown in FIG. 3. Figure 3 In this embodiment, the structure reward is the matching score corresponding to the structure dimension, the content reward is the matching score corresponding to the content dimension, the boundary reward is the matching score corresponding to the boundary dimension, the redundant penalty is the matching score corresponding to the redundant dimension, and the comprehensive reward is the reward score.

[0135] S204, gradient training is performed on the to-be-trained model according to the reward score and a preset loss function, to obtain a trained document recognition model.

[0136] For example, after obtaining the reward score, the to-be-trained model can be trained in combination with the reward score and a preset loss function, to obtain a document recognition model. In this embodiment, the preset loss function is not specifically limited. The training process of the model can be gradient training through back propagation, and the training process of the model is not specifically limited in this embodiment.

[0137] In this embodiment, by combining the reward strategy and the loss function, reinforcement learning of the model is realized, that is, the training of the model is influenced through reinforcement learning and auxiliary loss function, so that the model finally generates an outline structure that is clear, the content is complete, and the boundary is accurate, improving the accuracy of model training, and further improving the generation accuracy of the PPT.

[0138] In the embodiments of the present disclosure, a pre-collected to-be-trained document is obtained and prompt word information is written, the to-be-trained document and the prompt word information are input into a preset large model, and the recognition result of the to-be-trained document is determined. The recognition result includes outline information and content blocks, the outline information includes multiple titles in the document, and the content blocks represent the document content corresponding to the titles. The recognition result of the to-be-trained document is used as a label for model training, a to-be-trained model is trained based on a preset reward strategy, and a trained document recognition model is obtained. The reward strategy can be used to match the recognition result of the to-be-trained document and the output result of the to-be-trained model in different dimensions, and the outline information and the content blocks output by the document recognition model can be used to automatically generate a PPT. By using the large model, training in combination with different models is realized, manual label setting is not required, and the efficiency of model training is improved. By setting the reward function, the to-be-trained model can be trained in RFT, the to-be-trained model is improved in terms of pertinence and professionalism in generating the PPT, and further the generation efficiency and accuracy of the PPT are improved.

[0139] Figure 4 A flowchart of a training method of a document recognition model for generating a PPT is provided for an embodiment of the present disclosure, which is an optional embodiment based on the above-mentioned embodiment.

[0140] In this embodiment, based on the preset large model and prompt word information, the recognition result of the to-be-trained document is determined, including: based on the preset large model and prompt word information, recognizing and extracting the title in the to-be-trained document to obtain the outline information of the to-be-trained document; and determining the content block corresponding to the title according to the position information of the title in the to-be-trained document in the outline information.

[0141] As shown in Figure 4 The method comprises the following steps:

[0142] S401, acquiring a pre-acquired to-be-trained document, and based on a preset large model and prompt word information, recognizing and extracting the title in the to-be-trained document to obtain the outline information of the to-be-trained document.

[0143] Exemplarily, the to-be-trained document and the prompt word information are input into the preset large model, and the to-be-trained document is processed by the large model. For example, the arrangement mode of the text in the to-be-trained document can be recognized, or the to-be-trained document can be subjected to semantic recognition, etc. By recognizing and feature extracting the title in the to-be-trained document, all the titles in the to-be-trained document can be determined, and according to the extracted titles, the outline information can be determined. For example, the text in a single line and less than a preset number of characters can be regarded as a title, and according to the order of the title in the to-be-trained document, the outline information can be obtained. That is, the outline information can include all the recognized titles, and the titles can be arranged in order from top to bottom.

[0144] S402, determining the content block corresponding to the title according to the position information of the title in the to-be-trained document in the outline information.

[0145] Exemplarily, after the title is obtained, the position information of each title in the to-be-trained document is determined, for example, the line number of the title in the to-be-trained document can be determined. According to the position information of the title in the to-be-trained document, the text content corresponding to the title can be extracted from the to-be-trained document, so as to obtain the content block corresponding to the title. For example, the content between two adjacent titles can be regarded as the content block of the previous title.

[0146] In this embodiment, the title is recognized first, and then the content block is determined according to the title, which can avoid missing in the recognition result, improve the determination accuracy of the recognition result, and facilitate subsequent model training according to the recognition result.

[0147] In this embodiment, according to the position information of the title in the outline information in the to-be-trained document, the content block corresponding to the title is determined, including: according to the position information of the title in the outline information in the to-be-trained document, determining the document content to which the title belongs; extracting a start string and an end string from the document content to which the title belongs; wherein the start string represents a string at a first preset position in the document content, and the end string represents a string at a second preset position in the document content; and determining the content block corresponding to the title according to the start string and the end string.

[0148] Specifically, first, according to the position information of the title in the outline information in the to-be-trained document, the entire document content to which the title belongs is determined. For example, if it is determined according to the position information of the title that the title is not the last title, the content between the title and the next title is determined as the entire document content to which the title belongs; if the title is the last title, all the content after the title is determined as the entire document content to which the title belongs.

[0149] After obtaining the entire document content to which the title belongs, the entire document content can be taken as the content block of the title, that is, the content block can include the entire document content to which the title belongs.

[0150] The start string and the end string can also be extracted from the entire document content to which the title belongs, the start string representing a string at a first preset position in the document content, and the end string representing a string at a second preset position in the document content. For example, the first preset position is the position of the first few characters of the document content, and the second preset position is the position of the last few characters of the document content. According to the start string and the end string, the content block corresponding to the title is determined, that is, the content block can only contain the start string and the end string in the document content. For example, the entire document content to which the title "drug effect" belongs is "drug effect: sedation: by activating benzodiazepine receptors, producing a sedative effect. Anti-anxiety: can relieve anxiety symptoms, stabilize the mood of patients. Relax muscles: has a certain relaxing effect on muscles, and can be used for the treatment of muscle tension or spasm. Anti-convulsion: can inhibit the excessive excitement of the central nervous system, thereby playing an anti-convulsion role." The first preset position is the position of the first four characters, and the second preset position is the position of the last six characters. The content block can be "['start string: drug effect', 'end string: anti-convulsion effect. ']" The content block can indicate the start string and the end string.

[0151] The beneficial effects of such setting are that the content block is the part of the content to which the title belongs in the document, and only needs to contain the start string and the end string, so as to cover the entire content, effectively reducing the data amount in the content block, avoiding data redundancy in the content block, and affecting the subsequent model training efficiency.

[0152] In the embodiment, the method further includes: if the recognition result of the to-be-trained document includes the preset character, performing screening processing on the to-be-recognized document and the recognition result of the to-be-recognized document.

[0153] Specifically, for each to-be-trained document, after obtaining the recognition result output by the large model, the recognition result can be screened to determine whether the recognition result of the to-be-trained document meets the preset requirement. For example, the preset requirement is that the recognition result cannot contain the preset character. After obtaining the recognition result, it can be determined whether the recognition result contains the preset character. If yes, the to-be-recognized document and the recognition result of the to-be-recognized document are screened out and do not participate in subsequent model training. If no, the to-be-recognized document and the recognition result of the to-be-recognized document are retained and subsequent model training is continued.

[0154] The beneficial effect of such a setting is that the data in the recognition result is screened according to rules, and the recognition result containing special characters is filtered out, so as to screen out data meeting the requirement, improve the accuracy of subsequent model training, and reduce training errors.

[0155] In the embodiment, a plurality of to-be-trained documents are pre-collected, and each to-be-trained document corresponds to a recognition result. The method further includes: extracting a preset number of recognition results from the recognition results of all to-be-trained documents; determining a qualified rate of the preset number of recognition results according to a preset result qualified condition, wherein the preset result qualified condition is used to evaluate whether the recognition result is qualified, and the qualified rate represents a proportion of qualified recognition results in the preset number of recognition results; and if the qualified rate is equal to or greater than a preset proportion threshold, performing training on the to-be-trained model based on the to-be-trained document and the recognition result of the to-be-trained document according to a preset reward strategy to obtain a trained document recognition model.

[0156] Specifically, there are a plurality of to-be-trained documents, and the large model can output a recognition result for each to-be-trained document to obtain a set of recognition results. The set of recognition results can include all recognition results or only the recognition results retained after being screened based on a preset requirement.

[0157] A preset number of recognition results are randomly extracted from the set, for example, 10 recognition results can be extracted. For the preset number of recognition results, the qualified rate of the recognition results is determined, that is, the number of qualified recognition results in the preset number of recognition results is determined. Whether the recognition result is qualified can be automatically determined based on a preset judgment rule, or can be manually determined. The number of qualified recognition results is divided by the preset number to obtain the qualified rate.

[0158] A ratio threshold is set in advance, and the pass rate is compared with the ratio threshold. If the pass rate is equal to or greater than the preset ratio threshold, it means that there is no problem with the recognition result output by the large model. According to the set of recognition results, the training model can continue to be trained based on the preset reward strategy to obtain a trained document recognition model.

[0159] The beneficial effect of this setting is that the recognition results output by the large model can be randomly checked and evaluated to determine the pass rate of the recognition results, thereby avoiding errors in the large model processing process that lead to model training failure and improving the accuracy of model training.

[0160] In this embodiment, the method further includes: if the qualified rate is less than a preset ratio threshold, adjusting the prompt word information, and determining the recognition result of the document to be trained based on the preset large model and the adjusted prompt word information.

[0161] Specifically, if the pass rate is determined to be less than a preset ratio threshold, it indicates that there are errors in the output of the large model, and the prompt word information needs to be adjusted. For example, the prompt word information can be deleted or rewritten to obtain adjusted prompt word information. The training document and the adjusted prompt word information are input into the preset large model, and the recognition result of the training document is re-determined until the pass rate is equal to or greater than the preset ratio threshold.

[0162] The beneficial effect of this setting is that the prompt word information can be flexibly adjusted according to actual needs, the processing accuracy of the large model can be improved, and then the training accuracy of the subsequent model can be improved.

[0163] S403. According to the document to be trained and the recognition result of the document to be trained, based on the preset reward strategy, the model to be trained is trained to obtain a trained document recognition model; wherein, the preset reward strategy is used to match the recognition result of the document to be trained and the output result of the model to be trained on the document to be trained in preset dimensions, and the document recognition model is used to output the outline information and content blocks corresponding to the document, and the outline information and content blocks are used to generate PPT.

[0164] For example, this step may refer to the above-mentioned step S102 and will not be described in detail.

[0165] In the embodiments of the present disclosure, a pre-collected to-be-trained document is obtained and prompt word information is written, the to-be-trained document and the prompt word information are input into a preset large model, and a recognition result of the to-be-trained document is determined. The recognition result includes outline information and content blocks, the outline information includes multiple titles in the document, and the content blocks represent document content corresponding to the titles. The recognition result of the to-be-trained document is taken as a label for model training, a to-be-trained model is trained based on a preset reward strategy, and a trained document recognition model is obtained. The reward strategy can be used to match the recognition result of the to-be-trained document and an output result of the to-be-trained model in different dimensions, and the outline information and the content blocks output by the document recognition model can be used to automatically generate a PPT. By using the large model, training in combination with different models is realized, manual label setting is not needed, and the efficiency of model training is improved. By setting a reward function, the to-be-trained model can be trained in a targeted manner through RFT, the to-be-trained model is improved in terms of the pertinence and professionalism in generating a PPT, and the generation efficiency and accuracy of the PPT are further improved.

[0166] Figure 5 A flowchart of a training method of a document recognition model for generating a PPT is provided for the embodiments of the present disclosure, which is an optional embodiment based on the above-mentioned embodiments.

[0167] In the embodiment, the method further includes: obtaining a pre-collected to-be-verified document; inputting the to-be-verified document into the document recognition model to obtain outline information and content blocks corresponding to the to-be-verified document; wherein the content blocks include a start string and an end string in document content corresponding to a title, the start string represents a string at a first preset position in the document content, and the end string represents a string at a second preset position in the document content; extracting document content corresponding to the title in the outline information from the to-be-verified document according to the outline information and the content blocks corresponding to the to-be-verified document; and obtaining a PPT corresponding to the to-be-verified document based on a preset PPT rendering tool according to the document content corresponding to each title in the to-be-verified document.

[0168] As shown in Figure 5 , the method includes the following steps:

[0169] S501, obtaining a pre-collected to-be-verified document.

[0170] Exemplarily, after obtaining the document recognition model, the document recognition model can be directly applied, that is, a document that needs to be converted into a PPT is input into the document recognition model to obtain the PPT. The document recognition model can also be verified first, and then applied after it is determined that the document recognition model has no problem.

[0171] One or more documents can be pre-collected as to-be-verified documents, and the document recognition model is verified by the to-be-verified documents. The content of the to-be-verified documents can be different from the to-be-trained documents.

[0172] S502, input the to-be-verified document into the document recognition model to obtain the outline information and the content block corresponding to the to-be-verified document; wherein the content block includes the start string and the end string in the document content corresponding to the title, the start string represents the string at the first preset position in the document content, and the end string represents the string at the second preset position in the document content.

[0173] Exemplarily, the to-be-verified document is input into the document recognition model, and the document recognition model is used to perform feature extraction processing on the to-be-verified document to obtain the outline information and the content block corresponding to the to-be-verified document. The outline information includes multiple titles in the to-be-verified document, and the content block can only include the start string and the end string in the document content corresponding to the title, the start string represents the string at the first preset position in the document content, and the end string represents the string at the second preset position in the document content.

[0174] S503, according to the outline information and the content block corresponding to the to-be-verified document, extracting the document content corresponding to the title in the outline information from the to-be-verified document.

[0175] Exemplarily, the content block only includes the start string and the end string, and does not include the entire document content to which the title belongs. Therefore, the content block needs to be supplemented according to the to-be-verified document, so that the content block can include the entire document content corresponding to the title.

[0176] For example, each title is determined from the outline information in sequence, and then the content block corresponding to the title is found. According to the content block, the start string and the end string in the document content to which the title belongs are determined. The start string and the end string are found in the to-be-verified document, and the start string is located before the end string. The document content between the start string and the end string in the to-be-verified document is filled into the content block to obtain a complete content block, that is, the document content to which each title belongs.

[0177] S504, according to the document content corresponding to each title in the to-be-verified document, obtaining the PPT corresponding to the to-be-verified document based on the preset PPT rendering tool.

[0178] Exemplarily, after obtaining the document content corresponding to each title, the document content corresponding to each title is input into the preset PPT rendering tool, and the PPT corresponding to the to-be-verified document is rendered by the PPT rendering tool.

[0179] The user views and evaluates the PPT corresponding to the to-be-verified document. If the user considers that the PPT corresponding to the to-be-verified document has no problem, the document recognition model can be put online for application of PPT generation. If the user considers that the PPT corresponding to the to-be-verified document has a problem, the model training can be performed again until the document recognition model can be put online.

[0180] In this embodiment, the document recognition model is verified, so that the effect of the document recognition model is guaranteed, the generation efficiency and accuracy of the PPT are improved, the actual PPT generation requirement is met, and the experience of the user in applying the model is improved.

[0181] In this embodiment, according to the document content corresponding to each title in the to-be-verified document, the PPT corresponding to the to-be-verified document is obtained based on a preset PPT rendering tool. Specifically, for each title in the to-be-verified document, the number of characters in the document content corresponding to the title is determined. According to the number of characters and the document content corresponding to the title, the target content corresponding to the title is determined. The target content corresponding to each title is input into the preset PPT rendering tool to obtain the PPT corresponding to the to-be-verified document.

[0182] Specifically, before the document content corresponding to each title is input into the PPT rendering tool, it can be determined whether the number of characters in the document content corresponding to the title meets a preset number requirement. For example, it can be determined whether the number of characters in the document content corresponding to the title is too small or too large.

[0183] According to the number of characters in the document content corresponding to the title and the specific document content corresponding to the title, the document content corresponding to the title can be adjusted, and the adjusted document content is determined as the target content. For example, if the number of characters in the document content corresponding to the title is too large, the characters in the document content can be appropriately reduced. If the number of characters in the document content corresponding to the title is too small, the characters in the document content can be appropriately increased.

[0184] The document content of each title corresponds to its own target content. If the document content is not adjusted, the document content is the target content. The target content corresponding to each title is input into the preset PPT rendering tool to obtain the PPT corresponding to the to-be-verified document.

[0185] The beneficial effects of such a setting are that the uniformity of the document content is improved by adjusting the number of characters in the document content, the generated PPT conforms to the specification, and the situation that the number of characters in the PPT exceeds the page or the number of characters is too small is avoided, thereby improving the browsing experience of the user on the PPT.

[0186] In this embodiment, the target content corresponding to the title is determined according to the number of characters and the document content corresponding to the title. Specifically, if the number of characters exceeds a preset number threshold, the document content corresponding to the title is adjusted in structure to obtain the target content corresponding to the title.

[0187] Specifically, a quantity threshold is preset, and the character quantity in the document content corresponding to the title is compared with the preset quantity threshold. If the character quantity is less than or equal to the quantity threshold, the document content does not need to be adjusted, and the document content is the target content.

[0188] If the character quantity exceeds the quantity threshold, the document content corresponding to the title can be adjusted in structure, for example, the document content can be divided into lines or paragraphs according to the period in the document content, and the adjusted target content corresponding to the title is obtained.

[0189] The beneficial effect of such setting is that if the data quantity in the complete content block exceeds the limit, the structured arrangement of the matched and filled content block can be performed, for example, each sentence in the content block can be regarded as a small paragraph, and then rendered into a PPT, so as to avoid uneven distribution of content in the PPT, improve the aesthetic degree of the PPT, and improve the browsing experience of the user.

[0190] In the embodiment, after the document recognition model is verified, the document recognition model can be put into operation. The process of formal application of the model is similar to the process of model verification. When a PPT needs to be generated according to a document, the document can be determined as a to-be-applied document. The to-be-applied document is input into the document recognition model to obtain the outline information and the content block corresponding to the to-be-applied document.

[0191] According to the outline information and the content block corresponding to the to-be-applied document, the document content corresponding to the title in the outline information is extracted from the to-be-applied document. According to the document content corresponding to each title in the to-be-applied document, the PPT corresponding to the to-be-applied document is obtained based on the preset PPT rendering tool.

[0192] When the PPT is generated, for each title in the to-be-applied document, the character quantity of the document content corresponding to the title can be determined. According to the character quantity and the document content corresponding to the title, the target content corresponding to the title is determined. The target content corresponding to each title is input into the preset PPT rendering tool to obtain the PPT corresponding to the to-be-verified document.

[0193] When the target content is determined, if the character quantity exceeds the preset quantity threshold, the document content corresponding to the title can be adjusted in structure to obtain the target content corresponding to the title.

[0194] In the embodiments of the present disclosure, a pre-collected training document is obtained and prompt word information is written, the training document and the prompt word information are input into a preset large model, and a recognition result of the training document is determined. The recognition result includes outline information and content blocks, the outline information includes multiple titles in the document, and the content blocks represent document content corresponding to the titles. The recognition result of the training document is taken as a label for model training, a training model is trained based on a preset reward strategy, and a trained document recognition model is obtained. The reward strategy can be used to match the recognition result of the training document and an output result of the training model on the training document in different dimensions, and the outline information and the content blocks output by the document recognition model can be used to automatically generate a PPT. By using the large model, training in combination with different models is realized, manual label setting is not required, and the efficiency of model training is improved. By setting a reward function, the training model can be trained in a targeted manner through RFT, the targeting and professionalism of the training model in generating a PPT in a vertical field are improved, and the generation efficiency and accuracy of the PPT are further improved.

[0195] Figure 6 A structural block diagram of a training device for a document recognition model for generating a PPT is provided for the embodiments of the present disclosure. For ease of illustration, only parts related to the embodiments of the present disclosure are shown. For details, refer to Figure 6 The training device 600 for the document recognition model for generating a PPT includes a document recognition unit 601 and a model training unit 602.

[0196] The document recognition unit 601 is configured to obtain a pre-collected training document, determine a recognition result of the training document based on a preset large model and prompt word information, wherein the prompt word information is used to assist the large model in understanding an input document, the recognition result includes outline information and content blocks, the outline information includes at least one title, the title represents a content architecture of the document, and the content blocks represent document content corresponding to the title.

[0197] The model training unit 602 is configured to train a training model based on a preset reward strategy according to the training document and the recognition result of the training document, and obtain a trained document recognition model, wherein the preset reward strategy is used to match the recognition result of the training document and an output result of the training model on the training document in a preset dimension, and the document recognition model is used to output outline information and content blocks corresponding to a document, and the outline information and the content blocks are used to generate a PPT.

[0198] Figure 7 A structural block diagram of a training device for a document recognition model for generating a PPT is provided for the embodiments of the present disclosure, as Figure 7As shown, the training device 700 of the document recognition model for generating the PPT comprises a document recognition unit 701 and a model training unit 702, wherein the model training unit 702 comprises a result output module 7021, a reward determination module 7022 and a model training module 7023.

[0199] The result output module 7021 is configured to input the to-be-trained document into the to-be-trained model to obtain an output result, wherein the output result represents outline information and content blocks recognized by the to-be-trained model from the to-be-trained document.

[0200] The reward determination module 7022 is configured to obtain a reward score of the to-be-trained document based on a preset reward strategy according to the recognition result and the output result of the to-be-trained document, wherein the reward score represents the recognition accuracy of the to-be-trained model on the to-be-trained document.

[0201] The model training module 7023 is configured to perform gradient training on the to-be-trained model according to the reward score and a preset loss function to obtain a trained document recognition model.

[0202] In one example, the preset reward strategy comprises a matching rule in a preset dimension, and the matching rule is used to match the recognition result and the output result of the to-be-trained document. The reward determination module 7022 comprises:

[0203] A dimension matching sub-module is configured to determine a matching score of the to-be-trained document in the preset dimension based on the matching rule in the preset dimension according to the recognition result and the output result of the to-be-trained document, wherein the matching score represents the consistency of the recognition result and the output result of the to-be-trained document in the preset dimension.

[0204] A score determination sub-module is configured to determine the reward score of the to-be-trained document according to the matching scores of the to-be-trained document in all preset dimensions.

[0205] In one example, the preset dimension is a structure dimension, and the dimension matching sub-module is specifically configured to:

[0206] determine the outline information in the recognition result of the to-be-trained document as a first outline, and determine the outline information in the output result of the to-be-trained document as a second outline;

[0207] traverse the titles in the first outline and the titles in the second outline respectively to determine the number of matched titles in the first outline and the second outline;

[0208] determine the matching score of the to-be-trained document in the structure dimension according to the number of matched titles in the first outline and the second outline.

[0209] In one example, the preset dimension is a content dimension; the dimension matching submodule is specifically configured to:

[0210] determine, from the recognition result of the to-be-trained document, a title corresponding to a content block as a first title, and determine, from the output result of the to-be-trained document, a title corresponding to a content block as a second title;

[0211] match the first title and the second title to obtain a matching score of the to-be-trained document in the content dimension.

[0212] In one example, the preset dimension is a boundary dimension; the dimension matching submodule is specifically configured to:

[0213] if there is a same title in the recognition result and the output result of the to-be-trained document, determine, from the recognition result of the to-be-trained document, a content block of the title as a first content block, and determine, from the output result of the to-be-trained document, a content block of the title as a second content block;

[0214] determine, according to characters in the first content block and characters in the second content block, a matching score of the to-be-trained document in the boundary dimension.

[0215] In one example, the preset dimension is a redundancy dimension; the dimension matching submodule is specifically configured to:

[0216] determine a similarity between the recognition result and the output result of the to-be-trained document; wherein the similarity represents a number of redundant characters between the recognition result and the output result;

[0217] determine, according to the similarity, a matching score of the to-be-trained document in the redundancy dimension.

[0218] In one example, the document recognition unit 701 comprises:

[0219] an outline determination module configured to determine and extract a title in the to-be-trained document based on a preset large model and prompt word information to obtain outline information of the to-be-trained document;

[0220] a content block determination module configured to determine a content block corresponding to the title according to position information of the title in the to-be-trained document in the outline information.

[0221] In one example, the content block determination module comprises:

[0222] a content determination submodule configured to determine document content to which the title belongs according to position information of the title in the to-be-trained document in the outline information.

[0223] a string extraction submodule configured to extract a start string and an end string from content of a document to which the title belongs, wherein the start string represents a string at a first preset position in the content of the document, and the end string represents a string at a second preset position in the content of the document;

[0224] a content block determination submodule configured to determine a content block corresponding to the title according to the start string and the end string.

[0225] In one example, the device further comprises:

[0226] a result screening unit configured to perform screening processing on the to-be-recognized document and the recognition result of the to-be-recognized document if the recognition result of the to-be-trained document includes a preset character.

[0227] In one example, a plurality of to-be-trained documents are pre-acquired, and each to-be-trained document corresponds to a recognition result; the device further comprises:

[0228] a result extraction unit configured to extract a preset number of recognition results from recognition results of all to-be-trained documents.

[0229] a qualified rate determination unit configured to determine a qualified rate of the preset number of recognition results according to a preset result qualified condition, wherein the preset result qualified condition is used to evaluate whether a recognition result is qualified, and the qualified rate represents a proportion of qualified recognition results in the preset number of recognition results.

[0230] a qualified rate comparison unit configured to perform, if the qualified rate is equal to or greater than a preset proportion threshold, the training of the to-be-trained model according to the to-be-trained document and the recognition result of the to-be-trained document based on a preset reward strategy to obtain a trained document recognition model.

[0231] In one example, the device further comprises:

[0232] a prompt word adjustment unit configured to adjust the prompt word information if the qualified rate is less than the preset proportion threshold, and determine the recognition result of the to-be-trained document based on a preset large model and the adjusted prompt word information.

[0233] In one example, the device further comprises:

[0234] a document acquisition unit configured to acquire a to-be-verified document pre-acquired.

[0235] a document input unit, configured to input the document to be verified into the document recognition model to obtain outline information and a content block corresponding to the document to be verified; wherein the content block includes a start string and an end string in the document content corresponding to the title, the start string representing a string located at a first preset position in the document content, and the end string representing a string located at a second preset position in the document content;

[0236] a content extraction unit, configured to extract, from the document to be verified, document content corresponding to a title in the outline information, based on the outline information and content blocks corresponding to the document to be verified;

[0237] The PPT generating unit is configured to obtain the PPT corresponding to the document to be verified based on the document content corresponding to each title in the document to be verified and based on a preset PPT rendering tool.

[0238] In one example, the PPT generation unit includes:

[0239] a quantity determination module, configured to determine, for each title in the document to be verified, the number of characters in the document content corresponding to the title;

[0240] a target determination module, configured to determine target content corresponding to the title based on the number of characters and the document content corresponding to the title;

[0241] The PPT rendering module is used to input the target content corresponding to each title into a preset PPT rendering tool to obtain the PPT corresponding to the document to be verified.

[0242] In one example, the target determination module includes:

[0243] The structure adjustment submodule is used to perform structural adjustment on the document content corresponding to the title if the number of characters exceeds a preset threshold value to obtain the target content corresponding to the title.

[0244] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device.

[0245] Figure 8 A structural block diagram of an electronic device provided in an embodiment of the present disclosure, such as Figure 8 As shown, the electronic device 800 includes: at least one processor 802; and a memory 801 communicatively connected to the at least one processor 802; wherein the memory stores instructions that can be executed by the at least one processor 802, and the instructions are executed by the at least one processor 802 to enable the at least one processor 802 to execute the training method of the document recognition model for generating PPT disclosed in the present invention.

[0246] The electronic device 800 further includes a receiver 803 and a transmitter 804. The receiver 803 is configured to receive instructions and data transmitted by other devices, and the transmitter 804 is configured to transmit instructions and data to external devices.

[0247] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0248] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which comprises a computer program stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to perform the scheme provided in any of the above embodiments.

[0249] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0250] As shown in Figure 9 The device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0251] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, and the like; an output unit 907, such as various types of displays, speakers, and the like; a storage unit 908, such as a magnetic disk, an optical disk, and the like; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0252] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the training method of the document recognition model for generating PPT. For example, in some embodiments, the training method of the document recognition model for generating PPT can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the training method of the document recognition model for generating PPT described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the training method of the document recognition model for generating PPT by any other appropriate means, such as by means of firmware.

[0253] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0254] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0255] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0256] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0257] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0258] The computer system can include clients and servers. This relationship can be. The servers are generally remote from the users and can be accessed via the Internet using a communication network. The relationship can be a client-server relationship over a communications network, and as such, the servers can be accessed by the clients using computer programs. The servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are mainframe products in the cloud computing service system, and solve the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The servers can also be servers of a distributed system, or servers combined with a blockchain.

[0259] It should be understood that the various forms of flow shown above can be reordered, steps added or removed. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which are not limited herein.

[0260] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A training method for a document recognition model for generating a PPT, comprising: Obtaining a pre-collected document to be trained, and determining a recognition result for the document to be trained based on a preset large model and prompt word information; wherein the prompt word information is used to assist the large model in understanding the input document, and the recognition result includes outline information and content blocks, wherein the outline information includes at least one title, the title representing the content structure of the document, and the content block representing the document content corresponding to the title; According to the document to be trained and the recognition result of the document to be trained, based on a preset reward strategy, the model to be trained is trained to obtain a trained document recognition model; wherein, the preset reward strategy is used to match the recognition result of the document to be trained and the output result of the model to be trained on the document to be trained in preset dimensions, and the document recognition model is used to output the outline information and content blocks corresponding to the document, and the outline information and content blocks are used to generate a PPT.

2. The method according to claim 1, wherein The method of training the model to be trained based on the document to be trained and the recognition result of the document to be trained and based on a preset reward strategy to obtain a trained document recognition model includes: Inputting the document to be trained into the model to be trained to obtain an output result; wherein the output result represents the outline information and content blocks recognized by the model to be trained from the document to be trained; According to the recognition result and output result of the document to be trained, based on a preset reward strategy, a reward score for the document to be trained is obtained; wherein the reward score represents the recognition accuracy of the model to be trained on the document to be trained; According to the reward score and the preset loss function, gradient training is performed on the model to be trained to obtain a trained document recognition model.

3. The method according to claim 2, wherein: The preset reward strategy includes matching rules under preset dimensions, and the matching rules are used to match the recognition results and output results of the document to be trained; The step of obtaining a reward score for the document to be trained based on the recognition result and the output result of the document to be trained and a preset reward strategy includes: Determining a matching score for the document to be trained under the preset dimension based on the recognition result and output result of the document to be trained and the matching rule under the preset dimension; wherein the matching score represents the degree of consistency between the recognition result and the output result of the document to be trained under the preset dimension; The reward score of the document to be trained is determined based on the matching scores of the document to be trained under all preset dimensions.

4. The method according to claim 3, wherein: The preset dimension is a structural dimension; determining the matching score of the document to be trained under the preset dimension based on the recognition result and the output result of the document to be trained and the matching rule under the preset dimension includes: Determining the outline information in the recognition result of the document to be trained as a first outline, and determining the outline information in the output result of the document to be trained as a second outline; Traversing the titles in the first outline and the titles in the second outline respectively, and determining the number of matching titles in the first outline and the second outline; The matching score of the document to be trained under the structural dimension is determined according to the number of matching titles in the first outline and the second outline.

5. The method according to claim 3 or 4, wherein: The preset dimension is a content dimension; determining the matching score of the document to be trained under the preset dimension based on the recognition result and the output result of the document to be trained and the matching rule under the preset dimension includes: Determining a title corresponding to a content block from the recognition result of the document to be trained as a first title, and determining a title corresponding to a content block from the output result of the document to be trained as a second title; The first title and the second title are matched to obtain a matching score of the document to be trained under the content dimension.

6. The method according to any one of claims 3 to 5, wherein: The preset dimension is a boundary dimension; determining the matching score of the document to be trained under the preset dimension based on the recognition result and the output result of the document to be trained and the matching rule under the preset dimension includes: If the same title exists in the recognition result and the output result of the document to be trained, a content block of the title is determined from the recognition result of the document to be trained as the first content block, and a content block of the title is determined from the output result of the document to be trained as the second content block; A matching score of the document to be trained in the boundary dimension is determined according to the characters in the first content block and the characters in the second content block.

7. The method according to any one of claims 3 to 6, wherein The preset dimension is a redundant dimension; determining the matching score of the document to be trained under the preset dimension based on the recognition result and the output result of the document to be trained and the matching rule under the preset dimension includes: Determining the similarity between the recognition result and the output result of the document to be trained; wherein the similarity represents the number of redundant characters between the recognition result and the output result; According to the similarity, a matching score of the document to be trained under the redundant dimension is determined.

8. The method according to any one of claims 1 to 7, wherein The step of determining the recognition result of the document to be trained based on the preset large model and prompt word information includes: Based on the preset large model and prompt word information, the title in the document to be trained is identified and extracted to obtain the outline information of the document to be trained; The content block corresponding to the title is determined according to the position information of the title in the outline information in the document to be trained.

9. The method according to claim 8, wherein The step of determining the content block corresponding to the title according to the position information of the title in the outline information in the document to be trained includes: Determining the document content to which the title belongs based on the position information of the title in the outline information in the document to be trained; Extracting a starting character string and an ending character string from the document content to which the title belongs; wherein the starting character string represents a character string located at a first preset position in the document content, and the ending character string represents a character string located at a second preset position in the document content; The content block corresponding to the title is determined according to the starting character string and the ending character string.

10. The method according to any one of claims 1 to 9, further comprising: If the recognition result of the document to be trained includes preset characters, the document to be recognized and the recognition result of the document to be recognized are screened out.

11. The method according to any one of claims 1 to 10, wherein a plurality of documents to be trained are collected in advance, each document to be trained corresponding to a recognition result; the method further comprising: Extracting a preset number of recognition results from all the recognition results of the documents to be trained; Determining a pass rate of the preset number of recognition results according to a preset result pass condition; wherein the preset result pass condition is used to evaluate whether the recognition result is qualified, and the pass rate represents the proportion of qualified recognition results in the preset number of recognition results; If the qualified rate is equal to or greater than the preset ratio threshold, the training model is trained based on the document to be trained and the recognition result of the document to be trained, based on the preset reward strategy, to obtain a trained document recognition model.

12. The method according to claim 11, further comprising: If the qualified rate is less than a preset ratio threshold, the prompt word information is adjusted, and the recognition result of the document to be trained is determined based on the preset large model and the adjusted prompt word information.

13. The method according to any one of claims 1 to 12, further comprising: Obtain pre-collected documents to be verified; Inputting the document to be verified into the document recognition model to obtain outline information and a content block corresponding to the document to be verified; wherein the content block includes a start string and an end string in the document content corresponding to the title, the start string representing a string located at a first preset position in the document content, and the end string representing a string located at a second preset position in the document content; Extracting the document content corresponding to the title in the outline information from the document to be verified according to the outline information and content blocks corresponding to the document to be verified; According to the document content corresponding to each title in the document to be verified, based on a preset PPT rendering tool, a PPT corresponding to the document to be verified is obtained.

14. The method according to claim 13, wherein The step of obtaining a PPT corresponding to the document to be verified based on the document contents corresponding to each title in the document to be verified and using a preset PPT rendering tool includes: For each title in the document to be verified, determining the number of characters in the document content corresponding to the title; Determining target content corresponding to the title based on the number of characters and the document content corresponding to the title; The target content corresponding to each title is input into a preset PPT rendering tool to obtain the PPT corresponding to the document to be verified.

15. The method according to claim 14, wherein The determining the target content corresponding to the title according to the number of characters and the document content corresponding to the title includes: If the number of characters exceeds a preset threshold, structural adjustment is performed on the document content corresponding to the title to obtain the target content corresponding to the title.

16. A training device for a document recognition model for generating a PPT, comprising: A document recognition unit is configured to obtain a pre-collected document to be trained and determine a recognition result for the document to be trained based on a preset large model and prompt word information; wherein the prompt word information is used to assist the large model in understanding the input document, and the recognition result includes outline information and content blocks, wherein the outline information includes at least one title, wherein the title represents the content structure of the document, and the content block represents the document content corresponding to the title; The model training unit is used to train the model to be trained based on the document to be trained and the recognition result of the document to be trained, based on a preset reward strategy, to obtain a trained document recognition model; wherein the preset reward strategy is used to match the recognition result of the document to be trained and the output result of the document to be trained by the model to be trained in preset dimensions, and the document recognition model is used to output the outline information and content blocks corresponding to the document, and the outline information and content blocks are used to generate a PPT.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.

19. A computer program product, wherein The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 15 when executed by a processor.