PPT and speech draft generation method based on artificial intelligence

By analyzing the user's input text and PPT outline, adjusting the PPT content, and combining PPT to generate corresponding speeches, the problem of poor speech results caused by simple PPT provided by users is solved, and PPT generation with higher quality and occasion-compliant PPT, as well as corresponding speeches are achieved.

CN119988653AInactive Publication Date: 2025-05-13HEFEI MADAO INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510070735.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, although the method of generating speeches based on PPT can correspond to each other, if the PPT provided by the user is too simple during the speech process, it may lead to poor speech effects and affect the performance of the speaker.

Method used

By obtaining the user's input text, analyzing the outline of the PPT, generating the initial version of the PPT, and adjusting the initial version of the PPT based on the user's condition information to ensure that the final generated PPT meets the requirements of the speech occasion. At the same time, corresponding speeches are generated based on the final version of PPT and the user's input text.

Benefits of technology

The quality of the generated PPT is improved, making it more in line with practical application scenarios, and enhancing the effect of the speech; at the same time, it ensures the synchronization between the speech and PPT, and improves the visibility and attractiveness of the speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988653A_ABST
    Figure CN119988653A_ABST
Patent Text Reader

Abstract

The invention discloses a PPT and speech draft generation method based on artificial intelligence, relates to the technical field of artificial intelligence, and solves the technical problems that in the speech process of a speaker, if the PPT provided by a user is too simple, the speech effect cannot reach the expectation, and the speech of the speaker is affected. The method comprises the steps of obtaining an input text of a user, and determining a speech type and a PPT type according to the input text of the user; determining a template of the PPT according to the type of the PPT; analyzing the outline of the PPT according to the input text of the user; generating a first edition PPT based on the outline of the PPT; sending the first edition PPT to the user for confirmation, and obtaining condition information of the user; adjusting the initial version PPT according to the condition information to obtain a final version PPT; generating a speech draft of the user according to the final version PPT and the input text of the user; the PPT quality can be improved; and the uniformity of the PPT and the speech is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and relates to PPT and speech script generation technology, and specifically is a PPT and speech script generation method based on artificial intelligence. Background Art

[0002] PPT plays a vital role in many fields such as modern business, education, and academic reports. PPT presents complex information and data in an intuitive and concise way through text, pictures, animations and other forms, making it easier for the audience to understand and receive. Speech manuscripts are an important tool for speakers to organize their thoughts and content. By writing speech manuscripts, speakers can grasp the theme and key points of the speech more clearly, ensuring that the content of the speech is clear and logically rigorous. In the actual speech process, PPT and speech manuscripts are usually used in combination. PPT is responsible for presenting visual information and multimedia elements, while speech manuscripts are responsible for providing detailed text descriptions and explanations. The combination of the two can make the speech more vivid, interesting, and easy to understand. PPT and speech manuscripts complement and promote each other; together they constitute an important part of an efficient speech.

[0003] The prior art (invention patent application with publication number CN118964649A) discloses a method for generating a manuscript based on artificial intelligence, which includes: obtaining a PPT file; performing title recognition on the PPT file based on a title recognition model to obtain title data; performing paragraph extraction on the PPT file based on a paragraph extraction model to obtain paragraph text data; performing image data extraction on the PPT file based on a processing library to obtain picture data and chart data; performing format conversion processing on the picture data and chart data to obtain a first image file and a second image file; determining a placeholder matching the data to be filled in from a speech template; filling the data to be filled into the placeholder matching the fill data in the speech template to obtain a target speech; in the prior art, a speech is generated based on PPT, which can make the PPT and the speech correspond to each other; however, when the speaker is giving a speech, not only does the speech need to correspond to the PPT, but the PPT also needs to be suitable for the occasion of the speech. If the PPT provided by the user is too simple, it may result in the speech effect not meeting expectations, affecting the speaker's speech.

[0004] The present invention provides a PPT and speech script generation method based on artificial intelligence to solve the above technical problems. Summary of the invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a PPT and speech script generation method based on artificial intelligence, which is used to solve the problem of generating a speech script based on PPT in the prior art, and can make the PPT and the speech script correspond to each other; however, when the speaker is giving a speech, not only the speech script needs to correspond to the PPT, but also the PPT needs to be suitable for the occasion of the speech. If the PPT provided by the user is too simple, it may result in the effect of the speech not meeting the expectations, affecting the technical problem of the speaker's speech.

[0006] To achieve the above object, the first aspect of the present invention provides a method for generating PPT and speech scripts based on artificial intelligence, comprising:

[0007] Step S1: obtaining a user's input text, determining a speech type and a PPT type according to the user's input text; and determining a PPT template according to the PPT type;

[0008] Step S2: Analyze the outline of the PPT according to the user's input text; generate a preliminary version of the PPT based on the outline of the PPT;

[0009] Step S3: Send the first version of the PPT to the user for confirmation, and obtain the user's condition information; adjust the first version of the PPT according to the condition information to obtain the final version of the PPT;

[0010] Step S4: Generate the user's speech draft based on the final version PPT and the user's input text.

[0011] Preferably, determining the speech type and PPT type according to the user's input text includes:

[0012] Retrieve the user's input text; extract keywords from the input text to obtain key features; obtain several speech topics; calculate the correlation between the key features and several speech topics;

[0013] Sort the relevance in descending order to obtain a speech relevance ranking table; determine the speech topic according to the speech relevance ranking table; determine the speech type and PPT type according to the speech topic.

[0014] The present invention extracts keywords from a user's input text to obtain key features; calculates the correlation between the key features and a number of speech topics, and screens out the user's speech type and PPT type according to the correlation; and can determine the speech type and PPT template according to different speech topics, so that the generated PPT conforms to the actual application scenario, which is beneficial to improving the quality of the PPT.

[0015] Preferably, extracting keywords from the input text to obtain key features includes:

[0016] Retrieving input text, preprocessing the input text to obtain a number of words; after marking the parts of speech of the words, removing stop words to obtain a number of candidate keywords;

[0017] Take the candidate keywords as nodes and construct a directed weighted graph G = (V, E); use the co-occurrence relationship to construct the edge between any two points; where V is the node set consisting of the candidate keywords; E is the edge set, representing the relationship between the nodes;

[0018] Set the weight of each node in the directed weighted graph by the formula Iterate the propagation of node weights to obtain the final score of the node; where d represents the damping factor, M(i) represents the set of all other nodes pointing to node i; L(j) represents the number of other nodes that node j can reach; w j represents the current score of node j;

[0019] The nodes are sorted according to their final scores to obtain an important sorting table; candidate keywords corresponding to the first n nodes in the important sorting table are selected as keywords of the input text; the keywords are integrated to obtain key features of the input text.

[0020] It should be noted that M(i) represents the set of all other nodes pointing to node i, where pointing means that there is some form of association between two words; the number of iterations is the number of times the formula converges, and the condition for convergence is that the change between two consecutive rounds is less than the preset threshold; the number of selected keywords is artificially set based on actual experience.

[0021] The present invention preprocesses an input text to obtain a number of words, marks the parts of speech of the words, removes stop words to obtain candidate keywords, analyzes final scores of the candidate keywords through a formula, and screens n candidate keywords as keywords of the input text according to the final scores; the keywords in the input text can be accurately screened out; and a foundation is laid for subsequent analysis of user speech topics.

[0022] Preferably, the calculation of the correlation between the key features and several speech topics includes:

[0023] Retrieve keywords from key features to obtain several speech topics; use word embedding models to analyze feature vectors of keywords and speech topics; where word embedding models include Word2Vec or Glove;

[0024] By formula Calculate the cosine similarity between the keyword and the speech topic; where A·B represents the dot product of the keyword’s feature vector and the speech topic’s feature vector; ‖A‖ represents the modulus of the keyword’s feature vector; ‖B‖ represents the modulus of the speech topic’s feature vector;

[0025] The correlation between the key features and the speech topic is calculated using a weighted average method according to the set keyword weights; wherein the keyword weights are determined according to the number of keywords in the key features.

[0026] The present invention analyzes the feature vectors of the keywords and speech topics in the key features, calculates the cosine similarity between the keywords and the speech topics; and uses the weighted average method to calculate the correlation between the two; the speech topic of the input text can be analyzed according to the correlation between the key features of the input text and the speech topic, laying a foundation for the subsequent determination of the PPT template.

[0027] Preferably, determining the template of the PPT according to the type of the PPT includes:

[0028] Retrieve the type of PPT; select several PPT templates of corresponding types from the PPT type template library according to the type of PPT;

[0029] Obtain the user's historical PPT, and integrate the user's historical PPT and PPT template into a style input sequence; call the template selection model, input the style input sequence into the template selection model, and obtain the corresponding PPT template; wherein the template selection model is built based on an artificial intelligence model.

[0030] The present invention selects several PPT templates of corresponding types from a type template library according to the type of PPT, analyzes the user's historical PPTs, and uses an artificial intelligence model to select a PPT template that meets the user's style from several PPT templates; the user's style can be considered while selecting the PPT template, which is conducive to improving the quality and uniqueness of the PPT.

[0031] It should be noted that the PPT type template library is composed of a number of PPT templates collected by the user and stored in a storage medium.

[0032] Preferably, the template selection model is constructed based on an artificial intelligence model, including:

[0033] Acquire a standard data set; wherein the standard data set includes standard input data consistent with the content attributes of the style input sequence, and standard output data consistent with the content attributes of the PPT template;

[0034] Divide the standard data set into a training set, a validation set, and a test set; use the training set to train the artificial intelligence model; use the validation set to adjust the internal parameters of the artificial intelligence model; use the test set to test the trained artificial intelligence model and obtain test indicators;

[0035] Get the indicator threshold and determine whether the test indicator is greater than the indicator threshold; if yes, mark the trained artificial intelligence model as the template selection model; if no, retrain the artificial intelligence model; the artificial intelligence model includes a convolutional neural network model or an Inception series model.

[0036] It should be noted that the test indicators include accuracy, recall rate, F1 score and stability; the indicator thresholds are set by expert assessment; the proportion of the standard data set divided into training set, validation set and test set is obtained based on experiments. When the trained artificial intelligence model does not meet the requirements, the proportion of the standard data set is redivided and the proportion of the training set is appropriately increased.

[0037] Preferably, analyzing the outline of the PPT according to the user's input text includes:

[0038] Retrieving the user's input text and key features; marking the key features in the input text; analyzing the paragraphs of the input text according to the marking results to obtain paragraph features;

[0039] Merge paragraphs with the same paragraph features to obtain the paragraph topics of the input text; based on the paragraph topics of the input text, use the large language model to generate the outline of the PPT.

[0040] The present invention marks the key features of the input text in the input text, analyzes the paragraphs of the input text according to the marking results to obtain paragraph features; analyzes the paragraph themes of the text according to the paragraph features, and generates a PPT outline using a large language model; and can divide the main body of the input text, laying a foundation for the subsequent generation of PPT.

[0041] Preferably, the step of generating a preliminary version of the PPT based on the PPT outline includes:

[0042] Retrieve the outline of the PPT; determine the title of the PPT according to the outline of the PPT; generate the number of pages of the PPT according to the title of the PPT;

[0043] Retrieve the input text corresponding to the PPT title name; enter the input text into the position corresponding to the title name, integrate all the pages of PTT in the order of the title names in the input text, and obtain the first version of the PPT.

[0044] The present invention determines the title name of the PPT according to the outline of the PPT, generates the number of pages of the PPT according to the title name of the PPT, inputs the input text into the corresponding position, integrates the number of pages of the PPT, and obtains the first version of the PPT; the title name of the PPT and the text can be made to correspond to each other, which is beneficial to improving the practicability of the PPT.

[0045] Preferably, adjusting the initial version of the PPT according to the condition information to obtain the final version of the PPT includes:

[0046] Retrieve the user's conditional information; obtain the basic information in the first version of the PPT; the conditional information includes the number of pictures, the number of animations, and the number of PPT pages; the data type of the basic information is consistent with the conditional information;

[0047] Determine whether the basic information in the preliminary version of the PPT is greater than the conditional information; if so, mark the preliminary version of the PPT as the final version of the PPT; if not, adjust the preliminary version of the PPT until it meets the conditions.

[0048] Preferably, the step of generating the user's speech draft according to the final version PPT and the user's input text includes:

[0049] Retrieve the final version of the PPT and the user's input text; extract image information from the final version of the PPT, generate explanatory text based on the image information; and insert the explanatory text into the position corresponding to the input text;

[0050] Use artificial intelligence tools to check and correct the grammar and spelling in the input text, and expand the input text to obtain the corresponding speech; compare and match the speech with each page of the PPT to obtain a matching result; annotate the speech according to the matching result; among them, artificial intelligence tools include chat assistants or ProWritingAid.

[0051] The present invention generates explanatory text according to the image in the final version of PPT, inserts the explanatory text into the position corresponding to the input text, uses artificial intelligence tools to correct the grammar and spelling of the input text, and expands the input text to obtain the corresponding speech manuscript; the generated speech manuscript can correspond to the PPT, which is conducive to maintaining the synchronization of the speech manuscript and the PPT.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1. The present invention extracts keywords from the user's input text to obtain key features; calculates the correlation between the key features and several speech topics, and screens out the user's speech type and PPT type according to the correlation; can determine the speech type and PPT template according to different speech topics, so that the generated PPT conforms to the actual application scenario, which is beneficial to improving the quality of PPT; pre-processes the input text to obtain several words, and marks the parts of speech of several words, removes stop words to obtain candidate keywords, analyzes the final scores of several candidate keywords through formulas, and screens n candidate keywords as keywords of the input text according to the final scores; analyzes the feature vectors of the keywords and speech topics in the key features, calculates the cosine similarity between the keywords and the speech topics; and calculates the correlation between the two using a weighted average method; can analyze the speech topic of the input text according to the correlation between the key features of the input text and the speech topic, laying a foundation for the subsequent determination of the PPT template; selects several PPT templates of the corresponding type from the type template library according to the type of PPT, analyzes the user's historical PPT, and selects a PPT template that conforms to the user's style from several PPT templates using an artificial intelligence model; can consider the user's style while selecting the PPT template, which is beneficial to improving the quality and uniqueness of the PPT.

[0054] 2. The present invention marks the key features of the input text in the input text, analyzes the paragraphs of the input text according to the marking results, and obtains the paragraph features; analyzes the paragraph themes of the text according to the paragraph features, and generates a PPT outline using a large language model; can divide the main body of the input text, laying a foundation for the subsequent generation of PPT; determines the title of the PPT according to the outline of the PPT, generates the number of pages of the PPT according to the title of the PPT, and inputs the input text into the corresponding position, integrates the number of pages of the PPT, and obtains the first version of the PPT; can make the title of the PPT correspond to the text, which is beneficial to improving the practicality of the PPT; generates explanatory text according to the image in the final version of the PPT, inserts the explanatory text into the corresponding position of the input text, and uses artificial intelligence tools to correct the grammar and spelling of the input text, and expands the input text to obtain the corresponding speech manuscript; can make the generated speech manuscript correspond to the PPT, which is beneficial to maintaining the synchronization of the speech manuscript and the PPT. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 It is a schematic diagram of the overall steps of the present invention;

[0057] Figure 2 A schematic diagram of the steps of determining the PPT template of the present invention;

[0058] Figure 3 A schematic diagram of the steps of generating PPT and speech scripts in the present invention. DETAILED DESCRIPTION

[0059] The technical scheme of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] See also Figure 1 The first aspect of the present invention provides a method for generating PPT and speech scripts based on artificial intelligence, comprising:

[0061] Step S1: obtaining a user's input text, determining a speech type and a PPT type according to the user's input text; and determining a PPT template according to the PPT type;

[0062] Step S2: Analyze the outline of the PPT according to the user's input text; generate a preliminary version of the PPT based on the outline of the PPT;

[0063] Step S3: Send the first version of the PPT to the user for confirmation, and obtain the user's condition information; adjust the first version of the PPT according to the condition information to obtain the final version of the PPT;

[0064] Step S4: Generate the user's speech draft based on the final version PPT and the user's input text.

[0065] See also Figure 2 , obtain the user's input text, pre-process the input text to obtain several words; after marking the parts of speech of several words, remove stop words to obtain several candidate keywords; take the candidate keywords as nodes, build a directed weighted graph G = (V, E); use the co-occurrence relationship to construct the edge between any two points; where V is a node set consisting of candidate keywords; E is an edge set, representing the relationship between nodes; set the weight of each node in the directed weighted graph, through the formula Iterate the propagation of node weights to obtain the final score of the node; where d represents the damping factor, M(i) represents the set of all other nodes pointing to node i; L(j) represents the number of other nodes that node j can reach; w jRepresents the current score of node j; sorts the nodes according to their final scores to obtain an important sorting table; selects candidate keywords corresponding to the first n nodes in the important sorting table as keywords of the input text; integrates the keywords to obtain the key features of the input text.

[0066] It should be noted that d is a damping factor, which represents the probability of transferring from one node to another during the iteration process. d is a real number between 0 and 1, usually set to 0.85; in specific cases, it can be adjusted according to actual conditions.

[0067] For example: suppose that the input text T of user A is obtained, the input text is preprocessed to obtain several words, the parts of speech of several words are marked, and several candidate keywords are obtained after removing stop words, the final scores of the candidate keywords are analyzed and calculated, and they are sorted, the first 4 candidate keywords are selected as the keywords of the input text, and these 4 keywords are integrated into the key features of the input text.

[0068] Get several speech topics; use word embedding models to analyze the feature vectors of keywords and speech topics; the word embedding models include Word2Vec or Glove; through the formula Calculate the cosine similarity between the keyword and the speech topic; where A·B represents the dot product of the keyword's feature vector and the speech topic's feature vector; ‖A‖ represents the modulus of the keyword's feature vector; ‖B‖ represents the modulus of the speech topic's feature vector; calculate the correlation between the key features and the speech topic using the weighted average method based on the set keyword weights; where the keyword weights are determined based on the number of keywords in the key features; sort the correlations in descending order to obtain a speech-related ranking table; determine the speech topic based on the speech-related ranking table; determine the speech type and PPT type based on the speech topic.

[0069] For example: suppose the speech topic of user A's input text T is determined, and the word embedding model is used to analyze the feature vectors of 5 keywords and the speech topic; the cosine similarity between the 4 keywords and the speech topic feature vectors is analyzed and calculated; according to the number of keywords, the keyword weights of the 4 keywords are set from high to low as 0.4; 0.3; 0.2; 0.1; the correlation between the key features of the input text and the speech topic is calculated based on the weighted average of the keyword weights, and the speech type is obtained as activity introduction, and the PPT type is activity introduction.

[0070] According to the type of PPT, several PPT templates of corresponding types are selected from the PPT type template library; the user's historical PPT is obtained, and the user's historical PPT and PPT template are integrated into a style input sequence; the template selection model is called, and the style input sequence is input into the template selection model to obtain the corresponding PPT template.

[0071] For example: suppose that the PPT template is analyzed and determined based on the historical PPT of user A; the PPT of activity introduction type and the historical PPT of the user are integrated into a style input sequence; the template selection model is called, and the style input sequence is input into the template selection model, and the corresponding PPT template is obtained as PPT template 09 of activity introduction type.

[0072] It is worth noting that the template selection model is built based on an artificial intelligence model, including:

[0073] Acquire a standard data set; wherein the standard data set includes standard input data consistent with the content attributes of the style input sequence, and standard output data consistent with the content attributes of the PPT template;

[0074] Divide the standard data set into a training set, a validation set, and a test set; use the training set to train the artificial intelligence model; use the validation set to adjust the internal parameters of the artificial intelligence model; use the test set to test the trained artificial intelligence model and obtain test indicators;

[0075] Get the indicator threshold and determine whether the test indicator is greater than the indicator threshold; if yes, mark the trained artificial intelligence model as the template selection model; if no, retrain the artificial intelligence model; the artificial intelligence model includes a convolutional neural network model or an Inception series model.

[0076] It should be noted that the test indicators include accuracy, recall rate, F1 score and stability; the indicator thresholds are set by expert assessment; the proportion of the standard data set divided into training set, validation set and test set is obtained based on experiments. When the trained artificial intelligence model does not meet the requirements, the proportion of the standard data set is redivided and the proportion of the training set is appropriately increased.

[0077] See also Figure 3 , retrieve the user's input text and key features; annotate the key features in the input text; analyze the paragraphs of the input text according to the annotated results to obtain paragraph features; merge paragraphs with the same paragraph features to obtain the paragraph themes of the input text; based on the paragraph themes of the input text, use the large language model to generate a PPT outline.

[0078] Retrieve the outline of the PPT; determine the title name of the PPT according to the outline of the PPT, and generate the number of pages of the PPT according to the title name of the PPT; retrieve the input text corresponding to the title name of the PPT, and enter the input text into the position corresponding to the title name; integrate the page numbers of all PTTs in the order of the title names in the input text to obtain the first version of the PPT.

[0079] It should be noted that large language models include BERT and its variants, the GPT series, RoBERTa, etc.; users can choose the appropriate large language model according to their own needs.

[0080] For example: suppose that the outline of the PPT is analyzed based on the input text and key features of user A, and the key features are annotated in the input text; the paragraphs of the input text are analyzed according to the annotated results to obtain the paragraph features; the paragraphs with the same paragraph features are merged to obtain the paragraph themes of the input text. The outline of the PPT is generated based on the themes, which are: the significance of holding the event, the time of the event, the venue of the event, and the content of the event;

[0081] Determine the title of the PPT according to the outline of the PPT, and enter the input text into the position corresponding to the title name; integrate the pages of all PTTs in the order of the title names in the input text to obtain the first version of the PPT.

[0082] Retrieve the user's conditional information; obtain the basic information in the first version of the PPT; the conditional information includes the number of pictures, the number of animations, and the number of PPT pages; the data type of the basic information is consistent with the conditional information; determine whether the basic information in the first version of the PPT is greater than the conditional information; if yes, mark the first version of the PPT as the final version of the PPT; if not, adjust the first version of the PPT until it meets the conditions.

[0083] It should be noted that when adjusting the initial version of the PPT, the initial version of the PPT can be sent to the user, who can adjust the initial version of the PPT himself, or the user can use the PPT intelligent production tool to generate pictures and so on to adjust the initial version of the PPT.

[0084] For example: Assume that the condition information of user A is obtained, the basic information of the first version of the PPT is obtained, and whether the first version of the PPT meets the requirements is determined based on the basic information. It is found that the first version of the PPT has one less picture than the standard, so the first version of the PPT is analyzed. User A added a picture to the part of the event venue; the first version of the PPT is marked as the final version of the PPT.

[0085] Retrieve the final version of the PPT and the user's input text; extract image information from the final version of the PPT, and generate explanatory text based on the image information; and insert the explanatory text into the corresponding position of the input text; use artificial intelligence tools to check and correct the grammar and spelling in the input text, and expand the input text to obtain the corresponding speech; compare and match the speech with each page of the PPT to obtain a matching result; annotate the speech according to the matching result; wherein the artificial intelligence tool includes a chat assistant or ProWritingAid.

[0086] For example: suppose a speech is generated for user A, image information is extracted from the final version of the PPT, explanatory text is generated based on the image information, the image of the event venue is explained, the reasons and benefits of choosing the event venue are explained, and explanatory text is generated, and the explanatory text is inserted into the input text; the grammar and spelling in the input text are checked and corrected, and the content of the input text is supplemented to obtain the speech; the speech is matched with the final version of the PPT, and the speech content corresponding to each PPT is marked differently.

[0087] It is worth mentioning that matching the speech with the PPT and annotating the speech will enable the speaker to switch the PPT pages in time during the speech, which is conducive to improving the audience's viewing experience and satisfaction.

[0088] Part of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is a formula closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.

[0089] The working principle of the present invention is as follows: the present invention obtains the user's input text, determines the speech type and PPT type according to the user's input text; determines the PPT template according to the PPT type; analyzes the PPT outline according to the user's input text; generates a first version of the PPT based on the PPT outline; sends the first version of the PPT to the user for confirmation, and obtains the user's condition information; adjusts the first version of the PPT according to the condition information to obtain a final version of the PPT; and generates the user's speech manuscript according to the final version of the PPT and the user's input text.

[0090] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for generating PPT and speech scripts based on artificial intelligence, characterized in that: include: Step S1: obtaining a user's input text, and determining a speech type and a PPT type according to the user's input text; Determine the PPT template according to the type of PPT; Step S2: Analyze the outline of the PPT according to the user's input text; generate a preliminary version of the PPT based on the outline of the PPT; Step S3: Send the first version of the PPT to the user for confirmation and obtain the user's condition information; Adjust the first version of PPT according to the condition information to get the final version of PPT; Step S4: Generate the user's speech draft based on the final version PPT and the user's input text.

2. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: Determining the speech type and PPT type according to the user's input text includes: Retrieve the user's input text; extract keywords from the input text to obtain key features; obtain several speech topics; calculate the correlation between the key features and several speech topics; Sort the relevance in descending order to obtain a speech relevance ranking table; determine the speech topic according to the speech relevance ranking table; determine the speech type and PPT type according to the speech topic.

3. The method for generating PPT and speech based on artificial intelligence according to claim 2, characterized in that: The key features are obtained by extracting the key words from the input text, including: Retrieving input text, preprocessing the input text to obtain a number of words; after marking the parts of speech of the words, removing stop words to obtain a number of candidate keywords; Take the candidate keywords as nodes and construct a directed weighted graph G = (V, E); use the co-occurrence relationship to construct the edge between any two points; where V is the node set consisting of the candidate keywords; E is the edge set, representing the relationship between the nodes; Set the weight of each node in the directed weighted graph by the formula Iterate the propagation of node weights to obtain the final score of the node; where d represents the damping factor, M(i) represents the set of all other nodes pointing to node i; L(j) represents the number of other nodes that node j can reach; w j represents the current score of node j; The nodes are sorted according to their final scores to obtain an important sorting table; candidate keywords corresponding to the first n nodes in the important sorting table are selected as keywords of the input text; the keywords are integrated to obtain key features of the input text.

4. The method for generating PPT and speech based on artificial intelligence according to claim 2, characterized in that: The key features of the calculations are related to several presentation topics, including: Retrieve keywords from key features to obtain several speech topics; use word embedding models to analyze feature vectors of keywords and speech topics; where word embedding models include Word2Vec or Glove; By formula Calculate the cosine similarity between the keyword and the speech topic; where A·B represents the dot product of the keyword’s feature vector and the speech topic’s feature vector; ‖A‖ represents the modulus of the keyword’s feature vector; ‖B‖ represents the modulus of the speech topic’s feature vector; The correlation between the key features and the speech topic is calculated using a weighted average method according to the set keyword weights; wherein the keyword weights are determined according to the number of keywords in the key features.

5. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: Determining the template of the PPT according to the type of the PPT includes: Retrieve the type of PPT; select several PPT templates of corresponding types from the PPT type template library according to the type of PPT; Obtain the user's historical PPT, and integrate the user's historical PPT and PPT template into a style input sequence; call the template selection model, input the style input sequence into the template selection model, and obtain the corresponding PPT template; wherein the template selection model is built based on an artificial intelligence model.

6. The method for generating PPT and speech based on artificial intelligence according to claim 5, characterized in that: The template selection model is constructed based on an artificial intelligence model and includes: Acquire a standard data set; wherein the standard data set includes standard input data consistent with the content attributes of the style input sequence, and standard output data consistent with the content attributes of the PPT template; Divide the standard data set into a training set, a validation set, and a test set; use the training set to train the artificial intelligence model; use the validation set to adjust the internal parameters of the artificial intelligence model; use the test set to test the trained artificial intelligence model and obtain test indicators; Get the indicator threshold and determine whether the test indicator is greater than the indicator threshold; if yes, mark the trained artificial intelligence model as the template selection model; if no, retrain the artificial intelligence model; the artificial intelligence model includes a convolutional neural network model or an Inception series model.

7. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: The step of analyzing the outline of the PPT according to the user's input text includes: Retrieving the user's input text and key features; marking the key features in the input text; analyzing the paragraphs of the input text according to the marking results to obtain paragraph features; Merge paragraphs with the same paragraph features to obtain the paragraph topics of the input text; based on the paragraph topics of the input text, use the large language model to generate the outline of the PPT.

8. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: The generating of the first version of the PPT based on the PPT outline includes: Retrieve the outline of the PPT; determine the title of the PPT according to the outline of the PPT; generate the number of pages of the PPT according to the title of the PPT; Retrieve the input text corresponding to the PPT title name; enter the input text into the position corresponding to the title name, integrate all the pages of PTT in the order of the title names in the input text, and obtain the first version of the PPT.

9. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: The first version of the PPT is adjusted according to the condition information to obtain the final version of the PPT, including: Retrieve the user's conditional information; obtain the basic information in the first version of the PPT; the conditional information includes the number of pictures, the number of animations, and the number of PPT pages; the data type of the basic information is consistent with the conditional information; Determine whether the basic information in the preliminary version of the PPT is greater than the conditional information; if so, mark the preliminary version of the PPT as the final version of the PPT; if not, adjust the preliminary version of the PPT until it meets the conditions.

10. The method for generating PPT and speech based on artificial intelligence according to claim 1, characterized in that: The generating of the user's speech draft according to the final version PPT and the user's input text includes: Retrieve the final version of the PPT and the user's input text; extract image information from the final version of the PPT, generate explanatory text based on the image information; and insert the explanatory text into the position corresponding to the input text; Use artificial intelligence tools to check and correct the grammar and spelling in the input text, and expand the input text to obtain the corresponding speech; compare and match the speech with each page of the PPT to obtain a matching result; annotate the speech according to the matching result; among them, artificial intelligence tools include chat assistants or ProWritingAid.

Citation Information

Patent Citations

  • Manuscript generation method and device based on artificial intelligence, computer equipment and medium

    CN118964649A

Cited By

  • Presentation manuscript intelligent dubbing method and device, computer equipment and storage medium

    CN121191492A