Multi-terminal converged presentation display method and system

By recognizing real-time speech audio on mobile devices and generating display control commands, combined with synchronized adjustments on the cloud and local devices, the problem of discrepancies between the playback progress of the speech script and the actual speech progress has been solved, improving the speech effect and audience experience.

CN117493593BActive Publication Date: 2026-07-21SHANGHAI TECH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI TECH UNIV
Filing Date
2023-11-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In traditional speeches, the playback progress of the speech script often differs from the actual delivery of the speech, affecting the effectiveness of the speech and the audience's understanding.

Method used

By performing speech recognition on the user's real-time speech audio on the mobile device, matching the speech content, generating display control commands, and synchronously adjusting the speech display page through the cloud and local terminals, the system ensures that the speech content is consistent with the actual speech content.

Benefits of technology

It synchronizes the content of the speech draft with the actual speech, improving the speech effect and audience comprehension, and provides personalized polishing and expression prompts, enhancing the interactivity and appeal of the speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117493593B_ABST
    Figure CN117493593B_ABST
Patent Text Reader

Abstract

The application relates to a multi-terminal fusion speech method and system. The method comprises the following steps: the mobile terminal displays the speech script; the mobile terminal receives and identifies real-time speech audio of a user, matches the identified speech audio with the speech script, generates and sends corresponding display control instructions; the mobile terminal adjusts the display page of the speech script according to the display control instructions; the cloud terminal receives and forwards the corresponding display control instructions; and the local terminal receives the display control instructions and adjusts the display page of the speech script according to the display control instructions. Through the multi-terminal fusion mode, the speech is provided with all-round support based on the personalized user demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimedia processing technology, and in particular to a method and system for displaying presentation slides across multiple devices. Background Technology

[0002] In traditional speeches, speakers typically need external devices like page turners to navigate the slides. If the speaker forgets to turn a page or turns it too early, the playback progress becomes inconsistent with the content, negatively impacting the presentation. Furthermore, when the playback progress is out of sync with the content, the audience, relying solely on audio, cannot accurately grasp the main points, severely affecting their listening experience. Therefore, a multi-device integrated method and system for presenting speech slides is needed. Summary of the Invention

[0003] This invention provides a multi-terminal integrated method and system for displaying presentation slides. This addresses the problem in existing technologies where, without human intervention, the playback progress of the presentation slides often deviates from the actual presentation progress.

[0004] This invention provides a multi-terminal integrated speech presentation method, wherein the speech presentation is stored in the cloud, on a mobile device, and on a local device. The method includes: the mobile device displaying the speech presentation; the mobile device receiving and recognizing the user's real-time speech audio, matching the recognized speech audio with the speech presentation, generating and sending corresponding presentation control instructions; the mobile device adjusting the presentation page of the speech presentation according to the presentation control instructions; the cloud receiving and forwarding the corresponding presentation control instructions; and the local device receiving the presentation control instructions and adjusting the presentation page of the speech presentation according to the presentation control instructions.

[0005] In one embodiment of the present invention, the mobile terminal receives and recognizes the user's real-time speech audio, matches the recognized speech audio with the speech manuscript, and generates and sends corresponding display control instructions, including: the mobile terminal receives the user's real-time speech audio and, based on a speech recognition model, translates the speech audio into speech text; the mobile terminal calculates the matching degree between the speech text and each sentence in the speech manuscript based on a text similarity algorithm, and selects the sentence with the highest matching degree as the target sentence; the mobile terminal generates and sends corresponding display control instructions based on the position of the target sentence on the speech manuscript's display page.

[0006] In one embodiment of the present invention, the mobile terminal receives real-time speech audio from a user and translates the speech audio into speech text based on a speech recognition model, including: the mobile terminal receives real-time speech audio from a user and inputs the speech audio into an acoustic model to extract acoustic features of the speech audio; the mobile terminal inputs the acoustic features into a language model and processes the acoustic features based on audio decoding and search algorithms to obtain the speech text; wherein, the speech recognition model includes an acoustic model and a language model connected in sequence.

[0007] In one embodiment of the present invention, the mobile terminal adjusts the display page of the speech according to the display control instruction, including: the mobile terminal adjusts the speech to scroll to the display page of the target statement based on the display control instruction.

[0008] In one embodiment of the present invention, the mobile terminal adjusts the display page of the speech according to the display control instruction, and further includes: the mobile terminal obtaining the display page that the speech should be displayed at the current moment based on the expression strategy of the speech; wherein, the expression strategy is obtained by preprocessing the speech; the mobile terminal determines whether the display page to which the target statement belongs is the display page that should be displayed at the current moment, and generates a prompt message when it is not the display page that should be displayed at the current moment.

[0009] In one embodiment of the present invention, the mobile terminal adjusts the display page of the speech according to the display control instruction, and further includes: the mobile terminal obtaining the expression strategy corresponding to the display page to which the target statement belongs, and displaying the expression strategy on the interface of the mobile terminal; wherein, the expression strategy is obtained by preprocessing the speech.

[0010] In one embodiment of the present invention, after the mobile terminal receives the user's real-time speech audio, the method further includes: the mobile terminal performing noise reduction processing on the speech audio.

[0011] In one embodiment of the present invention, the speech draft is obtained through preprocessing. The preprocessing process includes: the local terminal acquiring an initial speech draft and speech draft requirements, and sending the initial speech draft and speech draft requirements to the cloud; the cloud calling a trained data processing model to polish the initial speech draft according to the speech draft requirements, obtaining a polished speech draft and expression prompts, and sending the polished speech draft and expression prompts to the local terminal; wherein, the data processing model is a ChatGPT4 model, and the expression prompts are modification strategies and expression strategies for the polished speech draft; the local terminal modifying the polished speech draft according to the expression prompts, and after modification, sending the expression strategies and modified speech draft to the cloud; the cloud receiving and forwarding the modified speech draft and expression strategies to the mobile terminal.

[0012] In one embodiment of the present invention, the multi-terminal integrated speech method further includes: the mobile terminal displaying the playback progress of the speech on the speech display page.

[0013] In another aspect of the present invention, a multi-terminal integrated speech presentation system is also provided. The system includes a mobile terminal, a cloud terminal, and a local terminal. The cloud terminal and the mobile terminal are communicatively connected, and the cloud terminal and the local terminal are also communicatively connected. The speech presentation is stored in the cloud terminal, the mobile terminal, and the local terminal. The mobile terminal includes: a speech presentation module for displaying the speech presentation; an instruction generation module for receiving and recognizing the user's real-time speech audio, matching the recognized speech audio with the speech presentation, generating and sending corresponding presentation control instructions; a page adjustment module for adjusting the presentation page of the speech presentation according to the presentation control instructions; and a first communication module for sending presentation control instructions to the cloud terminal. The cloud terminal includes: a second communication module for receiving presentation control instructions sent by the mobile terminal and forwarding them to the local terminal. The local terminal includes: a synchronization control module for adjusting the presentation page of the speech presentation according to the presentation control instructions; and a third communication module for receiving presentation control instructions sent by the cloud terminal.

[0014] This invention proposes a multi-terminal integrated speech presentation method and system. On the mobile device, real-time audio of the user's speech is recognized and matched with the speech manuscript to determine the audio's position within the manuscript. The mobile device controls the speech manuscript to adjust to the corresponding display page and sends display control commands to the cloud. The cloud sends these commands to the local device, enabling both the local and mobile devices to synchronously adjust the speech manuscript's display page. This ensures that the speech manuscript content presented locally matches the user's actual speech. This solves the problem in existing technologies where, without human intervention, the playback progress of the speech manuscript often deviates from the actual speech progress. By viewing the corresponding display page on the mobile device, users can understand the current speech content and deliver their speeches accordingly, greatly improving the presentation effect. Attached Figure Description

[0015] Figure 1 A flowchart illustrating a multi-terminal integrated speech presentation method provided in an embodiment of the present invention;

[0016] Figure 2 The diagram shows a multi-terminal integration process during the presentation phase in one embodiment of the present invention.

[0017] Figure 3 The diagram shows a flowchart of multi-terminal fusion in the preprocessing stage according to an embodiment of the present invention.

[0018] Figure 4This is a diagram showing the overall architecture of a multi-terminal integrated speech presentation method according to an embodiment of the present invention.

[0019] Figure 5 This is a schematic diagram illustrating the configuration of presentation requirements in one embodiment of the present invention;

[0020] Figure 6 The diagram shown is a schematic diagram of multi-level speech rate control in one embodiment of the present invention;

[0021] Figure 7 This is a schematic diagram illustrating the switching of the presentation slides page in one embodiment of the present invention;

[0022] Figure 8 The diagram shown is a structural block diagram of a multi-terminal integrated speech presentation system provided in an embodiment of the present invention.

[0023] Figure 9 The diagram shown is a structural schematic of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0025] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0026] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0027] The inventors discovered that in traditional academic presentations, speakers often face challenges such as preparing extensive materials, managing time effectively, and maintaining audience interest. Furthermore, speakers struggle to effectively convey information from multiple sources, often resulting in monotonous readings of prepared texts or slides. While existing tools and software can help speakers create slides and manage timings, and provide some prompts for their delivery, they often fail to adapt well to different environments or meet individual needs. This invention provides a multi-platform integrated presentation method that solves the problem in existing technologies where the playback progress of the presentation script often deviates from the actual presentation progress without human intervention. Users can view the corresponding presentation page on their mobile devices to understand the current presentation content and can then deliver their own presentations accordingly, greatly improving the presentation effect. In addition, the mobile interface displays corresponding prompts to guide users through actions, preventing monotonous readings and enhancing audience engagement. Furthermore, by viewing the content displayed locally, listeners can more quickly identify the location of the current speech content within the prepared text. This avoids situations where, due to the speaker's accent, listeners may struggle to understand the specific content solely based on audio, thus improving the listening experience. In addition, preprocessing the speech can allow for personalized polishing based on user needs.

[0028] Please see Figure 1 and Figure 2 The method for presenting a speech across multiple platforms includes the following steps:

[0029] S1. The mobile device displays the speech manuscript.

[0030] Mobile devices include, but are not limited to, smartphones or tablets. The speech draft is pre-saved on the mobile device. The speech draft can be a document (such as a Word document) or a slideshow. For ease of description, this invention will use a slideshow as an example. Other document formats are presented similarly to slideshows and will not be elaborated upon here. Specifically, when a user gives a speech, the speech draft will be displayed on the mobile device's interface. The user holds the mobile device while giving the speech, and the content displayed on the mobile device prompts the user with the current speech content and related expression prompts. This allows the user to take corresponding actions based on the expression prompts or guide the audience to respond accordingly, effectively enhancing the atmosphere and effectiveness of the speech. The mobile device can use the Kotlin and Compose frameworks to implement an Android App, thereby completing the mobile device framework and enabling the corresponding display interface to be presented on the mobile device.

[0031] S2. The mobile terminal receives and recognizes the user's real-time speech audio, matches the recognized speech audio with the speech manuscript, and generates and sends corresponding display control instructions.

[0032] During a presentation, the mobile device receives the user's real-time audio and processes it into text, identifying the corresponding text. The text is then matched against the prepared script to determine the position of the user's audio within the script. Display control commands are generated and sent to the cloud. These commands instruct the mobile and local devices to synchronously display the prepared script based on the user's real-time audio.

[0033] In one embodiment of the present invention, the mobile terminal receives and identifies the user's real-time speech audio, matches the identified speech audio with the speech transcript, generates and sends corresponding display control instructions, including:

[0034] The mobile device receives the user's real-time speech audio and, based on a speech recognition model, translates the speech audio into speech text.

[0035] The mobile device uses a text similarity algorithm to calculate the matching degree between the speech text and each sentence in the speech draft, and selects the sentence with the highest matching degree as the target sentence.

[0036] The mobile device generates and sends corresponding display control commands based on the position of the target statement on the presentation page.

[0037] After receiving the user's real-time speech audio, the mobile device inputs the audio into a speech recognition model for speech recognition, thus translating the audio into speech text. Then, based on a text similarity algorithm, the matching degree between the speech text and each sentence in the speech script is calculated, and the sentence with the highest matching degree is selected as the target sentence. Since users may not read the speech script exactly as it appears, there may be discrepancies between the actual speech and the script. In this case, the text matching algorithm can identify the most matching sentence in the speech script, enabling accurate positioning of the script. The mobile device then scrolls the speech script to the display page containing the target sentence. To facilitate synchronized display of the speech script on both the local and mobile devices, the mobile device generates display control commands based on the position of the target sentence within the speech script and sends these commands to the cloud. The speech recognition models include, but are not limited to, DeepSpeech, Jasper, PaddleSpeech, and the SpeechRecognizer framework. The text similarity algorithms include, but are not limited to, cosine similarity, BM25, and WMD (Word Mover's Distance). In one embodiment, the speech recognition model is the standard speech recognition framework SpeechRecognizer, which is built into the Android system, and the text similarity algorithm is the BM25 algorithm.

[0038] In one embodiment of the present invention, the mobile terminal receives the user's real-time speech audio and, based on a speech recognition model, translates the speech audio into speech text, including:

[0039] The mobile terminal receives the user's real-time speech audio and inputs the speech audio into the acoustic model to extract the acoustic features of the speech audio;

[0040] The mobile device inputs the acoustic features into the language model, processes the acoustic features based on audio decoding and search algorithms, and obtains the speech text; wherein, the speech recognition model includes an acoustic model and a language model connected in sequence.

[0041] After receiving the user's real-time speech audio, the mobile terminal inputs the audio into an acoustic model and extracts acoustic features based on acoustic characteristics. These acoustic features are then input into a language model to obtain the probabilities of possible word sequences. Using an existing dictionary, the word sequences are decoded to obtain the final text representation, which is then used as the translated speech text. Considering the potential presence of audience noise and recording interference in the speech audio, to improve speech recognition accuracy, in one embodiment of this invention, after receiving the user's real-time speech audio, the mobile terminal further performs noise reduction processing on the speech audio.

[0042] S3. The mobile terminal adjusts the display page of the speech according to the display control command.

[0043] The mobile app uses display control commands to move the speech script to a position matching the audio, and then displays the corresponding page on the mobile interface. This allows the speech script to be precisely positioned based on the user's speech, enabling the user to deliver their speech according to the content presented on the mobile app. The speech script can be used to switch pages via scrolling or by jumping directly to the corresponding page; the specific method is not limited here.

[0044] In one embodiment of the present invention, the mobile terminal adjusts the display page of the speech manuscript according to the display control command, including: the mobile terminal adjusting the slides of the speech manuscript to scroll to the display page of the target statement based on the display control command. The mobile terminal adjusts the slides of the speech to scroll to the display page of the target statement according to the display control command, thereby ensuring that the speaker's speaking speed matches the slide playback progress.

[0045] In one embodiment of the present invention, the mobile terminal adjusts the display page of the speech according to the display control command, further comprising:

[0046] The mobile device obtains the display page that the speech should be displayed at the current moment based on the expression strategy of the speech; wherein, the expression strategy is obtained by preprocessing the speech.

[0047] The mobile device determines whether the display page to which the target statement belongs is the display page that should be displayed at the current time, and generates a prompt message when it is not the display page that should be displayed at the current time.

[0048] Please see Figure 6During the presentation, the mobile app retrieves the content of the speech script to be displayed at the current moment based on the user's pre-set presentation duration. However, considering that the actual presentation progress may differ from the estimated progress, the mobile app needs to determine whether the target statement belongs to the page that should be displayed at the current moment. If it does, the mobile app displays the page that should be displayed at the current moment on the mobile interface for the speaker to view. If it does not belong, a speech speed prompt will be issued based on the position of the target statement. Specifically, the mobile app retrieves the paragraph that should be displayed at the current moment based on the user's set presentation duration. If the paragraph containing the target statement is before the paragraph that should be displayed at the current moment, a speech speed warning will be issued, prompting the user to slow down. If the paragraph containing the target statement is after the paragraph that should be displayed at the current moment, and the target statement is within the page that should be displayed, a speech speed warning will be issued, prompting the user to slightly speed up. If the paragraph containing the target statement is after the paragraph that should be displayed at the current moment, and the target statement is not within the page that should be displayed, a speech speed warning will be issued, prompting the user to speed up. This multi-level speech rate control ensures that the user's final presentation time matches the preset duration. Furthermore, to provide a more intuitive reminder, a downward arrow appears on the current screen when the speech rate is too slow, and an upward arrow appears when the speech rate is too fast. This allows users to easily determine whether their speech pace is too fast or too slow by following the arrow's direction.

[0049] Further, please refer to Figure 4 and Figure 7 The mobile interface dynamically displays the speech script, presentation strategies, and the content of the speech. When the speech script is a slideshow, the mobile interface also displays slide thumbnails and the content of the currently displayed slide. To enhance the speaker's experience, the speaker can switch slides by swiping up and down or by clicking on a viewfinder (i.e., a slide thumbnail). By switching the speech script on the mobile device, the speaker can quickly locate the desired section of the speech and generate presentation control commands on the mobile device, which are sent to the local device via the cloud to ensure that the playback progress of the speech script is synchronized between the local and mobile devices.

[0050] In one embodiment of the present invention, the multi-terminal integrated presentation method further includes: the mobile terminal displaying the playback progress of the presentation script on the presentation script's display page. The playback progress can be displayed on the display page in the form of a progress bar, so that the speaker can understand the progress of the presentation.

[0051] S4. The cloud receives and forwards the corresponding display control instructions.

[0052] After receiving the display control command from the mobile device, the cloud server indicates that the presentation page on the mobile device has changed. Therefore, the cloud server sends the display control command to the local server so that the local server can synchronize and adjust the presentation page with the mobile device. The cloud server can be implemented using a dedicated server or a server cluster consisting of multiple servers.

[0053] S5. The local terminal receives the display control instruction and adjusts the display page of the speech according to the display control instruction.

[0054] After receiving display control instructions forwarded from the cloud, the local device acts as the playback medium for the presentation. Based on the content of the instructions, it adjusts the presentation page and plays the adjusted presentation through a large screen or projector, ensuring synchronized playback progress on both the mobile and local devices. The local device can be, but is not limited to, various desktop computers, laptops, smartphones, and tablets. The local device can use the Yeoman generator to create an Office project and utilize the Office.js, TypeScript, and React frameworks to implement it as a PowerPoint add-in. Communication between the local device and the cloud can be achieved using axios.

[0055] In one embodiment of the present invention, the speech draft is obtained through preprocessing, the preprocessing process including:

[0056] The local terminal obtains the initial speech draft and speech draft requirements, and sends the initial speech draft and speech draft requirements to the cloud;

[0057] The cloud-based system calls a pre-trained data processing model to refine the initial speech based on the speech requirements, resulting in a refined speech and expression prompts. The refined speech and expression prompts are then sent to the local machine. The data processing model is the ChatGPT4 model, and the expression prompts are the modification and expression strategies for the refined speech.

[0058] The local terminal modifies the polished speech based on the expression prompts, and then sends the expression strategy and the modified speech to the cloud.

[0059] The cloud receives and forwards the revised speech and presentation strategy to the mobile device.

[0060] Please see Figures 3 to 5The initial speech draft refers to the preliminary stage of the speech. Although it contains some content, it is insufficient to support the completion of the speech. Users upload the initial speech draft to their local device, select the necessary polishing factors from a set of preset factors, and input the text content of the initial speech draft into the desktop interface. Polishing factors include verbal support factors, non-verbal support factors, and visual expression support factors. Verbal factors include, but are not limited to, tone of voice, speaking speed, pronunciation style, and draft size. Non-verbal factors include, but are not limited to, eye contact, facial expressions, composure, gestures, and posture. Visual expression support factors include, but are not limited to, page switching methods such as swiping or jumping. Furthermore, the desktop interface also has a time adjustment function, allowing users to select an appropriate speech duration. Speech requirements include polishing factors and speech duration. After setting the speech duration, users can choose to enhance the draft, thereby sending the initial speech draft and speech requirements to the cloud. The cloud calls the GPT-4 interface and uses the trained ChatGPT4 model to polish the initial speech draft according to the speech requirements and generate corresponding expression prompts. The expression prompts are modification and expression strategies generated by the ChatGPT4 model for the polished speech. Modification strategies can include suggestions for adding or deleting certain fields or words in the speech, while expression strategies can be automatically generated oral, non-oral, and visual expression suggestions based on the user's speaking needs. For example, when speaking statement A, increase the volume; when speaking statement B, use certain body language, etc. After polishing, the cloud sends the polished speech and corresponding expression prompts to the desktop. The user can then modify the polished speech on the desktop based on the expression prompts generated by ChatGPT4. The visual effects of the speech can also be optimized based on the expression prompts. Furthermore, new words can be marked in the speech, and during the modification process, the local device can switch between the initial and modified speech using a toggle button to understand the specific changes made. After modification, a confirmation message is sent to the cloud, along with the modified speech and the expression strategies generated by ChatGPT4. After receiving confirmation, the cloud platform sends the revised speech and presentation strategy to the mobile device. This revised speech serves as the final version for the user's final presentation. Furthermore, the cloud platform can use the Python programming language and the Flask framework to call the OpenAI API to access the ChatGPT4 model.

[0061] Please see Figure 8This multi-terminal integrated speech presentation system includes a mobile terminal, a cloud terminal, and a local terminal. The cloud terminal and the mobile terminal are communicatively connected, as are the local terminal. The speech presentation is stored in the cloud terminal, the mobile terminal, and the local terminal. The mobile terminal includes a speech presentation module 110, an instruction generation module 120, a page adjustment module 130, and a first communication module 140. The speech presentation module 110 is used to display the speech presentation. The instruction generation module 120 is used to receive and recognize the user's real-time speech audio, match the recognized audio with the speech presentation, generate and send corresponding presentation control instructions. The page adjustment module 130 is used to adjust the presentation page of the speech presentation according to the presentation control instructions. The first communication module 140 is used to send presentation control instructions to the cloud terminal. The cloud terminal includes a second communication module 150, used to receive presentation control instructions sent from the mobile terminal and forward them to the local terminal. The local terminal includes a synchronization control module 160 and a third communication module 170. The synchronization control module 160 is used to adjust the presentation page of the speech presentation according to the presentation control instructions. The third communication module 170 is used to receive display control commands sent from the cloud.

[0062] Specific limitations regarding the multi-terminal integrated speech presentation system can be found in the limitations of the multi-terminal integrated speech presentation method described above, and will not be repeated here. Each module in the aforementioned multi-terminal integrated speech presentation system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware format or independent of it, or stored in the memory of the computer device in software format, so that the processor can call the corresponding operations of each module.

[0063] It should be noted that, in order to highlight the innovative aspects of this invention, this embodiment does not include modules that are not closely related to solving the technical problems proposed by this invention, but this does not mean that there are no other modules in this embodiment.

[0064] Please see Figure 9 The electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a multi-terminal presentation program.

[0065] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 12 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 1. Furthermore, the memory 12 can include both internal and external storage units of the electronic device 1. The memory 12 can be used not only to store application software and various types of data installed on the electronic device 1, such as code for multi-terminal integrated presentation slides, but also to temporarily store data that has been output or will be output.

[0066] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the electronic device 1, connecting various components of the electronic device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., multi-terminal integrated presentation programs) and calls data stored in the memory 12 to perform various functions and process data of the electronic device 1.

[0067] The processor 13 executes the operating system of the electronic device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the multi-terminal integrated presentation method described above.

[0068] For example, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a presentation module 110, an instruction generation module 120, a page adjustment module 130, a first communication module 140, a second communication module 150, a synchronization control module 160, and a third communication module 170.

[0069] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The software functional module stored in the storage medium includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute some of the functions of the multi-terminal integrated presentation method described in the various embodiments of this application.

[0070] In summary, the multi-terminal integrated speech presentation method and system disclosed in this invention performs speech recognition on the user's real-time speech audio on the mobile terminal, matches the recognized audio content with the speech manuscript, and thus determines the position of the speech audio within the manuscript. The mobile terminal controls the speech manuscript to adjust to the corresponding display page and sends display control commands to the cloud. The cloud sends the display control commands to the local terminal, enabling both the local and mobile terminals to synchronously adjust the speech manuscript display page, thereby facilitating the matching of the speech manuscript content presented on the local terminal with the user's actual speech content. This solves the problem in existing technologies where the playback progress of the speech manuscript is inconsistent with the actual speech progress without human intervention. By using a large model like ChatGPT4 and combining it with a multi-terminal integration approach, this invention provides comprehensive support for the speech manuscript during the preprocessing stage, including manuscript polishing, duration prediction, and configuration of supporting factors, making the final generated speech manuscript more closely aligned with the speech requirements and meeting the user's personalized needs. Furthermore, during actual presentations, the mobile device offers remote slide control, multi-level speech rate adjustment, dynamic manuscript display, and embedded expression prompts, allowing speakers to adapt their presentation style based on the prompts. This invention integrates mobile, cloud, and local devices, significantly improving presentation effectiveness and audience acceptance through human-computer collaboration. Therefore, this invention effectively overcomes the shortcomings of existing technologies and possesses high industrial application value.

[0071] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for displaying speech manuscripts across multiple terminals, characterized in that, The speech transcript is stored in the cloud, on a mobile device, and locally. The method includes: The mobile device displays the speech transcript; The mobile terminal receives and recognizes the user's real-time speech audio, matches the recognized speech audio with the speech script, and generates and sends corresponding display control commands. The mobile device adjusts the display page of the speech according to the display control command; The cloud receives and forwards the corresponding display control commands; The local terminal receives the display control instruction and adjusts the display page of the speech according to the display control instruction; The mobile terminal adjusts the display page of the speech according to the display control command, including: The mobile device obtains the display page that the speech should be displayed at the current moment based on the expression strategy of the speech; wherein, the expression strategy is obtained by preprocessing the speech. The mobile terminal determines whether the display page to which the target statement belongs is the display page that should be displayed at the current time, and generates a prompt message when it is not the display page that should be displayed at the current time; The speech transcript was obtained through preprocessing, and the preprocessing process included: The local terminal obtains the initial speech draft and speech draft requirements, and sends the initial speech draft and speech draft requirements to the cloud; The cloud-based system calls a pre-trained data processing model to refine the initial speech based on the speech requirements, resulting in a refined speech and expression prompts. The refined speech and expression prompts are then sent to the local machine. The data processing model is the ChatGPT4 model, and the expression prompts are the modification and expression strategies for the refined speech. The local terminal modifies the polished speech based on the expression prompts, and then sends the expression strategy and the modified speech to the cloud. The cloud receives and forwards the revised speech and presentation strategy to the mobile device.

2. The multi-terminal integrated speech presentation method according to claim 1, characterized in that, The mobile terminal receives and recognizes the user's real-time speech audio, matches the recognized speech audio with the speech transcript, generates and sends corresponding display control commands, including: The mobile device receives the user's real-time speech audio and, based on a speech recognition model, translates the speech audio into speech text. The mobile device uses a text similarity algorithm to calculate the matching degree between the speech text and each sentence in the speech draft, and selects the sentence with the highest matching degree as the target sentence. The mobile device generates and sends corresponding display control commands based on the position of the target statement on the presentation page.

3. The multi-terminal integrated speech presentation method according to claim 2, characterized in that, The mobile terminal receives the user's real-time speech audio and, based on a speech recognition model, translates the speech audio into speech text, including: The mobile terminal receives the user's real-time speech audio and inputs the speech audio into the acoustic model to extract the acoustic features of the speech audio; The mobile device inputs the acoustic features into the language model, processes the acoustic features based on audio decoding and search algorithms, and obtains the speech text; wherein, the speech recognition model includes an acoustic model and a language model connected in sequence.

4. The multi-terminal integrated speech presentation method according to claim 2, characterized in that, The mobile terminal adjusts the display page of the speech according to the display control command, and further includes: the mobile terminal adjusts the speech to scroll to the display page of the target statement according to the display control command.

5. The multi-terminal integrated speech presentation method according to claim 4, characterized in that, The mobile terminal adjusts the display page of the speech according to the display control command, and further includes: the mobile terminal obtains the expression strategy corresponding to the display page to which the target statement belongs, and displays the expression strategy on the interface of the mobile terminal; wherein, the expression strategy is obtained by preprocessing the speech.

6. The multi-terminal integrated speech presentation method according to claim 1, characterized in that, After receiving the user's real-time speech audio, the mobile terminal further includes: the mobile terminal performing noise reduction processing on the speech audio.

7. The multi-terminal integrated speech presentation method according to claim 1, characterized in that, The multi-terminal integrated presentation method further includes: the mobile terminal displaying the playback progress of the presentation on the presentation's display page.

8. A multi-terminal integrated speech presentation system, characterized in that, The system includes a mobile terminal, a cloud terminal, and a local terminal. The cloud terminal and the mobile terminal are communicatively connected, and the cloud terminal and the local terminal are also communicatively connected. The presentation script is stored in the cloud terminal, the mobile terminal, and the local terminal. The mobile device includes: The speech presentation module is used to display the speech. The instruction generation module is used to receive and recognize the user's real-time speech audio, match the recognized speech audio with the speech script, generate and send corresponding display control instructions; The page adjustment module is used to adjust the display page of the speech manuscript according to the display control instructions; The first communication module is used to send display control commands to the cloud. The cloud includes: The second communication module is used to receive display control commands sent by the mobile terminal and forward them to the local terminal; The local terminal includes: A synchronization control module is used to adjust the display page of the speech manuscript according to the display control instructions; The third communication module is used to receive display control commands sent from the cloud; The mobile terminal adjusts the display page of the speech according to the display control command, including: The mobile device obtains the display page that the speech should be displayed at the current moment based on the expression strategy of the speech; wherein, the expression strategy is obtained by preprocessing the speech. The mobile terminal determines whether the display page to which the target statement belongs is the display page that should be displayed at the current time, and generates a prompt message when it is not the display page that should be displayed at the current time; The speech transcript was obtained through preprocessing, and the preprocessing process included: The local terminal obtains the initial speech draft and speech draft requirements, and sends the initial speech draft and speech draft requirements to the cloud; The cloud-based system calls a pre-trained data processing model to refine the initial speech based on the speech requirements, resulting in a refined speech and expression prompts. The refined speech and expression prompts are then sent to the local machine. The data processing model is the ChatGPT4 model, and the expression prompts are the modification and expression strategies for the refined speech. The local terminal modifies the polished speech based on the expression prompts, and then sends the expression strategy and the modified speech to the cloud. The cloud receives and forwards the revised speech and presentation strategy to the mobile device.