Teleprompting method and apparatus, AR glasses and storage medium
By displaying prompt text on AR glasses and collecting user voice information in real time to control the switching of prompt content, the shortcomings of AR glasses in prompting are solved, realizing real-time matching and seamless prompting function, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-14
- Publication Date
- 2026-03-27
AI Technical Summary
The existing AR glasses functionality still needs further development, especially in terms of prompting, which is insufficient and cannot effectively assist users in real-time prompting during speeches, performances and other activities.
By acquiring the prompt text and displaying the current prompt content on the AR glasses, the system collects the user's voice information in real time, controls the switching of prompt content based on the voice information, including determining the next prompt content and the switching delay time, and uses speech recognition and semantic understanding technologies to match the user's reading progress.
It enables real-time matching and seamless integration of AR glasses during the prompting process, improving the user experience and meeting users' real-time prompting needs.
Smart Images

Figure CN115811580B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of display devices, and in particular to a teleprompter method and device, AR glasses and a storage medium. BACKGROUND
[0002] In related technologies, a teleprompter is usually arranged at a position that can be directly viewed by a user's eyes, so as to facilitate the user to view the text information displayed on the teleprompter, and to provide assistance to the user in the process of speaking, performing, hosting, etc.
[0003] With the continuous development of technology, the application of augmented reality (AR) glasses is also becoming more and more widespread. At present, the functions of AR glasses still need to be further developed. SUMMARY
[0004] The present disclosure provides a teleprompter method and device, AR glasses and a storage medium, to develop the functions of AR glasses, so as to at least realize teleprompting using AR glasses. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a teleprompter method is provided, the method comprising: obtaining prompt text and displaying current prompt content in the prompt text through augmented reality (AR) glasses; obtaining current voice information of a user; and controlling the current prompt content displayed to switch to next prompt content according to the current voice information.
[0006] In some embodiments, the obtaining of the current voice information of the user comprises: pre-obtaining voice feature information of the user; collecting audio information in real time; and obtaining the current voice information of the user from the audio information collected in real time according to the voice feature information.
[0007] In some embodiments, the controlling of the current prompt content displayed to switch to the next prompt content according to the current voice information comprises: determining the next prompt content and a switching delay time according to the current voice information; and controlling the current prompt content displayed to switch to the next prompt content after the switching delay time is reached.
[0008] In some embodiments, the determining of the next prompt content according to the current voice information comprises: extracting a keyword according to the current voice information; determining a content segment in the prompt text according to the keyword; wherein the content segment is text content in the current prompt content, or is text content in second prompt content in the prompt text other than the current prompt content; and determining the next prompt content corresponding to the current voice information according to the content segment.
[0009] In some embodiments, the determining the next prompt content corresponding to the current voice information according to the content segment comprises: determining a text position of the content segment in the current prompt content or the second prompt content; and in a case where the text position is located within a first word range at the end of the current prompt content or the second prompt content, determining that text content of a preset number of words after the prompt text in the current prompt content or the second prompt content as the next prompt content.
[0010] In some embodiments, the method further comprises: in a case where the text position is located within a second word range at the end of the second prompt content, determining that the second prompt content as the next prompt content.
[0011] In some embodiments, the determining the switching delay time according to the current voice information comprises: determining speech speed information according to the current voice information; and determining the switching delay time according to the speech speed information and the text position.
[0012] In some embodiments, the determining the switching delay time according to the speech speed information and the text position comprises: determining a remaining word number of the current prompt content or the second prompt content after the text position according to the text position; and determining the switching delay time according to the remaining word number and the speech speed information.
[0013] In some embodiments, the determining that the text content of the preset number of words after the prompt text in the current prompt content or the second prompt content as the next prompt content comprises: determining the preset number of words according to the speech speed information; or determining the preset number of words according to a teleprompter display capacity of the AR glasses.
[0014] In some embodiments, the method further comprises: setting a text format of teleprompter display of the AR glasses to obtain the teleprompter display capacity.
[0015] In some embodiments, the obtaining the prompt text and displaying the current prompt content in the prompt text through the AR glasses comprises: uploading the prompt text to a teleprompter text library of the AR glasses for storage; obtaining the prompt text from the teleprompter text library; dividing the prompt text into at least one text unit according to a teleprompter display capacity of the AR glasses, taking a first text unit as the current prompt content, and displaying the current prompt content through the AR glasses.
[0016] In some embodiments, the obtaining the prompt text from the teleprompter text library comprises: performing semantic understanding on the current voice information to obtain the corresponding prompt text from the teleprompter text library.
[0017] According to a second aspect of the embodiments of the present disclosure, a word prompting device is further provided, the device comprising: a display unit configured to acquire prompt text and display current prompt content in the prompt text through AR glasses; an information acquisition unit configured to acquire current voice information of a user; and a control unit configured to control the current prompt content displayed to switch to next prompt content according to the current voice information.
[0018] According to a third aspect of the embodiments of the present disclosure, an AR glass is further provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the word prompting method as described above.
[0019] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer readable storage medium is further provided, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the word prompting method as described above.
[0020] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0021] The word prompting method provided in the embodiments of the present disclosure acquires prompt text and displays current prompt content in the prompt text through augmented reality (AR) glasses, acquires current voice information of a user, and controls the current prompt content displayed to switch to next prompt content according to the current voice information. Thus, the function of the AR glasses is further developed, and the AR glasses can be used to prompt words for the user, and the AR glasses can control the current prompt content displayed to switch to next prompt content in real time according to the current voice information of the user, keep the matching between the word prompting and the reading of the user, seamlessly connect, and improve the user experience.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings incorporated in the specification and forming a part of it, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure without limiting it in an inappropriate manner.
[0024] Figure 1 A flowchart of a word prompting method according to an embodiment of the present disclosure is shown;
[0025] Figure 2 A flowchart of a sub-step of S2 of the word prompting method according to an embodiment of the present disclosure is shown.
[0026] Figure 3 A structural diagram of a teleprompter device shown for an embodiment of the present disclosure;
[0027] Figure 4 A structural diagram of AR glasses shown for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.
[0029] Unless otherwise required by the context, throughout the specification and claims, the term "comprising" is to be interpreted as meaning "including but not limited to". In the description of the specification, the terms "some embodiments", "an embodiment", "one embodiment", etc. are intended to mean that the particular feature, structure, material, or characteristic being described is included in at least one embodiment or example of the present disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. In addition, the particular features, structures, materials, or characteristics described can be included in any suitable way in one or more embodiments or examples.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. The terms "first", "second" are used for description purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0031] In view of the further development of the functions of the existing AR glasses in the related art, the present inventors have found that the existing teleprompter device is usually arranged at a position that can be directly viewed by the user's eyes to provide assistance to the user during activities such as speaking, hosting, performing, etc., while the AR glasses are usually carried by the user and meet the condition of being directly viewed by the eyes. Therefore, the present embodiments provide a teleprompter method, device, AR glasses and storage medium. For example, the teleprompter method, device, electronic device and storage medium provided by the present embodiments are described in detail below with reference to the drawings.
[0032] It should be noted that the word prompting method of the embodiments of the present disclosure can be executed by a word prompting device of the embodiments of the present disclosure, which can be implemented in software and / or hardware, and can be configured in AR glasses, wherein the AR glasses can install and run the program of the word prompting device.
[0033] Figure 1 is a flowchart of a word prompting method according to an exemplary embodiment.
[0034] As shown in Figure 1 The word prompting method provided by the embodiments of the present disclosure includes but is not limited to the following steps:
[0035] S1: Obtain the prompt text, and display the current prompt content in the prompt text through the augmented reality (AR) glasses.
[0036] The prompt text can be a script that the user needs to prompt, and its format is usually text. It can be understood that the type of text can be various, such as Chinese, English or other text, which is set according to the user's needs, and the present disclosure does not make specific limitations.
[0037] The word prompting method provided by the embodiments of the present disclosure can be applied to various scenes according to the user's needs. For example, in the case of a user singing a song, the prompt text can be the lyrics; in the case of a user delivering a speech, the prompt text can be the speech; in the case of a user hosting a program, the prompt text can be the script, etc.
[0038] It can be known that the AR glasses include a display device, which can display images, text, etc.
[0039] It should be noted that the content in the prompt text includes many texts, which can be displayed on the display device of the AR glasses, but in the case of a large number of texts, it can not be displayed on the display device of the AR glasses. Therefore, at a certain moment, the AR glasses can display part or all of the content in the prompt text, at this time, the part or all of the content is the current prompt content.
[0040] For example, in the case of a prompt text being lyrics, the display device of the AR glasses can display two lines of lyrics at a time, at a certain moment, the AR glasses display two lines of lyrics, then at this moment, the current prompt content is the two lines of lyrics displayed by the AR glasses.
[0041] It should be noted that the above examples are only illustrative and not specific to the embodiments of the present disclosure.
[0042] S2: Obtain the current voice information of the user.
[0043] It is understood that, in this embodiment of the disclosure, the user's current voice information can be obtained before the user begins to speak or during the speaking process, indicating that the user is currently speaking.
[0044] For example, to obtain a user's current voice information, you can use a microphone set on the AR glasses to collect the user's voice information; or you can use other devices to collect the user's voice information and then send the collected user's current voice information to the AR glasses.
[0045] S3: Based on the current voice information, control the display of the current prompt content to switch to the next prompt content.
[0046] It is understandable that the current voice information may include the user's voice reading the prompt text, the user's voice, etc., and based on the user's voice reading the prompt text, it is possible to determine which part of the prompt text the user is reading.
[0047] Furthermore, by obtaining the user's position in the prompt text based on the current voice information, determining the next prompt content, and switching accordingly, the AR glasses can update the content of the prompt text displayed to follow the user's reading, thus providing real-time assistance to the user in reading the prompt text.
[0048] The prompting method provided in this embodiment acquires prompt text and displays the current prompt content in the prompt text through augmented reality (AR) glasses. It also acquires the user's current voice information and, based on this information, controls the display of the current prompt content to switch to the next prompt content. This further develops the functionality of AR glasses, enabling them to provide prompts for users. Furthermore, the AR glasses can control the display of the current prompt content to switch to the next prompt content in real time based on the user's current voice information, maintaining a seamless match between the prompt and the user's reading, thus improving the user experience.
[0049] like Figure 2 As shown, in some embodiments, S2 in this disclosure embodiment: obtaining the user's current voice information includes:
[0050] S21: Pre-acquire the user's voice characteristics information.
[0051] The voice feature information includes at least one of timbre, voiceprint, and pitch. The voice feature information provides distinctiveness for the recognition of the user's voice, so as to distinguish the user's voice from other sounds in the environment.
[0052] In the embodiments of the present disclosure, the voice feature information of the user is acquired in advance, and the corresponding voice recognition software or program can be arranged in the AR glasses, and the voice feature information of the user is identified and stored, so that the AR glasses can have a certain recognition degree for the voice of the user.
[0053] S22: Real-time collection of audio information.
[0054] In the embodiments of the present disclosure, the microphone arranged in the AR glasses is used to collect the voice in real time, and the collected voice can include the voice of the user or the voice of other persons in the environment.
[0055] S23: According to the voice feature information, the current voice information of the user is acquired from the real-time collected audio information.
[0056] In the embodiments of the present disclosure, the voice feature information of the user with a certain recognition degree is acquired in advance, and the real-time collected audio information is compared with the voice feature information, so that the voice of the user can be identified, and the voice of the user is extracted to acquire the current voice information of the user.
[0057] In some embodiments, S3: according to the current voice information, the current prompt content displayed is switched to the next prompt content, including: according to the current voice information, the next prompt content and the switching delay time are determined; after the switching delay time is reached, the current prompt content displayed is switched to the next prompt content.
[0058] It can be understood that, according to the above, when the content included in the prompt text is too much and the display device of the AR glasses cannot display all the content, the AR glasses only display part of the prompt text, and at this time, the part of the prompt text displayed by the AR glasses can be the current prompt content.
[0059] It should be understood by those skilled in the art that the teleprompter function needs to be switched to the next prompt content after the current teleprompter display is ended, so as to follow the user and meet the needs of the user in real time.
[0060] For example, the next prompt content can be part of the prompt text divided in advance, or can be part of the prompt text determined according to the current voice information. The embodiments of the present disclosure do not make specific limitations.
[0061] In addition, in the embodiments of the present disclosure, the next prompt content and the switching delay time are determined according to the current voice information. For example, the semantic recognition of the current voice information can acquire the specific content of the prompt text read by the user according to the current voice information, and then the position read by the user is determined, and then the next prompt content is switched after or before the current prompt content is read by the user.
[0062] In the case of switching the next prompt content when the user finishes reading the current prompt content, it is necessary to determine the time at which the switching is performed. In the embodiment of the present disclosure, the switching delay time can be determined according to the current voice information, and then the current prompt content displayed is switched to the next prompt content after the switching delay time is reached.
[0063] In some embodiments, determining the next prompt content according to the current voice information includes: extracting a keyword from the current voice information; determining a content segment in the prompt text according to the keyword; wherein the content segment is text content in the current prompt content, or is text content in a second prompt content in the prompt text other than the current prompt content; and determining the next prompt content corresponding to the current voice information according to the content segment.
[0064] In the embodiment of the present disclosure, voice recognition is performed according to the current voice information, and the text content is converted, and then the keyword in the text content is extracted. After the keyword is obtained, the content segment in the prompt text can be determined according to the keyword.
[0065] It can be understood that the content segment can be text content in the current prompt content, or text content in a second prompt content in the prompt text other than the current prompt content. For example, when the user reads according to the current prompt content, the current voice information of the user is obtained, the keyword is extracted, and the content segment in the prompt text is determined, which can be the current prompt content. In the case where the user does not read according to the current prompt content, for example, the user's thinking is confused, or the reading method is changed, and the text content not in the current prompt content is read, at this time, the current voice information of the user is obtained, the keyword is extracted, and the content segment in the prompt text is determined, which can be text content in a second prompt content other than the current prompt content.
[0066] It should be noted that the second prompt content is text content with the same number of words as the current prompt content, or close to the same number of words. The second prompt content can be displayed entirely on the display device of the AR glasses. Specifically, the prompt text can include multiple second prompt contents, and each second prompt content can be displayed entirely on the display device of the AR glasses.
[0067] In the embodiment of the present disclosure, the content segment in the prompt text where the content read by the user is located can be found according to the current voice information of the user, and then the next prompt content corresponding to the current voice information is determined according to the content segment.
[0068] In some embodiments, the next prompt content corresponding to the current voice information is determined according to the content segment, including: determining a text position of the content segment in the current prompt content or the second prompt content; and in a case where the text position is within a first number of words at the end of the current prompt content or the second prompt content, determining text content of a preset number of words of the prompt text after the current prompt content or the second prompt content as the next prompt content.
[0069] In the embodiments of the present disclosure, after the keyword is extracted according to the current voice information, the content segment in which the keyword is located in the prompt text is determined, and then the next prompt content corresponding to the current voice information is determined according to the content segment.
[0070] For example, the text position of the content segment in the current prompt content or the second prompt content is determined, and in a case where the text position is within a first number of words at the end of the current prompt content or the second prompt content, text content of a preset number of words of the prompt text after the current prompt content or the second prompt content is determined as the next prompt content.
[0071] It can be understood that the content segment is content in the prompt text, and the text position of the content segment in the prompt text can be determined according to the content segment, which can be in the current prompt content or in the second prompt content.
[0072] In a case where the text position is within a first number of words at the end of the current prompt content or the second prompt content, it can be understood that the reading direction of the prompt text is from front to back, and the end is the last word of the prompt text. For example, the end of the current prompt content and the second prompt content can be the last word in the reading direction of the current prompt content and the second prompt content.
[0073] In the embodiments of the present disclosure, in a case where the text position is within a first number of words at the end of the current prompt content or the second prompt content, text content of a preset number of words of the prompt text after the current prompt content or the second prompt content is determined as the next prompt content.
[0074] The preset number of words can be a number of words of a display paragraph divided by the user in advance, can be a number of words displayed after each switching set by the user, or can be a number of words determined according to the word display capacity of the display device of the AR glasses.
[0075] The first word number range can be set by the user according to the use habit, or be preconfigured in the AR glasses. When the text position is in the first word number range at the end of the current prompt content or the second prompt content, the end of the current prompt content or the second prompt content is approached, and switching needs to be performed. It can be understood that if the text position is at the end of the current prompt content, the current prompt content has been prompted, and if the next prompt content cannot be determined in time, the AR glasses cannot switch to the next prompt content, and the AR glasses cannot provide the user with a word prompt. Therefore, the next prompt content needs to be determined when only a few words are left before the current prompt content is prompted.
[0076] For example, the first word number range can be a few words or a dozen words, and can be set according to the performance of the AR glasses.
[0077] In some embodiments, when the text position is in a second word number range at the end of the second prompt content, the second prompt content is determined as the next prompt content.
[0078] When the text position is in the second word number range at the end of the second prompt content, the second prompt content is behind the text position and includes many words, that is, it can be understood that the second prompt content needs to be displayed to prompt the user at this time.
[0079] Therefore, the second prompt content is determined as the next prompt content to switch the current prompt content, so that the AR glasses display the second prompt content to read for the user.
[0080] In some embodiments, the switching delay time is determined according to the current voice information, including: determining the speech speed information according to the current voice information; and determining the switching delay time according to the speech speed information and the text position.
[0081] In the embodiments of the present disclosure, the speech speed information can be obtained according to the current voice information. For example, the current voice information can be recognized by voice recognition to convert the current voice information into text, and then the speech speed information is determined according to the time length and the number of words of the current voice information.
[0082] Based on the above, when the text position is in the first word number range at the end of the current prompt content, the text content of the preset word number after the current prompt content is determined as the next prompt content, the switching time needs to be confirmed to switch at the appropriate time, realize the real-time word prompt effect of the AR glasses, and improve the satisfaction of the user.
[0083] In some embodiments, the switching delay time is determined according to the speech speed information and the text position, including: determining, according to the text position, a remaining number of words of the current prompt content or the second prompt content after the text position; and determining, according to the remaining number of words and the speech speed information, the switching delay time.
[0084] It can be understood that, when determining the time for switching, the time required for the user to read the remaining field after the text position is determined according to the speech speed information, and the time is determined as the switching delay time, so that the next prompt content is displayed immediately after the user reads the remaining field after the text position, meeting the needs of the user for prompting.
[0085] In some embodiments, the text content of the preset number of words of the prompt text after the current prompt content or the second prompt content is determined as the next prompt content, including: determining the preset number of words according to the speech speed information; or determining the preset number of words according to the text display capacity of the AR glasses.
[0086] It can be understood that, the text content of the preset number of words of the prompt text after the current prompt content or the second prompt content is determined as the next prompt content, the preset number of words is determined.
[0087] In the embodiments of the present disclosure, the preset number of words can be determined according to the reading habits of the user, the speech speed information is analyzed, at least the next complete sentence read by the user is displayed, or at least several sentences are displayed, and a complete sentence is displayed at the end of the display.
[0088] In addition, in the embodiments of the present disclosure, the preset number of words can also be determined according to the text display capacity of the AR glasses. For example, because the size of the display area of the display device of the AR glasses is fixed, the maximum number of characters that can be accommodated can be calculated. Therefore, the preset number of words is determined according to the text display capacity of the AR glasses, which can maximize the display capacity of the AR glasses to display as much prompt text content as possible, and facilitate the user to use.
[0089] In some embodiments, the text format for text display of the AR glasses is set to obtain the text display capacity.
[0090] In the embodiments of the present disclosure, the text format of the character display of the display device of the AR glasses can be set, and the user can set it according to his own habits. Because the size of the display area of the display device of the AR glasses is fixed, when setting the text format, a smaller text format can be selected to make the AR glasses display as much prompt text content as possible. It can be understood that, the more prompt text content displayed by the AR glasses, the more content the user can preview, which facilitates the user to use and improves the user experience.
[0091] In some embodiments, the prompt text is acquired, and current prompt content in the prompt text is displayed through the AR glasses, including: uploading the prompt text to a teleprompter text library of the AR glasses for storage; acquiring the prompt text from the teleprompter text library; dividing the prompt text into at least one text unit according to a teleprompter display capacity of the AR glasses, taking a first text unit as the current prompt content, and displaying the current prompt content through the AR glasses.
[0092] In the embodiments of the present disclosure, before teleprompting is performed using the AR glasses, the prompt text needs to be uploaded to a teleprompter text library of the AR glasses for storage. For example, the user can upload the prompt text that needs to be teleprompted, or the prompt text acquired by using a crawler technology from a network can also be uploaded to the teleprompter text library. For example, in the case that the prompt text is a song lyric or other text that can be directly acquired from a network and the content is fixed, the prompt text can be stored in the teleprompter text library without repeated uploading.
[0093] It can be understood that the prompt text in the teleprompter text library can be gradually supplemented to enrich the functions of the AR glasses, so that the AR glasses can be applied to various scenes.
[0094] In some embodiments, the prompt text is acquired from the teleprompter text library, including: performing semantic understanding on current voice information, and acquiring corresponding prompt text from the teleprompter text library.
[0095] In the embodiments of the present disclosure, the AR glasses can also start teleprompting after the user finishes reading. For example, first, current voice information of the user is acquired, voice recognition is performed on the current voice information, the intention of the user is acquired, and the prompt text that needs to be teleprompted is determined from the teleprompter text library according to the intention of the user.
[0096] In an exemplary embodiment, when the user is the host, if the current voice information is "Next, let me introduce the guests attending the meeting," the user is identified as needing to introduce the guests, and the guest list is determined from the teleprompter text library as the prompt text, which is then displayed on the AR glasses' display device. In another exemplary embodiment, when the user is the speaker, if the current voice information is "I am number 23, and my speech title is 'My Father,'" the user is identified as needing to present the speech content of contestant number 23, "My Father," and the speech script of contestant number 23, "My Father," is determined from the teleprompter text library as the prompt text, which is then displayed on the AR glasses' display device. In yet another embodiment, when the user is the singer, if the current voice information is "Next, I will sing 'Zhang San's Summer,'" the user is identified as needing to sing the song "Zhang San's Summer," and the lyrics of "Zhang San's Summer" are determined from the teleprompter text library as the prompt text, which is then displayed on the AR glasses' display device, and so on.
[0097] Figure 3 This is a structural diagram of a prompting device according to an exemplary embodiment.
[0098] like Figure 3 As shown, the prompting device 10 includes: a display unit 11, an information acquisition unit 12, and a control unit 13.
[0099] Display unit 11 is used to acquire the prompt text and display the current prompt content in the prompt text through AR glasses.
[0100] The information acquisition unit 12 is used to acquire the user's current voice information.
[0101] The control unit 13 is used to control the display of the current prompt content to switch to the next prompt content based on the current voice information.
[0102] The prompting device provided in this embodiment includes a display unit 11 for acquiring prompt text and displaying the current prompt content in the prompt text through augmented reality (AR) glasses; an information acquisition unit 12 for acquiring the user's current voice information; and a control unit 13 for controlling the displayed prompt content to switch to the next prompt content based on the current voice information. This further develops the functionality of AR glasses, enabling them to provide prompts for users and switch the displayed prompt content to the next prompt content in real time based on the user's current voice information, maintaining a seamless match between the prompt and the user's reading, thus improving the user experience.
[0103] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0104] The word prompting device provided by the embodiments of the present disclosure can execute the word prompting method as described in some embodiments above, and has the same beneficial effects as the word prompting method described above, which will not be repeated here.
[0105] Figure 4 is a structural diagram of an AR glasses 100 for a word prompting method according to an exemplary embodiment.
[0106] Exemplarily, the AR glasses 100 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0107] As shown in Figure 4 The AR glasses 100 can include one or more of the following components: a processing component 101, a memory 102, a power supply component 103, a multimedia component 104, an audio component 105, an input / output (I / O) interface 106, a sensor component 107, and a communication component 108.
[0108] The processing component 101 generally controls the overall operation of the AR glasses 100, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 101 can include one or more processors 1011 to execute instructions to complete all or part of the steps of the methods described above. In addition, the processing component 101 can include one or more modules to facilitate interaction between the processing component 101 and other components. For example, the processing component 101 can include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.
[0109] The memory 102 is configured to store various types of data to support the operation of the AR glasses 100. Examples of these data include instructions for any application or method operating on the AR glasses 100, contact data, phonebook data, messages, pictures, videos, and the like. The memory 102 can be implemented by any type of volatile or non-volatile storage devices, or a combination thereof, such as SRAM (Static Random-Access Memory), EEPROM (Electrically Erasable Programmable read only memory), EPROM (Erasable Programmable Read-Only Memory), PROM (Programmable read-only memory), ROM (Read-Only Memory), magnetic storage, flash memory, magnetic or optical disks.
[0110] The power component 103 provides power for various components of the AR glasses 100. The power component 103 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the AR glasses 100.
[0111] The multimedia component 104 includes a touch display screen providing an output interface between the AR glasses 100 and a user. In some embodiments, the touch display screen can include an LCD (Liquid Crystal Display) and a TP (Touch Panel). The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 104 includes a front-facing camera and / or a rear-facing camera. The front-facing camera and / or the rear-facing camera can receive external multimedia data when the AR glasses 100 are in an operating mode, such as a photographing mode or a video mode. Each of the front-facing camera and the rear-facing camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0112] The audio component 105 is configured to output and / or input audio signals. For example, the audio component 105 includes a microphone (MIC) that is configured to receive an external audio signal when the AR glasses 100 are in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 102 or sent out via the communication component 108. In some embodiments, the audio component 105 also includes a speaker for outputting an audio signal.
[0113] The I / O interface 2112 provides an interface between the processing component 101 and peripheral interface modules, which can be a keyboard, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0114] The sensor component 107 includes one or more sensors for providing status assessments of various aspects of the AR glasses 100. For example, the sensor component 107 can detect an open / closed status of the AR glasses 100, relative positioning of components of the AR glasses 100, such as a display and a keypad of the AR glasses 100, a change in position of the AR glasses 100 or a component of the AR glasses 100, a presence or absence of user contact with the AR glasses 100, a orientation or acceleration / deceleration / rotation of the AR glasses 100, and a temperature change of the AR glasses 100. The sensor component 107 can include a proximity sensor that is configured to detect presence of a nearby object without any physical contact. The sensor component 107 can further include a light sensor, such as a CMOS (complementary metal-oxide semiconductor) or CCD (charge-coupled device) image sensor, for use in imaging applications. In some embodiments, the sensor component 107 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0115] The communication component 108 is configured to facilitate wired or wireless communication between the AR glasses 100 and other devices. The AR glasses 100 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an example embodiment, the communication component 108 receives broadcast signals or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 108 further includes an NFC (Near Field Communication) module to facilitate short-range communication. For example, the NFC module can be implemented based on RFID (Radio Frequency Identification) technology, IrDA (Infrared Data Association) technology, UWB (Ultra Wide Band) technology, BT (Bluetooth) technology, and other technologies.
[0116] In an example embodiment, the AR glasses 100 can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), controllers, microcontrollers, microprocessors or other electronic elements for performing the word prompting method described above. It should be noted that the implementation process and technical principles of the electronic device in the present embodiment are described above in the explanation of the word prompting method of the present disclosure, and will not be repeated here.
[0117] The AR glasses 100 provided by the embodiments of the present disclosure can perform the word prompting method as described in some of the above embodiments, and have the same beneficial effects as the word prompting method described above, which will not be repeated here.
[0118] In order to implement the above-mentioned embodiments, the present disclosure further provides a storage medium.
[0119] The instructions in the storage medium, when executed by the processor of the electronic device, enable the electronic device to perform the word prompting method as previously described. For example, the storage medium can be a ROM (Read Only Memory Image), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0120] To implement the above-described embodiments, the present disclosure also provides a computer program product, which, when executed by the processor of the electronic device, enables the electronic device to perform the word prompting method as previously described.
[0121] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present disclosure cover any and all variations of the present disclosure that come within the scope of the following claims and their equivalents. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0122] It should be understood that the present disclosure is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A teleprompter method, characterized in that, The method includes: Obtain the prompt text and display the current prompt content in the prompt text through augmented reality (AR) glasses; Obtain the user's current voice information; Extract keywords based on the current voice information; Based on the keywords, a content fragment in the prompt text is determined; wherein, the content fragment is the text content in the current prompt content, or the text content in a second prompt content other than the current prompt content, and the second prompt content is text content with the same or similar word count as the current prompt content; Determine the text position of the content fragment within the current prompt content or the second prompt content; If the text position is within a first number of characters from the end of the current prompt content or the second prompt content, the text content of the prompt text that is a preset number of characters after the current prompt content or the second prompt content is determined as the next prompt content; Based on the current speech information, determine the speech rate information; Based on the text position, determine the number of remaining characters of the current prompt content or the second prompt content after the text position; The switching delay time is determined based on the remaining number of words and the speech rate information; After the switching delay time is reached, the currently displayed prompt content is switched to the next prompt content.
2. The method as described in claim 1, characterized in that, The acquisition of the user's current voice information includes: Pre-acquiring user voice characteristics information; Real-time audio information acquisition; Based on the sound feature information, the user's current voice information is obtained from the real-time collected audio information.
3. The method as described in claim 1 or 2, characterized in that, The method further includes: If the text position is within the second character range of the end of the second prompt content, the second prompt content is determined to be the next prompt content.
4. The method as described in claim 1 or 2, characterized in that, The step of determining the next prompt content before determining the text content of a preset number of characters following the current prompt content or the second prompt content includes: The preset number of words is determined based on the speech rate information; Alternatively, the preset number of characters can be determined based on the teleprompter display capacity of the AR glasses.
5. The method as described in claim 4, characterized in that, The method further includes: The text format for prompting display on the AR glasses is set, and the prompting display capacity is obtained.
6. The method as described in claim 1 or 2, characterized in that, The step of acquiring the prompt text and displaying the current prompt content in the prompt text through AR glasses includes: The prompt text is uploaded to the prompt text library of the AR glasses for storage; Retrieve the prompt text from the prompting text library; Based on the prompting display capacity of the AR glasses, the prompt text is divided into at least one text unit, and the first text unit is used as the current prompt content and displayed through the AR glasses.
7. The method as described in claim 6, characterized in that, The step of obtaining the prompt text from the prompting text library includes: The current speech information is semantically understood, and the corresponding prompt text is obtained from the prompting text library.
8. A teleprompter device, characterized in that, The device includes: The display unit is used to acquire the prompt text and display the current prompt content in the prompt text through AR glasses; The information acquisition unit is used to acquire the user's current voice information; The control unit is configured to: extract keywords based on the current voice information; determine a content segment in the prompt text based on the keywords; wherein the content segment is the text content in the current prompt content, or the text content in a second prompt content other than the current prompt content, the second prompt content being text content with the same or similar word count as the current prompt content; determine the text position of the content segment in the current prompt content or the second prompt content; if the text position is within a first word count range from the end of the current prompt content or the second prompt content, determine the text content of a preset word count following the current prompt content or the second prompt content as the next prompt content; determine speech rate information based on the current voice information; determine the remaining word count of the current prompt content or the second prompt content after the text position based on the text position; determine a switching delay time based on the remaining word count and the speech rate information; and control the displayed current prompt content to switch to the next prompt content after the switching delay time is reached.
9. An AR glasses, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the prompting method as described in any one of claims 1 to 7.
10. A computer-readable storage medium comprising computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the prompting method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the prompting method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Lyric prompting method and device, storage medium and augmented reality equipment
CN110347865A
Text display method, teleprompter and teleprompter system
CN111259135A