Institute word display method, apparatus and device, computer readable medium and program product

By performing voice recognition and dynamic adjustment of user voice, combined with matching target prompt source information and rolling information generation, the problem of out-of-synchronization of inscription display in the prior art is solved, and the user experience and synchronization effect are improved.

CN120199249AInactive Publication Date: 2025-06-24HANGZHOU LINGBAN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510422160.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing inscription display method requires manual or timely scrolling of the document during the speech process, resulting in poor user experience and poor synchronization effect.

Method used

By performing voice recognition on the collected user voice, dynamically adjust the identification information, match the target prompt source information, generate scrolling information, and dynamically adjust the content of the teleprompt display.

Benefits of technology

It improves the inscription experience and inscription synchronization effect of users when speaking, so that the content of the teleprompt display can follow the user's speech progress without manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199249A_ABST
    Figure CN120199249A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a question word display method, device and equipment, a computer readable medium and a program product. A specific embodiment of the method comprises the following steps: performing voice recognition on collected user voice to obtain voice recognition information; dynamically adjusting the recognized voice recognition information to obtain adjusted text information; based on target prompt source information, matching the adjusted text information to obtain matching position information; generating rolling information according to the matching position information; and adjusting the prompt display content corresponding to the target prompt source information according to the rolling information. According to the embodiment, the question word experience and the question word synchronization effect when the user speaks are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to a method, apparatus, device, computer-readable medium, and program product for inscribing display. Background Art

[0002] An inscribing system is an auxiliary tool for helping users speak according to a pre-prepared manuscript, which is used to reduce the occurrence of forgetting words, skipping words, etc. Currently, when inscribing, the commonly used method is to manually or periodically scroll the manuscript during the speech.

[0003] However, when using the above method, there are often the following technical problems: the manual method requires distracted operation, resulting in poor user experience; the periodic scrolling method cannot scroll the manuscript according to the actual progress of the speech, and the synchronization effect is poor.

[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not constitute the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0005] The content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure provide a method, apparatus, device, computer-readable medium, and computer program product for inscribing display to solve one or more of the technical problems mentioned in the above background art section.

[0007] In a first aspect, some embodiments of the present disclosure provide an inscribing display method, the method including: performing speech recognition on the collected user speech to obtain speech recognition information; dynamically adjusting the recognized speech recognition information to obtain adjusted text information; matching the adjusted text information based on target prompt source information to obtain matching position information; generating scrolling information according to the matching position information; and adjusting the inscribing display content corresponding to the target prompt source information according to the scrolling information.

[0008] Optionally, performing speech recognition on the collected user speech to obtain speech recognition information, including: generating noise level information of the user speech; in response to determining that the noise level information meets a preset noise condition, adjusting noise reduction window information to obtain updated noise reduction window information; performing noise reduction processing on the user speech according to the updated noise reduction window information to obtain noise-reduced user speech; and performing speech recognition on the obtained noise-reduced user speech to obtain speech recognition information.

[0009] Optionally, performing dynamic adjustment on the recognized speech recognition information to obtain adjusted text information, including: performing dynamic correction processing on the recognized speech recognition information to obtain corrected speech recognition information; adding the obtained corrected speech recognition information to a historical record queue; correcting the corrected speech recognition information according to the historical record queue to obtain corrected speech recognition information; and generating adjusted text information according to the corrected speech recognition information.

[0010] Optionally, generating adjusted text information according to the corrected speech recognition information, including: determining the number of characters included in the corrected speech recognition information; in response to determining that the number of characters meets a preset character condition, performing sentence segmentation processing on the corrected speech recognition information to obtain segmented speech recognition information; and performing punctuation optimization processing on the segmented speech recognition information to obtain optimized speech recognition information as the adjusted text information.

[0011] Optionally, based on target prompt source information, performing matching on the adjusted text information to obtain matching position information, including: in response to determining that the adjusted text information meets a preset character change condition, determining the adjusted text information as text information to be matched; performing text truncation processing on the text information to be matched to obtain each truncated text; for each obtained truncated text, performing the following steps: determining whether there is a matching result corresponding to the truncated text in the cache; in response to determining that there is no matching result corresponding to the truncated text in the cache, determining the currently visible range text information of the target prompt source information; performing matching on the truncated text based on the currently visible range text information or the target prompt source information to obtain a matching position corresponding to the truncated text; and generating matching position information according to the obtained respective matching positions.

[0012] Optionally, matching the truncated text based on the above current visible range text information or the above target hint source information to obtain a matching position corresponding to the truncated text, including: matching the truncated text based on the above current visible range text information to obtain a matching result; in response to determining that the matching result indicates successful matching, determining the matching result as the matching position; in response to determining that the matching result indicates failed matching, matching the truncated text based on the above target hint source information to obtain a matching result as the matching position.

[0013] Optionally, the matching of the truncated text includes: performing an exact match on the truncated text to obtain a first matching result; in response to determining that the first matching result indicates failed matching, determining the longest common substring corresponding to the truncated text in the above current visible range text information or the above target hint source information; performing a match on the longest common substring to obtain a second matching result; in response to determining that the second matching result indicates failed matching, determining the trailing string corresponding to the truncated text; performing a match on the trailing string to obtain a third matching result; in response to determining that the third matching result indicates failed matching, extracting each phrase from the truncated text; performing a match on each phrase to obtain a fourth matching result; in response to determining that the fourth matching result indicates failed matching, performing a fuzzy match on the truncated text to obtain a fifth matching result; in response to determining that the fifth matching result indicates successful matching, determining the fifth matching result as the matching position corresponding to the truncated text.

[0014] Optionally, generating scrolling information according to the above matching position information includes: determining the distance between the above matching position information and the current display position; generating a scrolling step according to the distance, the time interval since the last adjustment of the teleprompter display content, and the scrolling speed; generating a scrolling position as the scrolling information according to the current display position and the above scrolling step.

[0015] Optionally, the method further includes: determining the user speech rate corresponding to the above user speech; in response to determining that the user speech rate and the above scrolling speed satisfy a first preset speed condition, updating the above scrolling speed according to the above scrolling speed and the preset maximum scrolling speed; in response to determining that the user speech rate and the above scrolling speed satisfy a second preset speed condition, updating the above scrolling speed according to the above scrolling speed and the preset minimum scrolling speed; in response to determining that the pause duration of the user speech satisfies a preset duration condition, updating the above scrolling speed to the preset speed.

[0016] Optionally, the method further includes: generating a next scrolling position as predicted scrolling information based on the average historical scrolling speed, the average historical scrolling acceleration, and the current display position; determining, according to the predicted scrolling information, visible range text information in the target prompt source information corresponding to the predicted scrolling information, where the visible range text information is used to preferentially match the text corresponding to the next user voice.

[0017] In a second aspect, some embodiments of the present disclosure provide a prompter display device, the device includes: an identification unit configured to perform speech recognition on the collected user voice to obtain speech recognition information; a first adjustment unit configured to dynamically adjust the recognized speech recognition information to obtain adjusted text information; a matching unit configured to match the adjusted text information based on target prompt source information to obtain matching position information; a generating unit configured to generate scrolling information according to the matching position information; a second adjustment unit configured to adjust the prompter display content corresponding to the target prompt source information according to the scrolling information.

[0018] Optionally, the identification unit is further configured to: generate noise level information of the user voice; in response to determining that the noise level information meets a preset noise condition, adjust noise reduction window information to obtain updated noise reduction window information; perform noise reduction processing on the user voice according to the updated noise reduction window information to obtain noise-reduced user voice; perform speech recognition on the obtained noise-reduced user voice to obtain speech recognition information.

[0019] Optionally, the second adjustment unit is further configured to: perform dynamic correction processing on the recognized speech recognition information to obtain corrected speech recognition information; add the obtained corrected speech recognition information to a historical record queue; correct the corrected speech recognition information according to the historical record queue to obtain corrected speech recognition information; generate adjusted text information according to the corrected speech recognition information.

[0020] Optionally, the second adjustment unit is further configured to: determine the number of characters included in the corrected speech recognition information; in response to determining that the number of characters meets a preset character condition, perform sentence segmentation processing on the corrected speech recognition information to obtain segmented speech recognition information; perform punctuation optimization processing on the segmented speech recognition information to obtain optimized speech recognition information as the adjusted text information.

[0021] Optionally, the matching unit is further configured to: in response to determining that the adjusted text information satisfies a preset character change condition, determine the adjusted text information as the text information to be matched; perform text truncation processing on the text information to be matched to obtain each truncated text; for each obtained truncated text, perform the following steps: determine whether there is a matching result corresponding to the truncated text in the cache; in response to determining that there is no matching result corresponding to the truncated text in the cache, determine the current visible range text information of the target prompt source information; based on the current visible range text information or the target prompt source information, match the truncated text to obtain a matching position corresponding to the truncated text; generate matching position information according to the obtained matching positions.

[0022] Optionally, the matching unit is further configured to: based on the current visible range text information, match the truncated text to obtain a matching result; in response to determining that the matching result indicates a successful match, determine the matching result as the matching position; in response to determining that the matching result indicates a failed match, based on the target prompt source information, match the truncated text to obtain a matching result as the matching position.

[0023] Optionally, the matching unit is further configured to: perform an exact match on the truncated text to obtain a first matching result; in response to determining that the first matching result indicates a failed match, determine the longest common substring corresponding to the truncated text in the current visible range text information or the target prompt source information; match the longest common substring to obtain a second matching result; in response to determining that the second matching result indicates a failed match, determine the trailing string corresponding to the truncated text; match the trailing string to obtain a third matching result; in response to determining that the third matching result indicates a failed match, extract each phrase from the truncated text; match each phrase to obtain a fourth matching result; in response to determining that the fourth matching result indicates a failed match, perform a fuzzy match on the truncated text to obtain a fifth matching result; in response to determining that the fifth matching result indicates a successful match, determine the fifth matching result as the matching position corresponding to the truncated text.

[0024] Optionally, the generating unit is further configured to: determine the distance between the matching position information and the current display position; generate a scrolling step size according to the distance, the time interval since the last adjustment of the prompt display content, and the scrolling speed; generate a scrolling position as scrolling information according to the current display position and the scrolling step size.

[0025] Optionally, the inscription display device further includes: a determination unit, a first update unit, a second update unit, and a third update unit. Among them, the determination unit is configured to determine the user speech rate corresponding to the above user speech; the first update unit is configured to update the above scrolling speed according to the above scrolling speed and the preset maximum scrolling speed in response to determining that the above user speech rate and the above scrolling speed meet the first preset speed condition; the second update unit is configured to update the above scrolling speed according to the above scrolling speed and the preset minimum scrolling speed in response to determining that the above user speech rate and the above scrolling speed meet the second preset speed condition; the third update unit is configured to update the above scrolling speed to the preset speed in response to determining that the pause duration of the user speech meets the preset duration condition.

[0026] Optionally, the inscription display device further includes: a position generation unit and an information determination unit. Among them, the position generation unit is configured to generate a next scrolling position as predicted scrolling information according to the average historical scrolling speed, the average historical scrolling acceleration, and the current display position. The information determination unit is configured to determine the visible range text information corresponding to the above predicted scrolling information in the above target prompt source information according to the above predicted scrolling information, where the above visible range text information is used to preferentially match the text corresponding to the next user speech.

[0027] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above first aspect.

[0028] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium on which a computer program is stored, where the program, when executed by a processor, implements the method described in any implementation manner of the above first aspect.

[0029] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, and the computer program, when executed by a processor, implements the method described in any implementation manner of the above first aspect.

[0030] The above - mentioned various embodiments of the present disclosure have the following beneficial effects: Through the inscription display method of some embodiments of the present disclosure, the inscription experience and inscription synchronization effect of users during speaking are improved. Specifically, the reasons for the poor user experience and synchronization effect are as follows: The manual method requires distracted operation, resulting in a poor user experience. The timed scrolling method cannot scroll the manuscript according to the actual progress of the speech, resulting in a poor synchronization effect. Based on this, in the inscription display method of some embodiments of the present disclosure, first, the user speech collected is subjected to speech recognition to obtain speech recognition information. Thus, the speech during the user's speech can be recognized. Then, the recognized speech recognition information is dynamically adjusted to obtain adjusted text information. Thus, the speech recognition result can be dynamically adjusted during the recognition process to improve the accuracy of the speech recognition result, facilitating subsequent text matching. Next, based on the target prompt source information, the above - adjusted text information is matched to obtain matching position information. Thus, the matching position of the adjusted text information can be located in the target prompt source information, and the matching position can be understood as the position corresponding to the current speech. Then, according to the above - mentioned matching position information, scrolling information is generated. Thus, the scrolling - related parameters required during the display process can be determined according to the matching position. Finally, according to the above - mentioned scrolling information, the inscription display content corresponding to the above - mentioned target prompt source information is adjusted. Thus, during the user's speech, the inscription display content can be dynamically adjusted according to the user's speech, enabling the inscription display content to follow the speech progress of the user without manual operation by the user, thereby improving the inscription experience and inscription synchronization effect of the user during speaking. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In combination with the accompanying drawings and with reference to the following specific embodiments, the above - mentioned and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0032] Figure 1 is a flowchart of some embodiments of the inscription display method according to the present disclosure;

[0033] Figure 2 is a schematic diagram of an application scenario of the inscription display method according to some embodiments of the present disclosure;

[0034] Figure 3 is a schematic structural diagram of some embodiments of the inscription display device according to the present disclosure;

[0035] Figure 4 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0037] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0038] It should be noted that concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0039] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0040] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0041] Regarding the operations of collecting, storing, using, etc. of the user's personal information (such as user voice, speech recognition information, target prompt source information) involved in the present disclosure, before performing the corresponding operations, relevant organizations or individuals shall fulfill obligations including conducting a personal information security impact assessment, fulfilling the obligation of informing the personal information subject, and obtaining the prior authorization and consent of the personal information subject.

[0042] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.

[0043] Figure 1 Flow 100 of some embodiments of the inscription display method according to the present disclosure is shown. The inscription display method includes the following steps:

[0044] Step 101, perform speech recognition on the collected user voice to obtain speech recognition information.

[0045] In some embodiments, the execution subject of the inscription display method (such as a head-mounted display device or a terminal device) may perform speech recognition on the collected user speech to obtain speech recognition information. Among them, the above-mentioned head-mounted display device may be a display device for imaging in front of the eyes of the wearing user. For example, the above-mentioned head-mounted display device may be, but not limited to: AR glasses, MR glasses, VR glasses. Preferably, the above-mentioned AR glasses may be binocular diffractive waveguide AR glasses. An audio collection device may be provided on the above-mentioned head-mounted display device. The audio collection device may include a microphone. The above-mentioned terminal device may be, but not limited to: a mobile phone, a tablet computer, a computer, a display device for inscription display. An audio collection device may be provided on the above-mentioned terminal device. For example, a user may wear a head-mounted display device during a speech and collect user speech through the audio collection device provided on the head-mounted display device. The user may also use a mobile phone or a tablet computer during a live broadcast and collect user speech through the audio collection device provided on the mobile phone or the tablet computer. The audio collection device may be pre-set with collection parameter information. The collection parameter information may include, but is not limited to, at least one of the following: sampling rate, audio channel, audio format, buffer size. For example, the collection parameter information may be: sampling rate: 16000HZ; audio channel: mono; audio format: PCM 16-bit; buffer size: 160ms. In practice, the above-mentioned execution subject may send the collected user speech to a speech recognition server and obtain the recognition result sent by the speech recognition server as the speech recognition information. For example, the above-mentioned execution subject may send the collected user speech to the speech recognition server through WebSocket real-time communication technology, HTTP polling or gRPC communication technology. In practice, the above-mentioned execution subject may also perform speech recognition using a pre-deployed local offline speech recognition engine. The above-mentioned speech recognition information may be the text of the recognized user speech.

[0046] In some optional implementation manners of some embodiments, the above-mentioned execution subject may perform speech recognition on the collected user speech through the following steps to obtain speech recognition information:

[0047] First, generate noise level information of the above-mentioned user speech. In practice, the above-mentioned execution subject may generate the root mean square of the above-mentioned user speech as the noise level information.

[0048] Second, in response to determining that the above-mentioned noise level information meets a preset noise condition, adjust the noise reduction window information to obtain updated noise reduction window information. Among them, the above-mentioned preset noise condition may be that the noise level information is greater than a preset threshold. The noise reduction window information may include the window size of moving average filtering. For example, the default window size may be 5. In practice, the default window size may be increased by a preset value to obtain the updated noise reduction window information. For example, the preset value may be 3.

[0049] Step 3: According to the updated noise reduction window information above, perform noise reduction processing on the user voice above to obtain a noise-reduced user voice. In practice, the updated noise reduction window information can be used to perform moving average filtering on the user voice above to perform noise reduction processing on the user voice and obtain a noise-reduced user voice.

[0050] Step 4: Perform speech recognition on the obtained noise-reduced user voice to obtain speech recognition information. Here, the method of performing speech recognition can refer to step 101 and will not be elaborated here. Thus, the surrounding human voices and mechanical noises can be better filtered out, improving the speech recognition effect.

[0051] Step 102: Dynamically adjust the recognized speech recognition information to obtain adjusted text information.

[0052] In some embodiments, the above execution entity can dynamically adjust the recognized speech recognition information to obtain adjusted text information. In practice, the above execution entity can perform sentence segmentation on the above speech recognition information to dynamically adjust the speech recognition information and obtain adjusted text information. For example, sentence segmentation can be performed using a bert or GPT model to divide a long text into short texts.

[0053] In some optional implementation manners of some embodiments, the above execution entity can dynamically adjust the recognized speech recognition information through the following steps to obtain adjusted text information:

[0054] Step 1: Perform dynamic correction processing on the recognized speech recognition information to obtain corrected speech recognition information. In practice, first, the above execution entity can perform speech recognition on the above user voice again to obtain comparison speech recognition information. Then, the similarity between the above speech recognition information and the above comparison speech recognition information can be generated. Then, in response to determining that the similarity is less than a preset similarity, the above speech recognition information can be corrected to obtain corrected speech recognition information. For example, the text part corresponding to the confidence level greater than the preset confidence level in the above speech recognition information can be determined first. Then, the above text part can be merged with the text different from the above text part in the above comparison speech recognition information to obtain corrected speech recognition information. For example, when the user says "The weather is nice today", it may be recognized as "The weather is nice today" for the first time and "The weather is really nice today" for the second time. By calculating the similarity of the two recognition results (about 0.7), partial correction is required. Since the confidence of the first recognition is relatively high, the system retains "The weather", and at the same time adopts "really nice today" in the second recognition. The finally obtained corrected speech recognition information is "The weather is really nice today".

[0055] In the second step, add the obtained corrected speech recognition information to the historical record queue. The historical record queue can be a queue with a fixed size. For example, the corresponding fixed size can be 5. When there are more than 5 records, the earliest added record to the queue can be automatically deleted.

[0056] In the third step, based on the above historical record queue, correct the above corrected speech recognition information to obtain corrected speech recognition information. In practice, first, the above execution entity can perform deduplication processing on each piece of corrected speech recognition information in the historical record queue. Then, based on each piece of corrected speech recognition information before the latest added corrected speech recognition information in the historical record queue, correct the latest added corrected speech recognition information to obtain corrected speech recognition information. For example, correction can be performed through a language model. The language model can be GPT, BERT, etc. The language model can detect low-probability words or phrases in the latest added corrected speech recognition information and replace them with candidate words with higher probabilities. For example, the historical record queue may contain "Is everyone ready?" When the user says "Let's go eat", the language model can analyze words such as "everyone" and "we" in the historical record queue that represent groups and infer that the current context may involve a collective activity. Therefore, "Let's go eat" can be optimized to "Let's go eat together" to make the expression more natural and complete. Thus, the accuracy of subsequent recognition can be improved.

[0057] In the fourth step, generate adjusted text information based on the above corrected speech recognition information. In practice, the above execution entity can determine the above corrected speech recognition information as the adjusted text information.

[0058] In some optional implementation manners of some embodiments, the above execution entity can generate adjusted text information based on the above corrected speech recognition information through the following steps:

[0059] In the first step, determine the number of characters included in the above corrected speech recognition information.

[0060] In the second step, in response to determining that the number of characters meets a preset character condition, perform sentence segmentation processing on the above corrected speech recognition information to obtain segmented speech recognition information. The above preset character condition can be that the number of characters is greater than a preset character quantity. For example, the preset character quantity can be 100. In practice, the bert or GPT model can be used for sentence segmentation processing to divide long texts into short texts.

[0061] In the third step, perform punctuation optimization processing on the above segmented speech recognition information to obtain optimized speech recognition information as the adjusted text information. In practice, the bert or GPT model can be used for punctuation optimization processing to optimize the position and type of punctuation. Thus, the pauses and sentence segments in natural speech can be recognized more accurately.

[0062] Step 103: Based on the target prompt source information, match the adjusted text information to obtain the matching position information.

[0063] In some embodiments, the above-mentioned execution entity may match the adjusted text information based on the target prompt source information to obtain the matching position information. Among them, the target prompt source information may be the speech file of the user's current speech. The speech file may be pre-uploaded by the user. For example, the speech file may be a speech script or a product introduction script for live broadcast. In practice, the above-mentioned execution entity may use a fuzzy matching algorithm to identify the line position of the adjusted text information in the target prompt source information as the matching position information.

[0064] In some optional implementation manners of some embodiments, the above-mentioned execution entity may match the adjusted text information based on the target prompt source information through the following steps to obtain the matching position information:

[0065] First step: In response to determining that the adjusted text information meets the preset character change condition, determine the adjusted text information as the text information to be matched. Among them, the above-mentioned preset character change condition may be that the number of characters in the adjusted text information is greater than the preset character change value. For example, the preset character change value may be 4. It should be noted that there may be multiple identical keywords in a piece of adjusted text information, such as "artificial intelligence". If the matching starts within the preset character change value of characters in the adjusted text information, there will be a matching failure, and the text content that the user actually reads will not be found, resulting in repeated jumping of the scroll. When the matching is performed after exceeding the preset character change value of characters, the situation of matching failure can be reduced.

[0066] Second step: Perform text truncation processing on the text information to be matched to obtain each truncated text. In practice, the text information to be matched is truncated through the punctuation marks of the text information to be matched to obtain each truncated text. For example, it can be truncated and divided with each punctuation mark as the boundary.

[0067] Third step: For each obtained truncated text, perform the following steps:

[0068] First sub-step: Determine whether there is a matching result corresponding to the truncated text in the cache. Thus, the historical matching results in the cache can be directly used for subsequent matching to avoid repeated calculations.

[0069] Second sub-step: In response to determining that there is no matching result corresponding to the truncated text in the cache, determine the current visible range text information of the target prompt source information. The current visible range text information may be the text content currently displayed on the screen.

[0070] The third sub-step is to match the truncated text based on the above-mentioned current visible range text information or the above-mentioned target hint source information to obtain the matching position corresponding to the truncated text.

[0071] The fourth step is to generate matching position information according to the obtained respective matching positions. In practice, the above-mentioned execution entity may determine the earliest matching position among the above-mentioned respective matching positions as the matching position information.

[0072] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform the following steps to match the truncated text based on the above-mentioned current visible range text information or the above-mentioned target hint source information to obtain the matching position corresponding to the truncated text:

[0073] The first step is to match the truncated text based on the above-mentioned current visible range text information to obtain a matching result. Thus, the matching can be preferentially performed within the current visible text range.

[0074] The second step is to, in response to determining that the above-mentioned matching result indicates a successful match, determine the above-mentioned matching result as the matching position.

[0075] The third step is to, in response to determining that the above-mentioned matching result indicates a failed match, match the truncated text based on the above-mentioned target hint source information to obtain a matching result as the matching position. Thus, when there is no matching text for the truncated text within the current visible text range, the matching can be performed within a larger range, thereby optimizing the matching strategy and shortening the matching time.

[0076] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform the following steps to match the truncated text:

[0077] The first step is to perform an exact match on the truncated text to obtain a first matching result. In practice, in the above-mentioned current visible range text information or the above-mentioned target hint source information, the Knuth-Morris-Pratt algorithm may be used to perform an exact match on the truncated text to obtain the matching line number as the first matching result. When the exact match fails, a null value or a preset failure value may be determined as the first matching result.

[0078] The second step is to, in response to determining that the above-mentioned first matching result indicates a failed match, determine the longest common substring corresponding to the truncated text in the above-mentioned current visible range text information or the above-mentioned target hint source information. The longest common substring may be the longest common string between the original text and the truncated text (the combination order of the strings is from front to back). In practice, the dynamic programming algorithm may be used to find the longest common substring.

[0079] In the third step, match the above-mentioned longest common substring to obtain a second matching result. In practice, in response to determining that the longest common substring is not empty, the line number of the longest common substring in the above-mentioned target prompt source information can be determined as the second matching result. In response to determining that the longest common substring is empty, a null value or a preset failure value can be determined as the second matching result.

[0080] In the fourth step, in response to determining that the above-mentioned second matching result indicates a matching failure, determine the trailing string corresponding to the above-mentioned truncated text. The trailing string can be the last preset number of characters of the truncated text.

[0081] In the fifth step, match the above-mentioned trailing string to obtain a third matching result. In practice, the above-mentioned trailing string can be matched in the above-mentioned currently visible range text information or the above-mentioned target prompt source information. When the above-mentioned trailing string is matched, the line number of the matched string in the above-mentioned target prompt source information can be determined as the third matching result. When the above-mentioned trailing string is not matched, a null value or a preset failure value can be determined as the third matching result.

[0082] In the sixth step, in response to determining that the above-mentioned third matching result indicates a matching failure, extract each phrase from the above-mentioned truncated text. In practice, the above-mentioned execution entity can extract each phrase with a preset part-of-speech combination from the truncated text. For example, the preset part-of-speech combination can include, but is not limited to: adjective + noun, verb + noun.

[0083] In the seventh step, match the above-mentioned each phrase to obtain a fourth matching result. In practice, the above-mentioned each phrase can be matched in the above-mentioned currently visible range text information or the above-mentioned target prompt source information. When a phrase is matched, the minimum line number of the matched each phrase in the above-mentioned target prompt source information can be determined as the fourth matching result. When no phrase is matched, a null value or a preset failure value can be determined as the fourth matching result.

[0084] In the eighth step, in response to determining that the above-mentioned fourth matching result indicates a matching failure, perform a fuzzy match on the above-mentioned truncated text to obtain a fifth matching result. The fifth matching result can be the line number obtained by fuzzy matching.

[0085] In the ninth step, in response to determining that the above-mentioned fifth matching result indicates a matching success, determine the above-mentioned fifth matching result as the matching position corresponding to the above-mentioned truncated text. Thus, a multi-level matching strategy can be adopted to ensure that the best matching position can be found in various situations.

[0086] Step 104, generate scrolling information according to the matching position information.

[0087] In some embodiments, the above-mentioned execution entity may generate scrolling information based on the above-mentioned matching position information. Among them, the above-mentioned scrolling information may be relevant parameter information required to scroll from the current inscription position to the matching position information. The current inscription position may be represented by the number of lines. In practice, the above-mentioned execution entity may determine the difference between the matching position information and the above-mentioned current inscription position as the interval distance. Then, the sum of the interval distance and the preset adjusted number of lines may be determined as the scrolling information. The preset adjusted number of lines may be used as the buffered number of lines for scrolling, which can make the displayed content of the inscription after scrolling not appear abruptly on the first line of the screen, but appear in the upper half of the screen. For example, the preset adjusted number of lines may be 1.

[0088] In some alternative implementation manners of some embodiments, the above-mentioned execution entity may generate scrolling information based on the above-mentioned matching position information through the following steps:

[0089] First step, determine the distance between the above-mentioned matching position information and the current display position. The current display position may be the number of lines currently displayed on the screen. In practice, the above-mentioned execution entity may determine the pixel distance between the above-mentioned matching position information and the current display position as the distance between the above-mentioned matching position information and the current display position.

[0090] Second step, generate a scrolling step based on the above-mentioned distance, the time interval from the last adjustment of the teleprompter display content, and the scrolling speed. Among them, the above-mentioned scrolling speed may be a basic scrolling speed coefficient, which can be used to adjust the overall scrolling speed. The above-mentioned execution entity may first determine the product of the above-mentioned time interval and the above-mentioned scrolling speed. Then, the minimum value may be taken between the above-mentioned product and a preset ratio. Finally, the product of the above-mentioned distance and the taken minimum value may be determined as the scrolling step.

[0091] Third step, generate a scrolling position as the scrolling information based on the current display position and the above-mentioned scrolling step. In practice, the sum of the pixel abscissa corresponding to the current display position and the above-mentioned scrolling step may be determined as the scrolling information. At this time, the scrolling information may be used as the pixel abscissa of the position to be displayed after scrolling. Thus, smooth scrolling can be achieved, avoiding abrupt jumps.

[0092] Optionally, the above-mentioned execution entity may also perform the following steps:

[0093] First step, determine the user speech rate corresponding to the above-mentioned user speech. Here, the user speech rate may be represented by the average number of words spoken per minute corresponding to the user speech.

[0094] Second step, in response to determining that the user speech rate and the scrolling speed satisfy the first preset speed condition, update the scrolling speed according to the scrolling speed and the preset maximum scrolling speed. The first preset speed condition may be that the user speech rate is greater than the product of the scrolling speed and the first preset coefficient. For example, the first preset coefficient may be 1.5. The preset maximum scrolling speed may be 2.0. In practice, the executing entity may determine the product of the scrolling speed and the second preset coefficient as the first adjusted scrolling speed. The second preset coefficient may be 1.2. Then, the minimum value of the first adjusted scrolling speed and the preset maximum scrolling speed may be determined as the updated scrolling speed to update the scrolling speed.

[0095] Third step, in response to determining that the user speech rate and the scrolling speed satisfy the second preset speed condition, update the scrolling speed according to the scrolling speed and the preset minimum scrolling speed. The first preset speed condition may be that the user speech rate is less than the product of the scrolling speed and the third preset coefficient. For example, the third preset coefficient may be 0.5. The preset minimum scrolling speed may be 0.5. In practice, the executing entity may determine the product of the scrolling speed and the fourth preset coefficient as the second adjusted scrolling speed. The fourth preset coefficient may be 0.8. Then, the maximum value of the second adjusted scrolling speed and the preset minimum scrolling speed may be determined as the updated scrolling speed to update the scrolling speed.

[0096] Fourth step, in response to determining that the pause duration of the user speech satisfies the preset duration condition, update the scrolling speed to the preset speed. The preset duration condition may be that the pause duration is greater than the preset duration. For example, the preset duration may be 1.5 seconds. The preset speed may be 0.5 or 0. Thus, the scrolling speed can be dynamically adjusted to adapt to the change of the speaker's speech rate.

[0097] Step 105, adjust the prompter display content corresponding to the target prompt source information according to the scrolling information.

[0098] In some embodiments, the executing entity may adjust the prompter display content corresponding to the target prompt source information according to the scrolling information. The prompter display content may be the text content in the target prompt source information that needs to prompt the user's speech progress. In practice, the executing entity may scroll the text of the target prompt source information displayed on the screen according to the scrolling information, and make the prompter display content corresponding to the matching position information prominently displayed on the screen.

[0099] As an example, the adjusted prompter display content may be as Figure 2 shown. The matching position information may be the second line in the target prompt source information. The prompter display content of the second line may be "We need to pay attention to the development trend of artificial intelligence", then the text of the second line can be enlarged on the screen for prominent display.

[0100] Optionally, the execution entity may generate the next scrolling position as predicted scrolling information based on the average historical scrolling speed, the average historical scrolling acceleration and the current display position. In practice, the product of the average historical scrolling speed and the preset prediction duration may be first determined as the first product. Then, the product of the average historical scrolling acceleration and the preset coefficient and the square of the preset prediction duration may be determined as the second product. Finally, the sum of the current display position and the first product and the second product may be determined as the next scrolling position, and the next scrolling position is the predicted scrolling information. Then, based on the predicted scrolling information, the visible range text information corresponding to the predicted scrolling information in the target prompt source information may be determined. The visible range text information corresponding to the predicted scrolling information may be the text content in the target prompt source information within the display range of the screen after the screen display content is scrolled according to the predicted scrolling information. The visible range text information may be used to preferentially match the text corresponding to the next user voice.

[0101] Optionally, the execution subject may, in response to detecting the gesture operation of the user corresponding to the user voice, perform image acquisition on the gesture operation to obtain various gesture images. The gesture operation may be an operation in which the user displays a gesture in front of the camera of the execution subject. In practice, the image of the user's hand may be continuously acquired by the camera set in the execution subject to obtain various gesture images. Then, the control type may be identified according to the various gesture images. In practice, the image matching may be performed by matching the various gesture images with the gesture images in the preset standard gesture image group set to identify the matching standard gesture image group, and the control type corresponding to the matching standard gesture image group may be determined as the identified control type. Each standard gesture image group may be an image of various standard gestures corresponding to the same control type. Finally, the prompt display content corresponding to the target prompt source information may be controlled according to the control type. The control logic corresponding to each control type may be preset. For example, the control type corresponding to the five-finger-closed gesture may be to pause the updating and adjustment of the prompt display content. The control type corresponding to the thumb and index finger open gesture may be to enlarge the inscription display content. The control types may include but are not limited to: pause, fast forward, rewind, reduce, enlarge, brighten, and dim. Thus, the user can control the display of the inscription display content through gesture interaction.

[0102] Optionally, the above execution entity may further perform the following steps:

[0103] First step, according to the above matching position information, extract the context text information from the above target prompt source information. In practice, the above execution entity can take the above matching position information as the center and read the text content from the first preset number of lines before to the second preset number of lines after in the above target prompt source information as the extracted context text information.

[0104] Second step, extract the abstract and each keyword from the above context text information as historical context information. In practice, the above execution entity can extract the abstract and each keyword from the above context text information through an abstract extraction algorithm and a keyword extraction algorithm respectively. Then, the extracted abstract and each keyword can be combined into historical context information

[0105] Third step, input the above matching position information, the text embedding vector sequence corresponding to the above target prompt source information, and the above historical context information into the input layer of a pre-trained content prediction model. Among them, the content prediction model can be a neural network model that takes the matching position information, the text embedding vector sequence corresponding to the target prompt source information, and the historical context information as input data and takes the predicted content of the next-time speech as output. The text embedding vector sequence corresponding to the above target prompt source information can be a sequence composed of the embedding vectors of each word. The above content prediction model can include an input layer, a graph construction layer, a graph attention layer, a dynamic memory enhancement layer, and a content prediction layer. The input layer can provide all the input information required by the model, including the matching position information representing the current speech position, the text embedding vector sequence representing the original text content, and the historical context information representing the historical context.

[0106] Fourth step, input the above text embedding vector sequence into the above graph construction layer to obtain graph structure data. The graph construction layer can be used to construct a fully connected graph based on the target prompt source information, where each node represents a word and the edge represents the relationship between words, facilitating the subsequent use of the graph attention mechanism to capture global dependencies. The node feature dimension can be equal to the word embedding dimension (such as 768 dimensions), and the edge weight can be initialized to 1, indicating that initially all words have the same association strength. Specifically, the graph construction layer can first construct a fully connected graph. Each node is connected to all other nodes, and the initial value of the edge weight is set to 1. Then, the text embedding vector can be assigned to the node features. After that, the adjacency matrix can be used to represent the relationship of the edges.

[0107] In the fifth step, input the above graph structure data and the above matching position information into the above graph attention layer to obtain updated node features. The graph attention layer can use a graph attention network (GAT) to process the graph structure data and calculate the importance weights of each node. Using the graph attention mechanism, according to the matching position information representing the current speaking position, dynamically adjust the weights between nodes to highlight the context information related to the current speaking position. The number of attention heads in the graph attention layer can be 8, the number of hidden units in each head can be 128, and the activation function is selected as LeakyReLU.

[0108] In the sixth step, input the above updated node features and the above historical context information into the above dynamic memory enhancement layer to obtain an updated memory state. The dynamic memory enhancement layer can introduce a dynamic memory module to store and update the historical context information, and dynamically update the memory state according to the current speaking position for predicting future content. The number of memory slots in the dynamic memory enhancement layer can be 64, the dimension of each memory slot can be 128, the update rule can adopt a gating mechanism (such as GRU), and finally an output gate can be used to control the reading of the memory state to obtain the updated memory state.

[0109] In the seventh step, input the above updated memory state into the above content prediction layer to obtain the predicted content. The content prediction layer can use a fully connected layer to predict the content that may appear next. When training the content prediction model, cross-entropy loss can be used as the loss function, and the Adam optimizer can be used to optimize the model. Specifically, the content prediction layer can flatten the updated memory state into a one-dimensional vector. Then, two fully connected layers can be used for mapping to obtain the predicted content.

[0110] In the eighth step, construct a query text corresponding to the above predicted content. In practice, keywords can be first extracted from the above predicted content as predicted keywords. Then a preset text template can be used to construct a query text corresponding to the above predicted keywords. For example, the preset text template can be "Please provide the latest research on XXX". Among them, the "XXX" part can be the string position for the keyword to be input. For example, the predicted keyword corresponding to the predicted content can be "climate change policy". The constructed query text can be "Please provide the latest research on climate change policy".

[0111] In the ninth step, obtain a sequence of query contents corresponding to the above query text. In practice, an API can be used to call an external search engine or an internal knowledge base to obtain a sequence of query contents corresponding to the above query text. The sequence of query contents can be the sorting result of the search results according to the relevance score.

[0112] Step 10: For each query content in the above query content sequence, generate summary information corresponding to the above query content. In practice, a text summarization algorithm (such as BERTSUM) can be used to generate summary information corresponding to the above query content.

[0113] Step 11: Obtain each entity information and each relationship information related to the above prediction content from the pre-constructed knowledge graph. The above knowledge graph can be an encyclopedia knowledge graph. In practice, each entity information and each relationship information related to the above prediction keywords can be obtained from the knowledge graph. Entity information can include the entity name and corresponding attributes. For example, for "climate change policy", relevant entities such as "Paris Agreement", "carbon emissions", and "renewable energy" can be found.

[0114] Step 12: Construct a local knowledge graph based on the above each entity information and each relationship information. In practice, the knowledge graph construction method of the graph database can be used to construct the above each entity information and each relationship information into a local knowledge graph corresponding to the prediction keywords.

[0115] Step 13: Extract each entity group in the above local knowledge graph. In practice, each entity group can be extracted through the relationship chain in the local knowledge graph, and each entity group includes each entity. For example, the extracted entity groups can include: climate change policy - Paris Agreement - international cooperation; carbon tax - carbon emissions - energy transition.

[0116] Step 14: Generate each associated topic according to the above each entity group. In practice, for each entity group, at least one topic that matches at least two entities in the above entity group can be determined from a preset topic set to obtain each associated topic. Here, the match can include exact text match and fuzzy text match. For example, each generated associated topic can include: the role of international cooperation in addressing climate change; the impact of carbon tax policies on economic development; the latest progress of renewable energy technologies. The preset topic set can be stored in a search engine.

[0117] Step 15: Sort the above each associated topic according to the above prediction content to obtain an associated topic sequence. In practice, the each associated topic can be sorted in descending order according to the similarity between the associated topic and the prediction keywords to obtain an associated topic sequence.

[0118] Step 16: For each associated topic in the above-mentioned associated topic sequence, generate prompt content corresponding to the above-mentioned associated topic as enhanced prompt content. In practice, the above-mentioned execution entity can use an API to call an external search engine or an internal knowledge base to obtain search content corresponding to the above-mentioned associated topic. Then, an abstract can be extracted from the above-mentioned search content as the prompt content corresponding to the above-mentioned associated topic, that is, the enhanced prompt content. Thus, the enhanced prompt content can be used as a further expansion of the direct query content to provide the user with richer prompt content.

[0119] Step 17: Display each generated abstract information and each enhanced prompt content. In practice, each of the above-mentioned abstract information and each of the above-mentioned enhanced prompt content can be displayed on one side of the current inscription display content.

[0120] The above Steps 1-17 are an inventive point of the embodiment of the present disclosure, which solves the technical problem of "traditional inscription display can only passively display the content of the speech draft. When it is necessary to explain content other than the speech draft, it cannot actively provide relevant background information or prompts for the speaker, resulting in the speaker having to prepare relevant content by himself or rely entirely on improvisation, resulting in a poor user experience". The factors that lead to a poor user experience are often as follows: traditional inscription display can only passively display the content of the speech draft. When it is necessary to explain content other than the speech draft, it cannot actively provide relevant background information or prompts for the speaker, resulting in the speaker having to prepare relevant content by himself or rely entirely on improvisation. If the above factors are solved, the effect of improving the user experience can be achieved. To achieve this effect, the present disclosure introduces advanced natural language processing technology, which can capture the global dependency relationships and semantic information in the original speech, combined with dynamic context modeling, and dynamically adjust the context weights according to the current speech position, and can more accurately predict the next content. Then, according to the predicted next content, relevant information can be automatically retrieved from external resources (such as knowledge graphs, search engines), and concise prompt content and enhanced prompt content can be generated. This dynamic loading function significantly improves the user's preparation efficiency, reduces the cognitive burden, can meet the user's personalized inscription needs, and improves the user experience.

[0121] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the inscription display method of some embodiments of the present disclosure, the inscription experience and inscription synchronization effect of users during speaking are improved. Specifically, the reasons for the poor user experience and synchronization effect are as follows: The manual method requires distracted operation, resulting in a poor user experience. The timed scrolling method cannot scroll the manuscript according to the actual progress of the speech, resulting in a poor synchronization effect. Based on this, in the inscription display method of some embodiments of the present disclosure, first, the collected user speech is subjected to speech recognition to obtain speech recognition information. Thus, the speech during the user's speech can be recognized. Then, the recognized speech recognition information is dynamically adjusted to obtain adjusted text information. Thus, the speech recognition result can be dynamically adjusted during the recognition process to improve the accuracy of the speech recognition result, facilitating subsequent text matching. Next, based on the target prompt source information, the above-mentioned adjusted text information is matched to obtain matching position information. Thus, the matching position of the adjusted text information can be located in the target prompt source information, and the matching position can be understood as the position corresponding to the current speech. Then, according to the above-mentioned matching position information, scrolling information is generated. Thus, the scrolling-related parameters required during the display can be determined according to the matching position. Finally, according to the above-mentioned scrolling information, the inscription display content corresponding to the above-mentioned target prompt source information is adjusted. Thus, during the user's speech, the inscription display content can be dynamically adjusted according to the user's speech, so that the inscription display content can follow the speech progress of the user without manual operation by the user, thereby improving the inscription experience and inscription synchronization effect of the user during speaking.

[0122] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an inscription display device. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0123] As Figure 3 shown, the inscription display device 300 of some embodiments includes: an identification unit 301, a first adjustment unit 302, a matching unit 303, a generation unit 304, and a second adjustment unit 305. Among them, the identification unit 301 is configured to perform speech recognition on the collected user speech to obtain speech recognition information; the second adjustment unit 302 is configured to dynamically adjust the recognized speech recognition information to obtain adjusted text information; the matching unit 303 is configured to match the above-mentioned adjusted text information based on the target prompt source information to obtain matching position information; the generation unit 304 is configured to generate scrolling information according to the above-mentioned matching position information; the second adjustment unit 305 is configured to adjust the inscription display content corresponding to the above-mentioned target prompt source information according to the above-mentioned scrolling information.

[0124] Optionally, the recognition unit 301 may further be configured to: generate noise level information of the above user voice; in response to determining that the noise level information meets a preset noise condition, adjust the noise reduction window information to obtain updated noise reduction window information; perform noise reduction processing on the above user voice according to the updated noise reduction window information to obtain noise-reduced user voice; perform speech recognition on the obtained noise-reduced user voice to obtain speech recognition information.

[0125] Optionally, the second adjustment unit 302 may further be configured to: perform dynamic correction processing on the recognized speech recognition information to obtain corrected speech recognition information; add the obtained corrected speech recognition information to the historical record queue; perform calibration on the corrected speech recognition information according to the historical record queue to obtain calibrated speech recognition information; generate adjusted text information according to the calibrated speech recognition information.

[0126] Optionally, the second adjustment unit 302 may further be configured to: determine the number of characters included in the calibrated speech recognition information; in response to determining that the number of characters meets a preset character condition, perform sentence segmentation processing on the calibrated speech recognition information to obtain segmented speech recognition information; perform punctuation optimization processing on the segmented speech recognition information to obtain optimized speech recognition information as the adjusted text information.

[0127] Optionally, the matching unit 303 may further be configured to: in response to determining that the adjusted text information meets a preset character change condition, determine the adjusted text information as the text information to be matched; perform text truncation processing on the text information to be matched to obtain each truncated text; for each obtained truncated text, perform the following steps: determine whether there is a matching result corresponding to the truncated text in the cache; in response to determining that there is no matching result corresponding to the truncated text in the cache, determine the current visible range text information of the target prompt source information; based on the current visible range text information or the target prompt source information, match the truncated text to obtain the matching position corresponding to the truncated text; generate matching position information according to the obtained respective matching positions.

[0128] Optionally, the matching unit 303 may further be configured to: match the truncated text based on the current visible range text information to obtain a matching result; in response to determining that the matching result indicates successful matching, determine the matching result as the matching position; in response to determining that the matching result indicates failed matching, match the truncated text based on the target prompt source information to obtain a matching result as the matching position.

[0129] Optionally, the matching unit 303 may be further configured to: perform an exact match on the truncated text to obtain a first matching result; in response to determining that the first matching result indicates a matching failure, determine the longest common substring corresponding to the truncated text in the current visible range text information or the target prompt source information; perform a match on the longest common substring to obtain a second matching result; in response to determining that the second matching result indicates a matching failure, determine the tail string corresponding to the truncated text; perform a match on the tail string to obtain a third matching result; in response to determining that the third matching result indicates a matching failure, extract each phrase from the truncated text; perform a match on each phrase to obtain a fourth matching result; in response to determining that the fourth matching result indicates a matching failure, perform a fuzzy match on the truncated text to obtain a fifth matching result; in response to determining that the fifth matching result indicates a successful match, determine the fifth matching result as the matching position corresponding to the truncated text.

[0130] Optionally, the generating unit 304 may be further configured to: determine the distance between the matching position information and the current display position; generate a scrolling step length according to the distance, the time interval from the last adjustment of the prompter display content, and the scrolling speed; generate a scrolling position as scrolling information according to the current display position and the scrolling step length.

[0131] Optionally, the prompter display device 300 may further include: a determining unit, a first updating unit, a second updating unit, and a third updating unit (not shown in the figure). Wherein, the determining unit is configured to determine the user speech rate corresponding to the user speech; the first updating unit is configured to, in response to determining that the user speech rate and the scrolling speed satisfy a first preset speed condition, update the scrolling speed according to the scrolling speed and the preset maximum scrolling speed; the second updating unit is configured to, in response to determining that the user speech rate and the scrolling speed satisfy a second preset speed condition, update the scrolling speed according to the scrolling speed and the preset minimum scrolling speed; the third updating unit is configured to, in response to determining that the pause duration of the user speech satisfies a preset duration condition, update the scrolling speed to a preset speed.

[0132] Optionally, the prompter display device 300 may further include: a position generating unit and an information determining unit (not shown in the figure). Wherein, the position generating unit is configured to generate a next scrolling position as predicted scrolling information according to the average historical scrolling speed, the average historical scrolling acceleration, and the current display position. The information determining unit is configured to determine the visible range text information corresponding to the predicted scrolling information in the target prompt source information according to the predicted scrolling information, where the visible range text information is used to preferentially match the text corresponding to the next user speech.

[0133] It can be understood that the various units described in the inscription display device 300 correspond to the respective steps in the method described in the reference Figure 1 description. Thus, the operations, features, and beneficial effects described above for the method also apply to the inscription display device 300 and the units included therein, and will not be elaborated here.

[0134] Reference is made below to Figure 4 , which shows a schematic structural diagram of an electronic device 400 (such as a head-mounted display device or a terminal device) suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0135] As Figure 4 shown, the electronic device 400 may include a processing device 401 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0136] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device 400 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 4 Each block shown in

[0137] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.

[0138] It should be noted that the computer-readable medium described in some embodiments of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0139] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0140] The computer-readable medium described above can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: perform speech recognition on the collected user speech to obtain speech recognition information; dynamically adjust the recognized speech recognition information to obtain adjusted text information; based on the target prompt source information, match the adjusted text information to obtain matching position information; generate scrolling information according to the matching position information; and adjust the prompter display content corresponding to the target prompt source information according to the scrolling information.

[0141] Computer program code for performing the operations of some embodiments of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0143] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an identification unit, a first adjustment unit, a matching unit, a generation unit, and a second adjustment unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the identification unit can also be described as "a unit that performs speech recognition on the collected user speech to obtain speech recognition information".

[0144] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0145] Some embodiments of the present disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any one of the above-mentioned caption display methods.

[0146] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A prompt display method, comprising: Performing voice recognition on the collected user voice to obtain voice recognition information; Dynamically adjusting the recognized speech recognition information to obtain adjusted text information; Based on the target prompt source information, matching the adjusted text information to obtain matching position information; Generate scroll information according to the matching position information; According to the scrolling information, the prompt display content corresponding to the target prompt source information is adjusted.

2. The method according to claim 1, wherein: The performing speech recognition on the collected user speech to obtain speech recognition information includes: generating noise level information of the user's speech; In response to determining that the noise level information satisfies a preset noise condition, adjusting the noise reduction window information to obtain updated noise reduction window information; According to the updated noise reduction window information, the user voice is subjected to noise reduction processing to obtain noise-reduced user voice; Perform speech recognition on the obtained noise-reduced user speech to obtain speech recognition information.

3. The method according to claim 1, wherein: The dynamically adjusting the recognized speech recognition information to obtain adjusted text information includes: Dynamically correcting the recognized speech recognition information to obtain corrected speech recognition information; Add the obtained corrected speech recognition information to the history record queue; Correcting the corrected speech recognition information according to the historical record queue to obtain corrected speech recognition information; According to the corrected speech recognition information, adjusted text information is generated.

4. The method according to claim 3, wherein: Generating adjusted text information according to the corrected speech recognition information includes: Determining the number of characters included in the corrected speech recognition information; In response to determining that the number of characters meets a preset character condition, segmenting the corrected speech recognition information to obtain segmented speech recognition information; Punctuation optimization processing is performed on the sentence segmentation speech recognition information to obtain optimized speech recognition information as adjusted text information.

5. The method according to claim 1, wherein: The step of matching the adjusted text information based on the target prompt source information to obtain matching position information includes: In response to determining that the adjusted text information satisfies a preset character change condition, determining the adjusted text information as text information to be matched; Performing text truncation processing on the text information to be matched to obtain various truncated texts; For each resulting truncated text, perform the following steps: Determine whether there is a matching result corresponding to the truncated text in the cache; In response to determining that there is no matching result corresponding to the truncated text in the cache, determining the current visible range text information of the target prompt source information; Based on the current visible range text information or the target prompt source information, the truncated text is matched to obtain a matching position corresponding to the truncated text; According to each obtained matching position, matching position information is generated.

6. The method according to claim 5, wherein: The matching of the truncated text based on the current visible range text information or the target prompt source information to obtain a matching position corresponding to the truncated text includes: Based on the text information of the current visible range, the truncated text is matched to obtain a matching result; In response to determining that the matching result indicates a successful match, determining the matching result as a matching position; In response to determining that the matching result indicates a matching failure, the truncated text is matched based on the target prompt source information to obtain a matching result as a matching position.

7. The method according to claim 5, wherein: The matching of the truncated text includes: Performing an exact match on the truncated text to obtain a first matching result; In response to determining that the first matching result indicates a matching failure, determining a longest common substring corresponding to the truncated text in the current visible range text information or the target prompt source information; Matching the longest common substring to obtain a second matching result; In response to determining that the second matching result indicates a matching failure, determining a tail character string corresponding to the truncated text; Matching the tail character string to obtain a third matching result; In response to determining that the third matching result indicates a matching failure, extracting individual phrases from the truncated text; Matching each of the phrases to obtain a fourth matching result; In response to determining that the fourth matching result indicates a matching failure, performing fuzzy matching on the truncated text to obtain a fifth matching result; In response to determining that the fifth matching result indicates a successful match, the fifth matching result is determined as a matching position corresponding to the truncated text.

8. The method according to claim 1, wherein: The step of generating scroll information according to the matching position information includes: Determining the distance between the matching position information and the current display position; Generate a scrolling step length according to the distance, the time interval between the last adjustment of the prompt display content and the scrolling speed; A scroll position is generated as scroll information according to the current display position and the scroll step length.

9. The method according to claim 8, wherein: The method further comprises: Determining a user speech rate corresponding to the user speech; In response to determining that the user speech speed and the scrolling speed meet a first preset speed condition, updating the scrolling speed according to the scrolling speed and a preset maximum scrolling speed; In response to determining that the user speech speed and the scrolling speed meet a second preset speed condition, updating the scrolling speed according to the scrolling speed and a preset minimum scrolling speed; In response to determining that the pause duration of the user's voice meets a preset duration condition, the scrolling speed is updated to a preset speed.

10. The method according to claim 1, wherein: The method further comprises: Generate a next scroll position as predicted scroll information according to the average historical scroll speed, the average historical scroll acceleration and the current display position; According to the predicted scrolling information, visible range text information corresponding to the predicted scrolling information in the target prompt source information is determined, wherein the visible range text information is used for preferentially matching text corresponding to the next user voice.

11. A prompt display device, comprising: A recognition unit, configured to perform speech recognition on the collected user speech to obtain speech recognition information; A first adjustment unit is configured to dynamically adjust the recognized speech recognition information to obtain adjusted text information; a matching unit configured to match the adjusted text information based on the target prompt source information to obtain matching position information; a generating unit, configured to generate scrolling information according to the matching position information; The second adjustment unit is configured to adjust the prompt display content corresponding to the target prompt source information according to the scrolling information.

12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.

13. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method, system and device for automatically scrolling subtitles based on voice rhythm

    CN112887779A

  • Intelligent word prompting method and device

    CN114999475A

  • Adaptive prompting method for traffic information broadcasting

    CN117874165A

  • Speech displaying system and method

    TW200611174A