Display control method of word mentioning content, terminal equipment and storage medium

By employing a dynamic context modulation semantic matching strategy in the terminal device, and combining the contextual similarity of candidate text blocks, the most similar text block is selected from multiple text blocks as the matching result. This solves the accuracy problem of teleprompters in cases of multiple text similarities or speech ambiguity, and achieves high-precision and high-robustness teleprompter content display.

CN121189334APending Publication Date: 2025-12-23ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511045268.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing teleprompters on terminal devices cannot reliably and accurately match electronic document content in cases of multiple text similarities or unclear speech, resulting in inaccurate teleprompter content and a poor user experience.

Method used

A semantic matching strategy based on dynamic context modulation is adopted to obtain the K candidate text blocks that are most semantically similar to the target audio segment from multiple pre-stored text blocks. The similarity of the candidate text blocks with the context text blocks is combined to determine the target similarity, and the candidate text block with the highest similarity is selected as the matching result.

Benefits of technology

It improves the accuracy and robustness of the prompting content, ensuring the precision and real-time nature of the text content that matches the target person's voice data in the electronic document.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189334A_ABST
    Figure CN121189334A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a prompt content display control method, terminal equipment and a storage medium. The method comprises the following steps: acquiring a target audio clip; obtaining K text blocks most similar to the semantics of the target audio clip from a plurality of pre-stored text blocks as K candidate text blocks; determining context text blocks of the candidate text blocks, and determining a context consistency coefficient according to the similarity between the target audio clip and each context text block and the similarity between the candidate text block and each context text block; according to the similarity between the target audio clip and the candidate text blocks and the context consistency coefficient, determining the target similarity between the target audio clip and the candidate text blocks, and determining the candidate text block corresponding to the highest target similarity in the K candidate text blocks as a target text block; and switching the displayed prompt content into the manuscript content related to the target text block in the electronic manuscript. And the accuracy of the displayed word mentioning content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of display control technology, and in particular to a method for displaying prompt content, a terminal device, and a storage medium. Background Technology

[0002] Currently, terminal devices have teleprompter functionality. When this function is enabled, the device uses an Automatic Speech Recognition (ASR) module to convert the target person's speech data into text. The recognized text is then matched against pre-stored electronic documents to identify the text content that matches the target person's speech data. The associated contextual content within the electronic document is then displayed as the teleprompter. However, string matching algorithms typically only achieve local matching and cannot effectively utilize contextual information. When multiple texts are similar or the speech is unclear, the string matching algorithm cannot reliably and accurately identify the text content that matches the target person's speech data within the electronic document, resulting in inaccurate teleprompter content and a poor user experience. Therefore, improving the accuracy of the displayed teleprompter content is a pressing issue that needs to be addressed. Summary of the Invention

[0003] This invention provides a method for controlling the display of prompt content, a terminal device, and a storage medium, aiming to improve the accuracy of the displayed prompt content.

[0004] In a first aspect, embodiments of the present invention provide a method for controlling the display of prompting content, comprising:

[0005] Acquire a target audio segment, wherein the target audio segment includes the voice data of the target person;

[0006] K text blocks that are most semantically similar to the target audio segment are obtained from a plurality of pre-stored text blocks as K candidate text blocks. The plurality of text blocks are obtained by pre-segmenting the electronic document, and K is an integer greater than or equal to 1.

[0007] For each candidate text block, the context text block of the candidate text block is determined, and the context consistency coefficient is determined based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block;

[0008] Based on the similarity between the target audio segment and the candidate text block and the context consistency coefficient, the target similarity between the target audio segment and the candidate text block is determined, and the candidate text block with the highest target similarity among the K candidate text blocks is determined as the target text block;

[0009] The displayed prompt will be switched to the text content in the electronic document that is related to the target text block.

[0010] In a second aspect, embodiments of the present invention also provide a terminal device, the terminal device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the display control method as described in the first aspect.

[0011] Thirdly, embodiments of the present invention also provide a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the display control method as described in the first aspect.

[0012] This invention provides a method for controlling the display of prompting content, a terminal device, and a storage medium. When determining the text content in an electronic document that matches the voice data of a target person, this invention does not simply select the text block with the highest similarity to the target audio segment as the final matched target text block. Instead, it performs matching based on a dynamic context modulation semantic matching strategy: first, it obtains K text blocks that are semantically most similar to the target audio segment from a pre-stored set of text blocks as K candidate text blocks. Then, for each candidate text block, it performs a more refined comprehensive matching that combines the context text blocks of the candidate text block, resulting in a more accurate and comprehensive matching of the calculated target audio segment. The target similarity between the target audio segment and the candidate text block not only integrates the similarity between the target audio segment and the candidate text block, but also integrates the similarity between the target audio segment and the candidate text block's context text block, as well as the similarity between the candidate text block and its context text block. This ensures the accuracy of the target similarity between the target audio segment and the candidate text block. In this way, the candidate text block with the highest target similarity among K candidate text blocks is determined as the target text block, achieving high-precision and high-robust matching between audio segments and text content. This improves the accuracy of determining the text content that matches the voice data of the target person in electronic documents, thereby improving the accuracy of the displayed prompt content. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1This is a schematic diagram of a scenario for implementing the prompting content display control method provided in the embodiments of the present invention;

[0015] Figure 2 This is a schematic diagram of another scenario for implementing the prompting content display control method provided in the embodiments of the present invention;

[0016] Figure 3 This is a flowchart illustrating a method for controlling the display of prompting content provided in an embodiment of the present invention;

[0017] Figure 4 yes Figure 3 A flowchart illustrating a sub-step of the method for controlling the display of teleprompter content;

[0018] Figure 5 This is a flowchart illustrating another method for controlling the display of prompting content provided in an embodiment of the present invention;

[0019] Figure 6 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0022] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0023] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0024] Currently, terminal devices have teleprompter functionality. When this function is enabled, the device uses an Automatic Speech Recognition (ASR) module to convert the target person's speech data into text. The recognized text is then matched against pre-stored electronic documents to identify the text content that matches the target person's speech data. The associated contextual content within the electronic document is then displayed as the teleprompter. However, string matching algorithms typically only achieve localized matching and cannot effectively utilize contextual information. When multiple texts are similar or the speech is unclear, the string matching algorithm cannot reliably and accurately identify the text content that matches the target person's speech data within the electronic document, resulting in inaccurate teleprompter content and a poor user experience.

[0025] To address the aforementioned problems, embodiments of the present invention provide a method for controlling the display of prompting content, a terminal device, and a storage medium. When determining the text content in an electronic document that matches the voice data of a target person, the present invention does not simply select the text block with the highest similarity to the target audio segment as the final matched target text block. Instead, it first obtains K text blocks that are semantically most similar to the target audio segment from a pre-stored pool of text blocks as K candidate text blocks. Then, for each candidate text block, a more refined comprehensive matching is performed, incorporating the context text blocks of the candidate text block, so that the calculated target audio segment matches the candidate text blocks. The target similarity not only integrates the similarity between the target audio segment and the candidate text block, but also the similarity between the target audio segment and the candidate text block's context text block, as well as the similarity between the candidate text block and its context text block. This ensures the accuracy of the target similarity between the target audio segment and the candidate text block. In this way, the candidate text block with the highest target similarity among K candidate text blocks is determined as the target text block. This achieves high-precision and high-robustness matching of audio segments and text content, improves the accuracy of determining the text content that matches the voice data of the target person in electronic documents, and thus improves the accuracy of the displayed prompting content.

[0026] In some embodiments, the method for controlling the display of prompting content provided in this invention can be applied to a terminal device, which may include a mobile terminal and a head-mounted display device, etc. The mobile terminal may include a mobile phone, tablet computer, laptop computer, and desktop computer, etc., and the head-mounted display device may include augmented reality (AR) glasses, AR helmets, mixed reality (MR) glasses, and MR helmets, etc. For example, the method for controlling the display of prompting content provided in this invention is applied to a head-mounted display device. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of a scenario for implementing the prompting content display control method provided in the embodiments of the present invention.

[0027] like Figure 1 As shown, the head-mounted display device 100 is worn on the head 11 of the target person. The head-mounted display device 100 includes a sound pickup device, a control device, and a display device. The display device includes an optical waveguide lens 101 and an optical engine (OEM). Figure 1 (Not shown). The sound pickup device is used to collect the voice data of the target person, and the control device is used to perform the following steps: acquiring the voice data of the target person and processing the voice data to obtain a target audio segment; obtaining K text blocks that are semantically most similar to the target audio segment from a pre-stored set of text blocks as K candidate text blocks; for each candidate text block, determining the context text block of the candidate text block, and determining the context consistency coefficient based on the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block; determining the target similarity between the target audio segment and the candidate text block based on the similarity between the target audio segment and the candidate text block and the context consistency coefficient, and determining the candidate text block with the highest target similarity among the K candidate text blocks as the target text block; controlling the display device to switch the displayed prompting content to the text content related to the target text block in the electronic document. For example, an optical engine (e.g., a microdisplay) is used to guide the prompting content, etc., onto a waveguide lens, so that the waveguide lens 101 guides the prompting content to the target person's eyes.

[0028] In some embodiments, the head-mounted display device 100 includes sensors for acquiring attitude data of the head-mounted display device 100. Figure 1 (Not shown). For example, the sensor that acquires attitude data of the head-mounted display device 100 may include an inertial measurement unit (IMU), which may include an accelerometer, a gyroscope, and / or a magnetometer. For example, the inertial measurement unit includes a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer.

[0029] In some embodiments, the method for controlling the display of prompting content provided in this invention can also be applied to a system composed of a head-mounted display device and a mobile terminal. For example, the method for controlling the display of prompting content provided in this invention is applied to a system composed of a head-mounted display device and a smartphone. Please refer to... Figure 2 , Figure 2 This is a schematic diagram of another scenario for implementing the prompting content display control method provided in the embodiments of the present invention. For example... Figure 2As shown, the head-mounted display device 100 is communicatively connected to the smartphone 200. The head-mounted display device 100 includes a microphone, a first wireless communication module, and a display device. The display device includes an optical waveguide lens 101 and an optical engine (OEM). Figure 1 (Not shown), the smartphone 200 includes a second wireless communication module. The first and second wireless communication modules may include Bluetooth or WiFi communication modules, etc.

[0030] For example, a head-mounted display device 100 is worn on the head 11 of a target person. The head-mounted display device 100 is used to perform the following steps: acquiring the target person's voice data through a sound pickup device and sending the voice data to a smartphone 200; the smartphone 200 is used to perform the following steps: receiving the voice data sent by the head-mounted display device 100, processing the voice data to obtain a target audio segment; obtaining K text blocks that are semantically most similar to the target audio segment from a pre-stored plurality of text blocks as K candidate text blocks; for each candidate text block, determining the context text block of the candidate text block, and determining the context consistency coefficient based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block; and based on the target audio segment... The similarity and contextual consistency coefficient between the target audio segment and the candidate text blocks are used to determine the target similarity between the target audio segment and the candidate text blocks. The candidate text block with the highest target similarity among the K candidate text blocks is determined as the target text block. This target text block is sent to the head-mounted display device 100. When the head-mounted display device 100 receives the target text block, it controls its own display device to switch the displayed prompt content to the text content related to the target text in the electronic document. Alternatively, the smartphone 200 sends the text content related to the target text in the electronic document to the head-mounted display device 100. When the head-mounted display device 100 receives the text content related to the target text, it controls its own display device to switch the displayed prompt content to the text content related to the target text in the electronic document. In this embodiment, the steps of the prompt content display control method that require high computing power are performed by a mobile terminal with high computing power, while the head-mounted display device only performs the two steps of voice data acquisition and prompt content switching that require low computing power. This can further reduce the synchronization delay between the prompt content and the voice of the target person and improve the real-time performance of the displayed prompt content.

[0031] In some embodiments, the method for controlling the display of prompting content provided in this invention can also be applied to a system consisting of a head-mounted display device and a server. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0032] For example, the head-mounted display device is used to perform the following steps: acquiring voice data of a target person through a sound pickup device and sending the voice data to a server; the server is used to perform the following steps: receiving the voice data sent by the head-mounted display device, processing the voice data to obtain a target audio segment; obtaining K text blocks that are semantically most similar to the target audio segment from a pre-stored set of text blocks as K candidate text blocks; for each candidate text block, determining the context text block of the candidate text block, and determining the context consistency coefficient based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block; and determining the context consistency coefficient based on the similarity between the target audio segment and the candidate text block. The similarity between the target audio segment and the candidate text block is determined by the similarity coefficient and the contextual consistency coefficient. The candidate text block with the highest similarity among the K candidate text blocks is identified as the target text block. This target text block is then sent to the head-mounted display device. Upon receiving the target text block, the head-mounted display device controls its own display device to switch the displayed prompt content to the text content related to the target text in the electronic document. Alternatively, the server sends the text content related to the target text in the electronic document to the head-mounted display device. Upon receiving the text content related to the target text, the head-mounted display device controls its own display device to switch the displayed prompt content to the text content related to the target text in the electronic document. In this embodiment, the computationally intensive steps of the prompt content display control method are performed by a server with high computing power, while the head-mounted display device only performs the two computationally intensive steps of voice data acquisition and prompt content switching. This further reduces the synchronization delay between the prompt content and the target person's voice, improving the real-time performance of the displayed prompt content.

[0033] The following will combine Figure 1 or Figure 2 The following describes in detail the method for controlling the display of prompting content provided by embodiments of the present invention. It should be noted that... Figure 1 or Figure 2 The scenarios described are only used to explain the method for controlling the display of prompting content provided in the embodiments of the present invention, but do not constitute a limitation on the application scenarios of the method for controlling the display of prompting content provided in the embodiments of the present invention.

[0034] Please see Figure 3 , Figure 3 This is a flowchart illustrating a method for controlling the display of prompt content provided in an embodiment of the present invention.

[0035] like Figure 3 As shown, the method for controlling the display of the prompt content includes steps S101 to S105.

[0036] Step S101: Obtain the target audio segment.

[0037] In this embodiment, the target audio segment includes the voice data of the target person. The target person can be a speaker, trainer, singer, host, etc. The target audio segment is obtained by segmenting continuous audio data collected in real time by the sound pickup device.

[0038] In some embodiments, before step S101, the method further includes: acquiring an electronic document; dividing the electronic document into multiple text blocks, the length of each text block being within a preset length range; and storing the multiple text blocks in a database. The preset length range can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. For example, the preset length range is 23 to 33; when the text block is in Chinese, the length of the text block is the number of characters in the text block; when the text block is not in Chinese, the length of the text block is the number of words in the text block. This embodiment pre-divides the electronic document into multiple text blocks and stores them, facilitating the rapid acquisition of text blocks and matching them with the target audio segment during the subsequent word prompting stage. This eliminates the need to spend additional time and computing resources dividing the electronic document, thus improving matching efficiency.

[0039] In some embodiments, segmenting an electronic document into multiple text blocks may include: cleaning the electronic document to obtain target text; splitting the target text into multiple sentences based on sentence termination characters; for each sentence, if the sentence length is within a preset length range, treating the sentence as a text block; or if the sentence length is not within the preset length range, dividing the sentence into multiple sub-sentences based on conjunctions in the sentence, and treating each sub-sentence as a text block, with the length of the multiple sub-sentences within the preset length range. The preset length range can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. For example, the preset length range is 23 to 33 characters or words. When the text block is in Chinese, the length of the text block is the number of characters in the text block; when the text block is not in Chinese, the length of the text block is the number of words in the text block. This embodiment can segment an electronic document into multiple text blocks with lengths within a set range, facilitating subsequent matching with audio segments and switching of prompting content.

[0040] For example, receiving user-uploaded electronic documents, removing irrelevant formatting tags (such as HTML) and special control characters, and uniformly encoding them (such as UTF-8) yields a clean target text. This clean target text is then segmented into a series of semantically relatively complete text chunks. Specifically, a preset length range of 23-33 words is determined, with a minimum of 23 and a maximum of 33 words. Sentence-ending characters (such as periods, question marks, and exclamation marks) are used to split the target text into multiple sentences. For each segmented sentence, if its length falls between 23 and 33 words, it is treated as a single text chunk. If the sentence length does not fall between 23 and 33 words, dependency parsing using spaCy is employed to identify coordinating conjunctions (such as And and But) or subordinating conjunctions (such as When and Before) as segmentation points. If, after using the above two methods, there are still a very small number of cases where the length does not fall between 23 and 33 words, forced segmentation based on 28 words is performed.

[0041] In some embodiments, segmenting an electronic document into multiple text blocks may include: cleaning the electronic document to obtain the target text, and splitting the target text into multiple text blocks based on sentence termination characters. This embodiment omits the step of finding conjunctions as segmentation points based on dependency parsing to segment the electronic document, reducing computational costs and enabling rapid segmentation of electronic documents even in resource-constrained head-mounted display devices. Furthermore, it does not affect the accuracy of subsequent audio segmentation matching with text blocks, ensuring the accuracy of the displayed prompts.

[0042] It is understandable that audio segments are obtained by segmenting continuous audio data obtained from the speech of a target person captured by a microphone. The segmentation quality of audio segments directly and profoundly affects the stability and accuracy of matching subsequent audio segments with text blocks in an electronic document. Related technologies primarily employ audio segmentation methods based on Voice Activity Detection (VAD) to segment continuous audio data obtained from the speech of a target person captured by a microphone.

[0043] However, VAD segmentation, relying solely on the presence or absence of sound, is prone to severing a complete semantic unit (such as a long sentence) during natural thinking or short pauses. This results in fragmented audio segments, making it difficult to ensure that subsequent matching segments contain stable and predictable semantic information. This affects the accuracy of the citations, and common filler words like "um" and "ah" are often recognized as valid speech by VAD, wasting valuable computational resources and resulting in low processing efficiency. Furthermore, different target speakers have different speaking styles; some speak quickly, while others speak slowly. VAD-based audio segmentation methods struggle to adapt to these varying styles to obtain high-quality segments. This leads to poor semantic matching accuracy between audio segments and text blocks in some scenarios. For example, for slow-speaking targets, VAD-determined audio segments contain less semantic information, resulting in poor semantic matching accuracy between subsequent audio segments and text blocks.

[0044] To address the aforementioned problems, embodiments of the present invention provide an audio segmentation strategy based on streaming semantic token counting, such as... Figure 4 As shown, step S101 includes sub-steps S1011 to S1013.

[0045] Sub-step S1011: Obtain audio units from the first buffer and assign sequence identifiers to the audio units.

[0046] In this embodiment, the audio unit segments and stores the continuous audio data acquired in real time by the sound pickup device into a first buffer. This audio data includes the voice data of the target person. The first buffer includes a First-In-First-Out (FIFO) buffer. A sequence identifier for each audio unit is used to uniquely identify it.

[0047] In some embodiments, before sub-step S1011, the method further includes: acquiring continuous audio data collected in real time by the pickup device; dividing the continuous audio data into multiple audio units, each audio unit having a preset length; and storing the multiple audio units sequentially into a first buffer according to the order of acquisition time to support streaming processing. The preset length can be set based on actual conditions, and this embodiment of the invention does not specifically limit it. For example, the preset length is 100 milliseconds.

[0048] Sub-step S1012: Perform semantic tokenization on the audio unit using a preset semantic tokenizer to obtain the semantic token of the audio unit, store the audio unit in the second buffer, and store the semantic token in the third buffer.

[0049] In this embodiment, the preset semantic tokenizer is a pre-trained lightweight deep learning model capable of real-time, supervised semantic tokenization of audio streams. The semantic tokenizer is quantized and deployed on efficient inference frameworks such as ONNX Runtime or NCNN. Because the semantic tokenizer is quantized and deployed on efficient inference frameworks such as ONNX Runtime or NCNN, it is suitable for edge devices such as head-mounted displays (e.g., AR glasses).

[0050] The semantic tokenizer employs a streaming processing mechanism, enabling it to perform incremental computation on the input audio units with extremely low latency and continuously output semantic tokens. Through supervised training, the semantic tokenizer learns the association between semantic tokens and linguistic units such as phonemes, syllables, or subwords, ensuring that the generated tokens capture the language content rather than merely reconstructing the physical properties of the sound. It also ensures that the output rate of semantic tokens is positively correlated with the semantic density of the spoken content.

[0051] In some embodiments, since the second buffer and the third buffer store the audio unit and the semantic token of the audio unit simultaneously, the storage of the audio unit and the semantic token of the audio unit is strictly synchronized, which makes it easy to accurately index the associated audio unit in the second buffer according to the sorting of the semantic tokens in the third buffer after the target audio unit is distributed.

[0052] In some embodiments, the semantic tokenizer includes an encoder and a residual vector quantizer. Performing semantic tokenization on an audio unit using the preset semantic tokenizer to obtain a semantic token for the audio unit may include: encoding the audio unit using the encoder to obtain a feature sequence; and performing semantic tokenization on the feature sequence using the residual vector quantizer to obtain a semantic token for the audio unit. The encoder may be a Transformer-based encoder or a Conformer-based encoder.

[0053] In response to the total number of semantic tokens in the third buffer being less than a preset token number threshold, the execution sub-step S1011 is returned to retrieve audio units from the first buffer and assign sequence identifiers to the audio units.

[0054] In this embodiment, a preset token quantity threshold defines the number of semantic tokens that a semantically complete audio segment should contain. If the total number of semantic tokens in the third buffer is less than the preset token quantity threshold, it indicates that the audio units stored in the second buffer cannot yet constitute a semantically complete audio segment. Therefore, audio units are continued to be obtained from the first buffer, a sequence identifier is assigned to each audio unit, and then the audio units are semantically tokenized using a preset semantic tokenizer to obtain their semantic tokens. The audio units are then stored in the second buffer, and the semantic tokens are stored in the third buffer, thereby increasing the number of semantic tokens in the third buffer and increasing the number of audio units in the second buffer. The preset token quantity threshold can be set according to actual conditions, and this embodiment does not specifically limit it. For example, the preset token quantity threshold may include 120, 150, or 180, etc.

[0055] Sub-step S1013: In response to the total number of semantic tokens in the third buffer being greater than or equal to a preset token number threshold, each audio unit in the second buffer is spliced ​​together according to the sequence identifier of each audio unit in the second buffer to obtain the target audio segment.

[0056] The audio segmentation strategy based on streaming semantic token counting provided in this embodiment can ensure that each target audio segment participating in the subsequent matching contains a stable and predictable amount of semantic information. This high-quality and consistent target audio segment provides a solid foundation for the accurate matching of subsequent target audio segments and text blocks, thereby further improving the accuracy of target audio segment and text block matching and the reliability in complex scenarios.

[0057] Furthermore, the audio segmentation strategy based on streaming semantic token counting can filter out silent and non-semantic noise segments in the audio data stream (these segments do not generate or generate very few semantic tokens), so that only audio units containing actual speech content are accumulated and processed, thereby reducing invalid computation, saving computing resources, improving the efficiency of matching target audio segments with text blocks, and further reducing the synchronization delay between the prompt content and the target person's speech, thus ensuring real-time synchronization between the prompt content and the target person's speech.

[0058] Furthermore, the audio segmentation strategy based on streaming semantic token counting can adapt to the speaking styles of different target speakers. Whether the target speaker speaks quickly or slowly and thoughtfully, the strategy will segment the target audio segment after they have spoken approximately the same amount of "content" (measured by the number of semantic tokens), thus triggering the matching of subsequent audio segments and text blocks. Similarly, common filler words in spoken language such as "um" and "ah," as well as natural pauses, will not interfere with the segmentation of audio segments. This inherent adaptability to individual differences improves robustness.

[0059] In some embodiments, the target audio segment is obtained by segmenting audio data based on a streaming semantic token counting audio segmentation strategy. This ensures that each target audio segment participating in subsequent matching contains a stable and predictable amount of semantic information. Such high-quality, consistent target audio segments provide a solid foundation for accurate matching of subsequent target audio segments and text blocks. Combining the semantic matching strategy based on dynamic context modulation with the semantic matching of target audio segments and text blocks can further improve the accuracy of matching target audio segments and text blocks and the reliability in complex scenarios.

[0060] In some embodiments, concatenating each audio unit in the second buffer sequentially according to its sequence identifier to obtain the target audio segment may include: determining the concatenation order of each audio unit based on its sequence identifier, and concatenating each audio unit in the second buffer sequentially according to that concatenation order to obtain the target audio segment. Specifically, the smaller the sequence identifier of an audio unit, the earlier it is concatenated; conversely, the larger the sequence identifier, the later it is concatenated.

[0061] In some embodiments, such as Figure 5 As shown, after sub-step S1013, step S101 further includes:

[0062] Sub-step S1014: Determine the last M semantic tokens stored in the third buffer.

[0063] In this embodiment, M is an integer greater than or equal to 1. M represents the semantic token overlap quantity, which defines the amount of overlap between two consecutive audio segments at the semantic token level. The semantic token overlap quantity M can be set based on actual conditions, and this embodiment does not impose a specific limitation on it. For example, the preset semantic token overlap quantity M = 50.

[0064] Sub-step S1015: Determine the sequence identifier of the audio unit associated with each of the M semantic tokens to obtain the sequence identifier set.

[0065] For example, obtain the first moment of each of the M semantic tokens stored in the third buffer, and starting from the audio unit at the end of the second buffer, compare the second moment of each audio unit stored in the second buffer with each first moment until a second moment that is the same as the first moment corresponding to each of the M semantic tokens is found. Then, obtain the sequence identifier of the audio unit corresponding to each found second moment, thereby obtaining the sequence identifier set.

[0066] Sub-step S1016: Delete audio units in the second buffer whose sequence identifiers are not located in the sequence identifier set and delete semantic tokens in the third buffer except for M semantic tokens.

[0067] This embodiment deletes audio units in the second buffer whose sequence identifiers are not located in the sequence identifier set, and deletes semantic tokens in the third buffer except for the M semantic tokens. This retains the audio units associated with the last M semantic tokens stored in the third buffer in the second buffer, and also retains the last M semantic tokens stored in the third buffer. This ensures that the accumulation of semantic tokens and audio units for the next target audio segment does not start from zero. Thus, multiple audio units at the beginning of the next target audio segment are identical to multiple audio units at the end of the current target audio segment, ensuring semantic overlap between two consecutive target audio segments and maintaining the continuity of the target audio segment's context, thereby preventing jumps in the prompting content. After sub-step S1016, sub-steps S1011 to S1016 are repeated until the prompting content displayed by the display module is the last text block of the electronic document, thereby continuously receiving new audio units to achieve continuous intelligent sampling and segmentation of the target person's speech.

[0068] In some embodiments, before step S101, the method further includes: acquiring an electronic document, dividing the electronic document into multiple text blocks, and assigning a text block identifier to each text block, wherein the length of each text block is within a preset length range; for each text block, determining at least one first text block that is forward adjacent to the text block and at least one second text block that is backward adjacent to the text block; determining a first similarity between the text block and each first text block and determining a second similarity between the text block and each second text block; using the first text block corresponding to a first similarity greater than or equal to a preset similarity threshold as the context text block of the text block, and / or using the second text block corresponding to a second similarity greater than or equal to a preset similarity threshold as the context text block of the text block; using the text block identifier of the context text block as the context pointer information of the text block, and determining a second embedding vector for each text block; storing the multiple text blocks, the second embedding vector of each text block in the multiple text blocks, the context pointer information of each text block in the multiple text blocks, and the similarity between each text block in the multiple text blocks and its corresponding context text block. In this embodiment, the electronic document is segmented into multiple text blocks, context pointer information between text blocks is automatically calculated and established, and a second embedding vector for each text block is determined. Then, each text block, the second embedding vector of each text block, the context pointer information, and the similarity between each text block and its corresponding context text block are stored in the database. This provides high-quality prior knowledge for the subsequent matching of audio segments and electronic documents, thereby improving the accuracy and efficiency of matching audio segments and electronic documents.

[0069] In some embodiments, storing multiple text blocks, the second embedding vector of each text block, the context pointer information of each text block, and the similarity between each text block and its corresponding context text block may include storing the multiple text blocks, the second embedding vector of each text block, the context pointer information of each text block, and the similarity between each text block and its corresponding context text block in a vector database. The vector database may include a FAISS database or a Milvus database, etc. This embodiment utilizes a vector database to store multiple text blocks, the second embedding vector of each text block, the context pointer information of each text block, and the similarity between each text block and its corresponding context text block, facilitating efficient and rapid retrieval of relevant prior knowledge from the vector database.

[0070] In some embodiments, determining the first similarity between a text block and a first text block may include: determining the embedding vector of the text block and the embedding vector of the first text block; determining the cosine similarity between the embedding vector of the text block and the embedding vector of the first text block, and determining the cosine similarity between the embedding vector of the text block and the embedding vector of the first text block as the first similarity between the text block and the first text block. Alternatively, based on the TF-IDF algorithm or the Best Match 25 algorithm, determining the word vectors of the text block and the word vectors of the first text block, determining the cosine similarity between the word vectors of the text block and the word vectors of the first text block, and determining the cosine similarity between the word vectors of the text block and the word vectors of the first text block as the first similarity between the text block and the first text block. Alternatively, determining a first number of intersection words between the text block and the first text block, and determining a second number of union words between the text block and the first text block, dividing the first number by the second number to obtain the first similarity between the text block and the first text block.

[0071] In some embodiments, determining the first similarity between a text block and a first text block may also include: determining multiple topics of the electronic document, and determining a first probability distribution vector of the text block on the multiple topics and a second probability distribution vector of the first text block on the multiple topics; determining the cosine similarity between the first probability distribution vector and the second probability distribution vector, and determining the cosine similarity between the first probability distribution vector and the second probability distribution vector as the first similarity between the text block and the first text block. Alternatively, determining the initial similarity between multiple text blocks, constructing a directed graph by using the text blocks as graph nodes and the initial similarity between the multiple text blocks as edges between graph nodes, determining the distance between the graph node corresponding to the text block in the directed graph and the graph node corresponding to the first text block in the directed graph, and determining the distance between the graph node corresponding to the text block in the directed graph and the graph node corresponding to the first text block in the directed graph as the first similarity between the text block and the first text block.

[0072] It should be noted that the specific method for determining the second similarity between the text block and the second text block can refer to the corresponding embodiment for determining the first similarity between the text block and the first text block in the foregoing embodiments, and will not be repeated here.

[0073] Step S102: Select the K text blocks that are most semantically similar to the target audio segment from the multiple pre-stored text blocks as K candidate text blocks.

[0074] In this embodiment, multiple text blocks are obtained by pre-segmenting the electronic document, and K is an integer greater than or equal to 1. For example, K = 5. This embodiment selects the K text blocks that are most semantically similar to the target audio segment from the pre-stored multiple text blocks as K candidate text blocks. This facilitates further matching of the K candidate text blocks that are most semantically similar to the target audio segment with the target audio segment in the subsequent process, thereby improving the accuracy of the matching.

[0075] In some embodiments, obtaining K text blocks that are semantically most similar to the target audio segment from a plurality of pre-stored text blocks as K candidate text blocks includes: determining a first embedding vector of the target audio segment and obtaining a pre-stored second embedding vector for each text block in the plurality of text blocks; determining the similarity between the first embedding vector and the second embedding vector of each text block in the plurality of text blocks; and obtaining K text blocks that are semantically most similar to the target audio segment from the plurality of text blocks as K candidate text blocks based on the similarity between the first embedding vector and each second embedding vector. In related technologies, the accuracy of automatic speech recognition is highly dependent on various factors, including ASR transcription errors, synonym substitutions, word order inversions, and slips of the tongue, which can lead to the inability to match accurate prompting content. This embodiment directly uses the embedding vector of the target audio segment and the embedding vector of the text blocks for matching, enabling semantic understanding of the text blocks and the speech of the target person. This solves the problem of inaccurate prompting content due to ASR transcription errors, synonym substitutions, word order inversions, or slips of the tongue, thus improving the accuracy of the displayed prompting content.

[0076] In some embodiments, obtaining the K text blocks that are most semantically similar to the target audio segment from multiple text blocks as K candidate text blocks based on the similarity between the first embedding vector and each second embedding vector may include: sorting the multiple text blocks in descending order of the similarity between the first embedding vector and each second embedding vector to obtain a text block queue, and obtaining the first K text blocks from the text block queue as K candidate text blocks.

[0077] In one embodiment, determining the first embedding vector of the target audio segment may include: converting the target audio segment into text data using an ASR model, and then embedding the text data using a preset text embedding model to obtain the first embedding vector. Each second embedding vector stored in the database is obtained by pre-embedding each segmented text block using the text embedding model. The preset text embedding model can be set based on actual conditions, and this embodiment does not specifically limit it. The preset text embedding model may include a Sentence-BERT model, a BERT model, a Doc2Vec model, or a Universal Sentence Encoder model, etc. This embodiment omits the step of calculating the embedding vector of the text block, thus reducing the synchronization delay between the prompting content and the target person's speech, and improving the real-time performance of displaying the prompting content.

[0078] In some embodiments, determining the first embedding vector of a target audio segment may include: embedding the target audio segment using a phono-text joint embedding model to obtain the first embedding vector. Each second embedding vector stored in the database is obtained in advance by embedding each segmented text block using the phono-text joint embedding model. In related technologies, automatic speech recognition requires a certain processing time, resulting in a significant delay from the target speaker's speech to the display of the prompting content on the terminal device. This makes it impossible to guarantee real-time synchronization between the prompting content and the target speaker's speech, resulting in low real-time performance. This embodiment determines the embedding vector of the target audio segment using a phono-text joint embedding model, eliminating the need for ASR transcription and thus bypassing this time-consuming step. Furthermore, the embedding vector of the text block is predetermined and stored, omitting the step of calculating the embedding vector of the text block. Therefore, the synchronization delay between the prompting content and the target speaker's speech is reduced, improving the real-time performance of the displayed prompting content.

[0079] Step S103: For each candidate text block, determine the context text block of the candidate text block, and determine the context consistency coefficient based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block.

[0080] In this embodiment, the context text block of the candidate text block includes text blocks whose similarity to the candidate text block is greater than or equal to a preset similarity threshold, and whose number of text blocks separated from the candidate text block is less than N.

[0081] In some embodiments, determining the context text block of a candidate text block may include: obtaining context pointer information corresponding to the candidate text block; and determining the context text block of the candidate text block from multiple text blocks based on the context pointer information. The context pointer information corresponding to the candidate text block may be predetermined and stored, or it may be determined in real time; this embodiment of the invention does not specifically limit this. Optionally, the context pointer information corresponding to the candidate text block may be predetermined and stored. This embodiment can quickly and accurately determine the context text block of a candidate text block based on the context pointer information of the candidate text block.

[0082] In some embodiments, determining the similarity between a target audio segment and a context text block may include: determining the similarity between a first embedding vector and a second embedding vector of the context text block, and determining the similarity between the first embedding vector and the second embedding vector of the context text block as the similarity between the target audio segment and the context text block. Similarly, determining the similarity between a target audio segment and a candidate text block may include: determining the similarity between a first embedding vector and a second embedding vector of the candidate text block, and determining the similarity between the first embedding vector and the second embedding vector of the candidate text block as the similarity between the target audio segment and the candidate text block. The similarity between the candidate text block and each context text block can be directly obtained from the database and does not need to be calculated.

[0083] In some embodiments, determining the context consistency coefficient based on the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block, may include: determining the importance weight of each context text block relative to the candidate text block; and determining the context consistency coefficient based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block. This embodiment can determine the context consistency coefficient more accurately.

[0084] In some embodiments, determining the importance weight of each context text block relative to a candidate text block includes: determining the importance weight of each context text block relative to a candidate text block based on a pre-stored importance weight relationship table, wherein the importance weight relationship table includes the importance weight of the context text block relative to the corresponding text block for each of a plurality of text blocks; or determining the distance between each context text block and a candidate text block, and determining the importance weight of each context text block relative to a candidate text block based on the distance between each context text block and a candidate text block. Specifically, the greater the distance between the context text block and the candidate text block, the smaller the importance weight of the context text block relative to the candidate text block; the closer the distance between the context text block and the candidate text block, the greater the importance weight of the context text block relative to the candidate text block.

[0085] In some embodiments, the importance weight of the context text block relative to the corresponding text block for each text block in a plurality of text blocks is a pre-set fixed value. This fixed value can be learned through machine learning methods based on a training sample dataset. For example, an end-to-end model for learning the importance weights is constructed, using the importance weights as trainable parameters of the model; a training sample dataset is obtained, where each training sample includes an audio segment sample, a candidate text block sample, a context text block sample of the candidate text block, and a first label, which is used to determine whether the audio segment sample and the candidate text block correctly match; a training sample is obtained from the training sample dataset as the current training sample; a context consistency coefficient is determined based on the similarity between the audio segment sample in the current training sample and each context text block in the current training sample, the similarity between the candidate text block in the current training sample and each context text block in the current training sample, and the importance weight of each context text block in the current training sample relative to the candidate text block; and then... The similarity between the audio segment sample in the current training sample data and the candidate text block in the current training sample data is multiplied by the context consistency coefficient to obtain the target similarity between the audio segment sample in the current training sample data and the candidate text block in the current training sample data. Based on the target similarity, a second label is determined. The second label is used to determine whether the audio segment sample in the current training sample data and the candidate text block in the current training sample data match correctly. Based on the second label and the first label in the current training sample data, the importance weight of each context text block in the current training sample data relative to the candidate text block is updated. The process of retrieving a training sample data from the training sample dataset as the current training sample data is repeated until every training sample data in the training sample dataset has been used. At this point, training stops, and a final set of importance weights is obtained.

[0086] In some embodiments, determining the context consistency coefficient based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block includes: for each context text block, multiplying the importance weight of the context text block relative to the candidate text block, the similarity between the context text block and the target audio segment, and the similarity between the context text block and the candidate text block to obtain M multiplication results, summing the M multiplication results, and then adding 1 to obtain the context consistency coefficient. The context consistency coefficient in this embodiment accurately reflects the comprehensive matching between the target audio segment and the candidate text block, and effectively suppresses matching that is locally similar but contextually dissimilar, enabling a more refined and accurate determination of the target similarity between the target audio segment and the candidate text block using the context consistency coefficient.

[0087] For example, the context consistency coefficient can be determined using the following first formula:

[0088]

[0089] in, It is the context consistency coefficient, where A is the target audio segment, and T is the context consistency coefficient. c It is a candidate text block. It is candidate text block T c The i-th context text block in M ​​context text blocks, where i is greater than or equal to 1 and less than or equal to M, w i It is the i-th context text block relative to the candidate text block T c Importance weights The target audio segment A and the candidate text block T c The similarity of the i-th context text block reflects whether the target audio segment A also supports context. It is candidate text block T c With candidate text block T c The similarity of the i-th context text block represents the similarity of candidate text block T. c With candidate text block T c The inherent link strength between the i-th context text blocks.

[0090] In some embodiments, determining the context consistency coefficient based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block includes: for each context text block, multiplying the importance weight of the context text block relative to the candidate text block, the similarity between the context text block and the target audio segment, and the similarity between the context text block and the candidate text block to obtain M multiplication results; taking the maximum value among the M multiplication results and adding 1 to obtain the context consistency coefficient. This embodiment determines the context consistency coefficient based on max pooling aggregation, reducing the interference of noise on the context consistency coefficient. If only one of the M context text blocks is highly relevant to the target audio segment, while the others are irrelevant or even misleading, max pooling aggregation can accurately capture this strongest support signal, avoiding it being averaged out by other weak signals or noise. This allows for a more refined and accurate determination of the target similarity between the target audio segment and the candidate text block through the context consistency coefficient.

[0091] For example, the context consistency coefficient can be determined using the following second formula:

[0092]

[0093] in, It is the context consistency coefficient, where A is the target audio segment, and T is the context consistency coefficient. c It is a candidate text block. It is candidate text block T c The i-th context text block in the M context text blocks, w i It is the i-th context text block relative to the candidate text block T c Importance weights The target audio segment A and the candidate text block T c The similarity of the i-th context text block, It is candidate text block T c With candidate text block T c The similarity of the i-th context text block.

[0094] In some embodiments, determining the context consistency coefficient based on the importance weight of each context text block relative to a candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block includes: determining a gating value based on the similarity between the target audio segment and the candidate text block; for each context text block, multiplying the importance weight of the context text block relative to the candidate text block, the similarity between the context text block and the target audio segment, and the similarity between the context text block and the candidate text block to obtain M multiplication results, summing the M multiplication results, multiplying by the gating value, and then adding 1 to obtain the context consistency coefficient. In this embodiment, the gating value (typically between 0 and 1) is determined by the similarity between the target audio segment and the candidate text block, which makes the contribution of the context controllable. When the similarity between the target audio segment and the candidate text block is very high, the gate value is close to 1, allowing the context information to play a full role. When the similarity between the target audio segment and the candidate text block is very low, the gate value is close to 0, thereby weakening the influence of the context. This effectively prevents the context information from being mistakenly used to enhance an incorrect candidate text block when the target audio segment and the candidate text block themselves are unreliable, thus increasing the stability of the algorithm.

[0095] For example, the context consistency coefficient can be determined using the following third formula:

[0096]

[0097] g = sigmoid(w g Sim(A,T) c )+b g ),

[0098] in, It is the context consistency coefficient, where A is the target audio segment, and T is the context consistency coefficient. c It is a candidate text block. It is candidate text block T c The i-th context text block in the M context text blocks, w i It is the i-th context text block relative to the candidate text block T c Importance weights, Sim(A,T) c ) is the target audio segment and candidate text block T c similarity, The target audio segment A and the candidate text block T c The similarity of the i-th context text block, It is candidate text block T c With candidate text block T cThe similarity of the i-th context text block, g is the gating value, and sigmoid is the activation function. The sigmoid function can adjust the similarity of w... g Sim(A,T) c )+b g The result is mapped to the range 0 to 1, w g It represents the gating weight, bg is the gating bias term, and w g The background and background are pre-set.

[0099] In some embodiments, determining the context consistency coefficient based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block includes: for each context text block, determining an attention weight based on the importance weight of the context text block relative to the candidate text block, the similarity between the target audio segment and the context text block, and the similarity between the candidate text block and the context text block; multiplying the attention weight of the context text block by the similarity between the target audio segment and the context text block to obtain M multiplication results; summing the M multiplication results and adding 1 to obtain the context consistency coefficient. This embodiment can automatically determine which contextual clues are more important in the current specific match based on real-time input (the similarity between the target audio segment and the context text block and the similarity between the candidate text block and the context text block) and assign them higher influence. This makes the aggregation process more flexible and adaptive, improves the accuracy of the context consistency coefficient, and enables subsequent determination of the target similarity between the target audio segment and the candidate text block more finely and accurately using the context consistency coefficient.

[0100] For example, the context consistency coefficient can be determined using the following fourth formula:

[0101]

[0102] in, It is the context consistency coefficient, where A is the target audio segment, and T is the context consistency coefficient. c It is a candidate text block. It is candidate text block T c The i-th context text block in the M context text blocks, w i It is the i-th context text block relative to the candidate text block T c Importance weights The target audio segment A and the candidate text block T c The similarity of the i-th context text block, It is candidate text block T c With candidate text block T c The similarity of the i-th context text block, ai These are the attention weights for the i-th context text block, and softmax is the activation function. Softmax can... The result is mapped to the range of 0 to 1.

[0103] Step S104: Based on the similarity and contextual consistency coefficient between the target audio segment and the candidate text block, determine the target similarity between the target audio segment and the candidate text block, and determine the candidate text block with the highest target similarity among the K candidate text blocks as the target text block.

[0104] For example, the target similarity between the target audio segment and the candidate text block is obtained by multiplying the similarity between the target audio segment and the candidate text block by the context consistency coefficient. For instance, the target similarity between the target audio segment and the candidate text block... Where Sim(A,T) c ) is the target audio segment and candidate text block T c similarity, It is the context consistency coefficient.

[0105] Step S105: Switch the displayed prompt content to the text content in the electronic document that is related to the target text block.

[0106] In this embodiment, when determining the text content that matches the voice data of a target person in an electronic document, it does not simply select the text block with the highest similarity to the target audio segment as the final target text block. Instead, it first obtains the K text blocks with the highest semantic similarity to the target audio segment from a pre-stored set of text blocks as K candidate text blocks. Then, for each candidate text block, a more refined comprehensive matching is performed, combining the candidate text block with its context text blocks. This ensures that the calculated target similarity between the target audio segment and the candidate text block not only integrates the similarity between the target audio segment and the candidate text block, but also integrates the similarity between the target audio segment and the candidate text block with its context text blocks, as well as the similarity between the candidate text block and its context text blocks. This guarantees the accuracy of the target similarity between the target audio segment and the candidate text block. By determining the candidate text block with the highest target similarity among the K candidate text blocks as the target text block, high-precision and robust matching of audio segments and text content is achieved, improving the accuracy of determining the text content that matches the voice data of a target person in an electronic document, thereby improving the accuracy of the displayed prompts.

[0107] In some embodiments, the text content related to the target text block includes the target text block itself, or the text content related to the target text block includes the next text block adjacent to the target text block. This embodiment can display the target text block, or display the target text block and the next text block adjacent to the target text block, as prompting content, enabling the target person to give a speech, sing, or tell a story more fluently based on the displayed prompting content, thus improving the user experience.

[0108] It should be noted that since the target audio segments are continuously generated, steps S101 to S105 are also executed in a continuous loop to achieve continuous tracking and prompting of the target person's speech. For example, when the target person begins to speak, after a short period of time (such as a few seconds), the first target audio segment is generated. Then, based on the first target audio segment, steps S101 to S105 are executed sequentially. After that, a new target audio segment is generated, and then based on the new target audio segment, steps S101 to S105 are executed sequentially.

[0109] In some embodiments, after step S104, the method further includes: responding to the presence of a text block forgotten by the target person among multiple text blocks, outputting a prompt message, the prompt message being used to indicate that a text block has been omitted from the electronic document. The output of the prompt message may include controlling the display device to output the prompt message. This embodiment outputs a prompt message when it is discovered that the target person has omitted content from the electronic document during their speech, thus preventing the target person from missing important content and improving the user experience.

[0110] In some embodiments, after step S104, the method further includes acquiring historical text blocks, which are text blocks corresponding to prompts displayed by the display device before the current moment; in response to the discontinuity between the text block identifier of the historical text block and the text block identifier of the target text block, it is determined that there are text blocks forgotten by the target person among the multiple text blocks. For example, if the historical text block is the fifth text block in the electronic document and the target text block is the eighth text block in the electronic document, then the target person has missed the sixth and seventh text blocks.

[0111] In some embodiments, in response to the presence of a forgotten text block among multiple text blocks, outputting a prompt message may include: marking the forgotten text block among the multiple text blocks; and outputting a prompt message in response to the marking duration of the marked text block being greater than or equal to a preset duration threshold. The preset duration threshold can be set based on actual circumstances, and this embodiment does not specifically limit it. For example, the preset duration threshold may include 5 minutes, 10 minutes, or 15 minutes. This embodiment can avoid frequent or continuous output of prompt messages indicating that text blocks have been omitted from electronic documents, thereby reducing the impact of the output prompt message on the target person's speech, singing, or narration.

[0112] In some embodiments, the prompting content switching method provided in this application further includes: in response to the target text block being a marked text block, removing the mark of the target text block and clearing the mark duration of the target text block.

[0113] In some embodiments, in response to the presence of a forgotten text block among multiple text blocks, outputting a prompt message may include: marking the forgotten text block among the multiple text blocks; and outputting a prompt message in response to the target text block being the last text block among the multiple text blocks, the prompt message indicating that a text block has been omitted from the electronic document. In this embodiment, when the prompt content has already displayed the last text block of the electronic document, if a text block in the electronic document has been omitted by the target, outputting a prompt message indicating that a text block has been omitted from the electronic document can prevent the target from missing important content and improve the user experience.

[0114] Please see Figure 6 , Figure 6 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.

[0115] like Figure 6 As shown, the terminal device 300 includes a processor 301 and a memory 302, which are connected by a bus 303, such as an I2C (Inter-integrated Circuit) bus.

[0116] Specifically, processor 301 provides computing and control capabilities to support the operation of the entire terminal device. Processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0117] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.

[0118] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the embodiments of the present invention, and does not constitute a limitation on the terminal device to which the embodiments of the present invention are applied. A specific terminal device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0119] The processor 301 is used to run a computer program stored in the memory 302, and when executing the computer program, implements any of the methods for displaying prompt content provided in the embodiments of the present invention.

[0120] In some embodiments, the processor 301 is configured to run a computer program stored in memory, and when executing the computer program, perform the following steps:

[0121] Acquire a target audio segment, wherein the target audio segment includes the voice data of the target person;

[0122] K text blocks that are most semantically similar to the target audio segment are obtained from a plurality of pre-stored text blocks as K candidate text blocks. The plurality of text blocks are obtained by pre-segmenting the electronic document, and K is an integer greater than or equal to 1.

[0123] For each candidate text block, the context text block of the candidate text block is determined, and the context consistency coefficient is determined based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block;

[0124] Based on the similarity between the target audio segment and the candidate text block and the context consistency coefficient, the target similarity between the target audio segment and the candidate text block is determined, and the candidate text block with the highest target similarity among the K candidate text blocks is determined as the target text block;

[0125] The displayed prompt will be switched to the text content in the electronic document that is related to the target text block.

[0126] In some embodiments, when the processor 301 determines the context text block of the candidate text block, it is configured to:

[0127] Obtain the context pointer information corresponding to the candidate text block;

[0128] Based on the context pointer information, the context text block of the candidate text block is determined from the plurality of text blocks.

[0129] In some embodiments, when the processor 301 retrieves K text blocks that are semantically most similar to the target audio segment from a plurality of pre-stored text blocks as K candidate text blocks, it is configured to:

[0130] Determine the first embedding vector of the target audio segment, and obtain the second embedding vector pre-stored for each of the plurality of text blocks;

[0131] Determine the similarity between the first embedding vector and the second embedding vector of each of the plurality of text blocks;

[0132] Based on the similarity between the first embedding vector and each of the second embedding vectors, K text blocks that are most semantically similar to the target audio segment are selected from the plurality of text blocks as K candidate text blocks.

[0133] In some embodiments, before acquiring the target audio segment, the processor 301 is further configured to:

[0134] The electronic document is obtained, the electronic document is divided into multiple text blocks, and each of the multiple text blocks is assigned a text block identifier. The length of each of the multiple text blocks is within a preset length range.

[0135] For each of the plurality of text blocks, at least one first text block that is forward adjacent to the text block and at least one second text block that is backward adjacent to the text block are determined from the plurality of text blocks, where N is an integer greater than or equal to 1;

[0136] Determine a first similarity between the text block and each of the first text blocks, and determine a second similarity between the text block and each of the second text blocks;

[0137] The first text block corresponding to the first similarity being greater than or equal to the preset similarity threshold is used as the context text block of the text block, and / or the second text block corresponding to the second similarity being greater than or equal to the preset similarity threshold is used as the context text block of the text block;

[0138] The text block identifier of the context text block is used as the context pointer information of the text block, and a second embedding vector is determined for each of the plurality of text blocks;

[0139] The system stores the plurality of text blocks, the second embedding vector of each of the plurality of text blocks, the context pointer information of each of the plurality of text blocks, and the similarity between each of the plurality of text blocks and its corresponding context text block.

[0140] In some embodiments, when the processor 301 determines the context consistency coefficient based on the similarity between the target audio segment and each of the context text blocks and the similarity between the candidate text block and each of the context text blocks, it is configured to:

[0141] Determine the importance weight of each of the context text blocks relative to the candidate text blocks;

[0142] The context consistency coefficient is determined based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block.

[0143] In some embodiments, the processor 301, when determining the importance weight of each of the context text blocks relative to the candidate text blocks, is configured to:

[0144] Determine the distance between each of the context text blocks and the candidate text blocks;

[0145] The importance weight of each context text block relative to the candidate text block is determined based on the distance between each context text block and the candidate text block.

[0146] In some embodiments, when acquiring a target audio segment, the processor 301 is configured to:

[0147] Audio units are obtained from the first buffer, and sequence identifiers are assigned to the audio units. The audio units are audio data that has been segmented and stored in the first buffer after being acquired in real time. The audio data includes the speech data.

[0148] The audio unit is semantically tokenized by a preset semantic tokenizer to obtain the semantic token of the audio unit. The audio unit is then stored in a second buffer and the semantic token is stored in a third buffer.

[0149] In response to the total number of semantic tokens in the third buffer being less than a preset token number threshold, the process returns to the step of obtaining audio units from the first buffer and assigning sequence identifiers to the audio units;

[0150] In response to the total number of semantic tokens in the third buffer being greater than or equal to the preset token number threshold, each audio unit in the second buffer is sequentially concatenated according to the sequence identifier of each audio unit in the second buffer to obtain the target audio segment.

[0151] In some embodiments, after concatenating each audio unit in the second buffer to obtain the target audio segment, the processor 301 is further configured to implement:

[0152] Determine the last M semantic tokens stored in the third buffer, where M is an integer greater than or equal to 1;

[0153] Determine the sequence identifier associated with each of the M semantic tokens to obtain a sequence identifier set;

[0154] Delete the audio unit whose sequence identifier is not located in the sequence identifier set in the second buffer, and delete the semantic tokens in the third buffer except for the M semantic tokens.

[0155] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the corresponding process in the aforementioned embodiment of the prompting content display control method, and will not be repeated here.

[0156] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, which can be executed by one or more processors to implement any of the prompting content display control methods provided in the specification of this invention.

[0157] The storage medium can be volatile or non-volatile. It can be an internal storage unit of the terminal device described in the foregoing embodiments, such as the hard drive or memory of the terminal device. Alternatively, it can be an external storage device of the terminal device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or FlashCard.

[0158] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0159] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0160] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for controlling the display of prompt content, characterized in that, The method includes: Acquire a target audio segment, wherein the target audio segment includes the voice data of the target person; K text blocks that are most semantically similar to the target audio segment are obtained from a plurality of pre-stored text blocks as K candidate text blocks. The plurality of text blocks are obtained by pre-segmenting the electronic document, and K is an integer greater than or equal to 1. For each candidate text block, the context text block of the candidate text block is determined, and the context consistency coefficient is determined based on the similarity between the target audio segment and each context text block and the similarity between the candidate text block and each context text block; Based on the similarity between the target audio segment and the candidate text block and the context consistency coefficient, the target similarity between the target audio segment and the candidate text block is determined, and the candidate text block with the highest target similarity among the K candidate text blocks is determined as the target text block; The displayed prompt will be switched to the text content in the electronic document that is related to the target text block.

2. The display control method according to claim 1, characterized in that, The step of determining the context text block of the candidate text block includes: Obtain the context pointer information corresponding to the candidate text block; Based on the context pointer information, the context text block of the candidate text block is determined from the plurality of text blocks.

3. The display control method according to claim 1, characterized in that, The step of obtaining the K text blocks that are most semantically similar to the target audio segment from a plurality of pre-stored text blocks as K candidate text blocks includes: Determine the first embedding vector of the target audio segment, and obtain the second embedding vector pre-stored for each of the plurality of text blocks; Determine the similarity between the first embedding vector and the second embedding vector of each of the plurality of text blocks; Based on the similarity between the first embedding vector and each of the second embedding vectors, K text blocks that are most semantically similar to the target audio segment are selected from the plurality of text blocks as K candidate text blocks.

4. The display control method according to claim 2 or 3, characterized in that, Before acquiring the target audio segment, the process also includes: The electronic document is obtained, the electronic document is divided into multiple text blocks, and each of the multiple text blocks is assigned a text block identifier. The length of each of the multiple text blocks is within a preset length range. For each of the plurality of text blocks, at least one first text block that is forward adjacent to the text block and at least one second text block that is backward adjacent to the text block are determined from the plurality of text blocks, where N is an integer greater than or equal to 1; Determine a first similarity between the text block and each of the first text blocks, and determine a second similarity between the text block and each of the second text blocks; The first text block corresponding to the first similarity being greater than or equal to the preset similarity threshold is used as the context text block of the text block, and / or the second text block corresponding to the second similarity being greater than or equal to the preset similarity threshold is used as the context text block of the text block; The text block identifier of the context text block is used as the context pointer information of the text block, and a second embedding vector is determined for each of the plurality of text blocks; The system stores the plurality of text blocks, the second embedding vector of each of the plurality of text blocks, the context pointer information of each of the plurality of text blocks, and the similarity between each of the plurality of text blocks and its corresponding context text block.

5. The display control method according to any one of claims 1-3, characterized in that, The step of determining the context consistency coefficient based on the similarity between the target audio segment and each of the context text blocks, and the similarity between the candidate text block and each of the context text blocks, includes: Determine the importance weight of each of the context text blocks relative to the candidate text blocks; The context consistency coefficient is determined based on the importance weight of each context text block relative to the candidate text block, the similarity between the target audio segment and each context text block, and the similarity between the candidate text block and each context text block.

6. The display control method according to claim 5, characterized in that, Determining the importance weight of each of the context text blocks relative to the candidate text blocks includes: Determine the distance between each of the context text blocks and the candidate text blocks; The importance weight of each context text block relative to the candidate text block is determined based on the distance between each context text block and the candidate text block.

7. The display control method according to any one of claims 1-3, characterized in that, The acquisition of the target audio segment includes: Audio units are obtained from the first buffer, and sequence identifiers are assigned to the audio units. The audio units are audio data that has been segmented and stored in the first buffer after being acquired in real time. The audio data includes the speech data. The audio unit is semantically tokenized by a preset semantic tokenizer to obtain the semantic token of the audio unit. The audio unit is then stored in a second buffer and the semantic token is stored in a third buffer. In response to the total number of semantic tokens in the third buffer being less than a preset token number threshold, the process returns to the step of obtaining audio units from the first buffer and assigning sequence identifiers to the audio units; In response to the total number of semantic tokens in the third buffer being greater than or equal to the preset token number threshold, each audio unit in the second buffer is sequentially concatenated according to the sequence identifier of each audio unit in the second buffer to obtain the target audio segment.

8. The display control method according to claim 7, characterized in that, After concatenating each audio unit in the second buffer to obtain the target audio segment, the method further includes: Determine the last M semantic tokens stored in the third buffer, where M is an integer greater than or equal to 1; Determine the sequence identifier associated with each of the M semantic tokens to obtain a sequence identifier set; Delete the audio unit whose sequence identifier is not located in the sequence identifier set in the second buffer, and delete the semantic tokens in the third buffer except for the M semantic tokens.

9. A terminal device, characterized in that, The terminal device includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for connecting and communicating between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of the method for displaying prompt content as described in any one of claims 1 to 8.

10. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the method for displaying prompt content according to any one of claims 1 to 8.

Citation Information

Cited By

  • Resume information extraction method and computing device

    CN120597873A