Word prompting control method and device, head-mounted word prompting equipment and storage medium

By combining a multi-microphone array with voiceprint and lip movement feature verification, the head-mounted teleprompter achieves high accuracy and secure teleprompter control, solving the problems of environmental noise and misidentification, and ensuring the accuracy and security of teleprompter scrolling.

CN121963760APending Publication Date: 2026-05-01ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUHAI MOJIE TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing head-mounted teleprompter devices are prone to misidentifying environmental noise, other people's speech, or background noise generated by the device itself as the user's voice, leading to incorrect scrolling of the teleprompter text. Furthermore, they cannot effectively distinguish between scenarios where the user is speaking to an audience and interacting with others, resulting in inaccurate teleprompter scrolling control.

Method used

The system uses a multi-microphone array for sound source localization and voiceprint similarity verification, combined with lip movement feature verification, to ensure the validity of the input speech; the prompting text scrolling is only executed when the text matching degree reaches a preset threshold.

Benefits of technology

It effectively filters out environmental noise and other people's voices, ensuring that the prompt scrolling is strictly synchronized with the user's voice, thus improving the accuracy and security of prompt control and preventing accidental triggering and illegal operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963760A_ABST
    Figure CN121963760A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, and provides a prompt control method and device, a head-mounted prompt device and a storage medium, and the method carries out the validity verification of the current input voice through a preset verification mode, guarantees that only the input voice meeting the requirement can be used for the judgment processing flow of prompt control, and improves the prompt control efficiency. And the probability of false triggering is reduced. After the current input voice is determined to be effective input, text matching is performed on the voice text of the current input voice and the paragraph to be prompted in the prompted text, and whether the voice content of the current input voice is consistent with the prompted content or not can be accurately judged by calculating the text matching degree; and mistaken scrolling caused by chat, cough or mistaken mouth of the wearer is avoided. According to the technical scheme, judgment is carried out through the preset matching degree threshold value, the prompt scrolling operation is executed only when the text matching degree is larger than or equal to the matching degree threshold value, it is ensured that prompt scrolling is strictly synchronized with the voice content of the current input voice, and therefore high-accuracy prompt control is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Teleprompter control method, device, head-mounted teleprompter and storage medium Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a teleprompter control method, device, head-mounted teleprompter device, and storage medium. Background Technology

[0002] With the widespread adoption of smart wearable devices, head-mounted teleprompters are increasingly used in scenarios such as speeches, live broadcasts, and meetings. Head-mounted teleprompters typically use speech recognition technology to detect what the user is reading and automatically scroll the teleprompter text, helping users express themselves more fluently. Traditional teleprompter control methods mainly rely on simple speech energy detection or keyword recognition, triggering the teleprompter scrolling when the user speaks.

[0003] However, the relevant prompting control methods are prone to misidentifying environmental noise, other people's voices, or background noise generated by the device itself as the user's voice, and cannot effectively distinguish the wearer's own voice from other speakers, resulting in the prompt text scrolling incorrectly; moreover, the relevant prompting control cannot effectively distinguish between scenarios where the user is speaking to an audience and interacting with others, resulting in inaccurate prompt scrolling control.

[0004] Therefore, improving the accuracy of the scrolling control of head-mounted teleprompter devices has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a teleprompter control method, device, head-mounted teleprompter device, and storage medium, aiming to solve the technical problem of low accuracy in teleprompter scrolling control of head-mounted teleprompter devices by related teleprompter control methods.

[0006] In a first aspect, this application provides a prompting control method, comprising: upon receiving current input speech, validating the current input speech based on a preset verification method; upon determining that the current input speech is valid input, performing text matching between the speech text of the current input speech and the prompting paragraph in the prompting text, and determining the text matching degree between the speech text and the prompting paragraph; and controlling the prompting text to scroll when the text matching degree is greater than or equal to a preset matching degree threshold.

[0007] Secondly, this application also provides a prompting control device, comprising: a validity verification module, configured to verify the validity of the current input speech based on a preset verification method when the current input speech is received; a text matching module, configured to perform text matching between the speech text of the current input speech and the prompting paragraph in the prompting text when the current input speech is determined to be valid input, and determine the text matching degree between the speech text and the prompting paragraph; and a prompting control module, configured to control the prompting text to scroll when the text matching degree is greater than or equal to a preset matching degree threshold.

[0008] Thirdly, this application also provides a head-mounted teleprompter device, which includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the teleprompter control method as described above.

[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the prompting control method described above.

[0010] This application provides a teleprompter control method, device, head-mounted teleprompter device, and storage medium. The method verifies the validity of the current input speech through a preset verification method, effectively filtering out unacceptable environmental noise, other people's voices, and invalid background noise. This ensures that only compliant input speech can be used in the teleprompter control judgment and processing flow, reducing the probability of false triggering from the source. After determining that the current input speech is valid, the speech text of the current input speech is matched with the segment to be teleprompted in the teleprompter text. By calculating the text matching degree, it can accurately determine whether the speech content of the current input speech is consistent with the teleprompter content, avoiding erroneous scrolling caused by the wearer's idle chatter, coughing, or slips of the tongue. A preset matching degree threshold is used for judgment; the teleprompter scrolling operation is only executed when the text matching degree is greater than or equal to the matching degree threshold, further improving the accuracy of teleprompter control and ensuring that the teleprompter scrolling is strictly synchronized with the speech content of the current input speech, thereby achieving highly accurate teleprompter control. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 is a flowchart illustrating a first embodiment of a prompting control method provided in this application; Figure 2 is a flowchart illustrating a second embodiment of a prompting control method provided in this application; Figure 3 is a structural schematic diagram illustrating a first embodiment of a prompting control device provided in this application; Figure 4 is a structural schematic block diagram illustrating a head-mounted prompting device provided in this application.

[0013] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0016] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0017] The teleprompter control method provided in this application is mainly applied to head-mounted teleprompter devices, such as augmented reality (AR) glasses, virtual reality (VR) glasses, and smart head-mounted helmets. This head-mounted teleprompter device aims to provide speakers, broadcasters, hosts, and other users who need to view teleprompter content in real time with a natural and seamless teleprompter interaction method.

[0018] The head-mounted teleprompter device incorporates a multi-microphone array to capture the audio signal of the user's (speaker's) input speech and supports sound source localization. For example, four microphones can be configured in a distributed layout (such as one microphone each on the left temple, right temple, the center of the top edge of the frame, and the center of the bottom edge of the frame) to achieve spatial positioning capabilities.

[0019] A camera is incorporated into this head-mounted teleprompter device to capture image sequences of the wearer's lip movements for lip movement feature extraction and verification. The camera's image acquisition area completely covers the wearer's mouth region.

[0020] In addition, the head-mounted teleprompter can display the teleprompter text through virtual display, physical display, or a combination of both.

[0021] Specifically, the virtual display method can be achieved through a built-in micro-display and optical system. By combining the micro-display and optical system, a virtual teleprompter text floating window is generated in front of the user's field of vision. The transparency, size, and position of this window can be adjusted to avoid obstructing the view.

[0022] Physical display can be achieved through an external physical display screen (such as a teleprompter, tablet, or monitor). The head-mounted teleprompter connects to the external physical display screen via a communication interface. The head-mounted teleprompter then sends teleprompter control commands and text content to the external physical display screen, which displays the teleprompter text. The communication between the head-mounted teleprompter and the external physical display screen can be via wireless communication methods such as Bluetooth or WiFi, or via wired communication methods such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), or other communication protocols.

[0023] Please refer to Figure 1, which is a flowchart of a first embodiment of a word prompting control method provided in this application.

[0024] As shown in Figure 1, the prompting control method includes steps S101 to S103.

[0025] S101. Upon receiving the current input voice, the validity of the current input voice is verified based on a preset verification method. In one embodiment, after the user wears and activates the head-mounted teleprompter device (taking AR glasses as an example), the ambient sound is collected in real time through the multi-microphone array built into the head-mounted teleprompter device, and the lip image is captured in real time through the camera built into the head-mounted teleprompter device.

[0026] When the multi-microphone array detects the current input speech, it first needs to verify the validity of the input speech to determine whether the speaker is the wearer of the head-mounted teleprompter device, thus eliminating interference from environmental noise, other people's voices, etc. If it is determined that the speaker is the wearer of the head-mounted teleprompter device, subsequent teleprompter control operations can be performed; if it is determined that the speaker is not the wearer of the head-mounted teleprompter device, the current input speech is filtered out and no response is given, thereby improving the accuracy of teleprompter control.

[0027] For example, validity verification methods may include spatial validity verification, voiceprint similarity verification, and lip movement feature verification, wherein lip movement feature verification includes semantic similarity verification and temporal axis deviation verification. The validity verification method for the current input speech may include at least one of spatial validity verification, voiceprint similarity verification, semantic similarity verification, and temporal axis deviation verification. One method may be selected as the validity verification method for the current input speech, or multiple methods may be selected in combination (e.g., any two, any three, or all combined) as the validity verification method for the current input speech.

[0028] When using a single validation method to validate the current input speech, the current input speech only needs to meet the pass condition of that validation method to be considered valid input.

[0029] When validating the current input speech using multiple validity verification methods, the current input speech must simultaneously meet the passing conditions of each of the combined validity verification methods to be considered valid input; if the current input speech does not meet the passing conditions of any validity verification method, it will be considered invalid input. In one embodiment, validating the current input speech based on preset verification methods includes: performing spatial validity verification on the sound source location of the current input speech based on a preset valid region; after the spatial validity verification passes, performing voiceprint similarity verification on the speaker's voiceprint template and the voiceprint features of the current input speech; after the feature similarity verification passes, performing semantic similarity verification on the speaker's lip movement feature sequence and the speech text of the current input speech, and performing time axis deviation verification on the lip movement feature sequence and the speech text; when both the semantic similarity verification and the time axis deviation verification pass, the current input speech is determined to be valid input.

[0030] The specific implementation methods for each validity verification provided in the embodiments of this application are described in detail below.

[0031] Furthermore, as shown in FIG2, based on the embodiment shown in FIG1 above, step S101 specifically includes steps S1011 to S1014.

[0032] S1011. Based on the preset effective area, the spatial validity of the sound source location of the current input speech is verified.

[0033] The preset effective area is set near the wearer's mouth. For example, the preset effective area can be set as a spherical area with a radius of 10-15cm centered on the wearer's mouth. The spatial position and range of this preset effective area in the world coordinate system can be adaptively adjusted according to parameters such as the wearing position of the head-mounted teleprompter and the wearer's head size.

[0034] The built-in processor system of a head-mounted teleprompter can calculate the location of a sound source in real time based on the audio signal of the current input speech collected by a multi-microphone array. The core of sound source localization is determining the location of the sound source by measuring the time difference of arrival (TDOA) of sound waves at each microphone. Sound waves travel at approximately 340 meters per second in air, and when a sound source emits sound, the sound waves travel at a certain speed to each microphone. Because the distance between the sound source and different microphones varies, the time it takes for the sound waves to arrive at each microphone also varies. By measuring these time differences, the location of the sound source can be inferred.

[0035] In one embodiment, at least two audio signals of the current input speech are acquired; based on the arrival time difference of each audio signal, the sound source location of the current input speech is calculated; it is determined whether the sound source location is located within the preset effective area; if the sound source location is located within the preset effective area, the spatial validity verification of the current input speech is determined to be successful.

[0036] To achieve sound source localization, a multi-microphone array, including at least two microphones, needs to be configured in a head-mounted teleprompter device (such as AR glasses). The number and arrangement of the microphones affect the accuracy and complexity of the localization. For example, multi-microphone array arrangements can include linear arrays, two-dimensional arrays, and three-dimensional arrays. In a linear array, multiple microphones are arranged along a straight line, suitable for one-dimensional spatial localization; in a two-dimensional array, multiple microphones are distributed on a plane, suitable for two-dimensional spatial localization; and in a three-dimensional array, multiple microphones are distributed in three-dimensional space, suitable for three-dimensional spatial localization.

[0037] To improve the accuracy of sound source location, it is preferable to configure a multi-microphone array consisting of a larger number of microphones (such as 3, 4, 8 or 16).

[0038] Each microphone in the multi-microphone array acquires the audio signal of the input speech through an independent audio acquisition circuit. After each microphone acquires the audio signal, preprocessing operations such as pre-emphasis, noise reduction, and filtering are performed on each acquired audio signal to eliminate interference from environmental noise and head-mounted teleprompter noise, thereby providing accurate sound source location.

[0039] Since the speed of sound is known (approximately 340 m / s under standard atmospheric pressure), by measuring the time difference of arrival of the current input speech at different microphones, the length difference of the sound propagation path can be calculated, thereby determining the position of the sound source of the current input speech relative to the multi-microphone array. In practical applications, due to limitations in noise and sampling accuracy, directly measuring the absolute time difference of arrival is very difficult. The cross-correlation method is typically used to estimate the time difference of arrival, i.e., calculating the cross-correlation function between the reference microphone signal and other microphone signals; the time shift corresponding to the peak value of this cross-correlation function is the estimated time difference of arrival.

[0040] Taking a multi-microphone array consisting of 4 microphones as an example, the 4 microphones ( , , , Simultaneously acquire four audio signals. Assume that sound source S emits the current input speech, and the times when the current input speech arrives at the four microphones are respectively... , , , .

[0041] First, it is necessary to calculate the time difference of arrival between each pair of microphones. Let's assume... Using the reference microphone, calculate the differences between other microphones and... Time difference between them:

[0042]

[0043]

[0044] in, , , , These represent the arrival of the currently input voice. , , , The timing of the four microphones; This indicates that the current input voice has reached the microphone. and reaching the microphone Time difference; This indicates that the current input voice has reached the microphone. and reaching the microphone Time difference; This indicates that the current input voice has reached the microphone. and reaching the microphone The time difference. These time differences represent the time difference between the arrival of the speech signal from the sound source S at the microphone and the arrival of the current input speech signal. , , Time relative to arrival at the microphone The time delay.

[0045] Assuming the speed of sound The speed is 340 meters per second. Convert the time difference to a distance difference:

[0046]

[0047]

[0048] in, Indicates the speed of sound. This indicates that the current input voice has reached the microphone. and reaching the microphone Time difference; This indicates that the current input voice has reached the microphone. and reaching the microphone Time difference; This indicates that the current input voice has reached the microphone. and reaching the microphone The time difference. This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference; This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference; This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference.

[0049] In three-dimensional space, based on the distance difference between the current input speech and each microphone, and combined with the geometric position of each microphone, the three-dimensional coordinates of the sound source S can be calculated using a geometric model (such as triangulation).

[0050] Specifically, assuming the coordinates of the four microphones are as follows: , , , The coordinates of the sound source S are S( Based on distance difference , , Establish the following system of equations:

[0051] in, microphone The three-dimensional coordinates microphone The three-dimensional coordinates microphone The three-dimensional coordinates microphone The three-dimensional coordinates Represents the three-dimensional coordinates of the sound source S. This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference; This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference; This indicates that the current input voice has reached the microphone. The current input voice reaches the microphone The distance difference.

[0052] The above system of equations is nonlinear and typically requires numerical methods to solve. Common methods include least squares and nonlinear optimization.

[0053] Taking the least squares method as an example, we define an objective function, such as the sum of the squares of the differences between all measured distance differences and the theoretical distance differences:

[0054]

[0055]

[0056] Minimize the above objective function using optimization algorithms (such as gradient descent, Newton's method, etc.). Thus, the three-dimensional coordinates of the sound source location can be solved. .

[0057] The calculated sound source location is compared with the preset effective area. For example, for a spherical preset effective area, it is determined whether the distance from the sound source S to the center point of the preset effective area is less than the preset radius; for a conical preset effective area, it is determined whether the azimuth and elevation angles of the sound source S are within the preset angle range, and whether the distance between the sound source S and the vertex of the conical effective area is less than the preset maximum distance.

[0058] If the sound source location is within the preset valid area, then the sound source location of the input speech is considered to be within the preset valid area. If valid, output a "spatial validity verification passed" signal; if the sound source location is not within the preset valid area, then the sound source location of the input speech is considered valid. Invalid, output a "Spatial validity verification failed" signal.

[0059] This embodiment, by configuring a multi-microphone array and utilizing the Time Difference of Arrival (TDOA) algorithm, can accurately calculate the sound source location of the current input speech and compare it with a preset effective area near the wearer's mouth. This enables spatial validity verification of the current input speech, effectively filtering out invalid noise from the environmental background, other speakers, and the head-mounted teleprompter itself. It ensures that only the input speech emitted by the wearer and located within the effective area of ​​the mouth can be recognized and processed, thus improving the accuracy and anti-interference capability of teleprompter control.

[0060] S1012. Verify the similarity between the speaker's voiceprint template and the voiceprint features of the current input speech.

[0061] In one embodiment, voiceprint registration and storage can be performed when the wearer first uses the head-mounted recording device. Generally, the speaker can be asked to read a pre-set text, such as a sequence of numbers, phrases, or sentences, in a quiet environment, and multiple speech samples can be collected.

[0062] The acquired speech samples are preprocessed, including noise reduction, DC bias removal, and gain adjustment. Feature vectors are extracted from the speech sample signals using speaker feature extraction algorithms (such as Mel-Frequency Cepstral Coefficient (MFCC), Linear Predictive Coding (LPC), and deep embedding features). The feature vectors from multiple speech samples are then clustered or averaged to generate a comprehensive speaker template. The generated speaker's speaker template is stored in encrypted form on the device's local storage or a cloud server.

[0063] Further, the voiceprint features of the current input speech are extracted; the voiceprint similarity between the voiceprint features of the current input speech and the speaker's voiceprint template is calculated; when the voiceprint similarity between the voiceprint features of the current input speech and the voiceprint template is greater than or equal to a preset voiceprint similarity threshold, the voiceprint similarity verification of the current input speech is determined to be successful.

[0064] The current input speech signal undergoes preprocessing. Pre-emphasis is applied to enhance high-frequency components. The speech signal is divided into short time frames, and a window function, such as a Hamming or Hanning window, is applied to each frame. The window function is then applied to each frame, multiplying each sample point by its value. A Fast Fourier Transform (FFT) is performed on each windowed frame to convert the time-domain signal to the frequency-domain signal. The squared amplitude of the frequency-domain signal is calculated to obtain the spectrum of the current input speech. The spectrum is then passed through a Mel filter bank to calculate the Mel spectrum. The logarithm of the Mel spectrum is taken, followed by a Discrete Cosine Transform (DCT) to obtain the MFCC feature vector, which serves as the speaker signature feature vector for the current input speech.

[0065] Feature similarity calculation methods, such as cosine similarity and Euclidean distance, are used to calculate the feature similarity between the pre-stored speaker's voiceprint template vector and the voiceprint feature vector of the current input speech. Taking cosine similarity as an example:

[0066] in, Represents the speaker's voiceprint template vector. This represents the speaker's feature vector for the current input speech. This represents the feature similarity between the speaker's voiceprint template vector and the voiceprint feature vector of the current input speech. The cosine similarity value is between -1 and 1; a higher similarity value indicates a higher similarity between the speaker's voiceprint template vector and the voiceprint feature vector of the current input speech.

[0067] A voiceprint similarity threshold can be set, such as a cosine similarity threshold of 0.8 and an Euclidean distance threshold of 0.5. If the feature similarity between the calculated speaker's voiceprint template vector and the voiceprint feature vector of the current input speech is greater than or equal to the voiceprint similarity threshold, then the current input speech is considered to match the voiceprint template, meaning that the current input speech was spoken by the wearer of the head-mounted teleprompter, and the feature similarity verification of the current input speech is deemed successful. Conversely, if the feature similarity between the calculated speaker's voiceprint template vector and the voiceprint feature vector of the current input speech is less than the voiceprint similarity threshold, then the current input speech is considered to not match the voiceprint template, meaning that the current input speech was not spoken by the wearer of the head-mounted teleprompter, and the feature similarity verification of the current input speech is deemed to have failed.

[0068] This embodiment achieves accurate verification of the speaker's identity by pre-registering and storing the speaker's voiceprint template, and extracting the voiceprint features of the current speech for similarity comparison each time a voice input is made. This effectively prevents others from accidentally or maliciously using the prompting function, ensuring that only authorized users can control the scrolling of the prompting text, thus improving the system's security and privacy and providing reliable identity authentication for prompting control.

[0069] S1013. Perform semantic similarity verification on the speaker's lip movement feature sequence and the speech text of the current input speech.

[0070] During semantic similarity verification, a sequence of lip movement images of the speaker (wearer) is captured using the built-in camera of a head-mounted teleprompter (which can be positioned at the front of the head-mounted recording device, with the acquisition area at least covering the wearer's mouth area). A pre-trained lip-reading model is then used to extract lip movement feature sequences. These lip movement feature sequences are then decoded into predicted lip movement text. The semantic similarity between the predicted lip movement text and the recognized speech text is calculated and compared to a preset semantic similarity threshold. Based on the comparison result, it is determined whether the lip movement content and the speech content are semantically consistent.

[0071] In one embodiment, a sequence of lip movement images of the speaker is acquired; based on a lip-reading recognition model, a lip movement feature sequence is extracted from the lip movement image sequence; the lip movement feature sequence is decoded into lip movement prediction text; the semantic similarity between the lip movement prediction text and the speech text is calculated; when the semantic similarity is greater than or equal to a preset semantic similarity threshold, the semantic similarity verification of the current input speech is determined to be successful.

[0072] Generally, the multi-microphone array and camera are started synchronously. While the multi-microphone array captures the current input speech, the camera simultaneously captures a sequence of images of the wearer's lip movements. The multi-microphone array and camera use a unified system clock source to ensure that the timestamps of the current input speech and lip movement images have a consistent time base.

[0073] For the current input speech, obtain its start and end timestamps. Considering the delays in sound propagation and signal processing, a compensation value (e.g., 30 milliseconds) can be set to compensate for the start and end times of the lip movement image sequence.

[0074] Based on the timestamp of the current input speech, find the image frame corresponding to the timestamp of the current input speech from the historical video stream captured by the camera. Considering the propagation delay of the speech signal, an image frame whose image acquisition time is earlier than the timestamp of the current input speech by a certain amount of time (e.g., a compensation value of 30 milliseconds) can be selected as the starting frame of the lip movement image sequence.

[0075] Similarly, the image frame corresponding to the end time of the current input speech acquisition, or the image frame whose image acquisition time is later than the end time of the current input speech acquisition (e.g., a compensation value of 30 milliseconds), is used as the end frame of the lip movement image sequence.

[0076] Using the compensation value as an additional time buffer helps ensure that the lip movement image sequence fully covers the time range related to the current input speech, thus avoiding image sequence truncation problems caused by time synchronization errors.

[0077] After determining the start and end frames, all image frames between these two frames are extracted from the video stream captured by the camera to form a complete sequence of lip motion images.

[0078] Choose a well-trained lip-reading model, such as a deep learning-based 3D-CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), or its variants (e.g., LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit)). For example, 3D-CNN can capture the spatiotemporal information of lip movements, while LSTM can handle the temporal dependencies of sequence data. The lip-reading model can be pre-trained on a large-scale lip-reading dataset to improve its generalization ability and recognition accuracy. In practical applications, the pre-trained lip-reading model can be fine-tuned according to the specific application scenario, optimizing the model parameters to adapt to the specific usage environment and speaker.

[0079] This embodiment acquires lip movement image sequences and decodes them into lip movement prediction text using a pre-trained lip reading recognition model. Then, it performs semantic similarity calculation with the speech recognition text, achieving dual verification of the authenticity of the speech content. This effectively prevents recording attacks or speech synthesis attacks, ensuring that the prompting control only responds to the wearer's actual lip movements and speech content. This enhances the security and anti-deception capabilities of the prompting control system, providing a solid guarantee for the reliable operation of the prompting function.

[0080] Before inputting the lip motion image sequence into the lip-reading recognition model, the lip motion image sequence is preprocessed. First, for each frame in the sequence, a face detection model (such as RetinaFace) is used to locate the face region. Each frame is input into the face detection model, which outputs the bounding box of the face (including the coordinates of the top left and bottom right corners) and the positions of facial key points (such as the eyes, nose, and corners of the mouth), thus extracting the face region. Then, a facial key point detection algorithm (such as Dlib's 68-point keypoint detector or a deep learning-based keypoint detection model) is used to locate lip key points within the face region, including the corners of the mouth, cupid's bow, and valley of the lips. Typically, 20 key points are used to describe the contour and shape of the lips.

[0081] For each frame of the image, extract the coordinates of the lip key points. Define a standard lip position, usually a rectangular region, the size and shape of which are determined according to the actual application requirements. For example, a square region with an aspect ratio of 1:1 can be defined as the standard lip position.

[0082] Select points from the lip keypoints that correspond to the keypoints in the standard lip position. For example, the upper left corner of the standard lip position can be selected from the left corner of the mouth, and the upper right corner can be selected from the right corner of the mouth, and so on. Store the coordinates of the selected lip keypoints and the coordinates of the keypoints in the standard lip position as two NumPy arrays respectively. Based on the correspondence between the lip keypoints and the standard lip position, calculate the affine transformation matrix using the affine transformation function of OpenCV (a cross-platform computer vision library). Use the calculated affine transformation matrix to transform the lip region of each frame of the image, aligning the lip region to the standard position.

[0083] Cropped lip images are generated based on the coordinates of the aligned standard lip positions. The cropped areas are ensured to include the complete lip contour and be of uniform size. The cropped lip images are then normalized, adjusting the image size to a uniform resolution (e.g., 96×96 pixels) and normalizing the pixel values ​​to [0,1] or [-1,1] to improve model training and inference efficiency. The cropped and normalized lip images are then arranged chronologically to form a continuous lip image sequence.

[0084] A sequence of lip movement images is input into a lip-reading model, which extracts feature vectors for each frame using convolutional and pooling layers. For example, when using a 3D-CNN model, the model extracts spatiotemporal features of lip movements through multiple convolutional and pooling layers, generating a high-dimensional feature vector sequence. This feature vector sequence is then processed by a time-series processing layer (such as LSTM or GRU) to capture the temporal features of lip movements. For instance, LSTM layers can effectively handle long-term dependencies in sequence data, generating a lip movement feature sequence containing temporal information for spatiotemporal lip movement analysis.

[0085] The lip movement feature sequence is input into a pre-trained decoder (such as a sequence-to-sequence (Seq2Seq) model based on an attention mechanism) for decoding. The decoder dynamically focuses on key parts of the lip movement feature sequence through an attention mechanism, and gradually generates text output, thereby mapping the lip movement feature sequence into a text sequence.

[0086] Attention-based sequence-to-sequence models employ an encoder-decoder architecture. The encoder is typically a recurrent neural network (RNN, such as LSTM or GRU) or a convolutional neural network (CNN) used to extract features from the lip movement sequence. The decoder is also typically an RNN (such as LSTM or GRU) or a Transformer, whose initial hidden states are usually initialized by the final hidden states of the encoder. The goal of the decoder is to progressively generate a text sequence based on the context vectors generated by the encoder.

[0087] Specifically, the lip movement feature sequence is input into the encoder, which processes the sequence frame by frame, updating its hidden states. For an RNN encoder, the hidden state at each step is passed to the next, and the final hidden state (or all hidden states) is output as a context vector. For a Transformer encoder, the lip movement feature sequence is divided into multiple sub-sequences, processed in parallel using a self-attention mechanism to generate a context vector. The initial input to the decoder is typically a special start marker (e.g., ...). <start>The sequence signifies the start of text generation. At each step of text generation, an attention mechanism is used to calculate the similarity between the decoder's current hidden state and the encoder's output context vector, generating attention weights that represent the importance of each frame in the input sequence to the current decoding step. The attention weights are then used to weight and sum the encoder's output context vector, generating a comprehensive context vector containing the most relevant input information for the current decoding step. The decoder combines the current hidden state and the comprehensive context vector, generating a probability distribution for the next word through a fully connected layer. The word with the highest probability is selected as the output of the current step and used as the input for the next step. This process is repeated until the generated text reaches a preset maximum length or an end marker (such as...) is generated. <end>).

[0088] The text sequence output by the decoder is typically a sequence of word indices, which represent the positions of words or characters predicted by the attention-based sequence-to-sequence model within a predefined vocabulary. This predefined vocabulary contains all possible words or characters and their corresponding indices. After the decoder outputs the text sequence (the sequence of word indices), this predefined vocabulary can be used to map the word index sequence to actual words or characters, generating lip movement prediction text.

[0089] The objects of speech similarity calculation are the lip movement predicted text and the speech text corresponding to the current input speech. Therefore, it is also necessary to convert the current input speech into speech text in order to perform semantic similarity verification.

[0090] An ASR (Automatic Speech Recognition) system can be used to convert the current input speech into speech-to-text. Specifically, conventional speech preprocessing operations such as noise reduction, DC bias removal, and gain adjustment are performed on the current input speech signal. The preprocessed speech signal is then pre-emphasized to enhance high-frequency components. The speech signal is divided into short time frames, and a window function, such as a Hamming or Hanning window, is applied to each short time frame. The window function is applied to each frame, and each sample point in each frame is multiplied by the window function value. A Fast Fourier Transform (FFT) is performed on each windowed speech signal to convert the time-domain signal to the frequency-domain signal. The squared amplitude of the frequency-domain signal is calculated to obtain the spectrum corresponding to the current input speech. The spectrum corresponding to the current input speech is then passed through a Mel filter bank to calculate the Mel spectrum. The logarithm of the Mel spectrum is taken, and then a Discrete Cosine Transform is performed to obtain the speech feature sequence of the current input speech.

[0091] Choose a well-trained ASR model, such as an end-to-end deep learning-based model. Input the speech feature sequence of the current input speech into the ASR model. The ASR model first encodes the input speech feature sequence through an encoder, converting it into a high-dimensional context vector representation. The encoder typically consists of multiple convolutional layers, self-attention layers, and feedforward networks, effectively capturing both local and global features of the speech signal. Then, the ASR model's decoder generates text word by word in an autoregressive manner based on the encoder's output. During decoding, the ASR model calculates the probability distribution of each word in the vocabulary and selects the word with the highest probability as the output of the current time step, ultimately outputting the original text sequence corresponding to the current input speech.

[0092] After the ASR model outputs the original text sequence, post-processing can be performed on the original text sequence to improve the readability and accuracy of the speech text. Among them, the post-processing includes language model rescoring and punctuation restoration, etc. Among them, language model rescoring uses a pre-trained language model to re-score the original text sequence output by the ASR model and selects the text that best conforms to language habits; punctuation restoration adds punctuation marks such as commas and periods to the text by means of rules or model prediction. Finally, the speech text corresponding to the current input speech is output. This speech text not only contains the recognized text content, but also records the timestamp information corresponding to each word.

[0093] After obtaining the lip movement prediction text and the speech text, preprocess the two texts to ensure the accuracy of semantic similarity calculation. The preprocessing steps may include: converting all texts to lowercase to eliminate case differences; removing punctuation marks and special characters from the texts; and normalizing entities such as numbers and dates, for example, uniformly representing "100” as "one hundred”.

[0094] Select a model that can effectively represent the semantics of the text, such as BERT (Bidirectional Encoder Representation from Transformers, pre-trained language representation model), GPT (Generative Pre-trained Transformer, generative pre-trained model).

[0095] Input the preprocessed speech text and lip movement prediction text into the pre-trained language model respectively, convert the lip movement prediction text and the speech text into semantic vectors respectively, and obtain the semantic vectors of the speech text and the semantic vectors of the lip movement prediction text. It can be understood that extracting the semantic vectors of the text by the language model is a conventional technical means in the art, and this embodiment will not be elaborated here.

[0096] Use feature similarity measurement methods such as cosine similarity and Euclidean distance to calculate the semantic similarity between the semantic vectors of the speech text and the semantic vectors of the lip movement prediction text.

[0097] For example, the cosine similarity calculation formula is:

[0098] Among them, represents the semantic similarity between the speech text and the lip movement prediction text, represents the semantic vector of the speech text, represents the semantic vector of the lip movement prediction text.

[0099] A semantic similarity threshold can be set, such as 0.8, which means that the speech text and the lip movement prediction text are considered to be semantically similar only when the semantic similarity is greater than or equal to 0.8.

[0100] If the calculated semantic similarity is greater than or equal to the preset semantic similarity threshold, the speech text and the predicted lip movement text are considered to be semantically matched, and the semantic similarity verification of the current input speech is deemed to have passed; if the calculated semantic similarity is less than the preset semantic similarity threshold, the speech text and the predicted lip movement text are considered to be semantically mismatched, and the semantic similarity verification of the current input speech is deemed to have failed.

[0101] This embodiment converts lip movement image sequences into lip movement prediction text and performs semantic similarity calculation with speech recognition text, thereby achieving multimodal verification of the authenticity of speech content. It can effectively distinguish between real lip movements and speech synchronization, as well as potential deceptive behaviors, and improve the anti-attack capability and reliability of the prompting system.

[0102] S1014. Perform time axis deviation verification on the speaker's lip movement feature sequence and the speech text of the current input speech.

[0103] During the time axis deviation verification process, a lip movement time axis is constructed based on the acquisition timestamps of the lip movement image sequence, and a speech time axis is constructed based on the input timestamps of the speech signal. The time axis deviation between these two time axes is calculated and compared with a preset deviation threshold. Based on the comparison result, it is determined whether the lip movement and speech occur synchronously at the same time point.

[0104] In one embodiment, based on the image acquisition time of the lip movement image sequence, the lip movement time axis corresponding to the lip movement feature sequence is determined; based on the signal input time of the current input speech, the speech time axis corresponding to the speech text is determined; the time axis deviation between the lip movement time axis and the speech time axis is calculated; when the time axis deviation is less than or equal to a preset deviation threshold, the time axis deviation of the current input speech is determined to be verified as passed.

[0105] When acquiring lip movement images, the acquisition timestamp of each image frame is recorded. The acquisition timestamp can be a relative time to the start of the video stream or an absolute time (such as UTC (Coordinated Universal Time)). Based on the acquisition timestamp of each image frame, a lip movement timeline corresponding to the lip movement feature sequence is constructed. The lip movement timeline is a time series that records the acquisition time corresponding to each image frame in the lip movement images.

[0106] Similarly, when acquiring the current input speech, the input timestamp of the speech signal is recorded. The input timestamp of the speech signal can be the start and end times of a speech segment, or it can be the sampling time sequence of the speech signal. Based on the input timestamp of the speech signal, a speech timeline corresponding to the speech text is constructed. The speech timeline records the start and end times of each word or phoneme in the speech text.

[0107] Align the lip movement timeline and the speech timeline to ensure they have the same time base. If the lip movement timeline and speech timestamp are relative times, convert them to absolute times under the same time base. After alignment, the lip movement timeline and speech timeline can be represented as two separate time series.

[0108] Calculate the time deviation between the aligned lip movement timeline and the speech timeline. The time deviation can be the average time difference between two time series, or the time difference at a specific time point (such as the start point or the end point).

[0109] For example, calculating the start time of the lip movement time axis. Start time of the speech timeline Deviation between:

[0110] in, This represents the time deviation between the start time of the lip movement time axis and the start time of the speech time axis. Indicates the start time of the speech timeline. This indicates the start time of the lip movement timeline.

[0111] Simultaneously, calculate the end time of the lip movement time axis. Start time of the speech timeline Deviation between:

[0112] in, This represents the time deviation between the end time of the lip movement time axis and the end time of the speech time axis. Indicates the end time of the speech timeline. This indicates the end time of the lip movement timeline. This indicates the total number of time points on the lip movement timeline or speech timeline.

[0113] Alternatively, the average time deviation between the lip movement time axis and the speech time axis can be calculated. That is, calculate the time axis deviation for each time pair on the lip movement time axis and the speech time axis, and divide the time axis deviation of all time pairs by the total number of all time pairs to obtain the average time deviation. .

[0114] A suitable time axis deviation threshold can be set according to the application scenario. This time axis deviation threshold represents the maximum allowable time deviation between the lip movement time axis and the speech time axis.

[0115] The calculated time axis deviation is compared with a preset time axis deviation threshold. If the time axis deviation is less than or equal to the threshold, it is considered that the current input speech and the lip movement of the wearer of the teleprompter device occur simultaneously. Therefore, the current input speech is from the wearer of the teleprompter device, and the time axis deviation verification of the current input speech is deemed successful. Conversely, if the time axis deviation is greater than the threshold, it is considered that the current input speech and the lip movement of the wearer of the teleprompter device do not occur simultaneously. Therefore, the current input speech is not from the wearer of the teleprompter device, and the time axis deviation verification of the current input speech fails.

[0116] For example, setting the time axis deviation threshold to 100 milliseconds means that the maximum allowable time deviation between the lip movement time axis and the speech time axis is 100 milliseconds, as shown in the time deviation value above ( , or average time deviation A time interval of less than or equal to 100 milliseconds is required to determine if the time axis deviation of the current input speech has passed verification.

[0117] This embodiment constructs a lip movement time axis and a speech time axis, and accurately calculates the time axis deviation between the two, thereby achieving rigorous verification of the synchronization between speech and lip movement. This effectively prevents recording playback attacks or speech synthesis attacks, ensuring that the prompting control only responds to the wearer's real-time, synchronous lip movements and speech, thus improving the security and anti-spoofing capabilities of the prompting control system.

[0118] S102. When the current input speech is determined to be valid input, the speech text of the current input speech is matched with the segment to be prompted in the prompting text to determine the text matching degree between the speech text and the segment to be prompted.

[0119] After validating the current input speech through at least one of the following methods: spatial validity verification, voiceprint similarity verification, lip movement feature verification (including semantic similarity verification), and time axis deviation verification, and determining that the current input speech is valid input, the speech text corresponding to the current input speech converted by ASR technology is obtained.

[0120] Head-mounted teleprompter devices load teleprompter text from local storage or cloud servers. This teleprompter text can be a speech manuscript, a live broadcast script, or meeting minutes, etc.

[0121] Based on the current prompting progress, determine the next prompting paragraph in the prompting text. This prompting paragraph can be the next line of text after the last line currently displayed in the prompting text, or it can be the next paragraph of the currently displayed text. For example, if the prompting text is displayed by scrolling line by line, and only one line is loaded each time, then the prompting paragraph is the next line of text after the currently displayed line. If the prompting text is displayed by scrolling paragraph by paragraph, if the current prompting paragraph has scrolled to the end of its paragraph, then the next prompting paragraph in the prompting text is the prompting paragraph; if the current prompting paragraph has not scrolled to the end of its paragraph, then the remaining content of the current prompting paragraph that has not yet been displayed is used as the prompting paragraph.

[0122] Extract the prompting paragraph from the prompting text. This prompting paragraph is the content that the speaker needs to read or express next. It is used to perform character-level text matching calculations with the speech text. If the extracted prompting paragraph is too long, the first K characters (e.g., 15 characters) can be truncated as the matching field to improve matching efficiency.

[0123] For example, a sliding window matching algorithm or a Longest Common Subsequence (LCS) matching algorithm can be used to perform word-by-word matching between the speech text and the proposed word segment. The sliding window matching algorithm uses the speech text as a sliding window, sliding across the proposed word segment and calculating the number of matching words at each position. The LCS matching algorithm uses a dynamic programming algorithm to calculate the longest common subsequence between the speech text and the proposed word segment.

[0124] During the word-by-word matching process, each character is strictly compared for similarity, including Chinese characters, letters, numbers, punctuation marks, and capitalization. If a character in the spoken text matches the character in the prompting segment, the character is counted as a match; if a character in the spoken text does not match the character in the prompting segment, the character is counted as a mismatch.

[0125] The number of successfully matched characters is counted. If sliding window matching is used, the maximum number of matched characters is taken; if LCS matching is used, the LCS length is taken. The total number of characters in the spoken text is counted, and the text matching degree between the spoken text and the proposed word segment is calculated using a proportional method.

[0126] The text matching score ranges from 0 to 100%, with a higher value indicating a higher degree of text matching.

[0127] S103. When the text matching degree is greater than or equal to the preset matching degree threshold, control the scrolling of the prompt text.

[0128] In one embodiment, the matching threshold can be dynamically set using a piecewise function based on the total number of characters in the recognized speech text.

[0129] In one embodiment, the total number of characters in the speech text is counted; based on the total number of characters in the speech text, the matching degree threshold corresponding to the current input speech is determined.

[0130] For example, if the recognized speech text is a single-character recognition text, that is, the recognized speech text contains only one character, then the matching threshold can be set to 100%, that is, the recognized character must be exactly the same as the first character of the segment to be promoted in the prompting text, to prevent short coughs, throat clearings or environmental noise from being mistakenly recognized as single characters, which would cause the prompting text to scroll erroneously.

[0131] When the recognized speech text contains 2 to 5 characters, a matching threshold of 60% can be set. For example, if 3 characters are recognized, at least 2 characters must match the prompting text. This allows for minor pronunciation errors, slips of the tongue, or recognition errors when the speaker reads short sentences, improving the fault tolerance of the prompting control system.

[0132] When the number of characters in the recognized speech text is greater than 5, a matching threshold of 80% can be set. For example, if 10 characters are recognized, at least 8 characters should match the prompt text. For long sentence readings, a higher matching score is required to ensure that the speaker is indeed reading the prompt content, rather than chatting or conversing with someone else.

[0133] Based on the above method, a matching threshold is dynamically set according to the total number of characters in the spoken text. When the calculated text matching score is greater than or equal to the threshold, it is determined that the speaker is currently reading the prompt, and the prompt text is scrolled to display the next prompt segment to the user. If the text matching score is less than the threshold, the current prompt remains paused to prevent misoperation due to misrecognition or partially similar speech.

[0134] This embodiment achieves precise tracking of the reading progress of the prompted content by matching the identified speech text with the prompted text word by word and using a dynamic threshold determination mechanism based on the number of words. It can intelligently distinguish between the wearer reading the prompted content and daily conversation or environmental noise, ensuring that the prompt scrolling is only triggered when the user accurately reads the preset content, thus improving the accuracy of prompt control and effectively preventing misoperation and accidental scrolling.

[0135] In one embodiment, when the text matching degree is greater than or equal to a preset matching degree threshold, the presence of subsequent input speech is detected within a preset delay confirmation window; if the subsequent input speech is not present within the delay confirmation window, the prompting text is controlled to scroll; if the subsequent input speech is detected within the delay confirmation window, the subsequent input speech and the current input speech are merged into a merged input speech; the text matching degree between the merged input speech and the prompting segment is calculated, and the prompting text is controlled to scroll when the text matching degree is greater than or equal to the preset matching degree threshold.

[0136] Once the matching degree between the spoken text and the prompted text reaches a preset threshold, the prompt scrolling operation can be delayed instead of being executed immediately. Instead, a delayed confirmation mechanism can be introduced to accommodate the speaker’s possible short pauses, thinking, or corrections during the reading process, thus avoiding erroneous scrolling caused by speech segmentation recognition.

[0137] A delayed confirmation window can be set, during which ambient sounds are continuously monitored to determine whether the speaker has completed voice input.

[0138] For example, the length of the delayed confirmation window can be preset to 500 milliseconds to ensure that the delayed confirmation window can effectively cover the natural pauses and thinking gaps during the user's reading, while avoiding affecting the prompting experience due to excessive waiting time. Generally, the length of the delayed confirmation window can be adaptively adjusted according to the speaker's speaking speed: when the speaker speaks quickly, the length of the delayed confirmation window can be appropriately shortened, such as setting its length to 300 milliseconds; when the speaker speaks slowly, the length of the delayed confirmation window can be appropriately lengthened, such as setting its length to 700 milliseconds, to adapt to the speaking habits of different speakers.

[0139] Within the delayed confirmation window, ambient sound is continuously monitored to detect the presence of subsequent voice input. A sound energy detection threshold can be set. If the ambient sound energy exceeds the threshold within the delayed confirmation window, it is determined that subsequent voice input has been detected; if no ambient sound energy exceeds the threshold within the delayed confirmation window, it is determined that no subsequent voice input has been detected.

[0140] If no further voice input occurs within the window, the current voice input is considered complete, and the prompting text can be scrolled. If further voice input is detected, it is concatenated with the current voice input to generate a merged voice input.

[0141] The speech text corresponding to subsequent speech inputs is concatenated with the speech text corresponding to the current input speech to obtain the speech text corresponding to the merged input speech. The text matching degree between the speech text corresponding to the merged input speech and the segment to be prompted is recalculated, and the matching degree threshold is updated according to the total number of characters in the speech text corresponding to the merged input speech. The recalculated text matching degree is compared with the updated matching degree threshold to determine whether to perform the prompting text scrolling operation, thereby improving the accuracy and smoothness of the prompting scrolling control. If the text matching degree is less than the matching degree threshold, the matching is considered to have failed, and the prompting text remains paused; if the text matching degree is greater than or equal to the matching degree threshold, the matching is considered to have succeeded, and the prompting scrolling operation is performed.

[0142] In one embodiment, the teleprompter can enter a pause state when it detects that the wearer of the head-mounted teleprompter has stopped speaking, remained silent for an extended period (e.g., more than 3 seconds), or actively issued a pause command (e.g., saying "pause"). In the pause state, the teleprompter control system continues to monitor the wearer's behavior. When a new input voice is detected that has passed all the aforementioned verifications, and the matched text of the detected input voice reaches a dynamic matching threshold with the text of the next segment to be teleprompted, the teleprompter text scrolling operation will resume. This ensures strict synchronization between the teleprompter scrolling and the speech, effectively distinguishing between audience-facing presentations and interactive scenarios.

[0143] This embodiment effectively handles natural pauses and sentence corrections during user reading by introducing a delayed confirmation window mechanism, avoiding erroneous scrolling caused by speech segmentation recognition. Simultaneously, by merging subsequent speech input and recalculating the matching degree, the accuracy and smoothness of the prompt scrolling are ensured. Furthermore, the behavior monitoring function in the prompt pause state intelligently distinguishes between speech and interactive scenarios, ensuring strict synchronization between the prompt scrolling and the speech content, thus enhancing the intelligence of the prompt control system and the user experience.

[0144] This embodiment provides a prompting control method. This method verifies the validity of the current input speech through a preset verification method, effectively filtering out unacceptable environmental noise, other people's voices, and invalid background noise. This ensures that only compliant input speech can be used in the prompting control process, reducing the probability of false triggering from the source. After determining that the current input speech is valid, the speech text of the current input speech is matched with the prompting segment in the prompting text. By calculating the text matching degree, it can accurately determine whether the speech content of the current input speech is consistent with the prompting content, avoiding erroneous scrolling caused by the wearer's idle chatter, coughing, or slips of the tongue. A preset matching degree threshold is used for judgment; the prompting scrolling operation is only executed when the text matching degree is greater than or equal to the threshold, further improving the accuracy of prompting control and ensuring that the prompting scrolling is strictly synchronized with the speech content of the current input speech, thereby achieving highly accurate prompting control.

[0145] Please refer to Figure 3, which is a schematic diagram of the structure of a first embodiment of a prompting control device provided in this application. The prompting control device is used to execute the aforementioned prompting control method.

[0146] As shown in Figure 3, the prompting control device 200 includes: a validity verification module 201, a text matching module 202, and a prompting control module 203.

[0147] The validity verification module 201 is used to verify the validity of the current input speech based on a preset verification method when the current input speech is received; the text matching module 202 is used to perform text matching between the speech text of the current input speech and the segment to be prompted in the prompting text when the current input speech is determined to be valid input, and to determine the text matching degree between the speech text and the segment to be prompted; the prompting control module 203 is used to control the scrolling of the prompting text when the text matching degree is greater than or equal to a preset matching degree threshold.

[0148] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the device and each module described above can be referred to the corresponding processes in the aforementioned teleprompter control method embodiments, and will not be repeated here.

[0149] The apparatus provided in the above embodiments can be implemented as a computer program that can run on the head-mounted teleprompter device shown in FIG4.

[0150] Please refer to Figure 4, which is a schematic block diagram of a head-mounted teleprompter device provided in an embodiment of this application. This head-mounted teleprompter device can be a server.

[0151] Referring to Figure 4, the head-mounted teleprompter device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0152] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any prompting control method.

[0153] The processor provides computing and control capabilities to support the operation of the entire head-mounted teleprompter device.

[0154] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When these computer programs are executed by a processor, the processor can perform any prompting control method.

[0155] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the head-mounted teleprompter device to which the present application is applied. A specific head-mounted teleprompter device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0156] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0157] In one embodiment, the processor is configured to run a computer program stored in a memory to perform the following steps: upon receiving current input speech, verifying the validity of the current input speech based on a preset verification method; when determining that the current input speech is valid input, performing text matching between the speech text of the current input speech and the prompting text segment to be prompted, and determining the text matching degree between the speech text and the prompting text segment; when the text matching degree is greater than or equal to a preset matching degree threshold, controlling the prompting text to scroll.

[0158] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the word prompting control methods provided in the embodiments of this application.

[0159] The computer-readable storage medium can be an internal storage unit of the head-mounted teleprompter device described in the foregoing embodiments, such as the hard drive or memory of the head-mounted teleprompter device. Alternatively, the computer-readable storage medium can be an external storage device of the head-mounted teleprompter device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the head-mounted teleprompter device.

[0160] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / end> < / start>

Claims

1. A method for controlling prompting, characterized in that, The method includes: upon receiving current input speech, validating the current input speech based on a preset verification method; when determining that the current input speech is valid input, performing text matching between the speech text of the current input speech and the segment to be prompted in the prompting text, and determining the text matching degree between the speech text and the segment to be prompted; and controlling the prompting text to scroll when the text matching degree is greater than or equal to a preset matching degree threshold.

2. The prompting control method according to claim 1, characterized in that, The validity verification of the current input speech based on the preset verification method includes at least one of the following: spatial validity verification of the sound source location of the current input speech based on a preset valid region; voiceprint similarity verification of the speaker's voiceprint template and the voiceprint features of the current input speech; semantic similarity verification of the speaker's lip movement feature sequence and the speech text of the current input speech; and time axis deviation verification of the speaker's lip movement feature sequence and the speech text of the current input speech.

3. The prompting control method according to claim 2, characterized in that, The step of verifying the spatial validity of the sound source location of the current input speech based on a preset effective area includes: acquiring at least two audio signals of the current input speech; calculating the sound source location of the current input speech based on the arrival time difference of each audio signal; determining whether the sound source location is located within the preset effective area; and if the sound source location is located within the preset effective area, then determining that the spatial validity verification of the current input speech has passed.

4. The prompting control method according to claim 2, characterized in that, The semantic similarity verification of the speaker's lip movement feature sequence and the speech text of the current input speech includes: acquiring the speaker's lip movement image sequence; extracting the lip movement feature sequence of the lip movement image sequence based on a lip reading recognition model; decoding the lip movement feature sequence into lip movement prediction text; calculating the semantic similarity between the lip movement prediction text and the speech text; and determining that the semantic similarity verification of the current input speech is passed when the semantic similarity is greater than or equal to a preset semantic similarity threshold.

5. The word prompting control method according to claim 4, characterized in that, The step of verifying the time axis deviation between the speaker's lip movement feature sequence and the speech text of the current input speech includes: determining the lip movement time axis corresponding to the lip movement feature sequence based on the image acquisition time of the lip movement image sequence; determining the speech time axis corresponding to the speech text based on the signal input time of the current input speech; calculating the time axis deviation between the lip movement time axis and the speech time axis; and determining that the time axis deviation verification of the current input speech is passed when the time axis deviation is less than or equal to a preset deviation threshold.

6. The prompting control method according to claim 1, characterized in that, Before controlling the scrolling of the prompting text when the text matching degree is greater than or equal to a preset matching degree threshold, the procedure includes: counting the total number of characters in the speech text; and determining the matching degree threshold corresponding to the current input speech based on the total number of characters in the speech text.

7. The prompting control method according to claim 1, characterized in that, After performing text matching between the current input speech and the prompting text to determine the text matching degree between the speech and the prompting text, the method further includes: when the text matching degree is greater than or equal to a preset matching degree threshold, detecting whether there is subsequent input speech within a preset delay confirmation window; if there is no subsequent input speech within the delay confirmation window, controlling the prompting text to scroll; if the subsequent input speech is detected within the delay confirmation window, merging the subsequent input speech and the current input speech into a merged input speech; calculating the text matching degree between the merged input speech and the prompting text, and controlling the prompting text to scroll when the text matching degree is greater than or equal to the preset matching degree threshold.

8. A prompting control device, characterized in that, The prompting control device includes: a validity verification module, used to verify the validity of the current input speech based on a preset verification method when the current input speech is received; a text matching module, used to perform text matching between the speech text of the current input speech and the prompting paragraph in the prompting text when the current input speech is determined to be valid input, and to determine the text matching degree between the speech text and the prompting paragraph; and a prompting control module, used to control the prompting text to scroll when the text matching degree is greater than or equal to a preset matching degree threshold.

9. A head-mounted teleprompter device, characterized in that, The head-mounted teleprompter device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the teleprompter control method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the prompting control method as described in any one of claims 1 to 7.