A wide-angle cast multi-modal AI quantum dot outdoor digital display device
By using a camera and multi-array microphones to collect information collaboratively, and combining a multimodal fusion strategy of speech and lip-reading recognition, the problem of speech recognition accuracy and continuity of outdoor digital display devices in complex environments has been solved, achieving stable multi-person interaction and efficient information transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI TOTAL SOLUTION ELEC CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing outdoor digital display devices are susceptible to noise in complex environments, and lack an effective compensation mechanism when confidence decreases, resulting in poor recognition accuracy and continuity, and making it impossible to achieve stable multi-person interaction.
The system employs a camera and multi-array microphones to collaboratively collect user information, establishes a dynamic recognition strategy based on speech confidence intervals, and utilizes multimodal fusion of speech and lip-reading recognition. By leveraging a confidence interval determination mechanism, it automatically triggers a remedial mechanism when speech confidence decreases, achieving character-level replacement and ensuring the recoverability and adaptability of the recognition process.
Stable multi-user interaction was achieved in complex environments, reducing the recognition error rate, improving the continuity and usability of the interaction, reducing the number of times users repeat their voices, and enhancing the reliability of the interaction in outdoor scenarios.
Smart Images

Figure CN121415781B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent display technology, specifically to a wide-angle projection multimodal AI quantum dot outdoor digital display device. Background Technology
[0002] With the rapid development of smart city construction, digital advertising, and public service terminals, outdoor digital display devices are gradually evolving towards intelligence and interactivity. Traditional outdoor display equipment only has a one-way display function and cannot achieve user-participatory interaction, which easily leads to problems such as low information transmission efficiency and weak user stickiness. Therefore, multimodal interactive display terminals based on interactive recognition of data such as voice, image, and gesture have become a major research direction in recent years.
[0003] In existing technologies, outdoor interaction typically relies on touchscreens or single-microphone voice recognition devices. However, touchscreen interaction requires user contact with the screen and is susceptible to hygiene conditions, screen damage, and environmental noise. While single-microphone voice recognition enables contactless operation, outdoor noise sources are diverse and their intensity is unstable. Vehicle noise, crowd noise interference, and wind noise can significantly reduce the accuracy of voice recognition, forcing users to repeat themselves multiple times, severely impacting the interactive experience. To improve recognition accuracy, some technologies use array microphones for beamforming and sound source focusing, but this still cannot fundamentally solve the problem of recognition accuracy when voice confidence decreases. Currently, some technologies attempt to introduce cameras for face detection to assist voice recognition. However, these technologies usually operate voice recognition and face detection in parallel, lacking a mechanism for dynamically switching recognition strategies based on confidence. This results in the voice result still being prioritized when voice confidence is low, leading to recognition errors. Summary of the Invention
[0004] To address the problem that existing speech recognition technologies are susceptible to environmental noise and lack effective compensation mechanisms when speech confidence is insufficient, this invention provides a wide-angle projection multimodal AI quantum dot outdoor digital display device.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] In one aspect, this application discloses a wide-angle multimodal AI quantum dot outdoor digital display device, including a camera, a multi-array microphone, a processor, and a UHD large screen.
[0007] The camera is used to capture images of the user's face;
[0008] Multi-array microphones are used to collect the user's audio signals;
[0009] The processor performs speech activity detection on the audio signal. If valid speech activity is detected, speech recognition is performed on the audio signal to obtain the initial text and speech confidence. The speech confidence is compared with a preset threshold range and the following decisions are made: (1) If the speech confidence is higher than the threshold range, the initial text is used as the recognition result; (2) If the speech confidence is within the threshold range, the initial text and speech confidence are divided into characters, and the characters of the initial text below the first threshold are replaced with the characters of the corresponding lip-shape recognition text and used as the recognition result; the lip-shape recognition text is the text whose confidence is higher than the second threshold after lip-shape recognition of the lip region, and the lip region is obtained by segmenting the face image through ROI; (3) If the speech confidence is lower than the threshold range, it is determined whether the confidence of the lip-shape recognition text is higher than the third threshold. If yes, the lip-shape recognition text is used as the recognition result; otherwise, the recognition failure is used as the recognition result; the recognition result generates the corresponding interaction strategy.
[0010] UHD large screens are used to execute interaction strategies and display corresponding content.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0012] 1. This application introduces a multimodal recognition mechanism into an outdoor digital display device, enabling stable human-computer interaction functions even when the user does not need to wear any devices and is in a natural interaction posture; in addition, it adopts a confidence interval determination mechanism, which does not immediately abandon recognition when the voice confidence decreases, but enters a multimodal recovery step, making the recognition process recoverable and adaptive, and meeting the usage needs of complex environments.
[0013] 2. Compared to traditional lip-reading recognition, which can only be used as an alternative to speech recognition, this application achieves a fine-grained substitution and supplementation mechanism through character-level fusion recognition. It can still retain the integrity and readability of the recognition information in complex contexts, fast speech speeds, and situations with speech occlusion. It reduces the recognition error rate while improving the continuity of natural interaction and also reduces the number of times users repeat their speech, thus effectively improving the usability of interaction in outdoor scenarios. Attached Figure Description
[0014] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:
[0015] Figure 1 This is an architectural diagram of a wide-angle projection multimodal AI quantum dot outdoor digital display device introduced in this invention;
[0016] Figure 2 For based on Figure 1A logical diagram illustrating the consistency determination between the text and the lip-shape recognition text;
[0017] Figure 3 For based on Figure 1 A schematic diagram illustrating the adjustment of interaction strategies based on user emotion recognition;
[0018] Figure 4 This is a schematic diagram of the appearance of a wide-angle projection multimodal AI quantum dot outdoor digital display device. Detailed Implementation
[0019] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.
[0020] In existing technologies, outdoor digital display devices generally rely on a single voice recognition or image recognition method for interaction, lacking a dynamic judgment mechanism for multi-source data. When voice recognition is affected by noise, distance, or multiple people, its confidence level often drops sharply. Traditional systems typically only make a "pass / reject" judgment based on the overall confidence level, leading to recognition interruption or requiring the user to re-enter, thus failing to achieve continuous interaction. Although some technologies have attempted to combine cameras for face detection or lip recognition, these only exist as a substitute module for voice recognition, without establishing an adaptive recognition logic based on confidence level hierarchy or achieving character-level fusion judgment. Therefore, it is difficult to maintain the stability of the interaction link during changes in recognition confidence.
[0021] To address the aforementioned issues, the technical solution proposed in this application utilizes a camera and multi-array microphones to collaboratively collect user information, establishing a dynamic recognition strategy based on speech confidence intervals. This enables a continuously adjustable multimodal fusion process between speech recognition, lip-reading recognition, and character segmentation, eliminating reliance on the reliability judgment of a single modality. When speech confidence decreases, the system automatically triggers a targeted remedial mechanism, using character-level replacement to finely fuse the speech and lip-reading recognition results. This preserves key information even in cases of speech distortion, rapid speech, or occlusion, ensuring the recoverability of the recognition process. Furthermore, by transforming the final recognition result into an interactive strategy, which is then executed and displayed on a quantum dot UHD large screen, this invention achieves a complete closed loop of "data acquisition—multimodal recognition—strategy generation—wide-angle display." This makes the interaction process not only recognizable but also executable, not only perceptible but also projectable, thus overcoming the problems of disconnect between recognition and display, reliance on a single modality for reliability, and unsustainable interaction in existing technologies.
[0022] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] like Figure 1 The image shows a wide-angle projection multimodal AI quantum dot outdoor digital display device, which mainly includes a camera, a multi-array microphone, a processor, and a UHD large screen.
[0024] The camera is used to capture images of the user's face.
[0025] After the device is powered on, the processor initializes and configures the camera, including setting exposure parameters, white balance mode, frame rate, and resolution, so that the camera can stably output image data under the current outdoor lighting conditions. After the camera completes parameter initialization, it continuously captures image frames of the area in front at a preset frame rate and sends each frame to the processor's image input buffer.
[0026] Multi-array microphones are used to capture the user's audio signals.
[0027] Before the multi-array microphones begin operation, the processor first reads the topology, number of channels, and array geometry parameters of the microphone array, and performs noise floor calibration and gain calibration to correct for outdoor background noise and sensor differences. After calibration, the multi-array microphones continuously acquire multi-channel audio signals from the environment using synchronous sampling, and send the raw audio data of each channel to the processor in real time.
[0028] The processor is specifically:
[0029] The system uses a camera to detect users. If a user is detected approaching, the system determines whether the user intends to interact based on the distance between the user and the screen or the duration of the user's gaze at the UHD screen. If so, the system drives a multi-array microphone to collect the user's audio signal.
[0030] The processor periodically receives image frames of the current scene from the camera and performs face detection and tracking in each frame. For each detected face target, the processor assigns a unique target number i and maintains the target's position information, size information, and historical trajectory information.
[0031] For the i-th user in a given frame, the processor determines the height of the bounding box of their face in the image. With preset reference height (For example, estimating the distance between the user and the UHD screen based on the height of a human face at a standard distance) Specifically, the following monocular estimation relationship can be used: ;
[0032] in, This is the distance scaling factor obtained through pre-calibration.
[0033] The processor will estimate the distance. Interaction distance threshold When comparing, At that time, it is determined that the user has approached the device.
[0034] Meanwhile, the processor uses facial landmark and head pose estimation algorithms to obtain the angle between the user's gaze direction and the screen normal. In consecutive image frames, when Less than the preset gaze threshold When a user is considered to be in a gaze state, the processor accumulates the time each user spends gazing at the UHD screen. For example, in each frame interval If the inside is satisfied ≤ Then according to Accumulate the distance. When the distance of at least one user is detected... Less than the interaction distance threshold or its gaze duration Greater than the preset time threshold At that time, the processor determines that the user has the intention to interact.
[0035] Among them, the interaction distance threshold The value range is 0.8~2.5m, preferably 1.2~2.0m. Gaze threshold The value range is 0° to 25°, and 20° is preferred in this embodiment. Time threshold. The value range is 0.5~2.5s, and 2s is preferred in this embodiment.
[0036] After determining that there is at least one user with interactive intent, the processor drives the multi-array microphone to focus on collecting and caching ambient audio from the current moment, providing audio input for subsequent speech recognition.
[0037] In a real-world scenario, when a user moves to within two meters of the screen, the camera detects an increase in the size of the user's face, and the processor determines that the user's distance is approaching a threshold. Simultaneously, the user looks up at the screen, aligning their head with the screen's orientation. At this point, the system accumulates the time the user has been looking at the screen. If the user continues to look at the screen for more than one or two seconds, the system considers that the user is preparing to interact. At this time, the system automatically activates the enhanced acquisition mode of the multi-array microphones, preparing to receive the user's upcoming voice command.
[0038] The multi-array microphones are only activated to enter acquisition mode when the system determines that the user is inclined to interact, fundamentally avoiding the problem of starting the speech recognition process when the user is not ready. This approach not only effectively reduces the continuous working time of the microphone array and minimizes invalid audio acquisition and background noise processing, but also improves the triggering accuracy of voice activity detection. This allows the system to enter the speech recognition process only when truly needed, thereby reducing energy consumption, improving overall processing efficiency, and significantly enhancing interaction stability and accuracy in complex outdoor environments.
[0039] When a user is detected approaching, if multiple users are approaching, the priority is calculated based on each user's position, line of sight, vocalization time, and lip movement frequency, and the user with the highest priority is selected as the interaction target.
[0040] The processor can determine the horizontal angle between the user's position and the center of the screen. Convert to position factor It is expressed as a function that decreases as the deviation angle increases: ; This is an angular scale parameter used to control the decay rate. In practical applications, distance can also be used. The user's location is used and normalized into a location factor. .
[0041] The angle of gaze can be determined by the duration of fixation. And normalize it into a viewing factor. .
[0042] Regarding the timing of speech, the processor records the timestamp of the user's first speech when it detects a speech segment. And based on the current moment Calculate the relative time difference and construct a time factor that elevates the earlier the sound is produced: ; is the time decay constant.
[0043] Regarding lip movement frequency, the processor estimates the lip movement frequency based on the deformation period (lip opening and closing) of the lip ROI in consecutive frames and normalizes it to obtain the lip movement factor. .
[0044] Based on this, the processor calculates priority for each user:
[0045] ;
[0046] in, For example, the weighting coefficient can be set to 0.25, 0.35, 0.25, or 0.15.
[0047] Processor selection The largest user is taken as the current interaction object, and the user's facial and voice data are mapped to the subsequent speech recognition and lip recognition processes.
[0048] Specifically, if multiple people approach the screen at the same time, such as two or three people standing in front of the screen simultaneously, the processor analyzes their position, gaze direction, whether they make vocalizations, and whether their lips are moving to determine who is most likely to actively interact with the device. This allows the system to select the most suitable person for voice capture in a multi-person environment, avoiding situations where the voices of other passersby are mistaken for commands.
[0049] If multiple users have the same priority, the processor recalculates the priority based on one or more of the secondary features, such as the order of speech time, peak volume of speech, and continuity of lip movements, and only uses the user with the highest priority as the interaction target; if they are still the same, the active prompting mechanism is triggered.
[0050] If there are multiple users with priority If the features are identical or extremely similar within a preset tolerance range, the processor initiates the secondary feature determination process. The processor calculates a secondary score for each user. :
[0051] ;
[0052] It is mapped from the order in which users speak (for example, the first user to speak is assigned a value of 1, and the value decreases sequentially for each subsequent user). The normalized peak volume factor (used to determine who speaks more clearly when multiple users speak at the same time); The lip motion continuity factor (range 0~1) is calculated based on the inter-frame continuity of lip region motion. , , The weighting coefficient can be set to 0.4, 0.35, or 0.25 in this embodiment.
[0053] The requirement emphasizes that if the duration during which the interacting object stops speaking exceeds the predefined waiting time, the processor recalculates the priority and determines the interacting object. The waiting time is typically set to 1-4 seconds, and 2.5 seconds is preferred in this embodiment.
[0054] This application reduces the probability of misselection and improves the accuracy of voice pickup in complex environments with multiple people nearby, thus enhancing the processor's target localization capabilities. Simultaneously, it reduces the entry of irrelevant voice into the recognition process, thereby lowering the processor load and improving recognition efficiency and interaction success rate. Overall, this solution achieves the technical effect of "automatically finding the right person among multiple people" in outdoor interaction scenarios, enabling the device to maintain a stable, efficient, and clear interactive experience even in high-traffic environments.
[0055] Having identified the current interaction target through the aforementioned priority calculation, the processor continuously monitors the audio signal corresponding to that target and tracks its voice activity status in real time. The processor maintains a "stop-speaking duration" parameter for the current interaction target. When the voice activity detection result is consistently silent, the processor accumulates this stop-speaking duration according to the system clock's time increment. If the stop-speaking duration exceeds the calibrated waiting time, the current interaction may have ended or the interaction intention may have been withdrawn. At this point, the processor triggers an interaction target reselection process, which involves re-invoking the aforementioned priority calculation steps based on features such as location, gaze, phonation time, and lip movement frequency. The processor recalculates the priority of all users in the current frame who meet the interaction conditions and selects the user with the highest priority as the new interaction target. If the original interaction target still obtains the highest priority after recalculation, the current interaction target remains unchanged; if it is a different user, the processor updates the interaction target identifier and binds the subsequent voice activity detection and speech recognition processes to the new interaction target.
[0056] Speech activity detection is performed on the audio signal of the selected interactive object. If valid speech activity is detected, speech recognition is performed on the audio signal to obtain the initial text and speech confidence score.
[0057] Based on the aforementioned priority calculation results, the processor determines the current interactive object and extracts the audio signal pointing towards that interactive object as the target audio signal through beamforming of a multi-array microphone. The processor performs frame-by-frame processing on the target audio signal, extracting speech activity detection features such as short-time energy and spectral features within each frame. Taking the k-th frame as an example, the short-time energy can be expressed as:
[0058] ;
[0059] in, Represented as the audio sample point within the k-th frame, Let n be the number of sampling points per frame, and n represent the index of the sampling point in the k-th frame. The processor will... Compared with noise background energy threshold When comparing, If other frequency domain characteristic conditions are met at the same time, the k-th frame is marked as a frame with "voice activity".
[0060] The processor aggregates the speech activity markers of consecutive frames. When several consecutive frames are marked as "speech activity", the start time of the speech segment is recorded; when several consecutive frames are marked as "no speech activity", the end time of the speech segment is recorded. The audio data between the start and end times is extracted into a complete speech segment to be recognized. If no "voice activity" frame is detected within the preset time window, it is assumed that there is no valid voice input, and the processor will not proceed to the subsequent voice recognition steps.
[0061] After detecting valid voice activity and capturing voice segments The processor then feeds the speech segment into a speech recognition model for decoding. In one embodiment, the speech recognition model can be an end-to-end model that combines a deep neural network-based acoustic model with a language model (e.g., an encoder-decoder structure or Wav2Vec), whose output is a sequence of characters or a sequence of sub-word units.
[0062] The processor first extracts features from the speech segment to obtain an acoustic feature vector sequence, which is then input into the encoder to calculate the posterior probability of each character (or word unit) c at each time t. Finally, the decoder obtains the initial text. The overall confidence index can be obtained by taking the logarithmic average of the posterior probabilities of each character along its decoding path, and then the normalized result can be used as the speech confidence index. .
[0063] In practical applications, the processor's computing power can be used to select the appropriate segmentation method, which can be implemented by word units, and then used for subsequent speech confidence assessment. If the character is within the threshold range, then character segmentation is performed; alternatively, character segmentation can be performed from the beginning, and the previously segmented characters can be used directly in subsequent steps.
[0064] Based on speech segments The silence ratio and signal-to-noise ratio are calculated, and the speech confidence is then corrected. The corrected speech confidence is then compared with a preset threshold range.
[0065] The processor will process the audio segments Frame segmentation is performed to obtain K frames of audio signal. Each frame contains N sampling points, where n is the index of the sampling point within the frame. For each frame of audio signal... Calculate short-time energy:
[0066] ;
[0067] like If the number of audio frames is 1, it is marked as a speech frame; otherwise, it is marked as a silence frame. Let the number of speech frames be 1. The number of silent frames is Then the total number of frames satisfies .
[0068] The formula for calculating the silent ratio is: .
[0069] in, ∈[0,1], the larger the value, the higher the proportion of silent segments in the speech segment and the shorter the effective duration of continuous speech.
[0070] The processor uses the energy of silent frames as a noise energy estimate to calculate the average noise power. The processor uses the energy of the speech frames as a signal power estimate to calculate the average signal power. Based on this, the average signal-to-noise ratio of the speech segment is calculated:
[0071] ;
[0072] processor Perform normalization to map it to the [0,1] interval, and obtain .
[0073] The corrected formula is as follows:
[0074] ;
[0075] in, Indicates the corrected speech confidence level; This represents the average signal-to-noise ratio of the audio signal. Indicates speech confidence; Indicates the quietness ratio; , , The values represent weights, and in this embodiment, 0.3, 0.5, and 0.2 are preferred.
[0076] Through the above steps, not only is the original confidence level given within the speech recognition model utilized... It also takes into account the signal-to-noise ratio of the external environment and the silence ratio of speech segments, and performs environmental adaptive correction on the speech confidence, so that the subsequent multimodal decision-making is more in line with the actual outdoor use scenario.
[0077] Corrected speech confidence With the preset threshold range Compare the results and make the following decision:
[0078] (1) If Then the initial text As the identification result .
[0079] (2) If Then, the initial text and speech confidence are divided into characters, and the characters of the initial text that are lower than the first threshold are replaced with the characters of the corresponding lip-shape recognition text, which are used as the recognition result; the lip-shape recognition text is the text whose confidence score is higher than the second threshold after lip-shape recognition of the lip region, and the lip region is obtained by ROI segmentation of the face image.
[0080] During the speech recognition stage, the processor can simultaneously obtain the speech confidence score for each character, thus representing the initial text as a sequence labeled with confidence scores. ; Representing speech characters The corresponding voice confidence level, This represents the number of characters in the language recognition text sequence. Simultaneously, the processor extracts the face region of the currently interacting object from the camera image and obtains the lip region through ROI segmentation. The lip image sequence is then input into the lip shape recognition model to obtain the lip shape recognition text sequence. and its corresponding character-level lip shape confidence. .
[0081] First threshold The value range is 0.40 to 0.75, with a typical value of 0.60, used to determine whether a speech character needs to be replaced with a lip-sync character; the second threshold The value can range from 0.50 to 0.85, with a typical value of 0.70, and is used to filter lip-shaped characters with high confidence.
[0082] The processor establishes an index mapping function between the speech character sequence and the lip-shape character sequence based on time alignment or position alignment: .in Representing speech characters Corresponding lip shape character , This indicates that the phonetic character has no valid correspondence in the lip-shape sequence. This indicates the number of characters in the lip-sync text sequence, i.e., the length of the character sequence output by the lip-sync model after CTC decoding.
[0083] If both conditions are met and If the lip shape is incorrect, replace it with the corresponding lip-shaped character; otherwise, retain the original lip-shaped character. For voice confidence, Confidence level for lip shape.
[0084] After completing the traversal of all characters, the processor uses the fused recognition result as the recognition result. .
[0085] The lip recognition model can adopt an architecture consisting of an input layer, a visual front-end, a temporal modeling layer, and an output layer.
[0086] The input layer is used to obtain the lip region from the camera image through face detection and ROI segmentation. It extracts T consecutive frames of lip images in chronological order, scaling each frame to a uniform size, and represents this as a four-dimensional tensor.
[0087] . For the number of frames; , Indicates the height and width of the image; The number of channels (1 or 3).
[0088] The visual front-end employs a 3D convolutional network (3D-CNN), comprising several 3D convolutional blocks. Each convolutional block may include a 3D convolutional layer (kernel size 3×3×3); batch normalization; ReLU activation function; and optional 3D max pooling (downsampling in both space and time). After L layers of 3D convolutional blocks, an intermediate feature tensor is obtained: Subsequently, through spatial dimensions ( , Perform global average pooling or flattening + linear transformation on the data to transform it into a feature sequence arranged in time:
[0089] .
[0090] The temporal modeling layer uses a bidirectional long short-term memory network (Bi-LSTM) for temporal modeling.
[0091] ;
[0092] , This indicates the hidden state of frames t and t-1.
[0093] The output layer uses a fully connected layer + softmax to output the probability distribution for each character:
[0094] ;
[0095] , These are learnable parameters.
[0096] During training, the Connectionist Tempora Classification (CTC) loss can be used to implicitly solve the time alignment problem; during inference, the lip-reading text is obtained through CTC decoding (optimal path or Beam Search). The corresponding character confidence score can be the probability of each output character at its alignment time step, or the average of the probabilities of that character over a time window.
[0097] (3) If If the confidence level of the lip-shape recognition text is higher than the third threshold, then the lip-shape recognition text is taken as the recognition result; otherwise, the recognition failure is taken as the recognition result.
[0098] Among them, the third threshold A value of 0.55 to 0.80 is acceptable, with a typical value of 0.65, and is used to determine whether the overall lip shape recognition result is reliable.
[0099] Based on the output of the lip-reading model, the processor calculates the overall confidence level of the lip-reading text corresponding to the speech segment. For example, the average of character-level lip shape confidence scores can be taken. Then the lip-shape recognition text As the identification result Otherwise, if the processor determines that neither the current speech nor the lip-reading modality can provide a reliable recognition result, it outputs a recognition failure flag, takes "recognition failure" as the recognition result of this round of speech segment, and notifies the subsequent interaction strategy module to adopt a prompting strategy, such as guiding the user to speak again or move closer to the device.
[0100] The recognition results are used to generate corresponding interaction strategies. Specifically, the processor processes the recognition results... The processor performs word segmentation, stop word filtering, and key information extraction. It then performs pattern matching on the sequence using intent trigger words, command words, and query words (such as "play," "check weather," "zoom in," "switch pages," etc.). Finally, it uses a lightweight semantic classification model (such as TextCNN, BiLSTM-Classifier, intent classifier, or rule engine) to perform intent recognition on the text and output the user intent category. ; This represents the probability of the semantic classification model for the i-th type of intent. Common intents include, but are not limited to:
[0101] Information query category (weather query, navigation query, etc.);
[0102] Content playback (video playback, ad switching);
[0103] Page operation functions (zoom in, zoom out, toggle);
[0104] System interaction functions (volume adjustment, brightness adjustment).
[0105] If the intent requires parameters (such as "play Shanghai weather video"), the processor extracts the parameters from the text, for example: location parameter: Shanghai; content parameter: weather; operation parameter: play; forming a structured parameter set P. If the parameters cannot be completely parsed from the text, the completion strategy is triggered (such as prompting the user "please add location information").
[0106] The system maintains a policy mapping table or policy template library, as shown in the table below:
[0107] Intent type Strategy Template Playback Play specified content strategy Query class Information display strategy Page manipulation class UI Action Strategy Control Class System parameter adjustment strategy
[0108] According to intent Select the corresponding strategy template and fill the parsed parameter P into the selected strategy template. Generate specific interaction strategy objects: ;in, Populate functions for the template. For example, a user says, "Play Shanghai weather." The intended message is: Playback class. Parameters: Location = Shanghai, Content = Weather. Strategy: ="Play Shanghai weather video" If the strategy requires triggering screen animations, effects, and UI prompts, the processor will synchronously generate a set of interactive strategy instructions such as UI display instructions, content rendering instructions, animation start instructions, and page jump instructions according to the strategy type.
[0109] The processor sends the generated instruction set to the UHD large screen. The UHD large screen executes the interaction strategy according to the instruction set: content presentation, UI switching, triggering play / pause, starting graphic display, displaying guide cards or prompts, and finally completing the interaction loop with the user.
[0110] UHD large screens employ a quantum dot-enhanced display structure. By exciting a quantum dot film or a quantum dot self-emissive layer with blue light, the quantum dots output spectrally stable red and green light over a wide viewing angle, thus reducing brightness decay and color shift caused by changes in viewing angle. Due to the narrow bandwidth, high luminous efficiency, and low angle dependence of quantum dots, the large screen exhibits consistent brightness and color performance within a 0°–170° range, achieving wide-angle projection capabilities in outdoor scenarios and ensuring users receive clear and uniform display content from various locations.
[0111] In practical applications, after obtaining valid recognition results, the processor will identify the text input interaction strategy module and generate an interaction strategy corresponding to the user's command through intent recognition and content parsing. For example, when the user says "Play the weather forecast," the screen will automatically switch to the weather information display interface; if the user says "Zoom in on the image," the interface zoom strategy will be triggered. Through the above process, the device can maintain stable and accurate interaction capabilities in complex environments such as outdoor noise and multi-person scenes, enabling users to interact with the system using natural voice.
[0112] The threshold range is calculated statistically from multiple rounds of interaction data and user feedback collected during the testing phase. The latest interaction data is statistically analyzed at preset time intervals, and the current threshold range is updated via a sliding window. The interaction data includes audio signals. Quietness ratio Signal-to-noise ratio Voice confidence And feedback information on whether the user is satisfied with the output.
[0113] The processor maintains a sliding window of size W (ranging from 50 to 100) for updating the threshold interval, which stores the most recent W interaction data. When new interaction data is generated, if the window is not full, it is appended to the window. If the window is full, the oldest data is popped out and the newest data is added, thus achieving real-time scrolling statistics.
[0114] Calculate the new threshold interval based on the speech confidence distribution within the window:
[0115] ;
[0116] ;
[0117] in, This represents the lower quantile (25%). This represents the high quantile (75%). This indicates the corrected speech confidence level. This means taking the p-th element after sorting the input data set in ascending order. Statistical operations for quantiles.
[0118] If user feedback indicates "dissatisfaction" or "error," the weight is adjusted to reduce the statistical validity of that sample. The threshold interval is recalculated every preset time period to ensure dynamic convergence of the threshold interval in response to changes in environmental noise, interactive behavior, etc. The preset time period ranges from 10 minutes to 240 minutes, with the specific value determined based on pedestrian traffic.
[0119] In practical applications, to avoid unnecessary computational waste, when comparing the speech confidence score with a preset threshold range, if the speech confidence score is within the threshold range, lip region recognition is only performed on the face images corresponding to characters in the initial text that are below the first threshold. For each character to be replaced, the time range of the corresponding speech segment is located, and the time range is mapped to an image frame sequence. The corresponding face image segment is obtained through the video frame timestamps. The processor performs ROI segmentation on each group of face image segments, extracts the lip region, and performs lip shape recognition. If the speech confidence score is below the threshold range, lip region recognition is performed on the entire face image corresponding to the initial text.
[0120] The processor performs lip region recognition only on local image segments corresponding to low-confidence characters, correcting a small number of unreliable characters through "local replacement." This allows the system to significantly reduce video processing load and improve processing speed while maintaining overall recognition accuracy. When the speech confidence level falls below a threshold, the processor automatically switches to performing lip recognition on the entire face image, ensuring reliable text results even when the speech is severely affected by noise. This achieves the effect of "dynamically adjusting processing intensity based on reliability," enabling the system to respond quickly and reduce unnecessary computational resource consumption in lightly noisy scenarios, while ensuring recognition reliability in heavily noisy scenarios. Overall, it achieves an optimal balance between computational load, real-time performance, and recognition accuracy, significantly improving the practicality and stability of multimodal interaction in outdoor scenarios.
[0121] The main scheme of this application has been introduced above. Below, we will describe the consistency determination of the lip-shape recognition text before replacing characters in the initial text below the first threshold with characters in the corresponding lip-shape recognition text. The specific steps are as follows:
[0122] The initial text with a speech confidence level higher than a threshold range within the initial preset time window for speech recognition of audio signals is used as the baseline text, and lip region recognition is simultaneously performed on the face images within the initial preset time window to obtain lip shape recognition text.
[0123] Calculate the similarity between the lip-shape recognition text and the corresponding baseline text. If the similarity exceeds the threshold, the lip region recognition is deemed reliable, and the initial text is corrected by lip-shape recognition text assistance.
[0124] Specifically, during the speech recognition stage, the processor divides the audio signal into multiple time windows and labels the average speech confidence of each window. When the speech confidence within a certain window exceeds a threshold range, the initial text for speech recognition corresponding to that window is defined as the baseline text. The processor synchronously extracts corresponding video frame sequences from the camera image sequence based on the time range of the window. It then performs the following operations on this sequence: face detection, ROI segmentation to obtain the lip region, and lip shape recognition model inference to obtain the lip shape recognition text sequence within the window and the corresponding lip shape confidence score. Lip shape recognition text with a confidence score higher than a second threshold is retained, forming a new filter sequence. If no character meets the condition, the lip shape recognition in this round is deemed ineffective, and lip shape correction is skipped directly.
[0125] To assess the stability between the visual and auditory modalities, the processor calculates their character similarity, which can be achieved using Levenshtein Distance, Jaccard character similarity, subsequence matching rate (LCS), or CTC path consistency comparison. Taking Levenshtein Distance as an example, the calculation formula is as follows:
[0126] ;
[0127] in, This indicates character similarity; the closer the similarity is to 1, the higher the consistency between the two modalities. Indicates the filter sequence; Indicates the base text.
[0128] Set the consistency threshold to 0.7; if the character similarity... If the result is greater than 0.7, the current lip-shape recognition result is considered stable and consistent with the speech modality. In this case, the subsequent lip-shape auxiliary correction operation (character replacement) is allowed. Otherwise, the lip-shape recognition is considered unreliable, character replacement is not performed, and the original speech text structure is maintained.
[0129] In existing multimodal fusion methods, lip-sync recognition can be flawed due to lighting, angle, or occlusion. This solution introduces a consistency check before replacement, using benchmark text from a high-confidence speech window as a reference. This ensures that lip-sync recognition only participates in correction when it exhibits stable consistency with the speech modality, significantly reducing the false replacement rate and improving overall recognition accuracy. Correction is performed only when the lip shape and speech are highly consistent, and local replacement is based on character confidence, avoiding unnecessary lip-sync analysis of the entire video segment. This effectively reduces processor load and enables the system to achieve higher response efficiency in real-time environments.
[0130] The above describes the specific steps for consistency determination of lip-reading text. The following section details the processor's further steps, including user emotion recognition based on facial images and audio signals, and the corresponding adjustment of interaction strategies. The specific steps are as follows:
[0131] The audio signal is subjected to first emotion feature extraction, and the facial image is subjected to second emotion feature extraction. The first emotion feature is the feature that reflects the user's speech rate, pause rate and volume changes, and the second emotion feature is the feature that reflects the user's eyebrow tension, eye lingering and mouth opening ratio.
[0132] An emotion index is calculated by weighting the first and second emotion features. The display speed and interaction mode of the displayed content in the interaction strategy are then adjusted based on the emotion index.
[0133] The processor divides the audio signal into frames and extracts the following audio features that reflect the user's emotional fluctuations:
[0134] Speech rate: Calculated by measuring the number of phonemes or effective speech segments per unit time.
[0135] Pause rate: Calculated by the percentage of silent segments;
[0136] Volume change: obtained based on short-time energy change.
[0137] The first emotional characteristic is formed.
[0138] Perform keypoint detection on the face image to obtain the coordinates of keypoints in the eyebrows, eyes, and lips, and calculate the following visual emotion features:
[0139] eyebrow tension : ; This indicates the change in the height of the eyebrows. It represents the distance between the center points of the two eyes.
[0140] Eye stay This refers to fixation time. Stability.
[0141] Mouth opening ratio : ; This indicates the vertical coordinates of the key points on the upper lip. This indicates the perpendicular coordinates of the key points on the lower lip. This represents the baseline width between the left and right corners of the mouth, used to normalize the opening and closing distance, making the mouth opening and closing ratio comparable for different users and at different distances.
[0142] This forms a second emotional characteristic.
[0143] The emotion index (ranging from 0 to 1) is calculated by weighting the normalized first and second emotion features. The emotion index can be divided into several ranges, for example: less than 0.3 indicates calmness, between 0.3 and 0.6 indicates focus, and greater than 0.6 indicates tension or excitement.
[0144] If the emotional index is high (such as tension, anger, excitement): slow down the screen display speed, provide more obvious visual guidance (such as highlighted arrows), switch to simplified mode or repeat confirmation mode, increase the button area, and reduce the amount of information.
[0145] If the mood index is low (e.g., calm, relaxed): improve interface responsiveness, display richer content, and enable enhanced interaction modes (e.g., animations, multi-turn Q&A).
[0146] If the user shows signs of fatigue or confusion (e.g., unusual eye movement, slow speech): activate accessibility prompts, reduce interaction steps, and provide clarifying feedback (e.g., "Do you need help?").
[0147] Finally, the processor sends the action list generated by the strategy to the UHD screen for execution.
[0148] This solution integrates emotional features from both audio and visual modalities, enabling the device to capture the user's state and automatically adjust its response based on the user's emotions. This makes the interaction more aligned with the user's psychological state, presenting a more user-friendly and human-centered interface, and making the interaction between the user and the device more natural, smooth, and enjoyable.
[0149] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. A wide-angle projection multimodal AI quantum dot outdoor digital display device, characterized in that, include: A camera is used to capture images of the user's face. Multi-array microphones are used to collect the user's audio signals; Processor, comprising: Speech activity detection is performed on the audio signal. If valid speech activity is detected, speech recognition is performed on the audio signal to obtain the initial text and speech confidence score. The speech confidence level is compared with a preset threshold range, and the following decision is made: (1) If the speech confidence is higher than the threshold range, the initial text is used as the recognition result; (2) If the speech confidence is within the threshold range, the initial text and the speech confidence are divided into characters. The characters of the initial text that are lower than the first threshold are replaced with the characters of the corresponding lip-shape recognition text. The replaced text is used as the recognition result. The lip-shape recognition text is the text whose confidence is higher than the second threshold after lip-shape recognition of the lip region. The lip region is obtained by segmenting the face image by ROI. (3) If the confidence level of the speech is lower than the threshold range, then determine whether the confidence level of the lip-reading text is higher than the third threshold. If yes, then the lip-reading text is taken as the recognition result; otherwise, the recognition failure is taken as the recognition result. Generate corresponding interaction strategies from the recognition results; Before replacing characters in the initial text that are below the first threshold with characters in the corresponding lip-sync text, the processor first performs a consistency check on the lip-sync text. The specific steps are as follows: The initial text with a speech confidence score higher than a threshold within the initial preset time window for speech recognition of the audio signal is used as the baseline text. Simultaneously, lip region recognition is performed on the face image within the initial preset time window to obtain lip shape recognition text. The speech confidence score of the initial preset time window is obtained by the processor dividing the audio signal into multiple time windows and marking the average speech confidence score of each window during the speech recognition stage. Calculate the similarity between the lip-shape recognition text and the corresponding reference text. If the similarity exceeds the threshold, the lip region recognition is deemed reliable, and the initial text is corrected by lip-shape recognition text assistance. UHD large screens are used to execute interaction strategies and display corresponding content.
2. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 1, characterized in that, Before comparing the speech confidence score with a preset threshold range, the processor also calculates the silence ratio and signal-to-noise ratio based on the audio signal, thereby correcting the speech confidence score, and then compares the corrected speech confidence score with the preset threshold range. The correction formula is as follows: ; in, Indicates the corrected speech confidence level; This represents the average signal-to-noise ratio of the audio signal. Indicates speech confidence; Indicates the quietness ratio; , , Indicates the weight.
3. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 2, characterized in that, The threshold range is obtained by statistical calculation of multiple rounds of interaction data and user feedback collected during the testing phase. The latest interaction data is statistically analyzed every preset time period, and the current threshold range is updated through a sliding window. The interactive data includes audio signals, silence ratio, signal-to-noise ratio, voice confidence, and feedback information on whether the user is satisfied with the output.
4. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 1, characterized in that, Before detecting voice activity in the audio signal, the processor detects the user through the camera. If the user is detected approaching, the processor determines whether the user has an intention to interact based on the distance between the user and the camera or the time the user looks at the UHD screen. If so, the processor drives the multi-array microphones to collect the user's audio signal.
5. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 4, characterized in that, When the processor detects a user approaching, if multiple users are approaching, it calculates the priority based on each user's position, gaze angle, vocalization time, and lip movement frequency, and selects the user with the highest priority as the interaction target.
6. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 5, characterized in that, If multiple users have the same priority, the processor recalculates the priority based on one or more of the secondary features, such as the order of speech time, peak volume of speech, and continuity of lip movements, and only uses the user with the highest priority as the interaction target; if they are still the same, the active prompting mechanism is triggered.
7. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 5, characterized in that, If the duration during which the processor determines an interactive object stops speaking exceeds the preset waiting time, the processor recalculates the priority and determines the interactive object.
8. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 1, characterized in that, When the processor compares the speech confidence score with a preset threshold range, if the speech confidence score is within the threshold range, it only performs lip region recognition on the face image corresponding to the characters of the initial text that are below the first threshold. If the speech confidence level is below the threshold range, then lip region recognition is performed on the entire face image corresponding to the initial text.
9. The wide-angle projection multimodal AI quantum dot outdoor digital display device according to claim 1, characterized in that, The processor also includes user emotion recognition based on facial images and audio signals, and corresponding adjustments to the interaction strategy. The specific steps are as follows: The audio signal is subjected to first emotion feature extraction, and the facial image is subjected to second emotion feature extraction. The first emotion feature is the feature that reflects the user's speech rate, pause rate and volume changes, and the second emotion feature is the feature that reflects the user's eyebrow tension, eye lingering and mouth opening ratio. An emotion index is calculated by weighting the first and second emotion features. The display speed and interaction mode of the displayed content in the interaction strategy are then adjusted based on the emotion index.
Citation Information
Patent Citations
Method for adjusting confidence coefficient threshold of voice recognition and electronic device
CN103578468A
Voice interaction method, device, equipment and system
CN115206306A
Keyword recognition method and device, electronic equipment and storage medium
CN116013260A
Video display method and device, equipment and storage medium
CN116055792A
Speech recognition method, device and equipment and computer readable storage medium
CN116805490A