LED multi-dimensional information issuing method based on voice intention recognition
By constructing a decoding word graph and embedding clarifying rhetorical questions, the user's clarification instructions are captured, the problem of misjudgment of homophones in speech recognition is solved, and synchronous clarification feedback between speech and LED screen is achieved, which improves semantic consistency and response stability.
Patent Information
- Application Number
- CN202510796658.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech recognition technology is prone to misjudgment or path drift when faced with homophonic or polysemous phrases, resulting in information content mismatch or release confusion, and a lack of effective semantic clarification feedback channels.
By constructing a decoding word graph, locating the semantic slots of LED multi-dimensional information, extracting high-confidence ambiguous candidate sets, embedding clarifying rhetorical questions for voice broadcast and synchronous display on the LED screen, capturing user clarification instructions, establishing a unique recognition path, and enhancing the bias weight of context-related words during the decoding process.
It improves the efficiency of clarifying ambiguity in speech recognition, enhances the context-awareness of recognition logic, achieves semantic consistency and response stability, and improves the continuity of understanding of continuous semantic instructions.
Smart Images

Figure CN120766682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a method for publishing LED multi-dimensional information based on speech intention recognition. Background Art
[0002] The field of speech recognition technology refers to the complete process of real-time collection, feature extraction, acoustic modeling, language modeling and semantic understanding of human speech signals through computers or embedded systems, enabling machines to recognize and understand the content and intention of human language.
[0003] Existing speech recognition technology primarily relies on acoustic and language models to generate a unified decoding path and perform maximum likelihood recognition for speech streams. This makes it difficult to perform fine-grained semantic discrimination in scenarios where multiple candidate paths have similar scores, and it also lacks user-directed feedback channels for semantic clarification. When user voice input contains homophones or polysemous phrases, misjudgment or path drift is prone to occur, resulting in information mismatch and confusing delivery. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and propose a method for LED multi-dimensional information release based on voice intention recognition.
[0005] In order to achieve the above objectives, the present invention adopts the following technical solution: a method for publishing LED multi-dimensional information based on voice intention recognition, comprising the following steps:
[0006] Decode the input speech signal to generate a decoded word graph, locate the preset LED multi-dimensional information semantic slots in the decoded word graph, extract similar competitive paths, and calculate a high-confidence ambiguous candidate set;
[0007] Based on the high-confidence ambiguous candidate set, two competing candidate words are extracted and embedded into a preset question template to generate a clarifying rhetorical question. Based on the clarifying rhetorical question, the voice broadcast is driven to synchronously display the question on the LED screen, and the user's clarification voice is captured to obtain the user's clarification instruction;
[0008] Based on the user clarification instruction, searching for a path matching the user clarification instruction in the competitive paths of the decoding word graph, establishing a unique identification path, parsing node information based on the unique identification path, extracting corresponding LED display color parameters and LED display font parameters, and establishing an LED display state vector;
[0009] When a new voice command is received, the current LED display state vector is called, and the corresponding command vocabulary is retrieved according to the LED display color parameters and the LED display font parameters, and a context-related vocabulary set is established. Based on the context-related vocabulary set, in the process of decoding the new voice command, bias weights are applied to the vocabulary belonging to the context-related vocabulary set in the language model to increase the prior probability and generate a context-biased instruction.
[0010] Preferably, the steps of obtaining the high-confidence ambiguous candidate set are:
[0011] Extract the Mel-frequency cepstral coefficients from the input speech signal, generate a decoding word graph for the corresponding time frame, traverse all path groups in the word graph that include the semantic slots of LED information, filter out path pairs with overlapping frame positions and similarity, and generate a set of semantically competitive path pairs;
[0012] Calculating a normalized logarithmic ambiguity based on the set of semantic competition path pairs;
[0013] According to the normalized logarithmic ambiguity, all path pairs whose normalized logarithmic ambiguity is greater than a corpus-level dynamic adaptive threshold are extracted, and the corresponding candidate words are combined and output to generate a high-confidence ambiguous candidate set.
[0014] Preferably, the steps for obtaining the clarification rhetorical question sentence are:
[0015] According to the high-confidence ambiguous candidate set, each group of candidate word combinations in the high-confidence ambiguous candidate set is retrieved one by one, the normalized logarithmic ambiguity within each group of candidate word combinations is sorted, and the two competing candidate word combinations that rank first in the normalized logarithmic ambiguity are selected as target combinations to generate a target candidate word combination;
[0016] Based on the target candidate word combination, the text contents of two competing candidate words in the target candidate word combination are sequentially extracted, a pre-set clarification question template is called, the text contents of the competing candidate words are respectively filled in the preset placeholder positions in the question template, and the grammar check of other text contents in the clarification question template is gradually completed to generate a clarification question text;
[0017] Based on the clarification question text, a complete clarification question voice file is synthesized, and the LED screen display content is formatted so that the LED screen text display content corresponds word for word to the voice file content, forming a clarification rhetorical question sentence pattern.
[0018] Preferably, the steps of obtaining the user's clarification instruction are:
[0019] According to the clarification interrogative sentence, the voice content of the clarification interrogative sentence is played word by word, and the LED screen display control interface is called to display the text content of the clarification interrogative sentence word by word in synchronization, so that the LED screen text display is consistent with the voice playing content in the time axis, forming a synchronous broadcast state.
[0020] Based on the synchronous broadcast state, the first valid silence segment after the end of playing the clarification interrogative sentence is taken as a starting point of capture, and the next silence segment is taken as an ending point of capture, the real-time clarification voice of the user in the silence gap is intercepted, the real-time clarification voice is segmented in a time domain window, noise is filtered, and feature extraction is performed, forming user clarification voice data to be recognized.
[0021] Based on the user clarification voice data to be recognized, the user clarification voice data is decoded frame by frame, the recognition text whose confidence reaches a preset recognition confidence standard is selected as an effective result as a recognition result, and a user clarification instruction is formed.
[0022] Preferably, the obtaining step of the unique recognition path is:
[0023] According to the user clarification instruction, the character sequence of the user clarification instruction is parsed word by word, and the core keyword text in the user clarification instruction is extracted, the core keyword text is labeled with semantic features piece by piece, and a semantic feature sequence of the user clarification instruction is generated.
[0024] Based on the semantic feature sequence of the user clarification instruction, the semantic feature sequence of the competing path in the decoding word graph is called, the semantic feature sequence of each competing path is compared character by character, the semantic feature sequence of the user clarification instruction is matched with the semantic feature sequence of each competing path in turn, and a semantic feature matching degree sequence is generated.
[0025] Based on the semantic feature matching degree sequence, the semantic feature matching degrees of each competing path are sorted, the competing path with the highest semantic feature matching degree is selected as the path completely matched with the user clarification instruction, and a unique recognition path is established.
[0026] Preferably, the obtaining step of the LED display state vector is:
[0027] Based on the unique recognition path, each node contained in the unique recognition path is accessed one by one, the preset text label and the corresponding node feature in the node are parsed node by node, the node feature information associated with the LED screen display content is extracted, and a node feature information set is generated.
[0028] According to the node feature information set, the color configuration codes and font configuration codes in the node feature information set are retrieved one by one, matched through the color coding table and the font coding table, and the color configuration codes and font configuration codes are converted into corresponding color RGB values and font size, style and thickness parameters respectively to generate an LED display parameter set;
[0029] Based on the LED display parameter set, the color RGB values, font size, font style and font thickness parameters in the set are called in sequence, and the color and font parameter combinations obtained by the calls are packaged into a vectorized data format to generate an LED display state vector.
[0030] Preferably, the steps of acquiring the context-related vocabulary set are:
[0031] After receiving a new voice command, the system performs frame recognition on the voice command, extracts all voice words and constructs a word vector for each voice word. It also extracts the color RGB values and font size of all text elements in the current LED display state vector to generate a voice word vector set and a screen text vector set.
[0032] Calculating the visual-semantic association between each speech word and all screen text elements based on the speech word vector set and the screen text vector set;
[0033] Based on the visual-semantic relevance of the speech word-units, all speech word-units whose visual-semantic relevance is higher than the average relevance of all word-units are screened, and the corresponding original texts are formed into a set to generate a context-related vocabulary set.
[0034] Preferably, the steps of obtaining the context bias instruction are:
[0035] Based on the context-related vocabulary set, extract the text content of each word in the context-related vocabulary set one by one, establish a context-related vocabulary text index list, and search one by one in the vocabulary library of the language model to determine whether the context-related vocabulary text exists in the language model vocabulary library, and generate a valid context-related vocabulary list;
[0036] Extracting, one by one, the original prior probability value corresponding to each valid context-related word in the language model according to the valid context-related word list, increasing the probability value according to a fixed ratio for each original prior probability value of the valid context-related word, updating the corresponding word probability parameters in the language model one by one, and generating a probability-enhanced context-related word prior probability set;
[0037] Based on the probability-enhanced context-associated vocabulary prior probability set, the audio feature sequence of the newly input voice instruction is decoded frame by frame. For each decoding step of the newly input voice instruction, the enhanced probability parameter in the probability-enhanced context-associated vocabulary prior probability set is called, the decoding score of the corresponding vocabulary is updated, and a context bias instruction is generated.
[0038] Compared with the prior art, the advantages and positive effects of the present invention are:
[0039] When processing speech signals, the present invention constructs a decoding word graph and locates semantic slots of multi-dimensional LED information. This allows for the extraction of similar candidate word combinations within ambiguous recognition paths. The posterior probability differences are further calculated to form ambiguous candidate sets. This introduces a clear basis for differentiation in scenarios with similar recognition accuracy but semantic conflict, alleviating recognition ambiguity caused by path divergence in the word graph. Clarification questions embedded with competing words are presented simultaneously via both speech and LED channels, making user feedback more targeted and improving the efficiency of ambiguous clarification. The user's clarifying speech is captured and a unique recognition path is established. Combining the color and font parameters of the nodes in the path, a semantically variable LED display state vector is constructed, achieving a semantic-to-visual linkage mapping. When new speech is input, the current display state vector parameters are used to reversely infer the command word candidate set for the current semantic context. This in turn increases the language model priority of related words during the decoding process, enabling continuous context-aware recognition logic. By sequentially processing steps such as explicit semantic ambiguity resolution, semantic linkage driven by visual parameters, and explicit intervention of contextual vocabulary, the system enhances the understanding continuity, semantic consistency, and response stability of continuous semantic instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the steps of the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0042] See also Figure 1 The present invention provides a technical solution, a method for publishing LED multi-dimensional information based on voice intention recognition, comprising the following steps:
[0043] Decode the input speech signal to generate a decoded word graph, locate the preset LED multi-dimensional information semantic slots in the decoded word graph, extract similar competitive paths, and calculate a high-confidence ambiguous candidate set;
[0044] Based on a high-confidence ambiguous candidate set, two competing candidate words are extracted and embedded into a preset question template to generate a clarifying rhetorical question. Based on the clarifying rhetorical question, the voice broadcast is driven and the LED screen displays the question synchronously. The user's clarification voice is captured to obtain the user's clarification instruction.
[0045] Based on the user's clarification instruction, a path matching the user's clarification instruction is retrieved from the competitive paths of the decoding word graph, and a unique identification path is established. Based on the unique identification path, the node information is parsed, the corresponding LED display color parameters and LED display font parameters are extracted, and an LED display state vector is established;
[0046] When a new voice command is received, the current LED display state vector is called, and the corresponding command vocabulary is retrieved according to the LED display color parameters and the LED display font parameters. A context-related vocabulary set is established. Based on the context-related vocabulary set, in the process of decoding the new voice command, bias weights are applied to the vocabulary belonging to the context-related vocabulary set in the language model to increase the prior probability and generate context-biased instructions.
[0047] The steps to obtain high-confidence ambiguous candidate sets are:
[0048] Extract the Mel-frequency cepstral coefficients from the input speech signal, generate a decoding word graph for the corresponding time frame, traverse all path groups in the word graph that include the semantic slots of LED information, filter out path pairs with overlapping frame positions and similarity, and generate a set of semantically competitive path pairs;
[0049] According to the set of semantic competition path pairs, the normalized logarithmic ambiguity is calculated. The calculation formula is:
[0050]
[0051] in, S ac1 is the original score of the acoustic model of path 1, S ac2 is the original score of the acoustic model of path 2, S LM1 is the original score of the language model of path 1, S LM2 is the original score of the language model of path 2, ∈ is a very small constant to prevent division by zero, R ac and R lm is the relative score difference ratio, A nla is the normalized logarithmic ambiguity;
[0052] According to the normalized logarithmic ambiguity, all path pairs whose normalized logarithmic ambiguity is greater than the corpus-level dynamic adaptive threshold are extracted, and the corresponding candidate words are combined and output to generate a high-confidence ambiguous candidate set.
[0053] Specifically, based on the received input speech signal, the original audio stream is first preprocessed, including pre-emphasis, framing and windowing. For example, the speech signal is collected at a sampling rate of 16kHz, the length of each frame is set to 25 milliseconds, the frame shift is set to 10 milliseconds, and the Hamming window is used for windowing. Then, a fast Fourier transform (FFT) is performed on each frame of the windowed speech signal to obtain its spectral characteristics. The spectral energy is then passed through a set of pre-designed Mel filter banks. This filter bank simulates the auditory characteristics of the human ear, and its center frequency is linearly distributed in the low frequency band. The high frequency band is logarithmically distributed, and the logarithmic energy of each filter output is calculated. Then, the Mel-frequency cepstral coefficient (MFCC) of each frame is obtained through discrete cosine transform (DCT). For example, 13-dimensional MFCC features including energy terms and their first-order and second-order differences are extracted to form a 39-dimensional feature vector. These acoustic feature vectors are used in combination with pre-trained acoustic models (such as acoustic models based on deep neural networks DNN or hidden Markov models HMM-GMM) and language models (such as N-gram language models or recurrent neural network language models R During the decoding process, a decoding word graph (Lattice) is generated, representing all possible word sequences and their corresponding acoustic and language scores. This Lattice is a weighted directed acyclic graph (DAG), where nodes represent word end times and edges represent recognized words and their acoustic and language model scores. The system then traverses all potential recognition paths in the decoding word graph, specifically searching for path groups that contain predefined LED information semantic slots (for example, slot types may include "color," "brightness," "display content," "font size," etc., which are clearly specified during system design). These path groups containing the target semantic slots are screened. The screening criteria are that the words or phrases corresponding to the semantic slots in the paths are roughly aligned in time frames, that is, their start and end timestamps differ within a very small tolerance (for example, no more than 50 milliseconds), and that these paths compete for the same type of semantic slot. For example, two paths identified as "display red" and "display blue" compete for the semantic slot "color." Through this screening, a set of semantically competing path pairs is ultimately generated.
[0054] formula: in, The benefit of the formula is that the normalized logarithmic ambiguity A nlaIt can effectively quantify the degree of ambiguity between two semantically competing paths. It takes into account the relative differences between the acoustic model scores and the language model scores, and amplifies the influence of smaller differences through the logarithmic function and the operation of adding one to the reciprocal, so that those path pairs with similar scores and thus more easily confused are given higher ambiguity values, while those path pairs with larger score differences are given lower ambiguity values; normalization processing (by dividing by the maximum absolute value) eliminates the influence of the original score scale on the ambiguity calculation, making the ambiguity calculated under different models or different conditions comparable; by multiplying the acoustic ambiguity and the language model ambiguity, the overall degree of ambiguity is comprehensively evaluated. Only when the two paths are relatively similar at the acoustic and linguistic levels will a higher A be obtained. nla values, thereby more accurately identifying high-confidence ambiguities that require user clarification.
[0055] S ac1 The steps to obtain the parameters are: ac1 This represents the acoustic model raw score for path one in a semantically competitive path pair. This score is calculated by the speech recognition system's acoustic model during decoding and reflects the degree of match between the acoustic features of the input speech signal and the pronunciation of the word sequence corresponding to path one. It is typically a log-likelihood value, with larger values (closer to 0) indicating a stronger match. This score is obtained by generating a decoded word graph during speech decoding, where each word in each path is assigned an acoustic score. The overall acoustic score for path one is the sum of the acoustic scores of all words it contains. For example, for the speech segment "set to red," path one might recognize "red." Its acoustic model calculates a score by matching this speech segment with the pronunciation model for "red." This score is obtained by decoding a test set of 10,000 annotated LED control speech clips using a DNN-HMM acoustic model trained using the Kaldi toolkit. For example, for path one (recognized as "red"), the acoustic model raw score is -85.6.
[0056] S ac2 The steps to obtain the parameters are: ac2 The raw score of the acoustic model of path 2 in the semantic competition path pair is obtained in the same way as S ac1 The same is true for competing path 2. For example, for the same speech segment “set to red”, path 2 incorrectly identifies it as “blue” and its acoustic model calculates a raw score of -88.2.
[0057] S LM1 The steps to obtain the parameters are: LM1The raw language model score for path one in a semantically competitive path pair is calculated by the speech recognition system's language model. It reflects the naturalness or probability of the word sequence corresponding to path one appearing in the target language (e.g., the instruction set for controlling an LED), and is a logarithmic probability value. This score is obtained by assigning a score to the word sequence of path one based on its internal statistical information (e.g., N-gram frequency or neural network prediction) during the decoding process. For example, the language model score for path one (identified as "red") in context (e.g., "Please set the color to red") is -2.5, calculated using a 3-gram language model trained on a corpus containing 50,000 LED control instruction texts.
[0058] S LM2 The steps to obtain the parameters are: LM2 represents the raw score of the language model of path 2 in the semantic competition path pair, which is obtained in the same way as S LM1 Exactly the same, but for competing path 2. For example, path 2 (recognizing “blue”) has a language model score of -3.1 in the same context.
[0059] The steps to obtain the ∈ parameter are as follows: ∈ is a very small constant set to prevent division by zero errors, and its value needs to be much smaller than the normal max(|S ac1 |,|S ac2 |) and max(|S LM1 |,|S LM2 |), but not so small as to affect the accuracy of floating-point calculations. It is set based on experience. For example, in common speech recognition systems, the absolute values of the acoustic score and language model score are greater than 1, so ∈ can be set to a very small positive number, such as 1×10 -9 .
[0060] Calculation process:
[0061] Given:
[0062] Path 1 (e.g., "red"): S ac1 =-85.6, S LM1 =-2.5;
[0063] Path 2 (e.g., "blue"): S ac2 =-88.2, S LM2 =-3.1;
[0064] Minimum constant: ∈=1×10 -9 ;
[0065] First, the relative score difference ratio R of the acoustic model score is calculated ac :
[0066]
[0067] Then, calculate the relative score difference ratio R of the language model score lm :
[0068]
[0069] Next, calculate the normalized logarithmic ambiguity A nla :
[0070]
[0071] Substitute R ac ≈0.029478 and R lm ≈0.193548:
[0072]
[0073]
[0074] A nla =(1+ln(1+33.9236))·(1+ln(1+5.1666));
[0075] A nla =(1+ln(34.9236))·(1+ln(6.1666));
[0076] Using the natural logarithms ln(34.9236) ≈ 3.5531 and ln(6.1666) ≈ 1.8192:
[0077] A nla =(1+3.5531)·(1+1.8192);
[0078] A nla =4.5531·2.8192;
[0079] A nla ≈12.835;
[0080] The results show that for the semantic competition path of path 1 ("red") and path 2 ("blue"), the normalized logarithmic ambiguity A nla The calculated result is approximately 12.835, which comprehensively reflects the similarity or confusion between them at the acoustic and language model levels.
[0081] According to the normalized logarithmic ambiguity A of each semantic competition path pair calculated in the previous step nla Value (for example, a path to calculate A nla =12.835), the system will extract these path pairs and their corresponding Anla value, and each A nla The value is compared with a preset "corpus-level dynamic adaptive threshold". The setting of this threshold is based on the statistical analysis results of large-scale, specific domain (such as smart home LED control commands) voice corpus. The specific setting process is as follows: first, corpus data containing at least 10,000 user voice commands are collected. These data are manually annotated to clarify the cases where the recognition results are typically ambiguous. The current speech recognition system is run on these corpus data to obtain the A values of all semantic competition path pairs. nla Value distribution, and then analyze these A nla values, especially those corresponding to the path pairs that were manually labeled as “highly ambiguous and often need clarification”. nla values, calculate the mean of these values and standard deviation Threshold T dynamic Can be set to Where k is an empirical coefficient, for example, set to 1.5. If we test a test corpus containing 500 known ambiguous samples, we find that when k = 1.5, T dynamic =10.5 (for example, Then T dynamic =7.0+1.5×2.33≈10.5) can better balance the recall rate and precision rate, that is, it can capture most of the real high-risk ambiguities without introducing too many unnecessary clarifications. Then, the threshold is determined. For each path pair, if the calculated A nla Value greater than this T dynamic (For example, 12.835>10.5), the path pair is determined to be high-confidence ambiguous. The system then extracts the two or more competing candidate words contained in this path pair (for example, "red" and "blue") as a word combination (such as ("red", "blue")) and adds it to a set. After traversing all semantically competing path pairs and completing the above comparison and extraction operations, the final set formed is the high-confidence ambiguous candidate set.
[0082] The steps to obtain the clarification question sentence are:
[0083] According to the high-confidence ambiguous candidate set, each candidate word combination in the high-confidence ambiguous candidate set is retrieved one by one, the normalized logarithmic ambiguity within each candidate word combination is sorted, and the two competing candidate word combinations that rank first in the normalized logarithmic ambiguity are selected as target combinations to generate a target candidate word combination;
[0084] Based on the target candidate word combination, the text contents of the two competing candidate words in the target candidate word combination are sequentially extracted, a pre-set clarification question template is called, the text contents of the competing candidate words are respectively filled in the preset placeholder positions in the question template, and the grammar verification of other text contents in the clarification question template is gradually completed to generate the clarification question text;
[0085] Based on the clarification question text, a complete clarification question voice file is synthesized, and the LED screen display content is formatted so that the LED screen text display content corresponds word for word to the voice file content, forming a clarification rhetorical question sentence.
[0086] Specifically, according to the high confidence ambiguous candidate set, the candidate set contains a series of candidate word pairs that are identified as having high ambiguity, such as [(("red", "blue"), 12.835), (("宋体", "黑体"), 11.500)], where each tuple contains a pair of competing candidate words and their corresponding normalized logarithmic ambiguity A nla The system will process this candidate set by first traversing each candidate word combination in this candidate set (i.e. each pair of competing candidate words and their A nla value), these A nla The values are collected and then the A nla The values are sorted in descending order. For example, if the candidate set contains A with ("red", "blue") nla A is 12.835 ("Songti", "Heiti") nla A is 11.500, ("increase", "decrease") nla is 10.200, then the sorted order is ("red", "blue"), then ("Songti", "Heiti"), and finally ("increase", "decrease"). After the sorting is completed, the system selects the two competing candidate words corresponding to the record with the highest normalized logarithmic ambiguity value, that is, the word pair in the first record in the sorted list. For example, if the first place after sorting is (("red", "blue"), 12.835), then the two competing candidate words "red" and "blue" are selected. This pair of selected competing candidate words constitutes the target of the subsequent clarification interaction, generating a target candidate word combination.
[0087] Based on the target candidate word combination, for example, the target candidate word combination determined in the previous step is ("red", "blue"), the system will extract the text content of the two competing candidate words from the combination, namely "red" and "blue". Next, the system calls a pre-set clarification question template library, which stores a variety of question formats for different semantic slots or situations. For example, for color-related ambiguity, there is a template: "Do you want to choose {word A} or {word B} as the color?" For font-related ambiguity, the template may be: "Should the font be set to {word A} or {word B}?" These templates are pre-written by interaction designers based on human-computer interaction design principles and a large amount of user research and stored in the system configuration file. They usually contain one or more placeholders, such as "{word A}" and "{word B}", for dynamically filling in specific candidate words. The system will select the most matching template based on the semantic slot type to which the current target candidate word combination belongs (for example, by decoding the semantic slot information in the word graph to determine whether the current clarification is about "color"). If a template for a specific slot does not exist, a general template is used, such as: "Do you mean {Word A} or {Word B}?" After selecting a template, the system will extract the text content of the two competing candidate words ("red" and "blue") and fill them into the corresponding placeholder positions in the template. For example, "red" is filled into {Word A} and "blue" is filled into {Word B}, forming a preliminary question: "Would you like to choose red or blue as the color?" The filled-in question text is then grammatically checked. This check process first checks whether the sentence structure is complete, such as whether it ends with the expected punctuation mark (such as a question mark). Secondly, a simple check of part of speech and collocation is performed. ,For example, if the template placeholder expects a noun, and the filled-in word is found to be a verb through dictionary query, ,then a backup template or correction logic may be triggered. ,However, in such LED control scenarios, the candidate words are mostly nouns or adjectives, ,with relatively few complex grammatical errors. ,The main verification is whether it is smooth and natural, and does not generate new ambiguities. ,For example, a predefined taboo word list is used to check whether the generated ,question contains inappropriate expressions, and rule-based text smoothing ,is used to adjust the conjunctions or word order to ensure that the ,finally generated clarification question text is grammatically and semantically clear.
[0088] Based on the clarification question text, for example, the text generated in the previous step is "Do you want to choose red or blue as the color?", the system first calls the text-to-speech (TTS) engine to synthesize the clarification question text into a voice signal. The engine is configured with a specific speaker (for example, the system's default "standard Mandarin female voice", the voice library ID is zh-CN-XiaoxiaoNeural), and sets a moderate speaking speed (for example, 1.1 times the normal speaking speed) and a natural tone. The TTS engine processes the input text string and outputs a digital audio data stream, usually in WAV format, 16kHz sampling rate, and 16-bit mono. This audio data is the complete clarification question voice file. While performing voice synthesis, the system also needs to format the original clarification question text for LED screen display content. This processing takes into account the physical characteristics of the LED screen, such as resolution (for example, 128x64 pixels), the number of displayable characters (for example, 16 Chinese characters per line, 4 lines in total), and whether it supports multi-color display (for example, monochrome yellow LED). The formatting process The process includes: automatic line wrapping according to the screen width to avoid improper interruption of words or phrases at the end of the line, selecting the appropriate font (for example, the system's built-in 16x16 dot matrix Songti) and font size, and for longer questions, paging or scrolling. The core is to ensure that the text content finally displayed on the LED screen corresponds to the content of the voice file synthesized by TTS in time. To achieve this synchronization, the TTS engine will synchronously output the pronunciation start and end timestamp information of each word or phrase (depending on the TTS capability and configuration) when synthesizing speech (for example, For example, a timestamp sequence containing [("you", 0.0, 0.2), ("yes", 0.2, 0.35), ("want", 0.35, 0.5), ...]) can be used by the LED display control program to highlight or display the corresponding text on the screen word by word based on this timestamp information when the voice playback reaches a certain word or phrase. For example, when the voice broadcasts "red", the word "red" on the LED screen will be displayed in reverse color or with different flashing effects. Through this close coordination of voice and visuals, a clarifying rhetorical question is formed.
[0089] The steps to obtain user clarification instructions are as follows:
[0090] According to the clarification-type rhetorical question sentence pattern, the voice content of the clarification-type rhetorical question sentence pattern is played word by word, and at the same time, the LED screen display control interface is called to synchronously display the text content of the clarification-type rhetorical question sentence pattern word by word, so that the text display on the LED screen and the voice playback content are consistent on the time axis, forming a synchronous broadcast state;
[0091] Based on the synchronous broadcast status, the capture starts with the first valid silence segment after the end of the clarification question and ends with the next silence segment. The user's real-time clarification speech during the silence interval is intercepted and the real-time clarification speech is segmented into time domain windows, noise filtered, and feature extracted to generate the user's clarification speech data to be identified.
[0092] Based on the user clarification voice data to be recognized, the user clarification voice data is decoded frame by frame, and the recognition text whose confidence level of the recognition result reaches the preset recognition confidence level standard is screened as a valid result to form a user clarification instruction.
[0093] Specifically, according to the clarification rhetorical question sentence pattern, the sentence pattern consists of a clarification question voice file and synchronized LED screen text display content. For example, the voice content is "Do you want to choose red or blue as the color?", and the LED screen also displays this question word by word. The system plays the content of the clarification question voice file word by word through an audio playback interface (for example, using an audio service such as PulseAudio or ALSA). At the same time, the system calls the LED screen display control interface (for example, a custom protocol interface that communicates with the LED driver chip via an SPI or I2C bus) to control the timing and method of displaying the text on the LED screen based on the timestamp information of each word or phrase previously generated by TTS. For example, when the voice plays the word "you", the word "you" is highlighted on the LED screen. Then, when the voice plays the word "yes", the word "yes" is highlighted, and the remaining words remain normally displayed or dimmed. This word-by-word synchronization method continues until the entire clarification question is played. In this way, the text display content on the LED screen and the voice playback content are strictly consistent on the timeline, forming a synchronized broadcast state.
[0094] Based on the synchronous broadcast state, once the voice playback and LED screen display of the clarification question sentence are completely completed, the system immediately starts Voice Activity Detection (VAD). The VAD continuously monitors the audio stream input by the microphone to identify speech segments and silence segments. The implementation of VAD can be based on multiple acoustic features such as energy threshold, short-term zero-crossing rate, and fundamental frequency detection. For example, when the audio energy for 200 consecutive milliseconds is lower than 1.5 times the preset ambient noise baseline energy and the short-term zero-crossing rate is lower than 0.8 times the average zero-crossing rate, it is determined to be the beginning of a silence segment. Otherwise, it is a speech segment. The system will use the end point of the first valid silence segment identified by the VAD module after the end of the clarification question (i.e., the moment when the user starts speaking) as the starting point for capturing the user's clarification speech. Here, "valid silence segment" refers to silence that lasts longer than a minimum threshold (e.g., 150 milliseconds) to filter out short pauses or noise. Subsequently, the system continues to monitor the audio stream until the VAD module detects the beginning point of the next silence segment (i.e., the moment when the user speaks). The system captures the audio data between these two time points, i.e., the user's real-time clarification during this period of silence. For example, if the user answers "red," this period of speech is captured completely. Next, the captured real-time clarification speech undergoes preprocessing. First, time-domain window segmentation is performed, dividing the continuous speech signal into frames of fixed length, for example, 25 milliseconds per frame with a frame shift of 10 milliseconds. A Hamming window is applied to each frame. Noise is then filtered using techniques such as spectral subtraction or Wiener filtering. The spectrum of each frame is processed based on an estimated background noise spectrum (for example, noise samples collected at system startup or when the user has been silent for an extended period). For example, if power-frequency noise near 50 Hz or device fan noise near 2 kHz is detected, the energy in these frequency bands is specifically attenuated. Finally, acoustic features are extracted from each frame of the denoised speech, typically Mel-frequency cepstral coefficients (MFCCs) and their first- and second-order differences. For example, a 39-dimensional feature vector is extracted to form the user's clarification speech data to be recognized.
[0095] Based on the clarified speech data of the user to be identified, the data is a series of acoustic feature vector sequences that have been preprocessed (framing, windowing, noise reduction, and feature extraction), such as a 39-dimensional MFCC feature vector every 10 milliseconds. The system calls the pre-trained speech recognition decoder to decode these feature vectors frame by frame. The decoder contains an acoustic model, a language model, and a pronunciation dictionary. The acoustic model is responsible for mapping acoustic features to phonemes or subword units, and the language model guides the search process according to the probability of the word sequence. The decoder uses Viterbi search or other efficient search algorithms (such as beam search) to find the optimal word sequence path in the decoding word graph or state network, and outputs one or more candidate recognition results and their corresponding confidence scores. This confidence score combines the acoustic model score and the language model score to reflect the reliability of the recognition result. The system will screen these Of these recognition results, only those recognition texts whose confidence reaches or exceeds a preset recognition confidence standard are retained as valid results. The setting of this recognition confidence standard is based on the evaluation of a large number of clarification speech recognition results, and aims to strike a balance between accepting correct recognition and rejecting incorrect recognition. For example, by testing a data set containing 5,000 user clarification voices, it was found that when the confidence threshold was 0.85, a recognition accuracy of 95% could be achieved while the false acceptance rate was less than 2%. 0.85 was set as the standard. If the confidence of a recognition result "red" is 0.92, which is higher than 0.85, it is considered valid. If the confidence of the recognition result "backboard" is 0.60, which is lower than 0.85, it may be discarded or marked as low confidence. Finally, these valid recognition texts that pass the confidence screening are integrated into user clarification instructions.
[0096] The steps to obtain the unique identification path are:
[0097] According to the user's clarification instruction, the character sequence of the user's clarification instruction is parsed word by word, and the core keyword text in the user's clarification instruction is extracted, and the semantic features of the core keyword text are marked one by one to generate a semantic feature sequence of the user's clarification instruction;
[0098] Based on the semantic feature sequence of the user's clarification instruction, the semantic feature sequence of the competing path in the decoding word graph is called, and the semantic feature sequence of each competing path is compared character by character. The semantic feature sequence of the user's clarification instruction is matched with the semantic feature sequence of each competing path in turn to generate a semantic feature matching degree sequence;
[0099] Based on the semantic feature matching degree sequence, the semantic feature matching degree of each competing path is ranked, and the competing path with the highest semantic feature matching degree is selected as the path that fully matches the user's clarification instruction to establish a unique identification path.
[0100] Specifically, according to the user's clarification instruction, for example, the clarification instruction text spoken by the user is "I choose red" or directly "red", the system first parses the text string word by word and decomposes it into a sequence of single characters, such as ['I', 'select', 'red', 'color'] or ['red', 'color']. Then, the system performs a core keyword extraction operation. The core keyword refers to the vocabulary that can directly correspond to the previous ambiguous options in the clarification context. It is usually identified through a predefined domain dictionary or a statistical keyword extraction algorithm (such as TF-IDF, but here it is more likely to be a simple word list matching because the vocabulary range of the clarification interaction is limited). For example, for color clarification, the dictionary contains "red", "blue", "green", etc. If the user instruction is "I choose red", then "red" is extracted as the core keyword. If the user instruction is "the red one ", then through certain fuzzy matching or synonym expansion (for example, the preset "red" is equivalent to "red"), "red" will also be identified. For each extracted core keyword text, the system will mark it with predefined semantic features. These semantic features are internal codes or labels set in advance for each possible LED control parameter value (such as color, font, etc.). For example, the semantic feature of "red" may be COLOR_RED, and the semantic feature of "Songti" is FONT_SIMSUN. These features are identifiers used within the system to uniquely identify specific parameter options. In this way, each core keyword in the user clarification instruction is converted into a representation of one or more semantic features, and finally a user clarification instruction semantic feature sequence is formed. For example, for the user instruction "red", its semantic feature sequence is [COLOR_RED].
[0101] Based on the semantic feature sequence of the user's clarification instruction, for example, the sequence obtained in the previous link is [COLOR_RED], the system will then call the competing paths (for example, path one represents "red", path two represents "blue") that were selected due to high confidence ambiguity in the decoding word graph (Lattice) generated previously when decoding the original speech signal. For each such competing path, the system will also pre-generate a corresponding semantic feature sequence for it. This generation method is similar to processing the user's clarification instruction, that is, extracting the core words in the path and marking their semantic features. For example, the semantic feature sequence of path one "red" in the decoding word graph is [COLOR_RED], and the semantic feature sequence of path two "blue" is [COLOR_BLUE]. Then, the system will perform a character-level precise comparison with the semantic feature sequence of the user's clarification instruction ([COLOR_RED]) and the semantic feature sequence of each competing path in the decoding word graph, or more accurately, a precise match of feature identifiers. For example, [C [COLOR_RED] is compared with [COLOR_RED] of path one, and [COLOR_RED] is compared with [COLOR_BLUE] of path two. The purpose of the comparison is to calculate a semantic feature matching degree. The simplest matching degree can be a complete match (score of 1) or an incomplete match / no match (score of 0). More complex edit distance or vector similarity-based methods can also be used (if the semantic features themselves are vector representations). However, here, since the semantic features are predefined atomic labels, exact matching is usually sufficient. Therefore, the matching degree of [COLOR_RED] with [COLOR_RED] is 1, and the matching degree of [COLOR_RED] with [COLOR_BLUE] is 0. After performing this matching operation on all competing paths, a matching degree list is obtained, where each element corresponds to a competing path and its semantic feature matching degree with the user's clarification instruction, forming a semantic feature matching degree sequence, for example, [(path one, 1), (path two, 0)].
[0102] Based on the semantic feature matching sequence, for example, the sequence generated in the previous step is [(path one, 1), (path two, 0)], where path one corresponds to "red" and path two corresponds to "blue", and the user's clarification instruction points to "red". The system will sort the items in this sequence in descending order according to their semantic feature matching values. In this example, since path one ("red") with a matching degree of 1 is higher than path two ("blue") with a matching degree of 0, the sorted result is still path one in front and path two in the back, that is, [(path one, 1), (path two, 0)]. After the sorting is completed, the system selects the competing path that ranks first in the semantic feature matching ranking result, that is, the path with the highest matching degree, as the path that fully matches the user's clarification instruction. In this example, path one ("red") is selected, and this selected path is established as the only identification path in the current interaction context that has been clarified and confirmed by the user, thus establishing a unique identification path.
[0103] The steps to obtain the LED display state vector are:
[0104] Based on the unique identification path, each node included in the unique identification path is visited one by one, the preset text label and corresponding node feature in the node are parsed node by node, the node feature information associated with the LED screen display content is extracted, and a node feature information set is generated;
[0105] According to the node feature information set, the color configuration code and font configuration code in the node feature information set are retrieved one by one, matched through the color coding table and the font coding table, and the color configuration code and font configuration code are converted into corresponding color RGB values and font size, style and thickness parameters respectively to generate an LED display parameter set;
[0106] Based on the LED display parameter set, the color RGB values, font size, font style, and font thickness parameters in the set are called in sequence, and the resulting color and font parameter combinations are encapsulated into a vectorized data format to generate an LED display state vector.
[0107] Specifically, based on the unique identification path, the path is a specific sequence in the decoding word graph (Lattice), which represents the final recognition result of the user's intention. For example, the path contains the word node sequence [<START_OF_SENTENCE> Please set the color to red.<END_OF_SENTENCE> ], the system will follow this unique identification path, starting from the starting node, and visit each word node included in the path one by one. For each visited node, the system will parse the preset information stored inside it, which includes the text label corresponding to the node (that is, the recognized word, such as "red") and the node features associated with the word. Node features are usually stored in the form of key-value pairs. For example, the node of the word "red" may contain the features {"slot_type":"COLOR","value_code":"C_RED","display_text":"red"}, where "slot_type" indicates the type of semantic slot filled by the word, and "value_code" is the internal "display_text" is the actual displayed text. The system will extract node feature information directly related to the content displayed on the LED screen. For example, for the word "red", the relevant feature information may be its parameter code C_RED and the corresponding display text "red". For functional words such as "please", "will", "set", and "for", they may not have a direct association with LED display parameters, or their associated information is used to control the overall display logic rather than specific content. The system collects all extracted feature information related to the LED display (such as parameter codes, display text, and possibly other codes such as brightness adjustment instructions) to form a node feature information set.
[0108] According to the node feature information set, for example, the set contains [{"type":"COLOR","code":"C_RED","text":"red"},{"type":"FONT_SIZE","code":"FS_LARGE","text":"large"}], the system will retrieve each feature information in this set one by one, specifically looking for the color configuration code (such as C_RED) and font configuration code (such as FS_LARGE) contained therein. Once these codes are found, the system will use the pre-established color coding table and font coding table for matching queries. The color coding table is a mapping relationship table that maps the internal color configuration code (such as C_RED) to a specific color representation, usually the RGB (red, green, and blue) three-channel value. For example, C_RED corresponds to (255, 0, 0), and C_BLUE corresponds to (0, 0, 255). This color coding table is defined during system design based on the color range supported by the LED hardware and user requirements. For example, if the LED supports 16 Similarly, the font encoding table maps internal font configuration codes (such as FS_LARGE) to specific font parameters, including font size (for example, FS_LARGE corresponds to 32 pixels high), font style (for example, STYLE_NORMAL corresponds to regular, STYLE_ITALIC corresponds to italic), and font weight (for example, WEIGHT_BOLD corresponds to bold, WEIGHT_NORMAL corresponds to regular). These font parameters are also defined according to the display capabilities of the LED screen and the preset font library. For example, the font size may have three options: small (16 pixels), medium (24 pixels), and large (32 pixels), and the font style and weight also have several preset values. Through table lookup operations, the system converts the abstract color configuration codes and font configuration codes into actual color RGB values that can be used for LED driving, as well as font size, style, and weight parameters. These converted parameters are then collected to form an LED display parameter set.
[0109] Based on the LED display parameter set, for example, the set now contains specific parameter values such as [{"color_rgb": (255, 0, 0)}, {"font_size": 32, "font_style": "NORMAL", "font_weight": "NORMAL"}], the system will call each parameter in this set in a predetermined order or logic (for example, color first and then font, or in the order in which they appear in the original instruction), first extracting the color RGB value, such as (255, 0, 0), then extracting the font size parameter, such as 32 (pixels), then extracting the font style parameter, such as "NORMAL" (regular), and finally extracting the font thickness parameter, such as "NORMAL" (regular). The system combines the specific color and font parameters obtained by these calls and encapsulates them into a standardized, easy-to-process and easy-to-transmit vectorized data format. The specific form of this vectorized data format depends on the subsequent processing The module requirement, for example, can be a structure or JSON object containing fixed fields, such as {"color": [255, 0, 0], "font": {"size": 32, "style": "NORMAL", "weight": "NORMAL"}, "text_content": "text content to be displayed"} (the text content is usually also obtained from the node feature information set), or a flattened numerical vector, in which the elements at specific positions represent specific parameters, for example, a vector of length N, the first 3 elements are RGB values, the fourth element is the font size, the fifth element is the font style encoding (such as 0 for NORMAL, 1 for ITALIC), and the sixth element is the font weight encoding (such as 0 for NORMAL, 1 for BOLD). In this way, the scattered display parameters obtained by parsing are integrated into a unified, structured data representation to generate the LED display state vector.
[0110] The steps for obtaining the context-related vocabulary set are:
[0111] After receiving a new voice command, the system performs frame recognition on the voice command, extracts all voice words and constructs a word vector for each voice word. It also extracts the color RGB values and font size of all text elements in the current LED display state vector to generate a voice word vector set and a screen text vector set.
[0112] Based on the speech word vector set and the screen text vector set, the visual-semantic association between each speech word and all screen text elements is calculated using the following formula:
[0113]
[0114] in, VS vsa (w) represents the visual-semantic relevance of the speech word w, M is the number of screen text elements, is the semantic similarity between the speech word w and the jth screen word, v w and are their respective word vectors, ZR j , G j 、B j is the color channel value of the jth screen word, T j is the font size of the jth screen word, T ref is the reference font size constant used to achieve normalization processing, V j is the visual impact of the j-th screen word;
[0115] Based on the visual-semantic correlation of speech words, all speech words whose visual-semantic correlation is higher than the average correlation of all words are screened, and the corresponding original texts are grouped together to generate a context-related vocabulary set.
[0116] Specifically, after receiving a new voice command, for example, the user says "turn up the brightness", the system first processes the input voice signal and converts it into text through a voice recognition process. This voice recognition process includes pre-emphasis, framing (for example, 25 milliseconds per frame, 10 milliseconds frame shift) and windowing (for example, Hamming window) of the audio signal, and then extracts the acoustic features of each frame (for example, 39 Mel-frequency cepstral coefficients and their differences). The system uses a pre-trained acoustic model (for example, a Transformer-based end-to-end acoustic model trained on a dataset containing 100,000 hours of general speech and 500 hours of LED control-specific speech, with input as an acoustic feature sequence and output as a phoneme or subword unit sequence) and a language model (for example, a 4-gram language model optimized for LED control commands, which is built based on a text corpus containing 100,000 LED control commands and related conversations) for decoding to obtain a recognized voice word sequence, such as ["turn up"]. "brightness", "adjust", "high", "some"], then, for each recognized speech word, a corresponding word vector is constructed. This is done through a pre-trained word vector model, for example, a Word2Vec model with a Skip-gram architecture. The model is trained on a general Chinese text corpus containing millions of words and a domain corpus containing 500,000 LED device operating instructions and product descriptions. During training, the word vector dimension is set to 100, the context window size is 5, the number of negative sampling samples is 10, the initial value of the learning rate is 0.025 and linearly decays to 0. Through this model, each speech word (such as "brightness") is mapped to a 100-dimensional real number vector. For example, the word vector of the word "brightness" is [-0.23, 0.51, ..., 1.05]. All these word vectors together constitute a speech word vector set. At the same time, information is extracted from the current LED display state vector maintained internally. The LED display state vector (for example, the content is
[0117] [{"text_content":"mode","color_rgb":[100, 100, 100],"font_size":16},
[0118] {"text_content": "current brightness", "color_rgb": [255, 0, 0], "font_size": 24}]) records each text element currently displayed on the LED screen and its visual attributes. The system traverses this state vector and extracts the actual text content of each text element (such as "mode", "current brightness"), its displayed color RGB value (such as [100, 100, 100], [255, 0, 0]) and font size (such as 16 pixels, 24 pixels). For the text content of each extracted screen text element, the aforementioned Word2Vec model is also used to construct a word vector for it, such as the word vector of "current brightness", forming the word vector of the screen text element and its corresponding color and font size information. After summarizing this information, a set of speech word element vectors and a set of screen text vectors are generated.
[0119] formula: in, The benefit of the formula is that it can identify new instruction words that are most closely related to the current screen content (especially visually prominent and semantically related elements) by quantifying the visual-semantic association between speech words and text elements on the screen, which provides key information for subsequent context bias. Specifically, V j The item integrates the visual impact of screen text. The color component reflects its brightness or saturation through the normalized Euclidean distance, and the font size component highlights the impact of large font size through the logarithmic function but avoids its excessive dominance. The term measures the semantic similarity between spoken and screen words, ensuring the semantic basis of the association; the max(0,...) operation ensures that only positive semantic associations are considered; the final weighted sum aggregates the influence of all screen elements on a single spoken word, thereby comprehensively evaluating its contextual relevance.
[0120] The M parameter is obtained as follows: M represents the total number of text elements currently displayed on the LED screen. This value is directly obtained from the current LED display state vector, which records information about all independent text blocks on the screen. M is obtained by traversing this vector and counting the number of entries containing text content. For example, if the LED screen currently displays two lines of text, the first line is "Current Mode: Energy Saving" and the second line is "Brightness: 50%", then there are two text elements, and M = 2.
[0121] v w The steps to obtain the parameters are: wThe word vector represents the specific speech word w in the new speech instruction currently being processed. The word vector is obtained by mapping the word w to a high-dimensional real number space through a pre-trained word embedding model (such as the Word2Vec model described in the previous step, with a dimension of 100). For example, for the word "red" in the new instruction, its corresponding 100-dimensional word vector v 红色 =[0.12,-0.34,...,0.56] is obtained from the constructed speech word element vector set.
[0122] The steps to obtain the parameters are: Represents the jth text element t displayed on the LED screen j The word vector of v w In the same way as the acquisition method, the text elements on the screen (such as "mode" and "brightness") are also converted into 100-dimensional word vectors through the same pre-trained Word2Vec model. For example, if the first text element on the screen is "mode", its word vector v 模式 =
[0123] [0.45, 0.02, ..., -0.11] is obtained from the constructed screen text vector set.
[0124] The steps to obtain the parameters are: The word vector v representing the phonetic word w w With the jth screen text element t j word vector The semantic similarity between them is calculated using cosine similarity, and the calculation formula is: The result ranges from -1 to 1, and the closer the value is to 1, the more similar the semantics are. For example, to calculate the cosine similarity between the speech word "red" and the screen word "color", if v 红色 =[0.1,0.8] and v 颜色 =[0.2,0.7], then v 红色 ·v 颜色 =0.1×0.2+0.8×0.7=0.02+0.56=0.58,
[0125] ZR j , G j 、B jThe parameters are obtained as follows: These represent the color channel values of the jth on-screen text element, namely the intensities of red (Red), green (Green), and blue (Blue). Each value ranges from 0 to 255. These values are extracted from the current LED display state vector for the jth text element. For example, if the first text element on the screen is displayed as pure red, its color channel values are ZR1 = 255, G1 = 0, and B1 = 0.
[0126] T j The steps to obtain the parameters are: T j Indicates the font size of the jth screen text element, usually in pixels (px). This value is extracted from the current LED display state vector for the jth text element. For example, if the font size of the first text element on the screen is 24 pixels, then T1 = 24.
[0127] T ref The steps to obtain the parameters are: T ref It is a reference font size constant used to normalize the font size. Its setting is based on the statistical analysis of the commonly used font sizes of the system LED display screen. A representative basic or smaller font size is selected as a reference. For example, by analyzing the 100 common display templates preset by the system, it is found that the smallest commonly used font size is 12 pixels, and 12 pixels and 16 pixels are the most frequently used auxiliary information font sizes. Therefore, T is selected. ref = 12 pixels as a reference.
[0128] Calculation process:
[0129] In the current new voice command, a word w is “brightness” and there are two text elements on the screen (M=2):
[0130] Element 1 (j=1): text t1=“MODE”, color ZR1=100, G1=100, B1=100 (gray), font size T1=16px.
[0131] Element 2 (j=2): text t2=“Current Brightness”, color ZR2=255, G2=0, B2=0 (red), font size T2=24px.
[0132] Reference font size T ref =12px.
[0133] Assume that the semantic similarity calculated by word vector is:
[0134] sim(v 亮度 ,v 模式 )=0.2;
[0135] sim(v 亮度 ,v当前亮度 )=0.9;
[0136] Calculate the visual impact V1 of the first screen text element:
[0137] Color part:
[0138] Normalized denominator:
[0139] Color impact:
[0140] Font size part:
[0141] V1=0.392·0.847≈0.332;
[0142] Calculate the visual impact V2 of the second screen text element:
[0143] Color part:
[0144] Color impact:
[0145] Font size part:
[0146] V2=0.577·1.0986≈0.634;
[0147] Calculate the visual-semantic correlation of the "brightness" of the speech word VS vsa (brightness):
[0148] VS vsa (brightness) = V1·max(0,sim(v 亮度 ,v 模式 ))+V2·max(0,sim(v 亮度 ,v 当前亮度 ));
[0149] VS vsa (brightness) = 0.332 max(0,0.2) + 0.634 max(0,0.9);
[0150] VS vsa (brightness) = 0.332 0.2 + 0.634 0.9;
[0151] VS vsa (brightness) = 0.0664 + 0.5706;
[0152] VS vsa (brightness) = 0.637;
[0153] The results show that the comprehensive visual-semantic correlation between the phonetic word "brightness" and all text elements on the current screen is 0.637, which reflects the strength of the association between the word "brightness" and the screen content (especially the visually prominent and semantically relevant "current brightness").
[0154] Based on the visual-semantic association previously calculated for each speech word in the new instruction, for example, for the speech instruction "turn up the brightness a little", the associations of each word may be: {"turn": 0.15, "brightness": 0.637, "turn up": 0.30, "some": 0.10}. The system first calculates the mean of these associations by adding up the visual-semantic associations of all words and then dividing it by the total number of words. In this example, the mean is (0.15+0.637+0.30+0.10) / 4=1.187 / 4≈0.29675. This mean is used as a dynamic screening threshold. Then, the system traverses each speech word in the instruction and compares its respective visual-semantic association with the calculated value. The calculated mean is compared. If the association degree of a word is greater than this mean, the word is selected. For example, the association degree of "brightness" is 0.637, which is greater than the mean of 0.29675, so "brightness" is selected. The association degree of "increase" is 0.30, which is also slightly greater than the mean of 0.29675 (depending on the comparison accuracy, if it is strictly greater than it is selected), and is also selected. The association degrees of "will" (0.15) and "some" (0.10) are both less than the mean, so they are not selected. Finally, the system collects the original texts (i.e., the words themselves) corresponding to all the selected speech words to form a set. For example, if "brightness" and "increase" are selected, the set formed is {"brightness", "increase"}, which is the context-related vocabulary set.
[0155] The steps for obtaining context bias instructions are:
[0156] Based on the context-related vocabulary set, the text content of each word in the context-related vocabulary set is extracted one by one, a context-related vocabulary text index list is established, and a search is performed item by item in the vocabulary library of the language model to determine whether the context-related vocabulary text exists in the language model vocabulary library, and a valid context-related vocabulary list is generated;
[0157] According to the list of valid context-related words, the original prior probability value corresponding to each valid context-related word in the language model is extracted one by one, the probability value of each valid context-related word is increased according to a fixed ratio, and the corresponding word probability parameters in the language model are updated one by one to generate a probability-enhanced context-related word prior probability set;
[0158] Based on the probability-enhanced context-associated vocabulary prior probability set, the audio feature sequence of the newly input voice command is decoded frame by frame. For each decoding step of the newly input voice command, the enhanced probability parameters in the probability-enhanced context-associated vocabulary prior probability set are called to update the decoding score of the corresponding vocabulary and generate a context bias instruction.
[0159] Specifically, based on the context-related vocabulary set, for example, the set is {"brightness", "adjust up"}, the system first extracts the text content of each word in this set one by one, that is, "brightness" and "adjust up", and then organizes these text contents into an ordered list, called the context-related vocabulary text index list, such as ["brightness", "adjust up"]. Next, the system accesses the vocabulary library of the language model loaded inside it, which is a data structure that contains all the words that the model can recognize and process and their related statistical information (such as frequency of occurrence, N-gram probability, etc.), for example, a dictionary implemented based on a hash table, with the key being the vocabulary text and the value being the vocabulary ID and probability information. The system traverses each vocabulary text ("brightness", "increase") in the context-related vocabulary text index list and searches each one of them in the vocabulary library of the language model. The purpose of the search is to determine whether there is an entry in the vocabulary library that is exactly the same as the current context-related vocabulary text. If the search is successful, that is, the corresponding vocabulary text is found in the vocabulary library, the vocabulary text is considered valid and added to a new list. If a context-related vocabulary text does not exist in the vocabulary library (for example, due to a segmentation error or the word is an unregistered word), the word is not considered valid. After searching and judging all context-related vocabulary texts, a valid context-related vocabulary list is finally formed.
[0160] According to the list of valid context-related words, for example, the list is ["brightness", "increase"], the system extracts the corresponding original prior probability value from the data structure of the language model for each valid context-related word in the list. This original prior probability refers to the single-word probability of the word, which reflects the prevalence of the word in the corpus of the training language model or the possibility of its appearance in a specific context. For example, the original prior probability of the word "brightness" may be P(brightness) = 0.005, and the original prior probability of the word "increase" may be P(increase) = 0.002. After obtaining these original probability values, the system increases the original prior probability value of each valid context-related word according to a preset fixed ratio. This fixed ratio is a hyperparameter, and its value is set based on experience tuning and experimental results. It aims to moderately improve the probability of occurrence of these context-related words while avoiding recognition errors caused by excessive bias. For example, if the fixed ratio is set to 1.5, the enhanced probability P ′(lexical) = P(lexical) * 1.5, so the enhanced probability of "brightness" becomes 0.0075, the system temporarily or during the current decoding session updates these enhanced probability values to the probability parameters of the corresponding lexicons in the language model, i.e. replaces the original prior probability values, for N-gram model, this updating process is directly modifying the corresponding entries in the lookup table storing N-gram probabilities, for neural network language model, it is adding a bias term to the logits before computing the output layer softmax, all the lexicons that have been enhanced and their new prior probability values together form the contextually related lexicon prior probability set enhanced by probability.
[0161] Based on the contextually related lexicon prior probability set enhanced by probability, which contains lexicons and their enhanced prior probabilities such as {"brightness": 0.0075, "turn up": 0.003}, the system starts to decode the audio feature sequence (for example, the 39-dimensional MFCC feature sequence extracted before) of the new input voice command frame by frame, the decoding process uses standard speech recognition decoding algorithms such as Viterbi search or beam search, at each step of the decoding, when the decoder needs to evaluate the score of the path extended to a certain lexicon (for example, lexicon w i ), the acoustic model score and the language model score are considered comprehensively, for the calculation of the language model score, if the current lexicon w i to be evaluated happens to exist in the contextually related lexicon prior probability set enhanced by probability, then the decoder will no longer use its original prior probability when calculating its language model probability, but call the enhanced probability parameter stored in the set for w i , (for example, if w i is "brightness", use 0.0075), if the lexicon to be evaluated is not in this enhanced set, then use its probability in the original language model normally, in this way, those lexicons related to the previous on-screen context will get higher overall scores in the decoding process because of the improvement of their language model probabilities, so they are more advantageous in the competition with other candidate words and are more likely to be recognized, the final recognition result generated after applying these enhanced probability parameters in the entire decoding process is the contextually biased instruction.
[0162] The above is only the preferred embodiment of the present application, not other forms of the present application limit, any skilled in the art of technical personnel may use the above disclosed technical content to change or modify as equivalent variation of equivalent embodiments applied to other fields, but without departing from the technical solution of the present application, according to the technical essence of the present application to the above embodiment of any simple modification, equivalent change and modification, still belongs to the protection scope of the technical scheme of the present application.
Claims
1. A method for LED multi-dimensional information release based on voice intention recognition, characterized in that: The following steps are involved: Decode the input speech signal to generate a decoded word graph, locate the preset LED multi-dimensional information semantic slots in the decoded word graph, extract similar competitive paths, and calculate a high-confidence ambiguous candidate set; Based on the high-confidence ambiguous candidate set, two competing candidate words are extracted and embedded into a preset question template to generate a clarifying rhetorical question. Based on the clarifying rhetorical question, the voice broadcast is driven to synchronously display the question on the LED screen, and the user's clarification voice is captured to obtain the user's clarification instruction; Based on the user clarification instruction, searching for a path matching the user clarification instruction in the competitive paths of the decoding word graph, establishing a unique identification path, parsing node information based on the unique identification path, extracting corresponding LED display color parameters and LED display font parameters, and establishing an LED display state vector; When a new voice command is received, the current LED display state vector is called, and the corresponding command vocabulary is retrieved according to the LED display color parameters and the LED display font parameters, and a context-related vocabulary set is established. Based on the context-related vocabulary set, in the process of decoding the new voice command, bias weights are applied to the vocabulary belonging to the context-related vocabulary set in the language model to increase the prior probability and generate a context-biased instruction.
2. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the high confidence ambiguous candidate set are: Extract the Mel-frequency cepstral coefficients from the input speech signal, generate a decoding word graph for the corresponding time frame, traverse all path groups in the word graph that include the semantic slots of LED information, filter out path pairs with overlapping frame positions and similarity, and generate a set of semantically competitive path pairs; Calculating a normalized logarithmic ambiguity based on the set of semantic competition path pairs; According to the normalized logarithmic ambiguity, all path pairs whose normalized logarithmic ambiguity is greater than a corpus-level dynamic adaptive threshold are extracted, and the corresponding candidate words are combined and output to generate a high-confidence ambiguous candidate set.
3. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the clarification rhetorical question sentence are: According to the high-confidence ambiguous candidate set, each group of candidate word combinations in the high-confidence ambiguous candidate set is retrieved one by one, the normalized logarithmic ambiguity within each group of candidate word combinations is sorted, and the two competing candidate word combinations that rank first in the normalized logarithmic ambiguity are selected as target combinations to generate a target candidate word combination; Based on the target candidate word combination, the text contents of two competing candidate words in the target candidate word combination are sequentially extracted, a pre-set clarification question template is called, the text contents of the competing candidate words are respectively filled in the preset placeholder positions in the question template, and the grammar check of other text contents in the clarification question template is gradually completed to generate a clarification question text; Based on the clarification question text, a complete clarification question voice file is synthesized, and the LED screen display content is formatted so that the LED screen text display content corresponds word for word to the voice file content, forming a clarification rhetorical question sentence pattern.
4. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the user clarification instruction are: According to the clarification-type rhetorical question sentence, the voice content of the clarification-type rhetorical question sentence is played word by word, and at the same time, the LED screen display control interface is called to synchronously display the text content of the clarification-type rhetorical question sentence word by word, so that the LED screen text display and the voice playback content are consistent on the time axis, forming a synchronous broadcast state; Based on the synchronous broadcast state, the first valid silent segment after the end of the clarification question is used as the capture starting point, and the next silent segment is used as the capture end point. The user's real-time clarification speech during the silent interval is intercepted, and the real-time clarification speech is segmented into time domain windows, noise filtered, and feature extracted to form the user's clarification speech data to be identified; Based on the user clarification voice data to be recognized, the user clarification voice data is decoded frame by frame, and recognition texts whose confidence levels of the recognition results meet a preset recognition confidence level are screened as valid results to form user clarification instructions.
5. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the unique identification path are: According to the user clarification instruction, the character sequence of the user clarification instruction is parsed word by word, and the core keyword text in the user clarification instruction is extracted, and the semantic features of the core keyword text are marked one by one to generate a user clarification instruction semantic feature sequence; Based on the semantic feature sequence of the user clarification instruction, calling the semantic feature sequence of the competing path in the decoding word graph, performing character comparison on the semantic feature sequence of each competing path one by one, and sequentially comparing and matching the semantic feature sequence of the user clarification instruction with the semantic feature sequence of each competing path to generate a semantic feature matching degree sequence; Based on the semantic feature matching degree sequence, the semantic feature matching degree of each competing path is ranked, and the competing path with the highest semantic feature matching degree ranking is selected as the path that fully matches the user clarification instruction to establish a unique identification path.
6. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the LED display state vector are: Based on the unique identification path, each node included in the unique identification path is accessed one by one, a preset text label and a corresponding node feature in each node are parsed node by node, node feature information associated with the LED screen display content is extracted, and a node feature information set is generated; According to the node feature information set, the color configuration codes and font configuration codes in the node feature information set are retrieved one by one, matched through the color coding table and the font coding table, and the color configuration codes and font configuration codes are converted into corresponding color RGB values and font size, style and thickness parameters respectively to generate an LED display parameter set; Based on the LED display parameter set, the color RGB values, font size, font style and font thickness parameters in the set are called in sequence, and the color and font parameter combinations obtained by the calls are packaged into a vectorized data format to generate an LED display state vector.
7. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps of obtaining the context-related vocabulary set are: After receiving a new voice command, the system performs frame recognition on the voice command, extracts all voice words and constructs a word vector for each voice word. It also extracts the color RGB values and font size of all text elements in the current LED display state vector to generate a voice word vector set and a screen text vector set. Calculating the visual-semantic association between each speech word and all screen text elements based on the speech word vector set and the screen text vector set; Based on the visual-semantic relevance of the speech word-units, all speech word-units whose visual-semantic relevance is higher than the average relevance of all word-units are screened, and the corresponding original texts are formed into a set to generate a context-related vocabulary set.
8. The LED multi-dimensional information publishing method based on voice intention recognition according to claim 1 is characterized in that: The steps for obtaining the context bias instruction are: Based on the context-related vocabulary set, extract the text content of each word in the context-related vocabulary set one by one, establish a context-related vocabulary text index list, and search one by one in the vocabulary library of the language model to determine whether the context-related vocabulary text exists in the language model vocabulary library, and generate a valid context-related vocabulary list; Extracting, one by one, the original prior probability value corresponding to each valid context-related word in the language model according to the valid context-related word list, increasing the probability value according to a fixed ratio for each original prior probability value of the valid context-related word, updating the corresponding word probability parameters in the language model one by one, and generating a probability-enhanced context-related word prior probability set; Based on the probability-enhanced context-associated vocabulary prior probability set, the audio feature sequence of the newly input voice instruction is decoded frame by frame. For each decoding step of the newly input voice instruction, the enhanced probability parameter in the probability-enhanced context-associated vocabulary prior probability set is called, the decoding score of the corresponding vocabulary is updated, and a context bias instruction is generated.