Intelligent keyboard multi-modal input signal fusion processing method and system

CN122654980APending Publication Date: 2026-08-28SHENZHEN JUPENG ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611020028.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]传统的智能键盘多模态输入信号融合处理方法,虽然能够对敲击按压触控滑动或者语音等不同形式的输入信号进行统一提取、组合与集中处理,但是,在用户使用键盘语音转换功能时,输入音频片段的起始位置可能受到静音段、缓冲器待读取音频帧变化迟滞以及按键连续触发时间关系的影响,存在音频起始时间与按键触发时间难以准确对应的问题,进而使按键输入序列和输入音频片段在是否属于同一次用户输入操作的判定中容易出现协同关系不清、拼接顺序不准确、混合输出与组合快捷指令触发依据不足的问题

Benefits of technology

本发明中,通过采集按键触发时间、按键输入序列、输入音频片段、待读取音频帧数变化序列以及音频读取触发时间,形成多模态信息集合,并进一步结合短时能量值、待读取音频帧数变化速率以及后一触发时间与按键触发截止时间的比较结果判断输入音频片段是否处于无效音频起始状态;在判定存在无效音频起始状态时,通过逆向回溯待读取音频帧数变化速率确定首次校正音频起始时间,计算音频起始基准可信时间范围,并基于调节后信号融合时间窗和用户意图判定置信度值确定协同关系判定结果或独立关系判定结果,从而实现对按键输入序列和输入音频片段的对应关系校正,起到了使多模态融合输入信号能够按照协同关联关系或时间先后关系生成,并支持文本字符、语音文本混合输出以及按键与语音组合触发键盘快捷指令的作用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654980A_ABST
    Figure CN122654980A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of signal processing, in particular to a kind of intelligent keyboard multimodal input signal fusion processing method and system.In the present application, by collecting key trigger time, key input sequence, input audio segment, the audio frame number change sequence to be read and audio reading trigger time, form multimodal information set, and combine short-time energy value, the audio frame number change rate to be read and key trigger cutoff time to judge invalid audio starting state;When there is invalid audio starting state, reverse backtracking is determined to first correction audio starting time, audio starting reference credible time range is calculated, and through adjusting post-signal fusion time window and user intention determination confidence value, the synergistic or independent relationship is determined, which plays a role in generating multimodal fusion input signal, supporting text character and speech text mixed output and triggering keyboard shortcut instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of signal processing technology, and in particular to a method and system for fusion processing of multimodal input signals from an intelligent keyboard. Background Technology

[0002] The multimodal input signal fusion processing method for smart keyboards refers to the technical process of extracting, combining, and analyzing various input signals generated by smart keyboards during use, including keystrokes, presses, touch controls, swipes, and voice input. It is mainly used to integrate these signals from different sources and in different forms received by the smart keyboard for centralized comparison and joint processing.

[0003] Traditional multimodal input signal fusion processing methods for smart keyboards can extract, combine, and centrally process different forms of input signals such as keystrokes, presses, touch swipes, and voice. However, when users use the keyboard voice conversion function, the starting position of the input audio segment may be affected by silence segments, the lag in changes of audio frames waiting to be read in the buffer, and the relationship between the continuous triggering time of keys. This makes it difficult to accurately correspond the audio start time with the key triggering time. Consequently, in determining whether the key input sequence and the input audio segment belong to the same user input operation, problems such as unclear coordination, inaccurate splicing order, and insufficient basis for triggering mixed output and combined shortcut commands may occur. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and to propose a method and system for multimodal input signal fusion processing of intelligent keyboards.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for multimodal input signal fusion processing of an intelligent keyboard, comprising the following steps: Collect key input data and audio segments of user voice input generated when the user operates the keyboard, and construct a multimodal information set of key input data and audio segments; Based on the multimodal information set, determine whether the input audio segment is in an invalid audio start state; When the input audio segment is in an invalid audio start state, the corresponding first correction audio start time is determined by referring to the multimodal information set, and the audio start reference reliable time range is calculated; Obtain a preset initial signal fusion time window, adjust the initial signal fusion time window based on the audio start reference confidence time range, and calculate different user intent judgment confidence values ​​by combining the multimodal information set and the first correction audio start time; The user intent determination confidence value is compared with the preset confidence threshold. Based on the comparison result, the secondary correction audio start time of the first correction audio start time is determined, and the multimodal information set is adjusted to generate a multimodal fusion input signal.

[0006] As a further aspect of the present invention, the step of collecting key input data generated when the user operates the keyboard and input audio segments of the user's voice, and constructing a multimodal information set of key input data and input audio segments includes: The system monitors key input data on the keyboard when the user presses a key. The key input data includes key trigger time and key input sequence. When the user uses the keyboard voice conversion function, the system collects audio segments of the user's voice input through the keyboard microphone. The input audio segment is passed to the hardware data buffer, the sequence of changes in the number of audio frames to be read is monitored during the passing process, and the audio reading trigger time of the input audio segment is read through the hardware data buffer. By integrating the key trigger time, the key input sequence, the input audio segment, the sequence of changes in the number of audio frames to be read, and the audio reading trigger time, a multimodal information set is obtained.

[0007] As a further aspect of the present invention, determining whether an input audio segment is in an invalid audio start state by referring to a multimodal information set includes: Obtain the initial speech signal and audio sampling time corresponding to the input audio segment in the multimodal information set, and calculate the short-time energy value of the initial speech signal according to the audio sampling time; The method involves obtaining the numerical value of each audio frame to be read and the corresponding audio frame sampling time in the audio frame number change sequence to be read from the multimodal information set, calculating the frame number difference between each audio frame number to be read and the adjacent audio frame number, calculating the time difference between each audio frame sampling time and the adjacent audio frame sampling time, and simultaneously calculating the ratio of the frame number difference to the time difference to obtain the rate of change of the audio frame number to be read. Obtain the trigger time of each key in the key input sequence in the multimodal information set and the next trigger time corresponding to each key trigger time. Calculate the sum of each key trigger time and the preset key trigger judgment bias duration to obtain the key trigger cutoff time corresponding to each key trigger time. Align the audio sampling time, the audio frame sampling time, and the trigger time of each button. If, within the preset time alignment range, the short-term energy value is less than the preset energy threshold, the rate of change of the audio frame to be read is less than the preset rate threshold, and the trigger time of each button corresponds to the next trigger time earlier than the button trigger end time, then the corresponding input audio segment is determined to be in an invalid audio start state.

[0008] As a further aspect of the present invention, when the input audio segment is in an invalid audio start state, determining the corresponding first correction audio start time by referring to the multimodal information set and calculating the audio start reference reliable time range includes: When the input audio segment is in the invalid audio start state, refer to the audio frame number change sequence to be read in the multimodal information set, trace back the audio frame number change rate to be read along the corresponding audio frame number sampling time of the audio frame number change sequence to be read, and select the audio frame number sampling time corresponding to when the audio frame number change rate to be read changes from less than the rate threshold to greater than or equal to the rate threshold as the first correction audio start time; Calculate the time deviation between the audio read trigger time and the first corrected audio start time in the multimodal information set; Using the initial audio correction start time as the starting point and the audio reading trigger time as the ending point, and using the time deviation value to limit the time interval between the starting point and the ending point, the reliable time range of the audio start reference is obtained.

[0009] As a further aspect of the present invention, the step of obtaining a preset initial signal fusion time window, adjusting the initial signal fusion time window based on the audio start reference confidence time range, and calculating different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time includes: Obtain a preset initial signal fusion time window, obtain the range start time and range end time of the audio start reference reliable time range, calculate the average time of the range start time and range end time to obtain the reliable range center time, calculate the difference between the reliable range center time and the first correction audio start time to obtain the time boundary adjustment value of the initial signal fusion time window, and use the time boundary adjustment value to shift the initial signal fusion time window to obtain the adjusted signal fusion time window. The first calibration audio start time is added to the preset coordination time determination offset duration to obtain the coordination determination cutoff time; The button trigger time corresponding to the fusion time window of the adjusted signal within the multimodal information set is obtained. When the button trigger time is later than or equal to the first calibration audio start time and earlier than or equal to the collaborative determination deadline, it is determined that there is a collaborative relationship between the button trigger time and the first calibration audio start time, and a collaborative relationship determination result is obtained. The difference between the collaborative determination deadline and the button trigger time is calculated to obtain the margin time. The ratio of the margin time to the collaborative time determination bias duration is calculated as the user intent determination confidence value corresponding to the collaborative relationship determination result. When the button trigger time is earlier than the first calibration audio start time or later than the collaborative determination deadline, it is determined that there is no collaborative relationship between the button trigger time and the first calibration audio start time, and an independent relationship determination result is obtained. The value of zero is taken as the user intent determination confidence value corresponding to the independent relationship determination result.

[0010] As a further aspect of the present invention, the step of comparing the user intent determination confidence value with a preset confidence threshold, determining the secondary correction audio start time based on the comparison result, adjusting the multimodal information set, and generating a multimodal fusion input signal includes: The system continuously monitors the number of audio frames to be read in the sequence of changes in the number of audio frames to be read in the multimodal information set. When the confidence value of the user intent determination is less than the preset confidence threshold and the number of audio frames to be read remains constant, additional audio sampling segments are collected through the keyboard and microphone to obtain the additional sampling time corresponding to the additional audio sampling segments. The initial audio correction start time is updated using the additional sampling time to obtain the secondary audio correction start time. The secondary audio correction start time is replaced with the initial audio correction start time, and the key trigger time is re-compared with the secondary audio correction start time and the collaborative judgment deadline. When the key trigger time is later than or the same as the secondary audio correction start time and earlier than or the same as the collaborative judgment deadline, it is determined that the key trigger time falls within the collaborative judgment time range defined by the secondary audio correction start time to the collaborative judgment deadline. The key input sequence and the input audio segment are determined to belong to the same user input operation, and the updated collaborative relationship judgment result is obtained. When the key trigger time is earlier than the secondary audio correction start time or later than the collaborative judgment deadline, it is determined that the key trigger time does not fall within the collaborative judgment time range defined by the secondary audio correction start time to the collaborative judgment deadline. The key input sequence and the input audio segment are determined to not belong to the same user input operation, and the updated independent relationship judgment result is obtained. When the updated collaborative relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are concatenated according to the collaborative correlation between key trigger time and secondary correction audio start time. When the updated independent relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are concatenated according to the temporal order of key trigger time and secondary correction audio start time to generate a multimodal fusion input signal. When the confidence value of the user intent determination is greater than or equal to the confidence threshold, the input audio segment and key input sequence in the multimodal information set are directly concatenated according to the collaborative relationship determination result or the independent relationship determination result to generate a multimodal fusion input signal. This signal is used to control the keyboard to output the text characters corresponding to the key input sequence and convert the input audio segment into speech text for mixed output, or to execute keyboard shortcut commands triggered by key and speech combination.

[0011] A smart keyboard multimodal input signal fusion processing system, the system comprising: The data acquisition module collects key input data and audio segments of user voice input generated when the user operates the keyboard, and constructs a multimodal information set of key input data and audio segments. The state determination module refers to the multimodal information set to determine whether the input audio segment is in an invalid audio start state; The time correction module determines the corresponding first correction audio start time by referring to the multimodal information set when the input audio segment is in an invalid audio start state, and calculates the audio start reference reliable time range. The confidence calculation module obtains a preset initial signal fusion time window, adjusts the initial signal fusion time window based on the audio start reference confidence time range, and calculates different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time. The signal generation module compares the user intent determination confidence value with the preset confidence threshold, determines the secondary correction audio start time based on the comparison result, adjusts the multimodal information set, and generates a multimodal fusion input signal.

[0012] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a multimodal information set is formed by collecting key trigger time, key input sequence, input audio segment, audio frame number change sequence to be read, and audio reading trigger time. Furthermore, it combines short-time energy value, audio frame number change rate to be read, and comparison results between the subsequent trigger time and key trigger cutoff time to determine whether the input audio segment is in an invalid audio start state. If an invalid audio start state is determined, the first corrected audio start time is determined by retrospectively tracing the audio frame number change rate to be read, calculating the audio start reference reliable time range, and determining the cooperative relationship determination result or independent relationship determination result based on the adjusted signal fusion time window and user intent determination confidence value. This achieves the correction of the correspondence between the key input sequence and the input audio segment, enabling the multimodal fusion input signal to be generated according to cooperative or temporal relationships, and supporting mixed output of text characters and speech text, as well as key and speech combination triggering of keyboard shortcuts. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the steps of the method of the present invention; Figure 2 This is a flowchart illustrating the steps and principles of the method of the present invention. Figure 3 This is a block diagram of the system of the present invention; Figure 4 This is a block diagram illustrating the module flow principle of the system of the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0015] Please see Figure 1-2 This invention provides a technical solution: a method for multimodal input signal fusion processing of an intelligent keyboard, comprising the following steps: S1: Collect key input data and audio segments of user voice input generated when the user operates the keyboard, and construct a multimodal information set of key input data and audio segments, including: The system monitors key input data on the keyboard when the user presses a key. The key input data includes key trigger time and key input sequence. When the user uses the keyboard voice conversion function, the system collects audio segments of the user's voice input through the keyboard microphone. The keyboard firmware is compiled and written to the smart keyboard using the embedded development software Keil MDK. This firmware controls key press acquisition, keyboard microphone acquisition, hardware data buffer read / write, and time recording. The smart keyboard includes key presses, a keyboard microphone, a hardware data buffer, and internal timing capabilities. When a user presses a key, the keyboard reads the key identifier and the internal keyboard timer for the key transitioning from an untriggered to a triggered state. The key identifier is the recognition record of the character key, function key, or voice conversion key corresponding to the pressed key. The key trigger time is the internal keyboard timer corresponding to the key transition state. Key identifiers are arranged from earliest to latest according to their trigger times to form a key input sequence, which is a record of key identifiers arranged chronologically. When the user uses the keyboard voice conversion function, the keyboard microphone reads the audio sample values ​​of the user's voice according to the continuous audio sampling time. The input audio segment is a voice record formed by arranging the continuous audio sample values ​​according to the audio sampling time. The key trigger time corresponds to the key input sequence, and the input audio segment corresponds to the audio sampling time, resulting in key input data and input audio segments.

[0016] The input audio segment is passed to the hardware data buffer, the sequence of changes in the number of audio frames to be read is monitored during the passing process, and the audio reading trigger time of the input audio segment is read through the hardware data buffer. The hardware data buffer temporarily stores the audio frames corresponding to the input audio segments. An audio frame is a reading unit formed by the sampling order of the input audio segment. The number of audio frames to be read is the number of audio frames in the hardware data buffer that have not yet been read. The audio frame sampling time is the internal keyboard timer corresponding to the time when the number of audio frames to be read is read. The audio read trigger time is the internal keyboard timer corresponding to the time when the hardware data buffer starts reading the input audio segment. The audio frames corresponding to the input audio segment are written to the hardware data buffer in the order they are entered. After each audio frame corresponding to an input audio segment is written, the number of audio frames in the hardware data buffer that have not yet been read is read. Each audio frame number to be read is bound to the corresponding audio frame sampling time and arranged from earliest to latest according to the audio frame sampling time. When the hardware data buffer starts reading the input audio segment, the corresponding internal keyboard timer is recorded, resulting in the sequence of changes in the number of audio frames to be read and the audio read trigger time.

[0017] By integrating key trigger time, key input sequence, input audio segment, audio frame number change sequence to be read, and audio reading trigger time, a multimodal information set is obtained; Key trigger time, key input sequence, input audio segment, audio frame number change sequence to be read, and audio reading trigger time are all aggregated under the same time base, which is the keyboard's internal timing. Key identifiers in the key input sequence are indexed to key trigger times; audio sample values ​​in the input audio segment are indexed to audio sample times; and the number of audio frames to be read in the audio frame number change sequence is indexed to the audio frame number sampling time. The audio reading trigger time is written to the reading record corresponding to the input audio segment. Key trigger time, audio sample time, audio frame number sampling time, and audio reading trigger time are arranged in chronological order to obtain a multimodal information set.

[0018] S2: Referring to the multimodal information set, determine whether the input audio segment is in an invalid audio start state, including: Obtain the initial speech signal and audio sampling time corresponding to the input audio segment in the multimodal information set, and calculate the short-time energy value of the initial speech signal according to the audio sampling time; The initial speech signal consists of continuous audio samples at the beginning of the input audio segment. The short-time energy value corresponds to the cumulative average energy of the initial speech signal within a short time division. The starting position of the input audio segment is determined by the first audio sampling time of the input audio segment in the multimodal information set. Continuous audio samples and their corresponding sampling times are read from the starting position of the input audio segment. Short-time divisions are performed according to the audio sampling time, with the interval between adjacent sampling times in the continuous audio sampling time serving as the division basis. The same number of continuous audio samples are taken in each short-time division. Each audio sample value within each short-time division is squared. The short-time energy value = the sum of the squares of the audio samples of the initial speech signal within the segment / the number of audio samples of the initial speech signal within the segment, yielding the short-time energy value of the initial speech signal.

[0019] The algorithm obtains the audio frame number value and corresponding audio frame sampling time of each audio frame number change sequence to be read in the multimodal information set, calculates the frame number difference between each audio frame number value and the adjacent audio frame number values, calculates the time difference between each audio frame sampling time and the adjacent audio frame sampling time, and calculates the ratio of the frame number difference to the time difference to obtain the rate of change of the audio frame number to be read. The frame difference is the difference between each audio frame number to be read and its adjacent frame number. The time difference is the time interval between the sampling time of each audio frame and the sampling time of its adjacent frames. The rate of change of the audio frame number corresponds to the change in the frame difference relative to the time difference. The audio frame numbers to be read are read through a hardware data buffer during the input audio segment. The audio frame sampling time is recorded by an internal keyboard timer. Each audio frame number to be read is read from earliest to latest according to its sampling time. Each audio frame number to be read is processed in pairs with its adjacent frames, and each audio frame sampling time is processed in pairs with its adjacent sampling time. Frame difference = Each audio frame number to be read - Adjacent audio frame number. Time difference = Each audio frame sampling time - Adjacent audio frame sampling time. Rate of change of the audio frame number to be read = Frame difference / Time difference.

[0020] Obtain the trigger time of each key in the key input sequence in the multimodal information set and the next trigger time corresponding to each key trigger time. Calculate the sum of each key trigger time and the preset key trigger judgment bias duration to obtain the key trigger cutoff time corresponding to each key trigger time. The next trigger time is the adjacent key trigger time after the current key trigger time. A preset key trigger determination offset duration is used to determine the key trigger cutoff time, which is the time boundary extending the current key trigger time by the preset key trigger determination offset duration. Each key trigger time in the key input sequence is obtained by acquiring user key presses through a key detection structure and read in ascending order. The adjacent key trigger time after each key trigger time is recorded as the next trigger time corresponding to that key trigger time. The preset key trigger determination offset duration is set before parameter writing. During setting, adjacent key trigger times classified as consecutive key triggers are read, the time interval between each group of adjacent key trigger times is calculated, all time intervals are arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset key trigger determination offset duration. Key trigger cutoff time = each key trigger time + preset key trigger determination offset duration, yielding the key trigger cutoff time corresponding to each key trigger time.

[0021] Align the audio sampling time, audio frame sampling time with each button trigger time. If, within the preset time alignment range, a short-term energy value is less than the preset energy threshold, the rate of change of the audio frame to be read is less than the preset rate threshold, and the next trigger time corresponding to each button trigger time is earlier than the button trigger end time, then the corresponding input audio segment is determined to be in an invalid audio start state. A preset energy threshold is used for comparison with short-time energy values, a preset rate threshold is used for comparison with the rate of change of the number of audio frames to be read, and a preset time alignment range is used to limit the time boundary for matching the audio sampling time, audio frame sampling time, and each key trigger time within the same time range. An invalid audio start state indicates that the starting position of the input audio segment meets the corresponding judgment conditions in terms of energy, the change of the audio frames to be read in the hardware data buffer, and the continuous key triggering relationship. The audio sampling time comes from the sampling time record when the keyboard microphone collects the input audio segment, the audio frame sampling time comes from the time record when the hardware data buffer reads the value of the number of audio frames to be read, and each key trigger time comes from the time record when the key detection structure reads the user's key press action. The audio sampling time, audio frame sampling time, and each key trigger time are placed under the same time reference. The preset energy threshold is set before the parameters are written. During setting, the input audio segment in the invalid audio start state is collected by the keyboard microphone, and the corresponding short-time energy value is obtained according to the short-time energy value calculation method. The short-time energy values ​​in the invalid audio start state are arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset energy threshold. The preset rate threshold is set before the parameters are written. During setting, the sequence of audio frame changes in the invalid audio start state is read. The corresponding audio frame change rate is obtained according to the calculation method for the audio frame change rate. The audio frame change rates in the invalid audio start state are arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset rate threshold. The preset time alignment range is also set before the parameters are written. During setting, the time interval between the audio sampling time and the key trigger time, and the time interval between the audio frame sampling time and the key trigger time, which are included in the same input operation, are read. These two types of time intervals are merged and arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset time alignment range. Within the preset time alignment range, if the short-term energy value is less than the preset energy threshold, the rate of change of the number of audio frames to be read is less than the preset rate threshold, and the subsequent trigger time corresponding to each key trigger time is earlier than the key trigger end time, it is considered as the current invalid audio start state determination result; if any of the following occurs: the short-term energy value is greater than or equal to the preset energy threshold, the rate of change of the number of audio frames to be read is greater than or equal to the preset rate threshold, and the subsequent trigger time corresponding to each key trigger time is later than or equal to the key trigger end time, it is considered as the current non-invalid audio start state determination result, thus obtaining the determination result of whether the input audio segment is in an invalid audio start state.

[0022] S3: When the input audio segment is in an invalid audio start state, the corresponding first correction audio start time is determined by referring to the multimodal information set, and the audio start reference reliable time range is calculated, including: When the input audio segment is in an invalid audio start state, refer to the audio frame number change sequence to be read in the multimodal information set, trace back the audio frame number change rate to be read along the corresponding audio frame number sampling time of the audio frame number change sequence to be read, and select the audio frame number sampling time corresponding to when the audio frame number change rate to be read changes from less than the rate threshold to greater than or equal to the rate threshold as the first correction audio start time; The rate threshold is used to identify the state where the rate of change of the audio frame count to be read changes from below the rate threshold to above the rate threshold. The initial audio correction start time is derived from the audio frame sampling time selected when tracing back the rate of change of the audio frame count to be read. The sequence of audio frame count changes to be read in the multimodal information set comes from the continuous reading of the audio frame count values ​​to be read from the hardware data buffer. The sequence of audio frame count changes to be read is used to locate the audio frame sampling time corresponding to the invalid audio start state. The rate of change of the audio frame count to be read is read point by point from late to early along the corresponding audio frame sampling time of the sequence of audio frame count changes to be read, and the rate of change of the audio frame count to be read is compared with the rate threshold. When the rate of change of the audio frame count to be read is less than the rate threshold, the rate of change of the audio frame count to be read corresponding to an earlier audio frame sampling time is read; when the rate of change of the audio frame count to be read changes from less than the rate threshold to greater than or equal to the rate threshold, reading to earlier audio frame sampling times is stopped, and the initial audio correction start time is obtained.

[0023] Calculate the time deviation between the audio read trigger time and the first corrected audio start time in the multimodal information set; The time deviation value corresponds to the time difference between the audio read trigger time and the first calibration audio start time. The audio read trigger time is obtained from the internal keyboard timing when the hardware data buffer begins reading the input audio segment. The first calibration audio start time is derived from the reverse backtracking result of the rate of change of the number of audio frames to be read. The audio read trigger time in the multimodal information set and the previously obtained first calibration audio start time are placed under the same time base. Time deviation value = audio read trigger time - first calibration audio start time, thus obtaining the time deviation value.

[0024] Starting from the first audio calibration start time and ending from the audio reading trigger time, and using the time deviation value to limit the time interval between the start and end, the reliable time range of the audio start reference is obtained. The reliable time range for the audio starting reference is defined by the first audio calibration start time and the audio read trigger time. The start time is the first audio calibration start time, and the end time is the audio read trigger time. The first audio calibration start time is derived from the reverse backtracking result of the rate of change of the number of audio frames to be read. The audio read trigger time is derived from the read time record of the hardware data buffer. The time deviation value is derived from the time difference between the audio read trigger time and the first audio calibration start time. The first audio calibration start time is used as the start point, the audio read trigger time as the end point, and the time deviation value as the time interval limit between the start and end points. The continuous time period from the first audio calibration start time to the audio read trigger time is used as the reliable time range for the audio starting reference. Therefore, the reliable time range for the audio starting reference is calculated as: Reliable time range for the audio starting reference = Reliable time range for the audio starting reference = Reliable time range for the audio starting reference

[0025] S4: Obtain the preset initial signal fusion time window, adjust the initial signal fusion time window based on the audio start reference confidence time range, and calculate different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time, including: The system obtains a preset initial signal fusion time window, acquires the start and end times of the audio start reference reliable time range, calculates the average time between the start and end times of the range to obtain the center time of the reliable range, calculates the difference between the center time of the reliable range and the start time of the first calibration audio to obtain the time boundary adjustment value of the initial signal fusion time window, and shifts the initial signal fusion time window using the time boundary adjustment value to obtain the adjusted signal fusion time window. The initial signal fusion time window is set according to the longest time interval between keyboard input and voice input that can be considered as the same input operation before and after the user presses the voice conversion key. The initial signal fusion time window is determined by the longest time interval between the keyboard input and voice input considered to belong to the same input operation before and after the user presses the voice conversion key. The center time of the confidence range corresponds to the midpoint time of the audio start reference confidence time range. The time boundary adjustment value corresponds to the offset of the center time of the confidence range relative to the first calibration audio start time. The adjusted signal fusion time window is derived from the translation of the initial signal fusion time window according to the time boundary adjustment value. The preset initial signal fusion time window is set before the parameters are written. During setting, the time interval between the keyboard input time and the voice input time that are classified as the same input operation before and after the user presses the voice conversion key is read. The keyboard input time comes from the time record of the key detection structure, and the voice input time comes from the audio sampling time record when the keyboard microphone collects the input audio segment. The time intervals are arranged from smallest to largest, and the value at the end of the arrangement is selected as the length of the initial signal fusion time window. The start time and end time of the audio start reference confidence time range are used in the center position calculation, and the first calibration audio start time is used in the offset calculation. Confidential range center time = (range start time + range end time) / 2. Time boundary adjustment value = Confidential range center time - first calibration audio start time. The adjusted signal fusion time window start time = the initial signal fusion time window start time + time boundary adjustment value. The adjusted signal fusion time window end time = the initial signal fusion time window end time + time boundary adjustment value. A time boundary adjustment value greater than zero indicates a backward shift, a time boundary adjustment value equal to zero indicates no shift, and a time boundary adjustment value less than zero indicates a forward shift, thus obtaining the adjusted signal fusion time window.

[0026] The first calibration audio start time is added to the preset coordination time determination offset duration to obtain the coordination determination cutoff time; the coordination time determination offset duration is determined based on the time interval between the user's key trigger time when using the keyboard voice conversion function and the first calibration audio start time, which is allowed to have a coordination association. The collaborative timing determination offset duration corresponds to the time interval between the key trigger time and the initial audio calibration start time, allowing for collaborative association. The collaborative timing determination cutoff time is derived from the time boundary after extending the collaborative timing determination offset duration from the initial audio calibration start time. The preset collaborative timing determination offset duration is set before parameter writing. During setting, the time interval between the key trigger time (which is considered collaboratively associated) and the initial audio calibration start time when the user uses the keyboard voice conversion function is read. The key trigger time is obtained by collecting the user's key presses through a key detection structure. The initial audio calibration start time is obtained by backtracking the rate of change of the number of audio frames to be read. The time intervals are arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset collaborative timing determination offset duration. The initial audio calibration start time and the preset collaborative timing determination offset duration are used in the time boundary calculation. The collaborative timing determination cutoff time = initial audio calibration start time + preset collaborative timing determination offset duration, yielding the collaborative timing determination cutoff time.

[0027] The button trigger time corresponding to the fusion time window of the adjusted signal within the multimodal information set is obtained. When the button trigger time is later than or equal to the first calibration audio start time and earlier than or equal to the collaborative judgment deadline, it is determined that there is a collaborative relationship between the button trigger time and the first calibration audio start time, and the collaborative relationship judgment result is obtained. The difference between the collaborative judgment deadline and the button trigger time is calculated to obtain the margin time. The ratio of the margin time to the collaborative time judgment bias duration is calculated as the user intent judgment confidence value corresponding to the collaborative relationship judgment result. When the button trigger time is earlier than the first calibration audio start time or later than the collaborative judgment deadline, it is determined that there is no collaborative relationship between the button trigger time and the first calibration audio start time, and the independent relationship judgment result is obtained. The value of zero is taken as the user intent judgment confidence value corresponding to the independent relationship judgment result. The cooperative relationship determination result corresponds to the relationship where the key trigger time and the first calibration audio start time are within the cooperative determination time range. The independent relationship determination result corresponds to the relationship where the key trigger time and the first calibration audio start time are not within the cooperative determination time range. The margin time corresponds to the time difference between the cooperative determination cutoff time and the key trigger time. The user intent determination confidence value is calculated from the margin time relative to the cooperative time determination bias duration. The key trigger time in the multimodal information set comes from the time record of the user's key action collected by the key detection structure. The adjusted signal fusion time window comes from the translation of the initial signal fusion time window according to the time boundary adjustment value. The key trigger time is compared with the window start time and the window end time of the adjusted signal fusion time window. When the key trigger time is later than or the same as the window start time of the adjusted signal fusion time window and earlier than or the same as the window end time of the adjusted signal fusion time window, the key trigger time enters the cooperative association comparison. The key trigger time entering the cooperative association comparison is compared with the first calibration audio start time and the cooperative determination cutoff time. When the button trigger time is later than or equal to the first audio calibration start time and earlier than or equal to the collaborative determination deadline, it is taken as the current collaborative relationship determination result. The margin time = collaborative determination deadline - button trigger time, and the user intent determination confidence value = margin time / collaborative time determination bias duration. When the button trigger time is earlier than the first audio calibration start time or later than the collaborative determination deadline, it is taken as the current independent relationship determination result, and the user intent determination confidence value = 0. The user intent determination confidence value is obtained.

[0028] S5: Compare the user intent determination confidence value with a preset confidence threshold, determine the secondary correction audio start time based on the comparison result, adjust the multimodal information set, and generate a multimodal fusion input signal including: The system continuously monitors the audio frame count values ​​in the audio frame count change sequence within the multimodal information set. When the user intent determination confidence value is less than a preset confidence threshold and the audio frame count remains constant, additional audio sampling segments are acquired via the keyboard and microphone to obtain the additional sampling time corresponding to the additional audio sampling segments. The determination process for the audio frame count to remain constant is as follows: within a preset constant monitoring time range, the audio frame count values ​​corresponding to multiple consecutive audio frame sampling times are read. If the read audio frame count values ​​are all the same, then the audio frame count in the audio frame count change sequence is determined to remain constant. The constant monitoring time range is used to continuously read the number of audio frames to be read and determine whether the number of audio frames remains constant. The preset confidence threshold is used to compare with the user intent judgment confidence value. The appended audio sampling segment is the audio sampling content collected by the keyboard and microphone after the input audio segment. The appended sampling time is the audio sampling time corresponding to the acquisition of the appended audio sampling segment. The preset constant monitoring time range is set before the parameters are written. When setting, the sampling time span of consecutive audio frames with the same number of audio frames to be read in the sequence of changes in the number of audio frames to be read is read. The number of audio frames to be read comes from the number of audio frames that have not yet been read in the hardware data buffer. The sampling time span of consecutive audio frames is arranged from smallest to largest, and the value at the end of the arrangement is selected as the preset constant monitoring time range. The preset confidence threshold is set before the parameters are written. When setting, the user intent judgment confidence value of the key input sequence and input audio segment belonging to the same user input operation is read. The user intent judgment confidence value is arranged from smallest to largest, and the value at the beginning of the arrangement is selected as the preset confidence threshold. Within a preset constant monitoring time range, the audio frame number values ​​are read from the sequence of audio frame number changes according to multiple consecutive audio frame number sampling times. If all read audio frame number values ​​are the same, it is considered a result indicating that the current audio frame number remains constant; if the read audio frame number values ​​differ, it is considered a result indicating that the current audio frame number does not remain constant. When the user intent determination confidence value is less than a preset confidence threshold and the audio frame number remains constant, the keyboard and microphone continue to collect audio samples after the input audio segment and record the corresponding additional sampling time to obtain the additional sampling time corresponding to the additional audio sampling segment.

[0029] The initial calibration audio start time is updated using the additional sampling time to obtain the secondary calibration audio start time. The secondary calibration audio start time is then replaced with the initial calibration audio start time. The key trigger time is then re-compared with the secondary calibration audio start time and the collaborative decision cutoff time. If the key trigger time is later than or equal to the secondary calibration audio start time and earlier than or equal to the collaborative decision cutoff time, it is determined that the key trigger time falls within the collaborative decision time range defined by the secondary calibration audio start time to the collaborative decision cutoff time. The key input sequence and the input audio segment are then determined to belong to the same user input operation, resulting in an updated collaborative relationship determination result. If the key trigger time is earlier than the secondary calibration audio start time or later than the collaborative decision cutoff time, it is determined that the key trigger time does not fall within the collaborative decision time range defined by the secondary calibration audio start time to the collaborative decision cutoff time. The key input sequence and the input audio segment are then determined to not belong to the same user input operation, resulting in an updated independent relationship determination result. The secondary calibration audio start time is derived from the update of the initial calibration audio start time using the appended sampling time. The collaborative judgment time range is limited by the secondary calibration audio start time to the collaborative judgment deadline. The updated collaborative relationship judgment result indicates that the key input sequence and the input audio segment belong to the same user input operation, while the updated independent relationship judgment result indicates that the key input sequence and the input audio segment do not belong to the same user input operation. The appended audio sampling segment is obtained by continuing to collect audio samples after the input audio segment using the keyboard microphone. The appended sampling time is obtained from the audio sampling time corresponding to the appended audio sampling segment. The appended sampling time corresponding to the appended audio sampling segment is written to the update position of the initial calibration audio start time, and the secondary calibration audio start time = appended sampling time. After the secondary calibration audio start time replaces the initial calibration audio start time, the key trigger time is re-compared with the secondary calibration audio start time and the collaborative judgment deadline time in terms of time boundaries. If the key trigger time is later than or equal to the start time of the secondary calibration audio but earlier than or equal to the collaborative determination deadline, the key trigger time is determined to fall within the collaborative determination time range defined by the start time of the secondary calibration audio to the collaborative determination deadline. The key input sequence and the input audio segment are then determined to belong to the same user input operation, and the updated collaborative relationship determination result is obtained. If the key trigger time is earlier than the start time of the secondary calibration audio or later than the collaborative determination deadline, the key trigger time is determined to not fall within the collaborative determination time range defined by the start time of the secondary calibration audio to the collaborative determination deadline. The key input sequence and the input audio segment are then determined to not belong to the same user input operation, and the updated independent relationship determination result is obtained.

[0030] When the updated collaborative relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are spliced ​​together according to the collaborative relationship between the key trigger time and the start time of the secondary correction audio. When the updated independent relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are spliced ​​together according to the temporal relationship between the key trigger time and the start time of the secondary correction audio to generate a multimodal fusion input signal. The multimodal fusion input signal is formed by concatenating input audio segments and key input sequences according to a cooperative relationship or temporal sequence. The input audio segments originate from the user's voice captured by the keyboard microphone, and the key input sequences originate from the user's key presses captured by the key detection structure. When the updated cooperative relationship determination result corresponds, the cooperative relationship between the key trigger time and the secondary correction audio start time is read, and the input audio segments and key input sequences are concatenated according to this cooperative relationship, using the key trigger time and the secondary correction audio start time as concatenation indices. When the updated independence relationship determination result corresponds, the key trigger time and the secondary correction audio start time are compared sequentially; if the key trigger time is earlier than the secondary correction audio start time, the key input sequence is concatenated before the input audio segments; if the key trigger time is later than the secondary correction audio start time, the input audio segments are concatenated before the key input sequences; if the key trigger time is the same as the secondary correction audio start time, they are concatenated according to the temporal order in the multimodal information set to generate the multimodal fusion input signal.

[0031] When the confidence value of the user intent judgment is greater than or equal to the confidence threshold, the input audio segment and key input sequence in the multimodal information set are directly spliced ​​together according to the collaborative relationship judgment result or the independent relationship judgment result to generate a multimodal fusion input signal. This signal is used to control the keyboard to output the text characters corresponding to the key input sequence and convert the input audio segment into speech text for mixed output, or to execute keyboard shortcut commands triggered by key and speech combination. Text characters are the character content corresponding to the key input sequence; speech text is the text content obtained by converting the input audio segment into text using the keyboard speech conversion function; keyboard shortcuts triggered by key and speech combinations are the keyboard instruction content corresponding to both the key input sequence and the input audio segment. The user intent determination confidence value is compared with a confidence threshold, which uses the same confidence comparison benchmark as a preset confidence threshold. When the user intent determination confidence value is greater than or equal to the confidence threshold, the input audio segment and key input sequence in the multimodal information set are concatenated based on the collaboration relationship determination result or the independence relationship determination result. When the user intent determination confidence value is less than the confidence threshold, the input audio segment and key input sequence in the multimodal information set are concatenated based on the additional sampling time corresponding to the additional audio sampling segment, the secondary correction audio start time, and the updated collaboration relationship determination result or the updated independence relationship determination result. The key identifiers in the key input sequence are obtained from the key detection structure's acquisition of the user's key press actions and arranged according to the key trigger time to form text characters; the input audio segments are obtained from the keyboard microphone's acquisition of the user's voice and converted into speech text through the keyboard speech conversion function; when the keyboard shortcut command is triggered by a combination of key presses and voice, the key input sequence and the input audio segments form the keyboard shortcut command content according to the combination trigger relationship, generating a multimodal fusion input signal.

[0032] Please see Figure 3-4 A smart keyboard multimodal input signal fusion processing system, the system comprising: The data acquisition module collects key input data and audio segments of user voice input generated when the user operates the keyboard, and constructs a multimodal information set of key input data and audio segments. The state determination module refers to the multimodal information set to determine whether the input audio segment is in an invalid audio start state; The time correction module determines the corresponding first correction audio start time by referring to the multimodal information set when the input audio segment is in an invalid audio start state, and calculates the audio start reference reliable time range. The confidence calculation module obtains a preset initial signal fusion time window, adjusts the initial signal fusion time window based on the audio start reference confidence time range, and calculates different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time. The signal generation module compares the user intent determination confidence value with the preset confidence threshold, determines the secondary correction audio start time based on the comparison result, adjusts the multimodal information set, and generates a multimodal fusion input signal.

[0033] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for fusion processing of multimodal input signals from an intelligent keyboard, characterized in that, Includes the following steps: Collect key input data and audio segments of user voice input generated when the user operates the keyboard, and construct a multimodal information set of key input data and audio segments; Based on the multimodal information set, determine whether the input audio segment is in an invalid audio start state; When the input audio segment is in an invalid audio start state, the corresponding first correction audio start time is determined by referring to the multimodal information set, and the audio start reference reliable time range is calculated; Obtain a preset initial signal fusion time window, adjust the initial signal fusion time window based on the audio start reference confidence time range, and calculate different user intent judgment confidence values ​​by combining the multimodal information set and the first correction audio start time; The user intent determination confidence value is compared with the preset confidence threshold. Based on the comparison result, the secondary correction audio start time of the first correction audio start time is determined, and the multimodal information set is adjusted to generate a multimodal fusion input signal.

2. The intelligent keyboard multimodal input signal fusion processing method according to claim 1, characterized in that, The process of collecting key input data and audio segments of user speech generated when the user operates the keyboard, and constructing a multimodal information set of key input data and audio segments includes: The system monitors key input data on the keyboard when the user presses a key. The key input data includes key trigger time and key input sequence. When the user uses the keyboard voice conversion function, the system collects audio segments of the user's voice input through the keyboard microphone. The input audio segment is passed to the hardware data buffer, the sequence of changes in the number of audio frames to be read is monitored during the passing process, and the audio reading trigger time of the input audio segment is read through the hardware data buffer. By integrating the key trigger time, the key input sequence, the input audio segment, the sequence of changes in the number of audio frames to be read, and the audio reading trigger time, a multimodal information set is obtained.

3. The intelligent keyboard multimodal input signal fusion processing method according to claim 2, characterized in that, The step of determining whether an input audio segment is in an invalid audio start state by referring to the multimodal information set includes: Obtain the initial speech signal and audio sampling time corresponding to the input audio segment in the multimodal information set, and calculate the short-time energy value of the initial speech signal according to the audio sampling time; The method involves obtaining the numerical value of each audio frame to be read and the corresponding audio frame sampling time in the audio frame number change sequence to be read from the multimodal information set, calculating the frame number difference between each audio frame number to be read and the adjacent audio frame number, calculating the time difference between each audio frame sampling time and the adjacent audio frame sampling time, and simultaneously calculating the ratio of the frame number difference to the time difference to obtain the rate of change of the audio frame number to be read. Obtain the trigger time of each key in the key input sequence in the multimodal information set and the next trigger time corresponding to each key trigger time. Calculate the sum of each key trigger time and the preset key trigger judgment bias duration to obtain the key trigger cutoff time corresponding to each key trigger time. Align the audio sampling time, the audio frame sampling time, and the trigger time of each button. If, within the preset time alignment range, the short-term energy value is less than the preset energy threshold, the rate of change of the audio frame to be read is less than the preset rate threshold, and the trigger time of each button corresponds to the next trigger time earlier than the button trigger end time, then the corresponding input audio segment is determined to be in an invalid audio start state.

4. The intelligent keyboard multimodal input signal fusion processing method according to claim 3, characterized in that, When the input audio segment is in an invalid audio start state, the process of determining the corresponding first correction audio start time by referring to the multimodal information set and calculating the audio start reference reliable time range includes: When the input audio segment is in the invalid audio start state, refer to the audio frame number change sequence to be read in the multimodal information set, trace back the audio frame number change rate to be read along the corresponding audio frame number sampling time of the audio frame number change sequence to be read, and select the audio frame number sampling time corresponding to when the audio frame number change rate to be read changes from less than the rate threshold to greater than or equal to the rate threshold as the first correction audio start time; Calculate the time deviation between the audio read trigger time and the first corrected audio start time in the multimodal information set; Using the initial audio correction start time as the starting point and the audio reading trigger time as the ending point, and using the time deviation value to limit the time interval between the starting point and the ending point, the reliable time range of the audio start reference is obtained.

5. The intelligent keyboard multimodal input signal fusion processing method according to claim 4, characterized in that, The process of obtaining a preset initial signal fusion time window, adjusting the initial signal fusion time window based on the audio start reference confidence time range, and calculating different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time includes: Obtain a preset initial signal fusion time window, obtain the range start time and range end time of the audio start reference reliable time range, calculate the average time of the range start time and range end time to obtain the reliable range center time, calculate the difference between the reliable range center time and the first correction audio start time to obtain the time boundary adjustment value of the initial signal fusion time window, and use the time boundary adjustment value to shift the initial signal fusion time window to obtain the adjusted signal fusion time window. The first calibration audio start time is added to the preset coordination time determination offset duration to obtain the coordination determination cutoff time; The button trigger time corresponding to the fusion time window of the adjusted signal within the multimodal information set is obtained. When the button trigger time is later than or equal to the first calibration audio start time and earlier than or equal to the collaborative determination deadline, it is determined that there is a collaborative relationship between the button trigger time and the first calibration audio start time, and a collaborative relationship determination result is obtained. The difference between the collaborative determination deadline and the button trigger time is calculated to obtain the margin time. The ratio of the margin time to the collaborative time determination bias duration is calculated as the user intent determination confidence value corresponding to the collaborative relationship determination result. When the button trigger time is earlier than the first calibration audio start time or later than the collaborative determination deadline, it is determined that there is no collaborative relationship between the button trigger time and the first calibration audio start time, and an independent relationship determination result is obtained. The value of zero is taken as the user intent determination confidence value corresponding to the independent relationship determination result.

6. The method for multimodal input signal fusion processing of an intelligent keyboard according to claim 5, characterized in that, The step of comparing the user intent determination confidence value with a preset confidence threshold, determining the secondary correction audio start time based on the comparison result, adjusting the multimodal information set, and generating a multimodal fusion input signal includes: The system continuously monitors the number of audio frames to be read in the sequence of changes in the number of audio frames to be read in the multimodal information set. When the confidence value of the user intent determination is less than the preset confidence threshold and the number of audio frames to be read remains constant, additional audio sampling segments are collected through the keyboard and microphone to obtain the additional sampling time corresponding to the additional audio sampling segments. The initial audio correction start time is updated using the additional sampling time to obtain the secondary audio correction start time. The secondary audio correction start time is replaced with the initial audio correction start time, and the key trigger time is re-compared with the secondary audio correction start time and the collaborative judgment deadline. When the key trigger time is later than or the same as the secondary audio correction start time and earlier than or the same as the collaborative judgment deadline, it is determined that the key trigger time falls within the collaborative judgment time range defined by the secondary audio correction start time to the collaborative judgment deadline. The key input sequence and the input audio segment are determined to belong to the same user input operation, and the updated collaborative relationship judgment result is obtained. When the key trigger time is earlier than the secondary audio correction start time or later than the collaborative judgment deadline, it is determined that the key trigger time does not fall within the collaborative judgment time range defined by the secondary audio correction start time to the collaborative judgment deadline. The key input sequence and the input audio segment are determined to not belong to the same user input operation, and the updated independent relationship judgment result is obtained. When the updated collaborative relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are concatenated according to the collaborative correlation between key trigger time and secondary correction audio start time. When the updated independent relationship determination result is obtained, the input audio segments and key input sequences in the multimodal information set are concatenated according to the temporal order of key trigger time and secondary correction audio start time to generate a multimodal fusion input signal. When the confidence value of the user intent determination is greater than or equal to the confidence threshold, the input audio segment and key input sequence in the multimodal information set are directly concatenated according to the collaborative relationship determination result or the independent relationship determination result to generate a multimodal fusion input signal. This signal is used to control the keyboard to output the text characters corresponding to the key input sequence and convert the input audio segment into speech text for mixed output, or to execute keyboard shortcut commands triggered by key and speech combination.

7. The intelligent keyboard multimodal input signal fusion processing method according to claim 5, characterized in that, The initial signal fusion time window is set based on the longest time interval between the keyboard input and voice input that can be considered to belong to the same input operation before and after the user presses the voice conversion key.

8. The method for multimodal input signal fusion processing of an intelligent keyboard according to claim 5, characterized in that, The offset duration for determining the collaborative time is determined based on the time interval between the key trigger time when the user uses the keyboard voice conversion function and the start time of the first audio correction, which allows for collaborative correlation.

9. The intelligent keyboard multimodal input signal fusion processing method according to claim 6, characterized in that, The process for determining that the number of audio frames to be read remains constant is as follows: within a preset constant monitoring time range, read the number of audio frames to be read corresponding to the sampling time of multiple consecutive audio frames. If the read number of audio frames to be read is the same, then it is determined that the number of audio frames to be read in the audio frame change sequence remains constant.

10. A smart keyboard multimodal input signal fusion processing system, characterized in that, The system includes: The data acquisition module collects key input data and audio segments of user voice input generated when the user operates the keyboard, and constructs a multimodal information set of key input data and audio segments. The state determination module refers to the multimodal information set to determine whether the input audio segment is in an invalid audio start state; The time correction module determines the corresponding first correction audio start time by referring to the multimodal information set when the input audio segment is in an invalid audio start state, and calculates the audio start reference reliable time range. The confidence calculation module obtains a preset initial signal fusion time window, adjusts the initial signal fusion time window based on the audio start reference confidence time range, and calculates different user intent determination confidence values ​​by combining the multimodal information set and the first correction audio start time. The signal generation module compares the user intent determination confidence value with the preset confidence threshold, determines the secondary correction audio start time based on the comparison result, adjusts the multimodal information set, and generates a multimodal fusion input signal.