A Multimodal Multi-Scene Speech Recognition Method

CN122575367APending Publication Date: 2026-08-14NANJING BAIYIN INTELLIGENT SOUND TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明的目的是为了解决现有技术中存在的设备周期振动和冲击噪声随生产节拍在声学信号中形成类似语音音节边界的伪边界,并在振动信号中同步增强,导致多模态语音识别系统误锁定伪音节边界、影响短语音指令识别准确性的缺点,而提出的一种基于多模态的多场景语音识别方法

Benefits of technology

本发明通过同步获取工业产线中的声学信号和振动信号,分别从声学信号中提取声学边界响应值和候选语音边界集合,从振动信号中提取振动冲击响应值、设备动作边界集合、设备节拍周期和节拍复现强度,使系统能够从声学模态和振动模态两个角度同时分析短语音指令边界与设备动作边界之间的关系;相较于仅对工业噪声进行整体降噪或者仅依据声学能量进行端点检测的方式,本发明能够进一步识别设备周期振动和冲击噪声在声学信号中形成的类似语音音节边界的伪边界,避免将设备动作边界简单作为真实语音边界处理,从而提高工业复杂声学环境下短语音指令边界确定的准确性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575367A_ABST
    Figure CN122575367A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal, multi-scenario speech recognition method, relating to the field of speech recognition technology. The method includes: acquiring acoustic and vibration signals from an industrial production line and constructing a synchronous multimodal frame sequence; extracting candidate speech boundaries based on acoustic short-time energy and acoustic spectral flux; extracting equipment action boundaries based on vibration short-time energy and determining the equipment beat cycle and beat reproduction intensity; further calculating the cross-modal pseudo-syllable locking value of the candidate speech boundaries; correcting the acoustic boundary response value based on this locking value; determining the target speech segment and performing speech recognition; and outputting the speech recognition result and recognition reliability. This invention can identify pseudo-syllable boundaries formed by equipment periodic vibration and impact noise, reducing the risk of mis-interpretation, mis-segmentation, and mis-triggering of short speech commands.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and in particular to a multimodal, multi-scenario speech recognition method. Background Technology

[0002] With the development of intelligent manufacturing, industrial automation, and human-machine collaboration technologies, voice recognition is gradually being applied in industrial production line scenarios such as automated assembly, stamping, packaging production, warehousing and sorting, robot collaboration, and equipment inspection. Operators can issue short voice commands such as start, pause, reset, confirm, clamp, release, lower, and stop, thereby reducing manual button and touch screen operations or manual confirmation operations, and improving the efficiency of human-machine interaction and operational safety on the production site. In order to improve the recognition capability in complex environments, existing voice recognition systems usually introduce multimodal data such as acoustic signals, vibration signals, and environmental parameters, and combine voice activity detection, endpoint detection, noise reduction processing, and voice recognition models to complete short command recognition.

[0003] However, the acoustic environment of industrial production lines differs significantly from that of ordinary indoor voice interaction environments. Industrial sites typically contain the operating noises of equipment such as conveying mechanisms, motors, fans, pneumatic clamps, stamping mechanisms, sealing mechanisms, cutting mechanisms, robot end effectors, and workpiece contact and unloading mechanisms. These sounds not only manifest as continuous background noise but also generate periodic vibration and impact noise with the production cycle. Especially at the boundaries of equipment movements, impact noise often creates short-term energy spikes and spectral abrupt changes in the acoustic signal, while simultaneously showing synchronous enhancement in the vibration signal. Due to the short duration and limited semantic context of short voice commands, when the impact noise of the equipment occurs adjacent to the start, end, or syllable start of the operator's voice command, it is easy to form pseudo-boundaries similar to speech syllable boundaries during the speech boundary extraction process.

[0004] Existing speech recognition methods often treat industrial noise as ordinary background noise and focus on improving the signal-to-noise ratio, speech enhancement, or speech activity detection accuracy, while rarely distinguishing the source differences between equipment action boundaries and true speech boundaries. On the other hand, existing multimodal speech recognition methods usually use the synchronous enhancement of acoustic and vibration modes as a basis for improving event credibility. However, in industrial production lines, equipment actions themselves can simultaneously generate airborne sound and structural vibration. In this case, the consistent enhancement of acoustic and vibration signals does not necessarily indicate the enhancement of true speech boundaries, but may instead indicate the enhancement of equipment periodic action boundaries. If the system still processes according to ordinary multimodal consistency logic, it may mistakenly lock the equipment action boundaries as pseudo-syllable boundaries in short speech commands, leading to premature or delayed truncation of speech segments, internal missegmentation, or incorrect triggering of control commands. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies where equipment periodic vibration and impact noise form pseudo-boundaries in acoustic signals similar to speech syllable boundaries with the production rhythm, and are synchronously amplified in vibration signals, causing multimodal speech recognition systems to mislock pseudo-syllable boundaries and affect the accuracy of short speech command recognition. Therefore, this invention proposes a multimodal, multi-scenario speech recognition method.

[0006] To address the problems existing in the prior art, the present invention adopts the following technical solution: A multimodal, multi-scene speech recognition method includes: S1. Acquire acoustic and vibration signals from the industrial production line, perform time synchronization and framing processing on the acoustic and vibration signals to obtain a synchronized multimodal frame sequence; S2. Generate acoustic boundary response values ​​based on the acoustic short-time energy and acoustic spectral flux in the synchronous multimodal frame sequence, and extract a candidate speech boundary set; S3. Generate vibration impact response values ​​based on the short-time vibration energy in the synchronous multimodal frame sequence, extract the equipment action boundary set, and determine the equipment cycle period and cycle reproduction intensity; S4. Calculate the cross-modal pseudo-syllable locking value of each candidate speech boundary based on the candidate speech boundary set, the acoustic boundary response value, the vibration shock response value, the device action boundary set, the device beat period, and the beat reproduction intensity. S5. Correct the acoustic boundary response value according to the cross-modal pseudo-syllable locking value, and determine the target speech segment by combining the speech activity posterior probability obtained through speech activity detection; S6. Perform speech recognition on the target speech segment and output the speech recognition result and recognition credibility.

[0007] Preferably, acquiring acoustic and vibration signals from an industrial production line includes: Acoustic signals formed by the operator's voice, equipment operation sounds, and environmental noise are acquired through a voice acquisition terminal; Vibration signals propagated through the structure are acquired by vibration sensors installed on the equipment frame, workbench, or actuator. The acoustic signal and the vibration signal are aligned according to the same time reference and divided into a synchronous multimodal frame sequence according to the same frame length and the same frame shift. Multiple consecutive synchronous multimodal frames constitute the current recognition window. The acoustic short-time energy, vibration short-time energy, and acoustic spectrum are calculated based on the synchronous multimodal frame sequence.

[0008] Preferably, acoustic boundary response values ​​are generated based on the acoustic short-time energy and acoustic spectral flux in the synchronous multimodal frame sequence, and a candidate speech boundary set is extracted, including: The increase in acoustic energy is obtained by the positive logarithmic change of the acoustic short-time energy of the current frame relative to the acoustic short-time energy of the previous frame. The acoustic spectral flux is obtained by calculating the cumulative positive change of the current frame's acoustic spectral amplitude relative to the previous frame's acoustic spectral amplitude. The acoustic energy rise and the acoustic spectral flux are normalized within the current identification window, and the normalized acoustic energy rise and the normalized acoustic spectral flux are superimposed to obtain the acoustic boundary response value. The local maxima of the acoustic boundary response value are used as candidate positions, and the boundary screening value is determined by the Otsu's method based on the distribution of the acoustic boundary response value within the current recognition window. Candidate positions that are not lower than the boundary screening value are determined as candidate speech boundaries, thus forming a set of candidate speech boundaries.

[0009] Preferably, the vibration impact response value is generated based on the short-time vibration energy in the synchronous multimodal frame sequence, the set of equipment action boundaries is extracted, and the equipment cycle period and cycle reproduction intensity are determined, including: The increase in vibration energy is obtained by the positive logarithmic change in the short-time vibration energy of the current frame relative to the short-time vibration energy of the previous frame. The vibration energy rise is normalized within the current identification window to obtain the vibration impact response value; The local maximum location of the vibration and shock response value is used as the candidate device action location. The device boundary screening value is determined by the maximum inter-class variance method according to the distribution of the vibration and shock response value in the current identification window. The candidate device action locations that are not lower than the device boundary screening value are determined as the device action boundary, forming a set of device action boundaries. The vibration impact response value is normalized by autocorrelation processing. The autocorrelation value with the largest value is selected from the autocorrelation values ​​at the non-zero hysteresis position. The hysteresis corresponding to the largest autocorrelation value is determined as the equipment cycle time, and the non-negative value of the largest autocorrelation value is determined as the cycle time recurrence intensity.

[0010] Preferably, calculating the cross-modal pseudo-syllable locking value for each candidate speech boundary includes: For any candidate speech boundary in the candidate speech boundary set, determine the device action boundary that is closest to the candidate speech boundary in time from the device action boundary set. The synchronization proximity is calculated based on the time distance between the candidate speech boundary and the nearest device action boundary, as well as the dispersion of the distances from each candidate speech boundary to the nearest device action boundary within the current recognition window.

[0011] Preferably, calculating the cross-modal pseudo-syllable locking value for each candidate speech boundary further includes: The beat reproduction weight is calculated based on the response intensity of the closest equipment action boundary in the vibration and shock response value, the vibration and shock response value of the equipment action boundary at the position of the previous equipment beat cycle, the vibration and shock response value of the equipment action boundary at the position of the next equipment beat cycle, and the beat reproduction intensity. When the position of the previous device cycle or the position of the next device cycle exceeds the current identification window, the vibration and impact response value corresponding to the position exceeding the window is recorded as zero.

[0012] Preferably, calculating the cross-modal pseudo-syllable locking value for each candidate speech boundary further includes: The vibration dominance is calculated based on the vibration impact response value at the device action boundary with the closest time distance and the acoustic boundary response value at the candidate speech boundary. The acoustic spectral flatness is calculated based on the ratio between the geometric mean and the arithmetic mean of the acoustic spectral power at the candidate speech boundary. The posterior probability of the candidate speech boundary is obtained by using a speech boundary model based on acoustic speech content. The non-speech boundary quantity is obtained by subtracting the posterior probability of the speech boundary, and the synchronization proximity, the beat reproduction weight, the vibration dominance, the acoustic spectrum flatness, and the non-speech boundary quantity are multiplied to obtain the cross-modal pseudo-syllable locking value of the candidate speech boundary.

[0013] Preferably, the acoustic boundary response value is corrected based on the cross-modal pseudo-syllable locking value, and the target speech segment is determined by combining the posterior probability of speech activity obtained through speech activity detection, including: The acoustic boundary response value at the corresponding candidate speech boundary is reduced according to the magnitude of the cross-modal pseudo-syllable locking value to obtain the corrected boundary response value; Based on the posterior probability of speech activity obtained through speech activity detection and the corrected boundary response value, a segment score for the candidate speech segment is constructed. The segment score is composed of the logarithmic summation of the posterior probabilities of speech activities in each frame within the candidate speech segment, and the logarithmic summation of the corrected boundary response values ​​at the start and end boundaries of the candidate speech segment. The speech segment with the highest score is selected from the candidate speech segments using dynamic programming, and then used as the target speech segment.

[0014] Preferably, speech recognition is performed on the target speech segment, and the speech recognition result and recognition confidence level are output, including: The acoustic features corresponding to the target speech segment are input into the speech recognition model to obtain candidate instruction texts and their recognition probabilities. The candidate instruction text with the highest recognition probability is determined as the speech recognition result; The recognition confidence level is calculated based on the recognition probability of the speech recognition result and the average value of the cross-modal pseudo-syllable locking value at the start and end boundaries of the target speech segment. Output the speech recognition result and the recognition reliability.

[0015] Preferably, the industrial production line includes at least one of an automated assembly line, a stamping line, a packaging line, a warehousing and sorting line, or a robotic collaborative production line; the voice recognition result includes at least one short voice command among start, pause, reset, confirm, clamp, release, lower, and stop.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention simultaneously acquires acoustic and vibration signals from industrial production lines, extracting acoustic boundary response values ​​and candidate speech boundary sets from the acoustic signals, and vibration impact response values, equipment action boundary sets, equipment beat cycles, and beat reproduction intensities from the vibration signals. This allows the system to simultaneously analyze the relationship between short speech command boundaries and equipment action boundaries from both acoustic and vibration modal perspectives. Compared to methods that only perform overall noise reduction on industrial noise or rely solely on acoustic energy for endpoint detection, this invention can further identify pseudo-boundaries in the acoustic signals formed by periodic vibrations and impact noise from equipment, resembling speech syllable boundaries. This avoids simply treating equipment action boundaries as true speech boundaries, thereby improving the accuracy of short speech command boundary determination in complex industrial acoustic environments. This invention calculates the synchronization proximity, beat reproduction weight, vibration dominance, acoustic spectral flatness, and non-speech boundary quantities to obtain the cross-modal pseudo-syllable locking value of the candidate speech boundary. This cross-modal pseudo-syllable locking value is then used to correct the acoustic boundary response value, thereby determining the target speech segment and outputting the speech recognition result and recognition reliability. Vibration signals are no longer merely auxiliary information to confirm the presence of speech, but are used to identify the pulling effect of equipment action beats on the speech boundary. This reduces the risk of multimodal synchronous enhancement being mistaken for real speech enhancement, and reduces the occurrence of short speech commands being prematurely truncated, delayed truncated, internally mis-segmented, or erroneously triggered. This improves the stability, security, and applicability of speech recognition in multi-scenario industrial production lines. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a multimodal, multi-scene speech recognition method according to an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] Example 1: This example provides a multimodal, multi-scenario speech recognition method applicable to automated assembly lines, stamping lines, packaging lines, warehousing and sorting lines, or robotic collaborative production lines. These industrial production lines experience periodic equipment vibration and impact noise. With the production cycle, these vibrations and noise create localized energy spikes in the acoustic signal, resembling speech syllable boundaries, and are synchronously amplified in the vibration signal. This can easily cause the multimodal speech recognition system to mistakenly identify the equipment action boundaries as pseudo-syllable boundaries of short speech commands. This example calculates the cross-modal pseudo-syllable locking value of candidate speech boundaries and corrects the acoustic boundary response value based on this value to determine the target speech segment and output the speech recognition result. In this embodiment, the acoustic signal is denoted as The vibration signal is denoted as ,in, The sampling time is represented by ; the discrete sampling sequence of the acoustic signal is denoted as . The discrete sampling sequence of the vibration signal is denoted as ,in, Indicates the sampling point number; Indicates the frame number; Indicates the sequence number of the sampling point within the frame; Indicates frame shift; Indicates the analysis window function; This indicates the number of points in the short-time Fourier transform. Indicates frequency index; Represents the imaginary unit; This represents a minimal computational constant used to prevent logarithmic and division anomalies. A value greater than zero can be taken as [value] in numerical calculations. In the following embodiments, the same parameter variable refers to only the same object; This embodiment uses an automated assembly line as an example for illustration; the automated assembly line includes a conveying mechanism, pneumatic clamps, a robot actuator, and a manual confirmation station; during equipment operation, the operator issues short voice commands through a voice acquisition terminal, the short voice commands including at least one of start, pause, reset, confirm, clamp, release, descent, and stop; the voice acquisition terminal is installed near the manual confirmation station, and vibration sensors are installed on the equipment frame, workstation surface, or actuator to acquire vibration signals generated by structural propagation.

[0020] The method in this embodiment includes the following steps; S1. Acquire acoustic and vibration signals to construct a synchronous multimodal frame sequence; Specifically, acoustic signals from industrial production lines are acquired through voice acquisition terminals. The acoustic signal This includes operator voice, equipment operating noise, and environmental noise; equipment operating noise includes the operating noise of the conveyor mechanism, the movement noise of the pneumatic clamps, the movement noise of the robot actuator, and the impact noise generated by the contact of the workpiece; Vibration signals are acquired by vibration sensors installed on the equipment frame, workbench, or actuator. The vibration signal This indicates structural propagation signals generated by equipment movement, structural transmission, workpiece contact, and equipment impact. acoustic signals and vibration signals To synchronize the time, so that the acoustic signal and vibration signals They share the same time base; time synchronization can be achieved through timestamps from the same acquisition controller or through a unified trigger signal at the acquisition end; after time synchronization is complete, the acoustic signals will be... and vibration signals The frames are divided into synchronous multimodal frames based on the same frame length and the same frame shift. In this embodiment, the frame length is set to 20ms to 40ms according to the requirements of short-term stationary speech analysis, and the frame shift is set to 10ms to 20ms; as a specific implementation, the frame length is 25ms and the frame shift is 10ms; multiple consecutive synchronous multimodal frames constitute the current recognition window; In this embodiment, the current recognition window refers to the continuous frame interval extracted from the synchronous multimodal frame sequence for a single short speech command recognition process; the current recognition window is not a fixed scene parameter, but is composed of continuous synchronous multimodal frames in the real-time cache, and is updated by sliding during the acquisition process; the current recognition window is used to limit the statistical range of subsequent acoustic energy normalization, vibration energy normalization, boundary screening, device beat cycle determination, candidate speech segment construction, and target speech segment determination; The length of the current recognition window is determined by the equipment action interval of the target production line and the duration of the short voice command. Specifically, during the system deployment or debugging phase, acoustic and vibration signals of the equipment in the target production line are collected during continuous operation. The frame interval between the local maximum positions of the vibration impact response corresponding to two adjacent equipment actions is counted, and target short voice command samples are collected, and the number of continuous frames of the target short voice command samples is counted. The length of the current recognition window is set to be no less than the larger of the frame interval corresponding to two adjacent equipment actions and the number of continuous frames of the target short voice command samples. For scenarios where the equipment action interval changes with the operating conditions, the length of the current recognition window is updated according to the boundary interval of the most recently recognized adjacent equipment actions, so that the current recognition window covers at least one short voice command to be recognized and its adjacent equipment action boundaries. In this way, the current recognition window can cover the entire time range of short voice commands and include the action information of adjacent devices used to determine the reproduction of device beats. This avoids the situation where the device beat cycle cannot be calculated due to the window being too short, or the action of irrelevant devices affecting normalization and boundary filtering due to the window being too long.

[0021] Calculate the acoustic short-time energy of the nth frame respectively. and the short-time energy of the nth frame vibration : ; ; in, This represents the energy intensity of the acoustic signal in the nth frame. This represents the energy intensity of the vibration signal in the nth frame; acoustic signals Corresponding discrete sampling sequence Perform a short-time Fourier transform to obtain the acoustic spectrum. : ; in, This represents the complex spectral value of the acoustic signal in the nth frame at the kth frequency index.

[0022] S2. Extract candidate speech boundaries based on acoustic energy surge and acoustic spectral flux; Specifically, based on acoustic short-time energy Calculate the increase in acoustic energy in the nth frame. : ; in, This represents the positive logarithmic change in the acoustic short-time energy of the current frame relative to the acoustic short-time energy of the previous frame; when the acoustic short-time energy of the current frame is less than or equal to the acoustic short-time energy of the previous frame, Set to 0; when the acoustic short-time energy of the current frame is greater than the acoustic short-time energy of the previous frame, Reflects the degree of sudden increase in acoustic energy; According to the acoustic spectrum Calculate the acoustic spectral flux of the nth frame. : ; in, This represents the cumulative positive change in the acoustic spectrum amplitude of the current frame relative to the acoustic spectrum amplitude of the previous frame; acoustic spectral flux. Used to characterize the degree of abrupt changes in the spectrum of acoustic signals; For the current recognition window and Normalization is performed separately for each sequence; for any sequence to be normalized Normalization operator for: ; in, Represents a sequence The minimum value within the current recognition window. Represents a sequence The maximum value within the current recognition window; Based on the normalized acoustic energy rise and normalized acoustic spectral flux Generate acoustic boundary response values : ; In this embodiment, the acoustic boundary response value This refers to the numerical value used to characterize whether the nth frame of the acoustic signal has the speech boundary candidate attribute; acoustic boundary response value. The acoustic boundary response value is determined jointly by the acoustic energy rise and the acoustic spectral flux. The acoustic energy rise reflects the sudden increase in the energy of the acoustic signal over time, while the acoustic spectral flux reflects the degree of change in the acoustic signal's spectrum. The larger the value, the more likely the nth frame is to correspond to the start or end point of a speech segment or the boundary of a syllable start point; For short speech commands, the onset of a real speech syllable is usually accompanied by a local energy increase and spectral changes, therefore the acoustic boundary response value... It can be used to extract candidate speech boundaries.

[0023] Based on the acoustic boundary response value within the current recognition window The distribution of the data is analyzed, and the boundary screening values ​​are automatically determined using the maximum inter-class variance method. Specifically, all [items] within the current recognition window As the data to be segmented, for each candidate filter value, The classes are divided into low-response and high-response categories. The proportions of low-response and high-response classes, as well as the mean values ​​of low-response and high-response classes, are calculated separately. The candidate screening value that maximizes the inter-class variance between the low-response and high-response classes is selected as the boundary screening value. ; acoustic boundary response value Local maxima are used as candidate positions, and the minimum threshold value is not lower than the boundary screening value. The candidate positions are determined as candidate speech boundaries, forming a set of candidate speech boundaries. : ; in, The element in the text is the frame number corresponding to the candidate speech boundary; In this embodiment, the candidate speech boundary refers to the boundary response value. The preliminary selection of possible speech boundary locations; candidate speech boundaries only indicate that the location has a local energy surge or spectral change in the acoustic signal, and do not necessarily represent the true speech boundary; the set of candidate speech boundaries. Each element in the algorithm is a frame number, which is used as a candidate speech boundary in subsequent steps to calculate the cross-modal pseudo-syllable locking value.

[0024] S3. Extract the device action boundary based on the sudden increase in vibration energy and the autocorrelation of the beat; Specifically, based on the short-time energy of vibration Calculate the increase in vibration energy in the nth frame. : ; in, This represents the positive logarithmic change in the short-time vibration energy of the current frame relative to the short-time vibration energy of the previous frame; when the short-time vibration energy of the current frame is less than or equal to the short-time vibration energy of the previous frame, Set to 0; when the short-time vibration energy of the current frame is greater than the short-time vibration energy of the previous frame, It reflects the degree of sudden increase in the vibration of the equipment structure; For the current recognition window After normalization, the vibration and shock response values ​​are obtained. : ; In this embodiment, the vibration and shock response value This refers to the numerical value used to characterize the impact intensity of equipment movement in the nth frame of the vibration signal; vibration impact response value. Obtained by normalizing the vibration energy rise, it reflects the sudden vibration rise generated in the structural propagation path during equipment clamping, releasing, punching, unloading, return, stopping, or robot end effector movements; vibration impact response value. The larger the value, the more likely the nth frame is to correspond to the structural vibration boundary caused by the device's operation; Based on the vibration and shock response values ​​within the current identification window The distribution of the equipment boundary screening values ​​is automatically determined using the Otsu's method. Specifically, all [items] within the current recognition window As the data to be segmented, for each candidate filter value, The devices are divided into low-impact response and high-impact response categories. The proportions of low-impact response and high-impact response categories, as well as their mean values, are calculated. The candidate value that maximizes the inter-class variance between the low-impact response and high-impact response categories is selected as the device boundary screening value. ; Vibration and shock response values The local maxima are used as candidate device action positions, and the values ​​are not lower than the device boundary screening value. The candidate device action positions are determined as device action boundaries, forming a set of device action boundaries. : ; in, The element in the array is the frame number corresponding to the device action boundary; In this embodiment, the equipment action boundary refers to the boundary based on the vibration and shock response value. Extracted equipment action locations; equipment action boundaries are used to characterize the boundary locations of equipment fixture closure, cylinder action, stamping contact, material unloading, conveyor stopping, robot gripper action, or workpiece contact events in the vibration signal; equipment action boundary set. Each element in the table is a frame number, which is used in subsequent steps to determine the device action boundary that is closest in time to the candidate speech boundary.

[0025] To determine whether the equipment's operation has a production cycle traction, the vibration and shock response value is... Normalized autocorrelation is performed to obtain the normalized autocorrelation function. : ; in, This is the frame lag. for The mean within the current recognition window, Indicates vibration and shock response value The degree of repeatability under lag; Before selecting the largest autocorrelation value from the non-zero hysteresis locations, the autocorrelation hysteresis search range must first be determined; specifically, when the equipment action boundary set... When there are at least two device action boundaries, the frame interval between adjacent device action boundaries is calculated in chronological order. The minimum value in the frame interval is taken as the lower limit of the autocorrelation hysteresis search, and the maximum value in the frame interval is taken as the upper limit of the autocorrelation hysteresis search. The autocorrelation value with the largest value is selected within the autocorrelation hysteresis search range. When the device action boundary set If there are fewer than two device action boundaries in the current identification window, it is assumed that there are no adjacent device action boundaries within the current identification window that can be used to determine the device cycle time, and the device cycle time is set to zero. Record it as 0, and the beat reproduction intensity Set to 0, and then set the weight for subsequent beat recurrence. Set to 0; When the device action boundary set When there are at least two device action boundaries, the hysteresis corresponding to the largest autocorrelation value within the autocorrelation hysteresis search range is determined as the device cycle time. The non-negativity of the largest autocorrelation value is then determined as the beat recurrence intensity. : ; ; in, This is the minimum frame interval between the action boundaries of adjacent devices. This represents the maximum frame interval between the action boundaries of adjacent devices; the device cycle time is determined in the above manner. The determination is constrained by the actual interval of the equipment action boundary, which can avoid mistaking the noise similarity of extremely short lag positions as the equipment cycle period; In this embodiment, the equipment cycle time This refers to the frame interval at which the boundaries of device actions within the current recognition window repeat according to the production cycle; the device cycle time. Instead of a preset fixed period, it is based on the vibration and shock response value. The normalized autocorrelation results, combined with the device action boundary set Determine the interval between the action boundaries of adjacent devices; Beat Reproduction Intensity This refers to the equipment operating according to the equipment cycle time. The intensity of repetition; the strength of beat repetition The larger the value, the more stable and repeatable the device actions are within the current recognition window; the beat reproduction intensity. When the value is 0, it means that there is no valid evidence of device beat reproduction within the current recognition window, and the subsequent beat reproduction weight does not contribute to the judgment of pseudo-syllables at the candidate speech boundary.

[0026] This step outputs the vibration and shock response value. Equipment action boundary set Equipment cycle time and beat reproduction intensity Vibration and shock response values Equipment action boundary set Equipment cycle time and beat reproduction intensity In step S4, the cross-modal pseudo-syllable locking value for each candidate speech boundary is calculated; if the device action boundary set If there are fewer than two device action boundaries, then the device cycle time is... Recorded as 0, beat recurrence intensity Set to 0, and the weight for subsequent beat reproduction. Take 0.

[0027] S4. Calculate the cross-modal pseudo-syllable locking value of candidate speech boundaries; Specifically, for the candidate speech boundary set Any candidate speech boundary in From the set of device action boundaries The device action boundary that is closest in time to the candidate speech boundary c is determined in the middle. : ; in, The frame number corresponding to the candidate speech boundary. Set of device action boundaries The frame number corresponding to the device action boundary in the code. To and The boundary of the device action closest in time; If the device action boundary set If empty, then set the cross-modal pseudo-syllable locking value L(c) of the candidate speech boundary c to 0, and proceed directly to step S5; if the device action boundary set If not empty, continue with the following calculations; Calculate candidate speech boundaries The device action boundary closest to the time distance Synchronization proximity between : ; in, The root mean square value of the distance from the boundary of each candidate speech within the current recognition window to the nearest device action boundary: ; In this embodiment, synchronization proximity Refers to candidate speech boundaries The device action boundary closest to its time distance The degree of proximity in time; the degree of synchronization proximity Used to determine whether the candidate speech boundary may be affected by device action boundaries; candidate speech boundary Boundary of equipment action The smaller the time distance, the closer the synchronization. Larger; candidate speech boundary Boundary of equipment action The greater the time distance, the closer the synchronization. The smaller; According to the equipment cycle time and beat reproduction intensity Calculate the beat reproduction weights corresponding to the candidate speech boundary c. ; When the beat reappears in intensity When it is 0, the beat reproduction weight =0; when the beat recurrence intensity is 0. When greater than 0, the cycle time recurrence weight Calculate as follows: ; in, Indicates the boundary of device action Vibration and shock response value at the position of the previous equipment cycle; Indicates the boundary of device action Vibration and shock response values ​​at the location; Indicates the boundary of device action The vibration and impact response value at the position of the next equipment cycle; Indicates the vibration and shock response value within the current identification window. The maximum value; when or When the current recognition window is exceeded, the vibration and shock response value corresponding to the position exceeded is recorded as 0; In this embodiment, the beat reproduction weight This refers to whether the device action boundary corresponding to the candidate speech boundary c follows the device clock cycle. Weights for repeated occurrences; weights for repeating beats By equipment action boundary Vibration shock response at the location, vibration shock response at the location of the previous equipment cycle, vibration shock response at the location of the next equipment cycle, and cycle reproduction intensity. Jointly determined; beat recurrence weight The larger the value, the more likely the device action near the candidate speech boundary is to be a recurring device action that occurs repeatedly with the production cycle, rather than a one-off random noise.

[0028] Based on the device action boundary with the closest time distance Vibration and shock response value at the location Acoustic boundary response value at candidate speech boundary c Calculate the vibration dominance : ; in, This represents the normalized acoustic boundary response value within the current recognition window; In this embodiment, vibration dominance This refers to the degree to which the boundary response near the candidate speech boundary c is dominated by device vibration; vibration dominance. Based on the device action boundary with the closest time distance The relative relationship between the vibration impact response value at point c and the acoustic boundary response value at candidate speech boundary c is determined; vibration dominance The larger the value, the more likely the candidate speech boundary is to originate primarily from structural vibrations caused by device movement, rather than the operator's speech itself; Based on the acoustic spectrum at the candidate speech boundary c Calculate acoustic spectral power : ; According to acoustic spectrum power The ratio between the geometric mean and the arithmetic mean is used to calculate the acoustic spectral flatness. : ; in, This refers to the number of frequency points. In this embodiment, acoustic spectral flatness The acoustic spectral flatness refers to the degree of dispersion of acoustic spectral power at the candidate speech boundary c along the frequency axis. Determined by the ratio between the geometric mean and arithmetic mean of the acoustic spectral power; acoustic spectral flatness. A higher value indicates a more broadband acoustic energy distribution at the candidate speech boundary, which better matches the broadband transient properties of device impulse noise; acoustic spectral flatness The lower the value, the more concentrated the acoustic energy at the candidate speech boundary is in a few frequency regions.

[0029] By using a speech boundary model based solely on acoustic speech content, the posterior probability of the candidate speech boundary c is obtained. The input to the speech boundary model is the acoustic features near the candidate speech boundary c, and the output is the probability that the candidate speech boundary c belongs to the real speech boundary; the acoustic features include at least one of Mel frequency cepstral coefficients, acoustic short-time energy, acoustic spectral flux, and acoustic spectral power. In this embodiment, the speech boundary model is a boundary binary classification model obtained before deployment or during the equipment debugging phase. The training samples of the boundary binary classification model include short speech command samples with real speech boundary annotations and non-boundary samples that do not belong to the real speech boundary. The real speech boundary annotations include the start boundary, end boundary, or syllable start boundary of the short speech command, and the non-boundary samples include speech smooth segments, background noise segments, and equipment impact sound segments. The boundary binary classification model can adopt one of the following: support vector machine, hidden Markov model, convolutional neural network, recurrent neural network, or end-to-end speech front-end model. Its input is the acoustic features near the candidate speech boundary c, and its output is the posterior probability of the candidate speech boundary c belonging to the real speech boundary. ; The speech boundary model does not input vibration and shock response values. Equipment action boundary set Equipment cycle time and beat reproduction intensity To ensure the posterior probability of speech boundaries This only indicates the degree to which the acoustic speech content supports the real speech boundary; therefore, when a candidate speech boundary is synchronized with the device action boundary, but its acoustic speech content does not support the real speech boundary, it can be subsequently determined by subtracting the posterior probability of the speech boundary. A larger non-speech boundary quantity is obtained; In this embodiment, the posterior probability of the speech boundary This refers to the probability that a candidate speech boundary c belongs to a real speech boundary based solely on its acoustic speech content; the posterior probability of a speech boundary. The output of the speech boundary model has a value ranging from 0 to 1; the speech boundary model does not accept vibration and shock response values ​​as input. Equipment action boundary set Equipment cycle time and beat reproduction intensity Therefore, the posterior probability of the speech boundary It does not reflect the timing information of device actions, but only the degree to which the acoustic speech content supports the boundaries of real speech; Based on the posterior probability of the speech boundary minus 1 We obtain the non-speech boundary quantities: ; In this embodiment, the non-speech boundary quantity refers to a quantity calculated by subtracting the posterior probability of the speech boundary from a given value. The obtained quantity is used to characterize the probability that the candidate speech boundary c does not belong to the real speech boundary; the posterior probability of the speech boundary The lower the value, the larger the non-speech boundary quantity, indicating that the candidate speech boundary lacks acoustic speech content support from the real speech boundary. Synchronization proximity Beat Recurrence Weight Vibration dominance Acoustic spectral flatness The cross-modal pseudosyllabic locking value of the candidate speech boundary c is obtained by product fusion of non-speech boundary quantities. : ; like If it is less than 0, then Recorded as 0; if If it is greater than 1, then It is denoted as 1 to ensure the correction of the boundary response value. It will not be less than 0 due to abnormal values; In this embodiment, the cross-modal pseudosyllable locking value This refers to a numerical value used to characterize the probability that a candidate speech boundary c will form a pseudo-syllable boundary due to the pull of device action beats; cross-modal pseudo-syllable locking value. From synchronization proximity Beat Recurrence Weight Vibration dominance Acoustic spectral flatness Determined jointly by non-speech boundary quantities; cross-modal pseudosyllable locking value The larger the value, the more likely the candidate speech boundary c is to be a pseudo-syllable boundary formed by device periodic vibration and impact noise; cross-modal pseudo-syllable locking value The smaller the value, the more likely the candidate speech boundary c is not a pseudo-syllable boundary formed by device beat traction.

[0030] S5. Release pseudo-syllable boundaries and determine the target speech segment based on the cross-modal pseudo-syllable locking value; Based on the cross-modal pseudosyllabic locking value of the candidate speech boundary c Correct the acoustic boundary response value at candidate speech boundary c. The corrected boundary response value is obtained. : ; In this embodiment, the boundary response value is corrected. This refers to the acoustic boundary response value based on the cross-modal pseudosyllable locking value L(c). The boundary response value obtained after reduction processing; the corrected boundary response value Used to replace the original acoustic boundary response value during the target speech segment determination process. When cross-modal pseudosyllable locking value When it is large, the corrected boundary response value corresponding to the candidate speech boundary c The reduction makes it difficult to use pseudo-syllable boundaries formed by device action beats as high-confidence speech segment boundaries; when the cross-modal pseudo-syllable locking value is reduced... When the value is small, the corrected boundary response value corresponding to the candidate speech boundary c Maintain a high level; The posterior probability of speech activity in the nth frame within the current recognition window is obtained through speech activity detection. Posterior probability of voice activity This represents the probability that the nth frame belongs to a speech frame, with a value ranging from 0 to 1; the input for speech activity detection is an acoustic signal. The extracted acoustic features include at least one of acoustic short-time energy, acoustic spectral flux, Mel spectral features, and Mel frequency cepstral coefficients; In one specific implementation, speech activity detection employs a trained binary classification model. The training samples of this model include acoustic samples with and without speech frame labels. The input is acoustic features, and the output is the probability that the current frame belongs to a speech frame. In another specific implementation, speech activity detection employs a statistical classifier based on acoustic short-time energy and acoustic spectral flux. The acoustic short-time energy and acoustic spectral flux of the current frame are input into the statistical classifier, and the output is the probability that the frame belongs to a speech frame. The above speech activity detection does not input vibration shock response values, equipment action boundary sets, equipment beat periods, and beat reproduction intensity, thus making the posterior probability of speech activity... It mainly reflects the degree of presence of speech activity in the acoustic signal; Set of candidate speech boundaries Candidate speech boundaries are arranged in chronological order to construct candidate speech segments; for any candidate speech segment U, the set of start and end boundaries of candidate speech segment U is denoted as . The set of frame numbers contained in the candidate speech segment U is denoted as The start and end boundaries of candidate speech segment U both come from the candidate speech boundary set. .

[0031] Based on the posterior probability of speech activity and corrected boundary response values Construct segment scores for candidate speech segments U : ; in, is the logarithmic summation term of the posterior probability of speech activity for each frame within the candidate speech segment U; is the candidate speech segment is the logarithmic summation term of the corrected boundary response values at the start and end boundaries; segment score simultaneously considers the reliability of speech activity within the speech segment and the reliability of the start and end boundaries of the speech segment; selects the speech segment with the highest segment score from the candidate speech segments through dynamic programming as the target speech segment : ; Specifically, arrange the candidate speech boundaries in the candidate speech boundary set in chronological order as , , ……, , where is the number of candidate speech boundaries; for any two candidate speech boundaries and that satisfy i < j, the continuous frame interval from to forms a candidate speech segment , and and are used as the start and end boundaries of the candidate speech segment ; calculate the segment score of each candidate speech segment ; Establish a dynamic programming state according to the chronological order of the candidate speech boundaries. Let represent the highest segment score that can be obtained before the j-th candidate speech boundary. Then is updated according to and the segment score of the candidate speech segment with as the end boundary. Record the start boundary and end boundary corresponding to the highest segment score during the update process; after traversing all candidate speech boundaries, determine the continuous frame interval between the recorded start boundary and end boundary as the target speech segment ; When this method only outputs one short speech command recognition result each time, the above dynamic programming is equivalent to selecting the candidate speech segment with the highest segment score and continuous frame sequence from all candidate speech segments as the target speech segment ; through the above method, the start point, end point, posterior probability of internal speech activity, and boundary correction response of the target speech segment all participate in segment selection, avoiding determining the short speech command segment only based on a single local boundary response.

[0032] In this embodiment, the target speech segment This refers to the continuous speech frame interval determined from the current recognition window after the pseudo-syllable boundaries have been released, which is used as the input to the speech recognition model; the target speech segment. The target speech segment is determined by a segment score, which considers both the posterior probability of speech activity in each frame within the candidate speech segment and the corrected boundary response value at the start and end boundaries of the candidate speech segment; The start and end boundaries do not use the high-locked candidate boundaries that are pulled by the device action beat, thereby reducing the impact of device impact noise on short voice command segmentation; If the candidate speech boundary set If the number of candidate speech boundaries is less than two, then the posterior probability of the speech activity is used. The distribution within the current recognition window is used to determine the speech activity screening value using the maximum inter-class variance method, and continuous frame intervals not lower than the speech activity screening value are selected as target speech segments. This process ensures that a recognizable target speech segment can still be obtained even when the number of candidate speech boundaries is insufficient. When candidate speech boundary set When empty, segment combination based on candidate speech boundaries is not performed; instead, the posterior probability of speech activity is used directly. The distribution within the current recognition window determines the speech activity filter value, and the longest consecutive frame interval not lower than the speech activity filter value is taken as the target speech segment. When there is no continuous frame interval not lower than the stated voice activity screening value, it is determined that there is no identifiable target voice segment within the current recognition window, and the target voice segment is... Record as empty.

[0033] S6. Perform short command recognition based on the target speech segment and output the recognition result; When the target speech segment When empty, the speech recognition model is not invoked for instruction recognition; instead, the unrecognized result and recognition confidence score of 0 are directly output. When the target speech segment is empty... If not empty, extract the target speech segment. The corresponding acoustic features are then input into the speech recognition model; the speech recognition model outputs candidate command texts and their recognition probabilities. In this embodiment, the speech recognition model is a trained short command recognition model, keyword recognition model, or automatic speech recognition model; the training samples of the speech recognition model include short voice command samples and non-command voice samples collected from the target industrial production line, with the short voice command samples bearing corresponding command text annotations; the input of the speech recognition model is the target voice segment. The corresponding acoustic features include at least one of Mel-frequency spectral features, Mel-frequency cepstral coefficients, acoustic short-time energy, and acoustic spectral flux; the output of the speech recognition model is candidate command texts and the recognition probability corresponding to each candidate command text; The candidate command text includes at least one short voice command selected from start, pause, reset, confirm, clamp, release, descend, and stop; the candidate command text with the highest recognition probability is determined as the voice recognition result. : ; in, Indicates the candidate instruction text. Indicates in the target speech segment Output candidate instruction text under acoustic characteristics conditions The recognition probability, This represents the speech recognition result with the highest recognition probability. Based on speech recognition results The recognition probability and the target speech segment The average value of the cross-modal pseudosyllable locking values ​​at the start and end boundaries is used to calculate the recognition confidence. : ; in, Represents the target speech segment The set of start and end boundaries, Represents the target speech segment The average value of cross-modal pseudosyllable locking at the start and end boundaries; if the target speech segment The boundary does not come from the candidate speech boundary set. Then the cross-modal pseudosyllabic locking value at the corresponding boundary is recorded as 0; In this embodiment, the credibility of identification is determined. It refers to the method used to characterize speech recognition results. A numerical value indicating the degree of reliability that can be adopted by industrial production line control systems; identification reliability. Based on speech recognition results The recognition probability and target speech segments The average value of the cross-modal pseudo-syllable locking value at the start and end boundaries is determined; when the recognition probability of the speech recognition result is high and the cross-modal pseudo-syllable locking value at the start and end boundaries of the target speech segment is low, the recognition reliability is high. The recognition reliability is relatively high when there are high cross-modal pseudo-syllable locking values ​​at the start and end boundaries of the target speech segment. Reduced; This step outputs the speech recognition results. and recognition credibility Speech recognition results It can be used for voice interaction control in industrial production lines, with high recognition reliability. It can be used to prompt manual confirmation, control system verification, or safety interlock judgment.

[0034] In this embodiment, if the device action boundary set If empty, then the cross-modal pseudo-syllable locking value of the candidate speech boundary c. Recorded as 0; if the device action boundary set If there are fewer than two device action boundaries, then the device cycle time is... Recorded as 0, beat recurrence intensity Set to 0, and set the beat recurrence weight. =0; if the beat recurrence intensity is 0; If it is 0, then the beat recurrence weight is... =0; if the candidate speech boundary set is 0; If the number of candidate speech boundaries is insufficient or empty, then the posterior probability of the speech activity is used. Identify target speech segment If the target speech segment If empty, the output will be an unrecognized result and a recognition confidence of 0. Through the above abnormal situation handling, this method can maintain process closure when the device action is not obvious, the device rhythm is unstable, the candidate speech boundary is insufficient, or no effective speech segment is detected, thus avoiding the generation of undefined intermediate quantities or erroneous output control commands.

[0035] Through the above steps, this embodiment can identify pseudo-syllable boundaries formed by equipment periodic vibration and impact noise during multimodal speech recognition in industrial production lines, and reduce the impact of pseudo-syllable boundaries on the target speech segment determination process by using cross-modal pseudo-syllable locking values, thereby reducing the situation where short speech commands are truncated, mis-segmented, or mis-triggered by equipment action boundaries.

[0036] Example 2: This example uses a stamping production line as an example to illustrate the implementation of this method in another industrial scenario. In the stamping production line, the downward movement of the press slide, the contact of the die, the unloading of the workpiece, and the positioning of the conveying mechanism will all generate obvious impact noise and structural vibration. The operator can issue short voice commands such as pause, stop, reset, or confirm through a voice acquisition terminal next to the stamping production line. In this embodiment, the acoustic signal Vibration signals are collected by a voice acquisition terminal located next to the stamping station. The vibration data is collected by vibration sensors installed on the press frame or the press station table; due to the distinct production rhythm of the press operation, the vibration and impact response values ​​are... Repeated peak values ​​appear in adjacent stamping cycles; the equipment cycle time can be obtained through the normalized autocorrelation processing in step S3. and beat reproduction intensity ; When the operator issues a pause or stop command before or after the punch press operation, the impact noise from the punch press may occur at the acoustic boundary response value. Local maxima are formed in the middle and enter the candidate speech boundary set. At this point, if the candidate speech boundary matches the device action boundary set... If the boundary times of the equipment actions are close, and there are vibration and shock response reproductions at the positions of the preceding and following equipment cycle periods, then the cycle reproduction weight... Increase; if the vibrational impact response at this boundary is stronger than the acoustic boundary response, then the vibrational dominance... Increase; if the acoustic spectrum flatness at this boundary High and speech boundary posterior probability If the value is lower, the cross-modal pseudosyllable locking value increases; Step S5 reduces the acoustic boundary response value at the candidate speech boundary based on the cross-modal pseudo-syllable locking value. The corrected boundary response value is obtained. This prevents the pseudo-syllable boundaries formed by the impact sound of the punch press from being used as highly reliable start and end boundaries of the target speech segment; thereby reducing the possibility of the impact sound of the punch press being misidentified as a short instruction segment; in this embodiment, the remaining data processing flow is the same as in Embodiment 1.

[0037] Example 3: This example uses a packaging production line as an example to illustrate the implementation of this method in a voice recognition scenario for packaging equipment. In the packaging production line, the sealing mechanism, cutting mechanism, pushing mechanism, palletizing mechanism, and conveyor stop will move according to a fixed or nearly fixed production rhythm, generating periodic impact noise and structural vibration. Operators usually issue short voice commands such as start, pause, reset, and confirm during equipment operation. In this embodiment, the voice acquisition terminal is installed near the control panel of the packaging equipment, and the vibration sensor is installed on the frame where the sealing mechanism or cutting mechanism is located; acoustic signal This includes operator voice, noise from packaging material friction, noise from the sealing mechanism, noise from the cutting blade, and noise from the conveying mechanism; vibration signals. Structural vibrations transmitted to the frame, including those from the sealing mechanism, cutting mechanism, and pushing mechanism; During the operation of the packaging equipment, the cutting mechanism generates broadband transient impact sound, which may be close to the start time of the operator's short voice command. Step S2 extracts the acoustic energy surge and acoustic spectrum change corresponding to the impact sound as candidate speech boundaries. Step S3 extracts the vibration surge corresponding to the cutting mechanism's action as the equipment action boundary, and obtains the equipment beat period and beat reproduction intensity through autocorrelation processing. Step S4 calculates the cross-modal pseudo-syllable locking value of the candidate speech boundary based on the synchronization proximity, beat reproduction weight, vibration dominance, acoustic spectrum flatness, and non-speech boundary quantity. If the candidate speech boundary is a pseudo-syllable boundary caused by the action of the cutting mechanism, it usually exhibits the following characteristics: the timing is close to the boundary of the device action; there is a vibration impact response reproduction at the positions of the preceding and following device beat cycles; the vibration response accounts for a relatively high proportion of the acoustic boundary response; the acoustic spectral power has high spectral flatness; the posterior probability of the speech boundary obtained solely based on the acoustic speech content is low; therefore, the cross-modal pseudo-syllable locking value is high; step S5 reduces the response of the candidate speech boundary according to the cross-modal pseudo-syllable locking value, thereby releasing the pseudo-syllable boundary locking; in this embodiment, the remaining data processing flow is the same as in embodiment one.

[0038] Example 4: This example uses a robot collaborative production line as an example to illustrate the implementation of this method in a robot motion noise scenario. In a robot collaborative production line, the robot joint start-up, end-effector clamping, end-effector release, workpiece placement, and safety fence interlocking actions may all generate acoustic impact and structural vibration. Operators may issue short voice commands such as stop, descent, clamping, release, or confirmation when the robot end-effector approaches the workpiece, the gripper closes, or the workpiece is placed. In this embodiment, the acoustic signal Vibration signals were acquired by a voice acquisition terminal located near the robot workstation. Vibration sensors installed on the robot base, fixture base, or workbench acquire data. The structural vibrations caused by the robot's movements and the airborne impact sounds are synchronized in time. When this synchronization occurs near short voice commands, it can easily cause the candidate voice boundary to be close to the device's movement boundary. After processing according to steps S1 to S6, the system can identify the cross-modal synchronization enhancement caused by robot actions as candidate pseudo-syllable boundaries, and correct the acoustic boundary response value based on the cross-modal pseudo-syllable locking value. For real speech boundaries, the cross-modal pseudo-syllable locking value is low because its posterior probability is high and it may not have device beat reproduction weight. For robot action boundaries, the cross-modal pseudo-syllable locking value is high because its vibration dominance, synchronization proximity, and acoustic spectrum flatness are high and it may have beat reproduction attributes. Through the above differences, the system can reduce the pull of robot action noise on the boundaries of short speech commands. In this embodiment, the remaining data processing flow is the same as in Embodiment 1.

[0039] Example 5 illustrates how this method works when the boundaries of equipment movements are unclear or the equipment cycle time is unstable. In some industrial production lines, equipment movements may not have a stable rhythm, or the number of equipment movements within the current recognition window may be low, leading to a decrease in the normalized autocorrelation function. The autocorrelation value is low at the non-zero lag position; at this time, step S3 determines the non-negative value of the autocorrelation value with the largest value within the autocorrelation lag search range as the beat recurrence intensity. If the autocorrelation value is less than or equal to 0, then the beat recurrence intensity... =0; if the device action boundary set is 0; If there are fewer than two device action boundaries, then the device cycle time is... Recorded as 0, beat recurrence intensity Set to 0; When the beat reappears in intensity When the value is 0, the beat reproduction weight in step S4 The value is 0. Since the cross-modal pseudo-syllable locking value is obtained by fusing the product of synchronization proximity, beat reproduction weight, vibration dominance, acoustic spectrum flatness and non-speech boundary quantity, the cross-modal pseudo-syllable locking value is 0. At this time, this method will not determine the candidate speech boundary as a pseudo-syllable boundary formed by device beat, nor will it incorrectly reduce the acoustic boundary response value of the real speech boundary due to the absence of device beat. In the set of device action boundaries When the value is empty, step S4 directly sets the cross-modal pseudo-syllable locking value of the candidate speech boundary c to 0; at this time, step S5 can still be based on the acoustic boundary response value. Posterior probability of speech activity Identify target speech segment This processing ensures that when the equipment impact is not significant or the vibration sensor does not detect the effective equipment action boundary, this method can still degenerate into a speech segment determination method based on acoustic boundary response and posterior probability of speech activity. When candidate speech boundary set If the value is empty, step S5 directly uses the posterior probability of the speech activity. Identify target speech segment When the target speech segment When empty, step S6 outputs an unrecognized result and a recognition confidence of 0, and does not output short voice commands for industrial control. Through this embodiment, it is possible to avoid erroneously releasing real voice boundaries or outputting control commands when there is a lack of equipment beat evidence, candidate voice boundaries, or effective target voice segments, thereby improving the stability of the method in multi-scenario industrial voice recognition.

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multimodal, multi-scene speech recognition method, characterized in that, Includes the following steps: S1. Acquire acoustic and vibration signals from the industrial production line, perform time synchronization and framing processing on the acoustic and vibration signals to obtain a synchronized multimodal frame sequence; S2. Generate acoustic boundary response values ​​based on the acoustic short-time energy and acoustic spectral flux in the synchronous multimodal frame sequence, and extract a candidate speech boundary set; S3. Generate vibration impact response values ​​based on the short-time vibration energy in the synchronous multimodal frame sequence, extract the equipment action boundary set, and determine the equipment cycle period and cycle reproduction intensity; S4. Calculate the cross-modal pseudo-syllable locking value of each candidate speech boundary based on the candidate speech boundary set, the acoustic boundary response value, the vibration shock response value, the device action boundary set, the device beat period, and the beat reproduction intensity. S5. Correct the acoustic boundary response value according to the cross-modal pseudo-syllable locking value, and determine the target speech segment by combining the speech activity posterior probability obtained through speech activity detection; S6. Perform speech recognition on the target speech segment and output the speech recognition result and recognition credibility.

2. The multimodal, multi-scene speech recognition method according to claim 1, characterized in that, Acquiring acoustic and vibration signals from industrial production lines, including: Acoustic signals formed by the operator's voice, equipment operation sounds, and environmental noise are acquired through a voice acquisition terminal; Vibration signals propagated through the structure are acquired by vibration sensors installed on the equipment frame, workbench, or actuator. The acoustic signal and the vibration signal are aligned according to the same time reference and divided into a synchronous multimodal frame sequence according to the same frame length and the same frame shift. Multiple consecutive synchronous multimodal frames constitute the current recognition window. The acoustic short-time energy, vibration short-time energy, and acoustic spectrum are calculated based on the synchronous multimodal frame sequence.

3. The multimodal, multi-scene speech recognition method according to claim 2, characterized in that, Acoustic boundary response values ​​are generated based on the acoustic short-time energy and acoustic spectral flux in the synchronous multimodal frame sequence, and a candidate speech boundary set is extracted, including: The increase in acoustic energy is obtained by the positive logarithmic change of the acoustic short-time energy of the current frame relative to the acoustic short-time energy of the previous frame. The acoustic spectral flux is obtained by calculating the cumulative positive change of the current frame's acoustic spectral amplitude relative to the previous frame's acoustic spectral amplitude. The acoustic energy rise and the acoustic spectral flux are normalized within the current identification window, and the normalized acoustic energy rise and the normalized acoustic spectral flux are superimposed to obtain the acoustic boundary response value. The local maxima of the acoustic boundary response value are used as candidate positions, and the boundary screening value is determined by the Otsu's method based on the distribution of the acoustic boundary response value within the current recognition window. Candidate positions that are not lower than the boundary screening value are determined as candidate speech boundaries, thus forming a set of candidate speech boundaries.

4. The multimodal, multi-scene speech recognition method according to claim 2, characterized in that, Based on the short-time vibration energy in the synchronous multimodal frame sequence, a vibration impact response value is generated, the equipment action boundary set is extracted, and the equipment cycle period and cycle reproduction intensity are determined, including: The increase in vibration energy is obtained by the positive logarithmic change in the short-time vibration energy of the current frame relative to the short-time vibration energy of the previous frame. The vibration energy rise is normalized within the current identification window to obtain the vibration impact response value; The local maximum location of the vibration and shock response value is used as the candidate device action location. The device boundary screening value is determined by the maximum inter-class variance method according to the distribution of the vibration and shock response value in the current identification window. The candidate device action locations that are not lower than the device boundary screening value are determined as the device action boundary, forming a set of device action boundaries. The vibration impact response value is normalized by autocorrelation processing. The autocorrelation value with the largest value is selected from the autocorrelation values ​​at the non-zero hysteresis position. The hysteresis corresponding to the largest autocorrelation value is determined as the equipment cycle time, and the non-negative value of the largest autocorrelation value is determined as the cycle time recurrence intensity.

5. The multimodal, multi-scene speech recognition method according to claim 4, characterized in that, Calculate the cross-modal pseudo-syllable locking value for each candidate speech boundary, including: For any candidate speech boundary in the candidate speech boundary set, determine the device action boundary that is closest to the candidate speech boundary in time from the device action boundary set. The synchronization proximity is calculated based on the time distance between the candidate speech boundary and the nearest device action boundary, as well as the dispersion of the distances from each candidate speech boundary to the nearest device action boundary within the current recognition window.

6. The multimodal, multi-scene speech recognition method according to claim 5, characterized in that, Calculating the cross-modal pseudo-syllable locking value for each candidate speech boundary also includes: The beat reproduction weight is calculated based on the response intensity of the closest equipment action boundary in the vibration and shock response value, the vibration and shock response value of the equipment action boundary at the position of the previous equipment beat cycle, the vibration and shock response value of the equipment action boundary at the position of the next equipment beat cycle, and the beat reproduction intensity. When the position of the previous device cycle or the position of the next device cycle exceeds the current identification window, the vibration and impact response value corresponding to the position exceeding the window is recorded as zero.

7. A multimodal, multi-scene speech recognition method according to claim 6, characterized in that, Calculating the cross-modal pseudo-syllable locking value for each candidate speech boundary also includes: The vibration dominance is calculated based on the vibration impact response value at the device action boundary with the closest time distance and the acoustic boundary response value at the candidate speech boundary. The acoustic spectral flatness is calculated based on the ratio between the geometric mean and the arithmetic mean of the acoustic spectral power at the candidate speech boundary. The posterior probability of the candidate speech boundary is obtained by using a speech boundary model based on acoustic speech content. The non-speech boundary quantity is obtained by subtracting the posterior probability of the speech boundary, and the synchronization proximity, the beat reproduction weight, the vibration dominance, the acoustic spectrum flatness, and the non-speech boundary quantity are multiplied to obtain the cross-modal pseudo-syllable locking value of the candidate speech boundary.

8. The multimodal, multi-scene speech recognition method according to claim 7, characterized in that, The acoustic boundary response value is corrected based on the cross-modal pseudo-syllable locking value, and the target speech segment is determined by combining the posterior probability of speech activity obtained through speech activity detection, including: The acoustic boundary response value at the corresponding candidate speech boundary is reduced according to the magnitude of the cross-modal pseudo-syllable locking value to obtain the corrected boundary response value; Based on the posterior probability of speech activity obtained through speech activity detection and the corrected boundary response value, a segment score for the candidate speech segment is constructed. The segment score is composed of the logarithmic summation of the posterior probabilities of speech activities in each frame within the candidate speech segment, and the logarithmic summation of the corrected boundary response values ​​at the start and end boundaries of the candidate speech segment. The speech segment with the highest score is selected from the candidate speech segments using dynamic programming, and then used as the target speech segment.

9. A multimodal, multi-scene speech recognition method according to claim 8, characterized in that, The target speech segment is subjected to speech recognition, and the speech recognition result and recognition confidence level are output, including: The acoustic features corresponding to the target speech segment are input into the speech recognition model to obtain candidate instruction texts and their recognition probabilities. The candidate instruction text with the highest recognition probability is determined as the speech recognition result; The recognition confidence level is calculated based on the recognition probability of the speech recognition result and the average value of the cross-modal pseudo-syllable locking value at the start and end boundaries of the target speech segment. Output the speech recognition result and the recognition reliability.

10. A multimodal, multi-scene speech recognition method according to claim 1, characterized in that, The industrial production line includes at least one of automated assembly line, stamping line, packaging line, warehousing and sorting line, or robotic collaborative line; the voice recognition result includes at least one short voice command among start, pause, reset, confirm, clamp, release, lower, and stop.