Voice annotation method and device, medium and product

By processing speech, video, and physiological signals through multimodal data fusion and feature fusion networks, the problem of low speech annotation accuracy in complex acoustic environments is solved, and high accuracy and robust annotation are achieved in noisy environments.

CN121565149APending Publication Date: 2026-02-24CHINA MOBILE GROUP JIANGSU +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511798217.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in speech annotation under complex acoustic environments. Factors such as background noise, reverberation, and speech rate variations affect the accuracy of annotation, while accents and unclear pronunciations make recognition difficult.

Method used

After acquiring speech, video, and physiological signal data, and performing data preprocessing and feature extraction, feature fusion is performed using a feature fusion network with enhanced gating multimodal fusion mechanism and cross-attention mechanism. Combined with multiple pre-set task decoders, speech-to-text, speaker recognition, emotion recognition, language recognition, and intent understanding tasks are performed.

Benefits of technology

It improves the accuracy of speech annotation in noisy environments, enhances the robustness of the model to various noises and interferences, adapts to complex environments, and has dynamic feedback optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565149A_ABST
    Figure CN121565149A_ABST
Patent Text Reader

Abstract

The invention discloses a voice annotation method and device, a medium and a product, and the method comprises the steps: obtaining to-be-recognized voice information and at least one modal data of synchronously collected video information, text information and physiological signal data; performing data preprocessing and feature extraction on the to-be-recognized voice information and the at least one modal data to obtain a modal data feature of each modal data; inputting the modal data features into a pre-trained feature fusion network based on an enhanced gating multi-modal fusion mechanism and a cross attention mechanism for feature fusion to obtain target fusion features; inputting the target fusion features into a plurality of preset task decoders to obtain a plurality of target task processing results; wherein the target tasks comprise a voice-to-text task, a speaker recognition task, an emotion recognition task, a language recognition task and an intention understanding task. According to the technical scheme of the embodiment of the invention, the voice annotation accuracy and the annotation model robustness can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a voice annotation method, device, medium, and product. Background Technology

[0002] Accurate and efficient annotation of speech data is a crucial prerequisite for subsequent model training, data analysis, and application development. However, current methods that only identify and annotate information in a single modality like speech have limited performance in complex acoustic environments. In particular, background noise, reverberation, and variations in speech rate significantly affect annotation accuracy, while differences in accents and unclear pronunciation can also lead to recognition difficulties. Summary of the Invention

[0003] This invention provides a speech annotation method, device, medium, and product that can improve the accuracy of speech annotation in noisy and complex environments, and enhance the robustness of the annotation model to various noises and interferences.

[0004] In a first aspect, embodiments of the present invention provide a speech annotation method, the method comprising:

[0005] Acquire at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized;

[0006] Data preprocessing and feature extraction are performed on the speech information to be recognized and at least one modal data respectively to obtain the modal data features of each modal data;

[0007] Modal data features are input into a pre-trained feature fusion network based on an enhanced gated multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features;

[0008] The target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; among them, the target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

[0009] Secondly, embodiments of the present invention also provide a voice annotation device, the device comprising:

[0010] The data acquisition module is used to acquire at least one modal data among the speech information to be recognized and video information, text information and physiological signal data collected synchronously with the speech information to be recognized;

[0011] The data preprocessing module is used to perform data preprocessing and feature extraction on the speech information to be recognized and at least one modal data respectively, so as to obtain the modal data features of each modal data.

[0012] The data feature fusion module is used to input modal data features into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features.

[0013] The data decoding module is used to input the target fusion features into multiple preset task decoders to obtain multiple target task processing results; among them, the target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

[0014] Thirdly, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech annotation method as described in any of the embodiments of the present invention.

[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech annotation method as described in any of the embodiments of the present invention.

[0016] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the speech annotation method as described in any of the embodiments of the present invention.

[0017] In this embodiment of the invention, at least one modal data is acquired from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized. Data preprocessing and feature extraction are performed on the speech information to be recognized and the at least one modal data respectively to obtain modal data features for each modal data. The modal data features are input into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism for feature fusion to obtain target fusion features. The target fusion features are input into multiple preset task decoders to obtain multiple target task processing results. The target tasks include speech-to-text tasks, speaker recognition tasks, emotion recognition tasks, language recognition tasks, and intent understanding tasks. The technical solution of this embodiment of the invention solves the problem of low speech annotation accuracy in noisy environments based on single speech data, and can improve the speech annotation accuracy in complex noisy environments, as well as enhance the robustness of the annotation model to various noises and interferences. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart of a speech annotation method provided in an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of the structure of a voice annotation device provided in an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0023] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0024] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the relevant content of the solution.

[0025] Figure 1 This is a flowchart illustrating a speech annotation method provided in an embodiment of the present invention. This embodiment is applicable to scenarios involving the recognition and annotation of speech data. The method can be executed by a speech annotation device, which can be implemented in software and / or hardware and integrated into a computer device.

[0026] like Figure 1 As shown, the speech annotation method includes the following steps:

[0027] S110. Acquire at least one modal data from the speech information to be recognized and video information, text information and physiological signal data collected synchronously with the speech information to be recognized.

[0028] The speech information to be recognized refers to the speech information that needs to be recognized and labeled. This speech information can be collected by a device with audio signal acquisition capabilities. Typically, environmental noise is also collected during the acquisition process. If the environment contains complex noise, it will affect the accuracy of the recognition results.

[0029] In this embodiment, while collecting the voice information to be recognized, one of the following will also be collected simultaneously: video information, text information, and physiological signal data.

[0030] The video information can be the sound source of the speech information to be identified, such as the speaker's video information, which can be obtained by recording the video while the speaker is speaking using an image information acquisition device. The text information can be the speaker's speech or meeting minutes, etc. The physiological signal data can be at least one of the speaker's eye movement signals, electroencephalogram (EEG) signals, or heart rate data.

[0031] For example, in a meeting, the speech to be recognized can be the speaker's voice, the video information can be a video recorded during the speaker's speech, the text information can be the speaker's speech outline or presentation slides, and the physiological signal data can be physiological signals collected by the speaker through a wearable device during the speech, such as heart rate data.

[0032] During each speech annotation process, the speech to be recognized can be associated with data from at least one modality different from the speech modality, and simultaneously input into the model for speech recognition analysis. For example, this could be the speech information to be recognized and video information collected synchronously with it, the speech information to be recognized and text information collected synchronously with it, or the speech information to be recognized and physiological signal data collected synchronously with it. Alternatively, it could be the speech information to be recognized and video and text information collected synchronously with it.

[0033] S120. Perform data preprocessing and feature extraction on the speech information to be recognized and at least one modal data respectively to obtain the modal data features of each modal data.

[0034] Data preprocessing is performed separately on the speech information to be recognized and at least one modality data. The goal of this method is to align and standardize the data of different modalities in time, so that the raw data can be transformed into a unified, clean and efficient model input, while retaining the key features of each modality.

[0035] Data preprocessing of the speech information to be recognized and at least one modality data may include the following steps:

[0036] First, perform data alignment processing on the speech information to be recognized and at least one modal data in the time dimension; second, perform data noise suppression on the time-aligned speech information to be recognized and at least one modal data to obtain denoised speech information to be recognized and at least one modal data; third, perform data normalization processing on the denoised speech information to be recognized and at least one modal data.

[0037] Specifically, in the first step of data alignment processing of the speech information to be recognized and at least one modal data in the time dimension, the timestamps of the speech information to be recognized and at least one modal data can be extracted; then, based on the timestamps, the time offset between the speech information to be recognized and at least one modal data is determined based on the cross-correlation method, and the speech information to be recognized and at least one modal data are aligned in time based on the time offset.

[0038] It's important to note that, firstly, when acquiring data from different modalities, hardware triggering and synchronization control can be implemented through the signal acquisition system corresponding to each modality, enabling synchronous acquisition from multiple sensors. For example, audio can be acquired using a 4-channel MEMS microphone array at a sampling rate of 16kHz and 24-bit precision, triggered by a high-level GPIO (General Purpose Input / Output) signal with rising-edge synchronization. Video can be acquired using a 1080p camera at a frame rate of 30fps; the trigger signal for video acquisition is the frame synchronization signal of the MIPI CSI interface. Physiological signals can be acquired through sensors with a synchronization accuracy of ±2ms.

[0039] In the second step, noise suppression is performed on the speech information to be recognized. An improved Wiener filtering algorithm can be used, expressed as follows: ;in, This represents the denoised speech data. Represents the Wiener filter function. This represents the estimated signal-to-noise ratio. Alternatively, minimum variance distortion-free response and generalized sidelobe cancellers can be used, or the speech information to be recognized can be filtered based on sound source localization.

[0040] When removing noise from video information, text information, and physiological signal data, data cleaning methods corresponding to each modality data can be used for denoising. For video information such as Gaussian filtering to delete redundant frames; for text information such as removing special characters, HTML tags, correcting redundant spaces (e.g., correcting "yin wei" to "because"), and unifying case; for physiological signal data, outliers can be removed, etc.

[0041] After data synchronization and denoising and other processes, feature extraction can be performed on each modality data.

[0042] Specifically, when extracting features from the speech information to be recognized and at least one modality data to obtain the modality data features of each modality data, for the speech data to be recognized, the Mel-Frequency Cepstral Coefficients (MFCC) of the speech data to be recognized can be calculated; and, the speech data to be recognized is filtered through a preset filter bank to obtain the corresponding filter bank features; and, the speech data to be recognized is input into a pre-trained feature extraction learning network to obtain the corresponding deep feature vector of the speech data to be recognized; the MFCC, filter bank features, and deep feature vector are concatenated as the modality data features corresponding to the speech data to be recognized.

[0043] That is, the speech feature extraction steps can be expressed as:

[0044] Step 1: Extract Mel-Frequency Cepstral Coefficients (MFCC). Use an improved Mel-Frequency Cepstral Coefficient algorithm:

[0045] ; where represents the m-th MFCC coefficient, represents the power spectrum of the k-th frequency bin, K represents the total number of frequency bins, and m represents the cepstral coefficient index.

[0046] Step 2: Calculate Fbank features. . Where represents the m-th filter bank feature, represents the frequency domain signal, represents the frequency response of the m-th Mel filter, and N represents the number of points of the Fast Fourier Transform (FFT).

[0047] Step 3: Deep learning feature extraction. ; where represents the deep feature vector, represents the Convolutional Neural Network (CNN) encoder, Represents a spectrogram.

[0048] In the process of extracting features from the speech information to be recognized and at least one modal data to obtain the modal data features of each modal data, the modal data feature extraction for video information can be performed by recognizing lip movement features in the video information through a preset optical flow algorithm; extracting facial expression features in the video information through a pre-trained facial recognition network; and recognizing gesture features in the video information through a pre-trained pose estimation network. The lip movement features, facial expression features, and gesture features are then concatenated as the modal data features of the video information.

[0049] The visual feature extraction steps can be represented as follows:

[0050] Step 1: Extraction of lip movement features. ;in, Indicates lip movement characteristics, This represents the optical flow algorithm. This represents the lip area. Step 2: Facial expression feature extraction. ;in, Indicates facial expression features, This represents a facial recognition network. This represents the facial region. Step 3: Gesture recognition feature extraction. ;in, Indicates gesture characteristics, This represents a pose estimation network. This indicates the hand area.

[0051] In the process of extracting features from the speech information to be recognized and at least one modality of data to obtain the modality features of each modality, the text feature extraction step for the text information can be step 1: word vector embedding. .in, Representing word vectors, This represents the Bidirectional Encoder Representations from Transformers (BERT) embedding function, and Text represents the input text. Step 2: Semantic feature encoding. ;in, Represents semantic features, This indicates a Transformer encoder.

[0052] In the process of extracting features from the speech information to be recognized and at least one modal data to obtain the modal data features of each modal data, the modal data feature extraction steps for physiological signal data may include signal filtering, artifact removal, signal segmentation, and feature extraction through a convolutional neural network.

[0053] S130. Input the modal data features into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features.

[0054] The feature fusion network based on the enhanced gating multimodal fusion mechanism and the cross-attention mechanism includes an input layer, a temporal awareness layer, a modal importance evaluation layer, a gating fusion layer, a cross-modal residual connection layer, a cross-attention layer, and an output layer.

[0055] The speech feature input of the input layer can be represented as: Where T represents the time series length, and dV represents the speech feature dimension (dV=512). The visual feature input of the input layer can be represented as... Where dL represents the visual feature dimension (dL=256). The text feature input of the input layer can be represented as... Where dT represents the text feature dimension .

[0056] The temporal awareness layer is a temporal information feature extraction layer that uses a Long Short-Term Memory (LSTM) network as the temporal memory unit. Its hidden layer has 256 dimensions and two layers. The temporal gating computation process can be represented as follows:

[0057] ;in, , , express Activation function. Wherein, Indicates the timing gating weights, Represents the time-series weight matrix. This indicates the hidden state at the previous moment. This represents the multimodal features at the current moment.

[0058] The modal importance assessment layer is a multilayer perceptron used for importance assessment, with 1024-dimensional input features, 521-dimensional hidden layers, and 3-dimensional output. The network structure is represented as follows: .

[0059] Context information encoding: Generated using a Transformer encoder, with a dimension of 256. The importance evaluation process for each modality's data features can be represented as follows:

[0060] ;

[0061] ;

[0062] .

[0063] in, , , These represent the importance weights of the speech, visual, and text modalities, respectively. Indicates importance assessment network, Indicates contextual information.

[0064] The gated fusion layer is a network used for gated weight calculation, in which, , ; , ; , Its hidden state transition network can be represented as: , ; , ; , .

[0065] The specific implementation of the cross-modal residual connection layer can be represented as follows:

[0066] ;in, It is a three-layer fully connected network: .

[0067] Used to add cross-modal residual connections on top of gated fusion: .in, , Indicates the fusion weight. This represents the cross-modal residual connection function. The fusion weights can be set to... , (Derived through grid search optimization).

[0068] The output layer can be based on a cross-attention mechanism. The cross-attention mechanism is a multi-head cross-attention mechanism, represented as:

[0069] ;

[0070] ;

[0071] .

[0072] Finally, the fusion output is obtained as follows:

[0073] The output feature dimension is .

[0074] The activation function can be a Gaussian Error Linear Unit (GELU).

[0075] S140. Input the target fusion features into multiple preset task decoders to obtain multiple target task processing results.

[0076] The target tasks include speech-to-text conversion, speaker recognition, emotion recognition, language recognition, and intent understanding.

[0077] The preset task decoder can be a Transformer decoder, and the input to the decoder is the fused features extracted in the previous step. The preset task decoder can Encoded as .in, Furthermore, the preset task decoder will then... Decoding Specifically, it can be expressed as: .

[0078] The technical solution of this invention involves acquiring at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information; preprocessing and feature extraction are performed on the speech information to be recognized and the at least one modal data respectively to obtain modal data features for each modal data; the modal data features are input into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism for feature fusion to obtain target fusion features; the target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; wherein, the target tasks include speech-to-text tasks, speaker recognition tasks, emotion recognition tasks, language recognition tasks, and intent understanding tasks. This technical solution solves the problem of low speech annotation accuracy in noisy environments based on single speech data, and can improve speech annotation accuracy in complex noisy environments, as well as enhance the robustness of the annotation model to various noises and interferences.

[0079] Furthermore, during model training and parameter updates, after analysis by the feature fusion network based on the enhanced gating multimodal fusion mechanism and the cross-attention mechanism, and the pre-defined task decoder, speech annotation results can be obtained. Of course, the confidence levels of the annotation results for each task can also be obtained. Then, based on the confidence levels of the annotation results, the parameters of the feature fusion network and the pre-defined task decoder can be updated using reinforcement learning methods.

[0080] Specifically, updating the parameters of the feature fusion network and the preset task decoder based on the strong chemical method is a process of updating the model through a dynamic feedback optimization mechanism based on the model task label results.

[0081] Specifically, the complete implementation process of the dynamic feedback algorithm:

[0082] First, a multidimensional confidence assessment is performed. Based on the prediction results Y_pred of each preset task decoder, the multimodal feature HFusion, and the historical feedback data History (historical reference or data label), the prediction entropy confidence, modal consistency confidence, language model confidence, and historical performance confidence are calculated. The comprehensive confidence result is obtained by weighting the above confidence calculation results.

[0083] Then, the calculated overall confidence level is compared with a preset lower threshold to determine whether to update the parameters. If the calculated overall confidence level is less than the preset lower threshold, then updating the parameters of the feature fusion network and the preset task decoder can be triggered.

[0084] Parameter optimization strategies can include standard gradient descent, adversarial training, and reinforcement learning updates. After updating the parameters, a validation process should be performed to determine whether the updated parameters will bring positive benefits to the model during application.

[0085] Figure 2 This is a schematic diagram of the speech annotation device provided in an embodiment of the present invention. This embodiment is applicable to speech recognition scenarios. The device can be implemented by software and / or hardware and integrated into a computer device.

[0086] like Figure 2 As shown, the speech annotation device includes: a data acquisition module 210, a data preprocessing module 220, a data feature fusion module 230, and a data decoding module 240.

[0087] The system includes a data acquisition module 210, which acquires at least one modality of data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized; a data preprocessing module 220, which performs data preprocessing and feature extraction on the speech information to be recognized and at least one modality of data respectively to obtain modality data features for each modality; a data feature fusion module 230, which inputs the modality data features into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain target fusion features; and a data decoding module 240, which inputs the target fusion features into multiple preset task decoders to obtain multiple target task processing results. The target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

[0088] The technical solution of this invention involves acquiring at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information; preprocessing and feature extraction are performed on the speech information to be recognized and the at least one modal data respectively to obtain modal data features for each modal data; the modal data features are input into a pre-trained feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism for feature fusion to obtain target fusion features; the target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; wherein, the target tasks include speech-to-text tasks, speaker recognition tasks, emotion recognition tasks, language recognition tasks, and intent understanding tasks. This technical solution solves the problem of low speech annotation accuracy in noisy environments based on single speech data, and can improve speech annotation accuracy in complex noisy environments, as well as enhance the robustness of the annotation model to various noises and interferences.

[0089] In one alternative implementation, the data preprocessing module 220 is specifically used for:

[0090] Data alignment processing is performed on the speech information to be recognized and at least one modality data in the time dimension;

[0091] Data noise suppression is performed on the time-aligned speech information to be recognized and at least one modal data to obtain the denoised speech information to be recognized and at least one modal data.

[0092] The noise-reduced speech information to be recognized and at least one modality data are subjected to data normalization processing.

[0093] In an optional implementation, the data preprocessing module 220 can also be used for:

[0094] Extract the timestamps of the speech information to be recognized and at least one modality data;

[0095] Based on the timestamp, the time offset between the speech information to be recognized and at least one modal data is determined using the cross-correlation method, and the speech information to be recognized and at least one modal data are aligned in time based on the time offset.

[0096] In an optional implementation, the data preprocessing module 220 can also be used for:

[0097] Calculate the Mel-frequency cepstral coefficients of the speech data to be recognized; and,

[0098] The speech data to be recognized is filtered using a preset filter bank to obtain the corresponding filter bank features; and...

[0099] The speech data to be recognized is input into a pre-trained feature extraction learning network to obtain the deep feature vector corresponding to the speech data to be recognized.

[0100] Mel frequency cepstral coefficients, filter bank features, and deep feature vectors are concatenated to form the modal data features corresponding to the speech data to be recognized.

[0101] In an optional implementation, the data preprocessing module 220 can also be used for:

[0102] The lip movement features in video information are identified using a pre-defined optical flow algorithm; and...

[0103] Facial expression features are extracted from video information using a pre-trained facial recognition network; and,

[0104] Recognize hand gesture features in video information using a pre-trained pose estimation network;

[0105] Lip movement features, facial expression features, and gesture features are spliced ​​together to form the modal data features of the video information.

[0106] In one alternative implementation, the feature fusion network based on the enhanced gating multimodal fusion mechanism and the cross-attention mechanism includes: an input layer, a temporal awareness layer, a modal importance evaluation layer, a gating fusion layer, a cross-modal residual connection layer, a cross-attention layer, and an output layer.

[0107] In one optional implementation, the speech annotation device further includes a model training device for:

[0108] The parameters of the feature fusion network and the pre-defined task decoder are updated based on the strong chemical method.

[0109] The speech annotation device provided in the embodiments of the present invention can execute the speech annotation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0110] In a specific example, the intelligent customer service system can adopt the voice annotation method described in the above embodiments. Specifically, the microphone used for voice acquisition in the intelligent customer service system is a 4-channel linear array with a sampling rate of 16kHz; the camera parameters for video information acquisition are 1080p, 30fps, and a frame length of 25ms; that is, simultaneous acquisition of 4-channel audio and video.

[0111] First, sound source localization (accuracy ±2°) and MVDR beamforming (target direction gain +8dB) can be performed on the acquired speech data. Then, data alignment preprocessing and multimodal feature extraction can be performed on the acquired speech and video data to obtain data features of the two modalities.

[0112] The obtained data features from the two modalities are input into a feature fusion network based on an enhanced gating multimodal fusion mechanism and a cross-attention mechanism for feature fusion to obtain target fused features. These target fused features are then input into multiple pre-defined task decoders to obtain multiple target task processing results. The Transformer annotation model can include a layer encoder and a 6-layer decoder. Ultimately, annotation results for speech-to-text, speaker recognition, emotion recognition, language recognition, and intent understanding are achieved.

[0113] In addition, the speech annotation method described in the above embodiments can also be applied to smart home devices. Based on the limitations of computing resources of smart home devices, the trained feature fusion network based on the enhanced gating multimodal fusion mechanism and the cross-attention mechanism and the preset task decoder can be lightweighted, which can also achieve accurate speech annotation.

[0114] The speech annotation method provided in this embodiment can significantly improve annotation accuracy based on the fusion of multimodal data features, has higher robustness to adapt to various complex environments and noise conditions, can achieve continuous optimization through dynamic feedback mechanism, can meet the latency requirements of real-time applications, can greatly reduce labor and deployment costs through model lightweight processing, and can support multiple application scenarios and device platforms with higher scalability.

[0115] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 3 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present invention is shown. Figure 3The computer device 12 shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as intelligent controllers and servers, mobile phones, and other terminal devices.

[0116] like Figure 3 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).

[0117] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0118] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0119] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 3 Not shown; usually referred to as a "hard drive"). Although Figure 3 As not shown, disk drives for reading and writing to removable non-volatile disks (e.g., "floppy disks") and optical disc drives for reading and writing to removable non-volatile optical discs (e.g., CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0120] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.

[0121] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 3 As not shown, it can be used in conjunction with computer device 12 with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0122] Processing unit 16 executes various functional applications and data processing by running programs stored in system memory 28, such as implementing the speech annotation method provided in this embodiment, which includes:

[0123] Acquire at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized;

[0124] Data preprocessing and feature extraction are performed on the speech information to be recognized and at least one modal data respectively to obtain the modal data features of each modal data;

[0125] Modal data features are input into a pre-trained feature fusion network based on an enhanced gated multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features;

[0126] The target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; among them, the target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

[0127] This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech annotation method as provided in any embodiment of the present invention, the method comprising:

[0128] Acquire at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized;

[0129] Data preprocessing and feature extraction are performed on the speech information to be recognized and at least one modal data respectively to obtain the modal data features of each modal data;

[0130] Modal data features are input into a pre-trained feature fusion network based on an enhanced gated multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features;

[0131] The target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; among them, the target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

[0132] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0133] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0134] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0135] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0136] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0137] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speech annotation method provided in any embodiment of this application.

[0138] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0139] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

[0140] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A speech annotation method, characterized in that, include: Acquire at least one modal data from the speech information to be recognized and video information, text information, and physiological signal data collected synchronously with the speech information to be recognized; Data preprocessing and feature extraction are performed on the speech information to be recognized and the at least one modal data respectively to obtain the modal data features of each modal data; The modal data features are input into a pre-trained feature fusion network based on an enhanced gated multimodal fusion mechanism and a cross-attention mechanism to perform feature fusion and obtain the target fused features. The target fusion features are input into multiple preset task decoders to obtain multiple target task processing results; wherein, the target tasks include speech-to-text task, speaker recognition task, emotion recognition task, language recognition task, and intent understanding task.

2. The method according to claim 1, characterized in that, The data preprocessing of the speech information to be recognized and the at least one modal data includes: The speech information to be recognized and the at least one modal data are aligned in the time dimension. Data noise suppression is performed on the time-aligned speech information to be recognized and the at least one modal data to obtain the noise-reduced speech information to be recognized and the at least one modal data; The noise-reduced speech information to be identified and the at least one modal data are subjected to data normalization processing.

3. The method according to claim 2, characterized in that, The data alignment process for the speech information to be recognized and the at least one modal data in the time dimension includes: Extract the timestamps of the speech information to be recognized and the at least one modality data; Based on the timestamp, the time offset between the speech information to be recognized and the at least one modal data is determined using a cross-correlation method, and the speech information to be recognized and the at least one modal data are aligned in time based on the time offset.

4. The method according to claim 1, characterized in that, Feature extraction is performed on the speech information to be recognized and the at least one modal data respectively to obtain the modal data features of each modal data, including: Calculate the Mel-frequency cepstral coefficients of the speech data to be recognized; and, The speech data to be recognized is filtered using a preset filter bank to obtain corresponding filter bank features; and... The speech data to be recognized is input into a pre-trained feature extraction learning network to obtain the deep feature vector corresponding to the speech data to be recognized. The Mel frequency cepstral coefficients, the filter bank features, and the depth feature vector are concatenated to form the modal data features corresponding to the speech data to be recognized.

5. The method according to claim 1, characterized in that, The step of extracting features from the speech information to be recognized and the at least one modal data to obtain modal data features for each modal data includes: The lip movement features in the video information are identified using a preset optical flow algorithm; and... Facial expression features are extracted from the video information using a pre-trained facial recognition network; and... Gesture features in the video information are identified using a pre-trained pose estimation network; The lip movement features, facial expression features, and gesture features are concatenated to form the modal data features of the video information.

6. The method according to claim 1, characterized in that, The feature fusion network based on the enhanced gating multimodal fusion mechanism and the cross-attention mechanism includes: The system consists of an input layer, a temporal awareness layer, a modal importance assessment layer, a gated fusion layer, a cross-modal residual connection layer, a cross-attention layer, and an output layer.

7. The method according to claim 1, characterized in that, Also includes: The parameters of the feature fusion network and the preset task decoder are updated based on the strong chemical method.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the speech annotation method as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the speech annotation method as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech annotation method as described in any one of claims 1-7.