Rejection method and device, storage medium, equipment and vehicle
Through the multi-modal feature fusion method of audio and video signals, the problem of machine error response noise or non-human machine instructions is solved, and the accuracy and adaptability of speech recognition are improved.
Patent Information
- Application Number
- CN202311843830.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, machines are prone to incorrectly responding to noise or non-human machine instructions in a continuous dialogue scenario between machines and users, resulting in poor interaction experience and task execution errors, and the existing recognition methods are less accurate.
By acoustically encoding and identifying the audio signal, acoustic coded features and semantic text features are extracted, and combined with the visual features of the video signal, input a pre-trained multimodal model for rejection, and fine-tuning the initial multimodal model using sample audio, text query and NLP intention to achieve multimodal feature fusion.
It improves the accuracy and adaptability of speech recognition, reduces the shortcomings and wrong judgments of single modal recognition, and provides more comprehensive and accurate results of rejection.
Smart Images

Figure CN120236569A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a rejection recognition method, apparatus, storage medium, device, and vehicle. Background Art
[0002] In the scenario of continuous conversation between a machine and a user, if the machine incorrectly responds to noise or non-human instructions in the continuous conversation, it will not only have a great negative impact on the interaction experience, but may also cause the machine to execute incorrect instructions, affecting, for example, driving safety.
[0003] In related scenarios, noise or non-human instructions can be excluded based on voice and semantic analysis, such as identifying noise in audio, detecting grammar errors in conversations, understanding semantic meanings in conversations, etc. Its accuracy is relatively low, and in these scenarios, the microphone is always on, and the proportion of non-human queries is very high. If non-human instructions are responded to, it will affect the execution of existing tasks and cause interference to users. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a rejection recognition method, apparatus, storage medium, device, and vehicle.
[0005] According to a first aspect of an embodiment of the present disclosure, a rejection recognition method is provided, including:
[0006] Performing acoustic encoding and recognition on the obtained audio signal respectively to obtain an acoustic encoding feature and a semantic text feature;
[0007] Performing visual feature extraction on the video signal obtained simultaneously with the audio signal to obtain a target visual feature;
[0008] Inputting the target visual feature, the semantic text feature, and the acoustic encoding feature into a pre-trained multi-modal large model to obtain a rejection recognition result output by the multi-modal large model;
[0009] Wherein, the multi-modal large model is obtained by fine-tuning an initial multi-modal large model pre-trained with tokens corresponding to sample audio, text queries, NLP intents, and tokens corresponding to sample videos, and the pre-trained initial multi-modal large model is obtained by pre-training the initial multi-modal large model with the sample audio, text queries corresponding to the sample audio, and NLP intents.
[0010] According to a second aspect of an embodiment of the present disclosure, a rejection recognition apparatus is provided, including:
[0011] An encoding and recognition module, configured to perform acoustic encoding and recognition on the obtained audio signal respectively to obtain an acoustic encoding feature and a semantic text feature;
[0012] An extraction module, configured to extract visual features from a video signal acquired simultaneously with the audio signal to obtain target visual features;
[0013] An input module, configured to input the target visual features, the semantic text features, and the acoustic encoding features into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model;
[0014] Wherein, the multi-modal large model is obtained by fine-tuning an initially pre-trained multi-modal large model with tokens corresponding to a sample audio, a text query, an NLP intention, and tokens corresponding to a sample video, wherein the initially pre-trained multi-modal large model is pre-trained with the sample audio, the text query corresponding to the sample audio, and the NLP intention.
[0015] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method according to any one of the first aspects are implemented.
[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0017] A processor;
[0018] A memory for storing processor-executable instructions;
[0019] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspects.
[0020] According to a fifth aspect of the embodiments of the present disclosure, there is provided a vehicle, including:
[0021] A processor;
[0022] A memory for storing processor-executable instructions;
[0023] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspects.
[0024] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0025] The obtained audio signal is acoustically encoded and recognized to obtain acoustic encoding features and semantic text features; the visual features of the video signal obtained simultaneously with the audio signal are extracted to obtain target visual features; the target visual features, semantic text features, and acoustic encoding features are input into a pre-trained multi-modal large model to obtain a rejection result; the multi-modal large model is fine-tuned from an initially pre-trained multi-modal large model using tokens corresponding to the sample audio, text queries, NLP intents, and tokens corresponding to the sample video, where the initially pre-trained multi-modal large model is pre-trained from the initial multi-modal large model using the sample audio, text queries corresponding to the sample audio, and NLP intents. By performing multi-modal fusion on the target visual features corresponding to the video signal, semantic text features corresponding to the audio signal, and acoustic encoding features of multiple different modalities, a rejection result is obtained, and its processing method is more similar to the human way of processing information, which can not only improve the adaptability and flexibility of speech recognition to the scenario, but also improve the accuracy of speech recognition.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0028] Figure 1 is a flowchart of a rejection method shown according to an exemplary embodiment.
[0029] Figure 2 is a schematic diagram of an implementation scenario of speech recognition shown according to an exemplary embodiment.
[0030] Figure 3 is a schematic diagram of a rejection method shown according to an exemplary embodiment.
[0031] Figure 4 is a flowchart of training a multi-modal large model shown according to an exemplary embodiment.
[0032] Figure 5 is a schematic diagram of pre-training a multi-modal large model shown according to an exemplary embodiment.
[0033] Figure 6 is a schematic diagram of fine-tuning a multi-modal large model shown according to an exemplary embodiment.
[0034] Figure 7 is a schematic diagram of implementing Figure 1 step S13 in
[0035] Figure 8 It is a schematic diagram of a rejection recognition method shown according to an exemplary embodiment.
[0036] Figure 9 It is a schematic diagram of a rejection recognition method shown according to an exemplary embodiment.
[0037] Figure 10 It is a block diagram of a rejection recognition device shown according to an exemplary embodiment.
[0038] Figure 11 It is a block diagram of a device 800 for rejection recognition shown according to an exemplary embodiment.
[0039] Figure 12 It is a block diagram of a device 1900 for rejection recognition shown according to an exemplary embodiment.
[0040] Figure 13 It is a schematic diagram of a functional block diagram of a vehicle shown according to an exemplary embodiment. Detailed implementation
[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0043] Before introducing a rejection recognition method provided by the present disclosure, the technical means in related scenarios will be introduced first. In related scenarios, the text of a voice request and the confidence of the corresponding word output are obtained, confidence features are generated based on the text and the corresponding confidence of the word output, the confidence features of the context are combined to generate target confidence features, and a trained semantic rejection recognition model is used to predict the target confidence features to obtain a rejection recognition result. Thus, through the semantic rejection recognition method, the text and the confidence of the corresponding word output can be used to generate confidence features, and a rejection recognition result can be obtained by predicting based on the confidence features of the above context, improving the accuracy of semantic rejection recognition. Thus, it is possible to avoid using noise or non-human machine instructions as input. However, the post-fusion method adopted in this technical solution is relatively simple. The features of each modality are independently encoded, and only local features of each modality can be captured, with relatively large limitations. Moreover, there are only text and voice modalities. At the same time, the text modality uses pre-trained or shallow neural networks for feature extraction, and the semantic understanding ability is limited. By simply combining acoustic signals in the audio, the characteristics of audio in complex scenarios cannot be distinguished.
[0044] In view of this, the present disclosure provides a rejection recognition method, aiming to more similarly simulate the process of human understanding and processing of different modality data such as text, video, and audio, and automatically align the semantics between different modalities, so as to better understand the semantic associations between different modalities, thereby providing more accurate information processing and analysis, and thus reasonably and flexibly rejecting noise and non-human machine instructions, and thus solving the problem of insufficient rejection recognition ability of a single modality, as well as the limitations of the post-fusion method of simply splicing independent encoders.
[0045] Figure 1 It is a flowchart of a rejection recognition method shown according to an exemplary embodiment. It can be noted that the rejection recognition method provided by the embodiments of the present disclosure is applied to a human-machine dialogue scenario, where the "machine" here can be, for example, terminal devices such as mobile phones, tablet computers, wearable devices, message receiving and sending devices, game consoles, medical devices, robots, and in-vehicle devices of vehicles. However, the rejection recognition method in the embodiments of the present disclosure can be executed either by the above-mentioned terminal devices, that is, the rejection recognition service deployed on the terminal device executes the rejection recognition method in the embodiments of the present disclosure, or by a server communicatively connected to, for example, mobile phones, tablet computers, wearable devices, message receiving and sending devices, game consoles, medical devices, and in-vehicle devices of vehicles, that is, the rejection recognition service deployed on the server executes the rejection recognition method in the embodiments of the present disclosure. As Figure 2As shown, the microphones configured on terminal devices such as mobile phones, tablet computers, wearable devices, messaging devices, game consoles, medical devices, and in-vehicle units of vehicles can acquire audio signals, and the cameras can acquire video signals. Then, the simultaneously acquired audio signals and video signals are uploaded to the cloud server. After the cloud server executes the rejection recognition method provided in this disclosure to obtain the rejection recognition result, it returns to the terminal device, and then the terminal device executes the rejection recognition or user instructions. Refer to Figure 1 As shown, the rejection recognition method in the embodiments of this disclosure includes the following steps.
[0046] In step S11, the acquired audio signal is respectively subjected to acoustic encoding and recognition to obtain an acoustic encoding feature and a semantic text feature.
[0047] In the embodiments of this disclosure, the audio signal can be obtained by performing audio extraction on the audio data collected by the microphone. Furthermore, acoustic encoding can extract the acoustic features of the audio signal, so that the features of the user's voice in many factors such as tone, intonation, volume, and speech rate can be extracted. In one implementation manner, a classification model, that is, model P(y|x), can be trained based on voice features, where y represents the acoustic encoding feature and x is the voice recorded by the microphone. For the final multi-modal fusion model, in model P(y|x), y represents whether to reject recognition. It can be noted that acoustic encoding can use digital signal processing methods, such as Fourier transform, Mel-frequency cepstral coefficients (MFCC), etc., to convert the audio signal into a representation form suitable for machine learning models to process.
[0048] In the embodiments of this disclosure, the acoustic encoding features can include features such as prosodic and spectrum. Among them, the prosodic encoding features can include information on dimensions such as the pitch, intonation, energy, and rhythm change of the user's voice, which is manifested as the "cadence" perceived by the human auditory system. For example, it can include fundamental frequency (F0), speech rate, and energy. The spectrum encoding features can include features of the user's voice in dimensions such as spectrum, power spectrum, cepstrum, and spectral envelope.
[0049] In the embodiments of this disclosure, an acoustic encoder based on SVM, decision tree, or neural network can be used to directly obtain the audio signal from the user's voice, and then extract a real-valued vector representation from the audio signal and input it into the downstream task.
[0050] Furthermore, when recognizing the audio signal, the audio signal can be converted into text form, and then combined with a deep learning model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), etc., to convert the audio signal into an understandable text representation form.
[0051] Refer toFigure 3 As shown, an audio encoder can be used to perform acoustic encoding on the acquired audio signal to obtain acoustic encoding features corresponding to the audio signal, and a text editor can be used to perform text extraction on the acquired audio signal, and then the extracted text features can be recognized to obtain semantic text features corresponding to the audio signal.
[0052] In step S12, visual feature extraction is performed on the video signal acquired simultaneously with the audio signal to obtain target visual features.
[0053] Continue to refer to Figure 3 As shown, a visual editor can be used to perform visual feature extraction on the video signal acquired simultaneously with the audio signal to obtain target visual features.
[0054] In the embodiments of the present disclosure, deep learning models such as convolutional neural networks (CNNs) can be used to extract useful features from video signals. For example, objects such as the lip shape of the user, the orientation of the face when the user is speaking, and whether there are other users in the direction of the face orientation.
[0055] In step S13, the target visual features, the semantic text features, and the acoustic encoding features are input into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model.
[0056] Among them, the multi-modal large model is fine-tuned from an initial pre-trained multi-modal large model using tokens corresponding to sample audio, text queries, NLP intents, and tokens corresponding to sample videos. Among them, the initial pre-trained multi-modal large model is pre-trained from an initial multi-modal large model using the sample audio, text queries corresponding to the sample audio, and NLP intents.
[0057] Among them, the multi-modal large model aligns the target visual features, the semantic text features, and the acoustic encoding features and then performs feature fusion to obtain a rejection result. The rejection result can either represent the speech corresponding to the audio signal as a valid sound or the speech corresponding to the audio signal as an invalid sound.
[0058] Among them, tokens are the smallest text units representing sample audio or sample videos, which are tokens that the multi-modal large model can recognize. Tokens can be represented by words, phrases, numbers, or special characters corresponding to sample audio or sample videos.
[0059] In the embodiments of the present disclosure, the target visual feature, the semantic text feature, and the acoustic coding feature can be combined as features of different modalities, so that the understanding and processing processes of humans for different modality data such as text, video, and audio can be more similarly simulated, which can help the model better understand the meaning of the audio signal, thereby improving the accuracy of speech recognition.
[0060] In the embodiments of the present disclosure, multimodal feature fusion can be performed through, for example, weighted average, feature splicing, attention, or gating mechanism to obtain a rejection result.
[0061] In one implementation, deep learning models such as Transformer can be used for multimodal fusion to achieve better results. Thereby, the long-distance dependencies of features of multiple modalities can be effectively captured, and parallel computing can be performed for multiple segments of user voices, with high computing efficiency.
[0062] In the embodiments of the present disclosure, the rejection result can be defined as a binary classification task based on the query intent and context. After feature fusion, the binary classification task is performed in the classification layer. Among them, the intent is the semantic structured representation obtained after the query is processed by the NLU module, which can be in the form of a list such as domain, intent, slot, and the context includes the user query intent and system reply in the conversation history.
[0063] The above technical solution can perform acoustic coding and recognition on the obtained audio signal to obtain an acoustic coding feature and a semantic text feature; extract visual features from the video signal obtained simultaneously with the audio signal to obtain a target visual feature; input the target visual feature, the semantic text feature, and the acoustic coding feature into a pre-trained multimodal large model for multimodal fusion to obtain a rejection result predicted by the multimodal large model; the multimodal large model is pre-trained by sample audio, a text query corresponding to the sample audio, and NLP intent, and then fine-tuned by tokens corresponding to the sample audio, the text query, NLP intent, and tokens corresponding to the sample video. By performing multimodal fusion on the target visual feature corresponding to the video signal, the semantic text feature corresponding to the audio signal, and features of multiple different modalities of the acoustic coding feature to obtain a rejection result, the processing method is more similar to the human information processing method, which can not only improve the adaptability and flexibility of speech recognition to the scene, but also improve the accuracy of speech recognition. Thereby, the problem of insufficient rejection recognition ability of a single modality and the limitation of the post-fusion method of simply splicing independent encoders can be solved.
[0064] Moreover, by performing multi-modal fusion on the target visual feature, the semantic text feature, and the acoustic encoding feature to obtain a rejection result, a more comprehensive and accurate judgment can be provided to the user, complementing each other and reducing the deficiencies and misjudgments of a single modality.
[0065] See Figure 4 As shown, the multi-modal large model is trained in the following manner:
[0066] In step S31, audio feature extraction is performed on the sample audio to obtain sample audio features.
[0067] In the embodiments of the present disclosure, audio feature extraction can extract features of the user in dimensions such as tone, intonation, volume, and speech rate from the sample audio. In the implementation manner of the present disclosure, audio feature extraction can be performed based on spectral features (such as MFCC), based on time-domain features (such as the short-time energy of the audio), or based on sound events (such as the start / end of the sound).
[0068] Among them, the toolkit for audio feature extraction can be, for example, LIBROSA.
[0069] In step S32, the sample audio features are converted through word segmentation to obtain the tokens corresponding to the sample audio.
[0070] In the embodiments of the present disclosure, see Figure 5 As shown, based on the audio Tokenizer, the sample audio features can be converted through word segmentation. Among them, the audio Tokenizer can be a visual Tokenizer, so that the sample audio features can be converted into a format suitable for model processing. For example, the sample audio features can be segmented into small segments of a fixed size (for example, each small segment is the audio feature of one frame), and then these small segments are converted into a specific word representation form (such as a word or a phrase), that is, the tokens corresponding to the sample audio are obtained, and further the model input representation that can be recognized by the model is obtained.
[0071] In step S33, the text query corresponding to the sample audio is converted through text conversion to obtain the tokens corresponding to the text query.
[0072] Continue to see Figure 5 As shown, the text query Query corresponding to the sample audio is input into the text Tokenizer. Through the text conversion of the text Tokenizer, the tokens corresponding to the text query Query can be obtained, and further the model input representation that can be recognized by the model is obtained.
[0073] In step S34, the tokens corresponding to the sample audio, the tokens corresponding to the text query, and the NLP intent are input into the initial multimodal large model, and the initial multimodal large model is pre-trained to obtain the pre-trained initial multimodal large model.
[0074] Continue to refer to Figure 5 As shown, the NLP intent can be processed by the paraphrasing Adapter to obtain an encoded model input representation that the model can recognize. Then, the tokens corresponding to the sample audio, the tokens corresponding to the text query, and the NLP Figure 3 are used as model inputs and input into the initial multimodal large model for model pre-training to obtain the pre-trained initial multimodal large model.
[0075] In step S35, according to the sample video corresponding to the sample audio, tokens corresponding to the sample video are generated.
[0076] In the embodiments of the present disclosure, tags (Tag) can be added to the sample video, and then the tags (Tag) corresponding to the sample video are converted into a part of the prompt to generate the tokens corresponding to the sample video.
[0077] In step S36, the pre-trained initial multimodal large model is fine-tuned by the tokens corresponding to the sample audio, the tokens corresponding to the text query, the NLP intent, and the tokens corresponding to the sample video to obtain the multimodal large model.
[0078] In the embodiments of the present disclosure, refer to Figure 6 As shown, the tokens corresponding to the sample video are used as the model input of the pre-trained initial multimodal large model, and are input into the pre-trained initial multimodal large model together with the tokens corresponding to the sample audio, the tokens corresponding to the text query, and the NLP Figure 1 to obtain a classification result corresponding to the sample audio. The classification result can be to classify the sample audio into two categories: valid audio and invalid audio, and the pre-trained initial multimodal large model is fine-tuned to obtain the multimodal large model.
[0079] The above technical solution fine-tunes the pre-trained initial multimodal large model through the tokens of the sample video. The trained multimodal large model can provide more comprehensive and accurate judgments, and can complement each other when predicting audio and video, reducing the deficiencies and misjudgments of a single modality.
[0080] In one implementation, in step S35, the generating the tokens corresponding to the sample video according to the sample video corresponding to the sample audio includes:
[0081] Extract video features from the sample video corresponding to the sample audio to obtain sample video features.
[0082] In the embodiments of the present disclosure, key features are extracted from the input sample video to generate corresponding sample video features. Among them, the key features may include optical character recognition (OCR), object detection / tracking (e.g., using models such as YOLO, Faster R-CNN, etc.), scene recognition, person recognition, etc. Extracting key features can improve the understanding and analysis capabilities of the model.
[0083] Input the sample video features into the modality adapter to obtain the tokens corresponding to the sample video output by the modality adapter.
[0084] Continue to refer to Figure 6 As shown, the modality adapter can convert the sample video features into the form of tokens acceptable to the model. The modality adapter can be implemented based on the architecture of Transformer.
[0085] In the embodiments of the present disclosure, the modality adapter can convert the sample video features into a sequence of vectors (i.e., tokens), and these tokens can be used as the model input of the multimodal large model for fine-tuning the multimodal large model.
[0086] Optionally, refer to Figure 7 As shown, in step S13, the inputting of the target visual feature, the semantic text feature, and the acoustic encoding feature into the pre-trained multimodal large model to obtain the rejection result output by the multimodal large model includes:
[0087] In step S131, the semantic text feature and the acoustic encoding feature are fused to obtain a speech-semantic rejection classification.
[0088] In the embodiments of the present disclosure, refer to Figure 8 As shown, the speech can be input into the speech Encoder to obtain the acoustic encoding feature output by the speech Encoder, and the text can be input into the text Encoder to obtain the semantic text feature output by the text Encoder. Then, the acoustic encoding feature output by the speech Encoder and the semantic text feature output by the text Encoder are fused, such as weighted, gate control, etc., to obtain a speech-semantic rejection classification. Jointly modeling the speech and semantic features enables the model to comprehensively utilize the speech and semantic information, learn more sufficient features, and make better decisions.
[0089] Among them, the speech semantic rejection classification can be obtained based on probabilistic logical reasoning, and a given output can be inferred from the input semantic text features and the acoustic coding features. In this step, the semantic text features and the acoustic coding features can be input into different levels of a multimodal large language model (MLLM) for processing, such as the semantic level, the pragmatic level, and the acoustic level, etc., so as to comprehensively understand the meaning and manifestation form of the speech signal.
[0090] In step S132, according to the speech semantic rejection classification, a speech semantic recognition result is obtained.
[0091] The speech semantic rejection classification is a classification that classifies the corresponding audio signal as a recognizable user instruction or an unrecognizable user instruction. And it is represented by probability. The larger the probability value, the higher the possibility of recognition and the higher the confidence; the smaller the probability value, the lower the possibility of recognition and the lower the confidence.
[0092] It can be understood that multiple semantic text features can be recognized from the same audio signal, and each semantic text feature can also have a corresponding confidence. The multimodal model can combine this confidence to obtain the final speech semantic recognition result. The speech semantic recognition result here can be an accurate result of whether to reject recognition.
[0093] In step S133, the semantic features in the speech semantic recognition result, the target visual features, and the semantic text features are input into a pre-trained multimodal large model to obtain a rejection result output by the multimodal large model.
[0094] Continue to refer to Figure 8 As shown, after visually understanding the video signal, target visual features are obtained. By performing large model semantic understanding on the text, semantic features can be obtained. Furthermore, the semantic features in the speech semantic recognition result, the target visual features, and the semantic text features are input into a pre-trained multimodal large model. Through the multimodal decision of the multimodal large model, a rejection result of whether to reject recognition can be obtained.
[0095] In the embodiments of the present disclosure, based on the combination of target visual features, acoustic coding features, and semantic text features, a significant improvement will be obtained in the effect of rejection recognition. For open-domain scenarios, such as the content of chatting with others and chatting with a voice assistant cannot be distinguished. Only based on the acoustic coding features and semantic text features, it is impossible to determine whether the user is chatting with a machine or with a person. Combining the target visual features can effectively solve this technical problem.
[0096] The above technical solution enables the model to better understand the meaning and representation form of the audio signal, improving the accuracy of speech recognition. Multiple features can be fused to improve the accuracy and robustness of speech recognition.
[0097] Optionally, perform acoustic encoding on the obtained audio signal to obtain acoustic encoding features, including:
[0098] Extract a real-valued vector representation for the obtained audio signal to obtain multiple audio feature vectors;
[0099] In the embodiments of the present disclosure, based on LFBE (Linear Frequency-Biased Energy), the audio signal can be converted from the time domain to the frequency domain, and multiple audio feature vectors can be extracted by calculating the energy values in the spectrum.
[0100] Specifically, the audio signal can be subjected to short-time Fourier transform (STFT) to convert the signal from the time domain to the frequency domain. Then, the result of STFT is decomposed by a Mel-filterbank to divide the spectrum into multiple frequency bands. The energy values of each frequency band are weighted, considering the frequency distribution of the signal. Finally, a logarithmic transformation is performed on the weighted energy values to obtain multiple LFBE feature vectors. The LFBE feature vector is the audio feature vector.
[0101] Encode each of the audio feature vectors into an acoustic feature vector of a preset dimension to obtain the acoustic encoding feature.
[0102] In the embodiments of the present disclosure, the audio feature vectors are encoded into acoustic feature vectors of a preset dimension, and the preset dimension is related to the network in the subsequent steps. For example, both the convolutional neural network (CNN) and the self-attention Transformer model have corresponding preset dimensions. In this way, after conversion to the preset dimension, the convolutional neural network (CNN) and the self-attention Transformer model can encode the acoustic feature vectors of the preset dimension to obtain the acoustic encoding feature corresponding to the audio signal.
[0103] Optionally, the extracting a real-valued vector representation for the obtained audio signal to obtain multiple audio feature vectors includes:
[0104] Perform pre-emphasis on the target signal in the audio signal to obtain a pre-emphasized audio signal.
[0105] In the embodiments of the present disclosure, pre-emphasis can enhance the high-frequency part of the audio signal as the target signal, making the energy of the high-frequency part equivalent to that of the low-frequency part, so that the subsequent FFT can better capture the high-frequency information of the signal.
[0106] Frame the pre-emphasized audio signal to obtain a discrete audio signal including a plurality of single audio frames.
[0107] In the embodiments of the present disclosure, the pre-emphasized audio signal is divided into segments of time series, and each segment is called an audio frame. The purpose of framing is to convert the audio signal into a discrete signal for subsequent digital signal processing.
[0108] Add a window function to each audio frame of the discrete audio signal to obtain a windowed audio signal corresponding to the audio frame.
[0109] In the embodiments of the present disclosure, a window function is added to each audio frame of the discrete audio signal. The window function can be, for example, a Hamming window or a rectangular window, so as to reduce the side lobes and ripples of the audio signal and make the audio signal smoother.
[0110] Perform a fast Fourier transform on the windowed audio signal corresponding to each audio frame to convert the windowed audio signal from a time-domain signal to a frequency-domain signal, obtaining an audio frequency-domain signal.
[0111] In the embodiments of the present disclosure, a fast Fourier transform is performed on the windowed audio signal corresponding to each audio frame to convert the windowed audio signal from a time-domain signal to a frequency-domain signal, so as to better reflect the frequency components and energy distribution of the signal.
[0112] Extract Mel frequency cepstral coefficient features from the audio frequency-domain signal to obtain an audio Mel spectrum.
[0113] In the embodiments of the present disclosure, Mel frequency cepstral coefficient features can be extracted through Mel filtering of a non-linear filter, so as to simulate the sound perception mechanism of the human ear, extract Mel frequency cepstral coefficient features, and realize the conversion of the frequency-domain signal into a Mel spectrum. The Mel spectrum can better reflect the frequency and intensity information of the signal.
[0114] Compress the energy values of the Mel frequency cepstral coefficient features in the audio Mel spectrum to obtain a plurality of audio feature vectors, where each audio frame in the audio signal corresponds to one of the audio feature vectors.
[0115] In the embodiments of the present disclosure, a log transformation is performed on the Mel spectrum to convert the energy into a relatively smooth numerical representation, which can better train and classify feature vectors. The energy value of the Mel frequency cepstral coefficient feature in the audio Mel spectrum is compressed to convert the amplitude value of the signal from a linear space to a logarithmic space. Thereby, weak signals in the signal are enhanced, and strong signals in the signal are weakened, so that the signal has better comparability in different amplitude ranges. In this way, the dynamic range of the signal is reduced, making the signal more stable and easier to process. At the same time, the log transformation can also reduce the noise and errors in the signal, thereby improving the signal-to-noise ratio and accuracy of the signal.
[0116] The above technical solution can obtain a feature vector for each frame by performing operations such as pre-emphasis, framing, windowing, FFT, Mel filtering, and log transformation on the audio signal. This feature vector includes the main information of the audio signal and can significantly improve the accuracy and robustness of semantic recognition.
[0117] Optionally, the semantic text features include: multiple types of encoded text features, decoded text features, and NLU semantic text features.
[0118] Identifying the obtained audio signal to obtain semantic text features, including:
[0119] Converting the audio signal into text to obtain multiple different sentences;
[0120] Encoding, decoding, and natural speech understanding are respectively performed on the multiple sentences to generate the encoded text features, the decoded text features, and the NLU semantic text features.
[0121] See Figure 9 As shown, the audio signal is first subjected to ASR feature extraction. Furthermore, based on the ASR features, n sentences with higher confidence can be selected based on the n-best mechanism and input into the text encoder for encoding to obtain ASR encoded text features. In the embodiments of the present disclosure, TextCNN, Transformer, or BERT can be used as the text encoder to encode the n sentences with higher confidence.
[0122] At the same time, the ASR features are decoded to obtain ASR decoded text features, and the semantic representations obtained by inputting the n sentences with higher confidence into the NLU (Natural Language Understanding) module for parsing are obtained as NLU semantic text features.
[0123] Optionally, encoding the multiple sentences to generate the encoded text features corresponding to the audio signal includes:
[0124] Determine the confidence of each of the multiple statements.
[0125] In an embodiment of the present disclosure, calculate the probability of each statement appearing in the audio signal. The probability of each statement appearing in the audio signal can be calculated through a speech recognition (ASR) model to obtain the confidence of the statement.
[0126] Select a target statement from the multiple statements according to the confidence.
[0127] In an embodiment of the present disclosure, based on the n-best mechanism, several statements with the highest confidence can be selected as the target statements.
[0128] Encode the target statement to generate the encoded text feature.
[0129] In an embodiment of the present disclosure, convert the target statement into a text feature, which can be used to describe the semantic content of the audio signal. The audio signal can be converted into text by using a speech recognition (ASR) model.
[0130] Optionally, extracting visual features from the video signal acquired simultaneously with the audio signal to obtain target visual features includes:
[0131] Extract visual features from the acquired video signal to obtain the extracted visual features.
[0132] According to the extracted visual features, determine the target action of the object in the video signal from the predefined actions.
[0133] In an embodiment of the present disclosure, corresponding Tags can be added to the extracted visual features. For example, when the embodiment of the present disclosure is applied to a vehicle, tags such as <video signal, Tag: lip movement, making a call, facing the co-pilot> can be added to the acquired extracted visual features, and then the base that has better performance in action classification (coarse-grained) can be fine-tuned to predict the Tags, so as to determine lip movement, line-of-sight recognition, and dialogue object determination of the object in the video signal.
[0134] Among them, the predefined actions include at least one of the following: lip movement, looking at the screen, making a call, and talking to a dialogue object.
[0135] Determine the target visual features according to the target action.
[0136] In the embodiments of the present disclosure, continuing to take the application of the embodiments of the present disclosure to vehicles as an example for illustration, by extracting visual features from video signals, it is determined whether there is lip movement. If there is no lip movement of the driver, it can be determined that other sound sources are making sounds at this time and no response is required. That is to say, if there is no lip movement of the driver, then the speech semantic recognition result in the foregoing embodiments does not need to be responded to; the determination of the dialogue object can refer to detecting whether there are other speaking objects, such as whether there is a user in the co-driver. The line-of-sight recognition can be to determine whether the speaking user is looking at the screen or looking at others. In the case of a conversation between people, they look at each other and do not look at the screen. During the conversation period, if a proportion of images higher than the threshold is determined to be that the user's line of sight is looking at the screen, then it is determined that the user is looking at the screen, and the rejection confidence is reduced.
[0137] Further, call recognition can be added. When a person talks to a device such as a mobile phone, the device will be placed next to the mouth. During the conversation period, if a frame of image recognizes that the mobile phone is making a call, it is determined that a call is being made, and the rejection result indicates that the voice is an invalid sound.
[0138] The embodiments of the present disclosure also provide a speech rejection device. Refer to Figure 10 As shown, the speech rejection device includes: an encoding and recognition module 1110, an extraction module 1120, and an input module 1130.
[0139] Among them, the encoding and recognition module 1110 is configured to perform acoustic encoding and recognition on the acquired audio signal respectively to obtain acoustic encoding features and semantic text features;
[0140] The extraction module 1120 is configured to extract visual features from the video signal acquired simultaneously with the audio signal to obtain target visual features;
[0141] The input module 1130 is configured to input the target visual features, the semantic text features, and the acoustic encoding features into a pre-trained multi-modal large model to obtain the rejection result output by the multi-modal large model;
[0142] Among them, the multi-modal large model is obtained by fine-tuning the pre-trained initial multi-modal large model with tokens corresponding to the sample audio, text queries, NLP intents, and tokens corresponding to the sample video. Among them, the pre-trained initial multi-modal large model is obtained by pre-training the initial multi-modal large model with the sample audio, text queries corresponding to the sample audio, and NLP intents.
[0143] Optionally, the input module 1130 is further configured to train the multi-modal large model in the following manner:
[0144] Extract audio features from the sample audio to obtain sample audio features;
[0145] Convert the sample audio features through word segmentation to obtain the tokens corresponding to the sample audio;
[0146] Perform text conversion on the text query corresponding to the sample audio to obtain the tokens corresponding to the text query;
[0147] Input the tokens corresponding to the sample audio, the tokens corresponding to the text query, and the NLP intent into the initial multi-modal large model, and pre-train the initial multi-modal large model to obtain the pre-trained initial multi-modal large model;
[0148] Generate the tokens corresponding to the sample video according to the sample video corresponding to the sample audio;
[0149] Fine-tune the pre-trained initial multi-modal large model through the tokens corresponding to the sample audio, the tokens corresponding to the text query, the NLP intent, and the tokens corresponding to the sample video to obtain the multi-modal large model.
[0150] Optionally, the input module 1130 is further configured to:
[0151] Extract video features from the sample video corresponding to the sample audio to obtain sample video features;
[0152] Input the sample video features into the modal adapter to obtain the tokens corresponding to the sample video output by the modal adapter.
[0153] Optionally, the multi-modal large model is configured to:
[0154] Fuse the semantic text features and the acoustic coding features to obtain a voice semantic rejection classification;
[0155] Obtain a voice semantic recognition result according to the voice semantic rejection classification;
[0156] Input the semantic features in the voice semantic recognition result, the target visual features, and the semantic text features into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model.
[0157] Optionally, the encoding and recognition module 1110 is configured to:
[0158] Extract a real-valued vector representation for the obtained audio signal to obtain a plurality of audio feature vectors;
[0159] Encode each of the audio feature vectors into an acoustic feature vector of a preset dimension to obtain the acoustic encoding feature.
[0160] Optionally, the encoding and recognition module 1110 is configured to:
[0161] Perform pre-emphasis on the target signal in the audio signal to obtain a pre-emphasized audio signal;
[0162] Frame the pre-emphasized audio signal to obtain a discrete audio signal including a plurality of single audio frames;
[0163] Add a window function to each audio frame of the discrete audio signal to obtain a window audio signal corresponding to the audio frame;
[0164] Perform a fast Fourier transform on the window audio signal corresponding to each audio frame, convert the window audio signal from a time-domain signal to a frequency-domain signal, and obtain an audio frequency-domain signal;
[0165] Extract mel-frequency cepstral coefficient features from the audio frequency-domain signal to obtain an audio mel-spectrum;
[0166] Compress the energy values of the mel-frequency cepstral coefficient features in the audio mel-spectrum to obtain a plurality of audio feature vectors, where each audio frame in the audio signal corresponds to one of the audio feature vectors.
[0167] Optionally, the semantic text features include: encoded text features, decoded text features, and NLU semantic text features of multiple types;
[0168] The encoding and recognition module 1110 is configured to:
[0169] Convert the audio signal into text to obtain a plurality of different sentences;
[0170] Encode, decode, and perform natural speech understanding on the plurality of sentences respectively to generate the encoded text features, the decoded text features, and the NLU semantic text features.
[0171] Optionally, the encoding and recognition module 1110 is configured to:
[0172] Determine the confidence of each of the plurality of sentences;
[0173] Select a target sentence from the plurality of sentences according to the confidence;
[0174] Encode the target sentence to generate the encoded text feature.
[0175] Optionally, the extraction module 1120 is configured to:
[0176] Extract visual features from the acquired video signal to obtain the extracted visual features;
[0177] Based on the extracted visual features, determine the target action of the object in the video signal from the pre-defined actions;
[0178] Based on the target action, determine the target visual features.
[0179] Optionally, the pre-defined actions include at least one of the following: lip movement, looking at the screen, making a call, talking to a conversation partner.
[0180] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0181] The embodiments of the present disclosure also provide a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method described in any one of the foregoing embodiments are implemented.
[0182] The embodiments of the present disclosure also provide an electronic device, including:
[0183] A processor;
[0184] A memory for storing instructions executable by the processor;
[0185] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method described in any one of the foregoing embodiments.
[0186] It can be noted that the electronic device can be configured as, for example, any one of a mobile phone, a tablet computer, a wearable device, and a vehicle-mounted computer, so as to recognize the user's speech. Thus, when the rejection result indicates that the speech is a valid sound, the speech is recognized, and then the user instruction corresponding to the speech is executed; when the rejection result indicates that the speech is an invalid sound, the speech is not recognized.
[0187] In the embodiments of the present disclosure, the electronic device can be configured as a server, so as to receive audio signals and video signals sent by, for example, a mobile phone, a tablet computer, a wearable device, and a vehicle-mounted computer, obtain a rejection result, and return the rejection result to the peer that uploaded the audio signal and video signal. So that the peer receiving the rejection result can, when the rejection result indicates that the speech is a valid sound, recognize the speech, and then execute the user instruction corresponding to the speech; when the rejection result indicates that the speech is an invalid sound, the speech is not recognized.
[0188] Figure 11FIG. 0 is a block diagram of an apparatus 800 for speech recognition according to an exemplary embodiment. For example, the apparatus 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0189] Referring Figure 11 , the apparatus 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output interface 812, a sensor component 814, and a communication component 816.
[0190] The processing component 802 generally controls the overall operation of the apparatus 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0191] The memory 804 is configured to store various types of data to support the operation of the apparatus 800. Examples of such data include instructions for any application or method operating on the apparatus 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0192] The power component 806 provides power to the various components of the apparatus 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 800.
[0193] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0194] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0195] The input / output interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0196] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0197] The communication component 816 is configured to facilitate communication, either wired or wirelessly, between the device 800 and other devices. The device 800 may access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0198] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described rejection recognition method.
[0199] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 804 including instructions, is also provided. The above instructions may be executed by the processor 820 of the device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, among others.
[0200] Figure 12 is a block diagram of a device 1900 for speech recognition shown according to an exemplary embodiment. For example, the device 1900 may be provided as a server. Referring to Figure 12 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-described rejection recognition method.
[0201] The device 1900 may further include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM, Unix TM , Linux TM , FreeBSD TM or the like.
[0202] Embodiments of the present disclosure also provide a vehicle, including:
[0203] a processor;
[0204] a memory for storing executable instructions executable by the processor;
[0205] wherein the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the foregoing embodiments.
[0206] It can be noted that the vehicle can obtain an audio signal through a microphone in the center console, and at the same time, obtain a video signal through a camera installed in the vehicle. Thus, in the case of communication signal coverage, the audio signal and the video signal can be uploaded to the server, and the server executes the rejection recognition method in the embodiments of the present disclosure. Also, in the case of no communication signal coverage, the vehicle's on-board computer executes the rejection recognition method in the embodiments of the present disclosure.
[0207] Figure 13 is a block diagram of a vehicle 600 shown according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, or may be a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0208] Referring to Figure 13 , the vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, the vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wired or wireless means.
[0209] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, etc.
[0210] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or may be a Beidou system or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.
[0211] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0212] The drive system 640 may include components that provide powered movement for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0213] Some or all of the functions of the vehicle 600 are controlled by the computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.
[0214] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0215] The memory 652 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0216] In addition to the instructions 653, the memory 652 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.
[0217] In an embodiment of the present disclosure, the processor 651 may execute the instructions 653 to complete all or part of the steps of the above-mentioned rejection recognition method.
[0218] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0219] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A rejection recognition method, characterized in that, Including: Performing acoustic encoding and recognition on the obtained audio signal respectively to obtain acoustic encoding features and semantic text features; Performing visual feature extraction on the video signal obtained simultaneously with the audio signal to obtain target visual features; Inputting the target visual features, the semantic text features, and the acoustic encoding features into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model; Wherein, the multi-modal large model is fine-tuned from a pre-trained initial multi-modal large model by using tokens corresponding to the sample audio, text queries, NLP intents, and tokens corresponding to the sample video. Among them, the pre-trained initial multi-modal large model is pre-trained from the initial multi-modal large model by using the sample audio, text queries corresponding to the sample audio, and NLP intents.
2. The method according to claim 1, characterized in that, The multi-modal large model is trained in the following manner: Performing audio feature extraction on the sample audio to obtain sample audio features; Converting the sample audio features through word segmentation to obtain tokens corresponding to the sample audio; Performing text conversion on the text queries corresponding to the sample audio to obtain tokens corresponding to the text queries; Inputting the tokens corresponding to the sample audio, the tokens corresponding to the text queries, and NLP intents into the initial multi-modal large model to pre-train the initial multi-modal large model to obtain a pre-trained initial multi-modal large model; Generating tokens corresponding to the sample video according to the sample video corresponding to the sample audio; Fine-tuning the pre-trained initial multi-modal large model by using the tokens corresponding to the sample audio, the tokens corresponding to the text queries, the NLP intents, and the tokens corresponding to the sample video to obtain the multi-modal large model.
3. The method according to claim 2, wherein The generating the tokens corresponding to the sample video according to the sample video corresponding to the sample audio includes: Performing video feature extraction on the sample video corresponding to the sample audio to obtain sample video features; Inputting the sample video features into a modality adapter to obtain the tokens corresponding to the sample video output by the modality adapter.
4. The method according to claim 1, wherein The inputting the target visual features, the semantic text features, and the acoustic encoding features into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model includes: Fusing the semantic text features and the acoustic encoding features to obtain a speech-semantic rejection classification; Obtaining a speech-semantic recognition result according to the speech-semantic rejection classification; Inputting the speech-semantic recognition result, the target visual features, and the semantic features in the semantic text features into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model.
5. The method according to claim 1, wherein Performing acoustic encoding on the obtained audio signal to obtain acoustic encoding features, including: Extracting a real-valued vector representation for the obtained audio signal to obtain a plurality of audio feature vectors; Encode each of the audio feature vectors into an acoustic feature vector of a preset dimension to obtain the acoustic encoding feature.
6. The method according to claim 5, wherein The extracting a vector representation in real number form from the obtained audio signal to obtain a plurality of audio feature vectors includes: Performing pre-emphasis on a target signal in the audio signal to obtain a pre-emphasized audio signal; Framing the pre-emphasized audio signal to obtain a discrete audio signal including a plurality of single audio frames; Adding a window function to each of the audio frames in the discrete audio signal to obtain a window audio signal corresponding to the audio frame; Performing a fast Fourier transform on the window audio signal corresponding to each audio frame to convert the window audio signal from a time-domain signal to a frequency-domain signal to obtain an audio frequency-domain signal; Performing Mel-frequency cepstral coefficient feature extraction on the audio frequency-domain signal to obtain an audio Mel spectrum; Performing energy value compression on the Mel-frequency cepstral coefficient features in the audio Mel spectrum to obtain a plurality of audio feature vectors, where each audio frame in the audio signal corresponds to one of the audio feature vectors.
7. The method according to claim 1, characterized in that The semantic text features include: encoded text features, decoded text features, and NLU semantic text features; Identifying the obtained audio signal to obtain semantic text features, including: Converting the audio signal into text to obtain a plurality of different sentences; Encoding, decoding, and performing natural speech understanding on the plurality of sentences respectively to generate the encoded text features, the decoded text features, and the NLU semantic text features.
8. The method according to claim 7, wherein Encoding the plurality of sentences to generate the encoded text features, including: Determining the confidence level of each of the plurality of sentences; Selecting a target sentence from the plurality of sentences according to the confidence level; Encoding the target sentence to generate the encoded text features.
9. The method according to any one of claims 1-8, characterized in that The extracting visual features from the video signal obtained simultaneously with the audio signal to obtain target visual features includes: Extracting visual features from the obtained video signal to obtain extracted visual features; Determining a target action of an object in the video signal from predefined actions according to the extracted visual features; Determining the target visual features according to the target action.
10. The method according to claim 9, characterized in that, The predefined actions include at least one of the following: lip movement, looking at the screen, making a phone call, and talking to a conversation partner.
11. A rejection recognition device, characterized in that, including: An encoding and recognition module configured to perform acoustic encoding and recognition on the obtained audio signal respectively to obtain an acoustic encoding feature and semantic text features; An extraction module configured to extract visual features from the video signal obtained simultaneously with the audio signal to obtain target visual features; An input module configured to input the target visual features, the semantic text features, and the acoustic encoding feature into a pre-trained multi-modal large model to obtain a rejection result output by the multi-modal large model; Among them, the multi-modal large model is obtained by fine-tuning the pre-trained initial multi-modal large model with tokens corresponding to the sample audio, text queries, NLP intents, and tokens corresponding to the sample video. Among them, the pre-trained initial multi-modal large model is obtained by pre-training the initial multi-modal large model with the sample audio, text queries corresponding to the sample audio, and NLP intents.
12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-10 are implemented.
13. An electronic device, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Among them, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-10.
14. A vehicle, characterized in that, Including: A processor; A memory for storing processor-executable instructions; Among them, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-10.