A semantic scene generation method and device in a noisy environment and a medium
Patent Information
- Application Number
- CN202310875397.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-07-17
AI Technical Summary
然而,这些方法在复杂的噪声环境中以及不同说话人的口音存在较大差异时,仍然存在一定弊端,并且,最终所识别到的语义场景是否基于上下文、是否具有对话前后逻辑很难得到保证,具有一定的局限性
[0045]对嘈杂环境语义场景中采集到的音频流进行音频处理,能够处理嘈杂环境中的噪声、说话人的口音差异等干扰因素。通过注意力机制模型和长短时记忆模型,能够捕捉音频序列中的上下文信息,保证对话的前后逻辑,从而提高对特定语义场景的识别和理解能力。利用模糊综合评价得到嘈杂环境语义场景中的目标上下文关系数据,通过该目标上下文关系数据,能够让语音交互系统在新场景下更快地适应和学习,生成更为符合用户需求的目标语义场景。
Smart Images

Figure CN116825094B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, specifically to a method, device, and medium for generating semantic scenes in noisy environments. Background Technology
[0002] In the current field of speech recognition and natural language processing, traditional speech recognition and natural language processing algorithms often struggle to accurately identify and understand specific semantic scenarios due to factors such as environmental noise, speaker accents, and background music.
[0003] Currently, the primary approach to recognizing and understanding specific semantic scenarios is through feature extraction and classification methods based on speech signals. For example, by analyzing the spectral, acoustic, and prosodic features of speech signals, different speakers, accents, and speech rates can be distinguished and identified. Simultaneously, machine learning algorithms can be used to classify and cluster large amounts of training data, thereby achieving automatic recognition and understanding of specific semantic scenarios. However, these methods still have certain drawbacks in complex noisy environments and when there are significant differences in accents among different speakers. Furthermore, it is difficult to guarantee whether the ultimately recognized semantic scenario is context-aware and possesses logical coherence within the dialogue, thus exhibiting certain limitations. Summary of the Invention
[0004] To address the aforementioned issues, this application proposes a semantic scene generation method for noisy environments, comprising:
[0005] For noisy environment semantic scenarios with different themes, multiple audio streams are collected for each noisy environment semantic scenario, and audio processing is performed on the multiple audio streams to obtain processed audio sequences; wherein, the audio processing includes filtering, feature extraction and feature classification;
[0006] Based on the attention mechanism model and the long short-term memory model, the contextual relationship data of the audio sequence is obtained, and the corresponding first contextual relationship dataset and second contextual relationship dataset are obtained respectively.
[0007] A fuzzy comprehensive evaluation is performed on the first context relation dataset and the second context relation dataset to obtain a comprehensive context relation dataset corresponding to the noisy environment semantic scene; wherein, the comprehensive context relation dataset includes the comprehensive evaluation results corresponding to the audio stream under each channel;
[0008] Based on the comprehensive evaluation results, target contextual relationship data corresponding to the noisy environment semantic scene is selected from the comprehensive contextual relationship dataset, and the contextual database corresponding to the noisy environment semantic scene is updated based on the target contextual relationship data to obtain the corresponding target semantic scene.
[0009] In one implementation of this application, a fuzzy comprehensive evaluation is performed on the first context relation dataset and the second context relation dataset to obtain a comprehensive context relation dataset corresponding to the noisy environment semantic scene, specifically including:
[0010] Obtain the evaluation index sets corresponding to the first context relationship dataset and the second context relationship dataset respectively, and determine the index weights corresponding to each evaluation index in the evaluation index set;
[0011] The evaluation indicators and their corresponding weights are normalized, and based on the normalized evaluation indicators and weights, the first evaluation result corresponding to the first context relationship dataset and the second evaluation result corresponding to the second context relationship dataset are determined respectively.
[0012] Based on the first evaluation result and the second evaluation result, a comprehensive evaluation result corresponding to the audio stream under each channel is determined, and the comprehensive evaluation results are integrated to obtain a comprehensive contextual relationship dataset corresponding to the noisy environment semantic scene.
[0013] In one implementation of this application, based on an attention mechanism model and a long short-term memory model, contextual relationship data of the audio sequence is obtained, resulting in a corresponding first contextual relationship dataset and a second contextual relationship dataset, specifically including:
[0014] The audio sequence is vectorized to obtain the corresponding audio feature vector, and the audio feature vector is encoded by a preset encoder to determine the first hidden state of the encoder at different first time steps;
[0015] Based on the attention mechanism model, according to the first hidden state, a first probability distribution of different target semantic scene labels corresponding to the audio sequence is determined, and according to the first probability distribution, a first contextual relation dataset corresponding to the audio sequence is obtained;
[0016] The audio sequence is subjected to speech recognition to obtain the corresponding text input sequence; wherein, the text input sequence is composed of audio feature vectors corresponding to different second time steps;
[0017] The text input sequence is input into the Long Short-Term Memory model, and based on the Long Short-Term Memory model, the second hidden state corresponding to the text input sequence at different second time steps is determined.
[0018] Based on the second hidden state, a second probability distribution of different target semantic scene labels corresponding to the audio sequence is predicted, and based on the second probability distribution, a second contextual relationship dataset corresponding to the audio sequence is obtained.
[0019] In one implementation of this application, determining a first probability distribution of different target semantic scene labels corresponding to the audio sequence based on the first hidden state specifically includes:
[0020] For each first time step, the first hidden state corresponding to the previous first time step is input into a preset decoder, so as to generate the decoding state corresponding to the first time step through the decoder;
[0021] The importance of the first hidden state corresponding to the first time step is calculated based on the first hidden state corresponding to different first time steps and the decoding state corresponding to the previous first time step.
[0022] For each first time step, the importance of each first hidden state at the first time step is weighted and summed, and the corresponding weighted summation result is concatenated with the decoding state corresponding to the first time step to obtain the first probability distribution of different target semantic scene labels corresponding to the audio sequence.
[0023] In one implementation of this application, audio processing is performed on the plurality of audio streams to obtain a processed audio sequence, specifically including:
[0024] The multiple audio streams are input into a convolutional neural network. The convolutional neural network performs convolution and downsampling operations on the audio streams, and averages the pooling results obtained after downsampling to obtain a global feature vector.
[0025] A fully connected operation is performed on the global feature vector to obtain a classification result. Based on the classification result, specified audio information of the noisy environment semantic scene is removed from the multiple audio streams. The specified audio information includes at least ambient sound and background noise.
[0026] Feature extraction and feature classification are performed on the audio stream from which the specified audio information has been removed to obtain the processed audio sequence.
[0027] In one implementation of this application, feature extraction and feature classification are performed on the audio stream from which the specified audio information has been removed to obtain a processed audio sequence, specifically including:
[0028] A short-time Fourier transform is performed on the audio stream from which the specified audio information has been removed to obtain the time-frequency information corresponding to the audio stream. The amplitude and phase of the time-frequency information are then classified to obtain the corresponding amplitude spectrum and phase spectrum, respectively.
[0029] The amplitude spectrum is input into the convolutional neural network, and the speech recognition result corresponding to the audio stream is output. Based on the speech recognition result, the audio sequence corresponding to the audio stream is determined.
[0030] In one implementation of this application, the amplitude spectrum is input into the convolutional neural network, and the speech recognition result corresponding to the audio stream is output. Based on the speech recognition result, the audio sequence corresponding to the audio stream is determined, specifically including:
[0031] The amplitude spectrum is input into the convolutional neural network, and the convolutional neural network extracts features from the amplitude spectrum to extract the frequency domain features corresponding to the audio stream;
[0032] The frequency domain features are classified to obtain the speech recognition result corresponding to the audio stream. The speech recognition result is then smoothed using a time suppression algorithm to obtain the processed audio sequence.
[0033] In one implementation of this application, after obtaining the corresponding target semantic scene, the method further includes:
[0034] Acquire audio streams of a specified scene in a noisy semantic scene, and perform audio processing on the specified scene audio streams to obtain a processed specified audio sequence;
[0035] The attention mechanism model is used to perform context recognition on the specified audio sequence and determine the first accuracy corresponding to the first context relationship data identified.
[0036] Using the Long Short-Term Memory model, context recognition is performed on the specified audio sequence, and the second accuracy corresponding to the recognized second context relationship data is determined.
[0037] The first accuracy rate and the second accuracy rate are compared with their corresponding accuracy rate thresholds, respectively, so as to filter out target context relationship data with an accuracy rate greater than the accuracy rate threshold from the first context relationship data and the second context relationship data;
[0038] The answer or rhetorical question slot value corresponding to the specified scene audio stream is determined using the target context relationship data.
[0039] This application provides a semantic scene generation device for noisy environments, the device comprising:
[0040] At least one processor;
[0041] And, a memory communicatively connected to the at least one processor;
[0042] The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a semantic scene generation method in a noisy environment as described above.
[0043] This application provides a non-volatile computer storage medium storing computer-executable instructions, which are configured to execute a semantic scene generation method in a noisy environment as described above.
[0044] The semantic scene generation method proposed in this application can bring the following benefits:
[0045] Audio processing of audio streams acquired in noisy semantic scenes can handle interference factors such as noise and speaker accent differences. Through attention mechanism models and long short-term memory models, contextual information in the audio sequence can be captured, ensuring the logical flow of the dialogue and thus improving the ability to recognize and understand specific semantic scenes. Target contextual relationship data in noisy semantic scenes is obtained using fuzzy comprehensive evaluation. This target contextual relationship data enables the voice interaction system to adapt and learn more quickly in new scenarios, generating target semantic scenes that better meet user needs. Attached Figure Description
[0046] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0047] Figure 1 A flowchart illustrating a semantic scene generation method in a noisy environment, provided as an embodiment of this application;
[0048] Figure 2 This is a schematic diagram of the structure of a semantic scene generation device in a noisy environment, provided as an embodiment of this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0051] like Figure 1 As shown in the embodiment of this application, a semantic scene generation method in a noisy environment includes:
[0052] S101: For noisy environment semantic scenarios with different themes, collect multiple audio streams under each noisy environment semantic scenario, and perform audio processing on the multiple audio streams to obtain the processed audio sequence; among which, audio processing includes filtering, feature extraction and feature classification.
[0053] For noisy environmental semantic scenarios with different themes, multiple video streams under the same theme can be acquired through multiple acquisition channels, and then the acquired video streams can be aggregated. Since noisy environments contain various interfering audio elements besides human voices, such as environmental noise, background noise, and accents, generating specific semantic scenarios in noisy environments requires audio processing of the acquired video streams to obtain a processed audio sequence. The purpose of audio processing is to remove interfering factors from the noisy environment; the process mainly includes filtering, feature extraction, and feature classification. In this embodiment, audio processing is based on an end-to-end CNN convolutional neural network, which can perform functions such as environmental noise reduction, noise sequencing, and human voice classification.
[0054] Specifically, the audio stream from multiple channels is input into a convolutional neural network. The convolutional neural network performs convolution operations and downsampling on the audio stream. During the convolution operation, multiple convolution kernels can be used according to different frequencies, which can be manually adjusted. The specific formula is as follows: Among them, H j W represents the output feature map. i,j This represents the weight of the i-th convolutional kernel, with the asterisk indicating the convolution operation. X i b represents the i-th channel of the input audio stream. jrepresents the bias value, and activation represents the activation function, which can be the ReLU function. Downsampling reduces the amount of data while preserving important feature information in the audio stream. At this point, pooling is required, and average pooling or max pooling can be used. The formula is: P i,j =max m,n H i,j,m,n , where P i,j This represents the pooling result of the j-th feature map in the i-th channel. After convolution and downsampling, the pooling result obtained after downsampling needs to be averaged to obtain the global feature vector. Where V represents the global feature vector and k represents the number of feature maps.
[0055] Furthermore, a fully connected operation is performed on the global feature vector to obtain the classification result. Here, the softmax function can be used to classify the output of the fully connected layer. The classification function is X = softmax(W1V + b), where X represents the classification result, W1 represents the weights of the fully connected layer, and b represents the bias value. Based on the obtained classification result, specified audio information related to noisy environmental semantic scenes can be removed from multiple audio streams. The specified audio information includes at least ambient sound and background noise.
[0056] After filtering the audio stream, a CNN model can be used to extract and classify features from the audio stream after removing specified audio information, thereby obtaining the processed audio sequence.
[0057] Specifically, a short-time Fourier transform is performed on the audio stream X(n) after removing specified audio information to obtain the corresponding time-frequency information X(m,k), where m represents the frequency and k represents the time frame. After obtaining the time-frequency information, the amplitude and phase of the time-frequency information are classified to obtain the corresponding amplitude spectrum |X(m,k)| and phase spectrum ∠X(m,k), respectively. The amplitude spectrum |X(m,k)| is input into a convolutional neural network for feature extraction, which outputs the speech recognition result corresponding to the audio stream X(n). Thus, based on the speech recognition result, the audio sequence T(n) corresponding to the audio stream can be obtained.
[0058] In the feature extraction process, the amplitude spectrum |X(m,k)| is input into a convolutional neural network (CNN). Through convolutional layers, pooling layers, and fully connected layers, the CNN extracts the frequency domain features corresponding to the audio stream. A softmax classifier classifies these frequency domain features to obtain the speech recognition result. After obtaining the speech recognition result, a time suppression algorithm is used to smooth it, ultimately yielding the processed audio sequence T(n). The audio sequence T(n) contains spectral features, acoustic features, prosodic features, features from different speakers, accents, and speech rates.
[0059] S102: Based on the attention mechanism model and the long short-term memory model, obtain the contextual relationship data of the audio sequence, and obtain the corresponding first contextual relationship dataset and second contextual relationship dataset respectively.
[0060] Attention mechanism models and Long Short-Term Memory (LSTM) models can improve the recognition and understanding of specific semantic scenes, capture the contextual information of audio sequences, and facilitate the generation of corresponding semantic scenes in noisy environments.
[0061] In one embodiment, the attention mechanism model can obtain contextual relationship data in an audio sequence in a sequence-to-sequence manner.
[0062] Specifically, the audio sequence T(n) is vectorized, mapping it to a fixed-length vector to obtain the corresponding audio feature vector. Then, a pre-defined encoder, such as LSTM or bidirectional LSTM, is used to encode the audio feature vector to determine the first hidden state of the encoder at different first time steps. For example, when using a BiLSTM as the encoder, the first hidden state can be represented as... in, Indicates the first hidden state, x t Let represent the audio feature vector at the t-th time step in the audio sequence T(n).
[0063] Furthermore, based on the attention mechanism model, the output of the encoder is mapped to the probability distribution of the target semantic scene labels. That is, at each first time step, the decoder, according to the first hidden state... First, determine the first probability distribution of different target semantic scene labels corresponding to the audio sequence. Based on the first probability distribution, the contextual relationship data that satisfies the target semantic scene labels can be determined, thus obtaining the first contextual relationship dataset. The specific implementation process can be carried out through the following steps:
[0064] First, for each first time step, the first hidden state corresponding to the previous first time step is... The input is fed into a preset decoder, which generates the decoding state corresponding to the first time step. This can be expressed as the following formula: in, Indicates the decoding state, y t-1 This indicates the decoding state at the previous time step. This represents the first hidden state of the LSTM in the previous first time step.
[0065] Secondly, based on the first hidden state corresponding to different first time steps The decoding state corresponding to the previous first time step of the first time step. Calculate the importance 'a' of the first hidden state corresponding to the first time step. t,i a t,i This represents the importance of the i-th hidden state of the encoder at the first time step t. The above process can be specifically expressed by the following formula:
[0066]
[0067] Where v is a learnable vector, and W1 and W2 are weight matrices.
[0068] Finally, at different first time steps, the importance of each first hidden state at the first time step is weighted and summed. The weighted sum and the corresponding decoded state at the first time step are then concatenated as vectors. A fully connected layer is then used to map the concatenated result to the first probability distribution of different target semantic scene labels. The formula is shown below:
[0069]
[0070] Among them, o t This represents the weighted summation result of the first hidden state at the current first time step. Let W represent the first probability distribution, W be the weight matrix, and [*,*] denote the vector concatenation operation.
[0071] In one embodiment, the LSTM model can obtain contextual relationship data in an audio sequence by modeling the audio sequence T(n).
[0072] Specifically, speech recognition is performed on the audio sequence T(n), converting T(n) into its corresponding text input sequence S(n). The text input sequence consists of audio feature vectors corresponding to different second time steps, which can be represented as X = [S1, S2, ..., S...]. T ], where S TLet T represent the audio feature vector at the t-th second time step in the text input sequence, where T represents the length of the text input sequence.
[0073] Furthermore, the text input sequence T(n) is input into the LSTM model. Based on the LSTM model, the text input sequence T(n) is modeled to determine the second hidden state corresponding to the text input sequence T(n) at different second time steps. Specifically, this can be represented as h t =LSTM(x t ,h t-1 ), where h t h represents the second hidden state of the LSTM model at the second time step t. t-1 This represents the second hidden state at the second time step t-1.
[0074] The second hidden state can be calculated through the following steps:
[0075] i t =σ(W i x t +U i h t-1 +b i )
[0076] f t =σ(W f x t +U f h t-1 +b f )
[0077] o t =σ(W o x t +U o h t-1 +b o )
[0078]
[0079]
[0080] h t =o t tanh(c t )
[0081] Among them, i t f t o t , These represent the input gate, forget gate, output gate, and cell state, respectively. σ(*) represents the sigmoid function, and tanh(*) represents the bitangent function.
[0082] Furthermore, after determining the second hidden state h t Then, based on the second hidden state, the second probability distribution of different target semantic scene labels corresponding to the audio sequence is predicted, and based on the second probability distribution, the second contextual relation dataset corresponding to the audio sequence is obtained. It should be noted that the second probability distribution can be obtained using the following formula: in, W represents the second probability distribution at the second time step t. y and b y This represents the weights and biases.
[0083] S103: Perform fuzzy comprehensive evaluation on the first context relation dataset and the second context relation dataset to obtain the comprehensive context relation dataset corresponding to the semantic scene of the noisy environment; wherein, the comprehensive context relation dataset includes the comprehensive evaluation results corresponding to the audio stream under each channel.
[0084] After obtaining the first and second context relation datasets, performing fuzzy comprehensive evaluation on them can more accurately and effectively identify the comprehensive context relation dataset PY(n) of semantic scenes in noisy environments.
[0085] The comprehensive contextual relationship dataset PY(n) consists of comprehensive contextual relationship data corresponding to the audio stream under each channel and the corresponding comprehensive evaluation result PY.
[0086] This application provides two evaluation methods: one is to evaluate the first contextual relation dataset obtained by the attention mechanism model, and the other is to evaluate the second contextual relation dataset obtained by the LSTM model. Before evaluation, the set of evaluation metrics A = a1, a2, ..., a... corresponding to each evaluation mode needs to be obtained in advance. n And the weights of each evaluation indicator in the evaluation indicator set.
[0087] Then, the evaluation indicators and their corresponding weights are normalized. The normalized evaluation indicators can be expressed as follows: S ij This represents the original evaluation result of the j-th evaluation indicator in the i-th evaluation method, max(S i ) and min(S i Let represent the minimum and maximum scores of all evaluation indicators in the i-th evaluation method, respectively. The normalized indicator weights can be expressed as:
[0088] After normalizing the evaluation indicators and their weights, the fuzzy comprehensive score for each evaluation indicator can be determined based on the normalized indicators and weights. The fuzzy comprehensive score can be obtained through... The calculation yields the fuzzy comprehensive score for each evaluation indicator. For each evaluation method, this can be achieved through... By summing the weighted scores of all evaluation indicators under this evaluation method, we can obtain the first evaluation result r1 corresponding to the first context relation dataset and the second evaluation result r2 corresponding to the second context relation dataset.
[0089] After obtaining the first evaluation result and the second result, according to The comprehensive evaluation result corresponding to the audio stream in each channel can be determined. Wherein, w i r represents the evaluation weight in the i-th evaluation method. i This represents the evaluation result in the i-th evaluation method.
[0090] Fuzzy comprehensive evaluation can assess the contextual relationship datasets obtained after contextual recognition of audio streams from different channels. The comprehensive evaluation result reflects the accuracy of the contextual relationship datasets corresponding to each channel's audio stream; the higher the score, the better the recognition effect for a specific semantic scene. Therefore, by integrating the comprehensive evaluation results corresponding to audio streams from different channels, a comprehensive contextual relationship dataset corresponding to semantic scenes in noisy environments can be obtained.
[0091] S104: Based on the comprehensive evaluation results, select the target context relation data corresponding to the noisy environment semantic scene from the comprehensive context relation dataset, and update the context database corresponding to the noisy environment semantic scene based on the target context relation data to obtain the corresponding target semantic scene.
[0092] Based on the comprehensive evaluation results, the highest-scoring target contextual relationship data can be selected from the comprehensive contextual relationship dataset with the same theme. This data serves as the contextual relationship for the noisy environment semantic scene, containing contextual slot values, answers, and annotation information. After manual review or automatic proofreading, the target contextual relationship data can be added to the context database corresponding to the noisy environment semantic scene. The context database stores multiple contextual information within the noisy environment semantic scene. Through the context database, the corresponding target semantic scene can be obtained. Furthermore, as the training process progresses, the voice interaction system can adapt and learn more quickly in new scenarios, generating multi-turn slot values and answers that better suit the user's context.
[0093] The above process can be used to train and generate target semantic scenes. Once the model has been trained to a certain extent, it can directly provide context-related audio streams through attention mechanism models and LSTM models.
[0094] Specifically, the process involves acquiring audio streams of a specified scene within a noisy semantic environment and processing them to obtain a processed audio sequence. On one hand, an attention mechanism model is used to perform context recognition on the specified audio sequence, determining the first accuracy rate corresponding to the identified first contextual relationship data. On the other hand, an LSTM model is used to perform context recognition on the specified audio sequence, determining the second accuracy rate corresponding to the identified second contextual relationship data. Both the attention mechanism model and the LSTM model have accuracy thresholds for the contextual relationship data. Therefore, the first accuracy rate is compared with the accuracy threshold of the attention mechanism model, and the second accuracy rate is compared with the accuracy threshold of the LSTM model. This allows for the selection of target contextual relationship data with accuracy rates greater than their corresponding accuracy thresholds from both the first and second contextual relationship data. For target contextual relationship data with high accuracy, further speech recognition is unnecessary; the context-related audio stream can be directly provided. In other words, the answer or question slot value corresponding to the specified scene audio stream can be directly provided using the target contextual relationship data.
[0095] The above are embodiments of the methods proposed in this application. Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.
[0096] Figure 2 This is a schematic diagram of the structure of a semantic scene generation device in a noisy environment, provided as an embodiment of this application. Figure 2 As shown, it includes:
[0097] At least one processor; and,
[0098] At least one processor-communication-connected memory; wherein,
[0099] The memory stores instructions that can be executed by at least one processor, and the instructions, when executed by at least one processor, enable at least one processor to:
[0100] For noisy environment semantic scenarios with different themes, multiple audio streams are collected for each noisy environment semantic scenario, and audio processing is performed on the multiple audio streams to obtain the processed audio sequence; among them, audio processing includes filtering, feature extraction and feature classification.
[0101] Based on the attention mechanism model and the long short-term memory model, the contextual relationship data of the audio sequence is obtained, and the corresponding first contextual relationship dataset and second contextual relationship dataset are obtained respectively.
[0102] A fuzzy comprehensive evaluation is performed on the first and second context relation datasets to obtain a comprehensive context relation dataset corresponding to the semantic scene of the noisy environment; wherein, the comprehensive context relation dataset includes the comprehensive evaluation results corresponding to the audio stream under each channel;
[0103] Based on the comprehensive evaluation results, target contextual relationship data corresponding to noisy environment semantic scenes are selected from the comprehensive contextual relationship dataset. Then, the contextual database corresponding to noisy environment semantic scenes is updated based on the target contextual relationship data to obtain the corresponding target semantic scenes.
[0104] This application provides a non-volatile computer storage medium storing computer-executable instructions, which are configured as follows:
[0105] For noisy environment semantic scenarios with different themes, multiple audio streams are collected for each noisy environment semantic scenario, and audio processing is performed on the multiple audio streams to obtain the processed audio sequence; among them, audio processing includes filtering, feature extraction and feature classification.
[0106] Based on the attention mechanism model and the long short-term memory model, the contextual relationship data of the audio sequence is obtained, and the corresponding first contextual relationship dataset and second contextual relationship dataset are obtained respectively.
[0107] A fuzzy comprehensive evaluation is performed on the first and second context relation datasets to obtain a comprehensive context relation dataset corresponding to the semantic scene of the noisy environment; wherein, the comprehensive context relation dataset includes the comprehensive evaluation results corresponding to the audio stream under each channel;
[0108] Based on the comprehensive evaluation results, target contextual relationship data corresponding to noisy environment semantic scenes are selected from the comprehensive contextual relationship dataset. Then, the contextual database corresponding to noisy environment semantic scenes is updated based on the target contextual relationship data to obtain the corresponding target semantic scenes.
[0109] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.
[0110] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0111] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0115] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0116] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0117] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0118] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0119] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating semantic scenes in noisy environments, characterized in that, The method includes: For noisy environment semantic scenarios with different themes, multiple audio streams are collected for each noisy environment semantic scenario, and audio processing is performed on the multiple audio streams to obtain processed audio sequences; wherein, the audio processing includes filtering, feature extraction and feature classification; Based on the attention mechanism model and the long short-term memory model, the contextual relationship data of the audio sequence is obtained, and the corresponding first contextual relationship dataset and second contextual relationship dataset are obtained respectively. A fuzzy comprehensive evaluation is performed on the first context relation dataset and the second context relation dataset to obtain a comprehensive context relation dataset corresponding to the noisy environment semantic scene; wherein, the comprehensive context relation dataset includes the comprehensive evaluation results corresponding to the audio stream under each channel; Based on the comprehensive evaluation results, target contextual relationship data corresponding to the noisy environment semantic scene is selected from the comprehensive contextual relationship dataset, and the contextual database corresponding to the noisy environment semantic scene is updated based on the target contextual relationship data to obtain the corresponding target semantic scene. Based on the attention mechanism model and the long short-term memory model, the contextual relationship data of the audio sequence is obtained, resulting in the corresponding first contextual relationship dataset and second contextual relationship dataset, specifically including: The audio sequence is vectorized to obtain the corresponding audio feature vector, and the audio feature vector is encoded by a preset encoder to determine the first hidden state of the encoder at different first time steps; Based on the attention mechanism model, according to the first hidden state, a first probability distribution of different target semantic scene labels corresponding to the audio sequence is determined, and according to the first probability distribution, a first contextual relation dataset corresponding to the audio sequence is obtained; The audio sequence is subjected to speech recognition to obtain the corresponding text input sequence; wherein, the text input sequence is composed of audio feature vectors corresponding to different second time steps; The text input sequence is input into the Long Short-Term Memory model, and based on the Long Short-Term Memory model, the second hidden state corresponding to the text input sequence at different second time steps is determined. Based on the second hidden state, a second probability distribution of different target semantic scene labels corresponding to the audio sequence is predicted, and based on the second probability distribution, a second contextual relationship dataset corresponding to the audio sequence is obtained.
2. The semantic scene generation method in a noisy environment according to claim 1, characterized in that, A fuzzy comprehensive evaluation is performed on the first context relation dataset and the second context relation dataset to obtain a comprehensive context relation dataset corresponding to the noisy environment semantic scene, specifically including: Obtain the evaluation index sets corresponding to the first context relationship dataset and the second context relationship dataset respectively, and determine the index weights corresponding to each evaluation index in the evaluation index set; The evaluation indicators and their corresponding weights are normalized, and based on the normalized evaluation indicators and weights, the first evaluation result corresponding to the first context relationship dataset and the second evaluation result corresponding to the second context relationship dataset are determined respectively. Based on the first evaluation result and the second evaluation result, a comprehensive evaluation result corresponding to the audio stream under each channel is determined, and the comprehensive evaluation results are integrated to obtain a comprehensive contextual relationship dataset corresponding to the noisy environment semantic scene.
3. The semantic scene generation method in a noisy environment according to claim 1, characterized in that, Based on the first hidden state, a first probability distribution of different target semantic scene labels corresponding to the audio sequence is determined, specifically including: For each first time step, the first hidden state corresponding to the previous first time step is input into a preset decoder, so as to generate the decoding state corresponding to the first time step through the decoder; The importance of the first hidden state corresponding to the first time step is calculated based on the first hidden state corresponding to different first time steps and the decoding state corresponding to the previous first time step. For each first time step, the importance of each first hidden state at the first time step is weighted and summed, and the corresponding weighted summation result is concatenated with the decoding state corresponding to the first time step to obtain the first probability distribution of different target semantic scene labels corresponding to the audio sequence.
4. The semantic scene generation method in a noisy environment according to claim 1, characterized in that, The multiple audio streams are subjected to audio processing to obtain a processed audio sequence, specifically including: The multiple audio streams are input into a convolutional neural network. The convolutional neural network performs convolution and downsampling operations on the audio streams, and averages the pooling results obtained after downsampling to obtain a global feature vector. A fully connected operation is performed on the global feature vector to obtain a classification result. Based on the classification result, specified audio information of the noisy environment semantic scene is removed from the multiple audio streams. The specified audio information includes at least ambient sound and background noise. Feature extraction and feature classification are performed on the audio stream from which the specified audio information has been removed to obtain the processed audio sequence.
5. The semantic scene generation method in a noisy environment according to claim 4, characterized in that, The audio stream from which the specified audio information has been removed is subjected to feature extraction and feature classification to obtain the processed audio sequence, specifically including: A short-time Fourier transform is performed on the audio stream from which the specified audio information has been removed to obtain the time-frequency information corresponding to the audio stream. The amplitude and phase of the time-frequency information are then classified to obtain the corresponding amplitude spectrum and phase spectrum, respectively. The amplitude spectrum is input into the convolutional neural network, and the speech recognition result corresponding to the audio stream is output. Based on the speech recognition result, the audio sequence corresponding to the audio stream is determined.
6. The semantic scene generation method in a noisy environment according to claim 5, characterized in that, The amplitude spectrum is input into the convolutional neural network, and the speech recognition result corresponding to the audio stream is output. Based on the speech recognition result, the audio sequence corresponding to the audio stream is determined, specifically including: The amplitude spectrum is input into the convolutional neural network, and the convolutional neural network extracts features from the amplitude spectrum to extract the frequency domain features corresponding to the audio stream; The frequency domain features are classified to obtain the speech recognition result corresponding to the audio stream. The speech recognition result is then smoothed using a time suppression algorithm to obtain the processed audio sequence.
7. The semantic scene generation method in a noisy environment according to claim 1, characterized in that, After obtaining the corresponding target semantic scene, the method further includes: Acquire audio streams of a specified scene in a noisy semantic scene, and perform audio processing on the specified scene audio streams to obtain a processed specified audio sequence; The attention mechanism model is used to perform context recognition on the specified audio sequence and determine the first accuracy corresponding to the first context relationship data identified. Using the Long Short-Term Memory model, context recognition is performed on the specified audio sequence, and the second accuracy corresponding to the recognized second context relationship data is determined. The first accuracy rate and the second accuracy rate are compared with their corresponding accuracy rate thresholds, respectively, so as to filter out target context relationship data with an accuracy rate greater than the accuracy rate threshold from the first context relationship data and the second context relationship data. The answer or rhetorical question slot value corresponding to the specified scene audio stream is determined using the target context relationship data.
8. A semantic scene generation device for noisy environments, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a semantic scene generation method in a noisy environment as described in any one of claims 1-7.
9. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: A semantic scene generation method in a noisy environment as described in any one of claims 1-7.
Citation Information
Patent Citations
Voice intention recognition method and device, storage medium and electronic equipment
CN114078477A
Intention recognition method and device, electronic equipment and storage medium
CN114611529A
Answer sequence evaluation
US20160125750A1
Speech recognition method and apparatus based on self-attention mechanism and memory network
WO2022121150A1