Voice control method and device based on artificial intelligence, equipment and storage medium
By preprocessing the collected voice signals and feature extraction of deep learning models, combined with the processing of Transformer encoder and multi-layer perceptron, standardized control instructions are generated, which solves the problems of low accuracy, poor real-time performance and insufficient generalization capabilities in the prior art, and achieves efficient and real-time voice control.
Patent Information
- Application Number
- CN202510219544.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing voice control technologies deal with non-standard voice, especially dysarthritis speech, they have low recognition accuracy, poor real-time performance, and insufficient generalization ability.
Using a speech control method based on artificial intelligence, the collected speech signals are preprocessed, and the Mel frequency cepspectral coefficients and differences are calculated, and converted into a Mel spectrum diagram. Then, the integrated deep learning model is used for feature extraction, attention-weighted fusion and dimensionality reduction, and the fusion feature representation is obtained. Subsequently, the fusion feature representation is blocked, linear projected and position encoding processing, and the encoding sequence is generated and input into the Transformer encoder, and a fixed-length vector representation is obtained through global average pooling. Finally, multi-layer perception machines are used for layered classification, standardized control instructions are generated and sent to smart devices.
Effectively handle non-standard voice, improve recognition accuracy and real-time, enhance the system's generalization capabilities, and is suitable for diversified voice input scenarios.
Smart Images

Figure CN120032647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of device control technology, and in particular to a voice control method, device, equipment and storage medium based on artificial intelligence. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, voice control has been widely used in smart homes, car systems, medical assistance equipment and other fields. This interactive method provides users with a more natural and convenient operation experience, greatly improving the usability of various smart devices. However, existing voice control technology still faces significant challenges in processing non-standard speech, especially dysarthric speech.
[0003] Traditional voice control methods are mainly optimized for standard speech, and often have poor recognition effects on speech with unclear pronunciation, unstable speech speed or accent. This leads to frequent command recognition errors or failure to recognize commands in actual applications, especially for special groups such as the elderly and patients with speech disorders, which seriously affects the user experience and the practicality of the system.
[0004] At present, some improved voice control methods attempt to improve recognition accuracy by increasing model complexity or introducing more voice processing steps. However, these methods often lead to slower system response speeds and are difficult to meet the needs of real-time interaction. In addition, although some methods perform well on certain types of non-standard voices, they lack generalization capabilities when faced with diverse voice inputs and are difficult to adapt to the needs of different users and different usage scenarios. Summary of the invention
[0005] The main purpose of the present invention is to solve the technical problems of low recognition accuracy, poor real-time performance and insufficient generalization ability of existing speech control methods when processing non-standard speech, especially dysarthric speech; A first aspect of the present invention provides a voice control method based on artificial intelligence, the voice control method based on artificial intelligence comprising: Preprocessing the collected speech signal to obtain a short-time frame sequence, and calculating the Mel-frequency cepstrum coefficients and differences of the speech signal according to the short-time frame sequence, and converting the calculation results into a Mel-frequency spectrum; Using a preset first integrated deep learning model and a second integrated deep learning model to extract features from the mel-spectrogram to obtain a joint feature vector, and performing attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; The fused feature representation is sequentially subjected to block processing, linear projection processing, and position encoding processing to obtain a coding sequence, and the coding sequence is input into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; A multi-layer perceptron is used to hierarchically classify the fixed-length vector representation to obtain instruction categories, and key parameters are extracted according to the instruction categories to generate standardized control instructions and send them to corresponding intelligent devices.
[0006] Optionally, in a first implementation of the first aspect of the present invention, the preprocessing of the collected speech signal to obtain a short-time frame sequence, and calculating the Mel-frequency cepstral coefficients and differences of the speech signal according to the short-time frame sequence, and converting the calculation results into a Mel-frequency spectrum diagram includes: Performing noise reduction and frame division processing on the collected speech signal to obtain a short-time frame sequence, and performing pre-emphasis and windowing processing and fast Fourier transform on the short-time frame sequence to obtain a frequency domain representation of the speech signal; Calculating Mel-frequency cepstral coefficients according to the frequency domain representation, and calculating first-order differences and second-order differences of the Mel-frequency cepstral coefficients; The Mel-frequency cepstrum coefficients and the first-order differences and the second-order differences are synthesized into a feature sequence, and the feature sequence is converted into a Mel-frequency spectrum graph.
[0007] Optionally, in a second implementation of the first aspect of the present invention, the use of a preset first integrated deep learning model and a second integrated deep learning model to perform feature extraction on the mel-spectrogram to obtain a joint feature vector, and performing attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation includes: Using a preset first integrated deep learning model to perform local feature extraction on the Mel-spectrogram to obtain a first feature set, wherein the first feature set includes short-term speech features and local spectrum patterns; Using a preset second integrated deep learning model to perform local feature extraction on the Mel-spectrogram to obtain a second feature set, wherein the second feature set includes a long-term dependency and a global speech structure; Zero-padding and concatenating the first feature set and the second feature set to obtain a joint feature vector; The multi-head attention mechanism is used to perform weighted fusion on the joint feature vector to obtain a fusion result, and the principal component analysis method is used to reduce the dimension of the fusion result to obtain a fusion feature representation.
[0008] Optionally, in a third implementation of the first aspect of the present invention, the first integrated deep learning model integrates a VGG16 network, a DenseNet201 network, and a GoogleNet network; The first feature set is obtained by extracting local features from the Mel-spectrogram using a preset first integrated deep learning model, and includes: Perform sliding window segmentation on the Mel-spectrogram to obtain multiple local feature regions, and use the VGG16 network to perform a convolution operation on each local feature region to obtain a first-level feature map; Performing maximum pooling on the first-level feature map to obtain a second-level feature map, and inputting the second-level feature map into a DenseNet201 network, performing feature extraction through a densely connected block in the DenseNet201 network, and obtaining a third-level feature map; The Inception module in the GoogleNet network is used to perform multi-way parallel convolution and pooling operations on the third-level feature map to obtain feature maps of different receptive fields, and the feature maps of different receptive fields are spliced in the channel dimension to obtain a multi-scale feature representation. Perform weighted sum feature aggregation on the multi-scale features to obtain a first feature set.
[0009] Optionally, in a fourth implementation manner of the first aspect of the present invention, the step of sequentially performing block processing, linear projection processing, and position encoding processing on the fused feature representation to obtain a coding sequence, and inputting the coding sequence into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling includes: Dividing the fused feature representation into blocks of equal size to obtain a feature block sequence, and performing linear projection on each feature block in the feature block sequence through a pre-trained weight matrix and a bias vector to obtain a projection feature; Generate a predefined position coding matrix, and add the position coding matrix to the projection feature element by element to obtain a coding sequence containing position information, wherein the number of rows of the position coding matrix is equal to the length of the projection feature, and each row corresponds to an element position in the projection feature; Inputting the encoded sequence into a preset Transformer encoder, processing it through a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder to obtain a context-aware feature sequence; A global average pooling operation is performed on the context-aware feature sequence to calculate the average value of all token representations in the context-aware feature sequence to obtain a vector representation of a fixed length.
[0010] Optionally, in a fifth implementation of the first aspect of the present invention, the inputting the encoded sequence into a preset Transformer encoder, and processing it through a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder to obtain a context-aware feature sequence includes: Inputting the encoding sequence into a preset Transformer encoder, performing multi-head self-attention calculation on the encoding sequence through a multi-head self-attention mechanism in the Transformer encoder, and obtaining an attention-weighted feature representation; Performing a multi-head splicing operation according to the attention-weighted feature representation to obtain a splicing feature, and performing a linear transformation on the splicing feature to obtain a linear transformation result; Inputting the linear transformation result into a feedforward neural network, performing two linear transformations and nonlinear activation function processing on the linear transformation result through the feedforward neural network, and obtaining output features of the feedforward network; Residual connection and layer normalization operations are performed on the output features of the feedforward network to obtain the output result of the current Transformer encoder, and the output result is used as the input of the next Transformer encoder until the last Transformer encoder to obtain a context-aware feature sequence.
[0011] Optionally, in a sixth implementation of the first aspect of the present invention, the step of using a multilayer perceptron to hierarchically classify the fixed-length vector representation to obtain an instruction category, extracting key parameters according to the instruction category, generating a standardized control instruction, and sending the instruction to the smart device includes: Using a preset first-layer multi-layer perceptron to identify the major categories of instructions on the fixed-length vector representation, to obtain the major categories of instructions; According to the major instruction categories, a preset second-layer multi-layer perceptron is used to perform fine-grained instruction recognition on the fixed-length vector representation to obtain a specific instruction category; The specific instruction category is analyzed in combination with the rule-based method and the sequence labeling model, key parameters in the specific instruction category are extracted, and standardized control instructions are generated and sent to the corresponding smart device based on the specific instruction category and the key parameters.
[0012] A second aspect of the present invention provides a voice control device based on artificial intelligence, the voice control device based on artificial intelligence comprising: A preprocessing module is used to preprocess the collected speech signal to obtain a short-time frame sequence, and calculate the Mel-frequency cepstrum coefficients and differences of the speech signal according to the short-time frame sequence, and convert the calculation results into a Mel-frequency spectrum; A feature extraction module, used to extract features from the mel-spectrogram using a preset first integrated deep learning model and a second integrated deep learning model to obtain a joint feature vector, and perform attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; A sequence encoding module is used to sequentially perform block processing, linear projection processing and position encoding processing on the fused feature representation to obtain a coding sequence, and input the coding sequence into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; The instruction recognition module is used to use a multi-layer perceptron to hierarchically classify the fixed-length vector representation to obtain instruction categories, extract key parameters according to the instruction categories, generate standardized control instructions and send them to corresponding smart devices.
[0013] The third aspect of the present invention provides an artificial intelligence-based voice control device, comprising: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected via lines; the at least one processor calls the instructions in the memory so that the artificial intelligence-based voice control device performs the steps of the above-mentioned artificial intelligence-based voice control method.
[0014] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the steps of the above-mentioned artificial intelligence-based voice control method.
[0015] The above-mentioned artificial intelligence-based voice control method, device, equipment and storage medium obtain a mel spectrum graph by preprocessing the collected voice signal; use two preset integrated deep learning models to extract features from the mel spectrum graph to obtain a joint feature vector; perform attention weighted fusion and dimension reduction on the joint feature vector to obtain a fused feature representation; perform block, linear projection and position encoding processing on the fused feature representation to obtain a coding sequence; input the coding sequence into the Transformer encoder, and obtain a fixed-length vector representation through global average pooling; use a multi-layer perceptron for hierarchical classification to obtain instruction categories and key parameters; generate standardized control instructions and send them to smart devices. This method can effectively process non-standard voices, improve recognition accuracy and real-time performance, enhance the system's generalization ability, and is suitable for diversified voice input scenarios.
[0016] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of a first embodiment of a voice control method based on artificial intelligence in an embodiment of the present invention; Figure 2 Schematic diagram of an embodiment of a voice control device based on artificial intelligence in an embodiment of the present invention; Figure 3 The figure is a schematic diagram of an embodiment of a voice control device based on artificial intelligence in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0021] To facilitate understanding of this embodiment, a voice control method based on artificial intelligence disclosed in an embodiment of the present invention is first introduced in detail. Figure 1 As shown, the method comprises the following steps: 101. Preprocess the collected speech signal to obtain a short-time frame sequence, calculate the Mel-frequency cepstrum coefficient and difference of the speech signal according to the short-time frame sequence, and convert the calculation result into a Mel-frequency spectrum diagram; In one embodiment of the present invention, the preprocessing of the collected speech signal to obtain a short-time frame sequence, and calculating the Mel-frequency cepstral coefficients and differences of the speech signal based on the short-time frame sequence, and converting the calculation results into a Mel-frequency spectrum diagram includes: performing noise reduction and frame division processing on the collected speech signal to obtain a short-time frame sequence, and performing pre-emphasis and windowing processing and fast Fourier transform on the short-time frame sequence to obtain a frequency domain representation of the speech signal; calculating the Mel-frequency cepstral coefficients based on the frequency domain representation, and calculating the first-order difference and second-order difference of the Mel-frequency cepstral coefficients; synthesizing the Mel-frequency cepstral coefficients and the first-order difference and the second-order difference into a feature sequence, and converting the feature sequence into a Mel-frequency spectrum diagram.
[0022] Specifically, spectral subtraction or Wiener filtering is first used to reduce noise to reduce the impact of background noise on speech signals. The denoised signal is divided into short time frames of 20-30 milliseconds through framing, with 10-15 milliseconds overlap between each frame. The purpose of framing is to analyze the characteristics of speech signals in a shorter period of time, because speech signals can be regarded as quasi-steady in a short period of time. Then, pre-emphasis processing is performed on each short time frame, and a first-order high-pass filter (y[n]= x[n] - αx[n-1], α is usually 0.95) is used to enhance the high-frequency part and compensate for the high-frequency attenuation of the speech signal. The pre-emphasized signal is windowed, usually using a Hamming window or a Hanning window, in order to reduce spectral leakage. Finally, a fast Fourier transform (FFT) is performed on each frame of the processed signal to convert the time domain signal into a frequency domain representation. The number of FFT points is usually selected as the power of 2 of the frame length, such as 512 or 1024 points. The frequency domain representation obtained after such processing contains the frequency component information of the speech signal in each short time frame.
[0023] Specifically, the Mel frequency cepstral coefficients are calculated according to the frequency domain representation, and the first-order difference and second-order difference of the Mel frequency cepstral coefficients are calculated. First, the linear spectrum is mapped to the Mel frequency scale using the formula mel(f) = 2595* log10(1 + f / 700). This mapping is closer to the auditory perception of the human ear. Then a Mel filter bank (usually 20-40 triangular filters) is applied to filter the power spectrum. The logarithm of the filtered result is taken, and then a discrete cosine transform (DCT) is performed, and the first 12-20 coefficients are usually retained as MFCCs. The calculation process of MFCC simulates the characteristics of the human auditory system and can effectively represent the key acoustic features of the speech signal. Then the first-order difference (delta) and second-order difference (delta-delta) of MFCC are calculated. The calculation formula of the first-order difference is dt = (ct+1 - ct-1) / 2, where ct is the MFCC at time t. The second-order difference is to apply the same formula to the first-order difference again. These dynamic features reflect the changes of MFCC over time and help capture the timing information of speech signals, especially for non-standard speech with unclear pronunciation or unstable speaking speed, these dynamic features can capture subtle changes.
[0024] Specifically, the Mel frequency cepstrum coefficients and the first-order differences and second-order differences are synthesized into a feature sequence, and the feature sequence is converted into a Mel frequency spectrogram. First, MFCC and its first-order and second-order differences are synthesized into a feature sequence. Assuming there are 13-dimensional MFCCs, plus 13-dimensional first-order differences and 13-dimensional second-order differences, a 39-dimensional feature vector is obtained. For each frame, these 39-dimensional feature vectors are connected to form a feature sequence. This feature sequence combines the static and dynamic features of the speech signal and provides rich acoustic information. Then, this feature sequence is converted into a Mel frequency spectrogram. The conversion process usually uses a short-time Fourier transform (STFT). First, the original signal is subjected to STFT to obtain a time-frequency representation of a complex value. Then, the amplitude of each time-frequency point is calculated, and the linear frequency axis is converted to a Mel frequency axis. Finally, the result is logarithmic to obtain a Mel frequency spectrogram. The Mel frequency spectrogram is a two-dimensional array, the horizontal axis represents time, the vertical axis represents Mel frequency, and the numerical value represents the energy of the corresponding time-frequency point. This representation intuitively displays the frequency components and energy distribution of the speech signal at different time points, enabling deep learning models to extract speech features more effectively, especially for the processing of non-standard speech.
[0025] 102. Use the preset first integrated deep learning model and the second integrated deep learning model to extract features from the mel-spectrogram to obtain a joint feature vector, and perform attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; In one embodiment of the present invention, the method of using a preset first integrated deep learning model and a second integrated deep learning model to extract features from the mel-spectrogram to obtain a joint feature vector, and performing attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation includes: using a preset first integrated deep learning model to extract local features from the mel-spectrogram to obtain a first feature set, wherein the first feature set includes short-term speech features and local spectral patterns; using a preset second integrated deep learning model to extract local features from the mel-spectrogram to obtain a second feature set, wherein the second feature set includes long-term dependencies and a global speech structure; zero-padding and concatenating the first feature set and the second feature set to obtain a joint feature vector; using a multi-head attention mechanism to perform weighted fusion on the joint feature vector to obtain a fusion result, and using a principal component analysis method to reduce the dimensionality of the fusion result to obtain a fused feature representation.
[0026] Specifically, the Mel spectrum map is subjected to local feature extraction using a preset first integrated deep learning model to obtain a first feature set. The first integrated deep learning model uses a combination of VGG16, GoogleNet and ResNet50. First, the Mel spectrum map is input into the VGG16 network and processed by 5 convolution blocks. Each convolution block contains 2-3 3x3 convolution layers, followed by a ReLU activation function, and then a 2x2 maximum pooling layer. The output of the last convolution block of VGG16 is passed to GoogleNet. GoogleNet uses 9 Inception modules, each of which uses 1x1, 3x3, 5x5 convolutions and 3x3 maximum pooling in parallel, and then concatenates the results. This structure allows the network to extract features at multiple receptive field sizes. ResNet50 contains multiple residual blocks, each of which has three layers of convolution (1x1, 3x3, 1x1) and a shortcut connection. Through residual learning, ResNet50 can effectively extract deeper features. Finally, the outputs of the three networks are globally averaged and pooled to compress the feature maps into one-dimensional vectors, which are then concatenated to form the first feature set. This feature set contains local time-frequency features of different scales, covering short-term pronunciation features and local spectral patterns.
[0027] The preset second integrated deep learning model is used to extract features from the Mel spectrogram, obtaining a second feature set. The second integrated deep learning model consists of DenseNet201, InceptionResNetV2, and Xception. DenseNet201 contains 4 dense blocks, and each layer within each dense block is directly connected to all previous layers. Specifically, the input of each layer is the concatenation of the outputs of all previous layers, and the number of output channels is usually 32 (growth rate). This dense connection structure strengthens feature reuse and alleviates the problem of gradient disappearance. InceptionResNetV2 combines the Inception structure and residual connections, containing 35 Inception-ResNet-A modules, 17 Inception-ResNet-B modules, and 11 Inception-ResNet-C modules. Inside each module, different-sized convolutional kernels are used to process the input in parallel, and then the results are added to the shortcut connection. The Xception network uses 36 depthwise separable convolutional blocks, and each block contains a pointwise convolution and a depth convolution. This structure separately learns the channel-wise correlation and spatial correlation, improving the computational efficiency. The outputs of the three networks are concatenated after dimensionality reduction by global average pooling, forming the second feature set, which contains the long-term dependence relationship and global structure information of the speech signal.
[0028] Zero-padding and concatenation are performed on the first feature set and the second feature set to obtain a joint feature vector. First, the dimensions of the two feature sets are determined. Assume the dimension of the first feature set is m and the dimension of the second feature set is n. If m < n, zero-padding is performed on the first feature set. The specific operation is to create a zero vector of length n, and then copy the m elements of the first feature set to the first m positions of this zero vector. If m > n, a similar zero-padding operation is performed on the second feature set. Zero-padding ensures that the two feature sets have the same dimension without introducing additional information interference. After padding, the two feature sets are concatenated in the feature dimension. The concatenation operation treats the two feature sets as two column vectors and then stacks them vertically to form a new longer column vector. This new column vector is the joint feature vector, and its dimension is m + n. The joint feature vector combines the outputs of the two integrated models, containing both the short-term local features extracted by VGG16, GoogleNet, and ResNet50, and the long-term global features extracted by DenseNet201, InceptionResNetV2, and Xception, providing a rich information basis for subsequent feature fusion.
[0029] The joint feature vector is weightedly fused using a multi-head attention mechanism to obtain a fusion result, and the fusion result is reduced in dimension using the principal component analysis method to obtain a fused feature representation. The multi-head attention mechanism first linearly projects the joint feature vector into the query, key, and value space. Specifically, three different weight matrices Wq, Wk, and Wv are used to transform the input vector X into query Q, key K, and value V, respectively. Then the attention weight is calculated: Attention(Q,K,V) = softmax(QKᵀ / √dk)V, where dk is the dimension of the key vector. Eight attention heads are used, each head calculates the attention score independently, and then the results of the eight heads are concatenated and the final attention output is obtained through a linear transformation. This mechanism allows the model to focus on different aspects of the features at the same time. Then, the principal component analysis (PCA) method is used to reduce the dimension of the fusion result. PCA first calculates the covariance matrix of the data, and then performs eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and corresponding eigenvectors. The eigenvectors corresponding to the first k largest eigenvalues are selected to form a projection matrix. Usually, k is chosen so that the retained principal components can explain 95% of the data variance. Finally, the fusion result is multiplied by the projection matrix to obtain the fusion feature representation after dimensionality reduction.
[0030] Furthermore, the first integrated deep learning model integrates a VGG16 network, a DenseNet201 network and a GoogleNet network; the use of the preset first integrated deep learning model to perform local feature extraction on the Mel spectrum map to obtain a first feature set includes: performing sliding window segmentation on the Mel spectrum map to obtain multiple local feature areas, and using the VGG16 network to perform a convolution operation on each local feature area to obtain a first-level feature map; performing maximum pooling on the first-level feature map to obtain a second-level feature map, and inputting the second-level feature map into the DenseNet201 network, performing feature extraction through the densely connected blocks in the DenseNet201 network to obtain a third-level feature map; using the Inception module in the GoogleNet network to perform multi-way parallel convolution and pooling operations on the third-level feature map to obtain feature maps of different receptive fields, and splicing the feature maps of different receptive fields in the channel dimension to obtain a multi-scale feature representation; performing weighted sum feature aggregation on the multi-scale features to obtain the first feature set.
[0031] Specifically, the first integrated deep learning model integrates the VGG16 network, the DenseNet201 network, and the GoogleNet network. This combination takes advantage of the different advantages of the three networks to extract local features of the mel-spectrogram. First, the mel-spectrogram is segmented using a sliding window to obtain multiple local feature regions. The size of the sliding window is usually set to 64x64 pixels with a step size of 32 pixels, which ensures sufficient coverage of the mel-spectrogram while retaining local structural information. For each local feature region, a convolution operation is performed using the VGG16 network. The VGG16 network contains 13 convolutional layers, divided into 5 convolutional blocks. Each convolutional block contains 2-3 3x3 convolutional layers followed by a ReLU activation function. The step size of each convolution operation is 1 and the padding is 1 to maintain the spatial size of the feature map. Through these convolution operations, the VGG16 network extracts low-level features of the local feature region to form a first-level feature map.
[0032] Specifically, the first-level feature map is subjected to a maximum pooling operation to obtain a second-level feature map. The maximum pooling uses a 2x2 window with a step size of 2, which can halve the spatial dimension of the feature map while retaining the most significant features. The pooling operation not only reduces the computational complexity, but also increases the translation invariance of the features. The second-level feature map is input into the DenseNet201 network, and features are extracted through its densely connected blocks to obtain the third-level feature map. DenseNet201 contains 4 dense blocks, and each layer in each dense block is directly connected to all previous layers. Specifically, the input of each layer is the concatenation of the outputs of all previous layers, and the number of output channels is usually 32 (growth rate). The densely connected structure promotes feature reuse and gradient flow, enabling the network to learn richer feature representations. Between each dense block, a transition layer (1x1 convolution and 2x2 average pooling) is used to reduce the spatial size and number of channels of the feature map. This structure of DenseNet201 enables the network to extract more complex feature patterns while maintaining computational efficiency.
[0033] Specifically, the Inception module in the GoogleNet network is used to perform multi-way parallel convolution and pooling operations on the third-level feature map to obtain feature maps with different receptive fields. The Inception module contains four parallel branches: 1x1 convolution, 1x1 convolution followed by 3x3 convolution, 1x1 convolution followed by 5x5 convolution, and 3x3 maximum pooling followed by 1x1 convolution. This structure allows the network to extract features at multiple scales simultaneously, enhancing the diversity of feature extraction. 1x1 convolution is used to reduce the number of channels and reduce the amount of computation. 3x3 and 5x5 convolutions capture local patterns of different sizes, while the maximum pooling branch retains significant features. These feature maps with different receptive fields are spliced in the channel dimension to form a multi-scale feature representation. The splicing operation retains the different types of features extracted by each branch, so that the final feature representation contains rich multi-scale information.
[0034] Specifically, weighted sum feature aggregation is performed on the multi-scale feature representation to obtain the first feature set. The weight of the weighted sum can be determined by an attention mechanism or a learnable parameter. In the specific implementation, a 1x1 convolution is first applied to each feature map in the multi-scale feature representation to map it to the same number of channels. Then, the global average pooling value of each feature map is calculated as an indicator of its importance. These indicators are normalized using the softmax function to obtain the weight of each feature map. Finally, the weight is multiplied by the corresponding feature map, and all weighted feature maps are added to obtain the final first feature set. This weighted summation method allows the model to adaptively fuse features of different scales and highlight important feature information. The first feature set combines the local feature extraction capabilities of VGG16, the multi-scale feature reuse capabilities of DenseNet201, and the multi-path parallel feature extraction capabilities of GoogleNet.
[0035] 103. The fused feature representation is sequentially subjected to block processing, linear projection processing, and position encoding processing to obtain a coding sequence, and the coding sequence is input into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; In one embodiment of the present invention, the fused feature representation is sequentially subjected to block processing, linear projection processing and position encoding processing to obtain a coding sequence, and the coding sequence is input into a preset Transformer encoder, and a fixed-length vector representation is obtained through global average pooling, including: equal-sized blocks are performed on the fused feature representation to obtain a feature block sequence, and linear projection is performed on each feature block in the feature block sequence through a pre-trained weight matrix and a bias vector to obtain a projection feature; a predefined position encoding matrix is generated, and the position encoding matrix is added element by element to the projection feature to obtain a coding sequence containing position information, wherein the number of rows of the position encoding matrix is equal to the length of the projection feature, and each row corresponds to an element position in the projection feature; the coding sequence is input into a preset Transformer encoder, and processed through a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder to obtain a context-aware feature sequence; a global average pooling operation is performed on the context-aware feature sequence, and the average value of all token representations in the context-aware feature sequence is calculated to obtain a fixed-length vector representation.
[0036] Specifically, the fused feature representation is divided into blocks of equal size to obtain a feature block sequence, and each feature block in the feature block sequence is linearly projected through the pre-trained weight matrix and bias vector to obtain the implementation process of the projection feature. The fused feature representation is first converted into a two-dimensional matrix form through a reshape operation, where each row corresponds to a feature block. The size of the feature block is usually set to 16×16 or 32×32, depending on the total size of the original fused feature representation and the expected sequence length. The purpose of this step is to standardize feature blocks of different sizes to a uniform size to ensure that subsequent processing steps are performed on a consistent basis. Next, using the pre-trained weight matrix and the bias vector Perform a linear transformation on each feature block. Specifically, the feature block Projection is performed using the following linear transformation formula: ; in, The dimension is Feature block size, The dimension is ,and It is usually set to 512 or 768 to match the hidden layer size of the Transformer model. Through this linear projection operation, feature blocks of different sizes are mapped to a unified high-dimensional space, which not only enhances the expressive power of the features, but also provides a solid foundation for subsequent feature fusion and context understanding. The result of linear projection The key information of each feature block is retained while converting it into a format suitable for high-dimensional processing, thereby improving the performance and accuracy of the overall model. The implementation process of generating a predefined position encoding matrix and adding the position encoding matrix to the projected features element by element to obtain a coding sequence containing position information involves effectively injecting position information into the feature sequence so that the model can recognize the relationship between different positions. First, the position encoding matrix is generated by the following sine and cosine functions: ; ; Among them, pos represents the position, represents the dimension. The dimension of the generated position encoding matrix is seq length ,in seq_length is usually set to 512 or 1024 to cover the maximum sequence length in most application scenarios. Then, the position encoding matrix is combined with the projected feature Add element by element to get the coded sequence containing position information ; This operation ensures that each feature block not only contains its own information, but also carries its position information in the sequence, so that the subsequent Transformer encoder can use this position information for more accurate context perception and feature fusion. In this way, the model can not only understand the independent information of each feature block, but also effectively capture the relative position relationship between feature blocks, significantly improving the overall feature representation capabilities and the model's understanding depth. The encoded sequence is input into the preset Transformer encoder, and processed by the multi-head self-attention mechanism and feedforward neural network in the Transformer encoder to obtain the context-aware feature sequence. The implementation process first converts the position-encoded feature sequence into The input is fed into a Transformer encoder with multiple layers. Each Transformer encoder layer consists of a multi-head self-attention sublayer and a feed-forward neural network sublayer. In the multi-head self-attention mechanism, the input feature sequence is first linearly projected onto the query ,key Sum Space, the attention weight is calculated by the following formula: ; in, is the dimension of the key vector. Multiple attention heads (usually 8) calculate their own attention weights in parallel, and then concatenate the results and obtain the final output through linear transformation. Next, the feedforward neural network sublayer performs a nonlinear transformation on the features output by the self-attention mechanism, which is specifically implemented by the following formula: ; Layer normalization and residual connection are applied after each sub-layer to ensure the effective transfer of gradients and the stability of the model. Through multi-layer iterative processing, the Transformer encoder can gradually extract and fuse the contextual information in the feature sequence to generate a feature sequence with high context awareness. In the end, the feature sequence after processing by multiple encoder layers not only contains the local information of each feature block, but also fuses the global context association to form a rich and compact feature representation, which provides a strong feature foundation for subsequent classification or regression tasks. Perform a global average pooling operation on the context-aware feature sequence, calculate the average value of all token representations in the sequence, and obtain a fixed-length vector representation. The implementation process first represents the context-aware feature sequence output by the Transformer encoder as a matrix , whose dimension is [seq length, The global average pooling is performed by Sum the sequence length dimension and divide it by the sequence length to calculate the average value in each dimension. The formula is as follows: ; in, is the final fixed-length vector representation with dimension The purpose of this operation is to compress the variable-length sequence information into a fixed-length vector, retaining the global information and feature distribution of the entire sequence. Global average pooling not only effectively reduces the dimension of the data and reduces the computational complexity, but also avoids the inconsistency problem caused by the change in sequence length, ensuring that the model has consistent output feature representation under different input lengths. In this way, the obtained fixed-length vector v integrates the feature information of all positions in the sequence to form a highly generalized and information-rich global feature representation.
[0037] Furthermore, the step of inputting the coding sequence into a preset Transformer encoder, processing the coding sequence through a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder, and obtaining a context-aware feature sequence includes: inputting the coding sequence into a preset Transformer encoder, performing multi-head self-attention calculation on the coding sequence through the multi-head self-attention mechanism in the Transformer encoder, and obtaining an attention-weighted feature representation; performing a multi-head splicing operation according to the attention-weighted feature representation to obtain a splicing feature, and performing a linear transformation on the splicing feature to obtain a linear transformation result; inputting the linear transformation result into a feedforward neural network, performing two linear transformations and nonlinear activation function processing on the linear transformation result through the feedforward neural network, and obtaining an output feature of the feedforward network; performing residual connection and layer normalization operations on the output feature of the feedforward network to obtain an output result of the current Transformer encoder, and using the output result as the input of the next Transformer encoder, until the last Transformer encoder, and obtaining a context-aware feature sequence.
[0038] Specifically, the encoding sequence is input into the preset Transformer encoder, and the multi-head self-attention mechanism in the Transformer encoder performs multi-head self-attention calculation on the encoding sequence to obtain the attention-weighted feature representation. In the specific implementation, the input encoding sequence X is first transformed into query (Q), key (K) and value (V) through three different linear transformations: Q = XWq, K = XWk, V = XWv, where Wq, Wk, Wv are learnable weight matrices. Then, the attention weight is calculated: Q is multiplied by the transpose of K, divided by the scaling factor (usually the square root of the key vector dimension), and then passed through the softmax function to obtain the attention score. Finally, these scores are weighted and summed on V to obtain the attention-weighted feature representation. This process is performed in parallel on multiple attention heads, and each head uses a different set of Wq, Wk, and Wv, allowing the model to focus on different aspects of the input.
[0039] Specifically, a multi-head splicing operation is performed according to the attention-weighted feature representation to obtain the spliced features, and the spliced features are linearly transformed to obtain the linear transformation results. The multi-head splicing operation splices the outputs of all attention heads in the last dimension. Assuming that there are h attention heads and the output dimension of each head is dv, the dimension of the spliced features is h*dv. The splicing operation retains the different types of information extracted by each attention head. Next, a linear transformation is performed on the spliced features: O = Concat(head1, ..., headh)Wo, where Wo is a learnable weight matrix that maps the spliced features back to the hidden dimension d_model of the model. The purpose of this linear transformation is to integrate the information of multiple heads to obtain a unified representation.
[0040] Specifically, the linear transformation result is input into a feedforward neural network, and the linear transformation result is subjected to two linear transformations and nonlinear activation function processing through the feedforward neural network to obtain the output features of the feedforward network. The feedforward neural network contains two linear transformation layers with a nonlinear activation function (usually ReLU) in the middle. The first linear transformation usually expands the input dimension to 4 times d_model, and the second linear transformation compresses the dimension back to d_model. The specific implementation is: FFN(x) =max(0, xW1 + b1)W2 + b2, where W1, W2, b1, and b2 are learnable parameters. This feedforward network increases the nonlinear ability of the model, allowing it to capture more complex feature relationships.
[0041] Specifically, residual connections and layer normalization operations are performed on the output features of the feedforward network to obtain the output of the current Transformer encoder, and the output result is used as the input of the next Transformer encoder until the last Transformer encoder to obtain a context-aware feature sequence. The residual connection adds the input directly to the output of the sublayer: x + Sublayer(x), where Sublayer can be a multi-head self-attention layer or a feedforward network layer. This connection helps alleviate the gradient vanishing problem and allows information to flow more easily in deep networks. Layer normalization normalizes the features of each sample: LN(x) = α * (x - μ) / (σ + ε) + β, where μ and σ are the feature mean and standard deviation, and α and β are learnable scaling and offset parameters. Layer normalization helps stabilize the training of deep networks. This process is repeated in multiple layers of the Transformer encoder, and the output of each layer is used as the input of the next layer until the last layer. The final output is a context-aware feature sequence, which contains global context information for each position in the input sequence.
[0042] 104. Use a multi-layer perceptron to hierarchically classify the fixed-length vector representation to obtain the instruction category, extract key parameters based on the instruction category, generate standardized control instructions and send them to the corresponding smart device.
[0043] In one embodiment of the present invention, the use of a multilayer perceptron to perform hierarchical classification on the fixed-length vector representation to obtain instruction categories, and extracting key parameters based on the instruction categories to generate standardized control instructions and send them to smart devices includes: using a preset first-layer multilayer perceptron to identify major instruction categories on the fixed-length vector representation to obtain major instruction categories; based on the major instruction categories, using a preset second-layer multilayer perceptron to perform fine-grained instruction identification on the fixed-length vector representation to obtain specific instruction categories; combining a rule-based method and a sequence labeling model to analyze the specific instruction categories, extracting key parameters in the specific instruction categories, and generating and sending standardized control instructions to the corresponding smart devices based on the specific instruction categories and the key parameters.
[0044] Specifically, the fixed-length vector representation is identified by using the preset first-layer multilayer perceptron to obtain the instruction categories. In this step, the fixed-length vector representation is first input into the first-layer multilayer perceptron. The multilayer perceptron usually contains two to three fully connected layers, and the ReLU activation function is used between each layer. The input dimension of the first layer is the same as the dimension of the fixed-length vector, while the output dimension is usually set to half or one-quarter of the input dimension. The dimension of the intermediate layer (if any) can be further reduced. The output dimension of the last layer is equal to the number of predefined instruction categories. A softmax function is applied after the last layer to convert the output into a probability distribution. The category with the highest probability is selected as the identified instruction category. This hierarchical recognition method first divides the instructions into coarse-grained categories, such as "device control", "information query", etc., to provide context information for subsequent fine-grained recognition.
[0045] According to the identified instruction categories, the preset second-layer multilayer perceptron is used to perform fine-grained instruction recognition on the fixed-length vector representation to obtain specific instruction categories. In this step, the structure of the second-layer multilayer perceptron is similar to that of the first layer, but there is a dedicated multilayer perceptron for each instruction category. According to the instruction category obtained in the first step, the corresponding second-layer multilayer perceptron is selected for fine-grained recognition. The input of this multilayer perceptron is the concatenation of the original fixed-length vector representation and the one-hot encoding of the instruction category. The concatenation operation explicitly provides the category information to the network, which helps to more accurately identify fine-grained instructions. The output dimension of the second-layer multilayer perceptron is equal to the number of specific instruction categories under the instruction category. The softmax function is also used to convert the output into a probability distribution, and the category with the highest probability is selected as the identified specific instruction category. This two-stage classification method can handle a large number of instruction categories while maintaining a high recognition accuracy.
[0046] Combine the rule-based method and the sequence labeling model to analyze the specific instruction categories and extract the key parameters in the specific instruction categories. First, based on the identified specific instruction categories, predefined rules are applied to preliminarily determine the type of key parameters that need to be extracted. For example, for the "adjust temperature" type of instruction, the temperature value and possible room location need to be extracted. Then, the original speech transcription text is processed using a sequence labeling model (such as conditional random field CRF or bidirectional LSTM-CRF) to annotate the semantic roles of each word. The input of the sequence labeling model is a sequence of word vectors, and the output is the label of each word (such as B-temperature, I-temperature, B-position, I-position, etc.). In this way, the key parameters in the instruction can be accurately located and extracted. Finally, according to the rules and labeling results, the extracted parameter values are standardized, such as converting "twenty-five degrees" to the value 25.
[0047] Generate and send standardized control instructions to the corresponding smart devices based on the specific instruction category and the extracted key parameters. This step first selects the corresponding instruction template based on the specific instruction category. The instruction template is a predefined structured format that contains the instruction type and parameter placeholders. For example, the template for temperature adjustment may be "SET_TEMPERATURE {device ID} {temperature value}". Then, fill in the extracted key parameters in the corresponding positions of the template. If some necessary parameters are missing, default values can be used or interactive queries can be triggered. The generated standardized control instructions also need to be formatted to ensure that they comply with the interface specifications of the target smart device. Finally, the standardized control instructions are sent to the corresponding smart devices through the preset communication protocol (such as MQTT, HTTP, etc.). Possible network delays and errors need to be handled during the sending process, and a retry mechanism needs to be implemented if necessary. After receiving the instruction, the device will perform the corresponding operation and return the execution status. The system needs to parse this status and provide feedback to the user if necessary.
[0048] In this embodiment, the collected speech signal is preprocessed to obtain a mel spectrum graph; two preset integrated deep learning models are used to extract features from the mel spectrum graph to obtain a joint feature vector; the joint feature vector is subjected to attention weighted fusion and dimension reduction to obtain a fused feature representation; the fused feature representation is subjected to block, linear projection and position encoding processing to obtain a coding sequence; the coding sequence is input into the Transformer encoder, and a fixed-length vector representation is obtained by global average pooling; hierarchical classification is performed using a multi-layer perceptron to obtain instruction categories and key parameters; standardized control instructions are generated and sent to the smart device. This method can effectively process non-standard speech, improve recognition accuracy and real-time performance, enhance the system's generalization ability, and is suitable for diversified speech input scenarios.
[0049] The above describes the voice control method based on artificial intelligence in the embodiment of the present invention. The following describes the voice control device based on artificial intelligence in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a voice control device based on artificial intelligence includes: The preprocessing module 201 is used to preprocess the collected speech signal to obtain a short-time frame sequence, calculate the Mel-frequency cepstral coefficients and differences of the speech signal according to the short-time frame sequence, and convert the calculation results into a Mel-frequency spectrum. A feature extraction module 202 is used to extract features from the mel-spectrogram using a preset first integrated deep learning model and a second integrated deep learning model to obtain a joint feature vector, and perform attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; A sequence encoding module 203 is used to sequentially perform block processing, linear projection processing and position encoding processing on the fused feature representation to obtain a coding sequence, and input the coding sequence into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; The instruction recognition module 204 is used to use a multi-layer perceptron to hierarchically classify the fixed-length vector representation to obtain instruction categories, extract key parameters according to the instruction categories, generate standardized control instructions, and send them to corresponding smart devices.
[0050] In an embodiment of the present invention, the artificial intelligence-based voice control device runs the above-mentioned artificial intelligence-based voice control method, and the artificial intelligence-based voice control device obtains a mel spectrum by preprocessing the collected voice signal; uses two preset integrated deep learning models to extract features from the mel spectrum to obtain a joint feature vector; performs attention weighted fusion and dimension reduction on the joint feature vector to obtain a fused feature representation; performs block, linear projection and position encoding processing on the fused feature representation to obtain a coding sequence; inputs the coding sequence into a Transformer encoder, and obtains a fixed-length vector representation through global average pooling; uses a multi-layer perceptron for hierarchical classification to obtain instruction categories and key parameters; generates standardized control instructions and sends them to smart devices. This method can effectively process non-standard voices, improve recognition accuracy and real-time performance, enhance system generalization capabilities, and is suitable for diversified voice input scenarios.
[0051] above Figure 2 The artificial intelligence-based voice control device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The artificial intelligence-based voice control device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0052] Figure 3 3 is a schematic diagram of the structure of an artificial intelligence-based voice control device provided by an embodiment of the present invention. The artificial intelligence-based voice control device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the artificial intelligence-based voice control device 300. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the artificial intelligence-based voice control device 300 to implement the steps of the above-mentioned artificial intelligence-based voice control method.
[0053] The artificial intelligence-based voice control device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 3 The structure of the artificial intelligence-based voice control device shown does not constitute a limitation on the artificial intelligence-based voice control device provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0054] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the artificial intelligence-based voice control method.
[0055] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0056] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0057] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice control method based on artificial intelligence, characterized in that: The artificial intelligence-based voice control method comprises: Preprocessing the collected speech signal to obtain a short-time frame sequence, and calculating the Mel-frequency cepstrum coefficients and differences of the speech signal according to the short-time frame sequence, and converting the calculation results into a Mel-frequency spectrum; Using a preset first integrated deep learning model and a second integrated deep learning model to extract features from the mel-spectrogram to obtain a joint feature vector, and performing attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; The fused feature representation is sequentially subjected to block processing, linear projection processing, and position encoding processing to obtain a coding sequence, and the coding sequence is input into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; A multi-layer perceptron is used to hierarchically classify the fixed-length vector representation to obtain instruction categories, and key parameters are extracted according to the instruction categories to generate standardized control instructions and send them to corresponding intelligent devices.
2. The artificial intelligence-based voice control method according to claim 1, characterized in that: The preprocessing of the collected speech signal to obtain a short-time frame sequence, calculating the Mel-frequency cepstrum coefficients and differences of the speech signal according to the short-time frame sequence, and converting the calculation results into a Mel-frequency spectrum diagram includes: Performing noise reduction and frame division processing on the collected speech signal to obtain a short-time frame sequence, and performing pre-emphasis and windowing processing and fast Fourier transform on the short-time frame sequence to obtain a frequency domain representation of the speech signal; Calculating Mel-frequency cepstral coefficients according to the frequency domain representation, and calculating first-order differences and second-order differences of the Mel-frequency cepstral coefficients; The Mel-frequency cepstrum coefficients and the first-order differences and the second-order differences are synthesized into a feature sequence, and the feature sequence is converted into a Mel-frequency spectrum graph.
3. The artificial intelligence-based voice control method according to claim 1, characterized in that: The method of extracting features from the Mel-spectrogram using the preset first integrated deep learning model and the second integrated deep learning model to obtain a joint feature vector, and performing attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation includes: Using a preset first integrated deep learning model to perform local feature extraction on the Mel-spectrogram to obtain a first feature set, wherein the first feature set includes short-term speech features and local spectrum patterns; Using a preset second integrated deep learning model to perform local feature extraction on the Mel-spectrogram to obtain a second feature set, wherein the second feature set includes a long-term dependency and a global speech structure; Zero-padding and concatenating the first feature set and the second feature set to obtain a joint feature vector; The multi-head attention mechanism is used to perform weighted fusion on the joint feature vector to obtain a fusion result, and the principal component analysis method is used to reduce the dimension of the fusion result to obtain a fusion feature representation.
4. The artificial intelligence-based voice control method according to claim 3, characterized in that: The first integrated deep learning model integrates a VGG16 network, a DenseNet201 network, and a GoogleNet network; The first feature set is obtained by extracting local features from the Mel-spectrogram using a preset first integrated deep learning model, and includes: Perform sliding window segmentation on the Mel-spectrogram to obtain multiple local feature regions, and use the VGG16 network to perform a convolution operation on each local feature region to obtain a first-level feature map; Performing maximum pooling on the first-level feature map to obtain a second-level feature map, and inputting the second-level feature map into a DenseNet201 network, performing feature extraction through a densely connected block in the DenseNet201 network, and obtaining a third-level feature map; The Inception module in the GoogleNet network is used to perform multi-way parallel convolution and pooling operations on the third-level feature map to obtain feature maps of different receptive fields, and the feature maps of different receptive fields are spliced in the channel dimension to obtain a multi-scale feature representation. Perform weighted sum feature aggregation on the multi-scale features to obtain a first feature set.
5. The artificial intelligence-based voice control method according to claim 1, characterized in that: The step of sequentially performing block processing, linear projection processing, and position encoding processing on the fused feature representation to obtain a coding sequence, and inputting the coding sequence into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling includes: Dividing the fused feature representation into blocks of equal size to obtain a feature block sequence, and performing linear projection on each feature block in the feature block sequence through a pre-trained weight matrix and a bias vector to obtain a projection feature; Generate a predefined position coding matrix, and add the position coding matrix to the projection feature element by element to obtain a coding sequence containing position information, wherein the number of rows of the position coding matrix is equal to the length of the projection feature, and each row corresponds to an element position in the projection feature; Inputting the encoded sequence into a preset Transformer encoder, processing it through a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder to obtain a context-aware feature sequence; A global average pooling operation is performed on the context-aware feature sequence to calculate the average value of all token representations in the context-aware feature sequence to obtain a vector representation of a fixed length.
6. The artificial intelligence-based voice control method according to claim 5, characterized in that: The encoding sequence is input into a preset Transformer encoder, and processed by a multi-head self-attention mechanism and a feedforward neural network in the Transformer encoder to obtain a context-aware feature sequence, including: Inputting the encoding sequence into a preset Transformer encoder, performing multi-head self-attention calculation on the encoding sequence through a multi-head self-attention mechanism in the Transformer encoder, and obtaining an attention-weighted feature representation; Performing a multi-head splicing operation according to the attention-weighted feature representation to obtain a splicing feature, and performing a linear transformation on the splicing feature to obtain a linear transformation result; Inputting the linear transformation result into a feedforward neural network, performing two linear transformations and nonlinear activation function processing on the linear transformation result through the feedforward neural network, and obtaining output features of the feedforward network; Residual connection and layer normalization operations are performed on the output features of the feedforward network to obtain the output result of the current Transformer encoder, and the output result is used as the input of the next Transformer encoder until the last Transformer encoder to obtain a context-aware feature sequence.
7. The artificial intelligence-based voice control method according to claim 1, characterized in that: The method of using a multi-layer perceptron to hierarchically classify the fixed-length vector representation to obtain instruction categories, extracting key parameters according to the instruction categories, generating standardized control instructions, and sending them to the smart device includes: Using a preset first-layer multi-layer perceptron to identify the major categories of instructions on the fixed-length vector representation, to obtain the major categories of instructions; According to the major instruction categories, a preset second-layer multi-layer perceptron is used to perform fine-grained instruction recognition on the fixed-length vector representation to obtain a specific instruction category; The specific instruction category is analyzed in combination with the rule-based method and the sequence labeling model, key parameters in the specific instruction category are extracted, and standardized control instructions are generated and sent to the corresponding smart device based on the specific instruction category and the key parameters.
8. A voice control device based on artificial intelligence, characterized in that: The artificial intelligence-based voice control device comprises: A preprocessing module is used to preprocess the collected speech signal to obtain a short-time frame sequence, and calculate the Mel-frequency cepstrum coefficients and differences of the speech signal according to the short-time frame sequence, and convert the calculation results into a Mel-frequency spectrum; A feature extraction module, used to extract features from the mel-spectrogram using a preset first integrated deep learning model and a second integrated deep learning model to obtain a joint feature vector, and perform attention weighted fusion and dimensionality reduction on the joint feature vector to obtain a fused feature representation; A sequence encoding module is used to sequentially perform block processing, linear projection processing and position encoding processing on the fused feature representation to obtain a coding sequence, and input the coding sequence into a preset Transformer encoder to obtain a fixed-length vector representation through global average pooling; The instruction recognition module is used to use a multi-layer perceptron to hierarchically classify the fixed-length vector representation to obtain instruction categories, extract key parameters according to the instruction categories, generate standardized control instructions and send them to corresponding smart devices.
9. A voice control device based on artificial intelligence, characterized in that: The artificial intelligence-based voice control device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory so that the artificial intelligence-based voice control device performs the steps of the artificial intelligence-based voice control method as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the artificial intelligence-based voice control method as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Depression state prediction method and system based on voice multi-scale time domain perception
CN120783803A
Depression state prediction method and system based on voice multi-scale time domain perception
CN120783803B