Acquisition method of voice recognition model in electric power emergency consultation environment, voice recognition method, device, equipment, storage medium and program product

By extracting and fusion of various data types in the power emergency consultation environment and training deep learning models, the problems of accuracy and reliability of traditional technologies in complex environments are solved, and efficient and accurate speech recognition is achieved.

CN120048250APending Publication Date: 2025-05-27GUANGDONG POWER GRID CO LTD EMERGENCY & RISK MANAGEMENT CENTER
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510116433.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional single-modal voice processing technology is difficult to deal with noise interference and multi-speaker recognition in complex power emergency consultation environments, resulting in low accuracy and reliability of emergency information output by voice recognition.

Method used

By extracting and fusion audio, text, image and video data in historical power emergency consultation environments, deep learning models are trained to obtain speech recognition models, which can process multiple data types at the same time, enhancing the model's understanding of complex environments.

Benefits of technology

It improves the accuracy and reliability of the emergency information output by voice recognition, and can efficiently extract and accurately identify the voice information of multi-speakers in complex power emergency consultation environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048250A_ABST
    Figure CN120048250A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of emergency communication and information processing of an intelligent power system, and provides a voice recognition model acquisition method and device, a voice recognition method and device, equipment, a storage medium and a program product in a power emergency consultation environment. The method comprises the steps of obtaining an audio feature sample, a text feature sample, a field image feature sample and a field video feature sample according to an audio data sample, a text data sample, a field image data sample and a field video data sample in a historical electric power emergency consultation environment; obtaining a historical event feature sample according to historical power emergency event data in a historical power emergency consultation environment; and according to the audio feature sample, the text feature sample, the on-site image feature sample, the on-site video feature sample and the historical event feature sample, obtaining a multi-modal feature sample to train a to-be-trained model to obtain a speech recognition model. By adopting the method, the accuracy and the reliability of the emergency information output by voice recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of emergency communication and information processing in intelligent power systems, and particularly to a method for obtaining a speech recognition model in a power emergency consultation environment, a speech recognition method in a power emergency consultation environment, a device for obtaining a speech recognition model in a power emergency consultation environment, a speech recognition device in a power emergency consultation environment, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the increasing complexity and intelligence of power systems, the speed and accuracy of emergency response have become one of the key factors in ensuring the safe and stable operation of the power grid. During power emergency consultations, a large amount of voice information needs to be processed quickly and accurately in order to make correct decisions in a timely manner.

[0003] However, traditional single-modal speech processing technologies have limitations in dealing with problems such as noise interference and multi-speaker recognition in complex power emergency consultations, resulting in low accuracy and reliability of the emergency information output by speech recognition, and it is difficult to meet the requirements of the power system for high efficiency and high reliability of emergency communication. Summary of the Invention

[0004] Based on this, it is necessary to provide a method for obtaining a speech recognition model in a power emergency consultation environment, a speech recognition method in a power emergency consultation environment, a device for obtaining a speech recognition model in a power emergency consultation environment, a speech recognition device in a power emergency consultation environment, a computer device, a computer-readable storage medium, and a computer program product for the above technical problems.

[0005] In the first aspect, this application provides a method for obtaining a speech recognition model in a power emergency consultation environment, including:

[0006] Extract features from audio data samples, text data samples, on-site image data samples, and on-site video data samples in a historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples;

[0007] Extract features from historical power emergency event data in a historical power emergency consultation environment to obtain historical event feature samples;

[0008] Fuse the audio feature samples, the text feature samples, the on-site image feature samples, the on-site video feature samples, and the historical event feature samples to obtain multi-modal feature samples;

[0009] Train a deep learning model to be trained according to the multi-modal feature samples to obtain a speech recognition model.

[0010] In a second aspect, the present application also provides a speech recognition method in a power emergency consultation environment, including:

[0011] Extract features from the audio data, text data, on-site image data, and on-site video data in the target power emergency consultation environment to obtain audio features, text features, on-site image features, and on-site video features;

[0012] Extract features from the target power emergency event data in the target power emergency consultation environment to obtain target event features;

[0013] Perform feature fusion on the audio features, text features, on-site image features, on-site video features, and target event features to obtain multimodal features;

[0014] Input the multimodal features into a speech recognition model to obtain power emergency information in the target power emergency consultation environment; the speech recognition model is a model obtained according to the embodiment of the method for obtaining the speech recognition model in the above power emergency consultation environment.

[0015] In a third aspect, the present application also provides an apparatus for obtaining a speech recognition model in a power emergency consultation environment, including:

[0016] A feature sample acquisition module, configured to extract features from the audio data samples, text data samples, on-site image data samples, and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples;

[0017] The feature sample acquisition module is further configured to extract features from the historical power emergency event data in the historical power emergency consultation environment to obtain historical event feature samples;

[0018] A multimodal feature sample acquisition module, configured to perform feature fusion on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples to obtain multimodal feature samples;

[0019] A speech recognition model acquisition module, configured to perform model training on a deep learning model to be trained according to the multimodal feature samples to obtain a speech recognition model.

[0020] In a fourth aspect, the present application also provides a speech recognition apparatus in a power emergency consultation environment, including:

[0021] A feature acquisition module, configured to extract features from audio data, text data, on-site image data, and on-site video data in a target power emergency consultation environment, to obtain audio features, text features, on-site image features, and on-site video features;

[0022] The feature acquisition module is further configured to extract features from target power emergency event data in the target power emergency consultation environment, to obtain target event features;

[0023] A multi-modal feature acquisition module, configured to perform feature fusion on the audio features, the text features, the on-site image features, and the on-site video features, to obtain multi-modal features;

[0024] A speech recognition module, configured to input the audio features, text features, on-site image features, and on-site video features into a speech recognition model, to obtain power emergency information in the target power emergency consultation environment; the speech recognition model is a model obtained by the method embodiment of the acquisition method of the speech recognition model in the above-mentioned power emergency consultation environment.

[0025] In a fifth aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the above method.

[0026] In a sixth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and the computer program is executed by the processor to execute the above method.

[0027] In a seventh aspect, the present application further provides a computer program product. The computer program product includes a computer program, and the computer program is executed by the processor to execute the above method.

[0028] This application extracts features from audio data samples, text data samples, on-site image data samples, and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples; extracts features from historical power emergency event data in the historical power emergency consultation environment to obtain historical event feature samples; performs feature fusion on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples to obtain multi-modal feature samples; and trains a deep learning model to be trained according to the multi-modal feature samples to obtain a speech recognition model. This application obtains multi-modal feature samples based on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples in the historical power emergency consultation environment, trains a deep learning model to be trained, and obtains a speech recognition model, which can process multiple data types such as speech, text, and images simultaneously, and can enhance the understanding ability of the speech recognition model for complex power emergency consultation environments; moreover, historical event feature samples are also extracted for speech recognition, further enhancing the understanding ability of the speech recognition model for complex power emergency consultation environments, enabling comprehensive understanding and efficient processing of multi-source information in complex power emergency consultation environments, efficiently extracting and accurately identifying multi-speaker speech information in complex power emergency consultation environments, obtaining power emergency information in complex power emergency consultation environments, and improving the accuracy and reliability of the emergency information output by speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.

[0030] Figure 1 It is an application environment diagram of a method for obtaining a speech recognition model and a speech recognition method in a power emergency consultation environment in an embodiment;

[0031] Figure 2 It is a schematic flowchart of a method for obtaining a speech recognition model in a power emergency consultation environment in an embodiment;

[0032] Figure 3 It is a schematic flowchart of a speech recognition method in a power emergency consultation environment in an embodiment;

[0033] Figure 4 It is a structural block diagram of a device for obtaining a speech recognition model in a power emergency consultation environment in an embodiment;

[0034] Figure 5 is a structural block diagram of a voice recognition device in a power emergency consultation environment in an embodiment;

[0035] Figure 6 is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0036] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0037] The method for obtaining a voice recognition model in a power emergency consultation environment and the voice recognition method in a power emergency consultation environment provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The terminal 102 can obtain audio data samples, text data samples, on-site image data samples and on-site video data samples in the historical power emergency consultation environment for feature extraction, and then obtain a voice recognition model. The terminal 102 can also obtain audio data, text data, on-site image data and on-site video data in the target power emergency consultation environment, and then obtain the power emergency information in the target power emergency consultation environment. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0038] In an exemplary embodiment, as Figure 2 shown in the figure, a method for obtaining a voice recognition model in a power emergency consultation environment is provided. Taking the method applied to Figure 1 the terminal 102 in the figure as an example for description, it includes the following steps S201 to S204. Among them:

[0039] Step S201: Extract features from audio data samples, text data samples, on-site image data samples and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples and on-site video feature samples.

[0040] Audio data, text data, on-site image data, and on-site video data can be collected from various historical emergency consultation environments in the power system, and data annotation and data preprocessing can be carried out to obtain audio data samples, text data samples, on-site image data samples, and on-site video data samples in the historical power emergency consultation environment. The audio data includes, but is not limited to, dispatching meeting recordings, on-site report recordings, and fault handling call recordings. The text data includes, but is not limited to, relevant text records, such as meeting minutes and operation logs.

[0041] The collected audio data, text data, on-site image data, and on-site video data need to cover a variety of environmental conditions, such as different background noise levels, different numbers of speakers, different device types and qualities, etc., so as to improve the diversity of audio data samples, text data samples, on-site image data samples, and on-site video data samples, and to improve the generalization ability of the speech recognition model.

[0042] The collected audio data can be annotated to mark key information (such as commands, warning messages, and operation instructions), and corresponding labels can be added to the text data, on-site image data, and on-site video data.

[0043] The specific steps for data preprocessing of the collected audio data, text data, on-site image data, and on-site video data are as follows:

[0044] For audio data, audio noise reduction algorithms (such as spectral subtraction, wavelet transform, etc.) can be used to remove background noise and improve the clarity of the speech signal; then, audio files in different formats are uniformly converted to a standard format (such as WAV or MP3) to ensure the consistency of subsequent processing; long audio files are split into shorter segments, and the length of each segment is usually between a few seconds and dozens of seconds, which is convenient for model processing and training; the audio segments are normalized so that the amplitude of the audio segments is within a certain range to avoid signals that are too large or too small from affecting the model performance.

[0045] For text data, irrelevant characters and punctuation marks can be removed, the text is converted to lowercase, stemming and stop word filtering are carried out to improve the readability and processing efficiency of the text; then, the text is segmented into words or phrases to generate a word vector representation, which is convenient for the model to understand and process.

[0046] For on-site image data and on-site video data, the images and videos can be cropped and scaled to a unified size to ensure that the image sizes input into the model are the same; then, image denoising algorithms (such as Gaussian filtering algorithm or median filtering algorithm) are used to remove the noise in the images and videos and improve the image quality; the brightness, contrast, and saturation of the images and videos can be adjusted to enhance the visual effects of the images and videos and improve the recognition ability of the model.

[0047] It is possible to extract features from audio data samples, text data samples, on-site image data samples, and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples.

[0048] Step S202: Extract features from the historical power emergency event data in the historical power emergency consultation environment to obtain historical event feature samples.

[0049] It is possible to extract the features of the type, geographical location, response measures, time information, and influence scope of the historical power emergency event data in the historical power emergency consultation environment to obtain the event type code, geographical location code, response measure code, time information feature, and influence scope feature corresponding to each historical power emergency event. The above-mentioned event type code, geographical location code, response measure code, time information feature, and influence scope feature corresponding to each historical power emergency event are used as historical event feature samples.

[0050] Step S203: Perform feature fusion on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples to obtain multi-modal feature samples.

[0051] It is possible to align the time axes of the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples, and perform feature fusion on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples after the time axis alignment. For example, perform simple splicing or splicing after weighted averaging on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples after the time axis alignment to obtain multi-modal feature samples.

[0052] Step S204: According to the multi-modal feature samples, train the deep learning model to be trained to obtain a speech recognition model.

[0053] Before training the deep learning model to be trained according to the multi-modal feature samples, it is possible to perform standardization processing on the multi-modal feature samples to avoid a certain modal feature sample dominating the learning process of the model, reduce the computational complexity, and improve the training efficiency of the model.

[0054] The specific steps for standardizing multi-modal feature samples are as follows: Normalize the multi-modal feature samples so that they are distributed within the same range, avoiding a certain modal feature from dominating the model's learning process; Then, use the Principal Component Analysis (PCA) method or the t-Distributed Stochastic Neighbor Embedding (t-SNE) method to reduce the dimensionality of the normalized multi-modal feature samples, thereby reducing the computational complexity and improving the training efficiency of the model. The standardized multi-modal feature samples can be packaged into a format suitable for model training, such as TFRecord or HDF5, to facilitate efficient reading and processing by the model.

[0055] The standardized multi-modal feature samples can be divided into a training set and a validation set. Usually, the training set accounts for 80% and the validation set accounts for 20%. To improve the generalization ability of the model, data augmentation techniques such as random cropping, rotation, flipping, etc. can be used on the multi-modal feature samples, especially for on-site image feature samples and on-site video feature samples.

[0056] A deep learning model suitable for multi-modal data processing can be selected, such as a Multimodal Transformer model, a Multimodal Convolutional Neural Network (MMCNN) model, or a Multimodal Recurrent Neural Network (MMRNN) model, as the deep learning model to be trained. The output layer of the deep learning model to be trained can be a classifier (for classifying different emergency response measures) or a regressor (for predicting the quantity of emergency resource allocation), and can be used to generate power emergency information, which can include emergency response strategies and resource allocation plans.

[0057] According to the multi-modal feature samples, the specific steps for training the deep learning model to be trained to obtain a speech recognition model are as follows:

[0058] The loss function can be selected according to the task of the deep learning model to be trained. If the task is to classify different emergency response measures, the cross-entropy loss function can be used. If the task is to predict the quantity of resource allocation, the mean squared error (MSE) loss function can be used. If both classification and regression tasks are to be processed simultaneously, a multi-task loss function can be used, and multiple loss functions are weighted and summed. An L1 or L2 regularization term can be added to the loss function to prevent the model from overfitting.

[0059] Optimization algorithms can be used to optimize the parameters of the deep learning model to be trained. For example, the Stochastic Gradient Descent (SGD) algorithm or variants of the stochastic gradient descent algorithm, such as the Adaptive Moment Estimation (Adam) algorithm and the Root Mean Square Propagation (RMSprop) algorithm, can be used to optimize the parameters of the deep learning model to be trained. A learning rate decay strategy, such as step decay and exponential decay, can be used to improve the convergence speed and stability of model training. The Dropout technique can be used in some layers of the deep learning model to be trained to randomly discard a part of the neurons and improve the generalization ability of the deep learning model to be trained. An appropriate batch size, usually between 32 and 128, can be selected for the deep learning model to be trained to balance memory usage and training speed. An appropriate number of epochs can be set, usually determined based on the performance on the validation set to decide when to stop training.

[0060] After training the deep learning model to be trained based on the multi-modal feature samples, an initial speech recognition model can be obtained. The initial speech recognition model can be evaluated for performance, and based on the performance evaluation results, the model can be tuned to obtain the speech recognition model.

[0061] The performance evaluation metrics can be selected according to the tasks of the initial speech recognition model. For classification tasks, metrics such as Accuracy, Precision, Recall, or F1 Score can be used to evaluate the classification performance of the initial speech recognition model. For regression tasks, metrics such as Mean Squared Error (MSE), Root Mean Squared Error (RMSE), or Mean Absolute Error (MAE) can be used to evaluate the regression performance of the initial speech recognition model. For multi-task learning, the performance of classification and regression tasks can be considered comprehensively, and comprehensive metrics can be used for evaluation.

[0062] The performance of the initial speech recognition model can be evaluated based on the validation set, and k-fold cross-validation can be used to evaluate the stability and generalization ability of the initial speech recognition model. The performance of the initial speech recognition model can be monitored on the validation set, and training can be terminated early when the performance of the initial speech recognition model no longer improves to prevent overfitting.

[0063] Model tuning can be performed based on the performance evaluation results to obtain a speech recognition model. For example, methods such as Grid Search or Random Search can be used to try different combinations of hyperparameters to find the optimal hyperparameter configuration for the speech recognition model. The Bayesian optimization method can be used to gradually optimize the hyperparameters of the speech recognition model by constructing a prior distribution of hyperparameters to improve the search efficiency. The Bootstrap Aggregating (Bagging) method can be used to train multiple base models and make predictions by voting or averaging to improve the stability and generalization ability of the speech recognition model. The Boosting method, such as the AdaBoost method or the Gradient Boosting method, can be used to gradually reduce the error and improve the prediction ability of the speech recognition model by serially training multiple weak models.

[0064] The trained speech recognition model can be saved and deployed. Specifically, during the model training process, the model with the best performance on the validation set can be saved regularly. Model compression techniques (such as quantization and pruning) are used to reduce the size of the speech recognition model and improve the inference speed. The trained speech recognition model can be exported in a standard format (such as ONNX or TensorFlow SavedModel) for easy deployment on different platforms. The speech recognition model can be deployed to a server or cloud platform, providing an Application Programming Interface (API) for convenient invocation of subsequent speech recognition tasks.

[0065] In the above method for obtaining a speech recognition model in a power emergency consultation environment, based on the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples, and historical event feature samples in the historical power emergency consultation environment, multi-modal feature samples are obtained to train the deep learning model to be trained, and a speech recognition model is obtained. It can process multiple data types such as speech, text, and images simultaneously, and can enhance the understanding ability of the speech recognition model for complex power emergency consultation environments; moreover, historical event feature samples are also extracted for speech recognition, further enhancing the understanding ability of the speech recognition model for complex power emergency consultation environments, enabling comprehensive understanding and efficient processing of multi-source information in complex power emergency consultation environments, so as to efficiently extract and accurately identify the multi-speaker speech information in complex power emergency consultation environments, obtain the power emergency information in complex power emergency consultation environments, and improve the accuracy and reliability of the emergency information output by speech recognition.

[0066] In one embodiment, feature extraction is performed on the audio data samples, text data samples, on-site image data samples, and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples. The specific steps are as follows: Feature extraction is performed on the audio data samples in the historical power emergency consultation environment to obtain the Mel frequency cepstral coefficients, short-time energy, zero-crossing rate, and linear predictive coding of the audio data samples as audio feature samples; according to the pre-constructed word embedding model and sentence embedding model, feature extraction is performed on the text data samples in the historical power emergency consultation environment to obtain the sentence vectors of the text data samples as text feature samples; according to the pre-constructed convolutional neural network model, feature extraction is performed on the on-site image data samples and on-site video data samples in the historical power emergency consultation environment to obtain the feature maps, local feature points, and descriptors of the on-site image data samples and on-site video data samples as on-site image feature samples and on-site video feature samples.

[0067] Feature extraction can be performed on the audio data samples in the historical power emergency consultation environment to obtain the Mel-frequency cepstral coefficients, short-time energy, zero-crossing rate, and linear predictive coding of the audio data samples, which are used as audio feature samples.

[0068] The specific steps for obtaining the Mel-frequency cepstral coefficients of the audio data samples are as follows: perform pre-emphasis processing on the original audio signal in the audio data samples to enhance the high-frequency part and reduce the influence of high-frequency noise; divide the pre-emphasized audio signal into multiple short-time frames, usually with a frame length of 20-40 milliseconds and a frame shift of 10-20 milliseconds; multiply each frame of the signal by a Hamming window or a Hanning window to reduce the discontinuity at the frame boundaries; perform a Fast Fourier Transform (FFT) on each frame of the signal to obtain the frequency-domain signal; pass the frequency-domain signal through a set of Mel filters to convert it to the Mel-frequency scale, take the logarithm of the output of each filter to obtain the logarithmic energy; perform a Discrete Cosine Transform (DCT) on the logarithmic energy to obtain the Mel-frequency cepstral coefficients (MFCC).

[0069] Among them, the short-time energy and zero-crossing rate can provide additional information about the energy change and syllable boundaries of the speech signal; linear predictive coding (LPC) can capture the formant information of the speech signal.

[0070] According to the pre-constructed word embedding model, the words in the text data samples in the historical power emergency consultation environment can be converted into high-dimensional vector representations, and the term frequency-inverse document frequency of each word can be calculated to reflect the importance of the word in the text. Then, according to the pre-constructed sentence embedding model, the entire sentence in the text data sample can be converted into a fixed-length vector representation to obtain the sentence vector of the text data sample, which is used as the text feature sample. For longer texts, a document-to-vector (Doc2Vec) model or other document embedding methods can be used to convert the entire text into a vector representation. Among them, the word embedding model can be a word-to-vector (Word2Vec) model, a global word vector (GloVe) model, or a bidirectional encoder representation from transformers (BERT) model. The sentence embedding model can be a Sentence-BERT model.

[0071] Based on a pre - constructed Convolutional Neural Networks (CNN) model, feature extraction can be performed on the on - site image data samples and on - site video data samples in the historical power emergency consultation environment to obtain the feature maps, local feature points, and descriptors of the on - site image data samples and on - site video data samples, which are used as on - site image feature samples and on - site video feature samples. Among them, the local feature points of the on - site image data samples and on - site video data samples can be obtained according to the Scale - Invariant Feature Transform (SIFT) algorithm. The convolutional neural network model can be a Visual Geometry Group (VGG) model, a Residual Network (ResNet) model, or an Inception model. The descriptor can specifically be the Histogram of Oriented Gradients (HOG), which is used to describe the edge and texture information of the image.

[0072] In this embodiment, feature extraction is performed on the audio data samples in the historical power emergency consultation environment to obtain the Mel - Frequency Cepstral Coefficients (MFCCs), short - time energy, zero - crossing rate, and Linear Predictive Coding (LPC) of the audio data samples, which are used as audio feature samples; feature extraction is performed on the text data samples in the historical power emergency consultation environment to obtain the sentence vectors of the text data samples, which are used as text feature samples; feature extraction is performed on the on - site image data samples and on - site video data samples in the historical power emergency consultation environment to obtain the feature maps, local feature points, and descriptors of the on - site image data samples and on - site video data samples, which are used as on - site image feature samples and on - site video feature samples. Abundant multi - modal features can be extracted from the audio data samples, text data samples, on - site image data samples, and on - site video data samples, providing high - quality training data for subsequent model training tasks.

[0073] In one embodiment, feature extraction is performed on the historical power emergency event data in the historical power emergency consultation environment to obtain historical event feature samples. The specific steps are as follows: Encode the types, geographical locations, and response measures of each historical power emergency event in the historical power emergency event data in the historical power emergency consultation environment to obtain the event type code, geographical location code, and response measure code corresponding to each historical power emergency event; Extract the time information of each historical power emergency event in the historical power emergency event data to obtain the time information feature corresponding to each historical power emergency event; Quantify the impact range of each historical power emergency event in the historical power emergency event data to obtain the impact range feature corresponding to each historical power emergency event; Use the event type code, geographical location code, response measure code, time information feature, and impact range feature corresponding to each historical power emergency event as historical event feature samples.

[0074] The types (such as equipment failures, natural disasters, or human accidents, etc.), geographical locations, and response measures (such as repair team dispatching, material allocation, or communication coordination, etc.) of each historical power emergency event in the historical power emergency event data in the historical power emergency consultation environment can be converted into numerical or vector representations to obtain the event type code, geographical location code, and response measure code corresponding to each historical power emergency event.

[0075] The time information of each historical power emergency event in the historical power emergency event data, such as the occurrence time, date, and season corresponding to the historical power emergency event, can be extracted to obtain the time information feature corresponding to each historical power emergency event.

[0076] The impact range of each historical power emergency event in the historical power emergency event data, such as the number of affected users and geographical areas, can be quantified to obtain the impact range feature corresponding to each historical power emergency event;

[0077] The event type code, geographical location code, response measure code, time information feature, and impact range feature corresponding to each historical power emergency event can be used as historical event feature samples.

[0078] In this embodiment, using the event type code, geographical location code, response measure code, time information feature, and impact range feature corresponding to each historical power emergency event as historical event feature samples can extract rich emergency feature information from the historical power emergency event data in the historical power emergency consultation environment and provide high-quality training data for subsequent model training tasks.

[0079] In one embodiment, the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample are subjected to feature fusion to obtain a multi-modal feature sample. The specific steps are as follows: Align the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample to the same time axis; splice the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample after time axis alignment and adjust the importance weights of different feature samples to obtain a multi-modal feature sample.

[0080] The audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample can be aligned to the same time axis to ensure that the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample are temporally corresponding and consistent.

[0081] Splice the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample after time axis alignment, and use the attention mechanism to dynamically adjust the importance weights of different feature samples to obtain a multi-modal feature sample, which can improve the model's ability to capture key feature information.

[0082] In this embodiment, the audio feature sample, text feature sample, on-site image feature sample, on-site video feature sample, and historical event feature sample after time axis alignment are spliced and the importance weights of different feature samples are adjusted to obtain a multi-modal feature sample, providing high-quality training data for subsequent model training tasks.

[0083] In an exemplary embodiment, as Figure 3 shown, a speech recognition method in a power emergency consultation environment is provided. Taking the method applied to the Figure 1 terminal 102 as an example, it includes the following steps S401 to S404. Among them:

[0084] Step S401, extract features from the audio data, text data, on-site image data, and on-site video data in the target power emergency consultation environment to obtain audio features, text features, on-site image features, and on-site video features.

[0085] It is possible to collect audio data in the target power emergency consultation environment. Specifically, it is possible to obtain sounds from different directions captured by an array composed of multiple microphones, and the signal-to-noise ratio and positioning accuracy of this audio data are relatively high. It is possible to extract text information from paper documents through keyboard input or Optical Character Recognition (OCR) technology to obtain text data in the target power emergency consultation environment. It is possible to collect on-site image data and on-site video data captured in real time through a camera.

[0086] It is possible to perform data preprocessing on the audio data, text data, on-site image data, and on-site video data in the target power emergency consultation environment. For the specific processing steps, reference can be made to the relevant embodiments in the foregoing training stage.

[0087] It is possible to extract features from the audio data in the target power emergency consultation environment to obtain the Mel Frequency Cepstral Coefficients (MFCCs), short-time energy, zero-crossing rate, and Linear Predictive Coding (LPC) of the audio data as audio features.

[0088] For the specific processing steps to obtain the Mel Frequency Cepstral Coefficients of the audio data, reference can be made to the relevant embodiments in the foregoing training stage.

[0089] It is possible to convert the words in the text data in the target power emergency consultation environment into high-dimensional vector representations according to a pre-constructed word embedding model, calculate the term frequency-inverse document frequency of each word to reflect the importance of the word in the text, and then convert the entire sentence in the text data into a fixed-length vector representation according to a pre-constructed sentence embedding model to obtain the sentence vector of the text data as text features. For longer texts, a document-to-vector (Doc2Vec) model or other document embedding methods can be used to convert the entire text into a vector representation.

[0090] It is possible to extract features from the on-site image data and on-site video data in the target power emergency consultation environment according to a pre-constructed Convolutional Neural Networks (CNN) model to obtain feature maps, local feature points, and descriptors of the on-site image data and on-site video data as on-site image features and on-site video features. Among them, local feature points of the on-site image data and on-site video data can be obtained according to the Scale-Invariant Feature Transform (SIFT) algorithm.

[0091] Step S402: Extract features from the target power emergency event data in the target power emergency consultation environment to obtain target event features.

[0092] The type of the target power emergency event (such as equipment failure, natural disaster or human accident, etc.), geographical location and response measures (such as repair team dispatch, material allocation or communication coordination, etc.) in the target power emergency event data in the target power emergency consultation environment can be converted into numerical or vector representations to obtain the event type code and geographical location code corresponding to the target power emergency event.

[0093] The time information of the target historical power emergency event in the target power emergency event data can be extracted, such as the time, date and season corresponding to the target power emergency event, to obtain the time information feature corresponding to the target power emergency event. The influence range of the target power emergency event in the target power emergency event data can be quantified, such as the number of affected users and geographical area, to obtain the influence range feature corresponding to the target power emergency event; the event type code, geographical location code, time information feature and influence range feature corresponding to the target power emergency event can be used as the target event feature sample.

[0094] Step S403: Perform feature fusion on the audio feature, text feature, on-site image feature, on-site video feature and target event feature to obtain the multi-modal feature.

[0095] The audio feature, text feature, on-site image feature, on-site video feature and historical event feature can be aligned to the same time axis to ensure that the audio feature, text feature, on-site image feature, on-site video feature and historical event feature are corresponding and consistent in time.

[0096] The audio feature, text feature, on-site image feature, on-site video feature and historical event feature after time axis alignment are spliced, and the attention mechanism is used to dynamically adjust the importance weights of different features to obtain the multi-modal feature, which can improve the model's ability to capture key feature information.

[0097] Step S404: Input the multi-modal feature into the speech recognition model to obtain the power emergency information in the target power emergency consultation environment; the speech recognition model is the model obtained according to the embodiment of the acquisition method of the speech recognition model in the above power emergency consultation environment.

[0098] According to the embodiments of the method for obtaining the speech recognition model in the above-mentioned power emergency consultation environment, a speech recognition model can be obtained. Multimodal features can be input into the speech recognition model to obtain the recognition result output by the speech recognition model, and post-processing can be performed on this recognition result, such as removing redundant information and correcting recognition errors. Keywords can be extracted from the recognition result after post-processing, such as "fault", "power outage", and "repair", and specific commands and instructions can be recognized, such as "start the backup power supply" and "dispatch the repair team". The context of the recognition result can also be understood to obtain the power emergency information in the target power emergency consultation environment, ensuring the accuracy, integrity, and reliability of the power emergency information in the target power emergency consultation environment. The power emergency information can include emergency response strategies and resource allocation plans.

[0099] In this embodiment, according to the speech recognition model with strong understanding ability for complex power emergency consultation environments, multi-speaker speech information in complex power emergency consultation environments can be efficiently extracted and accurately recognized to obtain the power emergency information in complex power emergency consultation environments, improving the accuracy and reliability of the emergency information output by speech recognition.

[0100] To better understand the above method, the following elaborates in detail the application embodiments of the method for obtaining the speech recognition model in the power emergency consultation environment of this application and the speech recognition method in the power emergency consultation environment.

[0101] This application embodiment uses deep learning, natural processing technology, and multimodal data processing technology for speech recognition in the power emergency consultation environment. Among them, this application embodiment also uses a speech recognition model trained according to a multimodal large model, enabling computer devices to combine multiple information sources such as text and images while processing speech information, improving the accuracy of information extraction and context understanding ability; that is to say, this application embodiment realizes the efficient extraction and accurate recognition of multi-speaker speech information in complex environments by fusing multiple data sources such as speech, text, and images. Through these technologies, this application embodiment provides a solution for speech extraction and recognition in complex emergency consultation environments in the power system, and can effectively improve the emergency response ability and intelligent level of the power system. This application embodiment improves the processing ability of speech information and the decision-making support level of the power system in emergency situations through speech processing technology and multimodal data fusion technology.

[0102] The main innovations of the embodiments of this application include: adopting a multi-modal large model, which can process various data types such as speech, text, and images simultaneously, and realize the comprehensive understanding and efficient processing of multi-source information in complex power emergency consultation environments. Compared with traditional single-modal speech recognition technologies, the solution provided by the embodiments of this application not only significantly improves the accuracy and anti-noise ability of speech recognition through deep learning algorithms, but also can accurately distinguish different speakers in a multi-speaker environment, effectively solving the limitations of traditional technologies in complex environments. In addition, the solution provided by the embodiments of this application also provides a dedicated post-processing module for optimizing the recognition results, removing redundant information and correcting recognition errors, ensuring the accuracy and reliability of the final output information. Through these processes, the solution provided by the embodiments of this application can better meet the needs of power system emergency response, improve the timeliness and effectiveness of decision-making support, and thus provide strong technical support for the safe and stable operation of the power system.

[0103] The solution provided by the embodiments of this application includes technical links such as multi-modal feature extraction, joint training, real-time speech processing, and result optimization, which can effectively improve the accuracy and robustness of speech information processing in noisy power emergency scenarios, and provide strong technical support for the rapid emergency response and decision-making support of the power system. Specifically, the solution provided by the embodiments of this application mainly includes the following core steps:

[0104] Step 1: Data collection and preprocessing

[0105] 1.1 Data collection

[0106] Data source: Collect multi-modal data from various complex historical emergency consultation environments of the power system. The multi-modal data includes but is not limited to audio data, text data, on-site image data, and on-site video data. Among them, the audio data includes dispatching meeting recordings, on-site report recordings, and fault handling call recordings. The text data includes relevant text records (such as meeting minutes, operation logs). The on-site image data includes images taken on-site. The on-site video data includes videos taken on-site. Data diversity: Ensure that the collected multi-modal data covers various environmental conditions, such as different background noise levels, different numbers of speakers, different device types and qualities, etc., to improve the generalization ability of the model. Data annotation: Annotate the collected audio data, mark the key information (such as commands, warning messages, operation instructions, etc.), and add corresponding labels to the text data, on-site image data, and on-site video data for subsequent supervised learning of the model.

[0107] 1.2 Data preprocessing

[0108] Audio preprocessing:

[0109] Noise reduction: Use audio noise reduction algorithms (such as spectral subtraction, wavelet transform, etc.) to remove background noise and improve the clarity of the speech signal. Format conversion: Convert audio files in different formats into a standard format (such as WAV or MP3) to ensure the consistency of subsequent processing. Segmentation: Split long audio files into shorter segments, with each segment usually ranging from a few seconds to dozens of seconds in length, to facilitate model processing and training. Normalization: Normalize the audio signal so that its amplitude is within a certain range to avoid signals that are too large or too small from affecting the model performance.

[0110] Text preprocessing:

[0111] Cleaning: Remove irrelevant characters and punctuation marks, convert the text to lowercase, perform stemming and stop word filtering to improve the readability and processing efficiency of the text. Tokenization: Split the text into words or phrases and generate word vector representations to facilitate model understanding and processing.

[0112] Image preprocessing:

[0113] Cropping and scaling: Crop and scale the image to a unified size to ensure that the image sizes input to the model are consistent. Denoising: Use image denoising algorithms (such as Gaussian filtering, median filtering, etc.) to remove noise in the image and improve the image quality. Enhancement: Adjust the brightness, contrast, saturation, etc. of the image to enhance the visual effect of the image and improve the recognition ability of the model.

[0114] Step 2: Multimodal feature extraction

[0115] 2.1 Audio feature sample extraction

[0116] Acoustic feature extraction: Specifically, during the model training stage, feature extraction can be performed on audio data samples in the historical power emergency consultation environment to obtain the Mel-frequency cepstral coefficients, short-time energy, zero-crossing rate, and linear predictive coding of the audio data samples as audio feature samples.

[0117] Among them, the Mel-frequency cepstral coefficients (MFCC) of the audio data samples in the historical power emergency consultation environment can be obtained by means of pre-emphasis, framing, windowing, Fourier transform, Mel filter bank, and discrete cosine transform for feature extraction of the audio data samples.

[0118] Pre-emphasis: Perform pre-emphasis processing on the original audio signal in the audio data sample to enhance the high-frequency part and reduce the influence of high-frequency noise. Frame segmentation: Divide the audio signal into multiple short-time frames, usually with a frame length of 20 - 40 milliseconds and a frame shift of 10 - 20 milliseconds. Windowing: Multiply each frame of the signal by a Hamming window or a Hanning window to reduce the discontinuity at the frame boundaries. Fast Fourier Transform (FFT): Perform a fast Fourier transform on each frame of the signal to obtain the frequency-domain representation. Mel filter bank: Pass the frequency-domain signal through a set of Mel filters to convert it to the Mel frequency scale. Logarithmic energy: Take the logarithm of the output of each filter to obtain the logarithmic energy. Discrete Cosine Transform (DCT): Perform a discrete cosine transform on the logarithmic energy to obtain the MFCC coefficients.

[0119] The short-time energy and zero-crossing rate of the audio data sample can provide additional information about the energy variation of the speech signal and the syllable boundaries. The Linear Predictive Coding (LPC) features of the audio data sample can capture the formant information of the speech signal and are applicable to certain specific speech recognition tasks.

[0120] 2.2 Text Feature Sample Extraction

[0121] Based on the pre-constructed word embedding model and sentence embedding model, feature extraction can be performed on the text data sample in the historical power emergency consultation environment to obtain the sentence vector of the text data sample as the text feature sample.

[0122] Word vector representation: Word embedding: Use a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT) to convert the words in the text into high-dimensional vector representations. TF-IDF: Calculate the term frequency-inverse document frequency of each word to reflect the importance of the word in the document.

[0123] Sentence and document representation: Sentence embedding: Use models such as Sentence-BERT to convert the entire sentence into a fixed-length vector representation. Document embedding: For longer texts, methods such as Doc2Vec or other document embedding methods can be used to convert the entire document into a vector representation.

[0124] 2.3 Image Feature Sample Extraction

[0125] Based on the pre-constructed convolutional neural network model, feature extraction is performed on the on-site image data sample in the historical power emergency consultation environment to obtain the feature map, local feature points, and descriptors of the on-site image data sample as the on-site image feature sample.

[0126] Convolutional Neural Network (CNN):

[0127] Pre-trained model: Use pre-trained CNN models (such as VGG, ResNet, Inception, etc.) to extract high-level features of images. Feature map: Extract feature maps from the intermediate layers of the pre-trained model, which contain rich image information. Local features: SIFT: Scale-invariant feature transform, extract local feature points and their descriptors in the image. HOG: Histogram of oriented gradients, extract edge and texture information of the image.

[0128] In the same way, feature extraction can be performed on the on-site video data samples in the historical power emergency consultation environment to obtain the feature maps, local feature points and descriptors of the on-site video data samples, which are used as on-site video feature samples.

[0129] 2.4 Extraction of historical event feature samples

[0130] Encode the types, geographical locations and response measures of each historical power emergency event in the historical power emergency event data in the historical power emergency consultation environment to obtain the event type code, geographical location code and response measure code corresponding to each historical power emergency event; extract the time information of each historical power emergency event in the historical power emergency event data to obtain the time information feature corresponding to each historical power emergency event; quantify the impact range of each historical power emergency event in the historical power emergency event data to obtain the impact range feature corresponding to each historical power emergency event; use the event type code, geographical location code, response measure code, time information feature and impact range feature corresponding to each historical power emergency event as historical event feature samples.

[0131] Event type code: Convert historical power emergency event types (such as equipment failures, natural disasters, human accidents, etc.) into numerical or vector representations. Time feature: Extract the time information when the historical power emergency event occurs, such as hour, date, season, etc. Location feature: Encode the geographical location where the historical power emergency event occurs into a numerical or vector representation. Impact range: Quantify the impact range of the historical power emergency event, such as the number of affected users, geographical area, etc. Response measure code: Encode the response measures in the historical power emergency event (such as repair team dispatching, material allocation, communication coordination, etc.) into numerical or vector representations.

[0132] 2.5 Multi-modal feature sample fusion

[0133] Align the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples and historical event feature samples to the same time axis; splice the audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples and historical event feature samples after time axis alignment and adjust the importance weights of different feature samples to obtain multi-modal feature samples.

[0134] Feature alignment: Align features of different modalities to the same time axis to ensure their temporal correspondence. Feature concatenation: Concatenate feature vectors of different modalities into a high-dimensional feature vector as the input for subsequent models. Multimodal attention mechanism: Use the attention mechanism to dynamically adjust the importance of features of different modalities and improve the model's ability to capture key information.

[0135] 2.6 Standardization of Multimodal Feature Samples

[0136] Normalize the extracted multimodal feature samples so that they are distributed within the same range, preventing a certain modality's feature samples from dominating the model's learning process. Use methods such as principal component analysis (PCA) or t-SNE to reduce the dimensionality of high-dimensional features, reducing computational complexity and improving the model's training efficiency.

[0137] Through the above steps, rich feature samples can be extracted from audio data samples, text data samples, on-site image data samples, on-site video data samples, and historical power emergency event data samples in the historical power emergency consultation environment. These feature samples are then fused together to obtain multimodal feature samples, providing high-quality input data for subsequent model training and recognition tasks. These feature samples can not only capture the internal information of each modality's data samples but also enhance the model's understanding of complex emergency consultation environments through multimodal fusion, thereby improving the accuracy and efficiency of emergency response.

[0138] Step 3: Construction and Training of Speech Recognition Model

[0139] Based on the multimodal feature samples, the deep learning model to be trained can be trained to obtain a speech recognition model.

[0140] 3.1 Model Selection and Design

[0141] Multimodal Fusion Model:

[0142] Architecture selection: Select a deep learning model suitable for multimodal data processing, such as multimodal Transformer, multimodal convolutional neural network (CNN), multimodal recurrent neural network (RNN), etc. Feature fusion layer: Design a dedicated feature fusion layer for fusing feature vectors of different modalities. Common fusion methods include simple concatenation, weighted average, attention mechanism, etc. Output layer: Design an output layer for generating the final emergency response strategy and resource allocation plan. The output layer can be a classifier (for classifying different emergency response measures) or a regressor (for predicting the quantity of resource allocation).

[0143] 3.2 Data Preparation

[0144] Training Set and Validation Set Division: Divide the preprocessed multi-modal feature samples into a training set and a validation set. Usually, the training set accounts for 80% and the validation set accounts for 20%. Data Augmentation: To improve the generalization ability of the model, data augmentation techniques such as random cropping, rotation, flipping, etc. can be used, especially for image data.

[0145] 3.3 Model Training

[0146] Loss Function:

[0147] Classification Task: If the task is to classify different emergency response measures, the Cross-Entropy Loss function can be used. Regression Task: If the task is to predict the quantity of resource allocation, the Mean Squared Error (MSE) loss function can be used. Multi-Task Learning: If the model simultaneously processes classification and regression tasks, a multi-task loss function can be used, which is the weighted sum of multiple loss functions.

[0148] Optimization Algorithm:

[0149] Gradient Descent: Use Stochastic Gradient Descent (SGD) or its variants (such as Adam, RMSprop) to optimize the model parameters. Learning Rate Scheduling: Use learning rate decay strategies such as step decay, exponential decay, etc. to improve the convergence speed and stability of training.

[0150] Regularization:

[0151] L1 / L2 Regularization: Add L1 or L2 regularization terms to the loss function to prevent the model from overfitting.

[0152] Dropout: Use the Dropout technique in some layers of the model to randomly discard a part of the neurons and improve the generalization ability of the model.

[0153] Batch Training: Select an appropriate batch size, usually between 32 and 128, to balance memory usage and training speed.

[0154] Number of Epochs: Set an appropriate number of epochs, usually determined by the performance on the validation set to decide when to stop training.

[0155] 3.4 Model Evaluation

[0156] Performance Metrics:

[0157] Classification tasks: Evaluate the classification performance of the model using metrics such as Accuracy, Precision, Recall, and F1 Score. Regression tasks: Evaluate the regression performance of the model using metrics such as Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE). Multi-task learning: Consider the performance of classification and regression tasks comprehensively and use comprehensive metrics for evaluation.

[0158] Evaluation on the validation set:

[0159] Cross-validation: Use k-fold cross-validation (K-Fold Cross Validation) to evaluate the stability and generalization ability of the model. Early stopping: Monitor the performance of the model on the validation set and terminate the training early when the performance no longer improves to prevent overfitting.

[0160] 3.5 Model Tuning

[0161] Hyperparameter tuning:

[0162] Grid search: Use the Grid Search or Random Search method to try different combinations of hyperparameters and find the optimal hyperparameter configuration. Bayesian optimization: Use the Bayesian optimization method to gradually optimize hyperparameters by constructing a prior distribution of hyperparameters and improve the search efficiency.

[0163] Model ensemble:

[0164] Bagging: Use the Bootstrap Aggregating (Bagging) method to train multiple base models and make predictions by voting or averaging to improve the stability and generalization ability of the model. Boosting: Use the Boosting method, such as AdaBoost and Gradient Boosting, to serially train multiple weak models, gradually reduce the error, and improve the prediction ability of the model.

[0165] 3.6 Model Saving and Deployment

[0166] Model saving:

[0167] Save the best model: During the training process, regularly save the model with the best performance on the validation set. Model compression: Use model compression techniques (such as quantization and pruning) to reduce the size of the model and improve the inference speed.

[0168] Model deployment:

[0169] Export the model: Export the trained model into a standard format (such as ONNX, TensorFlow SavedModel) to facilitate deployment on different platforms. Service-ification: Deploy the model to a server or cloud platform and provide API interfaces for easy invocation by front-end applications.

[0170] Through the above steps, an efficient multi-modal fusion model can be constructed and trained. As a speech recognition model, this model can process multi-modal data in complex emergency consultation environments in the power system and generate accurate emergency response strategies and resource allocation plans. During the training process of the speech recognition model, through reasonable data preparation, loss function selection, optimization algorithms, regularization techniques, batch training, model evaluation and tuning, it is ensured that the speech recognition model has good generalization ability and prediction performance. Finally, the trained speech recognition model is saved and deployed to provide strong technical support for the emergency response of the power system.

[0171] Step 4: Real-time speech extraction and recognition

[0172] 4.1 Real-time data collection

[0173] Audio data collection:

[0174] It is possible to collect audio data in the target power emergency consultation environment. Specifically, it is possible to obtain sounds from different directions captured by an array composed of multiple microphones. The signal-to-noise ratio and positioning accuracy of this audio data are relatively high. The collected audio data is transmitted in real time to the processing system, such as terminal 102, through the network or locally.

[0175] Collection of text data, on-site image data, and on-site video data:

[0176] Text input: Extract text information from paper documents through keyboard input or OCR technology. Image and video input: Use a camera to capture on-site images and on-site videos in real time and transmit them to the processing system through the network.

[0177] 4.2 Real-time preprocessing

[0178] Audio data preprocessing:

[0179] Noise reduction: Apply real-time noise reduction algorithms, such as spectral subtraction and wavelet transform, to reduce the influence of background noise. Frame splitting: Split the real-time audio stream into short-time frames, usually with a frame length of 20 - 40 milliseconds and a frame shift of 10 - 20 milliseconds. Normalization: Normalize each frame of the audio signal to ensure that the signal amplitude is within a certain range.

[0180] Text data preprocessing:

[0181] Cleaning: Remove irrelevant characters and punctuation in real time, convert the text to lowercase, perform stemming, and filter stop words. Tokenization: Split the text into words or phrases and generate a word vector representation.

[0182] Preprocessing of on-site image data and on-site video data:

[0183] Cropping and scaling: Crop and scale real-time images to a unified size. Denoising: Apply image denoising algorithms in real time, such as Gaussian filtering, median filtering, etc., to improve image quality. Enhancement: Adjust the brightness, contrast, saturation, etc. of the image to enhance the visual effect of the image.

[0184] 4.3 Feature extraction

[0185] Audio feature extraction:

[0186] MFCC: Calculate the Mel-frequency cepstral coefficients (MFCC) of each audio frame in real time. Short-time energy and zero-crossing rate: Calculate the short-time energy and zero-crossing rate of each audio frame in real time. LPC: Calculate the linear predictive coding (LPC) features of each audio frame in real time.

[0187] Text feature extraction:

[0188] Word embedding: Use a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT) to convert words in the text into high-dimensional vector representations. Sentence embedding: Use models such as Sentence-BERT to convert the entire sentence into a fixed-length vector representation.

[0189] On-site image feature extraction:

[0190] CNN features: Use pre-trained convolutional neural networks (such as VGG, ResNet, Inception, etc.) to extract high-level features of the image. Local features: Extract local feature points and their descriptors in the image in real time, such as SIFT, HOG, etc.

[0191] 4.4 Multimodal feature fusion

[0192] Feature alignment: Align features of different modalities to the same time axis to ensure their temporal correspondence. Feature concatenation: Concatenate feature vectors of different modalities into a high-dimensional feature vector as the input for subsequent models. Multimodal attention mechanism: Use the attention mechanism to dynamically adjust the importance of features of different modalities and improve the model's ability to capture key information.

[0193] 4.5 Speech recognition

[0194] Model Loading: Load the trained multi-modal fusion model, i.e., the speech recognition model. Real-time Inference: Input the real-time extracted multi-modal features into the speech recognition for real-time inference to generate recognition results. Result Processing: Post-process the recognition results output by the speech model, such as removing redundant information and correcting recognition errors, to obtain the power emergency information in the target power emergency consultation environment, ensuring the accuracy and reliability of the finally output power emergency information in the target power emergency consultation environment.

[0195] 4.6 Emergency Information Extraction

[0196] Keyword Extraction: Extract keyword vocabularies from the recognition results, such as "fault", "power outage", "repair", etc. Command Recognition: Identify specific commands and instructions from the recognition results, such as "start the backup power supply", "dispatch the repair team", etc. Context Understanding: Combine multi-modal information to understand the context of the recognition results to ensure the accuracy and integrity of the information.

[0197] 4.7 Result Output

[0198] Real-time Display: Display the recognition results on the user interface in real time for easy viewing by the operators. Log Recording: Record the recognition results and the processing process into a log file for subsequent analysis and auditing. Alarm Notification: For critical emergency information, trigger an alarm notification in real time to remind relevant personnel to take actions.

[0199] Through the above steps, real-time speech extraction and recognition in the power emergency consultation environment can be achieved, extracting key information from the multi-modal data in the complex power emergency consultation environment of the power system and generating accurate recognition results. Links such as real-time data collection, preprocessing, feature extraction, multi-modal fusion, speech recognition, emergency information extraction, and result output ensure the efficiency and accuracy of the system, providing timely and effective support for the emergency response of the power system.

[0200] Generally speaking, the embodiment of this application provides a voice extraction and recognition solution for the power system in a complex emergency consultation environment based on a multimodal large model. By integrating various data sources such as voice, text, and images, and using deep learning technology, it realizes the efficient extraction and accurate recognition of multi-speaker voice information. This solution first conducts data collection and preprocessing work to ensure the quality and consistency of multimodal data; then, it adopts a series of feature extraction technologies, such as Mel Frequency Cepstral Coefficients, word embeddings, convolutional neural networks, etc., to extract rich features from each modal data; then, through multimodal feature fusion technologies, such as feature alignment, splicing, and multimodal attention mechanisms, it organically combines features of different modalities, further enhancing the model's understanding ability of complex scenarios. In addition, this solution also realizes the efficient construction and training of the model, and through reasonable design and optimization, ensures the generalization ability and prediction performance of the model; moreover, this solution performs excellently in real-time voice extraction and recognition, and can quickly and accurately extract key emergency information from multimodal data, providing timely and effective technical support for the emergency response of the power system. Generally speaking, through the deep integration and processing of multimodal data, this solution not only greatly improves the recognition accuracy of voice information in emergency scenarios, but also has strong real-time performance and adaptability, can effectively support the rapid emergency response and decision-making support of the power system, and significantly improves the efficiency and quality of emergency handling.

[0201] The advantages of the solution provided by the embodiment of this application are as follows:

[0202] 1. Multimodal data fusion improves recognition accuracy

[0203] The solution provided by the embodiment of this application fuses data of multiple modalities such as audio, text, and images, and uses a multimodal large model for feature extraction and fusion, significantly improving the recognition accuracy of voice information in a complex power emergency consultation environment. Traditional single-modal recognition methods have limitations in processing multi-source information, while the solution provided by the embodiment of this application can comprehensively understand the complex situations in the power emergency scenario through the comprehensive processing of multimodal data, effectively improving the accuracy and reliability of voice information extraction.

[0204] 2. Strong real-time performance and adaptability

[0205] The solution provided by the embodiment of this application realizes the full-process automated processing from data collection, preprocessing, feature extraction to real-time voice extraction and recognition, ensuring the real-time performance and efficiency of emergency response. Especially in a complex and changeable power emergency environment, through real-time data collection and dynamic adjustment modules, it can update response measures and resource allocation plans in a timely manner according to the actual situation, ensuring the effectiveness and timeliness of response measures. This strong real-time performance and adaptability enable the solution provided by the embodiment of this application to better support the emergency response needs of the power system.

[0206] 3. Strong data processing and preprocessing capabilities

[0207] The solution provided by the embodiments of this application adopts a variety of technologies in the data collection and preprocessing stages, such as audio noise reduction, text cleaning, image denoising, etc., ensuring the quality and consistency of the input data. These preprocessing steps not only improve the usability of the data but also lay a solid foundation for subsequent feature extraction and model training. Through efficient data processing and preprocessing, it is possible to better handle various complex data sources and ensure stability and reliability in practical applications.

[0208] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0209] Based on the same inventive concept, the embodiments of this application also provide an apparatus for obtaining a speech recognition model in a power emergency consultation environment for implementing the method for obtaining a speech recognition model in the power emergency consultation environment involved above. The solution provided by this apparatus for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the apparatus for obtaining a speech recognition model in a power emergency consultation environment provided below can refer to the limitations on the method for obtaining a speech recognition model in a power emergency consultation environment in the above text, and will not be repeated here.

[0210] In an exemplary embodiment, as Figure 4 shown, an apparatus for obtaining a speech recognition model in a power emergency consultation environment is provided, where:

[0211] A feature sample acquisition module 401 is configured to extract features from audio data samples, text data samples, on-site image data samples, and on-site video data samples in a historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples;

[0212] The feature sample acquisition module 401 is further configured to extract features from historical power emergency event data in a historical power emergency consultation environment to obtain historical event feature samples;

[0213] The multi-modal feature sample acquisition module 402 is configured to perform feature fusion on the audio feature samples, the text feature samples, the on-site image feature samples, the on-site video feature samples, and the historical event feature samples to obtain multi-modal feature samples;

[0214] The speech recognition model acquisition module 403 is configured to perform model training on a deep learning model to be trained according to the multi-modal feature samples to obtain a speech recognition model.

[0215] In one embodiment, the feature sample acquisition module 401 is further configured to: extract features from the audio data samples in the historical power emergency consultation environment to obtain the Mel frequency cepstral coefficients, short-time energy, zero-crossing rate, and linear predictive coding of the audio data samples as audio feature samples; extract the sentence vectors of the text data samples in the historical power emergency consultation environment as text feature samples according to the pre-constructed word embedding model and sentence embedding model; extract the feature maps, local feature points, and descriptors of the on-site image data samples and on-site video data samples in the historical power emergency consultation environment as on-site image feature samples and on-site video feature samples according to the pre-constructed convolutional neural network model.

[0216] In one embodiment, the feature sample acquisition module 401 is further configured to: encode the types, geographical locations, and response measures of each historical power emergency event in the historical power emergency event data in the historical power emergency consultation environment to obtain the event type code, geographical location code, and response measure code corresponding to each historical power emergency event; extract the time information of each historical power emergency event in the historical power emergency event data to obtain the time information feature corresponding to each historical power emergency event; quantify the impact range of each historical power emergency event in the historical power emergency event data to obtain the impact range feature corresponding to each historical power emergency event; and use the event type code, geographical location code, response measure code, time information feature, and impact range feature corresponding to each historical power emergency event as historical event feature samples.

[0217] In one embodiment, the multimodal feature sample acquisition module 402 is further configured to: align the audio feature sample, the text feature sample, the on-site image feature sample, the on-site video feature sample, and the historical event feature sample to the same time axis; splice the audio feature sample, the text feature sample, the on-site image feature sample, the on-site video feature sample, and the historical event feature sample after time axis alignment and adjust the importance weights of different feature samples to obtain multimodal feature samples.

[0218] Each module in the above device for obtaining a speech recognition model in a power emergency consultation environment can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0219] Based on the same inventive concept, an embodiment of the present application further provides a speech recognition device in a power emergency consultation environment for implementing the above-mentioned speech recognition method in a power emergency consultation environment. The solution provided by the device for solving the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the speech recognition device in a power emergency consultation environment provided below can refer to the limitations on the speech recognition method in a power emergency consultation environment in the above text, and will not be repeated here.

[0220] In an exemplary embodiment, as Figure 5 shown, a speech recognition device in a power emergency consultation environment is provided, where:

[0221] The feature acquisition module 501 is configured to extract features from audio data, text data, on-site image data, and on-site video data in a target power emergency consultation environment to obtain audio features, text features, on-site image features, and on-site video features;

[0222] The feature acquisition module 501 is further configured to extract features from target power emergency event data in a target power emergency consultation environment to obtain target event features;

[0223] The multimodal feature acquisition module 502 is configured to perform feature fusion on the audio features, the text features, the on-site image features, and the on-site video features to obtain multimodal features;

[0224] The voice recognition module 503 is configured to input the audio features, text features, on-site image features, and on-site video features into a voice recognition model to obtain power emergency information in the target power emergency consultation environment; the voice recognition model is the model obtained by the method embodiment of the voice recognition model in the above-mentioned power emergency consultation environment.

[0225] Each module in the above-mentioned voice recognition device in the power emergency consultation environment can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in the form of hardware or independent of the processor, or stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0226] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the method embodiments of obtaining the voice recognition model in the power emergency consultation environment and the data of the method embodiments of the voice recognition method in the power emergency consultation environment. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for obtaining a voice recognition model in a power emergency consultation environment and a voice recognition method in a power emergency consultation environment.

[0227] Those skilled in the art can understand that Figure 6 the structure shown in

[0228] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0229] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0230] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0231] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0232] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0233] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope recorded in the present application.

[0234] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for acquiring a speech recognition model in an electric power emergency consultation environment, characterized in that: The method comprises: Extract features from audio data samples, text data samples, on-site image data samples, and on-site video data samples in a historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples, and on-site video feature samples; Extract features from historical power emergency event data in the historical power emergency consultation environment to obtain historical event feature samples; Performing feature fusion on the audio feature samples, the text feature samples, the on-site image feature samples, the on-site video feature samples, and the historical event feature samples to obtain a multimodal feature sample; According to the multimodal feature samples, model training is performed on the deep learning model to be trained to obtain a speech recognition model.

2. The method according to claim 1, characterized in that The feature extraction of audio data samples, text data samples, on-site image data samples and on-site video data samples in the historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples and on-site video feature samples includes: Extracting features of audio data samples in a historical power emergency consultation environment, and obtaining Mel-frequency cepstrum coefficients, short-time energy, zero-crossing rate, and linear prediction coding of the audio data samples as audio feature samples; According to the pre-built word embedding model and sentence embedding model, feature extraction is performed on the text data samples in the historical power emergency consultation environment to obtain the sentence vectors of the text data samples as text feature samples; According to the pre-built convolutional neural network model, feature extraction is performed on the on-site image data samples and on-site video data samples in the historical power emergency consultation environment to obtain feature maps, local feature points and descriptors of the on-site image data samples and on-site video data samples as on-site image feature samples and on-site video feature samples.

3. The method according to claim 1, characterized in that The feature extraction of the historical power emergency event data in the historical power emergency consultation environment to obtain the historical event feature samples includes: Encode the type, geographical location and response measures of each historical power emergency event in the historical power emergency event data under the historical power emergency consultation environment to obtain the event type code, geographical location code and response measure code corresponding to each historical power emergency event; Extracting the time information of each historical power emergency event in the historical power emergency event data to obtain the time information characteristics corresponding to each historical power emergency event; Quantifying the impact range of each historical power emergency event in the historical power emergency event data to obtain the impact range characteristics corresponding to each historical power emergency event; The event type code, geographical location code, response measure code, time information characteristics and impact range characteristics corresponding to each historical power emergency event are used as historical event feature samples.

4. The method according to claim 1, characterized in that The step of fusing the audio feature samples, the text feature samples, the on-site image feature samples, the on-site video feature samples and the historical event feature samples to obtain a multimodal feature sample includes: Aligning the audio feature samples, the text feature samples, the on-site image feature samples, the on-site video feature samples, and the historical event feature samples to the same timeline; The audio feature samples, text feature samples, on-site image feature samples, on-site video feature samples and historical event feature samples that are aligned on the time axis are spliced ​​and the importance weights of different feature samples are adjusted to obtain multimodal feature samples.

5. A speech recognition method in an electric power emergency consultation environment, characterized in that: The method comprises: Extract features of audio data, text data, on-site image data and on-site video data in the target power emergency consultation environment to obtain audio features, text features, on-site image features and on-site video features; Extract features of target power emergency event data in the target power emergency consultation environment to obtain target event features; Performing feature fusion on the audio feature, the text feature, the on-site image feature, the on-site video feature and the target event feature to obtain a multimodal feature; The multimodal features are input into a speech recognition model to obtain power emergency information in a target power emergency consultation environment; the speech recognition model is a model obtained according to the method according to any one of claims 1 to 4.

6. A device for acquiring a speech recognition model in an electric power emergency consultation environment, characterized in that: The device comprises: A feature sample acquisition module is used to extract features from audio data samples, text data samples, on-site image data samples and on-site video data samples in a historical power emergency consultation environment to obtain audio feature samples, text feature samples, on-site image feature samples and on-site video feature samples; The feature sample acquisition module is also used to extract features from historical power emergency event data in a historical power emergency consultation environment to obtain historical event feature samples; A multimodal feature sample acquisition module, used for performing feature fusion on the audio feature sample, the text feature sample, the on-site image feature sample, the on-site video feature sample and the historical event feature sample to obtain a multimodal feature sample; The speech recognition model acquisition module is used to perform model training on the deep learning model to be trained according to the multimodal feature samples to obtain a speech recognition model.

7. A speech recognition device in an electric power emergency consultation environment, characterized in that: The device comprises: A feature acquisition module is used to extract features of audio data, text data, on-site image data and on-site video data in the target power emergency consultation environment to obtain audio features, text features, on-site image features and on-site video features; The feature acquisition module is also used to extract features of target power emergency event data in the target power emergency consultation environment to obtain target event features; A multimodal feature acquisition module, used for fusing the audio feature, the text feature, the on-site image feature and the on-site video feature to obtain a multimodal feature; A speech recognition module is used to input the audio features, text features, on-site image features and on-site video features into a speech recognition model to obtain power emergency information in a target power emergency consultation environment; the speech recognition model is a model obtained according to the method described in any one of claims 1 to 4.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Data identification method and system based on neural network model and application

    CN120452432A