Low-altitude base intelligent operation and maintenance system command center voice data processing method and device

By setting up a microphone array in the command center of the intelligent operation and maintenance system of the low-altitude base and performing data processing, the problem of drone voice recognition being affected by noise is solved, and efficient and accurate voice recognition and drone control are achieved.

CN120236565AInactive Publication Date: 2025-07-01HEZONG KAIDA POWER TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724981.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the drone is flying outdoors, it is affected by factors such as wind noise and mechanical noise, resulting in a decrease in voice recognition effect. The accuracy and efficiency of the speech recognition system of the command center of the existing low-altitude base intelligent operation and maintenance system need to be improved.

Method used

By setting up a microphone array in the command center of the low-altitude base intelligent operation and maintenance system, it collects voice data, and filters, segments, windows and detection processing, then performs frequency domain transformation and feature extraction to determine the pronouncing of voice feature sequence and word transfer, and finally generates drone control instructions.

Benefits of technology

It improves the accuracy and efficiency of voice recognition, ensures that the drone can respond to voice commands in a timely manner, and improves the accuracy and efficiency of drone control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236565A_ABST
    Figure CN120236565A_ABST
Patent Text Reader

Abstract

The invention provides a low-altitude base intelligent operation and maintenance system command center voice data processing method and device. The method comprises the following steps: acquiring voice data through a microphone array arranged in a command center of a low-altitude base intelligent operation and maintenance system; the microphone array is arranged according to a preset layout; preprocessing the voice data to obtain preprocessed data; performing feature extraction on the preprocessed data to obtain feature data; determining a voice feature sequence and a word transition probability according to the feature data; and obtaining a speech recognition result according to the speech feature sequence and the word transition probability. The voice recognition accuracy and efficiency of the command center of the low-altitude base intelligent operation and maintenance system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and also relates to a method and device for processing voice data in a command center of a low-altitude base intelligent operation and maintenance system. Background Art

[0002] Based on people's requirements for highly intelligent control, there are currently drones controlled by voice on the market. Although the drone voice recognition technology brings convenience to drone operation, there are still some significant drawbacks: when the drone flies outdoors, it may be affected by various factors such as wind noise and mechanical noise, resulting in a decline in the voice recognition effect; voice recognition requires a certain amount of processing time, which causes the drone to be unable to respond immediately after receiving a voice command, and the processing efficiency is low; in addition, the voice recognition system in the existing command center of the low-altitude base intelligent operation and maintenance system needs to improve the accuracy and recognition efficiency of user voice recognition. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for processing voice data in a command center of a low-altitude base intelligent operation and maintenance system to improve the efficiency and accuracy of voice recognition.

[0004] To solve the above technical problem, the technical solution of the present invention is as follows:

[0005] In a first aspect of the present invention, there is provided a method for processing voice data in a command center of a low-altitude base intelligent operation and maintenance system, including:

[0006] Collecting voice data through a microphone array placed in the command center of the low-altitude base intelligent operation and maintenance system; the microphone array is set according to a preset layout;

[0007] Preprocessing the voice data to obtain preprocessed data;

[0008] Extracting features from the preprocessed data to obtain feature data;

[0009] Determining a voice feature sequence and a word transition probability according to the feature data;

[0010] Obtaining a voice recognition result according to the voice feature sequence and the word transition probability;

[0011] Generating a drone control command according to the voice recognition result.

[0012] Optionally, preprocessing the voice data to obtain preprocessed data includes:

[0013] Performing filtering processing on the voice data to obtain filtered data;

[0014] Perform segmentation processing on the filtered data to obtain segmented data;

[0015] Perform windowing processing on the segmented data to obtain windowed data;

[0016] Perform detection processing on the windowed data to obtain preprocessed data.

[0017] Optionally, perform feature extraction on the preprocessed data to obtain feature data, including:

[0018] By Perform frequency-domain transformation on the preprocessed data to obtain frequency-domain data; where is the spectrum of the i-th frame of speech data in the preprocessed data, , is the i-th frame of preprocessed data, n is the sampling time, n = 0, 1, ……, N - 1, N is the frame length, i = 1, 2, ……; k is the index of the frequency component, and the value range is , j is the imaginary unit, satisfying , e is the base of the natural logarithm, is the frequency-domain data of the i-th frame of speech data;

[0019] Perform feature extraction on the frequency-domain data to obtain feature data.

[0020] Optionally, perform feature extraction on the frequency-domain data to obtain feature data, including:

[0021] According to , determine the log energy;

[0022] According to , determine the first coefficient;

[0023] According to , determine the second coefficient;

[0024] According to , determine the third coefficient;

[0025] According to the first coefficient, the second coefficient and the third coefficient, obtain feature data;

[0026] Where is the log energy output by the m-th triangular band-pass filter, N is the frame length, k is the index of the frequency component, and the value range is , is the m-th triangular band-pass filter, , M is the number of triangular band-pass filters, is the first coefficient at the a-th position, a = 0, 1, ……, L - 1, L is an integer from 12 to 16, is the second coefficient, is the first coefficient at the (a + 1)-th position, is the first coefficient at the (a - 1)-th position, is the third coefficient, is the second coefficient at the (a + 1)-th position, is the second coefficient at the (a - 1)-th position.

[0027] Optionally, according to the feature data, determining a speech feature sequence, including:

[0028] Obtaining a state transition probability matrix and an observation probability matrix of a preset network model;

[0029] According to the state transition probability matrix, the observation probability matrix, and the feature data, determining forward variables and backward variables;

[0030] According to the forward variables and the backward variables, determining the total probability;

[0031] According to the total probability, determining the speech feature sequence.

[0032] Optionally, according to the feature data, determining a word transition probability, including:

[0033] According to the feature data, determining an updated hidden state;

[0034] According to the updated hidden state, determining an output vector;

[0035] According to the output vector, obtaining a probability distribution;

[0036] According to the probability distribution, determining the word transition probability.

[0037] Optionally, according to the speech feature sequence and the word transition probability, obtaining a speech recognition result, including:

[0038] According to the speech feature sequence, obtaining an acoustic probability;

[0039] According to the word transition probability, obtaining a language probability;

[0040] According to the acoustic probability and the language probability, obtaining the speech recognition result.

[0041] In a second aspect of the present invention, there is provided a voice data processing device for a command center of a low-altitude base intelligent operation and maintenance system, including:

[0042] An acquisition module, configured to acquire voice data through a microphone array disposed in the command center of the low-altitude base intelligent operation and maintenance system; the microphone array is arranged according to a preset layout;

[0043] A processing module, configured to preprocess the voice data to obtain preprocessed data; extract features from the preprocessed data to obtain feature data; determine a voice feature sequence and a word transition probability according to the feature data; obtain a voice recognition result according to the voice feature sequence and the word transition probability; and generate a drone control instruction according to the voice recognition result.

[0044] In a third aspect of the present invention, there is provided a computing device, including: a processor and a memory storing a computer program, and when the computer program is run by the processor, the method described in the first aspect is executed.

[0045] In a fourth aspect of the present invention, there is provided a computer-readable storage medium storing instructions, and when the instructions are run on a computer, the computer is caused to execute the method described in the first aspect.

[0046] The above solutions of the present invention at least include the following beneficial effects:

[0047] In the above solution of the present invention, voice data is collected through a microphone array of the command center of the low-altitude base intelligent operation and maintenance system, and then the voice data is preprocessed and feature-extracted to obtain feature data. Then, a voice feature sequence and a word transition probability are determined according to the feature data. Finally, a voice recognition result is obtained according to the voice feature sequence and the word transition probability, improving the accuracy and efficiency of voice recognition and the timeliness of issuing drone control instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic flowchart of a method for processing voice data of the command center of the low-altitude base intelligent operation and maintenance system in an embodiment of the present invention;

[0049] Figure 2 is a schematic structural diagram of a device for processing voice data of the command center of the low-altitude base intelligent operation and maintenance system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0051] As Figure 1 shown, an embodiment of the present invention proposes a method for processing voice data of the command center of a low-altitude base intelligent operation and maintenance system, including the following steps:

[0052] Step 101, collect voice data through a microphone array placed in the command center of the intelligent operation and maintenance system for the low-altitude base; the microphone array is set according to a preset layout;

[0053] Step 102, preprocess the voice data to obtain preprocessed data;

[0054] Step 103, extract features from the preprocessed data to obtain feature data;

[0055] Step 104, determine a voice feature sequence and a word transition probability according to the feature data;

[0056] Step 105, obtain a voice recognition result according to the voice feature sequence and the word transition probability;

[0057] Step 106, generate a drone control instruction according to the voice recognition result.

[0058] In the method for processing voice data in the command center of the intelligent operation and maintenance system for the low-altitude base according to the embodiment of the present invention, voice data is collected through a microphone array, and then the voice data is preprocessed and feature-extracted to obtain feature data. Then, a voice feature sequence and a word transition probability are determined according to the feature data, and a voice recognition result is obtained according to the voice feature sequence and the word transition probability. Finally, a drone control instruction is generated according to the voice recognition result, improving the accuracy and efficiency of voice recognition.

[0059] In an optional embodiment of the present invention, collecting voice data through a microphone array placed in the command center of the intelligent operation and maintenance system for the low-altitude base in Step 101 includes:

[0060] Step 1011, obtain a preset layout;

[0061] Specifically, according to the preset layout of the microphones in the command center of the low-altitude base intelligent operation and maintenance system, it includes the microphone type, the layout of the microphone array, the number of microphones, and the spacing. The microphone type can be a dynamic microphone or a condenser microphone. Dynamic microphones are suitable for collecting voice data in noisy environments, while condenser microphones provide clearer sound quality and are suitable for collecting voice data in quiet environments. The layout of the microphone array can be a linear array, a circular array, a planar array, or a spatial array. Among them, the linear array is suitable for one-dimensional sound source localization (a linear array means that the microphones are arranged in a straight line, such as voice interaction devices like smart speakers and mobile phones), the circular array (the microphones are arranged in a circular layout, suitable for 360° omnidirectional sound pickup, such as remote meetings, smart assistants, etc.), and the planar array and spatial array can be used for two-dimensional or three-dimensional sound source localization (such as robots, smart home devices, VR / AR, etc.). The number and spacing of the microphones in the array will affect the accuracy and robustness of sound source localization. Increasing the number of microphones will improve data accuracy but also increase the computational complexity. Therefore, the number of microphones can be determined according to the accuracy required by the user. In a specific embodiment, if the user's requirement is drone voice recognition, the corresponding preset layout includes a dynamic microphone as the microphone type, a circular array as the layout of the microphone array, 6 microphones evenly distributed on the circumference, and the spacing between the microphones is between 20 mm and 60 mm in radius (preferably 39 mm).

[0062] Step 1012, determine the microphone array according to the preset layout;

[0063] Specifically, install the microphones on the device in use (such as the device in the command center of the low-altitude base intelligent operation and maintenance system) according to the preset positions and spacing. Ensure the accurate relative positions of the microphones to avoid performance degradation caused by installation errors. In addition, microphone calibration is also required to ensure that parameters such as the sensitivity and frequency response of each microphone are consistent, so as to improve the overall performance and accuracy of the microphone array.

[0064] Step 1013, collect voice data in real time through the microphone array.

[0065] Specifically, after the microphones are installed according to the preset layout, the microphone array can be used to collect voice data in the command center of the low-altitude base intelligent operation and maintenance system in real time. It should be noted that the voice data collected in real time is a sequence of voice signals.

[0066] In an alternative embodiment of the present invention, step 102 includes:

[0067] Step 1021, perform filtering processing on the voice data to obtain filtered data;

[0068] Specifically, through Filter the speech data to obtain filtered data, where is the filtered data, is the speech amplitude value at the nth sampling moment in the speech data, is the filter coefficient, and its value range is from 0.95 to 0.98, is the speech amplitude value at the (n - 1)th sampling moment in the speech data. By filtering the speech data, the high-frequency part in the speech data can be enhanced, and the high-frequency resolution can be improved.

[0069] Step 1022: Segment the filtered data to obtain segmented data;

[0070] Specifically, since the speech signal (i.e., the speech data) in the filtered data has short-term stationarity, that is, within a short period of time (usually 10 to 30 ms), the characteristics of the speech signal basically remain unchanged. Therefore, the filtered data can be divided into several frames to facilitate short-term analysis of each frame. In this embodiment, a window with a fixed size slides on the speech signal, and each time it slides a certain frame shift, so that the filtered data is segmented into multiple overlapping frames (the overlap rate can be 50%), that is, the segmented data includes multiple overlapping frames.

[0071] Step 1023: Window the segmented data to obtain windowed data;

[0072] Specifically, by windowing the segmented data, windowed data is obtained. Where is the windowing weight, , n is the sampling moment, n = 0, 1, ……, N - 1, N is the frame length, is the windowed data of the i-th frame, is the segmented data of the i-th frame. Windowing makes the segmented data gradually decay to zero at the edges of the frame, reduces spectral leakage, and improves the accuracy of the speech data.

[0073] Step 1024: Detect the windowed data to obtain preprocessed data.

[0074] Specifically, first, according to E i = ∑ n=0 N-1 [ s i n ] 2 、 Z i = 1 2 ∑ n=1 N-1 sgn s i n -sgn s i n-1 ] determine the short-time energy and short-time zero-crossing rate of the windowed data of the i-th frame; where is the short-time energy of the windowed data of the i-th frame, is the windowed data of the i-th frame, n is the sampling time, n = 0, 1, ……, N - 1, N is the frame length, i = 1, 2, ……; is the short-time zero-crossing rate of the windowed data of the i-th frame, sgn[∙] is the sign function. When then and when then . Here, the short-time energy reflects the amplitude of the speech data (or speech signal) in the windowed data. The energy of the speech segment is usually larger than that of the silent segment; the short-time zero-crossing rate reflects the frequency variation of the speech data (or speech signal) in the windowed data. The zero-crossing rate of unvoiced sounds is relatively high, the zero-crossing rate of voiced sounds is relatively low, and the zero-crossing rate of the silent segment is also relatively low.

[0075] Then, according to the first preset short-time energy threshold, the second preset short-time energy threshold, the preset short-time zero-crossing rate threshold, the calculated short-time energy and short-time zero-crossing rate, the starting point and the ending point are determined. Specifically: starting from the starting end of the windowed data, when the short-time energy exceeds the first preset short-time energy threshold, it is considered that the speech segment may be entered, and the frame number at this time is used as the starting point of the speech segment; then search backward. When the short-time energy is lower than the second preset short-time energy threshold (the second preset short-time energy threshold is less than the first preset short-time energy threshold) and the short-time zero-crossing rate is lower than the preset short-time zero-crossing rate threshold, it is considered that the speech segment ends, and the frame number at this time is used as the ending point of the speech segment. The purpose of determining the starting point and the ending point is to remove the silent segment and improve the efficiency of subsequent speech recognition.

[0076] Finally, the windowed data, the starting point, and the ending point are used as preprocessing data for subsequent speech recognition.

[0077] In an optional embodiment of the present invention, step 103 includes:

[0078] Step 1031, through performs a frequency-domain transformation on the preprocessing data to obtain frequency-domain data; where is the spectrum of the i-th frame of speech data in the preprocessing data, , is the i-th frame of preprocessing data, n is the sampling time, n = 0, 1, ……, N - 1, N is the frame length, i = 1, 2, ……; k is the index of the frequency component, and the value range is , j is the imaginary unit, satisfying , e is the base of the natural logarithm, approximately equal to 2.71828, is the frequency-domain data of the i-th frame of speech data. Speech data appears as a waveform that changes over time in the time domain, while in the frequency domain, it appears as the distribution of different frequency components. By converting the preprocessed data from the time domain to the frequency domain, its frequency characteristics can be revealed and subsequent feature extraction can be performed.

[0079] Step 1032: Extract features from the frequency-domain data to obtain feature data.

[0080] In an optional embodiment of the present invention, step 1032 includes:

[0081] Step 10321: Determine the log energy according to the frequency-domain data;

[0082] Step 10322: Determine the first coefficient according to the log energy;

[0083] Step 10323: Determine the second coefficient and the third coefficient according to the first coefficient;

[0084] Step 10324: Obtain feature data according to the first coefficient, the second coefficient, and the third coefficient.

[0085] Specifically, according to the log energy is obtained, and then according to the first coefficient is obtained, and then according to the second coefficient is obtained, according to the third coefficient is obtained, and finally the first coefficient, the second coefficient, and the third coefficient are concatenated in sequence to obtain the speech feature vector, that is, the feature data.

[0086] Wherein, is the log energy output by the m-th triangular band-pass filter, N is the frame length, k is the index of the frequency component, and the value range is , is the frequency-domain data of the i-th frame, is the m-th triangular band-pass filter, , M is the number of triangular band-pass filters, and M usually takes values from 22 to 26, is the first coefficient at the a-th position, a = 0, 1,..., L - 1, and L is generally an integer from 12 to 16, is the second coefficient, is the first coefficient at the a + 1-th position, is the first coefficient at the a - 1-th position, is the third coefficient, is the second coefficient at the a + 1-th position, is the second coefficient at the a - 1-th position.

[0087] In an alternative embodiment of the present invention, determining the speech feature sequence according to the feature data in step 104 includes:

[0088] Step 10411, obtaining the state transition probability matrix and the observation probability matrix of the preset network model;

[0089] Step 10412, determining the forward variable and the backward variable according to the state transition probability matrix, the observation probability matrix and the feature data;

[0090] Specifically, according to α t i =[ ∑ j=1 F α t-1 j α ji ] b i ( o t ) , the forward variable is determined, and according to β t i = ∑ j=1 F α ij b j ( o t+1 )] β t+1 j , the backward variable is determined;

[0091] Step 10413, determining the total probability according to the forward variable and the backward variable;

[0092] Specifically, according to , the total probability is determined.

[0093] Step 10414, determining the speech feature sequence according to the total probability.

[0094] Wherein, is the forward variable, indicating the probability of partial feature data and the hidden state at time t; t = 1, 2,..., T, T is the length of the feature vector in the feature data, F is the number of hidden states, and i = 1, 2,..., F; is the forward probability of the hidden state at time t - 1, is an element in the state transition probability matrix A, indicating the probability of transitioning from the hidden state to the hidden state , is an element in the observation probability matrix B, indicating the probability of generating the feature data under the hidden state ; is the backward variable, indicating the probability of generating the remaining feature data starting from the hidden state at time t, is an element in the state transition probability matrix A, representing the probability of transitioning from the hidden state to the hidden state . is an element in the observation probability matrix B, representing the probability of generating the feature data under the hidden state . is the probability of generating the remaining feature data starting from the hidden state at time t + 1. . is the total probability, O is the feature data, , are the parameters of the preset network model, that is, the state transition probability matrix and the observation probability matrix of the preset network model.

[0095] Here, the hidden state sequence that maximizes the total probability is used as the speech feature sequence.

[0096] In an optional embodiment of the present invention, determining the word transition probability according to the feature data in step 104 includes:

[0097] Step 10421: Determine the updated hidden state according to the feature data;

[0098] Specifically, by inputting the feature vector in the feature data into the following formula, the updated hidden state can be calculated:

[0099] ;

[0100] ;

[0101] ;

[0102] ;

[0103] ;

[0104] ;

[0105] where is the input gate, is the forget gate, is the output gate, is the candidate hidden state, is the cell state, is the updated hidden state, , , , The first weight matrices for the input gate, forget gate, output gate, and candidate hidden state respectively, which are used to map the input feature vector at the current moment to the corresponding gate or state, , , , The second weight matrices for the input gate, forget gate, output gate, and candidate hidden state respectively, which are used to map the hidden state at the previous moment to the corresponding gate or state, , , , The bias vectors for the input gate, forget gate, output gate, and candidate hidden state respectively, is the sigmoid activation function, is the hyperbolic tangent activation function, is element-wise multiplication. The updated hidden state is obtained based on the input feature data and the hidden state at the previous moment. The updated hidden state carries the context information of the previous moments and provides a basis for the calculation of the subsequent output vector.

[0106] Step 10422: Determine the output vector according to the updated hidden state;

[0107] Specifically, by inputting the updated hidden state into , the output vector is obtained, where is the output vector, is the weight matrix of the output space, is the updated hidden state, is the bias vector of the output space. The updated one is mapped to the vocabulary space, providing a basis for the subsequent probability distribution conversion.

[0108] Step 10423: Obtain the probability distribution according to the output vector;

[0109] Specifically, through , where is the probability of the v-th word in the output vocabulary at the t-th moment, is the v-th component of the output vector , and V is the size of the vocabulary.

[0110] According to, the probability distribution is obtained. The output vector is converted into a probability distribution through the softmax function, so that each word corresponds to a probability value, which is convenient for subsequent determination of word transition probabilities.

[0111] Step 10424: Determine the word transition probability according to the probability distribution.

[0112] Specifically, by the word transition probability is obtained, where is the i-th word in the vocabulary, is the probability corresponding to the j-th word in the output vector at the i-th moment in the probability distribution, is that the current word is the word transition probability.

[0113] In an alternative embodiment of the present invention, step 105 includes:

[0114] Step 1051, obtaining an acoustic probability according to the speech feature sequence;

[0115] Specifically, inputting the speech feature sequence into the formula , , an acoustic probability is obtained, where is the acoustic probability of the word , is the i-th word in the vocabulary, X is the speech feature sequence, is the word the number of phonemes contained in, j is the j-th phoneme in the phoneme sequence corresponding to the word , is the set of speech frames corresponding to the phoneme , l is the index of the speech frame, and the value range is the frame number in, is the feature vector representing the l-th frame in the speech feature sequence X, given the feature vector of the l-th frame, the posterior probability of the phoneme is the acoustic probability of the vocabulary W.

[0116] Step 1052, obtaining a language probability according to the word transition probability;

[0117] Specifically, by inputting the word transition probability into the formula a language probability is obtained, where is the language probability of the vocabulary W, W is the vocabulary, i is the index of the word in the vocabulary, and the value range is from 1 to g, is the i-th word in the vocabulary, is all the words in the vocabulary W before the i-th word , is the word transition probability.

[0118] Step 1053, obtaining a speech recognition result according to the acoustic probability and the language probability.

[0119] Specifically, by inputting the acoustic probability and the language probability into the formula , the joint probability of the vocabulary W is obtained. The vocabulary with the maximum joint probability is used as the final speech recognition result.

[0120] In an alternative embodiment of the present invention, step 106 of generating a drone control instruction includes:

[0121] Step 1061, determining keywords according to the speech recognition result;

[0122] Specifically, extract keywords of user intention and parameters from the speech recognition result. For example, in a specific embodiment, the speech recognition result is "take off and rise to 50 meters", then the keywords include: the intention is "take off and adjust altitude", and the parameter is "altitude = 50 meters". Determining the keywords to clarify the user intention is beneficial to the efficiency and accuracy of subsequent generation of drone control instructions.

[0123] Step 1062, generating a drone control instruction according to the keywords.

[0124] Specifically, the control instruction with the highest matching degree with the keywords can be searched from the preset control instruction library and, after format conversion, a drone control instruction recognizable by the drone is generated.

[0125] The intelligent cockpit in the drone intelligent operation and maintenance system is an operation platform integrating multiple functions, providing intelligent and convenient support for the operation and maintenance of drones. The intelligent cockpit can display various state information of the drone in real time, such as flight attitude, position, speed, battery power, sensor data, etc. Through the intuitive graphical interface and dashboard, the operator can clearly understand the operation status of the drone and discover potential problems in a timely manner. For example, when the drone is performing power line inspection, the cockpit can display the image of the transmission line captured by the drone and the flight parameters in real time, helping the operator judge whether there is a fault in the line and whether the drone is in a normal flight state. According to the task requirements and environmental information, the intelligent cockpit can automatically plan the flight path of the drone. Using map data, obstacle information, and task target points, etc., an optimized flight trajectory is generated to ensure that the drone can complete the task safely and efficiently. For example, when performing urban mapping tasks, the intelligent cockpit can plan a reasonable flight route according to the scope and terrain of the mapping area to avoid the drone colliding with buildings or other obstacles.

[0126] The method for processing voice data of the command center of the low-altitude base intelligent operation and maintenance system according to the embodiment of the present invention can provide support for the operator to control the operation status of the drone through voice commands, improving the work efficiency of the drone and the speech recognition effect. The specific process includes:

[0127] Step 111, collect voice data through a microphone array;

[0128] Based on the internal space of the command center of the low-altitude base intelligent operation and maintenance system, etc., determine that the microphone type is a capacitive microphone, the layout of the microphone array is a dual microphone array (one microphone is used to collect human voices, and the other has the function of background noise collection, which is convenient for collecting ambient noise. After subtracting the two signals through a differential amplifier and then amplifying, the noise interference is reduced), the number of microphones is two, the spacing is 80 mm, the dual microphone array forms 3 beams within the range of 0° - 180°, and each beam corresponds to a recording range of 60°. Install the microphones in the command center of the low-altitude base intelligent operation and maintenance system according to the preset positions and spacing to collect voice data in real time.

[0129] Step 112, preprocess;

[0130] Perform filtering processing, segmentation processing, windowing processing, and detection processing on the voice data to obtain preprocessed data, so as to improve the accuracy of the voice data and the subsequent recognition efficiency.

[0131] Step 113, feature extraction;

[0132] By performing a frequency-domain transformation on the preprocessed data, obtain frequency-domain data, and then perform feature extraction based on the frequency-domain data to obtain feature data including feature vectors for subsequent speech recognition.

[0133] Step 114, determine the speech feature sequence and word transition probability;

[0134] According to the obtained preset parameters and relevant formulas, calculate the speech feature sequence and word transition probability respectively, providing an objective data basis for subsequent determination of the speech recognition result and improving the accuracy of the speech recognition result.

[0135] Step 115, determine the speech recognition result;

[0136] Calculate the acoustic probability and language probability respectively through the speech feature sequence and word transition probability, and then calculate the joint probability of the vocabulary list. Take the vocabulary list with the maximum joint probability as the final speech recognition result.

[0137] Step 116, generate a drone control instruction.

[0138] First, determine the keywords according to the speech recognition result, then search for the control instruction with the highest matching degree with the keywords from the preset control instruction library, and after format conversion, generate a drone control instruction that can be recognized by the drone. For example, the operator can let the drone take off automatically and fly to the preset position through the voice command "Take off to the specified location".

[0139] The voice data processing method of the command center of the low-altitude base intelligent operation and maintenance system according to the embodiments of the present invention determines the voice recognition result and generates a drone control instruction through the real-time collected voice data and related calculations, and has the advantage of accurate voice recognition results.

[0140] As Figure 2 shown, an embodiment of the present invention provides a voice data processing device 200 for the command center of a low-altitude base intelligent operation and maintenance system, including:

[0141] An acquisition module 201, configured to acquire voice data through a microphone array placed in the command center of the low-altitude base intelligent operation and maintenance system; the microphone array is arranged according to a preset layout;

[0142] A processing module 202, configured to preprocess the voice data to obtain preprocessed data; extract features from the preprocessed data to obtain feature data; determine a voice feature sequence and a word transition probability according to the feature data; obtain a voice recognition result according to the voice feature sequence and the word transition probability; generate a drone control instruction according to the voice recognition result.

[0143] Optionally, preprocessing the voice data to obtain preprocessed data includes:

[0144] Performing filtering processing on the voice data to obtain filtered data;

[0145] Performing segmentation processing on the filtered data to obtain segmented data;

[0146] Performing windowing processing on the segmented data to obtain windowed data;

[0147] Performing detection processing on the windowed data to obtain preprocessed data.

[0148] Optionally, extracting features from the preprocessed data to obtain feature data includes:

[0149] By Performing frequency-domain transformation on the preprocessed data to obtain frequency-domain data; where is the spectrum of the i-th frame of voice data in the preprocessed data, , is the i-th frame of preprocessed data, n is the sampling time, n = 0, 1,..., N - 1, N is the frame length, i = 1, 2,...; k is the index of the frequency component, and the value range is , j is the imaginary unit, satisfying , e is the base of the natural logarithm, is the frequency-domain data of the i-th frame of voice data;

[0150] Extract features from the frequency-domain data to obtain feature data.

[0151] Optionally, extracting features from the frequency-domain data to obtain feature data includes:

[0152] According to , determine the logarithmic energy;

[0153] According to , determine the first coefficient;

[0154] According to , determine the second coefficient;

[0155] According to , determine the third coefficient;

[0156] Obtain feature data according to the first coefficient, the second coefficient, and the third coefficient;

[0157] Where is the logarithmic energy output by the mth triangular band-pass filter, N is the frame length, k is the index of the frequency component, and the value range is , is the mth triangular band-pass filter, , M is the number of triangular band-pass filters, is the first coefficient at the a-th position, a = 0, 1, ……, L - 1, and L is an integer from 12 to 16, is the second coefficient, is the first coefficient at the (a + 1)-th position, is the first coefficient at the (a - 1)-th position, is the third coefficient, is the second coefficient at the (a + 1)-th position, is the second coefficient at the (a - 1)-th position.

[0158] Optionally, determining a speech feature sequence according to the feature data includes:

[0159] Obtain the state transition probability matrix and the observation probability matrix of the preset network model;

[0160] Determine the forward variable and the backward variable according to the state transition probability matrix, the observation probability matrix, and the feature data;

[0161] Determine the total probability according to the forward variable and the backward variable;

[0162] Determine the speech feature sequence according to the total probability.

[0163] Optionally, determining the word transition probability according to the feature data includes:

[0164] Determine the updated hidden state according to the described feature data;

[0165] Determine the output vector according to the updated hidden state;

[0166] Obtain the probability distribution according to the output vector;

[0167] Determine the word transition probability according to the probability distribution;

[0168] Optionally, obtain the speech recognition result according to the speech feature sequence and the word transition probability, including:

[0169] Obtain the acoustic probability according to the speech feature sequence;

[0170] Obtain the language probability according to the word transition probability;

[0171] Obtain the speech recognition result according to the acoustic probability and the language probability;

[0172] The voice data processing device of the command center of the low-altitude base intelligent operation and maintenance system according to the embodiment of the present invention collects voice data through a microphone array, then preprocesses and extracts features from the voice data to obtain feature data, and then determines the speech feature sequence and word transition probability according to the feature data. According to the speech feature sequence and word transition probability, the speech recognition result is obtained, and finally, according to the speech recognition result, a drone control instruction is generated, which improves the accuracy and efficiency of speech recognition in the command center of the low-altitude base intelligent operation and maintenance system.

[0173] It should be noted that this device corresponds to the above method, and all implementation manners in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effects. Details are not described again in this embodiment.

[0174] The embodiment of the present invention also provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method described in any one of the above embodiments. All implementation manners in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effects. Details are not described again in this embodiment.

[0175] The embodiment of the present invention also provides a computer-readable storage medium, on which instructions are stored. When the instructions are run on a computer, the computer is made to execute the method described in any one of the above embodiments. All implementation manners in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effects. Details are not described again in this embodiment.

[0176] It should be noted that in the device and method of the present invention, it is obvious that each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel, crosswise or independently of each other.

[0177] It should be noted that in the above embodiments, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the implementation manners of the above embodiments is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0178] The above is the preferred implementation manner of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for processing voice data of the command center of a low-altitude base intelligent operation and maintenance system, characterized in that, including: collecting voice data through a microphone array placed in the command center of the intelligent operation and maintenance system for low-altitude bases; the microphone array is set according to a preset layout; preprocessing the voice data to obtain preprocessed data; extracting features from the preprocessed data to obtain feature data; determining a voice feature sequence and a word transition probability according to the feature data; obtaining a voice recognition result according to the voice feature sequence and the word transition probability; generating a drone control instruction according to the voice recognition result.

2. The method for processing voice data of the command center of the intelligent operation and maintenance system for low-altitude bases according to claim 1, characterized in that, Preprocessing the voice data to obtain preprocessed data, including: performing filtering processing on the voice data to obtain filtered data; performing segmentation processing on the filtered data to obtain segmented data; performing windowing processing on the segmented data to obtain windowed data; performing detection processing on the windowed data to obtain preprocessed data.

3. The method for processing voice data of the command center of the intelligent operation and maintenance system for low-altitude bases according to claim 1, wherein, Extracting features from the preprocessed data to obtain feature data, including: By performing a frequency-domain transformation on the preprocessed data to obtain frequency-domain data; wherein is the spectrum of the i-th frame of speech data in the preprocessed data, , is the i-th frame of preprocessed data, n is the sampling time, n = 0, 1, ……, N - 1, N is the frame length, i = 1, 2, ……; k is the index of the frequency component, and the value range is , j is the imaginary unit, satisfying , e is the base of the natural logarithm, is the frequency-domain data of the i-th frame of speech data; extracting features from the frequency-domain data to obtain feature data.

4. The method for processing voice data of the command center of the low-altitude base intelligent operation and maintenance system according to claim 3, wherein Extracting features from the frequency-domain data to obtain feature data, including: According to , determine the logarithmic energy; According to , determine the first coefficient; According to , determine the second coefficient; According to , determine the third coefficient; obtaining feature data according to the first coefficient, the second coefficient, and the third coefficient; wherein, is the logarithmic energy output by the m-th triangular band-pass filter, N is the frame length, k is the index of the frequency component, and the value range is , is the m-th triangular band-pass filter, , M is the number of triangular band-pass filters, is the first coefficient at the a-th position, a = 0, 1, ……, L - 1, and L is an integer from 12 to 16, is the second coefficient, is the first coefficient at the (a + 1)-th position, is the first coefficient at the (a - 1)-th position, is the third coefficient, is the second coefficient at the (a + 1)-th position, is the second coefficient at the (a - 1)-th position.

5. The method for processing voice data of the command center of the intelligent operation and maintenance system for low-altitude bases according to claim 1, wherein Determining a voice feature sequence according to the feature data, including: obtaining a state transition probability matrix and an observation probability matrix of a preset network model; determining a forward variable and a backward variable according to the state transition probability matrix, the observation probability matrix, and the feature data; determining a total probability according to the forward variable and the backward variable; determining a voice feature sequence according to the total probability.

6. The method for processing voice data of the command center of the low-altitude base intelligent operation and maintenance system according to claim 5, wherein, Determining a word transition probability according to the feature data, including: determining an updated hidden state according to the feature data; determining an output vector according to the updated hidden state; obtaining a probability distribution according to the output vector; determining a word transition probability according to the probability distribution.

7. The method for processing voice data of the command center of the intelligent operation and maintenance system for low-altitude bases according to claim 6, wherein, Obtaining a voice recognition result according to the voice feature sequence and the word transition probability, including: obtaining an acoustic probability according to the voice feature sequence; obtaining a language probability according to the word transition probability; obtaining a voice recognition result according to the acoustic probability and the language probability.

8. A voice data processing device for the command center of a low-altitude base intelligent operation and maintenance system, characterized in that, including: a collection module for collecting voice data through a microphone array placed in the command center of the intelligent operation and maintenance system for low-altitude bases; the microphone array is set according to a preset layout; a processing module for preprocessing the voice data to obtain preprocessed data; extracting features from the preprocessed data to obtain feature data; determining a voice feature sequence and a word transition probability according to the feature data; obtaining a voice recognition result according to the voice feature sequence and the word transition probability; generating a drone control instruction according to the voice recognition result.

9. A computing device, characterized in that, including: a processor and a memory storing a computer program, and when the computer program is run by the processor, it executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, storing an instruction, and when the instruction runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 7.