An intelligent speech recognition and instruction execution system based on artificial intelligence during duty

By introducing a frequency mapping function of dialect and tone fusion, optimizing the speech recognition algorithm, combined with artificial intelligence technology, the problem of low recognition rate in the duty environment of traditional speech recognition systems is solved, and efficient, accurate recognition and safe execution are achieved in noisy environments.

CN120048263BActive Publication Date: 2025-07-08GUANGZHOU AEBELL ELECTRICAL TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510530638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-08
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Traditional speech recognition technology is difficult to adapt to dialects and tones in different regions in the duty environment, resulting in low recognition rates, affecting the efficiency and safety of duty work, especially in noisy environments, which are prone to misidentification.

Method used

Using an intelligent speech recognition system based on artificial intelligence, the speech extraction algorithm and acoustic model are optimized by introducing a frequency mapping function of dialect and tone fusion, combining the hidden Markov acoustic model and the N-meta-language model to improve the accuracy and environmental adaptability of speech recognition, and the intention of the on-duty personnel is understood through the intention recognition algorithm. The execution unit verifies it before performing the operation.

Benefits of technology

It significantly improves the recognition rate and adaptability of the system in complex duty environments, ensures the accurate execution of instructions, improves the efficiency and safety of duty work, especially in emergencies, and reduces the possibility of misoperation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048263B_ABST
    Figure CN120048263B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of speech recognition and command execution. Specifically, it relates to an intelligent speech recognition and command execution system based on artificial intelligence during duty. It includes: a voice receiving and processing unit that captures voice signals in the environment and performs preliminary processing; a speech recognition unit that converts the preliminarily processed voice signals into text information; a language understanding unit that analyzes the text information to understand the intentions of the duty personnel; a decision-making unit that decides how to respond to the commands of the duty personnel according to the results of language understanding; and an execution unit that executes the operations determined by the decision-making unit. The design of the present invention effectively solves the problem of low recognition rate of traditional speech recognition systems when facing duty personnel with different regions and accents by introducing a frequency mapping function that combines dialects and timbres. Especially in complex and changeable duty environments, this optimization can significantly improve the system's adaptability to regional dialects and special timbres.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition and instruction execution, and more specifically, to an intelligent speech recognition and instruction execution system based on artificial intelligence during duty Background Art

[0002] In traditional speech recognition technology, the system often has difficulty accurately recognizing speech commands with regional accents or specific voices, especially in a noisy duty environment where background noise can seriously interfere with speech recognition, resulting in a low recognition rate. This limits the application effect of speech recognition technology in actual duty work; Duty personnel in different regions may use different dialects, or due to different personal pronunciation habits, the recognition of speech commands becomes complicated. The lack of an effective dialect and voice adaptation mechanism makes it difficult for existing systems to meet diverse needs; Traditional speech recognition technology performs poorly in complex environments, affecting the efficiency and quality of duty work; In some key duty tasks, such as traffic command and emergency response, incorrect instruction execution may lead to serious consequences. Therefore, an intelligent speech recognition and instruction execution system based on artificial intelligence during duty is designed. Summary of the Invention

[0003] The purpose of the present invention is to provide an intelligent speech recognition and instruction execution system based on artificial intelligence during duty to solve the problems of insufficient accuracy of speech recognition, diversity of dialects and voices, and poor environmental adaptability mentioned in the above background art.

[0004] To achieve the above object, the present invention provides an intelligent speech recognition and instruction execution system based on artificial intelligence during duty, including:

[0005] A speech reception and processing unit that captures speech signals in the environment and performs preliminary processing;

[0006] A speech recognition unit that converts the preliminarily processed speech signals into text information. During the conversion process, a frequency mapping function for integrating dialects and voices is introduced for the dialects of duty personnel in different regions and the voices of duty personnel;

[0007] A language understanding unit that analyzes the text information and understands the intentions of duty personnel;

[0008] A decision-making unit that decides how to respond to the instructions of duty personnel based on the results of language understanding;

[0009] An execution unit that executes the operations determined by the decision-making unit.

[0010] As a further improvement of this technical solution, the voice receiving and processing unit includes a voice receiving module and a voice processing module;

[0011] Among them, the voice receiving module captures sound through an array composed of multiple microphones;

[0012] The voice processing module analyzes the sounds received by different microphones and improves the spectral characteristics of the voice signal by increasing the amplitude of the high-frequency components.

[0013] As a further improvement of this technical solution, the voice recognition unit includes a feature extraction module, an acoustic module, a language module, a word order generation module, and a text generation module;

[0014] Among them, the feature extraction module extracts relevant features from the voice signal based on a voice extraction algorithm. By introducing a frequency mapping function for the fusion of dialect and timbre, the voice extraction algorithm is optimized, and the voice extraction algorithm is further optimized so that after introducing the frequency mapping function for the fusion of dialect and timbre, the voice extraction algorithm is better compatible with the hidden Markov acoustic model;

[0015] The acoustic module maps the feature vectors of the voice signal to phoneme units based on the hidden Markov acoustic model;

[0016] The language module predicts the probability of the next word based on the N-gram language model;

[0017] The word order generation module combines the hidden Markov acoustic model and the N-gram language model, uses a word order search algorithm to find the most likely word sequence, and generates a preliminary recognition result;

[0018] The text generation module evaluates the confidence of the preliminary recognition result, filters out the recognition results with low confidence, performs spelling check and correction on the recognition results, adjusts the recognition results according to the context information, and converts the final recognition result into text format.

[0019] As a further improvement of this technical solution, the voice extraction algorithm is:

[0020] ;

[0021] In view of the different dialects of on-duty personnel in different regions and the different timbres of on-duty personnel, a frequency mapping function for the fusion of dialect and timbre is introduced into the voice extraction algorithm:

[0022] ;

[0023] The voice extraction algorithm is further optimized so that the voice extraction algorithm after introducing the frequency mapping function for the fusion of dialect and timbre is better compatible with the acoustic model:

[0024] ;

[0025] Among them, represents the extracted speech features; represents the speech features extracted after introducing the frequency mapping function for dialect and timbre fusion; represents the speech features extracted after further optimization; represents the number of filter banks; represents the output of the filter; represents the index of the speech extraction algorithm; represents the filter index; represents according to the dialect features the calculated adjustment factor; represents according to the timbre features the calculated adjustment factor; represents the weight factor.

[0026] As a further improvement of this technical solution, the word order generation module uses a word order search algorithm to find the most likely word sequence, including the following steps:

[0027] S2.1. For the first feature vector of the speech signal, calculate the probability of each possible initial state;

[0028] S2.2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state;

[0029] S2.3. For each time point and each possible state, calculate the probability of transitioning from all possible previous states to the current state, and in combination with the duty scenario, introduce traffic terms in the duty scenario to optimize the probability calculation and further optimize the probability calculation so that the probability after introducing traffic terms is applicable to different application scenarios;

[0030] S2.4. Among all possible previous states, update the probability of the current state at the time point to this maximum probability, and at the same time update the backtracking pointer of the current state;

[0031] S2.5. At the last time point, select the state with the highest probability as the most likely end state;

[0032] S2.6. Starting from the finally determined end state, backtrack along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence;

[0033] S2.7. For the obtained most likely hidden state sequence, calculate the probability of the hidden state sequence in the language model, and in combination with the probabilities of the hidden Markov acoustic model and the N-gram language model, calculate the final comprehensive probability;

[0034] S2.8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result.

[0035] As a further improvement of this technical solution, in S2.3, the probability of calculating the transition from all possible previous states to the current state is:

[0036] ;

[0037] Combined with the duty scenario, introduce traffic terms in the duty scenario to optimize the probability calculation as:

[0038] ;

[0039] Further optimize the probability calculation so that the probability after introducing traffic terms is applicable to different application scenarios:

[0040] ;

[0041] Among them, represents the time point; represents the current state; represents the previous state; represents the total number of states; represents the probability that the most likely path reaches state at time point ; represents the probability after introducing traffic terms in the duty scenario; represents the probability applicable to different application scenarios; represents the probability that the most likely path reaches state at time point ; represents the probability of transitioning from state to state ; represents the probability of generating the observation value in state ; represents the probability of considering the traffic term in state in the duty scenario; represents the scenario-related weight.

[0042] As a further improvement of this technical solution, the language understanding unit includes a text recognition module, a sentence analysis module, and a dialogue management module;

[0043] Among them, the text recognition module preprocesses the generated text and identifies named entities in the text through regular expressions to identify specific types of information;

[0044] The sentence analysis module performs syntactic structure analysis on the sentence through a greedy decoding algorithm, understands the relationships between various words, determines the roles played by each part in the sentence, and based on the analysis results, identifies the specific intentions of the duty officers through an intention recognition algorithm;

[0045] The dialogue management module tracks the dialogue history and context information, provides more coherent responses, and introduces an attention mechanism into the model.

[0046] As a further improvement of this technical solution, the intention recognition algorithm is:

[0047] ;

[0048] In view of the urgency of the intentions of the duty officers when expressing dissatisfaction and in case of emergencies, the intention recognition algorithm is optimized:

[0049] ;

[0050] Among them, represents the feature vector extracted from the speech recognition result; represents the intention of the duty officer; represents the weight vector; represents the bias term; represents the feature vector of emergencies and expressing dissatisfaction; represents the matching degree between the current input feature and the features of emergencies and dissatisfaction; is used to adjust the influence degree on the final prediction result represents the base of the natural logarithm; represents the transpose of the vector.

[0051] As a further improvement of this technical solution, the decision-making unit decides how to respond to the instructions of the duty officers according to the results of language understanding, including the following steps:

[0052] S4.1. Match the intention of the user with a predefined rule set;

[0053] S4.2. Once the intention is correctly parsed and successfully matches the rule, generate a specific action plan and issue relevant instructions to the execution unit;

[0054] S4.3. Before performing any operation, the decision-making unit performs a confirmation step.

[0055] As a further improvement of this technical solution, the execution unit executes the operations determined by the decision-making unit, including the following steps:

[0056] S5.1. Verify again whether the instructions from the decision-making unit are clear and reasonable, and check whether the on-duty personnel who issue the instructions have the authority to perform the operation;

[0057] S5.2. Format the control signal according to the communication protocol of the target device;

[0058] S5.3. Establish a communication connection with the target device, and send the control signal to the target device through the established communication channel;

[0059] S5.4. Continuously monitor the status of the device to determine whether the device starts as expected.

[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0061] 1. In the intelligent voice recognition and instruction execution system during on-duty based on artificial intelligence, by introducing the frequency mapping function of dialect and timbre fusion, the problem of low recognition rate of traditional voice recognition systems when facing on-duty personnel with different regions and accents is effectively solved. Especially in complex and changeable on-duty environments, this optimization can significantly improve the adaptability of the system to regional dialects and special timbres, ensuring that even in noisy environments, the instructions of on-duty personnel can be accurately recognized and understood, thus improving the practicability and reliability of the system.

[0062] 2. In the intelligent voice recognition and instruction execution system during on-duty based on artificial intelligence, it can not only quickly and accurately understand the intentions of on-duty personnel, but also automatically make response decisions according to the pre-set rule set, and even add a confirmation step before performing key operations to ensure the safety of operations. Especially in emergency situations, the system can quickly identify and respond to the urgent needs of on-duty personnel, take appropriate measures in a timely manner, effectively avoid risks caused by human negligence or misjudgment, and greatly improve the efficiency and safety of on-duty work. In addition, the execution unit will verify the effectiveness and reasonableness of the instructions again before performing the operation, which further guarantees the accuracy of system operations and reduces the possibility of misoperations. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is the overall flow block diagram of the present invention;

[0064] The meanings of the various labels in the figure are as follows:

[0065] 1. Voice reception and processing unit; 2. Voice recognition unit; 3. Language understanding unit; 4. Decision-making unit; 5. Execution unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] Example: Please refer to Figure 1 As shown, a smart voice recognition and instruction execution system during duty based on artificial intelligence is provided, including:

[0068] The voice receiving and processing unit 1 captures the voice signals in the environment and performs preliminary processing;

[0069] In this embodiment, the voice receiving and processing unit 1 includes a voice receiving module and a voice processing module;

[0070] Among them, the voice receiving module captures sounds through an array composed of multiple microphones. The microphone array can provide spatial information, which helps subsequent sound source localization and noise reduction processing;

[0071] The voice processing module analyzes the sounds received by different microphones and improves the spectral characteristics of the voice signals by increasing the amplitude of the high-frequency components, making subsequent processing easier.

[0072] The voice recognition unit 2 converts the preliminarily processed voice signals into text information;

[0073] In this embodiment, the voice recognition unit 2 includes a feature extraction module, an acoustic module, a language module, a word order generation module, and a text generation module;

[0074] Among them, the feature extraction module extracts relevant features from the voice signals based on a voice extraction algorithm. By introducing a frequency mapping function for dialect and timbre fusion, the voice extraction algorithm is optimized, and the voice extraction algorithm is further optimized so that after introducing the frequency mapping function for dialect and timbre fusion, the voice extraction algorithm is better compatible with the hidden Markov acoustic model;

[0075] Furthermore, the original signal is converted into a low-dimensional feature vector capable of characterizing the voice characteristics through a voice extraction algorithm. The voice extraction algorithm is:

[0076] ;

[0077] Law enforcement officers in different regions may use different dialects, which have significant differences in phonetic features. By introducing a frequency mapping function that fuses dialect and timbre, these differences can be better captured and processed, improving the system's ability to recognize different dialects; the timbres of law enforcement officers (such as pitch, voice quality, etc.) vary, and these differences affect the spectral characteristics of speech signals. By introducing a frequency mapping function for timbre fusion, the speech extraction algorithm can be optimized to better adapt to the characteristics of different timbres, thereby improving the accuracy of recognition; by introducing a frequency mapping function that fuses dialect and timbre, the system can maintain high recognition performance under different environments and conditions, unaffected by dialect and timbre differences, which enables the system to operate stably in various law enforcement scenarios, enhancing the reliability and practicality of the system; in view of the different dialects of law enforcement officers in different regions and the different timbres of law enforcement officers, a frequency mapping function that fuses dialect and timbre is introduced into the speech extraction algorithm :

[0078] If there is an original speech signal , the dialect feature and the timbre feature are respectively represented as a set of parameters and , where represents the specific phoneme frequency distribution of a certain dialect, represents the average value or variance of the speaker's pitch;

[0079] ;

[0080] Among them, the frequency mapping function that fuses dialect and timbre is used to describe how to calculate a new feature vector based on the original speech signal , the dialect feature and the timbre feature . By fusing dialect and timbre information, the feature extraction process of the speech signal is improved, making the extracted features better reflect the uniqueness of the speaker, thereby improving the performance of the speech recognition system;

[0081] The dialect feature adjustment factor is (in the form of a product, each dialect feature will affect the final frequency mapping result, and the degree of influence is controlled by ):

[0082] ;

[0083] In the formula, represents the number of dimensions of the dialect feature , that is, the dialect feature vector The number of different dialect features included; Indicates a dialect feature The index in, used to traverse all dialect features;

[0084] Timbre feature adjustment factor Is (in the form of exponential decay, for each timbre feature The greater the deviation, the smaller its impact on the final result, thus achieving effective adaptation to different timbres):

[0085] ;

[0086] A speech recognition system usually consists of multiple components, such as an acoustic model, a language model, etc. The speech extraction algorithm after introducing the frequency mapping function for dialect and timbre fusion may be incompatible with the existing acoustic model. The existing acoustic model is trained based on general speech features. The new speech extraction algorithm may cause a change in the data distribution input to the acoustic model, resulting in a decline in the performance of the acoustic model. For example, if the speech extraction algorithm changes the representation of speech features and the acoustic model is not adjusted for this new representation, then the acoustic model may make mistakes when classifying and recognizing these speech features. Further optimize the speech extraction algorithm to make the speech extraction algorithm after introducing the frequency mapping function for dialect and timbre fusion better compatible with the acoustic model:

[0087] ;

[0088] Among them, Indicates the extracted speech features, which are one of the features extracted from the speech signal and used to represent the spectral characteristics of the speech signal; Indicates the speech features extracted after introducing the frequency mapping function for dialect and timbre fusion; Indicates the speech features extracted after further optimization; Indicates the number of filter banks. On the Mel frequency scale, a set of filters is used to capture the energy in different frequency ranges; Indicates the output of the filter, which is the log energy. Each filter captures the energy in a specific frequency range and takes the logarithm to compress the dynamic range; Indicates the index of the speech extraction algorithm; Indicates the filter index; Indicates the adjustment factor calculated according to the dialect feature ; Indicates the adjustment factor calculated according to the timbre feature ; Indicates the weight parameter that adjusts the influence degree of each dialect feature on the final result; Indicates adjusting each timbre feature A weight parameter for the degree of influence on the final result; Represents a reference value for a certain timbre feature, such as the average pitch of a speaker; Represents a weight factor used to balance the contributions of the standard Mel frequency mapping and the dialect-specific frequency mapping, The value range of which is from 0 to 1, determining the relative importance of the two mapping methods;

[0089] The acoustic module maps the feature vectors of the speech signal to phoneme units based on the hidden Markov acoustic model. The hidden Markov model (HMM) is a statistical model used to model time series data, especially those with hidden states. In speech recognition, these hidden states are usually phonemes or sub-phoneme units, and the observations are the feature vectors extracted from the speech signal;

[0090] The language module predicts the probability of the next word based on the N-gram language model. The N-gram model is a statistical model based on historical word sequences, and it predicts the probability of the next word by statistically counting the frequencies of historical word sequences;

[0091] The word order generation module combines the hidden Markov acoustic model and the N-gram language model, and uses a word order search algorithm (solving the most likely hidden state sequence through dynamic programming. At any moment, there is only one most likely path to reach the current state. Therefore, the algorithm will trace these maximum probability paths and finally select the path with the highest total probability as the most likely state sequence) to find the most likely word sequence. This step aims to select the best path from all possible candidate word sequences to generate a preliminary recognition result;

[0092] Furthermore, find the most likely word sequence such that the probability of generating this word sequence is the largest. This ensures that the system can accurately convert the speech signal into text information, reducing misrecognition and missed recognition. During the duty process, accurate speech recognition is crucial for executing correct instructions, especially in emergency situations, where incorrect recognition may lead to serious consequences; through the word order search algorithm, the optimal solution can be found in polynomial time, avoiding the high computational cost of exhaustively searching all possible word sequences. In practical applications, the system needs to process a large number of speech signals in a short time, and an efficient algorithm can ensure real-time performance and response speed. Combining the joint probability of the HMM and the N-gram model comprehensively considers speech features and language models, improving the accuracy of recognition; using the word order search algorithm to find the most likely word sequence includes the following steps:

[0093] S2.1. For the first feature vector of the speech signal, calculate the probability of each possible initial state ( , is the probability of the initial state, is the probability of observing in the initial state);

[0094] S2.2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state. At the initial moment, since there is no previous state, these pointers can be set to meaningless values;

[0095] S2.3. For each time point (starting from 2 until the last observation), and for each possible state, calculate the probability of transitioning from all possible previous states to the current state, and combine the duty scenarios (such as traffic command, patrol inspection, etc.), introduce traffic terms in the duty scenarios to optimize the probability calculation, improve the accuracy and rationality of recognition, and further optimize the probability calculation so that the probability after introducing traffic terms is applicable to different application scenarios;

[0096] Among them, the calculation of the probability of transitioning from all possible previous states to the state is:

[0097] ;

[0098] By introducing traffic terms in the duty scenarios, the system can give priority to recognizing these words, reducing the possibility of misrecognition. For example, in the traffic command scenario, the system can give priority to recognizing command words such as "stop", "go ahead", "turn left", etc., rather than other irrelevant words. In the duty scenario, some words may have multiple meanings. By introducing traffic terms, the system can better understand the context, reduce ambiguity, and improve the accuracy of recognition. For example, "turn left" usually refers to a vehicle turning in traffic command, while it may have different meanings in other scenarios; by introducing traffic terms, the system can ensure the semantic consistency of the recognition results. For example, in traffic command, the system can ensure that commands such as "stop at red light, go at green light" are reasonable, rather than unreasonable commands such as "go at red light"; combining the duty scenarios (such as traffic command, patrol inspection, etc.), the optimization of the probability calculation by introducing traffic terms in the duty scenarios is:

[0099] ;

[0100] ;

[0101] Among them, is defined as a normalized weight, specifically: for each state calculate its matching degree score with the traffic terms , calculate the normalized weight of each state ;

[0102] The optimized probability calculation method can be applied to various duty scenarios, such as traffic command, patrol inspection, emergency response, etc., improving the universality and practicality of the system. The system can dynamically adjust the probability calculation parameters according to different scenarios to ensure high recognition performance in various duty tasks; further optimize the probability calculation to make the probability after introducing traffic terms applicable to different application scenarios:

[0103] ;

[0104] ;

[0105] Among them, represents the time point, that is, the th observation value of the current processed voice signal; represents the current state, that is, one of the possible states at the time point ; represents the previous state, that is, one of the possible states at the time point ; represents the total number of states, that is, the number of all possible states; represents the probability that the most likely path reaches state at the time point ; represents the probability after introducing traffic terms in the duty scenario; represents the probability applicable to different application scenarios; represents the probability that the most likely path reaches state at the time point ; represents the probability of transitioning from state to state ; represents the probability of generating the observation value in state , where is the observation value at the time point ; represents the probability of considering the traffic term in state ; represents the scenario-related weight used to adjust the value of . When , it increases the importance of traffic terms. When , it reduces the importance of traffic terms; represents the matching degree score between state and the traffic term in the duty scenario; represents the traffic term in the duty scenario; An index indicating the weight of traffic terms in the calculation of the duty scenario, that is, one of all possible states; A parameter indicating the control of the scoring sensitivity;

[0106] S2.4. Among all possible previous states, update the probability of the current state at the time point to this maximum probability, and at the same time update the backtracking pointer of the current state;

[0107] S2.5. At the last time point, select the state with the highest probability as the most likely end state, and the probability of this state is the probability of the entire most likely hidden state sequence;

[0108] S2.6. Starting from the finally determined end state, backtrack along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence;

[0109] S2.7. For the obtained most likely hidden state sequence, calculate the probability of the hidden state sequence in the language model ( , is the conditional probability in the language model), and combine the probabilities of the hidden Markov acoustic model and the N-gram language model to calculate the final comprehensive probability ( );

[0110] S2.8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result;

[0111] The text generation module evaluates the confidence of the preliminary recognition result, filters out the recognition results with low confidence, performs spelling checks and corrections on the recognition results to improve the readability of the text, adjusts the recognition results according to the context information to ensure the consistency and rationality of the semantics, and converts the final recognition result into a text format.

[0112] The language understanding unit 3 analyzes the text information to understand the intention of the duty personnel;

[0113] In this embodiment, the language understanding unit 3 includes a text recognition module, a sentence analysis module, and a dialogue management module;

[0114] Among them, the text recognition module preprocesses the generated text, including removing stop words (deleting common words such as "de", "shi", "zai" in the text, which usually do not help much in understanding the meaning of sentences), word segmentation (dividing continuous text into individual words or phrases for further processing), stemming or lemmatization (reducing words to their basic forms to reduce the impact of lexical variants), and identifying named entities in the text through regular expressions (regular expressions are tools for describing patterns to match strings. By writing regular expressions, specific patterns can be defined to identify and extract named entities in the text), such as personal names, place names, organization names, etc., and identifying specific types of information such as time, date, quantity, etc.;

[0115] The sentence analysis module analyzes the syntactic structure of the sentence through the greedy decoding algorithm (the greedy decoding algorithm is a strategy for constructing a dependency tree word by word. In each step, the currently optimal dependency relationship is selected to gradually construct the dependency tree of the entire sentence), understands the relationships between various words, such as the subject-predicate-object structure, determines the roles played by each part in the sentence, such as the doer of the action (agent), the object of the action (patient), etc., and based on the analysis results, identifies the specific intention of the duty personnel through the intention recognition algorithm, which may be querying information, requesting services, expressing opinions, etc.;

[0116] Furthermore, the intention recognition algorithm vectorizes the text features and then uses the algorithm to learn the mapping relationship from the input features to the intention categories, thereby realizing the classification and recognition of text intentions; through the intention recognition algorithm, it is ensured that the system can accurately understand the intentions of the duty personnel. Whether the instruction is given in spoken form, through the intention recognition algorithm, the system can parse out the core content and purpose of the instruction; by quickly and accurately identifying the intentions of the duty personnel, the system can immediately generate corresponding responses or execute operations, reduce the response time, provide a more natural and smooth interaction experience, enable the duty personnel to interact with the system more easily, and reduce the operation complexity; the duty personnel can communicate with the system through natural language without having to remember complex command grammars, improving work efficiency; the intention recognition algorithm is as follows:

[0117] ;

[0118] In case of an emergency, the system can quickly recognize the emergency intentions of the on-duty personnel and immediately take corresponding measures to reduce the response time and ensure safety. In case of emergencies such as traffic accidents and fires, the system can quickly recognize instructions such as "emergency evacuation" and "call for rescue" and immediately activate the emergency response plan; when the on-duty personnel express dissatisfaction, the system can accurately recognize their negative emotions and specific demands, and take timely measures to ease the conflict and improve service quality. If the on-duty personnel are dissatisfied with a certain decision or operation, the system can recognize it in time and provide corresponding solutions or explanations to reduce misunderstandings and conflicts; by optimizing the intention recognition algorithm, the system can maintain a high recognition accuracy and reliability in case of emergencies and dissatisfaction, ensuring the smooth completion of tasks. The optimized algorithm can recognize the emotional characteristics of the on-duty personnel, such as anger and anxiety, and improve the recognition accuracy of emergencies and dissatisfaction; aiming at the urgency of the intentions when the on-duty personnel express dissatisfaction and in case of emergencies, the intention recognition algorithm is optimized:

[0119] ;

[0120] ;

[0121] Among them, represents the feature vector extracted from the speech recognition result; represents the intention of the on-duty personnel; represents the weight vector, and each element corresponds to the importance of one feature in the input feature vector ; represents the bias term, regarded as the initial value of the model, which helps to improve the flexibility of the model; represents the feature vector of emergencies and expressing dissatisfaction; represents the matching degree between the current input feature and the features of emergencies and dissatisfaction; is used to adjust the influence degree on the final prediction result. By adjusting value, a suitable balance point can be found between normal situations and emergencies; represents the base of the natural logarithm; represents the transpose of the vector; represents the weight of the pitch change; represents the change value of the pitch; represents the weight of the accelerated speech rate; represents the change value of the speech rate;

[0122] The dialogue management module tracks the dialogue history and context information, provides more coherent responses, and introduces an attention mechanism into the model to enable the model to better focus on the key information in the dialogue;

[0123] Among them, the dialogue management module is specifically as follows: maintain a dialogue history to store the text content, timestamp, and relevant metadata of each dialogue, set a context window to only retain the content of the most recent few rounds of dialogue to reduce the computational burden and maintain the relevance of the dialogue; extract key information such as time, place, person, event, etc. from the dialogue history to form context features, calculate the correlation between the current dialogue and the historical dialogue to ensure the relevance and coherence of the response; use a feedforward neural network to calculate the similarity score between the current dialogue features and the historical dialogue features; weight and sum the historical dialogue features according to the attention weights to obtain a context vector; combine the current dialogue features and the context vector as inputs to generate the final response.

[0124] The execution unit 5 executes the operations determined by the decision-making unit 4.

[0125] In this embodiment, the execution unit 5 executes the operations determined by the decision-making unit 4, including the following steps:

[0126] S5.1. Verify again whether the instruction from the decision-making unit 4 is clear and reasonable. For example, confirm which device is to be turned on and whether this operation complies with the current security protocol, and check whether the duty officer who issued the instruction has the authority to execute this operation to prevent unauthorized access.

[0127] S5.2. Format the control signal according to the communication protocol of the target device. Different devices may require different formats of signals. For example, some devices may accept digital signals, while others may require analog signals or specific network protocols, and ensure that the sent control signal will not cause damage to the device or system. For example, check whether the signal exceeds the tolerance range of the device.

[0128] S5.3. Establish a communication connection with the target device, send the control signal to the target device through the established communication channel, and wait for the device to return a confirmation message to ensure that the device has received the control signal. If the confirmation message is not received, try to resend or report a fault.

[0129] S5.4. Continuously monitor the status of the device to determine whether the device starts as expected, including reading the status register, sensor data, or other feedback information of the device. If the device fails to start properly, take other measures such as retrying to start, restarting the device, reporting a fault, etc.

[0130] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed.

Claims

1. An intelligent voice recognition and instruction execution system during duty based on artificial intelligence, characterized in that, Including: A voice receiving and processing unit (1), which captures voice signals in the environment and performs preliminary processing; A voice recognition unit (2), which converts the preliminarily processed voice signals into text information. During the conversion process, in view of the different dialects of duty officers in different regions and the different voices of duty officers, a frequency mapping function for integrating dialect and voice is introduced; The voice recognition unit (2) includes a feature extraction module, an acoustic module, a language module, a word order generation module, and a text generation module; Among them, the feature extraction module extracts relevant features from the voice signals based on a voice extraction algorithm. By introducing a frequency mapping function for integrating dialect and voice, the voice extraction algorithm is optimized, and the voice extraction algorithm is further optimized, so that after introducing the frequency mapping function for integrating dialect and voice, the voice extraction algorithm is better compatible with the hidden Markov acoustic model; The voice extraction algorithm is: ; In view of the different dialects of duty officers in different regions and the different voices of duty officers, a frequency mapping function for integrating dialect and voice is introduced into the voice extraction algorithm: ; The voice extraction algorithm is further optimized so that the voice extraction algorithm after introducing the frequency mapping function for integrating dialect and voice is better compatible with the acoustic model: ; Among them, represents the extracted voice features; represents the voice features extracted after introducing the frequency mapping function for dialect and timbre fusion; represents the voice features extracted after further optimization; represents the number of filter banks; represents the output of the filter; represents the index of the voice extraction algorithm; represents the filter index; represents according to the dialect features the calculated adjustment factor; represents the adjustment factor calculated according to the timbre features; represents the weight factor; A language understanding unit (3), which analyzes the text information and understands the intention of the duty officer; A decision-making unit (4), which decides how to respond to the instructions of the duty officer according to the result of language understanding; An execution unit (5), which executes the operations determined by the decision-making unit (4).

2. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 1, wherein: The voice receiving and processing unit (1) includes a voice receiving module and a voice processing module; Among them, the voice receiving module captures sounds through an array composed of multiple microphones; The voice processing module analyzes the sounds received by different microphones and improves the spectral characteristics of the voice signals by increasing the amplitude of the high-frequency components.

3. The artificial intelligence-based intelligent voice recognition and instruction execution system during duty according to claim 2, characterized in that: The acoustic module maps the feature vectors of the voice signals to phoneme units based on the hidden Markov acoustic model; The language module predicts the probability of the next word based on the N-gram language model; The word order generation module combines the hidden Markov acoustic model and the N-gram language model, and uses a word order search algorithm to find the most likely word sequence to generate a preliminary recognition result; The text generation module evaluates the confidence of the preliminary recognition result, filters out the recognition results with low confidence, performs spelling check and correction on the recognition results, adjusts the recognition results according to the context information, and converts the final recognition result into a text format.

4. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 3, characterized in that: The word order generation module uses a word order search algorithm to find the most likely word sequence, including the following steps: S2.

1. For the first feature vector of the voice signal, calculate the probability of each possible initial state; S2.

2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state; S2.

3. For each time point and each possible state, calculate the probability of transitioning from all possible previous states to the current state. Combining with the duty scenario, introduce traffic terms in the duty scenario to optimize the probability calculation, and further optimize the probability calculation so that the probability after introducing traffic terms is applicable to different application scenarios; S2.

4. Among all possible previous states, update the probability of the current state at the time point to this maximum probability, and at the same time update the backtracking pointer of the current state; S2.

5. At the last time point, select the state with the highest probability as the most likely end state; S2.

6. Starting from the finally determined end state, backtrack along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence; S2.

7. For the obtained most likely hidden state sequence, calculate the probability of the hidden state sequence in the language model, and combine the probabilities of the hidden Markov acoustic model and the N-gram language model to calculate the final comprehensive probability; S2.

8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result.

5. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 4, characterized in that: In the above S2.3, the calculation of the probability of transitioning from all possible previous states to the state is: ; Combining with the duty scenario, introducing traffic terms in the duty scenario to optimize the probability calculation is: ; Further optimizing the probability calculation so that the probability after introducing traffic terms is applicable to different application scenarios: ; Among them, represents a time point; represents the current state; represents the previous state; represents the total number of states; represents at the time point when the most likely path reaches the state probability; represents the probability after introducing traffic terms in the duty scenario; represents the probability applicable to different application scenarios; represents at the time point when the most likely path reaches the state probability; represents the probability of transitioning from the state to the state probability; represents the probability of generating the observation value under the state probability; represents the probability of considering the traffic terms in the duty scenario under the state probability; represents the scenario-related weight.

6. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 5, characterized in that: The language understanding unit (3) includes a text recognition module, a sentence analysis module, and a dialogue management module; Among them, the text recognition module preprocesses the generated text, and identifies named entities in the text through regular expressions to identify specific types of information; The sentence analysis module performs syntactic structure analysis on the sentence through a greedy decoding algorithm, understands the relationships between various words, determines the roles played by each part in the sentence, and based on the analysis results, identifies the specific intention of the duty officer through an intention recognition algorithm; The dialogue management module tracks the dialogue history and context information, provides a more coherent response, and introduces an attention mechanism in the model.

7. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 6, characterized in that The intention recognition algorithm is: ; In view of the urgency of the intentions when the duty officer expresses dissatisfaction and emergency situations, optimize the intention recognition algorithm: ; Among them, represents the feature vector extracted from the speech recognition result; represents the intention of the on-duty personnel; represents the weight vector; represents the bias term; represents the feature vector of the emergency situation and expressing dissatisfaction; represents the matching degree between the current input feature and the emergency situation and dissatisfaction feature; used to adjust the influence degree on the final prediction result; represents the base of the natural logarithm; represents the transpose of the vector.

8. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 7, characterized in that: The decision-making unit (4) decides how to respond to the instructions of the duty officer according to the results of language understanding, including the following steps: S4.

1. Match the intention of the user with a predefined rule set; S4.

2. Once the intention is correctly parsed and successfully matches the rule, generate a specific action plan and issue relevant instructions to the execution unit (5); S4.

3. Before performing any operation, the decision-making unit (4) performs a confirmation step.

9. The intelligent voice recognition and instruction execution system during duty based on artificial intelligence according to claim 8, characterized in that: The execution unit (5) executes the operations determined by the decision-making unit (4), including the following steps: S5.

1. Verify again whether the instructions from the decision-making unit (4) are clear and reasonable, and check whether the duty officer who issues the instructions has the authority to perform this operation; S5.

2. Format the control signal according to the communication protocol of the target device; S5.

3. Establish a communication connection with the target device and send a control signal to the target device through the established communication channel; S5.

4. Continuously monitor the status of the device to determine whether the device starts as expected.

Citation Information

Patent Citations

  • Machine tool instruction execution system and method based on voice interaction

    CN118197308A

  • Agricultural disease and pest dialect voice intelligent recognition method based on deep learning technology

    CN119296514A