Intelligent voice recognition and instruction execution system in duty process based on artificial intelligence
By introducing frequency mapping functions and word order search algorithms for dialect and tone fusion in the speech recognition system, the problem of low recognition rate in complex environments of traditional speech recognition technology is solved, and the adaptability and duty efficiency of different dialects and tone is improved.
Patent Information
- Application Number
- CN202510530638.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Traditional speech recognition technology has a low recognition rate when facing on-duty personnel with different regions and accents, and performs poorly in complex environments, which affects the efficiency and quality of on-duty work.
Design an intelligent speech recognition and instruction execution system during duty process based on artificial intelligence. By introducing a frequency mapping function of dialect and tone fusion, the speech extraction algorithm is optimized to make it better compatible with the Hidden Markov acoustic model, and combined with the Hidden Markov acoustic model and the N-meta-language model, the word order search algorithm is used to generate the final recognition results.
It significantly improves the system's adaptability to regional dialects and special tones, ensures that the instructions of on-duty personnel can be accurately identified and understood in a high noise environment, improves the practicality and reliability of the system, and improves the efficiency and safety of on-duty work by quickly and accurately understanding intentions and making automatic response decisions.
Smart Images

Figure CN120048263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition and command execution, and in particular to an intelligent speech recognition and command execution system during on-duty based on artificial intelligence. Background Art
[0002] In traditional speech recognition technology, the system often has difficulty in accurately recognizing voice commands with regional accents or specific timbres, especially in noisy duty environments, where background noise can cause serious interference to speech recognition, resulting in low recognition rates, which limits the application of speech recognition technology in actual duty work; duty personnel in different regions may use different dialects, or due to different personal pronunciation habits, the recognition of voice commands becomes complicated, and the lack of an effective dialect and timbre adaptation mechanism makes it difficult for existing systems to meet diverse needs; traditional speech recognition technology performs poorly in complex environments, affecting the efficiency and quality of duty work; in some key duty tasks, such as traffic control and emergency response, erroneous command execution may lead to serious consequences. Therefore, an intelligent speech recognition and command execution system based on artificial intelligence is designed during duty. Summary of the invention
[0003] The purpose of the present invention is to provide an intelligent voice recognition and command execution system based on artificial intelligence during duty, so as to solve the problems of insufficient accuracy of voice recognition, diversity of dialects and timbres, and poor environmental adaptability proposed in the above-mentioned background technology.
[0004] To achieve the above object, the present invention aims to provide an intelligent voice recognition and command execution system based on artificial intelligence during duty, comprising: A voice receiving and processing unit, which captures voice signals in the environment and performs preliminary processing; A speech recognition unit, which converts the preliminarily processed speech signal into text information. During the conversion process, a frequency mapping function that integrates the dialect and the timbre is introduced according to the different dialects and timbre of the on-duty personnel in different regions; A language understanding unit, which parses text information and understands the intention of the on-duty personnel; A decision-making unit, which determines how to respond to the instructions of the on-duty personnel according to the result of language understanding; An execution unit executes the operation determined by the decision making unit.
[0005] As a further improvement of the technical solution, the voice receiving and processing unit includes a voice receiving module and a voice processing module; Among them, the voice receiving module captures the sound through an array composed of multiple microphones; The speech processing module analyzes the sound received by different microphones and improves the spectral characteristics of the speech signal by increasing the amplitude of the high-frequency components.
[0006] As a further improvement of the technical solution, the speech recognition unit includes a feature extraction module, an acoustic module, a language module, a word order generation module and a text generation module; The feature extraction module extracts relevant features from the speech signal based on the speech extraction algorithm, optimizes the speech extraction algorithm by introducing a frequency mapping function that integrates dialects and timbre, and further optimizes the speech extraction algorithm so that after the frequency mapping function that integrates dialects and timbre is introduced, the speech extraction algorithm is better compatible with the hidden Markov acoustic model; The acoustic module maps the feature vector of the speech signal to the phoneme unit based on the hidden Markov acoustic model; The language module predicts the probability of the next word based on the N-gram language model; The word order generation module combines the hidden Markov acoustic model and the N-gram language model, uses the word order search algorithm to find the most likely word sequence, and generates a preliminary recognition result; The text generation module evaluates the confidence of the preliminary recognition results, filters out low-confidence recognition results, performs spelling check and correction on the recognition results, adjusts the recognition results according to contextual information, and converts the final recognition results into text format.
[0007] As a further improvement of this technical solution, the speech extraction algorithm is: ; In view of the different dialects and timbre of on-duty personnel in different regions, a frequency mapping function that integrates dialect and timbre is introduced into the speech extraction algorithm: ; The speech extraction algorithm is further optimized to make it more compatible with the acoustic model after the frequency mapping function that integrates dialect and timbre is introduced: ; in, represents the extracted speech features; It represents the speech features extracted after the frequency mapping function of the dialect and timbre fusion is introduced; represents the speech features extracted after further optimization; represents the number of filter banks; represents the output of the filter; An index representing the speech extraction algorithm; Represents the filter index; According to dialect characteristics Calculated adjustment factors; According to the timbre characteristics Calculated adjustment factors; Represents the weight factor.
[0008] As a further improvement of the technical solution, the word order generation module uses a word order search algorithm to find the most likely word sequence, including the following steps: S2.1. For the first eigenvector of the speech signal, calculate the probability of each possible initial state; S2.2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state; S2.3. For each time point and each possible state, calculate the probability of transitioning from all possible previous states to the current state, and combine the duty scene, introduce the traffic terms in the duty scene to optimize the probability calculation, and further optimize the probability calculation so that the probability after the introduction of traffic terms is suitable for different application scenarios; S2.4. Update the probability of the current state at a time point to the maximum probability among all possible previous states, and update the backtracking pointer of the current state at the same time; S2.5. At the last time point, select the state with the highest probability as the most likely end state; S2.6. Starting from the final determined end state, trace back along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence; S2.7. For the most likely hidden state sequence obtained, calculate the probability of the hidden state sequence in the language model, and combine the probabilities of the hidden Markov acoustic model and the N-gram language model to calculate the final comprehensive probability; S2.8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result.
[0009] As a further improvement of the technical solution, in S2.3, the probability of transferring from all possible previous states to the state is calculated as: ; Combined with the duty scene, the traffic terms in the duty scene are introduced to optimize the probability calculation as follows: ; The probability calculation is further optimized to make the probability after introducing traffic terms suitable for different application scenarios: ; in, Indicates a point in time; Indicates the current state; Indicates the previous state; Indicates the total number of states; Indicates at a point in time When , the most likely path to reach state The probability of Represents the probability after the traffic terms in the duty scene are introduced; Indicates the probability of being applicable to different application scenarios; Indicates at a point in time When , the most likely path to reach state probability; Indicates from the state Transfer to state probability; Indicates in status Generate observations The probability of Indicates in status Consider the traffic terms in the on-duty scenario The probability of Represents the scene-related weight.
[0010] As a further improvement of the technical solution, the language understanding unit includes a text recognition module, a sentence analysis module and a dialogue management module; The text recognition module pre-processes the generated text and recognizes named entities in the text through regular expressions to identify specific types of information; The sentence analysis module performs syntactic structure analysis on the sentence through a greedy decoding algorithm, understands the relationship between each word, determines the role played by each part in the sentence, and identifies the specific intention of the on-duty personnel through an intention recognition algorithm based on the analysis results; The dialogue management module tracks the dialogue history and context information, provides more coherent responses, and introduces an attention mechanism in the model.
[0011] As a further improvement of this technical solution, the intention recognition algorithm is: ; The intent recognition algorithm is optimized to address the urgent intentions of on-duty personnel when expressing dissatisfaction and in emergency situations: ; in, Represents the feature vector extracted from the speech recognition result; Indicate the intention of the officer on duty; represents the weight vector; represents the bias term; feature vectors indicating urgency and expressing dissatisfaction; Indicates the degree of matching between the current input features and the emergency and dissatisfaction emotional features; For adjustment The degree of influence on the final prediction results represents the base of natural logarithms; Represents the transpose of a vector.
[0012] As a further improvement of the technical solution, the decision-making unit decides how to respond to the on-duty personnel's instructions according to the result of language understanding, including the following steps: S4.1, matching the user's intent with a predefined set of rules; S4.2. Once the intent is correctly parsed and successfully matched with the rules, a specific action plan is generated and relevant instructions are given to the execution unit; S4.3. Before performing any operation, the decision-making unit performs a confirmation step.
[0013] As a further improvement of the technical solution, the execution unit executes the operation determined by the decision-making unit, including the following steps: S5.1. Verify again whether the instructions from the decision-making unit are clear and reasonable, and check whether the on-duty personnel who issued the instructions have the authority to perform the operation; S5.2, formatting the control signal according to the communication protocol of the target device; S5.3, establishing a communication connection with the target device, and sending a control signal to the target device through the established communication channel; S5.4. Continuously monitor the status of the equipment to determine whether the equipment is started as expected.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. In the intelligent voice recognition and command execution system based on artificial intelligence during duty, by introducing the frequency mapping function of dialect and timbre fusion, the problem of low recognition rate of traditional voice recognition system when facing duty personnel with different regions and accents is effectively solved. Especially in complex and changeable duty environments, this optimization can significantly improve the system's adaptability to regional dialects and special timbres, ensuring that the instructions of duty personnel can be accurately recognized and understood even in a noisy environment, thereby improving the practicality and reliability of the system.
[0015] 2. The intelligent voice recognition and command execution system based on artificial intelligence during the duty process can not only quickly and accurately understand the intentions of the duty personnel, but also automatically make response decisions according to the pre-set rule set, and even add confirmation steps before performing key operations to ensure the safety of operations. Especially in emergency situations, the system can quickly identify and respond to the urgent needs of duty personnel, take appropriate measures in a timely manner, effectively avoid risks caused by human negligence or misjudgment, and greatly improve the efficiency and safety of duty work. In addition, the execution unit will verify the validity and rationality of the instruction again before executing the operation, which further ensures the accuracy of system operation and reduces the possibility of misoperation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is the overall flow chart of the present invention; The meaning of each number in the figure is: 1. Speech receiving and processing unit; 2. Speech recognition unit; 3. Language understanding unit; 4. Decision-making unit; 5. Execution unit. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] Example: See Figure 1 As shown, an intelligent voice recognition and command execution system based on artificial intelligence during duty is provided, including: The speech receiving and processing unit 1 captures the speech signal in the environment and performs preliminary processing; In this embodiment, the voice receiving and processing unit 1 includes a voice receiving module and a voice processing module; The voice receiving module captures sound through an array of multiple microphones. The microphone array can provide spatial information, which is helpful for subsequent sound source positioning and noise reduction processing. The speech processing module analyzes the sounds received by different microphones and improves the spectral characteristics of the speech signal by increasing the amplitude of the high-frequency components, making subsequent processing easier.
[0019] The speech recognition unit 2 converts the initially processed speech signal into text information; In this embodiment, the speech recognition unit 2 includes a feature extraction module, an acoustic module, a language module, a word order generation module and a text generation module; The feature extraction module extracts relevant features from the speech signal based on the speech extraction algorithm, optimizes the speech extraction algorithm by introducing a frequency mapping function that integrates dialects and timbre, and further optimizes the speech extraction algorithm so that after the frequency mapping function that integrates dialects and timbre is introduced, the speech extraction algorithm is better compatible with the hidden Markov acoustic model; Furthermore, the original signal is converted into a low-dimensional feature vector that can characterize the speech characteristics through a speech extraction algorithm. The speech extraction algorithm is: ; On-duty personnel in different regions may use different dialects, which have significant differences in speech features. By introducing a frequency mapping function that integrates dialects and timbre, these differences can be better captured and processed, thereby improving the system's ability to recognize different dialects. The timbre (such as pitch, sound quality, etc.) of on-duty personnel varies, and these differences will affect the spectral characteristics of the speech signal. By introducing a frequency mapping function that integrates timbre, the speech extraction algorithm can be optimized to make it more adaptable to the characteristics of different timbres, thereby improving recognition accuracy. By introducing a frequency mapping function that integrates dialects and timbre, the system can maintain a high recognition performance under different environments and conditions, and is not affected by differences in dialects and timbre, which enables the system to operate stably in various on-duty scenarios, thereby improving the reliability and practicality of the system. In view of the different dialects and timbre of on-duty personnel in different regions, a frequency mapping function that integrates dialects and timbre is introduced into the speech extraction algorithm. : If there is an original speech signal , dialect features and timbre characteristics Represented as a set of parameters and ,in Represents the frequency distribution of specific phonemes in a dialect, Indicates the mean or variance of the speaker's pitch; ; Among them, the frequency mapping function that introduces the fusion of dialect and timbre Used to describe how to , dialect features and timbre characteristics To calculate the new eigenvector , by integrating dialect and timbre information, the feature extraction process of speech signals is improved, so that the extracted features can better reflect the uniqueness of the speaker, thereby improving the performance of the speech recognition system; Dialect feature adjustment factor is (through the product form, each dialect feature will have an impact on the final frequency mapping result, and the degree of influence is determined by control): ; In the formula, Indicates dialect features The number of dimensions, i.e., the dialect feature vector the number of different dialectal features contained in Indicates dialect features The index in is used to traverse all dialect features; Tone Characteristic Adjustment Factor is (through exponential decay, each timbre feature The greater the deviation, the smaller its impact on the final result, thus achieving effective adaptation to different timbres): ; Speech recognition systems usually include multiple links, such as acoustic models, language models, etc. The speech extraction algorithm after introducing the frequency mapping function that integrates dialects and timbre may be incompatible with the existing acoustic model. The existing acoustic model is trained based on general speech features. The new speech extraction algorithm may cause the data distribution input to the acoustic model to change, resulting in a decrease in the performance of the acoustic model. For example, if the speech extraction algorithm changes the representation of speech features, and the acoustic model is not adjusted for this new representation, then the acoustic model may make errors in classifying and recognizing these speech features. The speech extraction algorithm is further optimized to make the speech extraction algorithm after introducing the frequency mapping function that integrates dialects and timbre more compatible with the acoustic model: ; in, Represents the extracted speech feature, which is one of the features extracted from the speech signal and is used to represent the spectral characteristics of the speech signal; It represents the speech features extracted after the frequency mapping function of the dialect and timbre fusion is introduced; represents the speech features extracted after further optimization; Indicates the number of filter banks. On the Mel frequency scale, a set of filters will be used to capture the energy in different frequency ranges. Represents the output of the filter, which is logarithmic energy. Each filter captures the energy within a specific frequency range and takes the logarithm to compress the dynamic range; An index representing the speech extraction algorithm; Represents the filter index; According to dialect characteristics Calculated adjustment factors; According to the timbre characteristics Calculated adjustment factors; Indicates adjustment of each dialect feature Weight parameters that influence the final result; Indicates adjustment of each timbre characteristic Weight parameters that influence the final result; A baseline value that represents a timbre feature, such as the average pitch of a speaker; represents the weight factor used to balance the contribution of the standard mel frequency map and the dialect-specific frequency map, The value range of is from 0 to 1, which determines the relative importance of the two mapping methods; The acoustic module maps the feature vector of the speech signal to the phoneme unit based on the hidden Markov acoustic model. The hidden Markov model (HMM) is a statistical model used to model time series data, especially those with hidden states. In speech recognition, these hidden states are usually phonemes or subphoneme units, and the observation value is the feature vector extracted from the speech signal; The language module predicts the probability of the next word based on the N-gram model. The N-gram model is a statistical model based on historical word sequences, which predicts the probability of the next word by counting the frequency of historical word sequences. The word order generation module combines the hidden Markov acoustic model and the N-gram language model, and uses a word order search algorithm (using dynamic programming to solve the most likely hidden state sequence. At any time, there is only one most likely path to reach the current state, so the algorithm will track these maximum probability paths and finally select the path with the highest total probability as the most likely state sequence) to find the most likely word sequence. This step aims to select the best path from all possible candidate word sequences and generate preliminary recognition results; Furthermore, the most likely word sequence is found so that the probability of generating the word sequence is maximized, which ensures that the system can accurately convert voice signals into text information and reduce misrecognition and missed recognition. In the process of on-duty, accurate voice recognition is crucial to executing correct instructions, especially in emergency situations, where incorrect recognition may lead to serious consequences. Through the word order search algorithm, the optimal solution is found in polynomial time, avoiding the high computational cost of exhausting all possible word sequences. In practical applications, the system needs to process a large number of voice signals in a short time. Efficient algorithms can ensure real-time performance and response speed. Combined with the joint probability of the HMM and N-gram models, the voice features and language models are comprehensively considered to improve the accuracy of recognition. Using the word order search algorithm to find the most likely word sequence includes the following steps: S2.1. For the first eigenvector of the speech signal, calculate the probability of each possible initial state ( , is the probability of the initial state, It is observed in the initial state probability of ); S2.2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state. At the initial moment, since there is no previous state, these pointers can be set to meaningless values; S2.3. For each time point (starting from 2 until the last observation value) and each possible state, calculate the probability of transitioning from all possible previous states to the current state, and combine the duty scenarios (such as traffic control, patrol inspection, etc.), introduce the traffic terms in the duty scenarios to optimize the probability calculation, improve the accuracy and rationality of recognition, and further optimize the probability calculation so that the probability after the introduction of traffic terms is suitable for different application scenarios; Among them, the probability of transferring to the state from all possible previous states is calculated as: ; By introducing traffic terms in the on-duty scenario, the system can give priority to identifying these words and reduce the possibility of misidentification. For example, in the traffic command scenario, the system can give priority to identifying command words such as "stop", "go", and "turn left" instead of other irrelevant words. In the on-duty scenario, some words may have multiple meanings. By introducing traffic terms, the system can better understand the context, reduce ambiguity, and improve the accuracy of recognition. For example, "turn left" usually refers to vehicle turning in traffic command, but may have different meanings in other scenarios; by introducing traffic terms, the system can ensure the semantic consistency of the recognition results. For example, in traffic command, the system can ensure that commands such as "stop at red light and go at green light" are reasonable, and unreasonable commands such as "go at red light" will not appear; combined with on-duty scenarios (such as traffic command, patrol inspection, etc.), the introduction of traffic terms in the on-duty scenario optimizes the probability calculation as follows: ; ; Among them, Defined as a normalized weight, specifically: for each state Calculate the matching score with traffic terms , calculate each state The normalized weight of ; The optimized probability calculation method can be applied to a variety of duty scenarios, such as traffic control, patrol inspection, emergency response, etc., to improve the universality and practicality of the system. The system can dynamically adjust the probability calculation parameters according to different scenarios to ensure that high recognition performance can be maintained in various duty tasks; the probability calculation is further optimized so that the probability after the introduction of traffic terms is suitable for different application scenarios: ; ; in, Indicates the time point, that is, the first observations; Indicates the current state, that is, at a point in time One of the possible states when Indicates the previous state, that is, at the time point One of the possible states when Represents the total number of states, that is, the number of all possible states; Indicates at a point in time When , the most likely path to reach state probability; Represents the probability after the traffic terms in the duty scene are introduced; Indicates the probability of being applicable to different application scenarios; Indicates at a point in time When , the most likely path to reach state probability; Indicates from the state Transfer to state probability; Indicates in status Generate observations The probability of It's at the time Observed value at time ; Indicates in status Consider the traffic terms in the on-duty scenario probability; Represents scene-related weights, used to adjust When When increasing the importance of traffic terms, When reducing the importance of traffic terms; Indicates status Traffic terminology in on-duty scenarios The matching score of Indicates traffic terms in on-duty scenarios; Represents the index when calculating the weight of traffic terms in the duty scenario, that is, one of all possible states; A parameter representing the control of scoring sensitivity; S2.4. Among all possible previous states, update the probability of the current state at the time point to this maximum probability, and at the same time update the backtracking pointer of the current state; S2.5. At the last time point, select the state with the highest probability as the most likely end state, and the probability of this state is the probability of the entire most likely hidden state sequence; S2.6. Starting from the finally determined end state, backtrack along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence; S2.7. For the obtained most likely hidden state sequence, calculate the probability of the hidden state sequence in the language model ( , which is the conditional probability in the language model), and combine the probabilities of the hidden Markov acoustic model and the N-gram language model to calculate the final comprehensive probability ( ); S2.8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result; The text generation module evaluates the confidence of the preliminary recognition result, filters out the recognition results with low confidence, performs spelling check and correction on the recognition results to improve the readability of the text, adjusts the recognition results according to the context information to ensure semantic consistency and rationality, and converts the final recognition result into text format.
[0020] The language understanding unit 3 analyzes the text information to understand the intention of the on-duty personnel; In this embodiment, the language understanding unit 3 includes a text recognition module, a sentence analysis module, and a dialogue management module; Among them, the text recognition module preprocesses the generated text, including removing stop words (deleting common words such as "of", "is", "in" in the text, which usually do not help much in understanding the meaning of the sentence), word segmentation (dividing continuous text into individual words or phrases for further processing), stemming or lemmatization (restoring words to their basic forms to reduce the impact of word variants), and identifying named entities in the text through regular expressions (regular expressions are tools for describing patterns to match strings. By writing regular expressions, specific patterns can be defined to identify and extract named entities in the text), such as personal names, place names, organization names, etc., and identifying specific types of information such as time, date, quantity, etc.; The sentence analysis module performs syntactic structure analysis on the sentence through a greedy decoding algorithm (a greedy decoding algorithm is a strategy for building a dependency tree word by word. In each step, the current optimal dependency relationship is selected to gradually build a dependency tree for the entire sentence), understands the relationship between each word, such as the subject-predicate-object structure, etc., determines the role played by each part in the sentence, such as the executor of the action (agent), the object of the action (patient), etc., and identifies the specific intention of the on-duty personnel through an intention recognition algorithm based on the analysis results, which may be to query information, request services, express opinions, etc.; Furthermore, the intent recognition algorithm vectorizes the text features and uses the algorithm to learn the mapping relationship from the input features to the intent category, thereby realizing the classification and recognition of the text intent; the intent recognition algorithm ensures that the system can accurately understand the intent of the on-duty personnel. Regardless of whether the instructions are given in spoken form, the intent recognition algorithm can parse the core content and purpose of the instructions; by quickly and accurately identifying the intent of the on-duty personnel, the system can immediately generate corresponding responses or perform operations, reducing response time, providing a more natural and smooth interactive experience, allowing on-duty personnel to interact with the system more easily and reducing operational complexity; on-duty personnel can communicate with the system through natural language without having to remember complex command syntax, thereby improving work efficiency; the intent recognition algorithm is: ; In an emergency, the system can quickly identify the emergency intentions of the on-duty personnel and take corresponding measures immediately to reduce response time and ensure safety. In emergencies, such as traffic accidents and fires, the system can quickly identify instructions such as "emergency evacuation" and "call for rescue" and immediately initiate emergency plans. When the on-duty personnel express dissatisfaction, the system can accurately identify their negative emotions and specific demands, take timely measures to alleviate conflicts and improve service quality. If the on-duty personnel are dissatisfied with a decision or operation, the system can promptly identify and provide corresponding solutions or explanations to reduce misunderstandings and conflicts. By optimizing the intention recognition algorithm, the system can maintain high recognition accuracy and reliability in emergency situations and dissatisfaction to ensure the smooth completion of the task. The optimized algorithm can identify the emotional characteristics of on-duty personnel, such as anger and anxiety, and improve the recognition accuracy of emergency situations and dissatisfaction. In view of the urgency of the intentions of on-duty personnel when expressing dissatisfaction and emergency situations, the intention recognition algorithm is optimized: ; ; in, Represents the feature vector extracted from the speech recognition result; Indicate the intention of the officer on duty; Represents a weight vector, each element of which corresponds to an input feature vector The importance of a feature in Represents the bias term, which is regarded as the initial value of the model and helps to improve the flexibility of the model; feature vectors indicating urgency and expressing dissatisfaction; Indicates the degree of matching between the current input features and the emergency and dissatisfaction emotional features; For adjustment The degree of influence on the final prediction result is adjusted by The value of can find a suitable balance between normal and emergency situations; represents the base of natural logarithms; Represents the transpose of a vector; The weight representing the pitch change; Indicates the change value of pitch; The weight indicating the speed of speech is increased; Indicates the change value of speech speed; The dialogue management module tracks the dialogue history and context information to provide more coherent responses, and introduces an attention mechanism into the model so that the model can better focus on key information in the dialogue; Among them, the dialogue management module is specifically as follows: maintain a dialogue history record, store the text content, timestamp and related metadata of each dialogue, set a context window, and only retain the content of the most recent few rounds of dialogue to reduce the computational burden and maintain the relevance of the dialogue; extract key information from the dialogue history, such as time, place, people, events, etc., to form context features, calculate the correlation between the current dialogue and the historical dialogue, and ensure the relevance and coherence of the response; use a feedforward neural network to calculate the similarity score between the current dialogue features and the historical dialogue features; weighted sum the historical dialogue features according to the attention weight to obtain the context vector; combine the current dialogue features and the context vector as input to generate the final response.
[0021] The execution unit 5 executes the operation determined by the decision making unit 4; In this embodiment, the execution unit 5 executes the operation determined by the decision-making unit 4, including the following steps: S5.1. Verify again whether the instruction from the decision-making unit 4 is clear and reasonable, for example, confirm which device is to be turned on and whether the operation complies with the current security protocol, and check whether the on-duty personnel who issued the instruction have the authority to perform the operation to prevent unauthorized access; S5.2. Format the control signal according to the communication protocol of the target device. Different devices may require signals in different formats. For example, some devices may accept digital signals, while others may require analog signals or specific network protocols. Ensure that the control signal sent will not cause damage to the device or system. For example, check whether the signal exceeds the tolerance range of the device; S5.3. Establish a communication connection with the target device, send a control signal to the target device through the established communication channel, and wait for the device to return a confirmation message to ensure that the device has received the control signal. If the confirmation message is not received, try to resend or report a fault; S5.4. Continuously monitor the status of the device to determine whether the device starts as expected, including reading the device's status registers, sensor data, or other feedback information. If the device fails to start normally, take other measures, such as retrying the startup, restarting the device, reporting a fault, etc.
[0022] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only preferred examples of the present invention, and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.
Claims
1. An intelligent voice recognition and command execution system based on artificial intelligence during duty, characterized in that: include: A speech receiving and processing unit (1), wherein the speech receiving and processing unit (1) captures speech signals in the environment and performs preliminary processing; A speech recognition unit (2), wherein the speech recognition unit (2) converts the initially processed speech signal into text information. During the conversion process, a frequency mapping function for integrating the dialect and the timbre is introduced in view of the different dialects and timbre of the on-duty personnel in different regions; A language understanding unit (3), wherein the language understanding unit (3) parses text information and understands the intention of the on-duty personnel; A decision-making unit (4), wherein the decision-making unit (4) decides how to respond to the instructions of the on-duty personnel according to the result of language understanding; An execution unit (5), wherein the execution unit (5) executes the operation determined by the decision-making unit (4).
2. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 1 is characterized by: The speech receiving and processing unit (1) comprises a speech receiving module and a speech processing module; Among them, the voice receiving module captures the sound through an array composed of multiple microphones; The speech processing module analyzes the sound received by different microphones and improves the spectral characteristics of the speech signal by increasing the amplitude of the high-frequency components.
3. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 2 is characterized by: The speech recognition unit (2) comprises a feature extraction module, an acoustic module, a language module, a word order generation module and a text generation module; The feature extraction module extracts relevant features from the speech signal based on the speech extraction algorithm, optimizes the speech extraction algorithm by introducing a frequency mapping function that integrates dialects and timbre, and further optimizes the speech extraction algorithm so that after the frequency mapping function that integrates dialects and timbre is introduced, the speech extraction algorithm is better compatible with the hidden Markov acoustic model; The acoustic module maps the feature vector of the speech signal to the phoneme unit based on the hidden Markov acoustic model; The language module predicts the probability of the next word based on the N-gram language model; The word order generation module combines the hidden Markov acoustic model and the N-gram language model, uses the word order search algorithm to find the most likely word sequence, and generates a preliminary recognition result; The text generation module evaluates the confidence of the preliminary recognition results, filters out low-confidence recognition results, performs spelling check and correction on the recognition results, adjusts the recognition results according to contextual information, and converts the final recognition results into text format.
4. According to claim 3, the intelligent voice recognition and command execution system based on artificial intelligence during duty is characterized in that: The speech extraction algorithm is: ; In view of the different dialects and timbre of on-duty personnel in different regions, a frequency mapping function that integrates dialect and timbre is introduced into the speech extraction algorithm: ; The speech extraction algorithm is further optimized to make it more compatible with the acoustic model after the frequency mapping function that integrates dialect and timbre is introduced: ; in, represents the extracted speech features; It represents the speech features extracted after the frequency mapping function of the dialect and timbre fusion is introduced; represents the speech features extracted after further optimization; represents the number of filter banks; represents the output of the filter; An index representing a speech extraction algorithm; Represents the filter index; According to dialect characteristics Calculated adjustment factors; According to the timbre characteristics Calculated adjustment factors; Represents the weight factor.
5. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 4 is characterized in that: The word order generation module uses a word order search algorithm to find the most likely word sequence, including the following steps: S2.
1. For the first eigenvector of the speech signal, calculate the probability of each possible initial state; S2.
2. Set a backtracking pointer for each state, pointing to the previous state that is most likely to lead to the current state; S2.
3. For each time point and each possible state, calculate the probability of transitioning from all possible previous states to the current state, and combine the duty scene, introduce the traffic terms in the duty scene to optimize the probability calculation, and further optimize the probability calculation so that the probability after the introduction of traffic terms is suitable for different application scenarios; S2.
4. Update the probability of the current state at a time point to the maximum probability among all possible previous states, and update the backtracking pointer of the current state at the same time; S2.
5. At the last time point, select the state with the highest probability as the most likely end state; S2.
6. Starting from the final determined end state, trace back along the backtracking pointer of each state until returning to the initial state to obtain the most likely hidden state sequence; S2.
7. For the most likely hidden state sequence obtained, calculate the probability of the hidden state sequence in the language model, and combine the probabilities of the hidden Markov acoustic model and the N-gram language model to calculate the final comprehensive probability; S2.
8. Convert the most likely hidden state sequence into a word sequence to generate the final recognition result.
6. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 5 is characterized by: In S2.3, the probability of transitioning from all possible previous states to the state is calculated as: ; Combined with the duty scene, the traffic terms in the duty scene are introduced to optimize the probability calculation as follows: ; The probability calculation is further optimized to make the probability after introducing traffic terms suitable for different application scenarios: ; in, Indicates a point in time; Indicates the current state; Indicates the previous state; Indicates the total number of states; Indicates at a point in time When , the most likely path to reach state probability; Represents the probability after the traffic terms in the duty scene are introduced; Indicates the probability of being applicable to different application scenarios; Indicates at a point in time When , the most likely path to reach state probability; Indicates from the state Transfer to state probability; Indicates in status Generate observations probability; Indicates in status Consider the traffic terms in the on-duty scenario probability; Represents the scene-related weight.
7. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 6 is characterized by: The language understanding unit (3) includes a text recognition module, a sentence analysis module and a dialogue management module; The text recognition module pre-processes the generated text and recognizes named entities in the text through regular expressions to identify specific types of information; The sentence analysis module performs syntactic structure analysis on the sentence through a greedy decoding algorithm, understands the relationship between each word, determines the role played by each part in the sentence, and identifies the specific intention of the on-duty personnel through an intention recognition algorithm based on the analysis results; The dialogue management module tracks the dialogue history and context information, provides more coherent responses, and introduces an attention mechanism in the model.
8. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 7 is characterized in that: The intent recognition algorithm is: ; The intent recognition algorithm is optimized to address the urgent intentions of on-duty personnel when expressing dissatisfaction and in emergency situations: ; in, Represents the feature vector extracted from the speech recognition result; Indicate the intention of the officer on duty; represents the weight vector; represents the bias term; feature vectors indicating urgency and expressing dissatisfaction; Indicates the degree of matching between the current input features and the emergency and dissatisfaction emotional features; For adjustment The degree of influence on the final forecast results; represents the base of natural logarithms; Represents the transpose of a vector.
9. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 8 is characterized by: The decision-making unit (4) determines how to respond to the on-duty personnel's instructions based on the result of language understanding, including the following steps: S4.1, matching the user's intent with a predefined set of rules; S4.
2. Once the intent is correctly parsed and successfully matched with the rules, a specific action plan is generated and relevant instructions are sent to the execution unit (5); S4.
3. Before performing any operation, the decision-making unit (4) performs a confirmation step.
10. The intelligent voice recognition and command execution system based on artificial intelligence during duty according to claim 9 is characterized in that: The execution unit (5) executes the operation determined by the decision-making unit (4), including the following steps: S5.
1. Verify again whether the instruction from the decision-making unit (4) is clear and reasonable, and check whether the on-duty personnel who issued the instruction have the authority to perform the operation; S5.2, formatting the control signal according to the communication protocol of the target device; S5.3, establishing a communication connection with the target device, and sending a control signal to the target device through the established communication channel; S5.
4. Continuously monitor the status of the equipment to determine whether the equipment is started as expected.
Citation Information
Patent Citations
Speech recognition model training method and device, equipment and storage medium
CN118136001A
Machine tool instruction execution system and method based on voice interaction
CN118197308A
Telephone customer service processing method and system based on personalized robot
CN118433311A
Agricultural disease and pest dialect voice intelligent recognition method based on deep learning technology
CN119296514A
A spoken dialogue system, a spoken dialogue method and a method of adapting a spoken dialogue system
GB201905974D0
Cited By
Vehicle-mounted speech recognition law enforcement linkage system and method based on edge calculation
CN120452427A