Speech recognition system based on artificial intelligence algorithm

By adopting artificial intelligence algorithms and generative adversarial network framework in the speech recognition system, the problems of unstable data quality and lack of real-time performance monitoring of traditional speech recognition systems are solved, and efficient and robust speech recognition and real-time performance monitoring are achieved, improving the system's adaptability and user experience.

CN120015015AInactive Publication Date: 2025-05-16SHANGHAI TECHN INST OF ELECTRONICS & INFORMATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169882.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional speech recognition systems are susceptible to environmental noise and equipment quality differences during the data acquisition stage, resulting in unstable data quality, affecting model training effects, and lack real-time performance monitoring mechanisms, resulting in service interruption or impaired user experience.

Method used

A speech recognition system based on artificial intelligence algorithm is adopted, including data acquisition module, data processing module, adversarial training module, model building module, monitoring module and security module. Voice data is collected through USB microphone, data preprocessing is performed using Python scripts and PyDub, a generative adversarial network framework is built to generate adversarial samples, and a speech recognition model is built based on the deep neural network architecture. At the same time, a comprehensive threshold is set for real-time performance monitoring, and the policy is adjusted and performance reports are generated based on the comparison results.

Benefits of technology

It realizes the collection and processing of high-quality voice data, enhances the robustness and generalization capabilities of the voice recognition model, improves the system's adaptability to complex environments, and promptly detects and responds to system abnormalities through real-time performance monitoring, avoiding service interruptions and user experience losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015015A_ABST
    Figure CN120015015A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition systems, and discloses a voice recognition system based on an artificial intelligence algorithm, which comprises a data acquisition module, a data processing module, an adversarial training module, a model construction module, a monitoring module and a safety module. The data acquisition module is used for acquiring a voice sample by adopting audio recording equipment to obtain an original voice data set; and the data processing module is used for preprocessing the original voice data set by adopting an automatic script method to obtain a high-quality training data set. According to the voice recognition system based on the artificial intelligence algorithm, a USB microphone Blue Yeti is used for recording in a relatively quiet environment, high-quality voice data collection is achieved, the consistency and reliability of the data are ensured, all original voice data are imported and standardized in batches through a Python script and an open source audio processing library PyDub, and the voice recognition efficiency is improved. And data format unification and quality optimization are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition systems, and in particular to a speech recognition system based on an artificial intelligence algorithm. Background Art

[0002] Speech recognition system is a technology that converts human language into text and is widely used in smart assistants, telephone customer service systems, voice input methods and other fields.

[0003] However, traditional speech recognition systems are easily affected by factors such as environmental noise and differences in the quality of recording equipment during the data collection stage, resulting in uneven quality of raw data and affecting the model training effect. In addition, traditional models have limited speech recognition capabilities in complex environments and are easily affected by factors such as background noise and accent changes, resulting in reduced recognition accuracy. At the same time, the existing system lacks an effective real-time performance monitoring mechanism and is unable to detect and respond to system anomalies in a timely manner, resulting in service interruptions or impaired user experience. Summary of the invention

[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] To achieve the above object, the present invention provides the following technical solutions: A speech recognition system based on artificial intelligence algorithm, comprising: Data acquisition module, data processing module, adversarial training module, model building module, monitoring module and security module; The data acquisition module is used to acquire the voice samples by using an audio recording device to obtain an original voice data set; The data processing module is used to pre-process the original speech data set by adopting an automated script method to obtain a high-quality training data set; The adversarial training module is used to generate adversarial samples for high-quality training data by adopting a generative adversarial network method to obtain test data; The model building module is used to build a speech recognition model based on a deep neural network architecture, input test data into the model, use a forward propagation algorithm to calculate the activation value of each layer, and output the speech recognition result; The monitoring module is used to set a comprehensive threshold, compare it with the speech recognition result, and determine whether it exceeds the normal range after comparison to obtain the comparison result; The security module sets policies for adjustment based on the comparison results and generates a performance report.

[0006] As a further solution of the present invention: the audio recording device is used to collect the speech samples to obtain the original speech data set, and the specific steps are as follows: Voice samples were collected by using a USB microphone Blue Yeti; The recording process must be carried out in a relatively quiet environment without obvious background noise, and the recording duration is set between 15 seconds and 60 seconds per segment; Get the original speech dataset.

[0007] As a further solution of the present invention: the original speech data set is preprocessed by using an automated script method to obtain a high-quality training data set, and the specific steps are: Use Python scripts and the open source audio processing library PyDub to batch import all raw speech datasets; Convert all non-WAV audio files to standard mono 16-bit PCM-encoded WAV format, and adjust the sampling rate of all audio files to a unified standard value; Removes silence or low volume segments by detecting audio energy levels that are continuously below a set threshold; Apply an adaptive filter to suppress background noise in the audio and retain the vocal part, and use the peak normalization method to adjust the volume of each audio clip to keep it consistent within a certain range; Automatically segment long audio files at predetermined time intervals, and add corresponding text tags to each segment by calling the automatic speech recognition ASR service to form synchronized subtitle files.

[0008] As a further solution of the present invention: the adversarial samples are generated by using the generative adversarial network method on the high-quality training data to obtain the test data, and the specific steps are as follows: Construct a generative adversarial network framework, including a generator and a discriminator, and generate adversarial samples from high-quality training data. The expression is: ; in, Indicates adversarial examples, Is an index collection containing all selected The index of the adversarial example, represents a normal sample randomly drawn from the original high-quality training data, Is an index collection containing all selected The index of the original training sample, For test data.

[0009] As a further solution of the present invention: the speech recognition model is constructed based on the deep neural network architecture, by inputting test data into the model, using the forward propagation algorithm to calculate the activation value of each layer, and outputting the speech recognition result, the specific steps are: Build a speech recognition model based on the deep neural network architecture WaveNet; Test data obtained through adversarial sample generation and screening Convert to a format acceptable to the model; The forward propagation algorithm is used to calculate the activation value of each layer, passing through the neurons of each layer in turn, and the activation value is calculated using the activation function ReLU; For each test audio clip ∈ , pass it as input to the model, and calculate the activation value of each layer through the forward propagation algorithm until the final output probability distribution is obtained, which is expressed as: ; in, Indicates Test samples The final recognition result is is the set of all possible characters, is a given input After that, the characters Probability of occurrence; Combining the results of all test samples, the expression is: ; in, It is a collection containing the recognition results of all test samples.

[0010] As a further solution of the present invention: the comprehensive threshold is set, compared with the speech recognition result, and after the comparison, it is determined whether it exceeds the normal range to obtain the comparison result. The specific steps are: Set comprehensive thresholds based on historical data , the expression is: ; in, , , is the weight coefficient of each performance index; By using logging tools, we monitor the processing time and resource consumption of each request and obtain performance statistics. The output of the speech recognition system Compare with the actual label or user feedback to calculate the comprehensive score of each sample , the expression is: ; The comprehensive score corresponding to the current speech recognition result With the pre-set comprehensive threshold Compare and determine whether it exceeds the normal range.

[0011] As a further solution of the present invention: the strategy is set based on the comparison result to make adjustments and generate a performance report, and the specific steps are: Based on the comparison results; when Greater than When an error occurs, an alarm notification is triggered immediately, and an alarm message is sent to the administrator, indicating that there is a potential problem that needs attention and resolution. The alarm can be sent via email, SMS, or internal notification messages of the system. All abnormal situations and their occurrence time, the specific test samples involved, and related performance indicators are recorded in detail. when Less than or equal to If the current comprehensive score If it is within the normal range, the existing system configuration is maintained unchanged and the real-time performance of the system continues to be monitored; Regularly summarize performance monitoring data, including word error rate, sentence error rate, response time, and system resource usage.

[0012] As a further solution of the present invention: the activation value of each layer is calculated by using the forward propagation algorithm, and the activation value is calculated by using the activation function ReLU through the neurons of each layer in turn. The expression is: ; ; in It is The weight matrix of the layer, It is The bias vector of the layer, is the activation function, It is The activation value of the layer, is the character corresponding to The unnormalized prediction score of .

[0013] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the speech recognition system based on the artificial intelligence algorithm as described in the first aspect of the present invention is implemented.

[0014] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the speech recognition system based on an artificial intelligence algorithm as described in the first aspect of the present invention is implemented.

[0015] Compared with the prior art, the present invention has the following beneficial effects: By using a USB microphone Blue Yeti to record in a relatively quiet environment, high-quality voice data can be collected and data consistency and reliability can be ensured. By using Python scripts and the open source audio processing library PyDub to batch import and standardize all raw voice data, data format unification and quality optimization can be achieved, reducing the problem of invalid data occupying storage space. At the same time, it also greatly reduces the need for manual intervention, improves the efficiency and accuracy of data preparation, and provides reliable data support for building an efficient speech recognition model. By building a generative adversarial network framework, adversarial samples are generated from high-quality training data, which enhances the robustness of the model. This not only challenges the recognition ability of the existing model and promotes the improvement of the model's generalization ability, but also provides an effective means to verify and improve the performance of the model, thereby enhancing the system's ability to cope with complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Schematic diagram of a speech recognition system based on artificial intelligence algorithm. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and understandable, the specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0018] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0019] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0020] Example 1 See also Figure 1, which is the first embodiment of the present invention, and which provides a speech recognition system based on an artificial intelligence algorithm, comprising: a data acquisition module, a data processing module, an adversarial training module, a model building module, a monitoring module and a security module; A data collection module is used to collect voice samples by using an audio recording device to obtain an original voice data set; Voice samples were collected by using a USB microphone Blue Yeti; The recording process must be carried out in a relatively quiet environment without obvious background noise, and the recording duration is set between 15 seconds and 60 seconds per segment; Get the original speech dataset.

[0021] A data processing module, used for preprocessing the original speech data set by adopting an automated script method to obtain a high-quality training data set; Use Python scripts and the open source audio processing library PyDub to batch import all raw speech datasets; Convert all non-WAV audio files to standard mono 16-bit PCM-encoded WAV format, and adjust the sampling rate of all audio files to a unified standard value; Removes silence or low volume segments by detecting audio energy levels that are continuously below a set threshold; Apply an adaptive filter to suppress background noise in the audio and retain the vocal part, and use the peak normalization method to adjust the volume of each audio clip to keep it consistent within a certain range; Automatically segment long audio files at predetermined time intervals, and add corresponding text tags to each segment by calling the automatic speech recognition ASR service to form synchronized subtitle files.

[0022] An adversarial training module is used to generate adversarial samples for high-quality training data by adopting a generative adversarial network method to obtain test data; Construct a generative adversarial network framework, including a generator and a discriminator, and generate adversarial samples from high-quality training data. The expression is: ; in, Indicates adversarial examples, Is an index collection containing all selected The index of the adversarial example, represents a normal sample randomly drawn from the original high-quality training data, Is an index collection containing all selected The index of the original training sample, For test data.

[0023] The model building module is used to build a speech recognition model based on a deep neural network architecture. By inputting test data into the model, the forward propagation algorithm is used to calculate the activation value of each layer and output the speech recognition result. Build a speech recognition model based on the deep neural network architecture WaveNet; Test data obtained through adversarial sample generation and screening Convert to a format acceptable to the model; The forward propagation algorithm is used to calculate the activation value of each layer, passing through the neurons of each layer in turn, and the activation value is calculated using the activation function ReLU; For each test audio clip ∈ , pass it as input to the model, and calculate the activation value of each layer through the forward propagation algorithm until the final output probability distribution is obtained, which is expressed as: ; in, Indicates Test samples The final recognition result is is the set of all possible characters, is a given input After that, the characters Probability of occurrence; Combining the results of all test samples, the expression is: ; in, It is a set containing all the recognition results of test samples; The forward propagation algorithm is used to calculate the activation value of each layer, passing through the neurons of each layer in turn, and the activation value is calculated using the activation function ReLU. The expression is: ; ; in It is The weight matrix of the layer, It is The bias vector of the layer, is the activation function, It is The activation value of the layer, is the character corresponding to The unnormalized prediction score of .

[0024] The monitoring module is used to set a comprehensive threshold value, compare it with the speech recognition result, and determine whether it exceeds the normal range after comparison to obtain the comparison result; Set comprehensive thresholds based on historical data , the expression is: ; in, , , is the weight coefficient of each performance index; By using logging tools, we monitor the processing time and resource consumption of each request and obtain performance statistics. The output of the speech recognition system Compare with the actual label or user feedback to calculate the comprehensive score of each sample , the expression is: ; The comprehensive score corresponding to the current speech recognition result With the pre-set comprehensive threshold Compare and determine whether it exceeds the normal range.

[0025] The security module sets policies for adjustment based on the comparison results and generates performance reports; Based on the comparison results; when Greater than When an error occurs, an alarm notification is triggered immediately, and an alarm message is sent to the administrator, indicating that there is a potential problem that needs attention and resolution. The alarm can be sent via email, SMS, or internal notification messages of the system. All abnormal situations and their occurrence time, the specific test samples involved, and related performance indicators are recorded in detail. when Less than or equal to If the current comprehensive score If it is within the normal range, the existing system configuration is maintained unchanged and the real-time performance of the system continues to be monitored; Regularly summarize performance monitoring data, including word error rate, sentence error rate, response time, and system resource usage.

[0026] This embodiment also provides a computer device, which is suitable for a speech recognition system based on an artificial intelligence algorithm, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement a speech recognition system based on an artificial intelligence algorithm as proposed in the above embodiment.

[0027] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0028] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, a speech recognition system based on an artificial intelligence algorithm as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage device, a flash memory, a disk or an optical disk.

[0029] In summary, by using the USB microphone Blue Yeti to record in a relatively quiet environment, we can collect high-quality voice data and ensure the consistency and reliability of the data. By using Python scripts and the open source audio processing library PyDub to batch import and standardize all raw voice data, we can achieve data format unification and quality optimization, reduce the problem of invalid data occupying storage space, and greatly reduce the need for manual intervention. It also improves the efficiency and accuracy of data preparation, and provides reliable data support for building an efficient speech recognition model. By building a generative adversarial network framework, adversarial samples are generated from high-quality training data, which enhances the robustness of the model. This not only challenges the recognition ability of the existing model and promotes the improvement of the model's generalization ability, but also provides an effective means to verify and improve the performance of the model, thereby enhancing the system's ability to cope with complex environments.

[0030] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A speech recognition system based on artificial intelligence algorithm, characterized in that: include: Data acquisition module, data processing module, adversarial training module, model building module, monitoring module and security module; The data acquisition module is used to acquire the voice samples by using an audio recording device to obtain an original voice data set; The data processing module is used to pre-process the original speech data set by adopting an automated script method to obtain a high-quality training data set; The adversarial training module is used to generate adversarial samples for high-quality training data by adopting a generative adversarial network method to obtain test data; The model building module is used to build a speech recognition model based on a deep neural network architecture, input test data into the model, use a forward propagation algorithm to calculate the activation value of each layer, and output the speech recognition result; The monitoring module is used to set a comprehensive threshold, compare it with the speech recognition result, and determine whether it exceeds the normal range after comparison to obtain the comparison result; The security module sets policies for adjustment based on the comparison results and generates a performance report.

2. A speech recognition system based on artificial intelligence algorithm according to claim 1, characterized in that: The method of collecting speech samples by using an audio recording device to obtain an original speech data set includes the following specific steps: Voice samples were collected by using a USB microphone Blue Yeti; The recording process must be carried out in a relatively quiet environment without obvious background noise, and the recording duration is set between 15 seconds and 60 seconds per segment; Get the original speech dataset.

3. A speech recognition system based on artificial intelligence algorithm according to claim 2, characterized in that: The original speech data set is preprocessed by using an automated script method to obtain a high-quality training data set, and the specific steps are as follows: Use Python scripts and the open source audio processing library PyDub to batch import all raw speech datasets; Convert all non-WAV audio files to standard mono 16-bit PCM-encoded WAV format, and adjust the sampling rate of all audio files to a unified standard value; Removes silence or low volume segments by detecting audio energy levels that are continuously below a set threshold; Apply an adaptive filter to suppress background noise in the audio and retain the vocal part, and use the peak normalization method to adjust the volume of each audio clip to keep it consistent within a certain range; Automatically segment long audio files at predetermined time intervals, and add corresponding text tags to each segment by calling the automatic speech recognition ASR service to form synchronized subtitle files.

4. A speech recognition system based on artificial intelligence algorithm according to claim 3, characterized in that: The method of generating adversarial samples for high-quality training data by using a generative adversarial network method to obtain test data is specifically performed as follows: Construct a generative adversarial network framework, including a generator and a discriminator, and generate adversarial samples from high-quality training data. The expression is: ; in, Indicates adversarial examples, is an index collection containing all selected The index of the adversarial example, represents a normal sample randomly drawn from the original high-quality training data, is an index collection containing all selected The index of the original training sample, For test data.

5. A speech recognition system based on artificial intelligence algorithm according to claim 4, characterized in that: The speech recognition model is constructed based on the deep neural network architecture. The test data is input into the model, the activation value of each layer is calculated using the forward propagation algorithm, and the speech recognition result is output. The specific steps are as follows: Build a speech recognition model based on the deep neural network architecture WaveNet; Test data obtained through adversarial sample generation and screening Convert to a format acceptable to the model; The forward propagation algorithm is used to calculate the activation value of each layer, passing through the neurons of each layer in turn, and the activation value is calculated using the activation function ReLU; For each test audio clip ∈ , pass it as input to the model, and calculate the activation value of each layer through the forward propagation algorithm until the final output probability distribution is obtained, which is expressed as: ; in, Indicates Test samples The final recognition result is is the set of all possible characters, is a given input After that, the characters Probability of occurrence; Combining the results of all test samples, the expression is: ; in, It is a collection containing the recognition results of all test samples.

6. A speech recognition system based on artificial intelligence algorithm according to claim 5, characterized in that: The comprehensive threshold is set, compared with the speech recognition result, and after the comparison, it is determined whether it exceeds the normal range to obtain the comparison result. The specific steps are as follows: Set comprehensive thresholds based on historical data , the expression is: ; in, , , is the weight coefficient of each performance index; By using logging tools, we monitor the processing time and resource consumption of each request and obtain performance statistics. The output of the speech recognition system Compare with the actual label or user feedback to calculate the comprehensive score of each sample , the expression is: ; The comprehensive score corresponding to the current speech recognition result With the pre-set comprehensive threshold Compare and determine whether it exceeds the normal range.

7. A speech recognition system based on artificial intelligence algorithm according to claim 6, characterized in that: Based on the comparison results, the strategy is set for adjustment and a performance report is generated. The specific steps are as follows: Based on the comparison results; when Greater than When an error occurs, an alarm notification is triggered immediately, and an alarm message is sent to the administrator, indicating that there is a potential problem that needs attention and resolution. The alarm can be sent via email, SMS, or internal notification messages of the system. All abnormal situations and their occurrence time, the specific test samples involved, and related performance indicators are recorded in detail. when Less than or equal to If the current comprehensive score If it is within the normal range, the existing system configuration is maintained unchanged and the real-time performance of the system continues to be monitored; Regularly summarize performance monitoring data, including word error rate, sentence error rate, response time, and system resource usage.

8. The speech recognition system based on artificial intelligence algorithm according to claim 5, characterized in that: The forward propagation algorithm is used to calculate the activation value of each layer, and the activation value is calculated using the activation function ReLU through the neurons of each layer in turn. The expression is: ; ; in It is The weight matrix of the layer, It is The bias vector of the layer, is the activation function, It is The activation value of the layer, is the character corresponding to The unnormalized prediction score of .

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the speech recognition system based on the artificial intelligence algorithm described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the speech recognition system based on the artificial intelligence algorithm described in any one of claims 1 to 8 are implemented.