English speech signal processing method and device based on artificial intelligence

By introducing a dynamic correction mechanism based on artificial intelligence in the English speech recognition system, the problem of low recognition accuracy of existing systems when dealing with different accents is solved, and the system's adaptability and recognition accuracy are improved.

CN119964558AInactive Publication Date: 2025-05-09CHANGCHUN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133665.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing English speech recognition system processes voice signals with different accents, it is difficult to accurately classify accent types, and lacks a dynamic correction mechanism, resulting in a decrease in the accuracy and reliability of the recognition results.

Method used

Using an artificial intelligence-based method, the speech signal is classified and processed through the trained first model, the speech processing strategy is determined, and processed through the trained second model. When accent classification fails, a dynamic correction mechanism is performed to make real-time corrections by asking for user feedback.

Benefits of technology

It improves the accent recognition accuracy of the English speech recognition system, enhances the system's adaptability and flexibility, and solves the problem of low accent recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964558A_ABST
    Figure CN119964558A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an English speech signal processing method and device based on artificial intelligence, and relates to the technical field of speech signal recognition technologies. The method comprises the following steps: acquiring a voice signal; performing pronunciation classification processing on the voice signal through a trained first model, and determining a voice processing strategy according to a pronunciation classification processing result; and processing the voice signal based on the voice processing strategy through a trained second model. According to the invention, the problem of low accent recognition precision is solved, and the effect of improving the English speech signal recognition precision is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of speech signal recognition, and in particular, to an English speech signal processing method and device based on artificial intelligence. Background Art

[0002] In existing speech recognition technology, the processing of English speech signals faces many challenges. As a language widely used around the world, English has many accents and dialects, and the pronunciation varies significantly in different regions, which brings great difficulties to the speech recognition system.

[0003] For example, existing speech recognition systems often have difficulty accurately classifying accent types when processing English speech signals with different accents, especially when faced with uncommon or mixed accents. Moreover, when accent classification fails, existing systems lack an effective dynamic correction mechanism and are unable to make real-time adjustments based on user feedback, resulting in a decrease in the accuracy and reliability of recognition results.

[0004] There is currently no better solution to the above problems. Summary of the invention

[0005] The embodiments of the present invention provide an English speech signal processing method and device based on artificial intelligence, so as to at least solve the problem of low accent recognition accuracy in the related art.

[0006] According to one embodiment of the present invention, there is provided an English speech signal processing method based on artificial intelligence, comprising:

[0007] Acquire a voice signal, wherein the voice signal comprises an English voice signal;

[0008] Performing pronunciation classification processing on the speech signal through the trained first model, and determining a speech processing strategy according to the pronunciation classification processing result;

[0009] The speech signal is processed based on the speech processing strategy using the trained second model.

[0010] In an exemplary embodiment, performing pronunciation classification processing on the speech signal by using the trained first model, and determining the speech processing strategy according to the pronunciation classification processing result includes:

[0011] Performing accent classification processing on the speech signal by using the first model, wherein the speech classification processing includes the accent classification processing;

[0012] According to the accent classification processing results, the corresponding speech feature recognition algorithm is matched;

[0013] The speech signal is subjected to signal recognition processing according to a matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

[0014] In an exemplary embodiment, after performing accent classification processing on the speech signal by the first model, the method includes:

[0015] If the accent classification process results in a classification failure, performing a correction inquiry process, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the voice module to perform a voice correction inquiry;

[0016] Acquire the inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback;

[0017] Correct the accent classification results according to the feedback recognition results;

[0018] According to the correction processing result, the corresponding speech feature recognition algorithm is matched.

[0019] In an exemplary embodiment, after obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result by using the third model, the method further includes:

[0020] Determining feedback recognition information according to the feedback recognition result, wherein the feedback recognition information includes a speech element string and a speech element string structure;

[0021] Performing encoding preprocessing on the speech element string and the speech element string structure to obtain element string data and element string structure data;

[0022] Constructing an element string matrix according to the element string data and the element string structure data;

[0023] A correlation calculation is performed on the element string matrix, and when the correlation calculation result does not meet the first condition, it is determined that the feedback recognition result is abnormal.

[0024] According to another embodiment of the present invention, there is provided an English speech signal processing device based on artificial intelligence, comprising:

[0025] A signal acquisition module, used to acquire a voice signal, wherein the voice signal includes an English voice signal;

[0026] A classification module, used for performing pronunciation classification processing on the speech signal through the trained first model, and determining a speech processing strategy according to the pronunciation classification processing result;

[0027] A processing module is used to process the speech signal based on the speech processing strategy through a trained second model.

[0028] In an exemplary embodiment, performing pronunciation classification processing on the speech signal by using the trained first model, and determining the speech processing strategy according to the pronunciation classification processing result includes:

[0029] Performing accent classification processing on the speech signal by using the first model, wherein the speech classification processing includes the accent classification processing;

[0030] According to the accent classification processing results, the corresponding speech feature recognition algorithm is matched;

[0031] The speech signal is subjected to signal recognition processing according to a matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

[0032] In an exemplary embodiment, the apparatus further comprises:

[0033] A correction inquiry module, configured to, after performing the accent classification process on the speech signal by the first model, perform a correction inquiry process in the case where the accent classification process result is a classification failure, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the speech module to perform a speech correction inquiry;

[0034] A feedback collection and recognition module, used to obtain an inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback;

[0035] A correction module, used to correct the accent classification processing result according to the feedback recognition result;

[0036] The matching module is used to match the corresponding speech feature recognition algorithm according to the correction processing result.

[0037] In an exemplary embodiment, the apparatus further comprises:

[0038] A feedback recognition module, configured to determine feedback recognition information according to the feedback recognition result after obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result through a third model, wherein the feedback recognition information includes a speech element string and a speech element string structure;

[0039] A preprocessing module, used for performing encoding preprocessing on the speech element string and the speech element string structure to obtain element string data and element string structure data;

[0040] A matrix construction module, used for constructing an element string matrix according to the element string data and the element string structure data;

[0041] The judgment module is used to perform correlation calculation on the element string matrix, and determine that the feedback recognition result is abnormal when the correlation calculation result does not meet the first condition.

[0042] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any one of the above method embodiments when run.

[0043] According to yet another embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0044] The present invention introduces a dynamic correction mechanism to achieve real-time correction of accent classification results, thereby improving the adaptability and flexibility of the system. Therefore, the problem of low accent recognition accuracy can be solved, thereby achieving the effect of improving English speech recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of an English speech signal processing method based on artificial intelligence according to an embodiment of the present invention;

[0046] Figure 2 4 is a structural block diagram of an English speech signal processing device based on artificial intelligence according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.

[0048] In the following, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0049] In addition, in the present application, directional terms such as "up", "down", "left" and "right" may be defined including but not limited to the orientation relative to the schematic placement of the components in the drawings. It should be understood that these directional terms may be relative concepts, which are used for relative description and clarification, and may change accordingly according to the change in the orientation of the components in the drawings.

[0050] In this application, unless otherwise specified or limited, the term "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium. In addition, the term "coupling" can be a way of achieving electrical connection for signal transmission.

[0051] As used herein, "about," "substantially," or "approximately" includes the stated value and an average value that is within an acceptable range of variation from the particular value as determined by one of ordinary skill in the art taking into account the measurements in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).

[0052] In this embodiment, an English speech signal processing method based on artificial intelligence is provided. Figure 1 is a flow chart of an English speech signal processing method based on artificial intelligence according to an embodiment of the present invention, such as Figure 1 As shown, the process includes the following steps:

[0053] Step S11, acquiring a voice signal, wherein the voice signal includes an English voice signal;

[0054] In this embodiment, English speech signals are collected by a microphone or other device, and the speech signals are converted into electronic signals for further processing.

[0055] Among them, after the voice signal is collected, it needs to be preprocessed, such as removing noise, reducing volume, echo cancellation, etc., to improve the accuracy of voice recognition.

[0056] It should be noted that, in addition to English voice signals, voice signals in other languages ​​may also be included, such as Chinese voice signals, Japanese, German, French, Spanish, Portuguese, Russian and other languages, which are not limited here.

[0057] Step S12, performing pronunciation classification processing on the speech signal through the trained first model, and determining a speech processing strategy according to the pronunciation classification processing result.

[0058] In this embodiment, different processing strategies are selected for different pronunciation classification processing results to specifically process relevant classification results, thereby improving the accuracy of speech recognition and the precision of speech signal recognition.

[0059] Specifically, it is to extract features that reflect the essence of the preprocessed speech signal, such as phonemes, syllables, phrases, etc.; commonly used feature extraction methods include Mel-frequency cepstral coefficients (MFCC), which is based on short-time Fourier transform. By setting several bandpass filters within the frequency spectrum of the speech, the signal energy of the corresponding filter group is calculated, and then the corresponding cepstral coefficients are calculated through discrete cosine transform (DCT); of course, other methods can also be used for extraction, for example, using a multi-head attention mechanism network (i.e., the first model) or other deep learning models to train speech data to identify different pronunciation features, and then input the extracted features into the trained model, and the model will output the pronunciation classification results of the speech, such as accent type, pronunciation accuracy, etc.

[0060] Among them, the speech processing strategy includes algorithm selection, processing module, speech content to be processed, processing target, etc.

[0061] It should be noted that speech with different accents will cause great difficulties in speech recognition, especially English, so it is necessary to additionally output the accent type and perform subsequent processing based on the accent classification.

[0062] Step S13: Process the speech signal based on the speech processing strategy using the trained second model.

[0063] In this embodiment, after the speech processing strategy is obtained, recognition, simulation, repair and other processing are performed according to the speech processing strategy to meet different needs.

[0064] For example, if there is an accent problem in the speech signal, an accent correction algorithm can be used to correct it, for example, by adjusting the speech features through a generative adversarial network (GAN) or a variational autoencoder (VAE); or some deep learning models and methods can be used to deal with the accent problem, for example, by using accent embedding to achieve more flexible feature extraction, and by enhancing the model's recognition ability for different accents through adversarial learning and contrastive learning. It is also possible to refer to the deep speaker recognition framework, use a convolutional recurrent neural network as the front-end encoder, and introduce a discriminant loss function to enhance the discriminant ability of accent features; or recognize the accent through a deep speaker recognition framework, using a CRNN as the front-end encoder, which consists of a ResNet and a bidirectional GRU network to extract frame-level descriptors, and then use a bidirectional GRU to integrate the calculated local descriptors into global sentence-level features. In order to solve the overfitting problem, a speech recognition auxiliary task based on connection temporal classification (CTC) is added during training, and some powerful discriminant loss functions in facial recognition work are introduced to enhance the discriminant ability of accent features.

[0065] Through the above steps, by introducing a dynamic correction mechanism, real-time correction of accent classification results is achieved, the adaptability and flexibility of the system are improved, the problem of low accent recognition accuracy is solved, and the recognition accuracy of English speech signals is improved.

[0066] The execution entities of the above steps may be [base stations, terminals], etc., but are not limited thereto.

[0067] In an optional embodiment, performing pronunciation classification processing on the speech signal by using the trained first model, and determining the speech processing strategy according to the pronunciation classification processing result includes:

[0068] Step S121, performing accent classification processing on the speech signal through the first model, wherein the speech classification processing includes the accent classification processing;

[0069] Step S122, matching the corresponding speech feature recognition algorithm according to the accent classification processing result;

[0070] Step S123, performing signal recognition processing on the speech signal according to the matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

[0071] In this embodiment, different speakers have different vocal organs, accents, and speaking styles. The same speaker may have different pronunciations at different times and in different states. These differences make it difficult for the speech recognition system to accurately recognize the speech of everyone, especially since there are many accents and dialects in English, and accents in different regions may have significant differences in the pronunciation of vowels and consonants. In addition, the same concept may be expressed differently in different languages, and the grammatical structures of different languages ​​are also different, which brings difficulties to cross-language and cross-cultural recognition; therefore, when it comes to speech recognition, especially English speech signal recognition, special processing is required based on the accent.

[0072] Specifically, for different accents, the type of accent can be analyzed and identified by adjusting the feature extraction parameters or using a specially optimized model or directly mapping from the audio signal to the text sequence, and then the corresponding speech recognition library can be selected according to the accent type, thereby reducing the recognition difficulty and improving the recognition accuracy.

[0073] Among them, the result of accent classification processing is mainly to output accent classification labels or categories, and then select a specific recognition algorithm based on the labels and different characteristics of speech signals with different accents (such as pitch, timbre, speaking speed, etc.); then, a specific speech feature recognition algorithm is used to extract the key features of the speech signal (such as tone, phonemes, rhythm, etc.), and output the recognition results of the speech signal based on the key features (for example, converting speech into text, or recognizing instructions in speech, etc.).

[0074] In an optional embodiment, after performing accent classification processing on the speech signal by using the first model, the method includes:

[0075] Step S124, when the accent classification processing result is classification failure, performing correction inquiry processing, wherein the correction inquiry processing includes generating and sending an inquiry voice instruction to instruct the voice module to perform a voice correction inquiry;

[0076] Step S125, obtaining an inquiry feedback result, and performing feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user regarding the voice correction inquiry feedback;

[0077] Step S126, correcting the accent classification result according to the feedback recognition result;

[0078] Step S127, matching the corresponding speech feature recognition algorithm according to the correction processing result.

[0079] In this embodiment, when normal speech recognition fails, real-time correction is performed through simple inquiry to ensure the recognition result.

[0080] The voice inquiry command may be a command to ask some simple questions that can be used for accent recognition, such as "Where are you from", "Please briefly describe your life experience", "Where have you lived before", etc., or the user may be asked to repeat a specific word or sentence, and then the feedback voice signal is subjected to feature extraction (such as the aforementioned MFCC), and the feedback voice signal is recognized using a third model (such as a deep learning model), which may be a model specially trained for accent correction, and can more accurately recognize and classify the user's voice; based on the feedback recognition result of the third model, the initial accent classification result is corrected, for example, if the feedback voice signal shows that the user's actual accent is inconsistent with the initial classification, the classification result is updated to identify the possible accent area; and then the most suitable voice feature recognition algorithm is selected based on the corrected accent classification result. For example, for certain specific accents, feature extraction parameters may be adjusted or a more optimized model architecture may be selected (such as a generative adversarial network GAN or a variational autoencoder VAE).

[0081] In an optional embodiment, after obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result by using the third model, the method further includes:

[0082] Step S1251, determining feedback recognition information according to the feedback recognition result, wherein the feedback recognition information includes a speech element string and a speech element string structure;

[0083] Step S1252, performing encoding preprocessing on the voice element string and the voice element string structure to obtain element string data and element string structure data;

[0084] Step S1253, constructing an element string matrix according to the element string data and the element string structure data;

[0085] Step S1254, performing correlation calculation on the element string matrix, and determining that the feedback recognition result is abnormal when the correlation calculation result does not meet the first condition.

[0086] In this embodiment, the speech element string refers to the serialized representation of the recognized speech content, for example, the character sequence after the speech is converted into text, and the speech element string structure refers to the organization method of the speech elements, for example, the semantic structure, grammatical structure or rhythmic structure of the speech, etc.; generally, the element string and the element string structure will have different changes in the later stage of recognition of different accents (such as the absence of vowels and consonants, etc.), so at this time, whether the recognition feedback result is correct can be judged by identifying whether the speech element string and the speech element string structure are normal.

[0087] Specifically, the speech element string and the speech element string structure are encoded and converted into a format suitable for subsequent processing. The encoding process can be to convert a text sequence into a digital sequence (for example, using character encoding or word embedding), or to convert structural information into a vector or matrix form (for example, using structured data representation); the element string matrix can be a two-dimensional array, in which each row or column represents a speech element, and the values ​​in the matrix represent the relationship between the elements (for example, similarity, correlation, etc.), wherein the diagonal of the matrix can represent the characteristics of the element itself, and the off-diagonal elements represent the relationship between the elements; then the similarity or correlation coefficient between the elements in the matrix is ​​calculated (the correlation coefficient can be calculated based on the Pearson coefficient), and it is checked whether the structure of the matrix conforms to a preset pattern or rule. If the correlation coefficient or similarity value is within a preset range, the data is determined to be normal, otherwise it is judged to be abnormal, and so on.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0089] In the present embodiment, an English speech signal processing device based on artificial intelligence is also provided, and the device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made are not repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceived.

[0090] Figure 2 is a structural block diagram of an English speech signal processing device based on artificial intelligence according to an embodiment of the present invention, such as Figure 2 As shown, the device comprises:

[0091] A signal acquisition module 21, used to acquire a voice signal, wherein the voice signal includes an English voice signal;

[0092] A classification module 22, configured to perform pronunciation classification processing on the speech signal using the trained first model, and determine a speech processing strategy according to the pronunciation classification processing result;

[0093] The processing module 23 is used to process the speech signal based on the speech processing strategy through the trained second model.

[0094] In an optional embodiment, performing pronunciation classification processing on the speech signal by using the trained first model, and determining the speech processing strategy according to the pronunciation classification processing result includes:

[0095] Performing accent classification processing on the speech signal by using the first model, wherein the speech classification processing includes the accent classification processing;

[0096] According to the accent classification processing results, the corresponding speech feature recognition algorithm is matched;

[0097] The speech signal is subjected to signal recognition processing according to a matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

[0098] In an optional embodiment, the device further comprises:

[0099] A correction inquiry module, configured to, after performing the accent classification process on the speech signal by the first model, perform a correction inquiry process in the case where the accent classification process result is a classification failure, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the speech module to perform a speech correction inquiry;

[0100] A feedback collection and recognition module, used to obtain an inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback;

[0101] A correction module, used to correct the accent classification processing result according to the feedback recognition result;

[0102] The matching module is used to match the corresponding speech feature recognition algorithm according to the correction processing result.

[0103] In an optional embodiment, the device further comprises:

[0104] A correction inquiry module, configured to, after performing the accent classification process on the speech signal by the first model, perform a correction inquiry process in the case where the accent classification process result is a classification failure, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the speech module to perform a speech correction inquiry;

[0105] A feedback collection and recognition module, used to obtain an inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback;

[0106] A correction module, used to correct the accent classification processing result according to the feedback recognition result;

[0107] The matching module is used to match the corresponding speech feature recognition algorithm according to the correction processing result.

[0108] In an optional embodiment, the device further comprises:

[0109] A feedback recognition module, configured to determine feedback recognition information according to the feedback recognition result after obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result through a third model, wherein the feedback recognition information includes a speech element string and a speech element string structure;

[0110] A preprocessing module, used for performing encoding preprocessing on the speech element string and the speech element string structure to obtain element string data and element string structure data;

[0111] A matrix construction module, used for constructing an element string matrix according to the element string data and the element string structure data;

[0112] The judgment module is used to perform correlation calculation on the element string matrix, and determine that the feedback recognition result is abnormal when the correlation calculation result does not meet the first condition.

[0113] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0114] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0115] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0116] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0117] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0118] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0119] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0120] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0121] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0122] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0123] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An English speech signal processing method based on artificial intelligence, characterized in that: include: Acquire a voice signal, wherein the voice signal comprises an English voice signal; Performing pronunciation classification processing on the speech signal through the trained first model, and determining a speech processing strategy according to the pronunciation classification processing result; The speech signal is processed based on the speech processing strategy using the trained second model.

2. The system according to claim 1, characterized in that The method of performing pronunciation classification processing on the speech signal by using the trained first model and determining the speech processing strategy according to the pronunciation classification processing result comprises: Performing accent classification processing on the speech signal by using the first model, wherein the speech classification processing includes the accent classification processing; According to the accent classification processing results, the corresponding speech feature recognition algorithm is matched; The speech signal is subjected to signal recognition processing according to a matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

3. The system according to claim 2, characterized in that After performing accent classification processing on the speech signal by using the first model, the method includes: If the accent classification process results in a classification failure, performing a correction inquiry process, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the voice module to perform a voice correction inquiry; Acquire the inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback; Correct the accent classification results according to the feedback recognition results; According to the correction processing result, the corresponding speech feature recognition algorithm is matched.

4. The system according to claim 3, characterized in that After obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result through the third model, the method further includes: Determining feedback recognition information according to the feedback recognition result, wherein the feedback recognition information includes a speech element string and a speech element string structure; Performing encoding preprocessing on the speech element string and the speech element string structure to obtain element string data and element string structure data; Constructing an element string matrix according to the element string data and the element string structure data; A correlation calculation is performed on the element string matrix, and when the correlation calculation result does not meet the first condition, it is determined that the feedback recognition result is abnormal.

5. An English speech signal processing device based on artificial intelligence, characterized in that: include: A signal acquisition module, used to acquire a voice signal, wherein the voice signal includes an English voice signal; A classification module, used for performing pronunciation classification processing on the speech signal through the trained first model, and determining a speech processing strategy according to the pronunciation classification processing result; A processing module is used to process the speech signal based on the speech processing strategy through a trained second model.

6. The device according to claim 5, characterized in that The method of performing pronunciation classification processing on the speech signal by using the trained first model and determining the speech processing strategy according to the pronunciation classification processing result comprises: Performing accent classification processing on the speech signal by using the first model, wherein the speech classification processing includes the accent classification processing; According to the accent classification processing results, the corresponding speech feature recognition algorithm is matched; The speech signal is subjected to signal recognition processing according to a matched speech feature recognition algorithm, wherein the second model includes the speech feature recognition algorithm.

7. The device according to claim 6, characterized in that The device also includes: A correction inquiry module, configured to, after performing the accent classification process on the speech signal by the first model, perform a correction inquiry process in the case where the accent classification process result is a classification failure, wherein the correction inquiry process includes generating and sending an inquiry voice instruction to instruct the speech module to perform a speech correction inquiry; A feedback collection and recognition module, used to obtain an inquiry feedback result, and perform feedback recognition on the inquiry feedback result through a third model, wherein the inquiry feedback result includes a feedback voice signal of the user for the voice correction inquiry feedback; A correction module, used to correct the accent classification processing result according to the feedback recognition result; The matching module is used to match the corresponding speech feature recognition algorithm according to the correction processing result.

8. The device according to claim 7, characterized in that The device also includes: A feedback recognition module, configured to determine feedback recognition information according to the feedback recognition result after obtaining the inquiry feedback result and performing feedback recognition on the inquiry feedback result through a third model, wherein the feedback recognition information includes a speech element string and a speech element string structure; A preprocessing module, used for performing encoding preprocessing on the speech element string and the speech element string structure to obtain element string data and element string structure data; A matrix construction module, used for constructing an element string matrix according to the element string data and the element string structure data; The judgment module is used to perform correlation calculation on the element string matrix, and determine that the feedback recognition result is abnormal when the correlation calculation result does not meet the first condition.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.