Speech recognition method and system, electronic equipment and storage medium

Through the multidimensional speech recognition engine performance evaluation and dynamic switching of weight data, the problem of insufficient accuracy and adaptability of speech recognition in complex environments in the prior art is solved, and efficient and accurate speech recognition and improved voice interaction experience are achieved.

CN120126460APending Publication Date: 2025-06-10E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294578.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing speech recognition technology has low recognition accuracy and adaptability in complex environments such as high noise background, different accents and dialects, far-field voice, resulting in poor voice interaction experience.

Method used

By obtaining the audio data to be recognized, inputting the preset evaluation decision module to perform multidimensional speech recognition engine performance evaluation, determining the desired speech recognition engine, and dynamically determining the switching weight data through the preset switching mechanism, performing engine mixed output to improve the accuracy and adaptability of speech recognition.

Benefits of technology

It effectively improves the accuracy and adaptability of speech recognition, improves the voice interaction experience, and realizes efficient speech recognition in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126460A_ABST
    Figure CN120126460A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition method and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring to-be-recognized audio data; inputting the to-be-recognized audio data into a preset evaluation decision module for multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data of a plurality of dimensions; determining an expected speech recognition engine through a preset decision network according to the performance evaluation data of the plurality of dimensions; switching weight data are dynamically determined through a preset switching mechanism, engine mixed output is carried out according to the switching weight data, and a voice recognition result is obtained; wherein the switching weight data comprises a mixed output proportion of a historical speech recognition engine and the expected speech recognition engine. According to the embodiment of the invention, the accuracy and adaptability of voice recognition can be effectively improved, and the voice interaction experience is effectively improved. The method can be widely applied to the technical field of voice processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technologies, and in particular, to a speech recognition method, system, electronic device, and storage medium. Background Art

[0002] As one of the key technologies in the field of artificial intelligence, automatic speech recognition (ASR) technology has made remarkable progress in recent years. In related technologies, when the automatic speech recognition technology is actually applied, it is often necessary to accurately recognize speech in various complex environments, such as high-noise backgrounds, different accents and dialects, far-field speech, etc. However, the accuracy and adaptability of speech recognition are relatively low, and the speech interaction experience is poor.

[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a speech recognition method, system, electronic device, and storage medium, which can effectively improve the accuracy and adaptability of speech recognition and effectively improve the speech interaction experience.

[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a speech recognition method, and the method includes the following steps:

[0006] Obtain audio data to be recognized;

[0007] Input the audio data to be recognized into a preset evaluation and decision-making module for multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data in several dimensions;

[0008] Determine an expected speech recognition engine through a preset decision network according to the performance evaluation data in several dimensions;

[0009] Dynamically determine switching weight data through a preset switching mechanism, and perform engine hybrid output according to the switching weight data to obtain a speech recognition result; wherein, the switching weight data includes the hybrid output ratio of the historical speech recognition engine and the expected speech recognition engine.

[0010] In some embodiments, the performance evaluation data includes audio length evaluation data, behavior feature evaluation data, speech structure evaluation data, and population feature evaluation data;

[0011] Among them, the step of inputting the audio data to be recognized into a preset evaluation and decision-making module for multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data in several dimensions includes:

[0012] Performing length anomaly analysis on the audio data to be recognized through a preset audio length threshold to obtain the audio length evaluation data; wherein, the preset audio length threshold is determined by the average length and standard deviation of a preset audio data set;

[0013] Performing user behavior analysis on the audio data to be recognized through a preset user behavior model to obtain the behavior feature evaluation data;

[0014] Performing speech structure feature analysis on the audio data to be recognized through a first long short-term memory network model to obtain the speech structure evaluation data;

[0015] Performing group feature analysis on the audio data to be recognized through a second long short-term memory network model to obtain the group feature evaluation data.

[0016] In some embodiments, before performing the step of performing user behavior analysis on the audio data to be recognized through a preset user behavior model to obtain the behavior feature evaluation data, the method further includes:

[0017] Constructing a speech behavior data set; wherein, the speech behavior data set includes user speech data in different scenarios;

[0018] Performing feature extraction on the speech behavior data set to obtain first speech feature data; wherein, the first speech feature data includes first acoustic feature data, first language feature data, and statistical feature data;

[0019] Superimposing the first acoustic feature data, the first language feature data, and the statistical feature data to obtain first superimposed feature data;

[0020] Training a preset transformer model through the first superimposed feature data to construct the preset user behavior model.

[0021] In some embodiments, before performing the step of performing speech structure feature analysis on the audio data to be recognized through a first long short-term memory network model to obtain the speech structure evaluation data, the method further includes:

[0022] Constructing a speech structure data set; wherein, the speech structure data set includes user speech data containing different keywords and sentiment information;

[0023] Performing feature extraction on the speech structure data set to obtain second speech feature data; wherein, the second speech feature data includes second acoustic feature data and second language feature data;

[0024] Superimpose the second acoustic feature data and the second language feature data to obtain second superimposed feature data;

[0025] Input the second superimposed feature data into a preset long short-term memory network model for model training to construct the first long short-term memory network model.

[0026] In some embodiments, before performing the analysis of the group features of the audio data to be recognized through the second long short-term memory network model to obtain the group feature evaluation data, the method further includes:

[0027] Construct a group feature data set; wherein, the group feature data includes user voice data of different user groups;

[0028] Extract features from the group feature data set to obtain third voice feature data; wherein, the third voice feature data includes third acoustic feature data and third language feature data;

[0029] Superimpose the third acoustic feature data and the third language feature data to obtain third superimposed feature data;

[0030] Input the third superimposed feature data into a preset long short-term memory network model for model training to construct the second long short-term memory network model.

[0031] In some embodiments, the method further includes:

[0032] Obtain a system log file; wherein, the system log file includes user instructions and corresponding feedback information;

[0033] Search and analyze the system log file to obtain abnormal instruction data; wherein, the abnormal instruction data includes the user instructions with parsing problems and the corresponding feedback information;

[0034] Perform correction processing according to the abnormal instruction data and update the system knowledge system according to the correction result; wherein, the correction processing includes re-parsing instruction operations, correcting error operations, and parameter adjustment operations.

[0035] In some embodiments, the dynamically determining the switching weight data through a preset switching mechanism to perform engine hybrid output according to the switching weight data to obtain a speech recognition result includes:

[0036] Dynamically calculate the switching weight data through a preset linear interpolation algorithm;

[0037] Perform speech recognition on the audio data to be recognized through the historical speech recognition engine to obtain first speech recognition data;

[0038] Performing speech recognition on the audio data to be recognized through the expected speech recognition engine to obtain second speech recognition data;

[0039] Performing weighted average processing on the first speech recognition data and the second speech recognition data according to the switching weight data to obtain the speech recognition result.

[0040] To achieve the above object, on the other hand, an embodiment of the present application proposes a speech recognition system, which includes:

[0041] A first module for obtaining audio data to be recognized;

[0042] A second module for inputting the audio data to be recognized into a preset evaluation and decision-making module to perform multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data in several dimensions;

[0043] A third module for determining an expected speech recognition engine according to the performance evaluation data in several dimensions through a preset decision network;

[0044] A fourth module for dynamically determining switching weight data through a preset switching mechanism to perform engine hybrid output according to the switching weight data to obtain a speech recognition result; wherein, the switching weight data includes the hybrid output ratio of the historical speech recognition engine and the expected speech recognition engine.

[0045] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, which includes:

[0046] At least one processor;

[0047] At least one memory for storing at least one program;

[0048] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0049] To achieve the above object, on the other hand, an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0050] The embodiments of the present application at least include the following beneficial effects: The present application provides a voice recognition method, system, electronic device, and storage medium. This solution obtains the audio data to be recognized and inputs the audio data to be recognized into a preset evaluation and decision-making model for multi-dimensional performance evaluation of the voice recognition engine, obtaining performance evaluation data in several dimensions. Then, according to the performance evaluation data in several dimensions, the expected voice recognition engine is determined through a preset decision-making network. Then, the switching weight data is dynamically determined through a preset switching mechanism to perform engine hybrid data according to the switching weight data to obtain the voice recognition result, which can accurately implement voice recognition and effectively improve the voice interaction experience. Among them, the switching weight data includes the mixing output ratio of the historical voice recognition engine and the expected voice recognition engine. It is easy to understand that the embodiments of the present invention perform multi-dimensional performance evaluation of the voice recognition engine through a preset evaluation and decision-making module to determine the expected voice recognition engine that is more suitable for the audio data to be recognized, and perform hybrid data of the historical voice recognition engine and the expected voice recognition engine through the switching weight data dynamically determined by the preset switching mechanism, so as to realize the switching of the dynamic voice recognition engine, which can effectively improve the voice interaction experience while improving the accuracy and adaptability of voice recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of the steps of the voice recognition method provided by the embodiments of the present invention;

[0052] Figure 2 is a flowchart of the steps of inputting the audio data to be recognized into a preset evaluation and decision-making module for multi-dimensional performance evaluation of the voice recognition engine to obtain performance evaluation data in several dimensions;

[0053] Figure 3 is an internal process schematic diagram of the preset evaluation and decision-making module provided by the embodiments of the present invention;

[0054] Figure 4 is a flowchart of the construction steps of the preset user behavior model provided by the embodiments of the present invention;

[0055] Figure 5 is a flowchart of the construction steps of the first long short-term memory network model provided by the embodiments of the present invention;

[0056] Figure 6 is a flowchart of the construction steps of the second long short-term memory network model provided by the embodiments of the present invention;

[0057] Figure 7 is a schematic diagram of the working steps of the user feedback module provided by the embodiments of the present invention;

[0058] Figure 8It is a flowchart of steps provided by an embodiment of the present invention for dynamically determining switching weight data through a preset switching mechanism, and performing engine hybrid output according to the switching weight data to obtain a speech recognition result;

[0059] Figure 9 It is a schematic diagram of the overall process of the speech recognition method provided by an embodiment of the present invention;

[0060] Figure 10 It is a schematic structural diagram of the speech recognition system provided by an embodiment of the present invention;

[0061] Figure 11 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0062] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0063] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "when...", or "in response to determining".

[0064] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more, a plurality includes two or more, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0066] Before elaborating on the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first explained, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0067] Automatic Speech Recognition (ASR): It is a technology that converts human language signals into text information that can be processed by a computer, also known as speech-to-text conversion. Through specific algorithms and models, it can automatically analyze and process human speech signals and accurately convert them into digital text.

[0068] Long Short-Term Memory (LSTM) network model: It refers to the long short-term memory network, which is a special type of recurrent neural network (RNN). By introducing a special gating mechanism, it can alleviate the problem of gradient vanishing or gradient explosion that occurs in traditional RNNs when processing long sequence data.

[0069] As one of the key technologies in the field of artificial intelligence, Automatic Speech Recognition (ASR) technology has made remarkable progress in recent years. In related technologies, when the ASR technology is actually applied, it often needs to accurately recognize speech in various complex environments, such as high-noise backgrounds, different accents and dialects, and far-field speech. However, the accuracy and adaptability of speech recognition are relatively low, and the speech interaction experience is poor.

[0070] Exemplarily, for example, traditional ASR systems are often optimized for specific scenarios or languages and lack flexibility. In addition, when switching between multiple ASR engines, problems such as recognition interruptions may occur during the switching between different ASR engines, making it difficult to provide a smooth speech interaction experience. At the same time, in specific fields such as medical and legal, ASR systems need to accurately recognize professional terms and proper nouns, which puts higher requirements on the accuracy and professionalism of the model. At the same time, as the business scale grows, new knowledge emerges, and currently, ASR systems often have difficulty accurately recognizing these emerging knowledge and patterns and have a low ability to recognize professional terms.

[0071] In view of this, embodiments of the present application provide a voice recognition method, system, electronic device, and storage medium. This solution obtains the audio data to be recognized, inputs the audio data to be recognized into a preset evaluation decision model for multi-dimensional performance evaluation of the voice recognition engine, obtains performance evaluation data in several dimensions, and then determines the expected voice recognition engine through a preset decision network according to the performance evaluation data in several dimensions. Then, a preset switching mechanism is used to dynamically determine the switching weight data, so as to perform engine hybrid data according to the switching weight data to obtain a voice recognition result (including the hybrid output ratio of the historical voice recognition engine and the expected voice recognition engine), which can accurately implement voice recognition, effectively improve the accuracy and adaptability of voice recognition, and effectively improve the voice interaction experience.

[0072] The voice recognition method provided by the embodiments of the present application relates to the technical field of voice processing. The voice recognition method provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side may be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application that implements the voice recognition method, etc., but is not limited to the above forms.

[0073] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0074] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0075] Figure 1 is an optional flowchart of the speech recognition method provided by the embodiments of the present application. Figure 1 The method in may include but is not limited to steps S110 to S140.

[0076] Step S110: Obtain the audio data to be recognized.

[0077] Step S120: Input the audio data to be recognized into a preset evaluation decision module for multi-dimensional performance evaluation of the speech recognition engine, and obtain performance evaluation data in several dimensions.

[0078] Step S130: Determine the expected speech recognition engine through a preset decision network according to the performance evaluation data in several dimensions.

[0079] Step S140: Dynamically determine the switching weight data through a preset switching mechanism, and perform engine hybrid output according to the switching weight data to obtain the speech recognition result. Among them, the switching weight data includes the hybrid output ratio of the historical speech recognition engine and the expected speech recognition engine.

[0080] In the process of the present specific embodiment, the embodiment of the present invention first obtains the audio data to be recognized, and inputs the audio data to be recognized into a preset evaluation decision module for multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data in several dimensions. Specifically, the audio data to be recognized in the embodiment of the present invention refers to the audio data that needs to be subjected to speech recognition. Correspondingly, the preset evaluation decision module in the embodiment of the present invention is used to dynamically evaluate the performance of each speech recognition engine, so as to determine whether it is necessary to switch the speech recognition engine subsequently. At the same time, in the embodiment of the present invention, the multi-dimensional speech recognition performance evaluation of each speech recognition engine is carried out through the preset evaluation decision module to obtain corresponding performance evaluation data in several dimensions, such as key indicators including accuracy rate, response time, etc. Further, the embodiment of the present invention determines the expected speech recognition engine through the preset decision network according to the performance evaluation data in several dimensions. Specifically, the expected speech recognition engine in the embodiment of the present invention refers to the speech recognition engine that is determined to be more suitable for recognizing the audio data to be recognized through analysis of the performance evaluation data evaluated by the prediction evaluation decision module, thereby effectively improving the accuracy and adaptability of speech recognition. Among them, after the performance evaluation data in different dimensions is obtained, the embodiment of the present invention analyzes these data through the preset decision network to determine whether it is necessary to switch the ASR engine and the target ASR engine to be switched, that is, the expected speech recognition engine. Finally, the embodiment of the present invention dynamically determines the switching weight data through the preset switching mechanism, and performs engine hybrid output according to the switching weight data to obtain the speech recognition result. Specifically, the preset switching mechanism in the embodiment of the present invention refers to the mechanism of switching the speech recognition engine from the currently used speech recognition engine (historical speech recognition engine) to the expected speech recognition engine. Among them, the switching weight data in the embodiment of the present invention includes the hybrid output ratio of the historical speech recognition engine and the expected speech recognition engine, that is, in the speech recognition result, the weights corresponding to the recognition results of the historical speech recognition engine and the expected speech recognition result. Correspondingly, the preset switching mechanism in the embodiment of the present invention performs smooth switching through the switching weight data, and calculates the switching weight data dynamically to perform speech recognition hybrid output, so that it can seamlessly switch to the new ASR engine without interrupting the user interaction, effectively improving the continuity of speech recognition and the speech interaction experience.

[0081] In some embodiments of the present invention, the performance evaluation data includes audio length evaluation data, behavior feature evaluation data, speech structure evaluation data, and group feature evaluation data. Correspondingly, referring to Figure 2 , in the embodiment of the present invention, inputting the audio data to be recognized into the preset evaluation decision module for multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data in several dimensions includes, but is not limited to, the following steps:

[0082] Step S210: Perform length anomaly analysis on the audio data to be recognized through a preset audio length threshold to obtain audio length evaluation data. The preset audio length threshold is determined by the average length and standard deviation of a preset audio data set.

[0083] Step S220: Perform user behavior analysis on the audio data to be recognized through a preset user behavior model to obtain behavior feature evaluation data.

[0084] Step S230: Perform speech structure feature analysis on the audio data to be recognized through a first long short-term memory network model to obtain speech structure evaluation data.

[0085] Step S240: Perform group feature analysis on the audio data to be recognized through a second long short-term memory network model to obtain group feature evaluation data.

[0086] In this specific embodiment, the performance evaluation data evaluated by the preset evaluation decision model of the embodiment of the present invention includes audio length evaluation data, behavior feature evaluation data, speech structure evaluation data, and group feature evaluation data. Among them, the audio length evaluation data in the embodiment of the present invention refers to the audio length anomaly evaluation data obtained by evaluating whether the audio length is abnormal. Specifically, the embodiment of the present invention performs length anomaly analysis on the audio data to be recognized through a preset audio length threshold to obtain audio length evaluation data. The embodiment of the present invention detects abnormal situations by calculating the average length of speech audio. For example, if the length of an audio file is abnormally long while the recognition result is abnormally short, this may indicate that there is a problem with the engine when processing long sentences or complex sentences. Among them, the preset audio length threshold in the embodiment of the present invention is determined by the average length and standard deviation of a preset audio data set. Exemplarily, the embodiment of the present invention first obtains a certain number of normal audio files to construct an audio data set, and these files cover various normal usage scenarios. Correspondingly, based on the constructed audio data, the average length μ and standard deviation σ of all audio files are calculated. When the embodiment of the present invention obtains n audio files, the length of each file is L 1 ,L 2 ,...,L n , the calculation formulas for the average length μ and standard deviation σ are shown in the following formula (1):

[0087]

[0088] Correspondingly, the embodiment of the present invention sets one or more anomaly detection thresholds, that is, the preset audio length threshold, according to the calculated average length μ and standard deviation σ. For example, the embodiment of the present invention determines the preset audio height threshold by using the method of adding and subtracting several times the standard deviation from the mean value, as shown in the following formula (2):

[0089]

[0090] Wherein, k is an adjustable parameter in the formula, and usually 2 or 3 can be taken. Correspondingly, when the length L of an audio file satisfies any one of the following conditions in formula (3), it is determined to be abnormal:

[0091]

[0092] Correspondingly, for the newly received audio file (audio data to be recognized), the calculator in the embodiment of the present invention calculates the length and compares it with the preset audio length to determine whether it belongs to an abnormal situation, and obtains the audio length evaluation data.

[0093] In addition, the behavior feature evaluation data in the embodiment of the present invention refers to the evaluation data of the user's personal speech habits and features. Specifically, the embodiment of the present invention analyzes the user behavior of the audio data to be recognized through a preset user behavior model to obtain the behavior feature evaluation data. Among them, the preset user behavior model in the present invention example refers to a pre-trained Transformer model. Correspondingly, the embodiment of the present invention analyzes the personal speech habits and features through the preset user behavior model to determine whether the user behaves abnormally when using the corresponding ASR. For example, when a user usually speaks very fast, but the audio length is extremely long in a certain recognition, the system may consider this as an abnormal situation and consider switching to an engine more suitable for fast speech speed. At the same time, the speech structure evaluation data in the embodiment of the present invention refers to the evaluation data obtained by analyzing the structural features of the speech. Specifically, the embodiment of the present invention analyzes the speech structure features of the audio data to be recognized through a first long short-term memory network model to obtain the speech structure evaluation data. Among them, the first long short-term memory network model in the embodiment of the present invention refers to a pre-trained speech structure feature analysis model. Correspondingly, the embodiment of the present invention analyzes the peaks, valleys or other sound-related features of the audio data to be recognized through the first long short-term memory network model, so as to identify different types of keywords, such as modal particles, subject-predicate-object structures, adverbs, etc. Among them, the embodiment of the present invention analyzes the speech structure features of the audio data to be recognized, so as to help the system understand the semantics and emotions of the speech, obtain the corresponding speech structure evaluation data, and thus determine whether the current engine is suitable for the current scenario.

[0094] In addition, in the embodiments of the present invention, the group feature evaluation data refers to the evaluation data obtained by analyzing the group features of audio data to determine the group to which the audio belongs (such as different groups like the elderly, children, adults, etc.). Specifically, in the embodiments of the present invention, the second long short-term memory network model is used to perform group feature analysis on the audio data to be recognized, and the group feature evaluation data is obtained. Among them, the second long short-term memory network model in the embodiments of the present invention is a pre-trained group speech feature recognition model. Correspondingly, in the embodiments of the present invention, it is determined by judging from the average speech rates of different groups to meet the speech recognition requirements of different groups such as the elderly, children, adults, etc. At the same time, the system in the embodiments of the present invention will adjust the recognition parameters of the engine according to the speech rate characteristics of the user to improve the recognition accuracy. For example, the speech rates of children and the elderly may be slower, while that of young people may be faster, and the system will optimize according to these characteristics. As Figure 3 shown, in the embodiments of the present invention, the preset evaluation decision module performs multi-dimensional comprehensive evaluation based on audio length abnormality, individual speech habits and characteristics, speech structure characteristics, and group speech characteristics, so as to determine whether it is necessary to switch the speech recognition engine.

[0095] It should be noted that in some embodiments of the present invention, after obtaining the audio length evaluation data, behavior feature evaluation data, speech structure evaluation data, and group feature evaluation data through evaluation, the embodiments of the present invention construct a preset decision network and use an algorithm (such as the maximum expected utility criterion) to calculate the optimal decision path to determine whether to switch the corresponding engine. Exemplarily, the embodiments of the present invention first define the nodes of the preset decision network. For example, in the embodiments of the present invention, D1, D2, D3, and D4 respectively represent the detection results corresponding to the audio length evaluation data, behavior feature evaluation data, speech structure evaluation data, and group feature evaluation data. Among them, each detection result can be "normal" or "abnormal". At the same time, in the embodiments of the present invention, E represents the current engine state, which can be "running well" or "having problems"; A represents the action taken, that is, whether to switch the engine, and the options are "switch" or "do not switch"; U represents the final utility / cost, such as the cost of operation, safety, etc. Then, the embodiments of the present invention define the edges. Each detection node (D1, D2, D3, D4) directly points to the action node (A), indicating that the detection result will directly affect the decision. In addition, the engine state node (E) may also affect the decision-making process, so it can be connected to the action node (A). In addition, the action node (A) is connected to the utility node (U) because corresponding utility or cost will be brought after performing a specific action. Finally, the embodiments of the present invention create a graphical model through the above nodes and edges, thereby constructing a preset decision network. Among them, for each pair of associated nodes, the embodiments of the present invention define the conditional probability distribution. For example, for D1->A, it is necessary to specify the probability of selecting each action when D1 is in different states; for A->U, it is necessary to define the expected utility value after taking different actions.

[0096] Referring to Figure 4 , in some embodiments of the present invention, before performing user behavior analysis on the audio data to be recognized through a preset user behavior model to obtain behavior feature evaluation data, the speech recognition method provided by the embodiments of the present invention further includes but is not limited to the following steps:

[0097] Step S310: Construct a speech behavior dataset. The speech behavior dataset includes user speech data in different scenarios.

[0098] Step S320: Extract features from the speech behavior dataset to obtain first speech feature data. The first speech feature data includes first acoustic feature data, first language feature data, and statistical feature data.

[0099] Step S330: Superimpose the first acoustic feature data, the first language feature data, and the statistical feature data to obtain first superimposed feature data.

[0100] Step S340: Train the preset converter model with the first superimposed feature data to construct a preset user behavior model.

[0101] In this specific embodiment, the embodiment of the present invention first constructs a speech behavior data set. Specifically, the speech behavior data set in the embodiment of the present invention includes user speech data in different scenarios. Among them, in order to make the data cover various normal usage situations, the embodiment of the present invention collects a large amount of user speech data in different scenarios to construct a speech behavior data set. At the same time, the embodiment of the present invention annotates the features of each piece of speech data in the speech behavior data set, such as speech rate, pitch, pause, etc., and annotates the regular behavior patterns of each user. Then, the embodiment of the present invention extracts features from the speech behavior data set to obtain first speech feature data. Specifically, the first speech feature data in the embodiment of the present invention includes first acoustic feature data, first language feature data, and statistical feature data. Among them, the embodiment of the present invention extracts acoustic features such as Mel Frequency Cepstral Coefficients (MFCC), zero-crossing rate, and energy of each piece of user speech data in the speech behavior data set to obtain first speech feature data. At the same time, the embodiment of the present invention extracts language features such as speech rate, pause time, and pitch of each piece of user speech data in the speech behavior data set to obtain first language feature data. In addition, the embodiment of the present invention calculates statistical features such as the average speech rate, average pitch, and average pause time of the user corresponding to each piece of user speech data to obtain statistical feature data. Further, the embodiment of the present invention superimposes the first acoustic feature data, the first language feature data, and the statistical feature data to obtain first superimposed feature data. Specifically, the embodiment of the present invention stacks the extracted first acoustic feature data, first language feature data, and statistical feature data, as shown in the following formula (4):

[0102] F 2 =Concat(F v1 ,F l1 ,F s ) (4)

[0103] Wherein, Concat(·) in the formula is a function for superimposing in a specified dimension, F v1 represents the first acoustic feature data, F l1 represents the first language feature data, and F s represents the statistical feature data.

[0104] Further, in the embodiment of the present invention, the preset converter model is trained through the first superimposed feature data to construct a preset user behavior model. Specifically, in the embodiment of the present invention, the Transformer model is used to model the speech sequence features of the user and learn the normal behavior pattern of the user. Among them, the Transformer model can capture the long-distance dependencies in the input sequence when processing sequence data with its excellent capabilities, and at the same time effectively extract context information through the self-attention mechanism. Accordingly, in the process of constructing the entity recognition ability, the system will be trained with a large amount of labeled data. First, the preset converter model in the embodiment of the present invention encodes the input instructions, and captures the correlation between each element in the sequence through the multi-head self-attention mechanism, so as to obtain rich and comprehensive context information, as shown in the following formula (5):

[0105]

[0106] Among them, in the formula, Query is the query matrix, and Key and Value represent the content to be concerned about, is used for scaling to alleviate the problem of too large dot product, and softmax( ) represents the activation function.

[0107] Accordingly, the multi-head attention mechanism of the preset converter model in the embodiment of the present invention is as shown in the following formula (6):

[0108]

[0109] Among them, in the formula, Concat( ) represents the superimposing function. Finally, the embodiment of the present invention provides a non-linear change through a fully connected feed-forward network, as shown in the following formula (7):

[0110] FeedFarwordN(x)=max(0,xW 1 +b 1 )W 2 +b 2 (7)

[0111] Among them, in the formula, W 1 、b 1 、W 2 and b 2 are all variable parameters, and max( ) represents the maximum value function.

[0112] Accordingly, in the embodiment of the present invention, after the preset user behavior model is trained by inputting the first superimposed feature data into the constructed preset converter model, the speech data to be recognized is input into the preset user behavior model to reconstruct the input speech features, and the reconstruction error is calculated. When the reconstruction error exceeds the preset error threshold, it indicates that the speech recognition engine has an abnormal recognition.

[0113] Reference Figure 5 , in some embodiments of the present invention, before performing speech structure feature analysis on the audio data to be recognized through the first long short-term memory network model to obtain speech structure evaluation data, the speech recognition method provided by the embodiments of the present invention further includes but is not limited to the following steps:

[0114] Step S410: Construct a speech structure data set. The speech structure data set includes a number of user speech data containing different keywords and emotional information.

[0115] Step S420: Extract features from the speech structure data set to obtain second speech feature data. The second speech feature data includes second acoustic feature data and second language feature data.

[0116] Step S430: Superimpose the second acoustic feature data and the second language feature data to obtain second superimposed feature data.

[0117] Step S440: Input the second superimposed feature data into a preset long short-term memory network model for model training to construct the first long short-term memory network model.

[0118] In this specific embodiment, the embodiments of the present invention first construct a speech structure data set. Specifically, the speech structure data set in the embodiments of the present invention includes a number of user speech data containing different keywords and emotional information. Among them, the embodiments of the present invention obtain user speech data containing different keywords and emotions, and label the keyword types (such as modal particles, subject-predicate-object structures, adverbs, etc.) and emotional labels (such as positive, negative, neutral) for each piece of speech data, thereby constructing a speech structure data set. Then, the embodiments of the present invention extract features from the speech structure data set to obtain second speech feature data. Specifically, the second speech feature data in the embodiments of the present invention includes second acoustic feature data and second language feature data. Correspondingly, the embodiments of the present invention extract Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate, Energy, peaks and valleys from each piece of user speech data in the speech structure data set to obtain second acoustic feature data. In addition, the embodiments of the present invention extract the speech rate (Words Per Minute, WPM), pitch, and pause duration from each piece of user speech data to obtain second language feature data. Then, the embodiments of the present invention superimpose the second acoustic feature data and the second language feature data to obtain second superimposed feature data. Specifically, the embodiments of the present invention stack the extracted second acoustic feature data and second language feature data, as shown in the following formula (8):

[0119] F3 = Concat(F v2 , F l2 ) (8)

[0120] where Concat(·) is a function for stacking in a specified dimension, F v2 represents the second acoustic feature data, and F l2 represents the second language feature data.

[0121] Furthermore, in the embodiments of the present invention, the second stacked feature data is input into a preset long short-term memory network model for model training to construct a first long short-term memory network model. Specifically, the preset long short-term memory network model in the embodiments of the present invention refers to a pre-constructed original long short-term memory network model. Among them, in the embodiments of the present invention, the second stacked feature data is first input into the input layer of the preset long short-term memory network model for feature representation, as shown in the following formula (9):

[0122] H = embedding(F 3 ) (9)

[0123] where embedding(·) represents an embedding operation.

[0124] Next, in the embodiments of the present invention, the data after feature representation is input into a convolutional layer to extract local features, as shown in the following formula (10):

[0125] H c = CNN(H, k, p) (10)

[0126] where CNN(·) represents a convolution operation, k represents the size of the convolution kernel, and p represents the size of the blank area.

[0127] Then, in the embodiments of the present invention, pooling operation is performed on the local feature data extracted by the convolutional layer, as shown in the following formula (11):

[0128] H p = Pooling(H c , s) (11)

[0129] where Pooling(·) represents a pooling operation, and s represents the size of the pooling area.

[0130] Finally, in the embodiments of the present invention, the data is input into a long short-term memory (LSTM) network model to capture temporal dependencies, and keyword recognition and sentiment analysis are output through a fully connected layer, as shown in the following formula (12):

[0131]

[0132] Among them, in the formula, LSTM(·) represents the long short-term memory network, and max(·) represents the maximum value function, W 1 and b 1 and W 2 and b 2 are all variable parameters. Correspondingly, during the training process, the embodiments of the present invention construct a loss function to improve the recognition effect of the model, as shown in the following formula (13):

[0133]

[0134] Among them, in the formula, α and β represent weights, L kw and L emo are the loss functions of keyword and sentiment analysis respectively, N is the number of samples, C represents the number of keyword categories, E represents the number of sentiment categories, y ij and p ij are the label of the i-th sample belonging to the j-th class and the probability of the i-th sample belonging to the j-th class respectively.

[0135] Referring to Figure 6 , in some embodiments of the present invention, before performing the analysis of group characteristics on the audio data to be recognized by the second long short-term memory network model to obtain group characteristic evaluation data, the speech recognition method provided by the embodiments of the present invention further includes but is not limited to the following steps:

[0136] Step S510: Construct a group characteristic data set. Among them, the group characteristic data includes the user speech data of different user groups.

[0137] Step S520: Extract features from the group characteristic data set to obtain the third speech feature data. Among them, the third speech feature data includes the third acoustic feature data and the third language feature data.

[0138] Step S530: Superimpose the third acoustic feature data and the third language feature data to obtain the third superimposed feature data.

[0139] Step S540: Input the third superimposed feature data into a preset long short-term memory network model for model training to construct the second long short-term memory network model.

[0140] In this specific embodiment, the embodiment of the present invention first constructs a group feature dataset. Specifically, the group feature data in the embodiment of the present invention includes user speech data of different user groups, such as user speech data of different groups like the elderly, children, adults, etc. For example, the embodiment of the present invention collects speech data containing different keywords and emotions, including different groups such as the elderly, children, adults, etc., and labels each segment of speech data with keyword types (such as modal particles, subject-predicate-object structures, adverbs, etc.) and emotion labels (such as positive, negative, and neutral), and at the same time labels the corresponding user's age group (such as the elderly group, children group, adult group) and speech rate (such as fast, medium, slow). Then, the embodiment of the present invention extracts features from the group feature data to obtain the third speech feature data. Specifically, the third speech feature data in the embodiment of the present invention includes third acoustic feature data and third language feature data. Among them, the embodiment of the present invention extracts Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate, Energy, peaks and valleys from each segment of user speech data in the group feature dataset to obtain the third acoustic feature data. In addition, the embodiment of the present invention extracts the speech rate (Words Per Minute, WPM), pitch, and pause duration from each segment of user speech data to obtain the third language feature data. Then, the embodiment of the present invention superimposes the third acoustic feature data and the third language feature data to obtain the third superimposed feature data. Specifically, the embodiment of the present invention stacks the extracted third acoustic feature data and third language feature data as shown in the following formula (14):

[0141] F 4 = Concat(F v3 , F l3 ) (14)

[0142] Wherein, Concat(·) in the formula is a function for superimposing in a specified dimension, F v3 represents the third acoustic feature data, and F l3 represents the third language feature data.

[0143] Furthermore, the embodiment of the present invention inputs the third superimposed feature data into a preset long short-term memory network model for model training to construct the second long short-term memory network model. Specifically, the process of constructing the preset long short-term memory network model in the embodiment of the present invention is the same as the process of constructing the preset long short-term memory network model in the training process of the first long short-term memory network model, which will not be elaborated here. Correspondingly, the embodiment of the present invention constructs a loss function to improve the recognition effect of the model, as shown in the following formula (15):

[0144]

[0145] Among them, in the formula, α, β, and γ are weights, and L age represents the loss function of the age group, and L speed represents the loss function of the speech rate regression, and L kw represents the loss function of the keyword, N is the number of samples, A is the number of age group categories, and y ij , p ij , y i , respectively represent the label that the i-th sample belongs to the j-th class, the probability that the i-th sample belongs to the j-th class, the true speech rate of the i-th sample, and the predicted preset.

[0146] It should be noted that in some embodiments of the present invention, data is collected through a data collection module to provide high-quality raw data support for subsequent performance evaluation and model training. Among them, the functions of the data collection module in the embodiments of the present invention include, but are not limited to, the collection of user voice data and the recovery of the output data of the automatic speech recognition (ASR) engine. In addition, in order to ensure the wide representativeness and comprehensive coverage of the collected data, the data collection module also needs to pay special attention to and incorporate speech samples under various environmental noise conditions, as well as speech inputs from different regions with different accents. In this way, the robustness and adaptability of the system can be effectively improved, ensuring that it can still maintain high-performance in complex and changing actual application environments.

[0147] Referring to Figure 7 , in some embodiments of the present invention, the speech recognition method provided by the embodiments of the present invention further includes, but is not limited to, the following steps:

[0148] Step S610: Obtain the system log file. Among them, the system log file includes user instructions and corresponding feedback information.

[0149] Step S620: Search and analyze the system log file to obtain abnormal instruction data. Among them, the abnormal instruction data includes user instructions with parsing problems and corresponding feedback information.

[0150] Step S630: Perform correction processing according to the abnormal instruction data, and update the system knowledge system according to the correction result. Among them, the correction processing includes re-parsing the instruction operation, correcting the error operation, and parameter adjustment operation.

[0151] In this specific embodiment, in the actual deployed production environment, the system judges and processes user instructions through the user feedback module to meet user needs to the greatest extent. Although the system has been elaborately designed and optimized, there may still be problems in the parsing of some instructions. Therefore, to alleviate the above-mentioned instruction parsing problems, the embodiment of the present invention constructs a user feedback module, which directly collects user satisfaction evaluations through the user interface or indirectly obtains feedback information by analyzing user usage behaviors, thereby enhancing the robustness of the system while providing a smoother and seamless interaction environment for users. Specifically, the embodiment of the present invention first obtains the system log file to search and analyze the system log file to obtain abnormal instruction data. Among them, the embodiment of the present invention records the corresponding feedback information of all instruction sets in the system log file, so the system log file includes user instructions and corresponding feedback information. Correspondingly, the embodiment of the present invention regularly checks the system log file through the user feedback module to find user instructions with parsing problems and related feedback information therein to obtain abnormal instruction data. Further, the embodiment of the present invention performs correction processing according to the abnormal instruction data and updates the system knowledge system according to the correction result, thereby improving the knowledge system of the system and enhancing the robustness and stability of the system. Among them, the correction processing in the embodiment of the present invention includes re-parsing instruction operations, correcting error operations, and parameter adjustment operations. For example. When the user feedback module determines that a user instruction with incorrect parsing is found, the embodiment of the present invention reprocesses it, including operations such as re-parsing the instruction, correcting the error, and adjusting the parameters, so that the user instruction can be successfully recognized and executed, and reported to the operation system. At the same time, the embodiment of the present invention feeds back the re-parsed result and relevant feedback information to the operation system, and the operation system notifies relevant operation personnel to compound the processing result. Correspondingly, if the processing is successful, it passes. On the contrary, the embodiment of the present invention will perform warehousing record processing on high-frequency instructions to improve the knowledge system of the system.

[0152] It should be noted that in some embodiments of the present invention, for specific fields (such as the medical and legal fields), the system needs to process different professional terms and background knowledge. The embodiments of the present invention construct customized and domain-adaptive modules and train them through domain-specific corpora to improve the adaptability of the model to that field. Therefore, the embodiments of the present invention construct an adaptive module knowledge system and define a series of rules and steps to ensure the stability and scalability of the system. Exemplarily, the embodiments of the present invention first define knowledge parameters, define a description for each knowledge, including its name, type, background, etc. At the same time, define the corresponding knowledge format, including the unique code (idcode) of the knowledge, type (tycode), and other possible fields. Then, the embodiments of the present invention establish the mapping between knowledge and knowledge, that is, define which knowledge and which knowledge are of the same background knowledge, etc. The embodiments of the present invention improve the stability and scalability of the system by constructing an adaptive module.

[0153] Referring to Figure 8 , in some embodiments of the present invention, the switching weight data is dynamically determined through a preset switching mechanism, and the engine hybrid output is performed according to the switching weight data to obtain a speech recognition result, including but not limited to the following steps:

[0154] Step S710: Dynamically calculate the switching weight data through a preset linear interpolation algorithm.

[0155] Step S720: Perform speech recognition on the audio data to be recognized through the historical speech recognition engine to obtain the first speech recognition data.

[0156] Step S730: Perform speech recognition on the audio data to be recognized through the expected speech recognition engine to obtain the second speech recognition data.

[0157] Step S740: Perform weighted average processing on the first speech recognition data and the second speech recognition data according to the switching weight data to obtain a speech recognition result.

[0158] In this specific embodiment, the present invention first dynamically calculates the switching weight data through a preset linear interpolation algorithm. Specifically, in the embodiments of the present invention, the switching weight data changes with time, and the switching weight data is dynamically adjusted, so as to seamlessly switch to a new ASR engine, that is, the expected speech recognition engine, without interrupting the user interaction. For example, in the embodiments of the present invention, the initial time point t 0 and the end time point t f , and the initial weights of the old engine E old (historical speech recognition engine) and the new engine E new (expected speech recognition engine) are w old (t 0 ) = 1 and wnew (t 0 ) = 0. Correspondingly, the calculation formula of the preset linear interpolation algorithm in the embodiments of the present invention is as shown in the following formula (16):

[0159]

[0160] Wherein, in the formula, t represents the current time, and w new (t) represents the weight of the new engine E new at time t, and w old (t) represents the weight of the old engine E old at time t.

[0161] Meanwhile, in the embodiments of the present invention, the historical speech recognition engine is used to perform speech recognition on the audio data to be recognized, and the first speech recognition data is obtained, and the expected speech recognition engine is used to perform speech recognition on the audio data to be recognized, and the second speech recognition data is obtained. Specifically, for each segment of the audio data to be recognized, in the embodiments of the present invention, the historical speech recognition engine and the expected speech recognition engine are respectively used for speech recognition, so as to respectively obtain the recognition results of the new engine E new and the old engine E old . Finally, in the embodiments of the present invention, the first speech recognition data and the second speech recognition data are weighted and averaged according to the switching weight data to obtain the speech recognition result. Specifically, in the embodiments of the present invention, according to the recognition results of the new and old speech recognition engines, weighted average processing is performed according to the currently calculated switching weight data, as shown in the following formula (17):

[0162] O t = w old (t) * E old (Audio) + w new (t) * E new (Audio) (17)

[0163] Correspondingly, in the embodiments of the present invention, the speech recognition output is performed according to the speech recognition result obtained by the weighted average processing, so that it is possible to seamlessly switch to a new ASR engine without interrupting the user interaction, ensuring the continuity of speech recognition. In addition, in the embodiments of the present invention, the values of w new (t) and w old (t) are continuously updated through the weight adjustment function until t = t f , at this time, w old (t) = 0 and w new (t) = 1, and at this time the engine is completely switched to the new engine, completing the seamless switching of the speech recognition engine, effectively improving the user experience of speech recognition.

[0164] Next, in combination with a specific speech recognition scenario, the solution of the embodiment of the present invention will be introduced and described in detail:

[0165] Exemplarily, as Figure 9 shown, Figure 9 is a schematic diagram of the overall process of the speech recognition method provided by the embodiment of the present invention. Specifically, in the data collection stage of the embodiment of the present invention, various speech samples of the language processing and analysis platform are collected, including recordings under different dialects, speech rates, and background noise conditions. At the same time, for each recording, its metadata and the results output by the ASR engine are recorded. Next, in the evaluation and decision-making stage, the embodiment of the present invention uses an audio length anomaly detection mechanism to check whether there are abnormally short or long recognition results. At the same time, the embodiment of the present invention also analyzes personal speech habits. For example, if a user who often uses a fast speech rate suddenly slows down, this may indicate a problem with ASR recognition. In addition, the embodiment of the present invention also detects speech structure features, such as identifying more filler words or emotional fluctuations, which may indicate that the user has an abnormal emotion. And the embodiment of the present invention also adjusts parameters according to group characteristics. For example, the elderly may use a slower speech rate, while children may have different expression choices. Further, in the engine switching stage, when it is found that the currently used ASR engine performs poorly in some aspects (such as slow response speed or low accuracy), the embodiment of the present invention switches to a new engine that is more suitable for the current situation in a smooth transition manner. Then, the embodiment of the present invention analyzes and corrects the misrecognition cases in the record through the user feedback module. At the same time, the embodiment of the present invention customizes an adaptation module to train through a corpus in a specific field, improving the adaptability of the model to this field, and further improving the stability and scalability of the system.

[0166] It is easy to understand that the embodiment of the present invention can monitor and evaluate the performance of the ASR engine in real time, including key indicators such as accuracy and response time, and at the same time process real-time data streams from different engines, update the evaluation results in real time, and can respond more quickly to environmental changes and engine performance fluctuations, so as to achieve a better user experience. In addition, the embodiment of the present invention can predict and evaluate the speech effect more accurately and achieve more refined decision-making. This data-driven decision-making mechanism can better adapt to the complex and changeable speech recognition environment compared with the traditional decision-making methods based on rules or static parameters. At the same time, the seamless switching mechanism in the embodiment of the present invention can seamlessly switch to the best-performing ASR engine without affecting the user experience, so as to provide a smoother user experience and reduce delays and interruptions during the switching process. Correspondingly, the embodiment of the present invention can identify and adapt to different usage scenarios and user preferences, dynamically adjust its algorithm parameters, so as to maintain the best performance in various environments and conditions, and thus can more effectively improve the adaptability and performance of the system, and achieve continuous performance optimization and improvement of the user experience.

[0167] Please refer to Figure 10 , the embodiment of the present application also provides a voice recognition system, which can implement the above voice recognition method. The system includes:

[0168] The first module 810 is used to obtain the audio data to be recognized.

[0169] The second module 820 is used to input the audio data to be recognized into a preset evaluation and decision module for multi-dimensional performance evaluation of the voice recognition engine, and obtain performance evaluation data in several dimensions.

[0170] The third module 830 is used to determine the expected voice recognition engine through a preset decision network according to the performance evaluation data in several dimensions.

[0171] The fourth module 840 is used to dynamically determine the switching weight data through a preset switching mechanism, and perform engine hybrid output according to the switching weight data to obtain the voice recognition result. Among them, the switching weight data includes the hybrid output ratio of the historical voice recognition engine and the expected voice recognition engine.

[0172] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0173] The embodiment of the present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above voice recognition method. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0174] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0175] Please refer to Figure 11 , Figure 11 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0176] The processor 910 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0177] The memory 920 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 920 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 920 and are called by the processor 910 to execute the voice recognition method of the embodiments of this application;

[0178] The input / output interface 930 is used to implement information input and output;

[0179] The communication interface 940 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0180] The bus 950 transmits information between the various components of the device (such as the processor 910, the memory 920, the input / output interface 930, and the communication interface 940);

[0181] Among them, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are communicatively connected to each other inside the device through the bus 950.

[0182] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned voice recognition method is implemented.

[0183] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0184] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0185] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0186] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0188] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0189] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0190] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0191] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0192] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0194] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0195] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.

Claims

1. A speech recognition method, characterized in that: The method comprises the following steps: Obtain audio data to be recognized; Inputting the audio data to be recognized into a preset evaluation decision module to perform multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data of several dimensions; Determine the desired speech recognition engine through a preset decision network according to the performance evaluation data of several dimensions; The switching weight data is dynamically determined through a preset switching mechanism to perform engine mixed output according to the switching weight data to obtain a speech recognition result; wherein the switching weight data includes a mixed output ratio of a historical speech recognition engine and the expected speech recognition engine.

2. The method according to claim 1, characterized in that The performance evaluation data includes audio length evaluation data, behavior characteristic evaluation data, speech structure evaluation data and group characteristic evaluation data; The audio data to be recognized is input into a preset evaluation decision module to perform multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data of several dimensions, including: Performing length anomaly analysis on the audio data to be identified by using a preset audio length threshold to obtain the audio length evaluation data; wherein the preset audio length threshold is determined by the average length and standard deviation of a preset audio data set; Performing user behavior analysis on the audio data to be identified by using a preset user behavior model to obtain the behavior feature evaluation data; Performing speech structure feature analysis on the audio data to be recognized by using a first long short-term memory network model to obtain the speech structure evaluation data; The group characteristic analysis is performed on the audio data to be identified through a second long short-term memory network model to obtain the group characteristic evaluation data.

3. The method according to claim 2, characterized in that Before performing the user behavior analysis on the audio data to be identified by using a preset user behavior model to obtain the behavior feature evaluation data, the method further includes: Constructing a voice behavior data set; wherein the voice behavior data set includes user voice data in different scenarios; Performing feature extraction on the speech behavior data set to obtain first speech feature data; wherein the first speech feature data includes first acoustic feature data, first language feature data and statistical feature data; Superimposing the first acoustic feature data, the first language feature data, and the statistical feature data to obtain first superimposed feature data; The preset converter model is trained by using the first superimposed feature data to construct the preset user behavior model.

4. The method according to claim 2, characterized in that: Before performing the speech structure feature analysis of the audio data to be recognized by using the first long short-term memory network model to obtain the speech structure evaluation data, the method further includes: Constructing a speech structure data set; wherein the speech structure data set includes a plurality of user speech data containing different keywords and emotional information; Performing feature extraction on the speech structure data set to obtain second speech feature data; wherein the second speech feature data includes second acoustic feature data and second language feature data; Superimposing the second acoustic feature data and the second language feature data to obtain second superimposed feature data; The second superimposed feature data is input into a preset long short-term memory network model for model training to construct the first long short-term memory network model.

5. The method according to claim 2, characterized in that: Before performing the group characteristic analysis on the audio data to be identified by using the second long short-term memory network model to obtain the group characteristic evaluation data, the method further includes: Constructing a group characteristic data set; wherein the group characteristic data includes user voice data of different user groups; Performing feature extraction on the group feature data set to obtain third voice feature data; wherein the third voice feature data includes third acoustic feature data and third language feature data; Superimposing the third acoustic feature data and the third language feature data to obtain third superimposed feature data; The third superimposed feature data is input into a preset long short-term memory network model for model training to construct the second long short-term memory network model.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining a system log file; wherein the system log file includes user instructions and corresponding feedback information; Searching and analyzing the system log file to obtain abnormal instruction data; wherein the abnormal instruction data includes the user instruction with the problem of analysis and the corresponding feedback information; Correction processing is performed according to the abnormal instruction data, and the system knowledge system is updated according to the correction result; wherein the correction processing includes re-parsing instruction operations, correcting error operations and parameter adjustment operations.

7. The method according to claim 1, characterized in that The dynamically determining the switching weight data by the preset switching mechanism, and performing engine mixed output according to the switching weight data to obtain the speech recognition result, comprises: The switching weight data is obtained by dynamically calculating through a preset linear interpolation algorithm; Performing speech recognition on the audio data to be recognized by the historical speech recognition engine to obtain first speech recognition data; Performing speech recognition on the audio data to be recognized by the expected speech recognition engine to obtain second speech recognition data; The first speech recognition data and the second speech recognition data are weighted averaged according to the switching weight data to obtain the speech recognition result.

8. A speech recognition system, characterized in that: The system comprises: The first module is used to obtain audio data to be recognized; The second module is used to input the audio data to be recognized into a preset evaluation decision module to perform multi-dimensional speech recognition engine performance evaluation to obtain performance evaluation data of several dimensions; A third module is used to determine a desired speech recognition engine through a preset decision network according to the performance evaluation data of several dimensions; The fourth module is used to dynamically determine the switching weight data through a preset switching mechanism, so as to perform engine mixed output according to the switching weight data to obtain a speech recognition result; wherein the switching weight data includes the mixed output ratio of the historical speech recognition engine and the expected speech recognition engine.

9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.