An artificial intelligence-based precise voice recognition method and system

By constructing sets of voice cleaning features and user recognition features, performing multi-channel voice recognition and semantic recognition aggregation, the problem of low voice recognition accuracy in existing technologies is solved, achieving higher user voice recognition accuracy and a better human-computer interaction experience.

CN116884403BActive Publication Date: 2026-08-25SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311098967.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2026-08-25
Estimated Expiration
2043-08-29

AI Technical Summary

Technical Problem

Existing speech recognition technologies have low accuracy rates due to their low relevance to users, which affects the user's human-computer interaction experience.

Method used

By collecting voice data from target users, a set of voice cleaning features and user identification features is constructed. Multi-channel voice recognition is performed, and semantic recognition is aggregated. The multi-channel voice conversion results are integrated to output accurate voice recognition results.

Benefits of technology

It improves the accuracy of user speech recognition, reduces the impact of noise data on recognition results, and enhances the precision of user speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884403B_ABST
    Figure CN116884403B_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure provides a kind of precision voice recognition method and system based on artificial intelligence, it is related to speech recognition technology, the method provided in the embodiment of the present disclosure comprises: collecting user voice data, and generating collection environment identification;Build speech cleaning features, the speech cleaning features are obtained based on user database and collection environment identification matching;User recognition feature set is constructed;Multi-channel speech recognition is executed, and multi-channel speech recognition is obtained by user recognition feature set to data cleaning after user voice data recognition;Semantic recognition aggregation is carried out, and the semantic recognition aggregation is obtained by the aggregation of multi-channel speech conversion result;According to semantic recognition aggregation result, multi-channel speech conversion result is integrated and output voice recognition result.Can solve the technical problem that the existing speech recognition technology is lower due to the degree of association with user, and the accuracy of user speech recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to speech recognition technology, and more specifically, to an artificial intelligence-based method and system for accurate speech recognition. Background Technology

[0002] Speech recognition technology, a branch of artificial intelligence, is a crucial component of human-computer natural interaction technology. In actual speech recognition processes, because each user uses different language types and sentence structures, the same words can convey different meanings depending on the pauses. Furthermore, each user has unique speech characteristics, such as filler words, pause habits, and unclear pronunciation. Current speech recognition technologies rarely analyze these characteristics, often resulting in low accuracy and negatively impacting the user's human-computer interaction experience.

[0003] The shortcomings of existing speech recognition technology are that the low degree of association with the user results in low accuracy of user speech recognition. Summary of the Invention

[0004] Therefore, in order to solve the above-mentioned technical problems, the technical solutions adopted in the embodiments of this disclosure are as follows:

[0005] An AI-based precision speech recognition method includes the following steps: collecting user speech data of a target user and generating a collection environment identifier for the user speech data; constructing speech cleaning features, which are obtained by calling the user database of the target user and matching the user database with the collection environment identifier; constructing a user identification feature set, which is constructed by extracting user features from the user database; performing multi-channel speech recognition, which involves cleaning the user speech data based on the speech cleaning features and then recognizing the cleaned user speech data using the user identification feature set; performing semantic recognition aggregation, which involves obtaining multi-channel speech conversion results and aggregating the multi-channel speech conversion results; integrating the multi-channel speech conversion results based on the semantic recognition aggregation results, and outputting a speech recognition result based on the integrated result.

[0006] An AI-based precision speech recognition system includes: a user speech data acquisition module for acquiring user speech data of a target user and generating an acquisition environment identifier for the user speech data; a speech cleaning feature construction module for constructing speech cleaning features, which are obtained by calling the target user's user database and matching the user database with the acquisition environment identifier; a user identification feature set construction module for constructing a user identification feature set, which is constructed by extracting user features from the user database; and a multi-channel speech recognition system. The system includes a recognition execution module, which performs multi-channel speech recognition by cleaning the user's speech data based on the speech cleaning features and then recognizing the cleaned user's speech data using the user recognition feature set; a semantic recognition aggregation module, which performs semantic recognition aggregation by obtaining multi-channel speech conversion results and aggregating them; and a speech recognition result output module, which integrates the multi-channel speech conversion results based on the semantic recognition aggregation results and outputs a speech recognition result based on the integrated result.

[0007] Due to the adoption of the above technical solution, the technical progress achieved by this disclosure compared with the prior art is as follows:

[0008] (1) It can solve the technical problem that the accuracy of user speech recognition is low due to the low degree of correlation between the existing speech recognition technology and the user. By setting up multiple channels according to user characteristics to recognize and convert the user speech data of the target user, and integrating the multi-channel speech conversion results according to the semantic aggregation results, the speech recognition result of the target user can be obtained, which can improve the accuracy of user speech recognition.

[0009] (2) By cleaning the user voice data, the impact of noise data in the user voice data on the user voice recognition result can be reduced, thereby improving the accuracy of user voice recognition. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0011] Figure 1 This application provides a flowchart illustrating an artificial intelligence-based precision speech recognition method.

[0012] Figure 2This application provides a flowchart illustrating the process of obtaining multi-channel speech conversion results based on a first speech conversion result in an artificial intelligence-based precision speech recognition method.

[0013] Figure 3 This application provides a schematic diagram of the structure of an artificial intelligence-based precision speech recognition system. Detailed Implementation

[0014] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0015] Based on the above description, please refer to Figure 1 ,like Figure 1 As shown, this disclosure provides an artificial intelligence-based method for accurate speech recognition, including:

[0016] Collect user voice data of the target user and generate a collection environment identifier for the user voice data;

[0017] With the rise of smart homes, smart cars, smart assistants, and other intelligent devices and services, voice recognition technology will become one of the core technologies for realizing human-computer interaction. People can communicate with devices through voice to achieve functions such as intelligent control, information retrieval, and task execution, which can improve life and work efficiency and enhance people's life experience.

[0018] Therefore, improving the accuracy of speech recognition is a crucial issue. In the embodiments of this disclosure, data processing, including data collection and cleaning, is primarily achieved through intelligent means and devices. This intelligent processing method significantly improves the efficiency of data acquisition, thereby enhancing the intelligence of speech recognition. The method provided in this disclosure is used to accurately recognize user speech data, thereby improving the accuracy of user speech recognition.

[0019] First, user voice data of the target users is collected. The voice data collection method can be realized through intelligent means. For example, the data collection sensor can be trained to perform data denoising through artificial intelligence network. Different voice data denoising parameters can be obtained through training according to different voice data collection scenarios. When collecting voice data, the voice denoising parameters can be automatically matched according to the voice data collection scenario to realize automated initial data denoising of voice data, thereby improving the accuracy of data collection and saving data processing time.

[0020] The target user refers to any user to be subjected to speech recognition. Then, the user's speech data undergoes a collection environment analysis. The collection environment refers to the scenario in which the target user's speech data is collected, such as a shopping mall, factory workshop, subway, bus, or office. This collection environment analysis can identify unique sound features specific to the scenario; for example, in a factory workshop, there are sound features such as machine operation sounds; on a bus, there are sound features such as bus engine noise. Then, the user's speech data is labeled with the collection environment, generating user speech data with the collection environment label. Obtaining user speech data with the collection environment label provides raw data support for user speech data recognition.

[0021] Constructing voice cleaning features, wherein the voice cleaning features are obtained by calling the user database of the target user and matching the user database with the collection environment identifier;

[0022] A user database for the target user is constructed based on the target user's historical voice data and historical voice recognition results. The historical voice data comprises multiple sets of historical voice data from various acquisition environments, and the historical voice data and historical voice recognition results have a corresponding relationship. For example, background sounds, including pedestrian voices and advertising sounds, are extracted from the target user's multiple sets of historical voice data using an artificial intelligence network. Background sound extraction can be performed based on the actual situation. Then, the voice acquisition environment is intelligently identified and labeled based on the background sound extraction results. The multiple sets of historical voice data are then divided according to the acquisition environment type to obtain multiple sets of historical voice data, and multiple sets of historical voice recognition results corresponding to these sets of historical voice data are obtained. Finally, based on the principle of decision trees, the historical voice data and historical voice recognition results are stored separately according to the acquisition environment type to construct the user database.

[0023] Then, the user database of the target user is accessed, and the collection environment identifier is input into the user database to match historical voice data from the same collection environment, obtaining multiple sets of matched historical voice data and multiple sets of matched historical voice recognition results. Noise analysis of the collection environment is then performed based on these multiple sets of matched historical voice data and multiple sets of matched historical voice recognition results. For example, when the collection environment is a shopping mall, there are various noises such as shop advertisements, pedestrian conversations, and children playing. Then, voice cleaning features are obtained based on the noise analysis results. These voice cleaning features refer to the noise characteristics in the user's voice data. Obtaining these voice cleaning features provides support for the next step of data cleaning of the user's voice data.

[0024] Construct a user identification feature set, which is constructed by extracting user features from the user database;

[0025] First, user voice data feature analysis is performed on the historical voice data and historical speech recognition results in the user database. These user voice data features include various voice characteristics such as the user's language type, filler words, and pause habits. Based on the user voice data feature analysis results, a user identification feature set is constructed, which includes user language features, user sentence features, user pause habits, and user part-of-speech aggregation features. Constructing this user identification feature set provides support for the next step of performing multi-channel speech recognition on the user voice data.

[0026] Perform multi-channel speech recognition, which is to clean the user's speech data based on the speech cleaning features, and then recognize the cleaned user's speech data through the user recognition feature set.

[0027] like Figure 2 As shown, in one embodiment, it further includes:

[0028] A basic recognition channel is constructed, which is built based on the language features extracted from the user's recognition feature set and the language feature extraction results.

[0029] A statement analysis channel is constructed, which is built based on the statement usage features extracted from the user identification feature set.

[0030] A first recognition channel is constructed, wherein the first recognition channel is obtained by mixing the basic recognition channel with the sentence analysis channel;

[0031] The user's voice data after data cleaning is processed by the first recognition channel to perform channel recognition and output the first voice conversion result.

[0032] The multi-channel speech conversion result is obtained based on the first speech conversion result.

[0033] First, voice cleaning methods are matched based on voice cleaning features, and the user voice data is cleaned using these methods. The voice cleaning methods refer to various sound denoising techniques, and there is a correspondence between the voice cleaning methods and the voice cleaning features. Those skilled in the art can select appropriate cleaning methods based on the actual situation of the voice cleaning features. For example, the adaptive noise reduction function in audio processing software can be used to denoise the sound data. The noise reduction level can be automatically matched according to the noise level in the voice cleaning features to achieve sound denoising processing, obtaining cleaned user voice data. By cleaning the user voice data, the impact of noise data on the user voice recognition results can be reduced, thereby improving the accuracy of user voice recognition.

[0034] The user language features of the target user are obtained based on the user identification feature set. The user language features refer to the language type used in the user's historical voice data, such as local dialect, Mandarin, English, Japanese and other foreign languages, professional vocabulary in special fields, etc. Then, a basic recognition channel is constructed based on the user language features. The basic recognition channel is used to recognize the text in the user's voice data according to the user's language type.

[0035] Then, user sentence usage features are extracted from the user identification feature set. Sentence features refer to the grammatical structure and usage of sentences. Sentence structure mainly includes components such as subject, predicate, object, complement, adverbial, and attributive. Sentence usage mainly includes declarative sentences, interrogative sentences, imperative sentences, and exclamatory sentences. Sentence usage features refer to users' habits of using sentence features. For example, in Shandong dialect, people are accustomed to using subject-predicate inversion sentences. For instance, "It's going to rain" is expressed as "It's going to rain" in Shandong dialect; "Let's go see a movie" is expressed as "Let's go see a movie." Then, a sentence analysis channel is constructed based on the sentence usage features, which is used to identify sentence features in user speech data.

[0036] The basic recognition channel and the sentence analysis channel are fused to obtain a first recognition channel. This first recognition channel is used to recognize the text in the user's speech data, and then to perform sentence feature recognition on the text recognition result. The cleaned user speech data is then subjected to speech data recognition through the first recognition channel to obtain a first speech conversion result. This first speech conversion result is a text recognition result of the user speech data with sentence features, and it is added to the multi-channel speech conversion result. Obtaining the first speech conversion result provides support for recognizing user speech data at the sentence usage feature level.

[0037] In one embodiment, it further includes:

[0038] Configure a word - nature aggregation channel, which is obtained by extracting the word - nature features of the user based on the user identification feature set;

[0039] Construct a second recognition channel, which is formed by mixing the word - nature aggregation channel and the basic recognition channel;

[0040] Perform channel recognition on the user voice data after data cleaning through the second recognition channel, and output a second voice conversion result;

[0041] Obtain the multi - channel voice conversion result based on the first voice conversion result and the second voice conversion result.

[0042] First, extract the user's word - nature features in the user identification feature set. The word - nature includes content words and function words. Content words include nouns, verbs, adjectives, etc.; function words include adverbs, prepositions, conjunctions, etc. The word - nature feature refers to the feature of the word effect that the user wants to express when using the word. For example, for the two characters "interaction", in the sentence "Implement product interaction functions", "interaction" means a verb; in the sentence "Retrieve interaction data", "interaction" means a noun. Then configure a word - nature aggregation channel according to the extraction result of the user's word - nature features. The word - nature aggregation channel is used to analyze the features expressed by special words in the user voice data.

[0043] Perform channel fusion on the word - nature aggregation channel and the basic recognition channel to obtain a second recognition channel. The second recognition channel is used to recognize the text in the user voice data, and then perform word - nature feature recognition on the text recognition result. When recognizing the text recognition result through word - nature features, it is necessary to analyze in combination with the words before and after or the sentences before and after the word. For example, for the character "bao", in "make dumplings", it is used as a verb to describe the action of "bao"; in "schoolbag", it is used as a noun to describe an item. Then perform voice data recognition on the user voice data after data cleaning through the second recognition channel to obtain a second voice conversion result. The second voice conversion result is the text recognition result of the user voice data with word - nature features, and add the second voice conversion result to the multi - channel voice conversion result. By obtaining the second voice conversion result, it provides support for the recognition of the user voice at the word - nature feature level.

[0044] In one embodiment, it further includes:

[0045] Extract the user's pause features based on the user identification feature set, where the pause features carry pause credibility identifiers;

[0046] A third recognition channel is constructed, which is obtained by constructing a pause aggregation channel through the pause features and mixing the basic recognition channel with the pause aggregation channel;

[0047] The third recognition channel is used to recognize the cleaned user voice data channel, and the third voice conversion result is output.

[0048] The multi-channel speech conversion result is obtained based on the first speech conversion result, the second speech conversion result, and the third speech conversion result.

[0049] First, user pause features are extracted from the user identification feature set. These user pause features refer to the patterns of pauses a user makes while speaking. The pause features include necessary pauses and non-necessary pauses based on the user's personal habits. For example, in the sentence "I'm very happy today because my son got first place in his exam, so I'm treating everyone today," the pause after "first place" is necessary. However, based on the user's pausing habits, the sentence is divided into segments: I, today, very happy, is, because, son, exam, got, first place. The pauses in this sentence are all non-necessary and will be paused according to the user pause features. These user pause features include a pause confidence indicator, which characterizes the confidence level of the pause. For example, the confidence level of necessary pauses is the highest, while the confidence level of non-necessary pauses can be set according to the user pause features. The confidence level can be represented by a pause coefficient; the larger the pause coefficient, the higher the confidence level of the pause.

[0050] A pause aggregation channel is constructed based on user pause characteristics. This channel embeds a pause credibility analysis model, which is a neural network model in machine learning capable of continuous iterative optimization. The input data for the pause analysis model are pause points in speech data, and the output data is the pause credibility coefficient. A model dataset is constructed based on historical user pause characteristics with pause credibility identifiers. This dataset is divided into a model training dataset and a model validation dataset according to a data partitioning ratio. This partitioning ratio can be set by those skilled in the art based on actual conditions; for example, the model training dataset accounts for 90%, and the model validation dataset accounts for 10%. The pause credibility analysis model is then trained under supervision using the model training dataset. When the model approaches convergence, it is validated using the model validation dataset. A preset model validation accuracy index is used, which can be customized based on actual conditions; for example, the model output accuracy is 96%. When the model output accuracy exceeds the model validation accuracy index, the trained pause credibility analysis model is obtained. By constructing a pause credibility analysis model based on neural networks, the accuracy and efficiency of pause credibility identification in user speech data can be improved, thereby enhancing the accuracy and efficiency of user speech data recognition.

[0051] Then, the pause aggregation channel and the basic recognition channel are fused to obtain a third recognition channel. This third recognition channel is used to recognize text in the user's speech data, and then to identify pause features in the text recognition results. Based on the third recognition channel, the cleaned user speech data is used for speech data recognition, and the pause confidence analysis model in the pause aggregation channel is used to assign pause confidence levels to the pause positions, resulting in a third speech conversion result. This third speech conversion result is a text recognition result of the user speech data with pause features, and these pause features are labeled with pause confidence levels. The third speech conversion result is added to the multi-channel speech conversion result. Obtaining this third speech conversion result provides support for pause feature-level recognition of user speech.

[0052] Semantic recognition aggregation is performed, which is obtained by acquiring multi-channel speech conversion results and aggregating the multi-channel speech conversion results;

[0053] In one embodiment, it also includes:

[0054] A strong or weak correlation identifier for converting the third speech conversion result based on the pause credibility identifier;

[0055] When the semantic recognition aggregation is performed using the first speech conversion result, the second speech conversion result, and the third speech conversion result, the aggregation process is optimized using the strong and weak association identifiers.

[0056] The semantic recognition aggregation is completed based on the aggregation optimization results.

[0057] A multi-channel speech conversion result is obtained, including a first speech conversion result, a second speech conversion result, and a third speech conversion result. Based on the pause confidence indicator, the pause positions in the speech data text recognition result of the third speech conversion result are assigned strong or weak association indicators, where the pause position with higher confidence is associated with a stronger indicator, and the pause position with lower confidence is associated with a weaker indicator.

[0058] Then, the first speech conversion result, the second speech conversion result, and the third speech conversion result are subjected to semantic recognition aggregation. Semantic recognition aggregation refers to the word aggregation of the text recognition results in the user's speech data, and the recognition of the content and meaning of the word aggregation results. In the process of semantic recognition aggregation, the aggregation process is optimized by the strong and weak association identifiers. For example, in the sentence "Are you making dumplings?", users are used to pausing at "Are you making dumplings?" and "Are you making dumplings?". If the pause confidence identifier at this pause position is low, this pause position can be removed, and "Are you making dumplings?" and "Are you making dumplings?" can be merged into one sentence. Finally, the semantic recognition aggregation of the multi-channel semantic conversion results is completed based on the aggregation optimization results.

[0059] By setting a pause confidence identifier and optimizing the semantic recognition aggregation process of multi-channel speech conversion results based on the pause confidence identifier, the accuracy of the semantic recognition aggregation results can be improved, thereby improving the accuracy of user speech data recognition.

[0060] The multi-channel speech conversion results are integrated based on the semantic recognition aggregation results, and the speech recognition results are output based on the integrated results.

[0061] The speech data text recognition results in the multi-channel speech conversion results are integrated based on the semantic recognition results. Integration refers to fusing the speech data text recognition results with the sentence content to obtain the speech recognition result. For example, the speech data text recognition result is "I'm cooking in the kitchen on the second floor, the knife is in the storage room on the first floor, quickly get it to me, I need to cut potatoes." After integration with the semantic recognition result, it becomes "I'm cooking in the kitchen on the second floor," "The knife is in the storage room on the first floor," "Quickly get it to me," and "I need to cut potatoes."

[0062] By setting multiple channels according to user characteristics, the user voice data of the target user is recognized and converted, and the multi-channel voice conversion results are integrated according to the semantic aggregation results to output the voice recognition result of the target user. This solves the technical problem that the existing voice recognition technology has a low accuracy of user voice recognition due to the low degree of association with the user, and can improve the accuracy of user voice recognition.

[0063] In one embodiment, it further includes:

[0064] Construct a fuzzy sound association database for the target user based on the user database;

[0065] Interact with the configuration sensitivity of the target user, and generate fuzzy matching constraints for the fuzzy sound association database based on the configuration sensitivity;

[0066] Control the fuzzy sound association database to perform initial extended recognition of the user voice data through the fuzzy matching constraints;

[0067] Perform subsequent multi-channel voice recognition according to the initial extended recognition result. Extract the historical fuzzy sounds and historical fuzzy sound recognition results of the target user in the user database, and construct a fuzzy sound association database for the target user according to the historical fuzzy sounds and historical fuzzy recognition results. The historical fuzzy sounds refer to syllables or words that are easy to confuse and difficult to distinguish in the historical voice recognition data of the target user. For example: when the user reads as When the user emits the voice of "sì bù sì", the meaning expressed is "shì bù shì".

[0068] Set the configuration sensitivity of the fuzzy sound association database of the target user. The configuration sensitivity refers to the number of fuzzy sound expansions, which can be custom-set by those skilled in the art based on the actual situation. For example: when the configuration sensitivity is set to 2, it can be expanded to When the user emits or , it can be considered that the meaning expressed by the user is Generate fuzzy matching constraints for the fuzzy sound association database according to the configuration sensitivity. The fuzzy matching constraints refer to the expansion quantity limit of fuzzy sounds. For example: when the configuration sensitivity is set to 1, the quantity limit for fuzzy sound expansion is 1, and at this time, the fuzzy sound with the highest occurrence frequency of the word needs to be expanded.

[0069] Initial augmented recognition is performed on the user voice data according to fuzzy matching constraints, and the initial augmented recognition results are added to the fuzzy sound association database. Finally, subsequent multi-channel speech conversion results are recognized according to the initial augmented recognition results in the fuzzy sound association database. By constructing the fuzzy sound association database of the target user, the accuracy and efficiency of fuzzy sound recognition in the target user's speech recognition data can be improved, indirectly improving the accuracy and efficiency of user speech recognition.

[0070] In one embodiment, it further includes:

[0071] Verify the speech recognition result and generate a verification identifier;

[0072] Construct a strong association database through the verification identifier;

[0073] Perform integration compensation for subsequent multi-channel speech conversion results through the strong association database.

[0074] Perform multiple speech recognition tests on the target user to obtain multiple groups of user voice data and multiple groups of user speech recognition results, where the user voice data and the user speech recognition results have a corresponding relationship. Then, send the multiple groups of speech recognition results to the target user for speech recognition result judgment, and generate a verification identifier according to the voice data with a large gap between the voice and the correct recognition result in the correct speech recognition results. For example: when the user makes a sound of "go there", the semantics expressed at this time must be the meaning of "go to the market", without other ambiguities.

[0075] Obtain multiple groups of user voice data with verification identifiers and use them as strong association data, and construct a strong association database according to the multiple groups of strong association data. When the target user performs subsequent integration of multi-channel speech conversion results, at this time, call the strong association data in the strong association database to recognize the multi-channel speech conversion results. By constructing the strong association database, the accuracy and efficiency of strong association data recognition in the target user's voice data can be improved, thereby indirectly improving the accuracy and efficiency of user speech recognition.

[0076] In one embodiment, as Figure 3 shown, a precision speech recognition system based on artificial intelligence is provided, including a user voice data acquisition module, a speech cleaning feature construction module, a user recognition feature set construction module, a multi-channel speech recognition execution module, a semantic recognition aggregation module, and a speech recognition result output module. Among them:

[0077] The user voice data acquisition module is used to acquire the user voice data of the target user and generate a collection environment identifier for the user voice data;

[0078] A voice cleaning feature construction module is used to construct voice cleaning features, which are obtained by calling the user database of the target user and matching the user database with the collection environment identifier;

[0079] A user identification feature set construction module is used to construct a user identification feature set, which is constructed by extracting user features from the user database.

[0080] A multi-channel speech recognition execution module is used to perform multi-channel speech recognition, which is to identify the user speech data after data cleaning based on the speech cleaning features and then through the user recognition feature set.

[0081] A semantic recognition aggregation module is used to perform semantic recognition aggregation, which is obtained by acquiring multi-channel speech conversion results and aggregating the multi-channel speech conversion results.

[0082] A speech recognition result output module is used to integrate the multi-channel speech conversion results based on the semantic recognition aggregation results, and output the speech recognition result based on the integrated result.

[0083] In one embodiment, the system further includes:

[0084] A basic recognition channel construction module is used to construct a basic recognition channel, which is constructed based on the language features extracted from the user's recognition feature set.

[0085] A statement analysis channel construction module is used to construct a statement analysis channel, which is constructed based on the statement usage features extracted from the user identification feature set.

[0086] A first recognition channel construction module is used to construct a first recognition channel, which is obtained by mixing the basic recognition channel with the statement analysis channel.

[0087] The first speech conversion result output module is used to perform channel recognition on the cleaned user speech data through the first recognition channel and output the first speech conversion result.

[0088] A multi-channel speech conversion result acquisition module is used to obtain the multi-channel speech conversion result based on the first speech conversion result.

[0089] In one embodiment, the system further includes:

[0090] A part-of-speech aggregation channel configuration module is used to configure a part-of-speech aggregation channel, which is obtained by extracting the part-of-speech features of the user based on the user identification feature set.

[0091] The second recognition channel construction module is used to construct a second recognition channel, which is formed by mixing the part-of-speech aggregation channel with the basic recognition channel.

[0092] The second speech conversion result output module is used to perform channel recognition on the cleaned user speech data through the second recognition channel and output the second speech conversion result.

[0093] A multi-channel speech conversion result acquisition module is used to obtain the multi-channel speech conversion result based on the first speech conversion result and the second speech conversion result.

[0094] In one embodiment, the system further includes:

[0095] A pause feature extraction module is used to extract pause features of a user based on the user identification feature set, wherein the pause features are accompanied by a pause confidence identifier;

[0096] The third recognition channel construction module is used to construct a third recognition channel, which is obtained by constructing a pause aggregation channel through the pause feature and mixing the basic recognition channel with the pause aggregation channel;

[0097] The third speech conversion result output module is used to perform the third recognition channel to recognize the user speech data channel after data cleaning, and output the third speech conversion result.

[0098] A multi-channel speech conversion result acquisition module is used to obtain the multi-channel speech conversion result based on the first speech conversion result, the second speech conversion result, and the third speech conversion result.

[0099] In one embodiment, the system further includes:

[0100] A strong-weak association identification module, which is used to identify the strong-weak association of the third speech conversion result based on the pause confidence identifier.

[0101] An aggregation optimization module is used to optimize the aggregation process by using the strong and weak association identifier when performing semantic recognition aggregation through the first speech conversion result, the second speech conversion result, and the third speech conversion result.

[0102] A semantic recognition aggregation completion module is used to complete the semantic recognition aggregation based on the aggregation optimization results.

[0103] In one embodiment, the system further includes:

[0104] A fuzzy sound association database construction module is used to construct a fuzzy sound association database for the target user based on the user database.

[0105] A fuzzy matching constraint generation module is used to interact with the target user's configuration sensitivity and generate fuzzy matching constraints for the fuzzy sound association database based on the configuration sensitivity.

[0106] An initial expanded recognition execution module is used to control the fuzzy sound association database to perform initial expanded recognition of the user's voice data through the fuzzy matching constraints;

[0107] A subsequent multi-channel speech recognition module is used to perform subsequent multi-channel speech recognition based on the initial expanded recognition results.

[0108] In one embodiment, the system further includes:

[0109] A verification identifier generation module is used to verify the speech recognition result and generate a verification identifier.

[0110] A strongly correlated database construction module, wherein the strongly correlated database construction module is used to construct a strongly correlated database using the verification identifier;

[0111] An integration compensation module is used to perform integration compensation on the subsequent multi-channel speech conversion results through the strongly correlated database.

[0112] In summary, compared with the prior art, the embodiments of this disclosure have the following technical effects:

[0113] 1. By setting up multiple channels based on user characteristics to recognize and convert the user's voice data, and integrating the multi-channel voice conversion results according to semantic aggregation results, the voice recognition result of the target user can be output, which can improve the accuracy of user voice recognition.

[0114] 2. By constructing a pause credibility analysis model based on neural networks, the accuracy and efficiency of pause credibility identification in user speech data can be improved, thereby improving the accuracy and efficiency of user speech data recognition.

[0115] 3. By setting a reliable pause identifier and optimizing the semantic recognition aggregation process of multi-channel speech conversion results based on the reliable pause identifier, the accuracy of the semantic recognition aggregation results can be improved, thereby improving the accuracy of user speech data recognition.

[0116] 4. By constructing a fuzzy sound association database for target users, the accuracy and efficiency of fuzzy sound recognition in target user speech recognition data can be improved; by constructing a strongly associated database, the accuracy and efficiency of strongly associated data recognition in target user speech data can be improved, thereby further improving the accuracy and efficiency of user speech recognition.

[0117] The embodiments described above are merely illustrative of several implementations of this disclosure and should not be construed as limiting the scope of the invention. Therefore, those skilled in the art can make various types of substitutions, modifications, and alterations without departing from the scope of the concept as defined by the appended claims, and all such substitutions, modifications, and alterations fall within the protection scope of this disclosure.

Claims

1. A precise speech recognition method based on artificial intelligence, characterized in that, The method includes: Collect user voice data of the target user and generate a collection environment identifier for the user voice data; Constructing voice cleaning features, wherein the voice cleaning features are obtained by calling the user database of the target user and matching the user database with the collection environment identifier; Construct a user identification feature set, which is constructed by extracting user features from the user database; Perform multi-channel speech recognition, which is to clean the user's speech data based on the speech cleaning features, and then recognize the cleaned user's speech data through the user recognition feature set. Semantic recognition aggregation is performed, which is obtained by acquiring multi-channel speech conversion results and aggregating the multi-channel speech conversion results; The multi-channel speech conversion results are integrated based on the semantic recognition aggregation results, and the speech recognition results are output based on the integration results. The method further includes: A basic recognition channel is constructed, which is built based on the language features extracted from the user's recognition feature set and the language feature extraction results. A statement analysis channel is constructed, which is built based on the statement usage features extracted from the user identification feature set. A first recognition channel is constructed, wherein the first recognition channel is obtained by mixing the basic recognition channel with the sentence analysis channel; The user's voice data after data cleaning is processed by the first recognition channel to perform channel recognition and output the first voice conversion result. The multi-channel speech conversion result is obtained based on the first speech conversion result; The method further includes: Configure a part-of-speech aggregation channel, wherein the part-of-speech aggregation channel is obtained by extracting the user's part-of-speech features based on the user identification feature set; A second recognition channel is constructed, which is formed by mixing the part-of-speech aggregation channel with the basic recognition channel; The user's voice data after data cleaning is processed by the second recognition channel to perform channel recognition and output the second voice conversion result. The multi-channel speech conversion result is obtained based on the first speech conversion result and the second speech conversion result; The method further includes: Based on the user identification feature set, pause features of the user are extracted, wherein the pause features are accompanied by a pause credibility identifier; A third recognition channel is constructed, which is obtained by constructing a pause aggregation channel through the pause features and mixing the basic recognition channel with the pause aggregation channel; The third recognition channel is used to recognize the cleaned user voice data channel, and the third voice conversion result is output. The multi-channel speech conversion result is obtained based on the first speech conversion result, the second speech conversion result, and the third speech conversion result; The method further includes: A strong or weak correlation identifier for converting the third speech conversion result based on the pause credibility identifier; When the semantic recognition aggregation is performed using the first speech conversion result, the second speech conversion result, and the third speech conversion result, the aggregation process is optimized using the strong and weak association identifiers. The semantic recognition aggregation is completed based on the aggregation optimization results.

2. The method as described in claim 1, characterized in that, The method further includes: Construct a fuzzy sound association database for the target user based on the user database; The target user's configuration sensitivity is interacted with, and fuzzy matching constraints for the fuzzy sound association database are generated based on the configuration sensitivity; The fuzzy matching constraint controls the fuzzy sound association database to perform the initial augmentation recognition of the user's voice data; Subsequent multi-channel speech recognition is performed based on the initial expanded recognition results.

3. The method as described in claim 1, characterized in that, The method further includes: Verify the speech recognition result and generate a verification identifier; A strongly associated database is constructed using the verification identifier; Integration compensation is performed by integrating the results of subsequent multi-channel speech conversion using the strongly correlated database.

4. A precise speech recognition system based on artificial intelligence, characterized in that, The system comprises the following steps for performing any one of the artificial intelligence-based precision speech recognition methods described in claims 1-3: The user voice data acquisition module is used to acquire user voice data of the target user and generate an acquisition environment identifier for the user voice data. A voice cleaning feature construction module is used to construct voice cleaning features, which are obtained by calling the user database of the target user and matching the user database with the collection environment identifier; A user identification feature set construction module is used to construct a user identification feature set, which is constructed by extracting user features from the user database. A multi-channel speech recognition execution module is used to perform multi-channel speech recognition, which is to identify the user speech data after data cleaning based on the speech cleaning features and then through the user recognition feature set. A semantic recognition aggregation module is used to perform semantic recognition aggregation, which is obtained by acquiring multi-channel speech conversion results and aggregating the multi-channel speech conversion results. A speech recognition result output module is used to integrate the multi-channel speech conversion results based on the semantic recognition aggregation results, and output the speech recognition result based on the integrated result.

Citation Information

Patent Citations

  • Audio data processing method and device, storage medium and equipment

    CN111554300A

  • System and method for speech understanding via integrated audio and visual based speech recognition

    CN112204564A