A natural language data analysis method based on speech recognition

By segmenting and extracting features from the initial speech data, and combining time sequence and system function operation form matching, the accuracy problem of speech recognition technology under noise and non-standard pronunciation is solved, and high-accuracy user intent parsing is achieved in noisy environments.

CN120260553BActive Publication Date: 2025-11-07HANGZHOU HONGXIONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510579142.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-11-07
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing speech recognition technologies lack accuracy and generalization ability in the face of environmental noise interference and non-standard pronunciation, resulting in inaccurate recognition results.

Method used

By segmenting and extracting features from the initial speech data, analyzing and parsing user intent using intermediate speech data, and combining time sequence and system function operation form matching, the most credible user intent is selected to improve recognition accuracy.

Benefits of technology

It improves the accuracy of natural language recognition in noisy environments, ensuring the executability of user intentions and the correctness of recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260553B_ABST
    Figure CN120260553B_ABST
Patent Text Reader

Abstract

The application discloses a natural language data analysis method based on voice recognition, and relates to the technical field of voice recognition.The method comprises the following steps: collecting voice signals to obtain initial voice data; performing voice recognition on the initial voice data based on voice recognition technology to obtain initial recognition data; performing user intention analysis on the initial recognition data; if it is determined that the user intention cannot be understood, segmenting the initial voice data based on the initial recognition data to obtain intermediate voice data, analyzing the initial voice data based on the intermediate voice data to determine intermediate recognition data, performing user intention analysis on the intermediate recognition data to determine whether the system can correctly understand the user intention; if it is determined that the system can correctly understand the user intention, performing an operation corresponding to the user intention; and if it is determined that the system cannot correctly understand the user intention, outputting an alarm signal to remind the user to input again.The application has the effect of improving the recognition accuracy of voice recognition technology on natural language.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, in particular to a natural language data analysis method based on speech recognition. BACKGROUND

[0002] Speech recognition technology is a technology that converts human speech signals into computer-readable text or instructions. It analyzes and understands speech signals and converts them into text or commands to achieve human-computer interaction. For example, intelligent interaction function in vehicle driving, through reading the voice of the driver, to provide functional assistance for the driver, reduce the degree of distraction of the driver.

[0003] In related technologies, the collected speech information is processed by noise reduction filtering to improve information quality, and the processed speech information is recognized according to the established natural language understanding model to analyze the user's intention. However, in the prior art, when the collected speech information has environmental noise interference, the recognition accuracy of speech recognition will be reduced; when there is non-standard pronunciation, it will also cause weak generalization ability, which needs to be improved. SUMMARY

[0004] In order to improve the recognition accuracy of natural language by speech recognition technology, the present application provides a natural language data analysis method based on speech recognition.

[0005] The present application provides a natural language data analysis method based on speech recognition, which adopts the following technical scheme:

[0006] A natural language data analysis method based on speech recognition comprises:

[0007] Step S1, collecting and denoising the speech signal to obtain initial speech data;

[0008] Step S2, performing speech recognition on the initial speech data based on speech recognition technology to obtain initial recognition data;

[0009] Step S3, performing user intention analysis on the initial recognition data to determine whether the system can correctly understand the user's intention;

[0010] Step S4, if it is determined that the user's intention cannot be understood, segmenting the initial speech data based on the initial recognition data to obtain intermediate speech data, and analyzing the initial speech data based on the intermediate speech data to determine intermediate recognition data, and performing user intention analysis on the intermediate recognition data to determine whether the system can correctly understand the user's intention;

[0011] Step S5, if the system can correctly understand the user's intention, the operation corresponding to the user's intention is executed, if the system cannot correctly understand the user's intention, an alarm signal is output to remind the user to input again.

[0012] Preferably, step S41, the initial recognition data is converted into text data, and each character in the text data is divided separately to obtain text sub-data.

[0013] Step S42, based on the text sub-data, the initial voice data segment corresponding to each text sub-data is determined, and the text sub-data and the initial voice data segment corresponding thereto are grouped to obtain intermediate voice data.

[0014] Preferably, step S43, the initial voice data segment in the intermediate voice data is subjected to sound feature extraction to obtain first voice features.

[0015] Step S44, based on the first voice features, the initial voice data is subjected to sound feature separation to obtain feature voice data containing the first voice features and residual voice data not containing the first voice features.

[0016] Step S45, based on the first voice features, the feature voice data is subjected to voice recognition to obtain intermediate recognition data, and the intermediate recognition data of the plurality of text sub-data is subjected to statistical and user intention analysis to determine whether there is a user intention that the system can understand in the plurality of intermediate recognition data, and the corresponding execution operation is started for the user intention that can be understood.

[0017] Preferably, based on the time sequence, the first voice features corresponding to the character sub-data in the initial recognition data are sorted to obtain a feature sequence set with a time sequence.

[0018] The intermediate recognition data is subjected to user intention analysis, if it is determined that there is a correct user intention, the residual voice data is processed according to the feature sequence set to obtain second feature voice data and second residual voice data in the residual voice data, and the second feature voice data is subjected to voice recognition and user intention analysis until there is no residual voice data or the feature sequence set is matched.

[0019] Preferably, the parsed user intention is obtained, and the user intention is compared with the built-in system function operation form to determine whether there is an operation corresponding to the user intention in the system function operation form, if it is determined that there is, the recognition is correct and the corresponding operation is started, if it is determined that there is not, an alarm signal is output to remind the user to input again.

[0020] Preferably, step S46, if there are a plurality of intermediate voice data that can parse the user intention, the user intention parsed by each intermediate voice data is matched with the built-in system function operation form to determine the matching success item.

[0021] Step S47, compare the number of matching successful items of each intermediate voice data, and take the one with the largest number as the best matching result based on time sequence and execute.

[0022] Preferably, the matching successful items of each intermediate voice data are counted to obtain a first data value, and the user intent recognized by the intermediate voice data is counted to obtain a second data value.

[0023] The first data value and the second data are calculated by ratio, and the user intent with the largest ratio data is taken as the best recognition result, and the corresponding operation is performed according to the system function operation of the corresponding matching successful item.

[0024] Preferably, step S48, the system owner is prompted to input voice to obtain first priority voice data, and the first priority voice data is subjected to sound feature extraction to obtain first priority voice feature value.

[0025] Step S49, when the intermediate voice data is subjected to voice recognition, the first voice feature is matched with the first priority voice feature value, if the matching is successful, the user intent analyzed by the first priority voice feature value is taken as the first priority execution operation.

[0026] Preferably, based on the first priority voice data and the corresponding original text data, the first priority voice data is subjected to voice distortion to obtain distorted voice data, and the distorted voice data is subjected to voice recognition to obtain distorted recognition text.

[0027] The distorted recognition text is matched with the original text data, the distorted voice data that can be correctly recognized is taken as effective distorted voice data, and the effective distorted voice is subjected to sound feature extraction to obtain a first priority voice feature value set.

[0028] In summary, the present application has at least one of the following beneficial technical effects:

[0029] 1. By utilizing existing technologies to acquire, process, and recognize speech signals, the recognition results achieve the accuracy achievable by current technologies. User intent parsing is then performed on the initial recognition data to determine its correctness. If incorrect, the initial speech data is segmented based on the initial recognition data. The resulting intermediate-level speech data is then analyzed, ensuring that all intermediate-level recognition data is derived from this intermediate-level data. This reduces the impact of noise on the intermediate-level recognition data, thereby improving its accuracy. Furthermore, user intent parsing of the intermediate-level recognition data further enhances the accuracy of the parsing results, ultimately improving the speech recognition technology's accuracy in recognizing natural language.

[0030] 2. By converting the initial recognition data into text data, the segmentation becomes more accurate. By extracting sound features from the segmented intermediate speech data and using these sound features as a reference to separate the initial speech data, the feature speech data contains only the target sound features, improving speech quality and making the recognition results based on the feature speech data more accurate. User intent parsing is then performed on the recognized intermediate recognition data, making the judgment of the parsing results more correct and improving the accuracy of speech recognition technology in recognizing natural language.

[0031] 3. By matching the identified user intent with the system's function operation forms, the executable intent items in the identified user intent are determined. By counting the number of each executable intent based on intermediate speech data, and combining the time sequence and the proportion of the corresponding executable intent in the intermediate speech data for comprehensive judgment, the credibility of the user intent in the final selected intermediate speech data is maximized, thereby improving the accuracy of speech recognition in recognizing natural language. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating the steps of the natural language data analysis method based on speech recognition in this embodiment. Detailed Implementation

[0033] The following is in conjunction with the appendix Figure 1 This application will be described in further detail.

[0034] This application discloses a natural language data analysis method based on speech recognition.

[0035] Example: Figure 1 As shown, the present invention provides a natural language data analysis method based on speech recognition, comprising:

[0036] Step S1, collecting and denoising the voice signal to obtain initial voice data;

[0037] Step S2, performing voice recognition on the initial voice data based on voice recognition technology to obtain initial recognition data;

[0038] Step S3, performing user intention analysis on the initial recognition data to determine whether the system can correctly understand the user intention;

[0039] Step S4, if it is determined that the user intention cannot be understood, segmenting the initial voice data based on the initial recognition data to obtain intermediate voice data, analyzing the initial voice data based on the intermediate voice data to determine intermediate recognition data, and performing user intention analysis on the intermediate recognition data to determine whether the system can correctly understand the user intention;

[0040] Step S5, if it is determined that the system can correctly understand the user intention, performing an operation corresponding to the user intention, and if it is determined that the system cannot correctly understand the user intention, outputting an alarm signal to remind the user to input again.

[0041] In this embodiment, the voice signal is collected, processed and recognized by using the existing technology, so that the recognition result has the accuracy that can be achieved by the existing technology. Then, the user intention analysis is performed on the initial recognition data to determine the correctness of the initial recognition data. When it is determined that the initial recognition data is incorrect, the initial voice data is segmented based on the initial recognition data, and the initial voice data is analyzed based on the segmented intermediate voice data. Therefore, the intermediate recognition data obtained by the analysis is based on the intermediate voice data, which reduces the influence of the chaotic signal on the intermediate recognition data, and improves the accuracy of the intermediate recognition data. Then, the user intention analysis is performed on the intermediate recognition data, so that the analysis result is more accurate, and the recognition accuracy of the voice recognition technology for natural language is improved.

[0042] Exemplarily, in processing natural language data by using voice recognition technology, the voice signal is collected and processed by using the existing technology, so as to improve the quality of the natural language data, and the processed voice is recognized by using the existing voice recognition technology for the first time, so as to determine whether the existing technology can correctly recognize the user's intention of the voice signal. When the existing technology cannot accurately recognize, the initial recognition data obtained by the first recognition is used as a reference to segment the initial voice data. Since the initial recognition data is obtained by the existing technology, the recognition result has accuracy but not necessarily correctness. Each recognized character has its corresponding voice segment, so the independent intermediate voice data is obtained by extracting the voice segment, and the extracted intermediate voice data is used as a recognition basis to recognize and analyze the initial voice data, so as to improve the accuracy of recognition. Then, the user's intention of the intermediate recognition data is analyzed, so as to further determine the correctness of recognition, and ensure the executability of the user's intention.

[0043] Wherein, since the voice input is performed, one sound corresponds to multiple characters, so that when the characters are matched according to the sound, only the correctness of the sound can be determined, that is, the accuracy of the recognition result, and the correctness of the matched characters, that is, the correctness of the recognition result, cannot be determined.

[0044] In step S4, if it is determined that the user's intention cannot be understood, the initial voice data is segmented based on the initial recognition data to obtain intermediate voice data, and the initial voice data is analyzed based on the intermediate voice data to determine intermediate recognition data, and the user's intention of the intermediate recognition data is analyzed to determine whether the system can correctly understand the user's intention, including the following steps:

[0045] In step S41, the initial recognition data is converted into text data, and each character in the text data is divided separately to obtain text sub-data;

[0046] In step S42, the initial voice data segment corresponding to each text sub-data is determined based on the text sub-data, and the text sub-data and the corresponding initial voice data segment are grouped to obtain intermediate voice data.

[0047] In step S43, the initial voice data segment in the intermediate voice data is extracted to obtain first voice features;

[0048] In step S44, the initial voice data is separated based on the first voice features to obtain feature voice data containing the first voice features and residual voice data not containing the first voice features;

[0049] In step S45, the voice recognition is performed on the feature voice data based on the first voice feature to obtain intermediate recognition data, and the intermediate recognition data of the plurality of text sub-data is counted and user intention is analyzed to determine whether there is a user intention that can be understood by the system in the plurality of intermediate recognition data, and the corresponding execution operation is started for the user intention that can be understood.

[0050] In the embodiment, the initial recognition data is converted into text data to make the division more accurate, the sound feature is extracted from the divided intermediate voice data, and the voice feature is separated from the initial voice data based on the sound feature, so that the feature voice data only contains the target sound feature, the voice quality is improved, and the result obtained by recognizing the feature voice data is more accurate. The user intention analysis of the intermediate recognition data after recognition makes the judgment of the analysis result more correct, and improves the recognition accuracy of the voice recognition technology for natural language.

[0051] For example, when it is determined that the initial recognition data does not have an understandable user intention, the initial recognition data is converted into text data, and then the voice segment corresponding to each character in the text data is determined according to the corresponding character, for example, the voice data is "open **##", the text data is "open **", the voice data is segmented according to the text data, and the voice segments are "hit", "open", "*", "*", "#", "#", respectively. The sound feature of the voice segment "hit", "open", "*", "*", "#", "#" is extracted to obtain the corresponding voice feature, for example, the corresponding voice feature is "a", "b", "c", respectively. The initial voice data is separated according to the voice feature "a", "b", "c", for example, the voice data "open **##" is 5 seconds in total, the sound feature of 1-2 seconds is "a", the sound feature of 2-4 seconds is "b+other", and the sound feature of the 5th second is "c+other". According to the voice feature "a", 1-2 seconds is "a" itself, which is extracted, and it is determined whether "b+other" in 2-4 seconds can be converted into "a+other". If it is determined that it can be converted, the "a" part is extracted, and the same operation is performed on the 5th second, and then the feature voice data composed of "a" is obtained, and the feature voice data is re-recognized and analyzed to determine whether the user intention can be analyzed. The voice feature "b" and the voice feature "c" are analyzed in the same way, the interference of noise on the user voice signal is reduced, and the recognition accuracy is improved.

[0052] In step S45, speech recognition is performed on the feature speech data based on the first speech features to obtain intermediate recognition data, and the intermediate recognition data of the plurality of text sub-data is counted and user intention is analyzed to determine whether there is a user intention that can be understood by the system in the plurality of intermediate recognition data, and the corresponding execution operation is started for the user intention that can be understood, including the following steps:

[0053] In step S451, the first speech features corresponding to the text sub-data in the initial recognition data are sorted based on the time sequence to obtain a feature sequence set with a time sequence.

[0054] In step S452, the user intention of the intermediate recognition data is analyzed, and if it is determined that there is a correct user intention, the remaining speech data is processed according to the feature sequence set to obtain second feature speech data and second remaining speech data in the remaining speech data, and the second feature speech data is subjected to speech recognition and user intention analysis until there is no remaining speech data or the feature sequence set is matched.

[0055] In this embodiment, the first speech features that need to be determined are sorted by using the time sequence to determine the extraction order in the same group of speech feature extraction, which ensures the order of the recognition result, and further makes the recognition result of the same group more complete, and the analysis result of the user intention analysis of the recognition result of the same group is more accurate, thereby improving the accuracy of natural language recognition by speech recognition.

[0056] For example, when performing speech recognition on the feature speech data containing the first speech features, the first speech features are recognized according to the time sequence to ensure the order of the recognition result. Since the user will input corresponding voice instructions to the system after the system is awakened by voice operation function, the earlier the time, the higher the credibility of the input voice instruction. Therefore, when the number of instructions is the same, the credibility of the corresponding group of instructions is determined by the time sequence, thereby improving the correctness of the recognition.

[0057] After the initial speech data is recognized according to the first speech features, the recognized feature speech data is subjected to speech recognition, and when it is determined that the feature speech data corresponding to the first speech features can recognize the user intention, the remaining speech data is extracted based on the next speech feature according to the time sequence, and the extracted feature speech data is subjected to corresponding speech recognition to determine whether the user intention can be recognized, until the speech feature recognition is completed or the remaining speech data is recognized.

[0058] For example, there are first voice features A, B, C, and the feature voice data divided by the first voice features are Aa and noise X, Bb and noise Y. When analyzing the first voice feature A, first perform voice recognition on Aa to determine whether there is a user intent in Aa. When determining that there is, then sort B and C according to time sequence, if B first and C second, then perform voice extraction on noise X with the first voice feature B to obtain the feature voice data XB in noise X corresponding to B and noise Xb. Perform voice recognition on XB to determine whether there is a user intent in XB. If there is a user intent, then continue to perform operation C on Xb. If it is determined that there is no user intent, then perform operation C on noise X.

[0059] In step S4, if it is determined that the user intent cannot be understood, then the initial voice data is segmented based on the initial recognition data to obtain intermediate voice data, the initial voice data is analyzed based on the intermediate voice data to determine intermediate recognition data, and the user intent is parsed based on the intermediate recognition data to determine whether the system can correctly understand the user intent, including the following steps:

[0060] The parsed user intent is obtained, and the user intent is compared with the built-in system function operation table to determine whether there is an operation corresponding to the user intent in the system function operation table. If it is determined that there is, then the recognition is correct and the corresponding operation is started. If it is determined that there is not, then an alarm signal is output to remind the user to input again.

[0061] In this embodiment, after the voice is parsed for the user intent, the correctness of the parsed result is determined, thereby further determining the correctness of the recognition result and improving the recognition accuracy of the voice recognition system for natural language.

[0062] For example, when the voice is parsed for the user intent, the parsed user intent is understood by the voice recognition system, but the voice recognition system can understand and the function operation system can execute do not have a strict comparison relationship. Therefore, after the user intent is recognized, the executability of the user intent needs to be determined to ensure that the function operation system can normally execute the command. For example, the parsed user intent is "open the sunroof", but the actual car function does not have the function of "opening the sunroof", so it is indicated that the intent cannot be normally executed, and it is further indicated that the correctness of the recognition result is low.

[0063] In step S4, if it is determined that the user intent cannot be understood, then the initial voice data is segmented based on the initial recognition data to obtain intermediate voice data, the initial voice data is analyzed based on the intermediate voice data to determine intermediate recognition data, and the user intent is parsed based on the intermediate recognition data to determine whether the system can correctly understand the user intent, including the following steps:

[0064] Step S46, if there are multiple intermediate voice data that can all parse the user intent, then the user intent parsed by each intermediate voice data is matched with the built-in system function operation table to determine the matching success item;

[0065] Step S47, the number of matching success items of each intermediate voice data is compared, and the largest number is taken as the best matching result based on the time sequence and executed.

[0066] Step S471, the matching success items of each intermediate voice data are counted to obtain a first data value, and the user intent recognized by the intermediate voice data is counted to obtain a second data value;

[0067] Step S472, the first data value and the second data are ratio calculated, the user intent with the largest ratio data is taken as the best recognition result, and the corresponding system function operation of the matching success item is performed.

[0068] In this embodiment, the recognized user intent is matched with the system function operation table to determine the executable intent item in the recognized user intent, and the number of each executable intent is counted according to the intermediate voice data, and the time sequence and the proportion of the corresponding executable intent in the intermediate voice data are comprehensively judged, so that the credibility of the user intent of the intermediate voice data filtered by the final judgment reaches the maximum, and the recognition accuracy of the natural language by the voice recognition is improved.

[0069] For example, since a person's language characteristics are fixed, when the user inputs a voice instruction, the user's voice characteristics should also be single and fixed, so when the user intent of the intermediate voice data is parsed, the result of the parsing is also single. When the voice is distorted due to external interference, multiple results may be obtained. Therefore, when the result is multiple, the result needs to be filtered to ensure that the filtered result is most likely to be the same as the truth. The executable function quantity in the user intent is determined by the executable function quantity in the user intent, and the more executable functions in a group of user intents, the more likely it is to cover the user's demand. Similarly, since the user will preferentially input the instruction to be executed when performing voice recognition and voice instruction input, when the function quantity is the same, the time sequence of the corresponding group is earlier, and the possibility of meeting the user's demand is greater.

[0070] Similarly, since the user will prefer to input instructions, and the possibility of dialogue with the dialogue module corresponding to the speech recognition system in a noisy environment is low, the user's input voice is mostly the content of the function instruction, and then the executable intention in the recognized user intention is judged by the proportion, so as to judge the credibility of the user intention recognized by each group as the standard of accuracy, and then improve the accuracy of the recognition result.

[0071] Before starting the system for speech recognition, the following steps are also included:

[0072] Step S48, the system's host user is inputted by voice, and the first priority voice data is obtained, and the first priority voice feature value is extracted from the sound feature, and the first priority voice feature value is obtained;

[0073] Step S49, when the intermediate voice data is recognized by voice, the first voice feature is matched with the first priority voice feature value, if the matching is successful, the user intention analyzed by the first priority voice feature value is executed as the first priority operation.

[0074] In this embodiment, the voice of the host user is recorded and analyzed in advance, so that the voice recognition accuracy of the built-in speech recognition system for the host user is higher than that of other personnel, so when the speech recognition system cannot directly recognize the collected voice signal, the first voice feature analyzed is matched with the first priority voice feature value, so as to determine whether there is the voice of the host user, and the voice feature of the host user is preferentially used as the judgment reference to analyze and execute the user intention, so as to improve the recognition accuracy of the speech recognition for the natural language in the noisy environment.

[0075] For example, when the user starts the system for the first time, the user's voice can be recorded as the first priority for processing to ensure the safety of speech recognition, and the sound feature of the user's voice is extracted to identify the user's voice and preferentially recognize the user's instruction to further improve the accuracy of recognition.

[0076] For example, the application is applied to the intelligent interaction in vehicle driving, the voice of the vehicle owner is preferentially recorded as the basis to ensure the priority of the vehicle owner's instruction, the voice of the vehicle owner is preprocessed to enable the intelligent interaction system to accurately recognize the voice instruction of the vehicle owner, so that when the environment is not noisy, the intelligent interaction system can normally recognize the voice of the vehicle owner or others through voice recognition technology, when others have accent problems and cannot be normally recognized, if the vehicle owner is also present, the vehicle owner will input the instruction again to ensure the executability of the user's intention, and if the vehicle owner is not present, the user is reminded to input again.

[0077] When the environment is noisy, the chaotic sound causes the intelligent interaction system to be unable to normally recognize, and because the voice recognition accuracy of the vehicle owner is higher than that of other personnel, the voice of the vehicle owner is recognized, if the voice of the vehicle owner is recognized, the instruction of the vehicle owner has the first execution right, and it also indicates that the instruction group corresponding to the vehicle owner is more accurate, so the voice of the vehicle owner is taken as the first priority to process the collected voice signal, so that the natural language recognized by the voice recognition is more accurate.

[0078] In step S48, the voice of the system owner is recorded to obtain first-priority voice data, and the first-priority voice data is subjected to sound feature extraction to obtain first-priority voice feature values, and the following steps are further included:

[0079] In step S481, based on the first-priority voice data and the corresponding original text data, the first-priority voice data is subjected to voice distortion to obtain distorted voice data, and the distorted voice data is subjected to voice recognition to obtain distorted recognition text; wherein the voice distortion is to change the sound characteristics such as tone, pitch, volume, etc. of the user's voice.

[0080] In step S482, the distorted recognition text is matched with the original text data, the distorted voice data that can be correctly recognized is taken as effective distorted voice data, and the effective distorted voice is subjected to sound feature extraction to obtain a first-priority voice feature value set.

[0081] In this embodiment, the recorded user voice is subjected to conventional voice distortion, so that the distorted voice data can cover the possible range of user voice changes to a greater extent, and then the voice recognition technology is used to recognize the distorted voice, and the recognized text is compared with the original text to determine the accuracy of the voice recognition technology in recognizing the user voice, and then the distorted voice data is screened to improve the adaptability of the voice recognition technology to the user voice, and thus the recognition accuracy of the voice recognition technology to the user's natural language is improved.

[0082] Exemplarily, since the voice features of a person are basically in a fixed state, a noisy environment only superimposes a state on the voice features of the user, so when separating the collected voice information according to the voice features of the user, the voice of the user and the noise can be effectively separated.

[0083] However, due to different body states of the user, the voice will change to a certain extent, so by using big data, the change value of the voice characteristics under different body states is determined, for example, when sick, the voice will become hoarse, and hoarseness corresponds to the change of the tone, so by evaluating the common tone change range and simulating the original voice of the user according to the tone change range, the corresponding distorted voice data is obtained, and thus the whole voice change interval of the user is covered, thereby increasing the recognition accuracy of the voice recognition system and the natural language of the user.

[0084] Compared with the existing natural language data analysis method based on voice recognition, the application improves the accuracy of voice recognition for natural language recognition.

[0085] The above are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, therefore: any equivalent changes made according to the structure, shape, principle of the present application should be covered within the protection scope of the present application.

Claims

1. A method for natural language data analysis based on speech recognition, characterized in that, The method comprises the following steps: Step S1, collecting and denoising the voice signal to obtain initial voice data; Step S2, performing voice recognition on the initial voice data based on voice recognition technology to obtain initial recognition data; Step S3, performing user intention analysis on the initial recognition data to determine whether the system can correctly understand the user intention; Step S4, if it is determined that the user intention cannot be understood, segmenting the initial voice data based on the initial recognition data to obtain intermediate voice data, analyzing the initial voice data based on the intermediate voice data to determine intermediate recognition data, and performing user intention analysis on the intermediate recognition data to determine whether the system can correctly understand the user intention; Step S41, converting the initial recognition data into text data, and separately dividing each character in the text data to obtain text sub-data; Step S42, determining the initial voice data segment corresponding to each text sub-data based on the text sub-data, and grouping the text sub-data and the initial voice data segment corresponding thereto to obtain intermediate voice data; Step S43, extracting sound features from the initial voice data segment in the intermediate voice data to obtain first voice features; Step S44, separating sound features from the initial voice data based on the first voice features to obtain feature voice data containing the first voice features and residual voice data not containing the first voice features; Step S45, performing voice recognition on the feature voice data based on the first voice features to obtain intermediate recognition data, and performing statistics and user intention analysis on the intermediate recognition data of multiple text sub-data to determine whether there is a user intention that can be understood by the system in the multiple intermediate recognition data, and starting the corresponding execution operation for the user intention that can be understood; Step S451, based on the time sequence, sorting the first voice features corresponding to the character sub-data in the initial recognition data to obtain a feature sequence set with a time sequence; Step S452, performing user intention analysis on the intermediate recognition data, if it is determined that there is a correct user intention, processing the residual voice data according to the feature sequence set to obtain second feature voice data and second residual voice data in the residual voice data, and performing voice recognition and user intention analysis on the second feature voice data until there is no residual voice data or the feature sequence set is matched; Step S5, if it is determined that the system can correctly understand the user intention, performing the operation corresponding to the user intention, if it is determined that the system cannot correctly understand the user intention, outputting an alarm signal to remind the user to input again. 2.The voice recognition based natural language data analysis method of claim 1, wherein: Step S4 further comprises: obtaining the parsed user intention, and comparing the user intention with the built-in system function operation table to determine whether there is an operation corresponding to the user intention in the system function operation table, if it is determined that there is, the recognition is correct and the corresponding operation is started, if it is determined that there is not, an alarm signal is output to remind the user to input again.

3. The natural language data analysis method based on speech recognition according to claim 2, characterized in that: Step S4 further comprises: Step S46, if there are multiple intermediate voice data that can analyze the user intention, matching the user intention analyzed by each intermediate voice data with the built-in system function operation table to determine the matching success item; Step S47, compare the number of matching successful items of each intermediate voice data, and take the one with the largest number as the best matching result based on time sequence and execute.

4. The natural language data analysis method based on speech recognition according to claim 3, characterized in that: Step S47, further comprising: Count the matching successful items of each intermediate voice data to obtain a first data value, and count the user intent recognized by the intermediate voice data to obtain a second data value; Calculate the ratio of the first data value and the second data, take the user intent with the largest ratio data as the best recognition result, and perform the corresponding operation according to the system function operation of the corresponding matching successful item.

5. The natural language data analysis method based on speech recognition according to claim 4, characterized in that: Before starting the system for voice recognition, further comprising: Step S48, perform voice input on the system owner to obtain first priority voice data, and perform sound feature extraction on the first priority voice data to obtain first priority voice feature value; Step S49, when performing voice recognition on the intermediate voice data, match the first voice feature with the first priority voice feature value, if the matching is successful, take the user intent analyzed by the first priority voice feature value as the first priority execution operation.

6. The natural language data analysis method based on speech recognition according to claim 5, characterized in that, Further comprising: Based on the first priority voice data and its corresponding original text data, perform voice deformation on the first priority voice data to obtain distorted voice data, and perform voice recognition on the distorted voice data to obtain distorted recognition text; Match the distorted recognition text with the original text data, take the distorted voice data that can be correctly recognized as effective deformation voice data, and perform sound feature extraction on the effective deformation voice to obtain a first priority voice feature value set.

Citation Information

Patent Citations

  • Intention recognition method and device, storage medium and terminal

    CN110097886A