Intelligent speech recognition method and system based on NLP

By acquiring and aligning the key semantic markers of speech samples, constructing semantic coordinate system and distribution functions, the accuracy and abnormal detection problems of speech recognition technology in complex environments are solved, and more efficient speech recognition and interaction are achieved.

CN120375823AInactive Publication Date: 2025-07-25GUANGZHOU HUANXUN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510514855.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing speech recognition technology lacks recognition accuracy in complex environments, lacks in-depth understanding of semantics, and is difficult to deal with complex speech and abnormal detection, resulting in poor service quality and interaction effects in scenarios such as smart customer service and smart home.

Method used

By obtaining key semantic markers in standard speech samples, generating standard semantic regions, and semantic alignment with actual speech data, constructing semantic coordinate systems and fitting semantic distribution functions, and calculating semantic bias values to judge and identify exceptions.

Benefits of technology

It improves the accuracy and stability of speech recognition, can accurately understand semantics in complex environments, reduce noise interference, and detect abnormalities in a timely manner, improving the interaction quality of smart customer service and smart home.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375823A_ABST
    Figure CN120375823A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent speech recognition, and discloses an intelligent speech recognition method and system based on NLP. The method comprises the following steps: firstly, obtaining a standard voice sample of target voice, determining a key semantic mark and generating a standard semantic region; obtaining actual voice data to be recognized, and generating an actual semantic region; semantic alignment is carried out, an actual semantic coordinate system and a standard semantic coordinate system are constructed, and a semantic distribution function is fitted; and finally, calculating a semantic deviation value and comparing the semantic deviation value with a dynamic threshold value to judge an identification result. The system comprises a sample acquisition and processing module, an actual voice processing module, a semantic alignment module and the like. The NLP technology is utilized, the voice recognition accuracy is improved, complex environment voice can be processed, the semantic understanding ability is enhanced, abnormity can be effectively detected and recognized, and the method has wide application prospects in the fields of intelligent customer service, vehicle-mounted, home furnishing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent speech recognition, and specifically to an intelligent speech recognition method and system based on NLP. Background Art

[0002] In today's digital age, natural language processing (NLP) and speech recognition technologies have been widely penetrated into all fields of people's life and work. From the intelligent speech assistants used in daily life, to the voice command input in in-vehicle navigation systems, and then to the voice control of smart home devices, speech recognition technology has greatly improved the convenience of human-computer interaction. However, the existing speech recognition technologies still face many challenges.

[0003] Most traditional speech recognition methods are based on acoustic models and language models, mainly focusing on the acoustic features of speech and simple grammar rules. In complex real-world environments, the limitations of these methods are becoming increasingly prominent. For example, in a noisy environment, background noise will seriously interfere with the speech signal, resulting in inaccurate extraction of acoustic features, and thus affecting the recognition result. When encountering accents, dialects, or non-standard pronunciations, the recognition accuracy of the model trained based on standard pronunciation will drop significantly.

[0004] Existing speech recognition systems often lack in-depth understanding of semantics. They simply convert speech into text and cannot truly understand the meaning behind the speech content. In the scenario of intelligent customer service, the questions of customers may be expressed in various ways. If the speech recognition system cannot accurately understand the semantics, it cannot provide effective answers, reducing the service quality. In smart home control, vague or incomplete voice commands may cause the system to perform incorrect operations, bringing troubles to users.

[0005] With the development of artificial intelligence technology, although some speech recognition models based on deep learning have made certain progress, there are still problems of insufficient utilization of context information. The semantics in speech are often closely related to the context. Current models are difficult to fully capture and utilize this context dependence, resulting in poor performance when dealing with long text speech or speech with complex semantics.

[0006] Moreover, existing speech recognition systems also have defects in recognition anomaly detection. They usually lack effective mechanisms to judge whether the recognition result is accurate. Even if an error occurs in recognition, it may not be detected and corrected in time, which is unacceptable in some application scenarios with extremely high accuracy requirements. In summary, it is of great practical significance to develop an intelligent speech recognition method and system based on NLP that can overcome the above problems. Summary of the Invention

[0007] The object of the present invention is to provide an intelligent speech recognition method and system based on NLP to solve the problems raised in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solution: An intelligent speech recognition method based on NLP, the method includes:

[0009] Obtain at least one standard speech sample of the target speech, determine the key semantic markers in the standard speech sample, the number of the key semantic markers is at least three, and they are distributed at different positions in the speech context;

[0010] Extract the semantic intervals corresponding to the key semantic markers in the standard speech sample to generate a standard semantic region;

[0011] Obtain the actual speech data to be recognized, detect the key semantic markers in the actual speech data to generate an actual semantic region;

[0012] Perform semantic alignment between the standard semantic region and the corresponding actual semantic region;

[0013] Select three non - collinear actual semantic regions to construct an actual semantic coordinate system, and construct a standard semantic coordinate system through the corresponding standard semantic regions;

[0014] In the actual semantic coordinate system, fit the semantic distribution function of the actual speech data, and in the standard semantic coordinate system, fit the semantic distribution function of the standard speech sample;

[0015] Calculate the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function, and determine the dynamic threshold of the semantic deviation value;

[0016] When the semantic deviation value exceeds the dynamic threshold, it is determined that there is an abnormal recognition of the actual speech data, otherwise it is determined to be normal.

[0017] Preferably, the determination of the key semantic markers in the standard speech sample includes the following steps:

[0018] Perform frame - by - frame processing on the standard speech sample, and extract the acoustic feature vectors of each frame of speech;

[0019] Based on the clustering result of the acoustic feature vectors, filter out the frames with a frequency distribution dispersion higher than a preset value as candidate semantic frames;

[0020] In the candidate semantic frames, determine at least three semantic nodes across sentence boundaries through context - dependent relationship analysis as the key semantic markers.

[0021] Preferably, the generation of the standard semantic region includes the following steps:

[0022] Perform dynamic time warping on the speech segment where the key semantic marker is located to generate a standard time window;

[0023] Within the standard time window, extract the sequence of associated words before and after the key semantic marker to form a standard context chain;

[0024] Map the standard context chain to a continuous region in the multi-dimensional semantic space as the standard semantic region.

[0025] Preferably, the generation of the actual semantic region includes the following steps:

[0026] Perform noise suppression processing on the actual speech data to generate an enhanced speech signal;

[0027] Based on the enhanced speech signal, use a sliding window to traverse and detect the acoustic features of the key semantic marker;

[0028] By dynamically adjusting the window length, capture the boundaries of the key semantic marker on the time axis to form an actual context chain;

[0029] Convert the actual context chain into a discrete region in the multi-dimensional semantic space as the actual semantic region.

[0030] Preferably, the semantic alignment includes the following steps:

[0031] Perform semantic vector quantization encoding on the standard semantic region and the actual semantic region respectively;

[0032] Calculate the similarity matrix between the standard semantic vector and the actual semantic vector based on the attention mechanism;

[0033] In the similarity matrix, screen the pair of semantic vectors with the highest similarity and establish a one-to-one mapping relationship;

[0034] According to the mapping relationship, perform interpolation compensation on the unmatched semantic regions to complete the alignment.

[0035] Preferably, the construction of the actual semantic coordinate system includes the following steps:

[0036] Select three actual semantic regions as the first reference area, the second reference area, and the third reference area;

[0037] Take the semantic center of the first reference area as the origin, the line connecting the semantic centers of the first reference area and the second reference area as the first coordinate axis, and the line connecting the semantic centers of the first reference area and the third reference area as the second coordinate axis;

[0038] Normalize and generate the unit scale of the actual semantic coordinate system according to the semantic spans of the first coordinate axis and the second coordinate axis.

[0039] Preferably, the construction of the standard semantic coordinate system includes the following steps:

[0040] Take the standard semantic regions corresponding to the first reference region, the second reference region, and the third reference region as the first standard region, the second standard region, and the third standard region respectively;

[0041] Take the semantic center of the first standard region as the origin, the line connecting the semantic centers of the first standard region and the second standard region as the reference first axis, and the line connecting the semantic centers of the first standard region and the third standard region as the reference second axis;

[0042] Dynamically adjust the dimensional weights of the standard semantic coordinate system according to the semantic correlation degree of the reference first axis and the reference second axis.

[0043] Preferably, the fitting of the semantic distribution function includes the following steps:

[0044] In the actual semantic coordinate system, extract the boundary point set of the actual semantic region and generate the actual semantic distribution function through polynomial interpolation;

[0045] In the standard semantic coordinate system, extract the boundary point set of the standard semantic region and generate the standard semantic distribution function through piecewise linear fitting;

[0046] Among them, the extraction of the boundary point set is based on the gradient change threshold of semantic density.

[0047] Preferably, the determination of the dynamic threshold includes the following steps:

[0048] Randomly combine multiple standard speech samples to generate a mixed sample set;

[0049] In the mixed sample set, calculate the local deviation integral between the standard semantic distribution functions pairwise;

[0050] Statistically calculate the maximum value of all local deviation integrals as the initial value of the dynamic threshold;

[0051] Perform adaptive weighted correction on the initial value according to the noise level of the actual speech environment.

[0052] Preferably, the present invention further includes an intelligent speech recognition system based on NLP, and the system includes:

[0053] Sample acquisition and processing module: used to obtain at least one standard speech sample of the target speech, determine the key semantic markers in the standard speech sample, the number of the key semantic markers is at least three and they are distributed at different positions in the speech context, and extract the semantic intervals corresponding to the key semantic markers in the standard speech sample to generate a standard semantic region;

[0054] Actual speech processing module: used to obtain the actual speech data to be recognized, detect the key semantic markers in the actual speech data, and generate an actual semantic region;

[0055] Semantic alignment module: semantically align the standard semantic region with the corresponding actual semantic region;

[0056] Coordinate system construction module: select three non-collinear actual semantic regions to construct an actual semantic coordinate system, and construct a standard semantic coordinate system through the corresponding standard semantic regions;

[0057] Semantic function fitting module: in the actual semantic coordinate system, fit the semantic distribution function of the actual speech data, and in the standard semantic coordinate system, fit the semantic distribution function of the standard speech sample;

[0058] Deviation judgment module: calculate the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function, and determine the dynamic threshold of the semantic deviation value. When the semantic deviation value exceeds the dynamic threshold, it is determined that there is an abnormal recognition in the actual speech data, otherwise it is determined to be normal.

[0059] Compared with the prior art, the beneficial effects of the present invention are:

[0060] The present invention has many significant beneficial effects:

[0061] In terms of improving the accuracy of speech recognition, by determining the key semantic markers in the standard speech sample and generating a standard semantic region, and at the same time detecting the key semantic markers in the actual speech data to be recognized to generate an actual semantic region, this semantic-based processing method can grasp the speech content more accurately compared with the traditional method that only relies on acoustic features. For example, in the intelligent customer service scenario, in the face of the complex and diverse expressions of customers, the system can quickly locate the core problems through the key semantic markers, avoid misjudgments caused by similar but different semantics in speech, greatly improve the recognition accuracy, and provide more accurate services for customers.

[0062] In dealing with speech in complex environments, noise suppression is performed on actual speech data, effectively reducing the interference of environmental noise on speech signals. In a noisy in-vehicle environment, even with engine noise, road noise, etc., the system can generate enhanced speech signals, accurately detect key semantic markers, ensure the stable operation of the in-vehicle speech system, and enable the driver to smoothly control functions such as navigation and music playback through speech, enhancing the driving experience and safety.

[0063] The improvement in semantic understanding and processing capabilities is a major highlight. By semantic alignment, constructing a semantic coordinate system, and fitting a semantic distribution function, the system can deeply understand the semantics of speech and its context relationship. In the smart home control scenario, when the user issues a vague or incomplete instruction, such as "Make the lights in the living room brighter", the system can accurately understand the intention through semantic analysis, avoid incorrect operations, and achieve a more intelligent and user-friendly interaction.

[0064] In anomaly detection for recognition, the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function is calculated, and a dynamic threshold is determined. This enables the system to promptly detect anomalies during the recognition process. In the intelligent speech writing assistance system, if there is a recognition error in the speech input, the system can quickly detect it and prompt the user to ensure the accuracy of the output text and improve work efficiency.

[0065] The system also has good adaptability. The dynamic threshold is adaptively weighted and corrected according to the noise level of the actual speech environment, and can maintain stable performance in different environments. Whether in a quiet indoor environment or a noisy outdoor environment, it can accurately recognize speech, expand the application range, and provide reliable support for speech interaction in more fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is the working principle diagram of the intelligent speech recognition method described in the present invention;

[0067] Figure 2 is the flowchart of semantic alignment;

[0068] Figure 3 is the flowchart of constructing a standard semantic coordinate system;

[0069] Figure 4 is the flowchart of fitting a semantic distribution function. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Please refer to Figures 1-4 , the present invention provides a technical solution: an intelligent speech recognition method based on NLP, and the specific implementation solution is as follows:

[0072] Obtain standard speech samples and determine key semantic markers: Obtain at least one standard speech sample of the target speech, and these samples are speech data that are carefully selected and can represent the normal characteristics of the target speech. Determine the key semantic markers in the standard speech samples. The number of key semantic markers is at least three and they are distributed at different positions in the speech context. Key semantic markers are parts of the speech content with important semantic information, and their reasonable selection is crucial for the accuracy of subsequent speech recognition.

[0073] Generate standard semantic regions: Extract the semantic intervals corresponding to the key semantic markers in the standard speech samples to generate standard semantic regions. By processing the speech segments where the key semantic markers are located and combining the semantic information before and after them, determine a region containing the key semantic information, and this region will be used as the standard for subsequent comparison.

[0074] Obtain actual speech data and generate actual semantic regions: Obtain the actual speech data to be recognized, which is the speech content that needs to be recognized and judged. Detect the key semantic markers in the actual speech data and generate actual semantic regions. The generation process of the actual semantic regions is to analyze and process the actual speech data, find the parts corresponding to the standard semantic markers, and determine their locations.

[0075] Semantic alignment: Align the standard semantic regions with the corresponding actual semantic regions semantically. Through specific algorithms and processing methods, enable the standard semantic regions and the actual semantic regions to be accurately compared at the semantic level, ensuring that the semantic information of both is compared in the same dimension.

[0076] Construct a semantic coordinate system: Select three non-collinear actual semantic regions to construct an actual semantic coordinate system, and construct a standard semantic coordinate system through the corresponding standard semantic regions. The construction of the semantic coordinate system provides a spatial framework for subsequent fitting of the semantic distribution function, facilitating the quantification and analysis of semantic information.

[0077] Fitting semantic distribution function: In the actual semantic coordinate system, fit the semantic distribution function of actual speech data; in the standard semantic coordinate system, fit the semantic distribution function of standard speech samples. By fitting the semantic distribution function, semantic information can be transformed into a mathematical model to more intuitively display the distribution characteristics of semantics.

[0078] Judging the recognition result: Calculate the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function, and determine the dynamic threshold of the semantic deviation value. When the semantic deviation value exceeds the dynamic threshold, it is determined that there is an abnormal recognition of the actual speech data, otherwise it is determined to be normal. According to the comparison result between the semantic deviation value and the dynamic threshold, the recognition situation of the actual speech data is obtained to judge whether it conforms to the characteristics of the standard speech sample.

[0079] The present invention will be further described below in conjunction with Embodiments 1 to 6:

[0080] Embodiment 1:

[0081] In an actual application scenario, for example, in an intelligent customer service system, high requirements are placed on the accuracy and reliability of speech recognition. First, for the step of determining the key semantic markers in the standard speech sample, the standard speech sample is frame-divided. Assuming the duration of the standard speech sample is T seconds, it is frame-divided according to a certain frame length Δt, so that the speech sample is divided into multiple short frames. Each frame of speech has unique acoustic features, and these features are characterized by extracting the acoustic feature vectors of each frame of speech. The acoustic feature vectors contain information such as the frequency and amplitude of the speech.

[0082] Based on these acoustic feature vectors, clustering analysis is performed. The clustering algorithm will group frames with similar acoustic features into one category. In the clustering result, frames with a frequency distribution dispersion higher than a preset value are selected. The frequency distribution dispersion can be measured by calculating the standard deviation of the acoustic feature frequencies within the frame, etc. The preset value is determined based on a large number of experiments and experiences. These selected frames are used as candidate semantic frames, and they are more likely to contain key semantic information.

[0083] In the candidate semantic frames, the key semantic markers are determined through context-dependent relationship analysis. For example, using the dependency syntactic analysis technology in natural language processing, analyze the syntactic and semantic relationships between words in the candidate semantic frames. If a word has a close semantic association with other words in multiple sentences and this association crosses sentence boundaries, then the semantic node corresponding to this word is determined as a key semantic marker. Ensure that there are at least three key semantic markers and they are distributed at different positions to fully cover the semantic content of the speech.

[0084] When generating the standard semantic region, dynamic time warping is performed on the speech segment where the key semantic marker is located. The dynamic time warping algorithm adjusts the time axis according to the actual duration and features of the speech segment to match it with the standard duration or reference pattern, generating a standard time window. Within the standard time window, the sequence of associated words before and after the key semantic marker is extracted. For example, centered on the key semantic marker, N words are taken forward and backward respectively, and these words form a standard context chain, where the value of N is determined according to the actual situation and experimental results.

[0085] The standard context chain is mapped to a continuous region in the multi-dimensional semantic space as the standard semantic region. A word vector model, such as Word2Vec or GloVe, can be used to convert each word into a vector representation, and then these vectors are combined to form a continuous region in the multi-dimensional space. In this way, the standard semantic region contains the semantic information of the key semantic marker and its context, providing an accurate reference standard for subsequent speech recognition.

[0086] Example 2:

[0087] Taking the intelligent in-vehicle voice system as an example, after obtaining the actual speech data to be recognized, due to the existence of various noises in the in-vehicle environment, such as engine noise, road noise, etc., it is necessary to perform noise suppression processing on the actual speech data. A noise suppression algorithm based on deep learning can be used, such as using a deep neural network (DNN) or a convolutional neural network (CNN). These networks can effectively remove noise from the actual speech data and generate an enhanced speech signal by learning a large number of pairs of noisy speech and clean speech samples.

[0088] Based on the enhanced speech signal, a sliding window is used to traverse and detect the acoustic features of the key semantic marker. The size and step length of the sliding window are set according to the speech features and experimental results. For example, the window size is M frames and the step length is S frames, and both M and S are values optimized through multiple experiments. During the sliding window traversal, the acoustic features of the speech signal within each window are extracted and compared with the acoustic feature template of the pre-set key semantic marker.

[0089] When it is detected that there may be a key semantic marker, the window length is dynamically adjusted to capture the boundary of the key semantic marker on the time axis. If it is found that the matching degree between the acoustic features within the window gradually increases and then gradually decreases, the window can be dynamically enlarged or reduced according to the change of the matching degree until the accurate boundary of the key semantic marker is determined, forming an actual context chain.

[0090] Convert the actual context chain into discrete regions in a multi-dimensional semantic space as the actual semantic regions. Similarly, using the word vector model, each word in the actual context chain is converted into a vector representation, and these vectors form a discrete set of points in the multi-dimensional space, and these point sets constitute the actual semantic regions. Through such processing, the actual semantic regions contain the semantic information of the key semantic markers and their contexts in the actual speech data, providing a basis for subsequent comparison with the standard semantic regions.

[0091] Example 3:

[0092] In an intelligent conference voice recording system, semantic alignment is a very crucial step. Semantic vectorization encoding is performed on the standard semantic regions and the actual semantic regions respectively. Taking the use of the BERT (Bidirectional Encoder Representations from Transformers) model as an example, the text content in the standard semantic regions and the actual semantic regions is input into the BERT model, and the BERT model will perform in-depth semantic understanding and encoding on the text and output the corresponding semantic vectors.

[0093] Calculate the similarity matrix between the standard semantic vector and the actual semantic vector based on the attention mechanism. The attention mechanism can enable the model to pay more attention to the important parts in the semantic vectors when calculating the similarity. Suppose there is a standard semantic vector S = [s1, s2, …, s n and an actual semantic vector R = [r1, r2, …, r m , the element M ij of the similarity matrix can be obtained by calculating the cosine similarity between s i and r j and so on, that is where i = 1, 2, …, n, j = 1, 2, …, m.

[0094] In the similarity matrix, screen the semantic vector pairs with the highest similarity and establish a one-to-one mapping relationship. From the similarity matrix, find the column index corresponding to the maximum value in each row, and the row index corresponding to the maximum value in each column. Through these indexes, determine the vector pair with the highest similarity, thereby establishing a mapping relationship between the standard semantic vector and the actual semantic vector.

[0095] For the unmatched semantic regions, perform interpolation compensation according to the established mapping relationship. For example, if there is a standard semantic vector s k that has no directly matching vector in the actual semantic vectors, then an approximate vector can be generated using methods such as linear interpolation based on the vector pairs adjacent to s k and already matched to represent s kComplete semantic alignment for the corresponding parts in the actual semantic region. Through such semantic alignment processing, the differences between the standard semantic region and the actual semantic region can be compared more accurately, providing a reliable basis for subsequent speech recognition judgment.

[0096] Example 4:

[0097] In a smart home voice control system, constructing an actual semantic coordinate system is of great significance. Select three actual semantic regions as the first reference area, the second reference area, and the third reference area. The selection of these three regions should ensure non-collinear distribution to ensure that the constructed coordinate system has good spatial characteristics. For example, in a voice command "Turn on the living room lights and raise the air conditioner temperature", the semantic regions of "Turn on", "lights", and "raise" can be used as the first reference area, the second reference area, and the third reference area respectively.

[0098] Take the semantic center of the first reference area as the origin. The semantic center can be obtained by calculating the average value of all semantic vectors within this semantic region. Assume that the first reference area contains semantic vectors v1, v2, …, v p , and its semantic center Take the line connecting the semantic centers of the first reference area and the second reference area as the first coordinate axis. The direction and length of this line represent the semantic relationship between the two semantic regions. Similarly, take the line connecting the semantic centers of the first reference area and the third reference area as the second coordinate axis.

[0099] According to the semantic spans of the first coordinate axis and the second coordinate axis, normalize to generate the unit scale of the actual semantic coordinate system. The semantic span can be measured by calculating the distance between two semantic centers, such as using the Euclidean distance. Assume that the semantic center of the first reference area is O, the semantic center of the second reference area is A, and the semantic center of the third reference area is B. The semantic span of the first coordinate axis The semantic span of the second coordinate axis where d is the dimension of the semantic vector. By normalizing these semantic spans, for example, dividing them by a fixed constant (such as the maximum semantic span), the unit scale of the actual semantic coordinate system is obtained, enabling the positions and relationships of different semantic regions in the coordinate system to be more accurately quantified.

[0100] Example 5:

[0101] In a smart translation voice system, constructing a standard semantic coordinate system is an important basis for achieving accurate translation. The standard semantic regions corresponding to the first reference area, the second reference area, and the third reference area are used as the first standard area, the second standard area, and the third standard area respectively. These three standard areas are generated based on standard speech samples and are representative semantic regions.

[0102] Taking the semantic center of the first standard area as the origin, similar to constructing an actual semantic coordinate system, the origin position is obtained by calculating the average value of all semantic vectors in the first standard area. Taking the line connecting the semantic centers of the first standard area and the second standard area as the reference first axis, and the line connecting the semantic centers of the first standard area and the third standard area as the reference second axis.

[0103] According to the semantic correlation degrees of the reference first axis and the reference second axis, dynamically adjust the dimensional weights of the standard semantic coordinate system. The semantic correlation degree can be measured by calculating the semantic similarity between two semantic regions, for example, using the cosine similarity between semantic vectors. Assume that the semantic center vector of the first standard area is O s , the semantic center vector of the second standard area is A s , the semantic center vector of the third standard area is B s , the semantic correlation degree of the reference first axis the semantic correlation degree of the reference second axis According to the values of these semantic correlation degrees, different weights are assigned to the two dimensions of the standard semantic coordinate system respectively. If r OA is larger, it indicates that the semantic correlation between the first standard area and the second standard area is closer. Then the weight of the first dimension in the coordinate system can be appropriately increased, and vice versa, so as to realize the dynamic adjustment of the dimensional weights of the standard semantic coordinate system and make the coordinate system more in line with the semantic characteristics of the standard speech samples.

[0104] Embodiment 6:

[0105] In the intelligent speech writing assistance system, fitting the semantic distribution function and determining the dynamic threshold are key steps to ensure the accuracy of speech recognition. In the actual semantic coordinate system, extract the boundary point set of the actual semantic area. The extraction of the boundary point set is based on the gradient change threshold of semantic density. The semantic density can be measured by calculating the distribution density of semantic vectors in a certain area. Assume that in the actual semantic coordinate system, for a small area Ω, the semantic density where n is the number of semantic vectors in the area, and |Ω| is the volume of the area. The gradient of the semantic density represents the change rate of the semantic density in space. When exceeds the preset gradient change threshold, the corresponding point is considered a boundary point, and these boundary points form the boundary point set. Generate the actual semantic distribution function through polynomial interpolation. For example, using the Lagrange interpolation polynomial, for the boundary point set {(x1,y1),(x2,y2),…,(x m ,y m )}, the Lagrange interpolation polynomial where x is the independent variable, and P(x) is the actually fitted semantic distribution function.

[0106] In the standard semantic coordinate system, the boundary point set of the standard semantic region is extracted, also based on the gradient change threshold of semantic density. The standard semantic distribution function is generated by piecewise linear fitting. The standard semantic region is divided into multiple small segments. Within each small segment, a linear function y = ax + b is used to fit the boundary points, where a and b are determined by methods such as the least squares method, so as to obtain the standard semantic distribution function.

[0107] When determining the dynamic threshold, multiple standard speech samples are randomly combined to generate a mixed sample set. Suppose there are N standard speech samples in total. Each time, k samples are randomly selected from these N samples for combination to generate multiple mixed samples. In the mixed sample set, the local deviation integral between the standard semantic distribution functions is calculated pairwise. For two standard semantic distribution functions f(x) and g(x), the local deviation integral where [a, b] is the integration interval, and the integration interval can be determined according to the semantic range of the speech samples.

[0108] The maximum value of all local deviation integrals is statistically calculated as the initial value of the dynamic threshold. Then, according to the noise level of the actual speech environment, the initial value is adaptively weighted and corrected. The noise level can be obtained by sensor measurement or other means. Suppose the noise level is L and the adaptive weighting coefficient is w. The dynamic threshold T = wL × I max , where I max is the maximum value of all local deviation integrals. The dynamic threshold determined in this way can better adapt to different speech environments and improve the accuracy and reliability of speech recognition.

[0109] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0110] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An NLP-based intelligent speech recognition method, characterized in that, Including: Obtain at least one standard speech sample of the target speech, and determine the key semantic markers in the standard speech sample. The number of the key semantic markers is at least three and they are distributed at different positions in the speech context; Extract the semantic intervals corresponding to the key semantic markers in the standard speech sample to generate a standard semantic region; Obtain the actual speech data to be recognized, detect the key semantic markers in the actual speech data, and generate an actual semantic region; Perform semantic alignment between the standard semantic region and the corresponding actual semantic region; Select three non-collinear actual semantic regions to construct an actual semantic coordinate system, and construct a standard semantic coordinate system through the corresponding standard semantic regions; In the actual semantic coordinate system, fit the semantic distribution function of the actual speech data, and in the standard semantic coordinate system, fit the semantic distribution function of the standard speech sample; Calculate the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function, and determine the dynamic threshold of the semantic deviation value; When the semantic deviation value exceeds the dynamic threshold, it is determined that there is an abnormal recognition of the actual speech data, otherwise it is determined to be normal.

2. The method for intelligent speech recognition based on NLP according to claim 1, characterized in that, The determination of the key semantic markers in the standard speech sample includes the following steps: Perform frame processing on the standard speech sample, and extract the acoustic feature vectors of each frame of speech; Based on the clustering result of the acoustic feature vectors, filter out the frames with a frequency distribution dispersion higher than a preset value as candidate semantic frames; In the candidate semantic frames, determine at least three semantic nodes across sentence boundaries through context-dependent relationship analysis as the key semantic markers.

3. The intelligent speech recognition method based on NLP according to claim 2, wherein The generation of the standard semantic region includes the following steps: Perform dynamic time warping on the speech segment where the key semantic marker is located to generate a standard time window; Extract the sequence of front and back related words of the key semantic marker within the standard time window to form a standard context chain; Map the standard context chain to a continuous region in the multi-dimensional semantic space as the standard semantic region.

4. An NLP-based intelligent speech recognition method according to claim 3, characterized in that, The generation of the actual semantic region includes the following steps: Perform noise suppression processing on the actual speech data to generate an enhanced speech signal; Based on the enhanced speech signal, use a sliding window to traverse and detect the acoustic features of the key semantic markers; Capture the boundaries of the key semantic markers on the time axis by dynamically adjusting the window length to form an actual context chain; Convert the actual context chain into a discrete region in the multi-dimensional semantic space as the actual semantic region.

5. An intelligent speech recognition method based on NLP according to claim 4, characterized in that, The semantic alignment includes the following steps: Perform semantic vector quantization encoding on the standard semantic region and the actual semantic region respectively; Calculate the similarity matrix between the standard semantic vector and the actual semantic vector based on the attention mechanism; In the similarity matrix, filter out the pairs of semantic vectors with the highest similarity and establish a one-to-one mapping relationship; According to the mapping relationship, perform interpolation compensation on the unmatched semantic regions to complete the alignment.

6. The method for intelligent speech recognition based on NLP according to claim 5, wherein The construction of the actual semantic coordinate system includes the following steps: Select three actual semantic regions as the first reference area, the second reference area, and the third reference area; Taking the semantic center of the first reference area as the origin, the line connecting the semantic centers of the first reference area and the second reference area as the first coordinate axis, and the line connecting the semantic centers of the first reference area and the third reference area as the second coordinate axis; Normalize to generate the unit scale of the actual semantic coordinate system according to the semantic spans of the first coordinate axis and the second coordinate axis.

7. An NLP-based intelligent speech recognition method according to claim 6, characterized in that, The construction of the standard semantic coordinate system includes the following steps: Take the corresponding standard semantic areas of the first reference area, the second reference area, and the third reference area as the first standard area, the second standard area, and the third standard area respectively; Taking the semantic center of the first standard area as the origin, the line connecting the semantic centers of the first standard area and the second standard area as the reference first axis, and the line connecting the semantic centers of the first standard area and the third standard area as the reference second axis; Dynamically adjust the dimension weights of the standard semantic coordinate system according to the semantic correlation degrees of the reference first axis and the reference second axis.

8. An intelligent speech recognition method based on NLP according to claim 7, characterized in that The fitting of the semantic distribution function includes the following steps: In the actual semantic coordinate system, extract the boundary point set of the actual semantic area, and generate the actual semantic distribution function through polynomial interpolation; In the standard semantic coordinate system, extract the boundary point set of the standard semantic area, and generate the standard semantic distribution function through piecewise linear fitting; Among them, the extraction of the boundary point set is based on the gradient change threshold of semantic density.

9. An intelligent speech recognition method based on NLP according to claim 8, characterized in that, The determination of the dynamic threshold includes the following steps: Randomly combine multiple standard speech samples to generate a mixed sample set; In the mixed sample set, calculate the local deviation integral between the standard semantic distribution functions pairwise; Statistically calculate the maximum value of all local deviation integrals as the initial value of the dynamic threshold; Perform adaptive weighted correction on the initial value according to the noise level of the actual speech environment.

10. An NLP-based intelligent speech recognition system, characterized in that, Include: Sample acquisition and processing module: used to acquire at least one standard speech sample of the target speech, determine the key semantic markers in the standard speech sample, the number of the key semantic markers is at least three and they are distributed at different positions in the speech context, and extract the semantic intervals corresponding to the key semantic markers in the standard speech sample to generate a standard semantic area; Actual speech processing module: used to acquire the actual speech data to be recognized, detect the key semantic markers in the actual speech data, and generate an actual semantic area; Semantic alignment module: semantically align the standard semantic area with the corresponding actual semantic area; Coordinate system construction module: select three non-collinear actual semantic areas to construct an actual semantic coordinate system, and construct a standard semantic coordinate system through the corresponding standard semantic areas; Semantic function fitting module: in the actual semantic coordinate system, fit the semantic distribution function of the actual speech data, and in the standard semantic coordinate system, fit the semantic distribution function of the standard speech sample; Deviation judgment module: calculate the semantic deviation value between the standard semantic distribution function and the actual semantic distribution function, and determine the dynamic threshold of the semantic deviation value. When the semantic deviation value exceeds the dynamic threshold, it is determined that there is an abnormal recognition in the actual speech data, otherwise it is determined to be normal.