Intelligent call shorthand method, system and medium based on speech recognition

By optimizing voice signals and voice recognition data, combining the calculation of voice recognition efficiency accuracy evaluation index and the comparison of threshold, the shortcomings of accuracy and real-time in smart call shorthand are solved, and more efficient call content recording is achieved.

CN119724245BActive Publication Date: 2025-06-06CHINA UNICOM WO MUSIC & CULTURE CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510214306.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the implementation of smart call shorthand, the prior art lacks accuracy, real-timeness and adaptability to complex environments.

Method used

By optimizing voice signals, optimizing voice recognition data, and calculating the comparison of speech recognition effectiveness accuracy and threshold, real-time voice recognition and shorthand can be achieved.

Benefits of technology

It improves the accuracy and real-timeness of smart call shorthand, enhances the ability to adapt to complex environments, and meets the needs of quickly and accurately recording call content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724245B_ABST
    Figure CN119724245B_ABST
Patent Text Reader

Abstract

The present application provides a method, system and medium for intelligent call shorthand based on speech recognition. The method includes: obtaining the speech signal of the real-time call for preprocessing to obtain the optimized speech signal and performing speech framing to obtain the speech frame, then performing speech feature extraction, obtaining the real-time voiceprint feature vector and the real-time speech feature vector and processing them, obtaining the caller identity label data and the corresponding optimized speech recognition data, obtaining the speech recognition evaluation data and processing them, obtaining the speech recognition accuracy evaluation index, and finally performing a threshold comparison with the preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result; the present application realizes the intelligence and accuracy of real-time call speech recognition and shorthand by optimizing the speech signal, optimizing the speech recognition data and the calculation and threshold comparison of the speech recognition accuracy evaluation index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and in particular to a method, system and medium for intelligent call shorthand based on speech recognition. Background Art

[0002] In the modern communication environment, whether it is business meetings, customer service, legal evidence collection or daily important calls, the content of the call needs to be accurately recorded. The traditional manual recording method is inefficient and prone to errors, and it is difficult to meet the needs of quickly and accurately recording the content of the call. With the development of speech recognition technology, it has become possible to use this technology to achieve intelligent call shorthand, but the current methods still have shortcomings in terms of accuracy, real-time performance, and adaptability to complex environments.

[0003] In view of the above problems, effective technical solutions are urgently needed. Summary of the invention

[0004] The purpose of this application is to provide an intelligent call shorthand method, system and medium based on speech recognition, which can achieve the intelligence and accuracy of real-time call speech recognition and shorthand by optimizing speech signals, optimizing speech recognition data and calculating and comparing thresholds of speech recognition accuracy evaluation indexes.

[0005] The present application also provides an intelligent call shorthand method based on speech recognition, comprising the following steps:

[0006] Acquire the voice signal of the real-time call, and pre-process the voice signal to obtain the optimized voice signal;

[0007] Processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors;

[0008] Processing is performed according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data;

[0009] Acquire and process the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index;

[0010] The speech recognition accuracy evaluation index is compared with a preset speech recognition accuracy threshold, and the speech recognition state is determined according to the threshold comparison result.

[0011] Optionally, in the intelligent call shorthand method based on speech recognition described in the present application, the step of acquiring a voice signal of a real-time call and preprocessing the voice signal to obtain an optimized voice signal includes:

[0012] Acquire a voice signal of a real-time call, and extract background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data;

[0013] According to the type feature data, noise intensity data and frequency characteristic data, a preset filter parameter database is searched to obtain corresponding filter setting parameters, and noise suppression is performed on the speech signal according to the filter setting parameters to obtain a denoised speech signal;

[0014] Extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal;

[0015] The gain speech signal is processed by a preset pre-emphasis coefficient to obtain an optimized speech signal.

[0016] Optionally, in the intelligent call shorthand method based on speech recognition described in the present application, the processing according to the optimized speech signal to obtain speech frames, and performing speech feature extraction on the speech frames respectively to obtain real-time voiceprint feature vectors and real-time speech feature vectors, includes:

[0017] According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames;

[0018] Extracting a corresponding real-time voiceprint feature vector according to the speech frame;

[0019] Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

[0020] Optionally, in the intelligent call shorthand method based on speech recognition described in the present application, the processing according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain the caller identity label data and the corresponding optimized speech recognition data includes:

[0021] Comparing the real-time voiceprint feature vector with the historical voiceprint feature vectors in a preset voiceprint database for similarity, to obtain corresponding similarities;

[0022] If the similarity is greater than a preset similarity threshold, the corresponding caller identity tag data is obtained;

[0023] Inputting the real-time speech feature vector into a preset speech recognition model for processing to obtain real-time speech recognition data corresponding to the caller identity tag data;

[0024] Inputting the real-time speech feature vector into a preset language recognition model for processing to obtain probability scores corresponding to words in the real-time speech recognition data;

[0025] The words with the highest probability scores are output in order to obtain optimized speech recognition data corresponding to the caller identity tag data.

[0026] Optionally, in the intelligent call shorthand method based on speech recognition described in the present application, the acquiring and processing the speech recognition evaluation data of the optimized speech recognition data to obtain the speech recognition accuracy evaluation index includes:

[0027] Acquire speech recognition evaluation data of the optimized speech recognition data, including word accuracy, sentence accuracy, recall rate, recognition time data, and recognition word count data;

[0028] Inputting the word accuracy, sentence accuracy and recall rate into a preset speech recognition accuracy evaluation model for processing to obtain a speech recognition accuracy evaluation index;

[0029] Compare the recognition time length data with the recognition word count data to obtain a recognition time efficiency coefficient, query a preset time efficiency weight list according to the recognition time efficiency coefficient, and obtain a corresponding time efficiency weight coefficient;

[0030] The speech recognition accuracy evaluation index is obtained by processing the timeliness weight coefficient and the speech recognition accuracy evaluation index.

[0031] Optionally, in the intelligent call shorthand method based on speech recognition described in the present application, the speech recognition accuracy evaluation index is threshold-compared with a preset speech recognition accuracy threshold, and the speech recognition state is determined according to the threshold comparison result, including:

[0032] Comparing the speech recognition accuracy evaluation index with a preset speech recognition accuracy requirement evaluation index to obtain a speech recognition accuracy relative value;

[0033] Comparing the speech recognition accuracy relative value with a preset speech recognition accuracy threshold;

[0034] If it is less than the preset speech recognition accuracy threshold, the speech recognition state is abnormal and an early warning response is output;

[0035] If it is greater than or equal to the preset speech recognition accuracy threshold, the speech recognition state is normal, and is recorded and stored according to the caller identity tag data and the corresponding optimized speech recognition data.

[0036] In a second aspect, the present application provides an intelligent call shorthand system based on speech recognition, the system comprising: a memory and a processor, the memory comprising a program of an intelligent call shorthand method based on speech recognition, the program of the intelligent call shorthand method based on speech recognition being executed by the processor to implement the following steps:

[0037] Acquire the voice signal of the real-time call, and pre-process the voice signal to obtain the optimized voice signal;

[0038] Processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors;

[0039] Processing is performed according to the voiceprint feature vector and the speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data;

[0040] Acquire and process the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index;

[0041] The speech recognition accuracy evaluation index is compared with a preset speech recognition accuracy threshold, and the speech recognition state is determined according to the threshold comparison result.

[0042] Optionally, in the intelligent call shorthand system based on speech recognition described in the present application, the step of acquiring a voice signal of a real-time call and preprocessing the voice signal to obtain an optimized voice signal includes:

[0043] Acquire a voice signal of a real-time call, and extract background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data;

[0044] According to the type feature data, noise intensity data and frequency characteristic data, a preset filter parameter database is searched to obtain corresponding filter setting parameters, and noise suppression is performed on the speech signal according to the filter setting parameters to obtain a denoised speech signal;

[0045] Extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal;

[0046] The gain speech signal is processed by a preset pre-emphasis coefficient to obtain an optimized speech signal.

[0047] Optionally, in the intelligent call shorthand system based on speech recognition described in the present application, the processing according to the optimized speech signal to obtain speech frames, and performing speech feature extraction on the speech frames respectively to obtain real-time voiceprint feature vectors and real-time speech feature vectors include:

[0048] According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames;

[0049] Extracting a corresponding real-time voiceprint feature vector according to the speech frame;

[0050] Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

[0051] In the third aspect, the present application also provides a computer-readable storage medium, which stores a program for an intelligent call shorthand method based on speech recognition. When the program for an intelligent call shorthand method based on speech recognition is executed by a processor, the steps of an intelligent call shorthand method based on speech recognition as described in any one of the above items are implemented.

[0052] From the above, it can be seen that the intelligent call shorthand method, system and medium based on speech recognition provided by the present application realize the intelligence and accuracy of real-time call speech recognition and shorthand by optimizing speech signals, optimizing speech recognition data and calculating and comparing thresholds of speech recognition accuracy evaluation indexes.

[0053] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the embodiments of the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0055] Figure 1 A flowchart of an intelligent call shorthand method based on speech recognition provided in an embodiment of the present application;

[0056] Figure 2A flowchart of obtaining an optimized voice signal in an intelligent call shorthand method based on voice recognition provided in an embodiment of the present application;

[0057] Figure 3 A flowchart of obtaining caller identity tag data and corresponding optimized speech recognition data in an intelligent speech recognition-based shorthand method provided in an embodiment of the present application;

[0058] Figure 4 A flowchart of obtaining a speech recognition accuracy evaluation index for an intelligent call shorthand method based on speech recognition provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0060] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0061] Please refer to Figure 1 , Figure 1 The flowchart of an intelligent call shorthand method based on speech recognition in some embodiments of the present application. The intelligent call shorthand method based on speech recognition is used in a terminal device, such as a computer, a mobile phone terminal, etc. The intelligent call shorthand method based on speech recognition includes the following steps:

[0062] S11, obtaining a voice signal of a real-time call, and preprocessing the voice signal to obtain an optimized voice signal;

[0063] S12, processing the optimized voice signal to obtain voice frames, and extracting voice features from the voice frames to obtain real-time voiceprint feature vectors and real-time voice feature vectors;

[0064] S13, processing the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data;

[0065] S14, obtaining and processing speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index;

[0066] S15, performing a threshold comparison between the speech recognition accuracy evaluation index and a preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result.

[0067] It should be noted that in order to realize the intelligent recognition and shorthand of real-time voice calls, the voice signals of real-time calls should be collected first, and preprocessed through preprocessing methods such as denoising, gain and pre-emphasis to improve the quality of voice signals and provide data support for accurate voice recognition. Then, the voice signals are divided into frames according to the predetermined frame length and frame shift to obtain multiple voice frames, and voice features are extracted to obtain real-time voiceprint feature vectors and real-time voice feature vectors. The real-time voiceprint feature vectors are processed to obtain the caller identity label data. The real-time voice feature vectors are processed to obtain voice recognition data, and the language recognition model is used to optimize to obtain the corresponding optimized voice recognition data. Then, the voice recognition evaluation data of the optimized voice recognition data is obtained and processed to obtain the voice recognition effectiveness evaluation index, which is used to evaluate the recognition accuracy and efficiency of the optimized voice recognition data. Finally, the voice recognition status is determined by threshold comparison, including normal status or abnormal status, and further processed according to the voice recognition status.

[0068] Please refer to Figure 2 , Figure 2 The flowchart of obtaining an optimized voice signal of an intelligent call shorthand method based on voice recognition in some embodiments of the present application. According to an embodiment of the present invention, the method of obtaining a voice signal of a real-time call and preprocessing the voice signal to obtain an optimized voice signal includes:

[0069] S21, obtaining a voice signal of a real-time call, and extracting background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data;

[0070] S22, querying a preset filter parameter database according to the type feature data, noise intensity data and frequency characteristic data, obtaining corresponding filter setting parameters, and performing noise suppression on the speech signal according to the filter setting parameters to obtain a denoised speech signal;

[0071] S23, extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal;

[0072] S24, processing the gain speech signal through a preset pre-emphasis coefficient to obtain an optimized speech signal.

[0073] It should be noted that due to the uncertainty of the call environment, for example, a call may be made in a quiet indoor environment, a noisy outdoor environment, or other different environments, resulting in unstable call quality, which affects the accuracy of speech recognition. Therefore, in order to improve the quality of the voice signal of the real-time call, the background noise feature data is first extracted, including type feature data, noise intensity data, and frequency characteristic data. The preset filter parameter database is queried based on the extracted type feature data, noise intensity data, and frequency characteristic data to obtain the corresponding filter setting parameters, wherein the preset filter parameter database is obtained by obtaining the type feature data, noise intensity data, and frequency characteristic data of a large number of historical samples and the corresponding filter setting parameter analysis, and is provided by a preset intelligent call shorthand platform. The voice signal is subjected to noise suppression based on the filter setting parameters to obtain De-noised speech signal; at the same time, due to the different voices of the callers, the accuracy of speech recognition will also be affected. Therefore, technical personnel in this field can obtain the energy value of the speech signal through the energy calculation module according to the preset time, and query the preset gain parameter list according to the energy value to obtain the corresponding gain parameter, wherein the preset gain parameter list is provided by the preset intelligent call shorthand platform, and the signal gain of the denoised speech signal is adjusted according to the gain parameter to obtain a gain speech signal to ensure that the speech signal has a certain strength; finally, after completing the noise removal and gain adjustment, the speech signal is pre-emphasized according to the preset pre-emphasis coefficient, wherein the preset pre-emphasis coefficient can be determined by technical personnel in this field through spectrum analysis and energy attenuation law, and finally an optimized speech signal is obtained to provide high-quality data support for subsequent speech recognition.

[0074] According to an embodiment of the present invention, the processing according to the optimized voice signal to obtain voice frames, and extracting voice features from the voice frames to obtain real-time voiceprint feature vectors and real-time voice feature vectors includes:

[0075] According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames;

[0076] Extracting a corresponding real-time voiceprint feature vector according to the speech frame;

[0077] Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

[0078] It should be noted that in order to identify the identity of the caller and the corresponding call voice recognition data, the voice signal is first framed according to the optimized voice signal by preset frame length and frame shift parameters to obtain multiple voice frames, wherein, in this embodiment, the frame length is set to 30 milliseconds and the frame shift is set to 15 milliseconds. The corresponding real-time voiceprint feature vector is extracted according to the obtained voice frame to identify the identity of the caller; technicians in this field can extract the corresponding Mel-frequency cepstral coefficients by performing discrete cosine transform on the obtained voice frame after spectral analysis conversion and Mel filter group processing. The Mel-frequency cepstral coefficient is a spectral feature used to simulate human auditory perception and is used for voice. It has good characterization ability in terms of timbre, sound quality, etc.; at the same time, through critical band analysis, calculation of autocorrelation function, linear prediction analysis and perceptual weighted extraction of corresponding perceptual linear prediction coefficients, the perceptual linear prediction coefficient provides additional information from the perspective of critical band and spectrum envelope perceived by the human ear. Finally, according to the obtained Mel-frequency cepstrum coefficients and perceptual linear prediction coefficients, feature vectors are combined through direct concatenation or weighted summation to obtain the real-time speech feature vector corresponding to the real-time voiceprint feature vector, which effectively makes up for the shortcomings of each single feature in describing the complex acoustic characteristics of speech, and provides a richer and more accurate information basis for subsequent speech recognition.

[0079] Please refer to Figure 3 , Figure 3 The present invention is a flowchart of obtaining caller identity tag data and corresponding optimized speech recognition data of an intelligent call shorthand method based on speech recognition in some embodiments of the present application. According to an embodiment of the present invention, the processing according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain the caller identity tag data and the corresponding optimized speech recognition data includes:

[0080] S31, comparing the real-time voiceprint feature vector with the historical voiceprint feature vectors in a preset voiceprint database to obtain corresponding similarities;

[0081] S32, if the similarity is greater than a preset similarity threshold, obtaining corresponding caller identity tag data;

[0082] S33, inputting the real-time speech feature vector into a preset speech recognition model for processing, and obtaining real-time speech recognition data corresponding to the caller identity tag data;

[0083] S34, inputting the real-time speech feature vector into a preset language recognition model for processing, and obtaining probability scores corresponding to words in the real-time speech recognition data;

[0084] S35. Output the words with the highest probability scores in order to obtain optimized speech recognition data corresponding to the caller identity tag data.

[0085] It should be noted that in order to identify the identity of the caller and the corresponding optimized voice recognition data for outputting call records, first, the obtained real-time voiceprint feature vector is compared with the historical voiceprint feature vector in the preset voiceprint database for similarity to obtain the corresponding similarity, wherein the preset voiceprint database stores the voiceprint feature vectors of people of different identities, which is provided by the preset intelligent call shorthand platform and can be updated in real time. The obtained similarity is compared with the preset similarity threshold. If it is greater than, the corresponding caller identity label data is obtained. The caller identity label data is used to distinguish the caller corresponding to the optimized voice recognition data for easy reference. If it is less than or equal to the preset similarity threshold, the caller identity label data is marked as a stranger. If necessary, the preset voiceprint database is updated in time; then the real-time voice feature vector is input into the preset voice recognition model for processing to obtain the real-time voice recognition data corresponding to the caller identity label data, wherein The speech recognition model is obtained by obtaining speech feature vectors of a large number of historical samples and corresponding speech recognition data for training. Affected by the accuracy of speech recognition, there may be multiple real-time speech recognition data for recognition. Finally, the real-time speech feature vector is input into the preset language recognition model for processing to obtain the probability score corresponding to the word in the real-time speech recognition data. Among them, the language recognition model is obtained by obtaining speech feature vectors of a large number of historical samples and corresponding probability score training, and then the words with the highest probability score are output in sequence to obtain the optimized speech recognition data corresponding to the caller identity label data. For example, the real-time speech recognition data is recognized as two results, "I like to eat apples" and "I like to eat apples". The probability score of "I" is 9.5 points, the probability score of "I" is 2 points, the probability score of "apple" is 9 points, and the probability score of "flat" is 3 points. The output words are "I" and "apple", and the optimized speech recognition data is "I like to eat apples".

[0086] Please refer to Figure 4 , Figure 4 The present invention is a flowchart of obtaining a speech recognition accuracy evaluation index of a method for intelligent call shorthand based on speech recognition in some embodiments of the present application. According to an embodiment of the present invention, the step of obtaining speech recognition evaluation data of the optimized speech recognition data and processing the data to obtain a speech recognition accuracy evaluation index includes:

[0087] S41, obtaining speech recognition evaluation data of the optimized speech recognition data, including word accuracy, sentence accuracy, recall rate, recognition time data and recognition word count data;

[0088] S42, inputting the word accuracy, sentence accuracy and recall rate into a preset speech recognition accuracy evaluation model for processing to obtain a speech recognition accuracy evaluation index;

[0089] S43, comparing the recognition time length data with the recognition word count data to obtain a recognition time efficiency coefficient, querying a preset time efficiency weight list according to the recognition time efficiency coefficient, and obtaining a corresponding time efficiency weight coefficient;

[0090] S44, performing processing according to the timeliness weight coefficient and the speech recognition accuracy evaluation index to obtain a speech recognition accuracy evaluation index.

[0091] It should be noted that after completing speech recognition and optimization, it is necessary to evaluate whether the accuracy of recognition meets the requirements. First, obtain speech recognition evaluation data including word accuracy, sentence accuracy, recall rate, recognition time data and recognition word count data. Among them, word accuracy refers to the ratio of the number of correctly recognized words to the total number of words, sentence accuracy refers to the ratio of the number of correctly recognized sentences to the total number of sentences, and recall rate refers to the ratio of correctly recognized speech content to the actual speech content. For example, there are 100 valid information units (such as keywords, key sentences, etc.) in the actual call, and the speech recognition system only recognizes 70, then the recall rate is 70%; input the word accuracy, sentence accuracy and recall rate into the preset speech recognition accuracy evaluation model for processing to obtain the speech recognition accuracy evaluation index;

[0092] The calculation formula of the speech recognition accuracy evaluation index is:

[0093] ;

[0094] in, is the speech recognition accuracy evaluation index, , , They are word accuracy, sentence accuracy and recall, , , is a preset characteristic coefficient (the characteristic coefficient is obtained by querying a preset intelligent call shorthand platform);

[0095] Then, the recognition time efficiency coefficient is obtained by comparing the obtained recognition time data with the recognition word count data. For example, if the recognition time data is 10 minutes and the recognition word count data is 100, then 10 / 100=0.1 is the recognition time efficiency coefficient. According to the obtained recognition time efficiency coefficient, the preset time efficiency weight list is queried to obtain the corresponding time efficiency weight coefficient, wherein the time efficiency weight list is obtained by querying the preset intelligent call shorthand platform; finally, the time efficiency weight coefficient and the speech recognition accuracy evaluation index are processed to obtain the speech recognition accuracy evaluation index;

[0096] The calculation formula of the speech recognition accuracy evaluation index is:

[0097] ;

[0098] in, is the speech recognition accuracy evaluation index. , are the timeliness weight coefficient and speech recognition accuracy evaluation index respectively, It is the preset characteristic coefficient (the characteristic coefficient is obtained by querying the preset intelligent call shorthand platform).

[0099] According to an embodiment of the present invention, the step of performing a threshold comparison between the speech recognition accuracy evaluation index and a preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result, includes:

[0100] Comparing the speech recognition accuracy evaluation index with a preset speech recognition accuracy requirement evaluation index to obtain a speech recognition accuracy relative value;

[0101] Comparing the speech recognition accuracy relative value with a preset speech recognition accuracy threshold;

[0102] If it is less than the preset speech recognition accuracy threshold, the speech recognition state is abnormal and an early warning response is output;

[0103] If it is greater than or equal to the preset speech recognition accuracy threshold, the speech recognition state is normal, and is recorded and stored according to the caller identity tag data and the corresponding optimized speech recognition data.

[0104] It should be noted that in order to determine whether the speech recognition state is normal, the obtained speech recognition accuracy evaluation index is first compared with the preset speech recognition accuracy requirement evaluation index to obtain the speech recognition accuracy relative value. For example, the obtained speech recognition accuracy evaluation index is 7, and the speech recognition accuracy requirement evaluation index is 10, then 7 / 10=0.7 is the speech recognition accuracy relative value, and then the speech recognition accuracy relative value is compared with the preset speech recognition accuracy threshold. In this embodiment, the speech recognition accuracy threshold is set to (0.0.75) and [0.75, 1], which correspond to abnormal state and normal state respectively. For example, if the obtained speech recognition accuracy relative value is 0.7, the speech recognition state is abnormal, and an early warning reminder response is output at the same time. If the obtained speech recognition accuracy relative value is 0.8, the speech recognition state is normal, and it is recorded and stored according to the caller identity label data and the corresponding optimized speech recognition data.

[0105] It is worth mentioning that according to an embodiment of the present invention, it also includes:

[0106] Get the identity information of the person who reads the optimized speech recognition data, and query the preset permission list based on the identity information to obtain the corresponding access permission information:

[0107] If the access permission information is higher than the preset reading permission, the optimized voice recognition data that can be read by the access permission information is sent to the reader terminal for display;

[0108] If the access permission information is lower than the preset read permission, the optimized voice recognition data is refused to be read and a warning response is output to the management end.

[0109] It should be noted that in order to ensure the security of identified and stored call data, permissions should be set for access to the data, and only authorized personnel are allowed to read it. First, the identity information of the person who reads the optimized voice recognition data is obtained, and the preset permission list is queried based on the identity information to obtain the corresponding access permission information, where the preset permission list is obtained by querying the preset intelligent call shorthand platform, and the obtained access permission information is compared with the preset reading permission. If the access permission information is higher than the preset reading permission, the optimized voice recognition data that the access permission information allows to be read is sent to the reader's terminal for display. If the access permission information is lower than the preset reading permission, there may be illegal access, and the optimized voice recognition data is refused to be read and an early warning response is output to the management end.

[0110] It is worth mentioning that, according to an embodiment of the present invention, the method further includes: processing the optimized voice signal to obtain voice frames, extracting voice features from the voice frames to obtain real-time voiceprint feature vectors and real-time voice feature vectors; and then:

[0111] Obtaining first-order difference values ​​and second-order difference values ​​of real-time speech feature vectors of adjacent frames in the speech frame;

[0112] A real-time speech optimization feature vector is obtained by performing a combination process on the first-order difference value, the second-order difference value and the real-time speech feature vector.

[0113] It should be noted that in order to effectively improve the accuracy of speech recognition in complex calls, the real-time speech feature vector should be further optimized by obtaining the first-order difference value and the second-order difference value of the real-time speech feature vector of adjacent frames in the speech frame, wherein the first-order difference value is used to describe the rate of change of speech features between adjacent frames, and the first-order difference value is obtained by calculating the difference between the Mel-frequency cepstral coefficient and the perceptual linear prediction coefficient of each frame of speech and the corresponding coefficient of the previous frame; the second-order difference value further describes the acceleration of the change of speech features on the basis of the first-order difference, and is obtained by calculating the adjacent frame difference of the first-order difference value; technicians in this field can calculate the calculated The first-order difference value and the second-order difference value are combined with the original Mel-frequency cepstral coefficient and perceptual linear prediction coefficient to form a more comprehensive real-time speech optimization feature vector. This fused feature vector contains both the static acoustic feature information of the speech and its dynamic change information in the time dimension. It can describe the speech signal more accurately and comprehensively, and provide richer input data for subsequent speech recognition models, thereby improving the speech recognition system's ability to understand and recognize speech content, especially when processing continuous speech, complex speech environments and diverse speaker voices, significantly improving the performance of the intelligent call shorthand method.

[0114] The present invention also discloses an intelligent call shorthand system based on speech recognition, comprising a memory and a processor, wherein the memory comprises an intelligent call shorthand method program based on speech recognition, and when the intelligent call shorthand method program based on speech recognition is executed by the processor, the following steps are implemented:

[0115] Acquire the voice signal of the real-time call, and pre-process the voice signal to obtain the optimized voice signal;

[0116] Processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors;

[0117] Processing is performed according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data;

[0118] Acquire and process the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index;

[0119] The speech recognition accuracy evaluation index is compared with a preset speech recognition accuracy threshold, and the speech recognition state is determined according to the threshold comparison result.

[0120] It should be noted that in order to realize the intelligent recognition and shorthand of real-time voice calls, the voice signals of real-time calls should be collected first, and preprocessed through preprocessing methods such as denoising, gain and pre-emphasis to improve the quality of voice signals and provide data support for accurate voice recognition. Then, the voice signals are divided into frames according to the predetermined frame length and frame shift to obtain multiple voice frames, and voice features are extracted to obtain real-time voiceprint feature vectors and real-time voice feature vectors. The real-time voiceprint feature vectors are processed to obtain the caller identity label data. The real-time voice feature vectors are processed to obtain voice recognition data, and the language recognition model is used to optimize to obtain the corresponding optimized voice recognition data. Then, the voice recognition evaluation data of the optimized voice recognition data is obtained and processed to obtain the voice recognition effectiveness evaluation index, which is used to evaluate the recognition accuracy and efficiency of the optimized voice recognition data. Finally, the voice recognition status is determined by threshold comparison, including normal status or abnormal status, and further processed according to the voice recognition status.

[0121] According to an embodiment of the present invention, the step of acquiring a voice signal of a real-time call and preprocessing the voice signal to obtain an optimized voice signal includes:

[0122] Acquire a voice signal of a real-time call, and extract background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data;

[0123] According to the type feature data, noise intensity data and frequency characteristic data, a preset filter parameter database is searched to obtain corresponding filter setting parameters, and noise suppression is performed on the speech signal according to the filter setting parameters to obtain a denoised speech signal;

[0124] Extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal;

[0125] The gain speech signal is processed by a preset pre-emphasis coefficient to obtain an optimized speech signal.

[0126] It should be noted that due to the uncertainty of the call environment, for example, a call may be made in a quiet indoor environment, a noisy outdoor environment, or other different environments, resulting in unstable call quality, which affects the accuracy of speech recognition. Therefore, in order to improve the quality of the voice signal of the real-time call, the background noise feature data is first extracted, including type feature data, noise intensity data, and frequency characteristic data. The preset filter parameter database is queried based on the extracted type feature data, noise intensity data, and frequency characteristic data to obtain the corresponding filter setting parameters, wherein the preset filter parameter database is obtained by obtaining the type feature data, noise intensity data, and frequency characteristic data of a large number of historical samples and the corresponding filter setting parameter analysis, and is provided by a preset intelligent call shorthand platform. The voice signal is subjected to noise suppression based on the filter setting parameters to obtain De-noised speech signal; at the same time, due to the different voices of the callers, the accuracy of speech recognition will also be affected. Therefore, technical personnel in this field can obtain the energy value of the speech signal through the energy calculation module according to the preset time, and query the preset gain parameter list according to the energy value to obtain the corresponding gain parameter, wherein the preset gain parameter list is provided by the preset intelligent call shorthand platform, and the signal gain of the denoised speech signal is adjusted according to the gain parameter to obtain a gain speech signal to ensure that the speech signal has a certain strength; finally, after completing the noise removal and gain adjustment, the speech signal is pre-emphasized according to the preset pre-emphasis coefficient, wherein the preset pre-emphasis coefficient can be determined by technical personnel in this field through spectrum analysis and energy attenuation law, and finally an optimized speech signal is obtained to provide high-quality data support for subsequent speech recognition.

[0127] According to an embodiment of the present invention, the processing according to the optimized voice signal to obtain voice frames, and extracting voice features from the voice frames to obtain real-time voiceprint feature vectors and real-time voice feature vectors includes:

[0128] According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames;

[0129] Extracting a corresponding real-time voiceprint feature vector according to the speech frame;

[0130] Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

[0131] It should be noted that in order to identify the identity of the caller and the corresponding call voice recognition data, the voice signal is first framed according to the optimized voice signal by preset frame length and frame shift parameters to obtain multiple voice frames, wherein, in this embodiment, the frame length is set to 30 milliseconds and the frame shift is set to 15 milliseconds. The corresponding real-time voiceprint feature vector is extracted according to the obtained voice frame to identify the identity of the caller; technicians in this field can extract the corresponding Mel-frequency cepstral coefficients by performing discrete cosine transform on the obtained voice frame after spectral analysis conversion and Mel filter group processing. The Mel-frequency cepstral coefficient is a spectral feature used to simulate human auditory perception and is used for voice. It has good characterization ability in terms of timbre, sound quality, etc.; at the same time, through critical band analysis, calculation of autocorrelation function, linear prediction analysis and perceptual weighted extraction of corresponding perceptual linear prediction coefficients, the perceptual linear prediction coefficient provides additional information from the perspective of critical band and spectrum envelope perceived by the human ear. Finally, according to the obtained Mel-frequency cepstrum coefficients and perceptual linear prediction coefficients, feature vectors are combined through direct concatenation or weighted summation to obtain the real-time speech feature vector corresponding to the real-time voiceprint feature vector, which effectively makes up for the shortcomings of each single feature in describing the complex acoustic characteristics of speech, and provides a richer and more accurate information basis for subsequent speech recognition.

[0132] According to an embodiment of the present invention, the processing according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data includes:

[0133] Comparing the real-time voiceprint feature vector with the historical voiceprint feature vectors in a preset voiceprint database for similarity, to obtain corresponding similarities;

[0134] If the similarity is greater than a preset similarity threshold, the corresponding caller identity tag data is obtained;

[0135] Inputting the real-time speech feature vector into a preset speech recognition model for processing to obtain real-time speech recognition data corresponding to the caller identity tag data;

[0136] Inputting the real-time speech feature vector into a preset language recognition model for processing to obtain probability scores corresponding to words in the real-time speech recognition data;

[0137] The words with the highest probability scores are output in order to obtain optimized speech recognition data corresponding to the caller identity tag data.

[0138] It should be noted that in order to identify the identity of the caller and the corresponding optimized voice recognition data for outputting call records, first, the obtained real-time voiceprint feature vector is compared with the historical voiceprint feature vector in the preset voiceprint database for similarity to obtain the corresponding similarity, wherein the preset voiceprint database stores the voiceprint feature vectors of people of different identities, which is provided by the preset intelligent call shorthand platform and can be updated in real time. The obtained similarity is compared with the preset similarity threshold. If it is greater than, the corresponding caller identity label data is obtained. The caller identity label data is used to distinguish the caller corresponding to the optimized voice recognition data for easy reference. If it is less than or equal to the preset similarity threshold, the caller identity label data is marked as a stranger. If necessary, the preset voiceprint database is updated in time; then the real-time voice feature vector is input into the preset voice recognition model for processing to obtain the real-time voice recognition data corresponding to the caller identity label data, wherein The speech recognition model is obtained by obtaining speech feature vectors of a large number of historical samples and corresponding speech recognition data for training. Affected by the accuracy of speech recognition, there may be multiple real-time speech recognition data for recognition. Finally, the real-time speech feature vector is input into the preset language recognition model for processing to obtain the probability score corresponding to the word in the real-time speech recognition data. Among them, the language recognition model is obtained by obtaining speech feature vectors of a large number of historical samples and corresponding probability score training, and then the words with the highest probability score are output in sequence to obtain the optimized speech recognition data corresponding to the caller identity label data. For example, the real-time speech recognition data is recognized as two results, "I like to eat apples" and "I like to eat apples". The probability score of "I" is 9.5 points, the probability score of "I" is 2 points, the probability score of "apple" is 9 points, and the probability score of "flat" is 3 points. The output words are "I" and "apple", and the optimized speech recognition data is "I like to eat apples".

[0139] According to an embodiment of the present invention, the step of acquiring and processing speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index includes:

[0140] Acquire speech recognition evaluation data of the optimized speech recognition data, including word accuracy, sentence accuracy, recall rate, recognition time data, and recognition word count data;

[0141] Inputting the word accuracy, sentence accuracy and recall rate into a preset speech recognition accuracy evaluation model for processing to obtain a speech recognition accuracy evaluation index;

[0142] Compare the recognition time length data with the recognition word count data to obtain a recognition time efficiency coefficient, query a preset time efficiency weight list according to the recognition time efficiency coefficient, and obtain a corresponding time efficiency weight coefficient;

[0143] The speech recognition accuracy evaluation index is obtained by processing the timeliness weight coefficient and the speech recognition accuracy evaluation index.

[0144] It should be noted that after completing speech recognition and optimization, it is necessary to evaluate whether the accuracy of recognition meets the requirements. First, obtain speech recognition evaluation data including word accuracy, sentence accuracy, recall rate, recognition time data and recognition word count data. Among them, word accuracy refers to the ratio of the number of correctly recognized words to the total number of words, sentence accuracy refers to the ratio of the number of correctly recognized sentences to the total number of sentences, and recall rate refers to the ratio of correctly recognized speech content to the actual speech content. For example, there are 100 valid information units (such as keywords, key sentences, etc.) in the actual call, and the speech recognition system only recognizes 70, then the recall rate is 70%; input the word accuracy, sentence accuracy and recall rate into the preset speech recognition accuracy evaluation model for processing to obtain the speech recognition accuracy evaluation index;

[0145] The calculation formula of the speech recognition accuracy evaluation index is:

[0146] ;

[0147] in, is the speech recognition accuracy evaluation index, , , They are word accuracy, sentence accuracy and recall, , , is a preset characteristic coefficient (the characteristic coefficient is obtained by querying a preset intelligent call shorthand platform);

[0148] Then, the recognition time efficiency coefficient is obtained by comparing the obtained recognition time data with the recognition word count data. For example, if the recognition time data is 10 minutes and the recognition word count data is 100, then 10 / 100=0.1 is the recognition time efficiency coefficient. According to the obtained recognition time efficiency coefficient, the preset time efficiency weight list is queried to obtain the corresponding time efficiency weight coefficient, wherein the time efficiency weight list is obtained by querying the preset intelligent call shorthand platform; finally, the time efficiency weight coefficient and the speech recognition accuracy evaluation index are processed to obtain the speech recognition accuracy evaluation index;

[0149] The calculation formula of the speech recognition accuracy evaluation index is:

[0150] ;

[0151] in, is the speech recognition accuracy evaluation index. , are the timeliness weight coefficient and speech recognition accuracy evaluation index respectively, It is the preset characteristic coefficient (the characteristic coefficient is obtained by querying the preset intelligent call shorthand platform).

[0152] According to an embodiment of the present invention, the step of performing a threshold comparison between the speech recognition accuracy evaluation index and a preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result, includes:

[0153] Comparing the speech recognition accuracy evaluation index with a preset speech recognition accuracy requirement evaluation index to obtain a speech recognition accuracy relative value;

[0154] Comparing the speech recognition accuracy relative value with a preset speech recognition accuracy threshold;

[0155] If it is less than the preset speech recognition accuracy threshold, the speech recognition state is abnormal and an early warning response is output;

[0156] If it is greater than or equal to the preset speech recognition accuracy threshold, the speech recognition state is normal, and is recorded and stored according to the caller identity tag data and the corresponding optimized speech recognition data.

[0157] It should be noted that in order to determine whether the speech recognition state is normal, the obtained speech recognition accuracy evaluation index is first compared with the preset speech recognition accuracy requirement evaluation index to obtain the speech recognition accuracy relative value. For example, the obtained speech recognition accuracy evaluation index is 7, and the speech recognition accuracy requirement evaluation index is 10, then 7 / 10=0.7 is the speech recognition accuracy relative value, and then the speech recognition accuracy relative value is compared with the preset speech recognition accuracy threshold. In this embodiment, the speech recognition accuracy threshold is set to (0.0.75) and [0.75, 1], which correspond to abnormal state and normal state respectively. For example, if the obtained speech recognition accuracy relative value is 0.7, the speech recognition state is abnormal, and an early warning reminder response is output at the same time. If the obtained speech recognition accuracy relative value is 0.8, the speech recognition state is normal, and it is recorded and stored according to the caller identity label data and the corresponding optimized speech recognition data.

[0158] It is worth mentioning that according to an embodiment of the present invention, it also includes:

[0159] Get the identity information of the person who reads the optimized speech recognition data, and query the preset permission list based on the identity information to obtain the corresponding access permission information:

[0160] If the access permission information is higher than the preset reading permission, the optimized voice recognition data that can be read by the access permission information is sent to the reader terminal for display;

[0161] If the access permission information is lower than the preset read permission, the optimized voice recognition data is refused to be read and a warning response is output to the management end.

[0162] It should be noted that in order to ensure the security of identified and stored call data, permissions should be set for access to the data, and only authorized personnel are allowed to read it. First, the identity information of the person who reads the optimized voice recognition data is obtained, and the preset permission list is queried based on the identity information to obtain the corresponding access permission information, where the preset permission list is obtained by querying the preset intelligent call shorthand platform, and the obtained access permission information is compared with the preset reading permission. If the access permission information is higher than the preset reading permission, the optimized voice recognition data that the access permission information allows to be read is sent to the reader's terminal for display. If the access permission information is lower than the preset reading permission, there may be illegal access, and the optimized voice recognition data is refused to be read and an early warning response is output to the management end.

[0163] It is worth mentioning that, according to an embodiment of the present invention, the method further includes: processing the optimized voice signal to obtain voice frames, extracting voice features from the voice frames to obtain real-time voiceprint feature vectors and real-time voice feature vectors; and then:

[0164] Obtaining first-order difference values ​​and second-order difference values ​​of real-time speech feature vectors of adjacent frames in the speech frame;

[0165] A real-time speech optimization feature vector is obtained by performing a combination process on the first-order difference value, the second-order difference value and the real-time speech feature vector.

[0166] It should be noted that in order to effectively improve the accuracy of speech recognition in complex calls, the real-time speech feature vector should be further optimized by obtaining the first-order difference value and the second-order difference value of the real-time speech feature vector of adjacent frames in the speech frame, wherein the first-order difference value is used to describe the rate of change of speech features between adjacent frames, and the first-order difference value is obtained by calculating the difference between the Mel-frequency cepstral coefficient and the perceptual linear prediction coefficient of each frame of speech and the corresponding coefficient of the previous frame; the second-order difference value further describes the acceleration of the change of speech features on the basis of the first-order difference, and is obtained by calculating the adjacent frame difference of the first-order difference value; technicians in this field can calculate the calculated The first-order difference value and the second-order difference value are combined with the original Mel-frequency cepstral coefficient and perceptual linear prediction coefficient to form a more comprehensive real-time speech optimization feature vector. This fused feature vector contains both the static acoustic feature information of the speech and its dynamic change information in the time dimension. It can describe the speech signal more accurately and comprehensively, and provide richer input data for subsequent speech recognition models, thereby improving the speech recognition system's ability to understand and recognize speech content, especially when processing continuous speech, complex speech environments and diverse speaker voices, significantly improving the performance of the intelligent call shorthand method.

[0167] A third aspect of the present invention provides a readable storage medium, which stores a program for an intelligent call shorthand method based on speech recognition. When the program for an intelligent call shorthand method based on speech recognition is executed by a processor, the steps of an intelligent call shorthand method based on speech recognition as described in any one of the above items are implemented.

[0168] The present invention discloses an intelligent call shorthand method, system and medium based on speech recognition, which realizes the intelligence and accuracy of real-time call speech recognition and shorthand by optimizing speech signals, optimizing speech recognition data and calculating and comparing speech recognition accuracy evaluation indexes with threshold values.

[0169] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0170] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0171] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0172] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, the aforementioned program can be stored in a readable storage medium, and when the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, and other media that can store program codes.

[0173] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

Claims

1. An intelligent call shorthand method based on speech recognition, characterized in that: The following steps are involved: Acquire the voice signal of the real-time call, and pre-process the voice signal to obtain the optimized voice signal; Processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors; Processing is performed according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data; Acquire and process the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index; Performing a threshold comparison between the speech recognition accuracy evaluation index and a preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result; The step of obtaining and processing the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index includes: Acquire speech recognition evaluation data of the optimized speech recognition data, including word accuracy, sentence accuracy, recall rate, recognition time data, and recognition word count data; Inputting the word accuracy, sentence accuracy and recall rate into a preset speech recognition accuracy evaluation model for processing to obtain a speech recognition accuracy evaluation index; Compare the recognition time length data with the recognition word count data to obtain a recognition time efficiency coefficient, query a preset time efficiency weight list according to the recognition time efficiency coefficient, and obtain a corresponding time efficiency weight coefficient; The speech recognition accuracy evaluation index is obtained by processing the timeliness weight coefficient and the speech recognition accuracy evaluation index.

2. The intelligent call shorthand method based on speech recognition according to claim 1, characterized in that: The acquiring of the voice signal of the real-time call and preprocessing the voice signal to obtain the optimized voice signal includes: Acquire a voice signal of a real-time call, and extract background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data; According to the type feature data, noise intensity data and frequency characteristic data, a preset filter parameter database is searched to obtain corresponding filter setting parameters, and noise suppression is performed on the speech signal according to the filter setting parameters to obtain a denoised speech signal; Extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal; The gain speech signal is processed by a preset pre-emphasis coefficient to obtain an optimized speech signal.

3. The intelligent call shorthand method based on speech recognition according to claim 2 is characterized in that: The step of processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors includes: According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames; Extracting a corresponding real-time voiceprint feature vector according to the speech frame; Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

4. The intelligent call shorthand method based on speech recognition according to claim 3 is characterized in that: The processing according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data includes: Comparing the real-time voiceprint feature vector with the historical voiceprint feature vectors in a preset voiceprint database for similarity, to obtain corresponding similarities; If the similarity is greater than a preset similarity threshold, the corresponding caller identity tag data is obtained; Inputting the real-time speech feature vector into a preset speech recognition model for processing to obtain real-time speech recognition data corresponding to the caller identity tag data; Inputting the real-time speech feature vector into a preset language recognition model for processing to obtain probability scores corresponding to words in the real-time speech recognition data; The words with the highest probability scores are output in order to obtain optimized speech recognition data corresponding to the caller identity tag data.

5. The intelligent call shorthand method based on speech recognition according to claim 1, characterized in that: The step of comparing the speech recognition accuracy evaluation index with a preset speech recognition accuracy threshold and determining the speech recognition state according to the threshold comparison result includes: Comparing the speech recognition accuracy evaluation index with a preset speech recognition accuracy requirement evaluation index to obtain a speech recognition accuracy relative value; Comparing the speech recognition accuracy relative value with a preset speech recognition accuracy threshold; If it is less than the preset speech recognition accuracy threshold, the speech recognition state is abnormal and an early warning response is output; If it is greater than or equal to the preset speech recognition accuracy threshold, the speech recognition state is normal, and is recorded and stored according to the caller identity tag data and the corresponding optimized speech recognition data.

6. An intelligent conversation shorthand system based on speech recognition, characterized in that: The invention comprises a memory and a processor, wherein the memory comprises an intelligent call shorthand method program based on speech recognition, and the intelligent call shorthand method program based on speech recognition is executed by the processor to implement the following steps: Acquire the voice signal of the real-time call, and pre-process the voice signal to obtain the optimized voice signal; Processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors; Processing is performed according to the real-time voiceprint feature vector and the real-time speech feature vector to obtain caller identity tag data and corresponding optimized speech recognition data; Acquire and process the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index; Performing a threshold comparison between the speech recognition accuracy evaluation index and a preset speech recognition accuracy threshold, and determining the speech recognition state according to the threshold comparison result; The step of obtaining and processing the speech recognition evaluation data of the optimized speech recognition data to obtain a speech recognition accuracy evaluation index includes: Acquire speech recognition evaluation data of the optimized speech recognition data, including word accuracy, sentence accuracy, recall rate, recognition time data, and recognition word count data; Inputting the word accuracy, sentence accuracy and recall rate into a preset speech recognition accuracy evaluation model for processing to obtain a speech recognition accuracy evaluation index; Compare the recognition time length data with the recognition word count data to obtain a recognition time efficiency coefficient, query a preset time efficiency weight list according to the recognition time efficiency coefficient, and obtain a corresponding time efficiency weight coefficient; The speech recognition accuracy evaluation index is obtained by processing the timeliness weight coefficient and the speech recognition accuracy evaluation index.

7. The intelligent speech recording system based on speech recognition according to claim 6 is characterized in that: The acquiring of the voice signal of the real-time call and preprocessing the voice signal to obtain the optimized voice signal includes: Acquire a voice signal of a real-time call, and extract background noise feature data within a preset time period according to the voice signal, including type feature data, noise intensity data, and frequency characteristic data; According to the type feature data, noise intensity data and frequency characteristic data, a preset filter parameter database is searched to obtain corresponding filter setting parameters, and noise suppression is performed on the speech signal according to the filter setting parameters to obtain a denoised speech signal; Extracting an energy value within a preset time period according to the denoised speech signal, querying a preset gain parameter list according to the energy value to obtain a corresponding gain parameter, and adjusting the signal gain of the denoised speech signal according to the gain parameter to obtain a gain speech signal; The gain speech signal is processed by a preset pre-emphasis coefficient to obtain an optimized speech signal.

8. The intelligent speech recording system based on speech recognition according to claim 7 is characterized in that: The step of processing the optimized speech signal to obtain speech frames, and extracting speech features from the speech frames to obtain real-time voiceprint feature vectors and real-time speech feature vectors includes: According to the optimized speech signal, the speech signal is framed by a preset framing method to obtain corresponding speech frames; Extracting a corresponding real-time voiceprint feature vector according to the speech frame; Processing is performed according to the speech frame to obtain corresponding Mel-frequency cepstral coefficients and perceptual linear prediction coefficients, and combined processing is performed according to the Mel-frequency cepstral coefficients and the perceptual linear prediction coefficients to obtain a real-time speech feature vector corresponding to the real-time voiceprint feature vector.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program for an intelligent call shorthand method based on speech recognition. When the program for an intelligent call shorthand method based on speech recognition is executed by a processor, the steps of an intelligent call shorthand method based on speech recognition as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multi-person voiceprint recognition method and device based on telephone channel

    CN117457008A