A speech prosody recognition method, system, device and storage medium
By performing signal preprocessing and feature extraction on customer service call voice data, combined with prosody models and template threshold judgment, the problem of high error rate in dialect speech recognition was solved, and efficient speech prosody recognition was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING GARUI INTELLIGENT TECH GRP CO LTD
- Filing Date
- 2022-11-23
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, dialect speech recognition based on deep learning has a high error rate and low recognition efficiency, and requires a large amount of data and deeper model training, which leads to increased costs and business complexity.
By preprocessing customer service call voice signals, extracting feature matrices, and using prosodic models combined with standard Mandarin and dialect template thresholds for recognition, speech prosodic recognition is achieved by using the absolute value of the difference between the spectral feature matrix and the template threshold.
It improved the accuracy and efficiency of dialect speech recognition, reduced recognition costs, and simplified business processes.
Smart Images

Figure CN115762472B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, specifically to a speech prosody recognition method, system, device, and storage medium. Background Technology
[0002] With the development of cloud computing and big data technologies, customer service call centers in the telecommunications industry need to collect voice recordings of conversations during customer service calls, and then perform speech recognition and transcription of the collected voice recordings into text.
[0003] Existing technologies rely on mainstream Mandarin speech recognition, which suffers from high error rates when dealing with speech containing dialects. Current deep learning-based expert knowledge systems require additional computing power to process dialect speech and necessitate the collection and simulation of large amounts of dialect data, which are inherently small samples in real-world scenarios. Traditional deep learning solutions require massive amounts of data and deeper model training, leading to decreased business complexity and model flexibility, increased costs, and ultimately, low efficiency in recognizing speech with dialects. Summary of the Invention
[0004] To address this issue, embodiments of the present invention provide a speech prosody recognition method, system, device, and storage medium to solve the problems of high error rate and low recognition efficiency in existing technologies for speech recognition with dialects.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] According to a first aspect of the present invention, a speech prosody recognition method is provided, the method comprising:
[0007] The voice recordings from customer service calls are captured to obtain the conversational voice signal.
[0008] The dialogue voice signal is preprocessed to obtain a preprocessed dialogue voice file;
[0009] Feature extraction is performed on the preprocessed dialogue speech file, and the preprocessed dialogue speech file is vectorized based on the feature extraction results to obtain the corresponding feature matrix.
[0010] Based on the signal features of historical speech data, a prosodic model is generated. The feature matrix is then input into the prosodic model to obtain the model calculation results.
[0011] Based on the preset Mandarin template and the preset dialect template, the threshold values for the Mandarin template and the dialect template are calculated respectively.
[0012] Using the model calculation results, the Mandarin template threshold, and the dialect template threshold, the prosody recognition result corresponding to the feature matrix is obtained;
[0013] Based on the prosody recognition results, the feature matrix is processed by text mapping to obtain the dialogue word order text.
[0014] Further, the dialogue voice signal is preprocessed to obtain a preprocessed dialogue voice file, including:
[0015] The dialogue voice signal is subjected to a first beamforming process to obtain a first preprocessed signal;
[0016] The first preprocessed signal is subjected to a second beamforming process to obtain a second preprocessed signal;
[0017] The second preprocessed signal is used for spectrum signal control processing to obtain the dialogue voice file.
[0018] Further, feature extraction is performed on the preprocessed dialogue speech file, and the preprocessed dialogue speech file is vectorized based on the feature extraction results to obtain the corresponding feature matrix, including:
[0019] The dialogue audio file is segmented based on time sequence to obtain the segmented dialogue audio file;
[0020] Feature extraction is performed on each segment of the dialogue speech file to obtain the speech spectrum features corresponding to the segmented dialogue speech file. The speech spectrum features include spectral weight parameters. Signal delay parameters Dialect intensity parameters Where n is a positive integer greater than or equal to 0 and less than the total number of segments; using the aforementioned spectral weighting parameters The signal delay parameter and the dialect intensity parameters The first spectral feature matrix is calculated. The first spectral feature matrix The calculation formula is:
[0021]
[0022] in, Represents the first spectral feature matrix The nth element in the equation; s is the multivariate nonlinear fitting parameter.
[0023] Furthermore, based on the signal features of historical speech data, a prosodic model is generated. The feature matrix is input into the prosodic model to obtain the model calculation results, including:
[0024] The pitch, intensity, and duration features of the historical speech data are extracted to obtain the pitch feature data, intensity feature data, and duration feature data of the historical speech data.
[0025] The pitch feature data, intensity feature data, and duration feature data are all averaged to obtain the corresponding pitch feature parameters. Sound intensity characteristic parameters Harmonic length characteristic parameters ;
[0026] The first spectral feature matrix A is input into the prosody recognition model to calculate the prosody model calculation result X. The calculation formula for the prosody model calculation result X is as follows:
[0027]
[0028] Where m is determined by the length of the first spectral feature matrix A; j is a preset weighting parameter; and x is a preset parameter.
[0029] Furthermore, based on the preset Mandarin template and the preset dialect template, the Mandarin template threshold and the dialect template threshold are calculated respectively, including:
[0030] The preset Mandarin template audio file is segmented to obtain the segmented Mandarin template audio file;
[0031] Using the segmented Mandarin template speech file, the corresponding second spectral feature matrix B is calculated;
[0032] The Mandarin template threshold is calculated using the second spectral feature matrix B. The Mandarin template threshold The calculation formula is:
[0033]
[0034] in, The length of the second spectral feature matrix B is determined by its length. Represents the second spectral feature matrix The first in One element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented Mandarin template speech file;
[0035] The preset dialect template audio file is segmented to obtain the segmented dialect template audio file;
[0036] Using the segmented dialect speech template file, the corresponding third spectral feature matrix C is calculated;
[0037] The dialect template threshold is calculated using the third spectral feature matrix C. The dialect template threshold The calculation formula is:
[0038]
[0039] in, The length of the third spectral feature matrix C is determined by the length of the matrix. The third spectral feature matrix C represents the first... One element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented dialect template speech file.
[0040] Further, using the model calculation results, the Mandarin template threshold, and the dialect template threshold, the prosodic recognition result corresponding to the feature matrix is obtained, including:
[0041] Calculation results using the prosodic model and Mandarin template threshold The absolute value of the first difference is calculated. The absolute value of the first difference The calculation formula is:
[0042]
[0043] Calculation results using the prosodic model The absolute value of the second difference is calculated by combining the dialect template threshold. The absolute value of the second difference The calculation formula is:
[0044]
[0045] Determine the absolute value of the first difference Is it greater than the absolute value of the second difference? ;
[0046] If the absolute value of the first difference Greater than the absolute value of the second difference Then the first spectral feature matrix The prosody recognition result is dialect;
[0047] If the absolute value of the first difference Less than or equal to the absolute value of the second difference Then the first spectral feature matrix The prosody recognition result is Mandarin.
[0048] Further, based on the prosody recognition result, the feature matrix is subjected to text mapping processing to obtain the dialogue word order text, including:
[0049] Determine the prosody recognition result of the first spectral feature matrix A;
[0050] If the prosody recognition result of the first spectral feature matrix A is Mandarin, then the first spectral feature matrix A is subjected to the first text mapping process to obtain the first dialogue word order text D;
[0051] If the prosody recognition result of the first spectral feature matrix A is dialect, then the first spectral feature matrix A is subjected to a second text mapping process to obtain the second dialogue word order text. .
[0052] According to a second aspect of the present invention, a speech prosody recognition system is provided, the system comprising:
[0053] The voice signal acquisition module is used to acquire the voice from customer service calls and obtain the dialogue voice signal.
[0054] The speech signal preprocessing module is used to preprocess the dialogue speech signal to obtain a preprocessed dialogue speech file.
[0055] The feature matrix mapping module is used to extract features from the preprocessed dialogue speech file, and to vectorize the preprocessed dialogue speech file based on the feature extraction results to obtain the corresponding feature matrix.
[0056] The prosody model module is used to generate a prosody model based on the signal features of historical speech data. The feature matrix is input into the prosody model to obtain the model calculation result.
[0057] The template threshold generation module is used to calculate the Mandarin template threshold and the dialect template threshold respectively based on the preset Mandarin template and the preset dialect template.
[0058] The prosody recognition module is used to obtain the prosody recognition result corresponding to the feature matrix by using the model calculation result, the Mandarin template threshold and the dialect template threshold;
[0059] The word order text mapping module is used to perform text mapping processing on the feature matrix based on the prosody recognition results to obtain the dialogue word order text.
[0060] Further, the dialogue voice signal is preprocessed to obtain a preprocessed dialogue voice file, including:
[0061] The dialogue voice signal is subjected to a first beamforming process to obtain a first preprocessed signal;
[0062] The first preprocessed signal is subjected to a second beamforming process to obtain a second preprocessed signal;
[0063] The second preprocessed signal is used for spectrum signal control processing to obtain the dialogue voice file.
[0064] Further, feature extraction is performed on the preprocessed dialogue speech file, and the preprocessed dialogue speech file is vectorized based on the feature extraction results to obtain the corresponding feature matrix, including:
[0065] The dialogue audio file is segmented based on time sequence to obtain the segmented dialogue audio file;
[0066] Feature extraction is performed on each segment of the dialogue speech file to obtain the speech spectrum features corresponding to the segmented dialogue speech file. The speech spectrum features include spectral weight parameters. Signal delay parameters Dialect intensity parameters , where n is a positive integer greater than or equal to 0 and less than the total number of segments;
[0067] Using the aforementioned spectral weighting parameters The signal delay parameters and the dialect intensity parameters The first spectral feature matrix A is calculated, and the formula for calculating the first spectral feature matrix A is as follows:
[0068]
[0069] in, represents the nth element in the first spectral feature matrix A; s is a multivariate nonlinear fitting parameter.
[0070] Furthermore, based on the signal features of historical speech data, a prosodic model is generated. The feature matrix is input into the prosodic model to obtain the model calculation results, including:
[0071] The pitch, intensity, and duration features of the historical speech data are extracted to obtain the pitch feature data, intensity feature data, and duration feature data of the historical speech data.
[0072] The pitch feature data, intensity feature data, and duration feature data are all averaged to obtain the corresponding pitch feature parameters. Sound intensity characteristic parameters Harmonic length characteristic parameters ;
[0073] The first spectral feature matrix A is input into the prosody recognition model to calculate the prosody model calculation result X. The calculation formula for the prosody model calculation result X is as follows:
[0074]
[0075] Where m is determined by the length of the first spectral feature matrix A; j is a preset weighting parameter; and x is a preset parameter.
[0076] Furthermore, based on the preset Mandarin template and the preset dialect template, the Mandarin template threshold and the dialect template threshold are calculated respectively, including:
[0077] The preset Mandarin template audio file is segmented to obtain the segmented Mandarin template audio file;
[0078] Using the segmented Mandarin template speech file, the corresponding second spectral feature matrix is calculated. ;
[0079] Using the second spectral feature matrix The threshold of the Mandarin template is calculated. The Mandarin template threshold The calculation formula is:
[0080]
[0081] in, The length of the second spectral feature matrix B is determined by its length. Represents the second spectral feature matrix B. element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented Mandarin template speech file;
[0082] The preset dialect template audio file is segmented to obtain the segmented dialect template audio file;
[0083] Using the segmented dialect speech template file, the corresponding third spectral feature matrix is calculated. ;
[0084] The dialect template threshold is calculated using the third spectral feature matrix C. The dialect template threshold The calculation formula is:
[0085]
[0086] in, The length of the third spectral feature matrix C is determined by the length of the matrix. The third spectral feature matrix C represents the first... One element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented dialect template speech file.
[0087] Further, using the model calculation results, the Mandarin template threshold, and the dialect template threshold, the prosodic recognition result corresponding to the feature matrix is obtained, including:
[0088] The result X calculated using the prosodic model and the Mandarin template threshold The absolute value of the first difference is calculated. The absolute value of the first difference The calculation formula is:
[0089]
[0090] The result X calculated using the prosodic model and the dialect template threshold The absolute value of the second difference is calculated. The absolute value of the second difference The calculation formula is:
[0091]
[0092] Determine the absolute value of the first difference Is it greater than the absolute value of the second difference? ;
[0093] If the absolute value of the first difference Greater than the absolute value of the second difference Then the first spectral feature matrix The prosody recognition result is dialect;
[0094] If the absolute value of the first difference Less than or equal to the absolute value of the second difference If the prosody recognition result of the first spectral feature matrix A is Mandarin, then the result is Mandarin.
[0095] Further, based on the prosody recognition result, the feature matrix is subjected to text mapping processing to obtain the dialogue word order text, including:
[0096] Determine the first spectral feature matrix The prosody recognition results;
[0097] If the first spectral feature matrix If the prosody recognition result is Mandarin, then the first spectral feature matrix A is subjected to the first text mapping process to obtain the first dialogue word order text D;
[0098] If the prosody recognition result of the first spectral feature matrix A is dialect, then the first spectral feature matrix A is subjected to a second text mapping process to obtain the second dialogue word order text. .
[0099] According to a third aspect of the present invention, a speech prosody recognition device is provided, the device comprising: a processor and a memory;
[0100] The memory is used to store one or more program instructions;
[0101] The processor is configured to run one or more program instructions to perform the steps of a speech prosody recognition method as described in any of the preceding claims.
[0102] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program, when executed by a processor, implements the steps of the speech prosody recognition method as described in any of the preceding claims.
[0103] The embodiments of the present invention have the following advantages:
[0104] This invention discloses a speech prosody recognition method, system, device, and storage medium. First, the voice from a customer service call is captured to obtain a dialogue speech signal. Then, the dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file. Next, the preprocessed dialogue speech file is vectorized to obtain a corresponding feature matrix. The feature matrix is input into a trained prosody model to obtain the model calculation result. Then, based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model calculation result and template thresholds, the prosody recognition result of the feature matrix is obtained. Finally, based on the prosody recognition result, the feature matrix is processed by text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects. Attached Figure Description
[0105] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0106] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0107] Figure 1 A schematic diagram of the logical structure of a speech prosody recognition system provided in an embodiment of the present invention;
[0108] Figure 2 A schematic flowchart of a speech prosody recognition method provided in an embodiment of the present invention;
[0109] Figure 3 A flowchart illustrating the voice acquisition and signal preprocessing process is provided for embodiments of the present invention.
[0110] Figure 4 A schematic diagram of the feature matrix mapping process provided in an embodiment of the present invention;
[0111] Figure 5 This is a schematic diagram of the prosody recognition process based on a prosody model, provided as an embodiment of the present invention. Detailed Implementation
[0112] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0113] refer to Figure 1 This invention provides a speech prosody recognition system, which specifically includes: a speech signal acquisition module 1, a speech signal preprocessing module 2, a feature matrix mapping module 3, a prosody model module 4, a template threshold generation module 5, a prosody recognition module 6, and a word order text mapping module 7.
[0114] Furthermore, the speech signal acquisition module 1 is used to acquire the voice from customer service calls to obtain the dialogue speech signal; the speech signal preprocessing module 2 is used to preprocess the dialogue speech signal to obtain the preprocessed dialogue speech file; the feature matrix mapping module 3 is used to extract features from the preprocessed dialogue speech file, and vectorize the preprocessed dialogue speech file based on the feature extraction results to obtain the corresponding feature matrix; the prosody model module 4 is used to generate a prosody model based on the feature parameters of historical speech data, and input the feature matrix into the prosody model to obtain the model calculation result; the template threshold generation module 5 is used to calculate the Mandarin template threshold and the dialect template threshold according to the preset Mandarin template and the preset dialect template, respectively; the prosody recognition module 6 is used to obtain the prosody recognition result corresponding to the feature matrix using the model calculation result, the Mandarin template threshold and the dialect template threshold; and the word order text mapping module 7 is used to perform text mapping processing on the feature matrix according to the prosody recognition result to obtain the dialogue word order text.
[0115] This invention discloses a speech prosody recognition system. First, it collects audio from customer service phone calls to obtain a dialogue speech signal. Then, it preprocesses the dialogue speech signal to obtain a preprocessed dialogue speech file. Next, it performs vectorization processing on the preprocessed dialogue speech file to obtain a corresponding feature matrix. The feature matrix is then input into a trained prosody model to obtain the model's calculation result. Based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model's calculation result and the template thresholds, the prosody recognition result of the feature matrix is obtained. Finally, based on the prosody recognition result, the feature matrix is processed through text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects.
[0116] Corresponding to the speech prosody recognition system disclosed above, this invention also discloses a speech prosody recognition method. The following describes in detail a speech prosody recognition method disclosed in this invention, in conjunction with the speech prosody recognition system described above.
[0117] refer to Figure 2 The following describes the specific steps of a speech prosody recognition method provided by an embodiment of the present invention.
[0118] The voice signal acquisition module 1 acquires the voice from the customer service call to obtain the dialogue voice signal; the voice signal preprocessing module 2 performs signal preprocessing on the dialogue voice signal to obtain the preprocessed dialogue voice file.
[0119] refer to Figure 3The above steps specifically include: first, using a microphone to collect the conversational voice from the customer service call to obtain the conversational voice signal; then, using a microphone signal amplifier to perform a first beamforming process on the conversational voice signal to obtain a first preprocessed signal, wherein the first beamforming process is a microphone array beamforming process; next, using a microphone signal processor to perform a second beamforming process on the first preprocessed signal to obtain a second preprocessed signal, wherein the second beamforming process is a fixed beamforming process and an adaptive beamforming process; finally, using a microphone signal controller to perform spectrum signal control processing on the second preprocessed signal to process the electromagnetic wave timing signal into an electromagnetic wave frequency signal to obtain the conversational voice file.
[0120] This invention collects voice data from customer service calls and preprocesses the collected voice signals to convert the voice signals into radio wave timing signals, and then converts the radio wave timing signals into electromagnetic wave frequency signals.
[0121] The feature matrix mapping module 3 extracts features from the preprocessed dialogue speech file, and then vectorizes the preprocessed dialogue speech file based on the feature extraction results to obtain the corresponding feature matrix.
[0122] refer to Figure 4 The above steps specifically include: first, segmenting the dialogue speech file based on temporal features, dividing the dialogue speech file into several segments to obtain segmented dialogue speech files; then, extracting features from each segmented dialogue speech file to obtain the speech spectrum features corresponding to the segmented dialogue speech files, wherein the speech spectrum features include spectrum weight parameters. Signal delay parameters Dialect intensity parameters Where n is a positive integer greater than or equal to 0 and less than the total number of segments; using the above-mentioned spectral weighting parameters Signal delay parameters Dialect intensity parameters The first spectral feature matrix A is calculated, and the formula for calculating the first spectral feature matrix A is:
[0123]
[0124] in, represents the nth element in the first spectral feature matrix A; s is the multivariate nonlinear fitting parameter.
[0125] This invention, through vectorization of the preprocessed dialogue speech file, maps the dialogue speech file into a spectral feature matrix based on its signal characteristics.
[0126] The prosody model module 4 generates a prosody model based on the feature parameters of historical speech data. The feature matrix is then input into the prosody model to obtain the model calculation results.
[0127] The above steps specifically include: first, extracting the pitch, intensity, and duration features from historical speech data to obtain the pitch, intensity, and duration feature data of the historical speech data; then, calculating the mean of each pitch, intensity, and duration feature data to obtain the corresponding pitch feature parameters. Sound intensity characteristic parameters Harmonic length characteristic parameters The first spectral feature matrix A is input into the prosody recognition model to calculate the prosody model calculation result X. The formula for calculating the prosody model calculation result X is as follows:
[0128]
[0129] Where m represents the duration weight of the first spectral feature matrix, which is determined by the length of the first spectral feature matrix A; j is a preset Dirichlet weighting parameter; and x is a preset parameter.
[0130] The template threshold generation module 5 calculates the Mandarin template threshold and the dialect template threshold respectively based on the preset Mandarin template and the preset dialect template.
[0131] refer to Figure 5 The above steps specifically include: segmenting the preset Mandarin template speech file to obtain segmented Mandarin template speech files; calculating the corresponding second spectral feature matrix B using the segmented Mandarin template speech files; and calculating the Mandarin template threshold using the second spectral feature matrix B. Mandarin template threshold The calculation formula is:
[0132]
[0133] in, From the second spectral feature matrix The length is determined by, Represents the second spectral characteristic matrix The first in One element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented Mandarin template speech file;
[0134] The preset dialect template speech file is segmented to obtain segmented dialect template speech files; the corresponding third spectral feature matrix is calculated using the segmented dialect template speech files. ; Utilizing the third spectral feature matrix The dialect template threshold was calculated. Dialect template threshold The calculation formula is:
[0135]
[0136] in, From the third spectral feature matrix The length is determined by; Represents the third spectral characteristic matrix The first in One element; A positive integer greater than or equal to 0 and less than the total number of segments in the segmented dialect template speech file.
[0137] The prosody recognition module 6 uses the model calculation results, the Mandarin template threshold, and the dialect template threshold to obtain the prosody recognition results corresponding to the feature matrix.
[0138] refer to Figure 5 The above steps specifically include: calculating the results using a prosodic model. and Mandarin template threshold The absolute value of the first difference is calculated. The absolute value of the first difference The calculation formula is:
[0139]
[0140] Results calculated using prosodic models Dialect template threshold The absolute value of the second difference is calculated. Second difference absolute value The calculation formula is:
[0141]
[0142] Determine the absolute value of the first difference Is it greater than the absolute value of the second difference? If the absolute value of the first difference Greater than the absolute value of the second difference If the absolute value of the first difference is..., then the prosody recognition result of the spectral feature matrix is dialect; Less than or equal to the absolute value of the second difference If the spectral feature matrix is then used, the prosody recognition result is Mandarin.
[0143] The word order text mapping module 7 performs text mapping processing on the feature matrix based on the prosody recognition results to obtain the dialogue word order text.
[0144] The above steps specifically include: determining the first spectral feature matrix. The prosody recognition results; if the first spectral feature matrix If the prosody recognition result is Mandarin, then the first spectral feature matrix... Perform the first text mapping process to obtain the first dialogue word order text. If the first spectral feature matrix If the prosody recognition result is dialect, then the first spectral feature matrix... Perform a second text mapping process to obtain the second dialogue word order text. .
[0145] In this embodiment of the invention, the spectral feature matrix is processed using sequence and function transformation methods to extract feature parameters of pitch, duration, and intensity. Then, it is classified and identified as Mandarin and dialects according to threshold comparison. Two different feature encoding methods are then used to map the spectral feature matrix onto word order text, thereby realizing speech classification and recognition of dialects and Mandarin. Finally, the multimodal data of Mandarin and dialects are uniformly encoded, decoded, mapped, and output as a unified word order text.
[0146] This invention discloses a speech prosody recognition method. First, the voice from a customer service call is captured to obtain a dialogue speech signal. Then, the dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file. Next, the preprocessed dialogue speech file is vectorized to obtain a corresponding feature matrix. The feature matrix is input into a trained prosody model to obtain the model calculation result. Then, based on Mandarin and dialect templates, corresponding template thresholds are obtained. Using the model calculation result and the template thresholds, the prosody recognition result of the feature matrix is obtained. Finally, based on the prosody recognition result, the feature matrix is processed through text mapping to obtain the dialogue word order text. This invention effectively improves the recognition accuracy and efficiency for speech containing dialects.
[0147] In addition, embodiments of the present invention also provide a speech prosody recognition device, the device comprising: a processor and a memory; the memory being used to store one or more program instructions; the processor being used to run one or more program instructions to perform the steps of a speech prosody recognition method as described in any of the preceding embodiments.
[0148] In addition, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the speech prosody recognition method as described in any of the preceding claims.
[0149] In this embodiment of the invention, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0150] The various methods, steps, and logic diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor reads information from the storage medium and, in conjunction with its hardware, completes the steps of the above methods.
[0151] The storage medium can be memory, such as volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
[0152] Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.
[0153] Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0154] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0155] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using a combination of hardware and software. When applied as software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of computer programs from one place to another. Storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0156] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for recognizing speech prosody, characterized in that, The method includes: The voice recordings from customer service calls are captured to obtain the conversational voice signal. The dialogue voice signal is preprocessed to obtain a preprocessed dialogue voice file; Feature extraction is performed on the preprocessed dialogue speech file, and the preprocessed dialogue speech file is vectorized based on the feature extraction results to obtain the corresponding feature matrix. Based on the signal features of historical speech data, a prosodic model is generated. The feature matrix is then input into the prosodic model to obtain the model calculation results. Based on the preset Mandarin template and the preset dialect template, the threshold values for the Mandarin template and the dialect template are calculated respectively. Using the model calculation results, the Mandarin template threshold, and the dialect template threshold, the prosody recognition result corresponding to the feature matrix is obtained; Based on the prosody recognition results, the feature matrix is processed by text mapping to obtain the dialogue word order text.
2. The speech prosody recognition method as described in claim 1, characterized in that, The dialogue speech signal is preprocessed to obtain a preprocessed dialogue speech file, including: The dialogue voice signal is subjected to a first beamforming process to obtain a first preprocessed signal; The first preprocessed signal is subjected to a second beamforming process to obtain a second preprocessed signal; The second preprocessed signal is used for spectrum signal control processing to obtain the dialogue voice file.
3. The speech prosody recognition method as described in claim 2, characterized in that, Feature extraction is performed on the preprocessed dialogue speech file. Based on the feature extraction results, the preprocessed dialogue speech file is vectorized to obtain the corresponding feature matrix, including: The dialogue audio file is segmented based on time sequence to obtain the segmented dialogue audio file; Feature extraction is performed on each segment of the dialogue speech file to obtain the speech spectrum features corresponding to the segmented dialogue speech file. The speech spectrum features include the spectrum weight parameter t. n Signal delay parameter y n Dialect intensity parameter τ n , where n is a positive integer greater than or equal to 0 and less than the total number of segments; Using the aforementioned spectral weighting parameter t n The signal delay parameter y n and the dialect intensity parameter τ n The first spectral feature matrix A is calculated, and the formula for calculating the first spectral feature matrix A is as follows: A = {A(n)} A(n)=y n ×s(t n +τ n ) Where A(n) represents the nth element in the first spectral feature matrix A; s is a multivariate nonlinear fitting parameter.
4. The speech prosody recognition method as described in claim 3, characterized in that, Based on the signal features of historical speech data, a prosodic model is generated. The feature matrix is input into the prosodic model to obtain the model calculation results, including: The pitch, intensity, and duration features of the historical speech data are extracted to obtain the pitch feature data, intensity feature data, and duration feature data of the historical speech data. The pitch feature data, the intensity feature data, and the duration feature data are all averaged to obtain the corresponding pitch feature parameter ω, intensity feature parameter θ, and duration feature parameter v. The first spectral feature matrix A is input into the prosody recognition model to calculate the prosody model calculation result X. The calculation formula for the prosody model calculation result X is as follows: Where m is determined by the length of the first spectral feature matrix A; j is a preset weighting parameter; and x is a preset parameter.
5. The speech prosody recognition method as described in claim 4, characterized in that, Based on preset Mandarin templates and preset dialect templates, the Mandarin template threshold and dialect template threshold are calculated respectively, including: The preset Mandarin template audio file is segmented to obtain the segmented Mandarin template audio file; Using the segmented Mandarin template speech file, the corresponding second spectral feature matrix B is calculated; The Mandarin template threshold X is calculated using the second spectral feature matrix B. ′ The Mandarin template threshold X ′ The calculation formula is: Where, m ′ The length of the second spectral feature matrix B is determined by the length of B(n). ′ ) represents the nth element in the second spectral feature matrix B. ′ n elements; ′ A positive integer greater than or equal to 0 and less than the total number of segments in the segmented Mandarin template speech file; The preset dialect template audio file is segmented to obtain the segmented dialect template audio file; Using the segmented dialect speech template file, the corresponding third spectral feature matrix C is calculated; Using the third spectral feature matrix C, the dialect template threshold X is calculated. ″ The dialect template threshold X ″ The calculation formula is: Where, m ″ The length of the third spectral feature matrix C is determined by the length of C(n). ″ ) represents the nth element in the third spectral feature matrix C. ″ n elements; ″ It is a positive integer greater than or equal to 0 and less than the total number of segments in the segmented dialect template speech file.
6. The speech prosody recognition method as described in claim 5, characterized in that, Using the model calculation results, the Mandarin template threshold, and the dialect template threshold, the prosodic recognition result corresponding to the feature matrix is obtained, including: The result X calculated using the prosodic model and the Mandarin template threshold X ′ The absolute value of the first difference, C1, is calculated. The formula for calculating the absolute value of the first difference, C1, is as follows: C1=||X|-|X′|| The result X calculated using the prosodic model and the dialect template threshold X ″ The absolute value of the second difference, C2, is calculated using the following formula: C2=||X|-|X″|| Determine whether the absolute value of the first difference C1 is greater than the absolute value of the second difference C2; If the absolute value of the first difference C1 is greater than the absolute value of the second difference C2, then the prosody recognition result of the first spectral feature matrix A is dialect. If the absolute value of the first difference C1 is less than or equal to the absolute value of the second difference C2, then the prosody recognition result of the first spectral feature matrix A is Mandarin.
7. The speech prosody recognition method as described in claim 6, characterized in that, Based on the prosody recognition results, the feature matrix is processed by text mapping to obtain the dialogue word order text, including: Determine the prosody recognition result of the first spectral feature matrix A; If the prosody recognition result of the first spectral feature matrix A is Mandarin, then the first spectral feature matrix A is subjected to the first text mapping process to obtain the first dialogue word order text D; If the prosody recognition result of the first spectral feature matrix A is dialect, then the first spectral feature matrix A is subjected to a second text mapping process to obtain the second dialogue word order text D. ′ .
8. A speech prosody recognition system, characterized in that, The system includes: The voice signal acquisition module is used to acquire the voice from customer service calls and obtain the dialogue voice signal. The speech signal preprocessing module is used to preprocess the dialogue speech signal to obtain a preprocessed dialogue speech file. The feature matrix mapping module is used to extract features from the preprocessed dialogue speech file, and to vectorize the preprocessed dialogue speech file based on the feature extraction results to obtain the corresponding feature matrix. The prosody model module is used to generate a prosody model based on the signal features of historical speech data. The feature matrix is input into the prosody model to obtain the model calculation result. The template threshold generation module is used to calculate the Mandarin template threshold and the dialect template threshold respectively based on the preset Mandarin template and the preset dialect template. The prosody recognition module is used to obtain the prosody recognition result corresponding to the feature matrix by using the model calculation result, the Mandarin template threshold and the dialect template threshold; The word order text mapping module is used to perform text mapping processing on the feature matrix based on the prosody recognition results to obtain the dialogue word order text.
9. A speech prosody recognition device, characterized in that, The device includes: a processor and a memory; The memory is used to store one or more program instructions; The processor is configured to run one or more program instructions to perform the steps of a speech prosody recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of a speech prosody recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Prosodic information-combined Chinese dialect identification method
CN105810191A
Chinese spoken language distinguishing and synthesis type vocoder
CN1122936A