Real-time user call emotion recognition method and system based on asr
By collecting user call voice data in real time, and combining voice segmentation strategies and deep learning models, the problem of the inability to identify user emotions in real time in existing technologies has been solved, achieving accurate emotion recognition and personalized services, which is applicable to a variety of application scenarios.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING JIZHI TECH CO LTD
- Filing Date
- 2025-04-17
- Publication Date
- 2026-07-30
AI Technical Summary
Existing user call emotion recognition technologies rely on manual annotation or post-event analysis, which cannot capture and respond to users' emotional changes in real time. Furthermore, ASR systems have low recognition accuracy when faced with factors such as background noise, speech rate changes, and dialect accents. The applicability of text emotion analysis in dynamic continuous speech data is also limited.
By collecting user call voice data in real time, the system uses a comprehensive voice segmentation strategy combined with the first and second segmentation coefficients for segmentation processing, combines it with the ASR system for text conversion, and uses a deep learning model for emotion analysis, including the calculation and adjustment of parameters such as speech rate, volume intensity, and pitch change rate.
It enables real-time recognition of user emotions, improving the accuracy and adaptability of emotion recognition, and can provide personalized services in different scenarios, applicable to customer service, smart home and healthcare fields.
Smart Images

Figure CN2025089525_30072026_PF_FP_ABST
Abstract
Description
A method and system for real-time emotion recognition in user calls based on ASR Technical Field
[0001] This invention proposes a real-time emotion recognition method and system for user calls based on ASR, belonging to the field of emotion recognition technology. Background Technology
[0002] In today's communication environment, emotion recognition technology for user calls is gradually becoming one of the key technologies for improving customer service quality and enhancing human-computer interaction. Traditional emotion recognition methods often rely on manual annotation or post-event analysis, which is not only inefficient but also unable to capture and respond to users' emotional changes in real time. With the rapid development of Automatic Speech Recognition (ASR) technology, it has become possible to automatically recognize user call content and analyze user emotions through ASR technology.
[0003] However, directly applying ASR technology to emotion recognition faces numerous challenges. First, voice data during user calls often contains various background noises, speech rate variations, and regional accents, all of which significantly impact the accuracy of ASR systems. Second, even if an ASR system can accurately recognize the speech content and convert it to text, effectively extracting emotional features from this text data and performing accurate analysis remains a complex problem. Most existing text emotion analysis technologies are based on static text data, which has limited applicability to dynamic, continuous voice data such as real-time calls. Summary of the Invention
[0004] This invention provides a method and system for real-time user call emotion recognition based on ASR (Automatic Sentence Recognition) to solve the technical problems of the prior art. The technical solution adopted is as follows:
[0005] A method for real-time user call emotion recognition based on ASR, the method comprising:
[0006] The system collects voice data during user calls in real time, and uses a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to segment the voice data to obtain multiple voice data segments to be recognized.
[0007] The ASR speech recognition system is used to convert the speech data segment to be recognized into text, thereby obtaining the text data corresponding to the speech data segment to be recognized.
[0008] The text data is preprocessed by combining the speech data segment to be recognized to obtain the preprocessed text data.
[0009] The preprocessed text data is input into a trained deep learning model for sentiment analysis to obtain the sentiment recognition results corresponding to the preprocessed text data.
[0010] Furthermore, voice data during user calls is collected in real time, and the voice data is segmented using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to obtain multiple voice data segments to be recognized, including:
[0011] Real-time collection of voice data during user calls;
[0012] The speech data is initially segmented using a preset initial speech rate threshold to obtain the first speech data and the second speech data.
[0013] Wherein, the first speech data is speech data whose speed exceeds a preset initial speech speed threshold; the second speech data is speech data whose speed does not exceed the preset initial speech speed threshold.
[0014] The comprehensive speech segmentation strategy is used in conjunction with the first segmentation coefficient and the second segmentation coefficient to segment the first speech data and the second speech data respectively, so as to obtain the speech data segments to be recognized corresponding to the first speech data and the second speech data.
[0015] Furthermore, the comprehensive speech segmentation strategy includes:
[0016] Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data;
[0017] The first segment coefficient is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data, and the first segment coefficient is used to adjust the preset first segment speech rate threshold to obtain the adjusted first segment speech rate threshold.
[0018] The speech rate in the first speech data is compared with the adjusted first segment speech rate threshold to obtain the first target speech data and the second target speech data.
[0019] Wherein, the first target speech data is the first speech data whose speech rate exceeds the adjusted first segment speech rate threshold; the second target speech data is the first speech data whose speech rate does not exceed the adjusted first segment speech rate threshold;
[0020] Furthermore, the first target speech data and the second target speech data are the speech data segments to be identified.
[0021] Furthermore, the process of obtaining the first piecewise coefficient includes:
[0022] Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data;
[0023] The first segmentation coefficient corresponding to the first speech data is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data;
[0024] The first piecewise coefficient is obtained by the following formula:
[0025] Among them, S 01 denoted as the first segment coefficient; n represents the number of time units contained in the first speech data, and the value of the time unit ranges from 5s to 3min; K i K represents the number of syllables contained in the first speech data in the i-th unit of time; c This indicates a preset reference value for the number of syllables; P si P represents the rate of change of the fundamental tone at the beginning of the i-th unit of time; zi T represents the rate of change of the fundamental tone at the end of the i-th unit of time; di P represents the time length corresponding to the i-th unit of time; maxi P represents the maximum rate of change of the fundamental tone in the i-th unit of time; mini T represents the minimum rate of change of the fundamental tone in the i-th unit of time; mi B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum in the i-th unit of time. maxi B represents the maximum volume intensity occurring in the i-th unit of time; mini B represents the minimum volume intensity occurring in the i-th unit of time; zpi K represents the average volume intensity of all syllables occurring in the i-th unit of time; di This represents the number of syllables that appear during the time interval between the time of maximum volume intensity and the time of minimum volume intensity in the i-th unit of time.
[0026] The first segment speech rate threshold is adjusted using the first segment coefficient to obtain the adjusted first segment speech rate threshold; wherein, the adjusted first segment speech rate threshold is obtained by the following formula:
[0027] Among them, K t01 K represents the adjusted first segment speech rate threshold; 01 S represents the speech rate threshold for the first segment; 01 X represents the first piecewise coefficient; p01 This represents the average slope of the pitch change rate over n units of time corresponding to the first speech data.
[0028] Furthermore, the comprehensive speech segmentation strategy also includes:
[0029] Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data;
[0030] The second segment coefficient is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the second speech data, and the preset second segment speech rate threshold is adjusted by using the second segment coefficient to obtain the adjusted second segment speech rate threshold.
[0031] The speech rate in the second speech data is compared with the adjusted second segment speech rate threshold to obtain the third target speech data and the fourth target speech data.
[0032] The third target speech data is the second speech data whose speech rate exceeds the adjusted second segment speech rate threshold; the fourth target speech data is the second speech data whose speech rate does not exceed the adjusted second segment speech rate threshold.
[0033] Furthermore, the third target speech data and the fourth target speech data are the speech data segments to be identified.
[0034] Furthermore, the process of obtaining the second piecewise coefficient includes:
[0035] Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data;
[0036] The second segmentation coefficients corresponding to the second speech data are obtained by using the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data.
[0037] The second piecewise coefficient is obtained using the following formula:
[0038] Among them, S 02 represents the second segment coefficient; m represents the number of time units contained in the second speech data, and the value of each time unit is 20s-3mjn; K j K represents the number of syllables contained in the second speech data in the j-th unit of time; c This indicates a preset reference value for the number of syllables; T dj P represents the time length corresponding to the j-th unit of time; maxj P represents the maximum rate of change of the fundamental tone in the j-th unit of time; mjnj T represents the minimum rate of change of the fundamental tone in the j-th unit of time; mj B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum at the j-th unit of time. maxj B represents the maximum volume intensity occurring in the j-th unit of time; mjnj B represents the minimum volume intensity occurring in the j-th unit of time;zpj K represents the average volume intensity of all syllables occurring in the j-th unit of time; dj This represents the number of syllables that occur during the time interval between the maximum and minimum volume intensity of the j-th unit of time.
[0039] The second segment speech rate threshold is adjusted using the second segment coefficient to obtain the adjusted second segment speech rate threshold; wherein, the adjusted second segment speech rate threshold is obtained by the following formula:
[0040] Among them, K t02 K represents the adjusted second-segment speech rate threshold; 02 S represents the speech rate threshold for the second segment; 02 X represents the second piecewise coefficient; p02 This represents the average slope of the pitch change rate over m units of time corresponding to the second speech data.
[0041] Furthermore, the ASR speech recognition system is used to perform text conversion on the speech data segment to be recognized, obtaining the text data corresponding to the speech data segment to be recognized, including:
[0042] Extract multiple speech data segments to be recognized;
[0043] The unit time is adjusted to obtain the adjusted unit time length, and the adjusted unit time is used as the unit time period; wherein the value range of the unit time period is 10s-1min;
[0044] Each speech data segment to be recognized is subjected to speech preprocessing to obtain multiple preprocessed speech data segments to be recognized; wherein, the preprocessing includes, but is not limited to, noise suppression processing and echo cancellation processing;
[0045] Retrieve the speech rate parameters corresponding to each preprocessed speech data segment to be recognized;
[0046] Extract the time frame length corresponding to the preset frame segmentation process;
[0047] The overlap ratio is obtained using the speech rate parameter corresponding to each speech data segment to be recognized.
[0048] The overlap ratio is obtained by the following formula:
[0049] Among them, P bI Indicates the overlap ratio; P bl0 This represents the preset overlap ratio baseline value; k represents the number of time segments contained in each segment of speech data to be recognized; K xi K represents the number of syllables corresponding to the i-th time unit;wp K represents the average number of syllables per unit time corresponding to the parent speech data segment for each speech data segment to be recognized; c This indicates a preset reference value for the number of syllables;
[0050] According to the preset frame processing time frame length and overlap ratio, each preprocessed speech data segment to be recognized is processed into frames to obtain the audio data frame corresponding to each speech data segment to be recognized.
[0051] The audio data frame corresponding to each speech data segment to be recognized is input into the ASR speech automatic recognition system to generate text data corresponding to each speech data segment to be recognized.
[0052] Furthermore, the text data is preprocessed in conjunction with the speech data segment to be recognized to obtain preprocessed text data, including:
[0053] Retrieve the text data corresponding to each speech data segment to be recognized;
[0054] Remove leading and several whitespace characters, extra spaces, and special symbols that are not Chinese characters, numbers, or English characters from the text data to obtain the first sample data;
[0055] The first sample data is segmented using a word segmentation tool to obtain the second sample data after word segmentation.
[0056] The second sample data is subjected to error correction and modification processing to obtain the corrected text data.
[0057] The corrected and revised text data refers to the preprocessed text data.
[0058] Furthermore, the deep learning model is a convolutional neural network, and the structure of the deep learning model is as follows:
[0059] The input layer is used to receive preprocessed text data.
[0060] An embedding layer is used to map the index of each word in the text data to a dense vector space of fixed dimension;
[0061] Convolutional layers are used to extract features from the text data and obtain feature maps corresponding to the text data;
[0062] Pooling layers are used to reduce the dimensionality of the feature map;
[0063] The concatenation layer is used to concatenate the dimensionality-reduced feature maps to form feature vectors;
[0064] Fully connected layers are used to integrate features from feature vectors;
[0065] The output layer is used to output the emotion recognition results.
[0066] A real-time emotion recognition system for user calls based on ASR (Automatic Sentence Recognition), the system comprising:
[0067] The voice segmentation module is used to collect voice data during user calls in real time, and to segment the voice data using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to obtain multiple voice data segments to be recognized.
[0068] The text data acquisition module is used to convert the speech data segment to be recognized into text using the ASR speech automatic recognition system, and to acquire the text data corresponding to the speech data segment to be recognized.
[0069] The text data preprocessing module is used to preprocess the text data in conjunction with the speech data segment to be recognized, and obtain the preprocessed text data.
[0070] The emotion recognition module is used to input the preprocessed text data into a trained deep learning model for emotion analysis and to obtain the emotion recognition result corresponding to the preprocessed text data.
[0071] Beneficial effects of this invention:
[0072] This invention proposes a real-time emotion recognition method and system for user calls based on Automatic Speech Recognition (ASR). This method can collect voice data during user calls in real time, perform segmentation processing and text conversion, thereby achieving real-time emotion recognition. This helps to promptly capture changes in user emotions, providing a basis for subsequent decision-making and services. By integrating a voice segmentation strategy with an ASR automatic speech recognition system, this method can more accurately recognize user voice content and convert it into text information. Simultaneously, the application of a deep learning model also improves the accuracy of emotion recognition. This method can be applied to various scenarios, such as customer service, smart homes, and healthcare. By adjusting the segmentation coefficients and the parameters of the deep learning model, it can adapt to the emotion recognition needs of different scenarios. For the needs of different user groups and application scenarios, this method can provide more personalized customization services. Attached Figure Description
[0073] Figure 1 is a flowchart of the method described in this invention;
[0074] Figure 2 is a system block diagram of the system described in this invention. Detailed Implementation
[0075] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0076] This invention proposes a real-time emotion recognition method for user calls based on ASR (Automatic Sentence Recognition), as shown in Figure 1. The real-time emotion recognition method for user calls based on ASR includes:
[0077] S1. Collect voice data during user calls in real time, and use a comprehensive voice segmentation strategy combined with the first segmentation coefficient and the second segmentation coefficient to segment the voice data to obtain multiple voice data segments to be recognized.
[0078] S2. Use the ASR speech recognition system to convert the speech data segment to be recognized into text, and obtain the text data corresponding to the speech data segment to be recognized.
[0079] S3. Preprocess the text data by combining the speech data segment to be recognized to obtain the preprocessed text data;
[0080] S4. Input the preprocessed text data into the trained deep learning model for sentiment analysis and obtain the sentiment recognition result corresponding to the preprocessed text data.
[0081] The working principle of the above technical solution is as follows: First, the method collects voice data during user calls in real time to ensure the timeliness and accuracy of the data. Next, a comprehensive voice segmentation strategy is used, combining first and second segmentation coefficients to segment the voice data. These coefficients may be set based on characteristics such as voice volume, speech rate, and pauses to ensure that the segmented voice data segments more accurately reflect the user's emotional changes. Through segmentation, multiple voice data segments to be recognized are obtained, providing basic data for subsequent steps. An Automatic Speech Recognition (ASR) system is used to convert the voice data segments to be recognized into text. The ASR system can convert voice signals into text information, which is a fundamental step in achieving emotion recognition. In this process, the ASR system needs to overcome various challenges, such as noise interference and accent differences, to ensure the accuracy and reliability of text conversion. The text data is preprocessed in conjunction with the voice data segments to be recognized. Preprocessing steps may include removing irrelevant characters, standardizing text format, and correcting spelling errors. The preprocessed text data is clearer and more accurate, which is beneficial for subsequent emotion analysis steps. The preprocessed text data is then input into a trained deep learning model for emotion analysis. Deep learning models can learn emotional features in text and classify and identify them. Through the computation and analysis of deep learning models, emotion recognition results corresponding to preprocessed text data are obtained. These results may include various emotion categories such as anger, sadness, and happiness.
[0082] The above technical solution achieves the following results: This method can collect voice data during user calls in real time, perform segmentation processing and text conversion, thereby enabling real-time emotion recognition. This helps to promptly capture changes in user emotions, providing a basis for subsequent decision-making and services. By integrating a voice segmentation strategy with an ASR (Automatic Speech Recognition) system, this method can more accurately identify user voice content and convert it into text information. Simultaneously, the application of a deep learning model also improves the accuracy of emotion recognition. This method can be applied to various scenarios, such as customer service, smart homes, and healthcare. By adjusting the segmentation coefficients and parameters of the deep learning model, it can adapt to the emotion recognition needs of different scenarios. For the needs of different user groups and application scenarios, this method can provide more personalized customized services. For example, in customer service, service strategies and processes can be adjusted and optimized based on the user's emotional state; in smart homes, the operating modes and settings of home devices can be adjusted based on changes in the user's emotions.
[0083] In summary, the ASR-based real-time emotion recognition method for user calls offers advantages such as real-time performance, accuracy, flexibility, and personalization, providing users with a more intelligent, convenient, and efficient service experience.
[0084] In one embodiment of the present invention, voice data during a user's call is collected in real time, and the voice data is segmented using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to obtain multiple voice data segments to be recognized, including:
[0085] S101. Real-time collection of voice data during user calls;
[0086] S102. Initially segment the speech data using a preset initial speech rate threshold to obtain first speech data and second speech data; wherein, the first speech data is speech data whose speed exceeds the preset initial speech rate threshold; and the second speech data is speech data whose speed does not exceed the preset initial speech rate threshold.
[0087] S103. Using the comprehensive speech segmentation strategy, the first speech data and the second speech data are segmented by combining the first segmentation coefficient and the second segmentation coefficient to obtain the speech data segments to be recognized corresponding to the first speech data and the second speech data.
[0088] The working principle of the above technical solution is as follows: During a user's call, the system captures and collects voice data in real time. This step is the foundation for all subsequent processing, ensuring the integrity and real-time nature of the voice data. The system presets an initial speech rate threshold, which is used to distinguish between fast and slow speech data. The system then segments the speech data into two parts: first, speech rate exceeding the threshold, and second, speech rate not exceeding the threshold. This initial segmentation facilitates more refined processing of speech data at different speeds.
[0089] For the first and second speech data, the system applies a comprehensive speech segmentation strategy. This strategy combines two segmentation coefficients (a first segmentation coefficient and a second segmentation coefficient), which may be adjusted based on specific characteristics of the speech data (such as volume, pitch changes, pauses, etc.). By applying these coefficients and strategies, the system further segments the first and second speech data into multiple speech data segments to be recognized.
[0090] The above technical solution achieves the following effects: By initially segmenting speech data according to speech rate and then further refining it using a comprehensive speech segmentation strategy, the system can more accurately recognize speech content. This segmentation method helps reduce recognition errors caused by changes in speech rate. Real-time acquisition and processing of speech data ensures the system can quickly respond to and process speech information during user calls. Segmentation processing reduces the amount of data per recognition, helping to improve processing speed and efficiency. By setting and adjusting speech rate thresholds and segmentation coefficients, the system can adapt to the speech characteristics and habits of different users. This adaptability allows the system to maintain high recognition performance and user experience in various scenarios. This technical solution can handle various speech changes during user calls, such as speech rate and volume. This enables the system to maintain stable recognition performance in noisy environments, different languages, or dialects, and other complex scenarios.
[0091] In summary, this technical solution achieves accurate, efficient, and highly adaptable processing of user call voice data by acquiring voice data in real time, performing preliminary segmentation using an initial speech rate threshold, and applying a comprehensive voice segmentation strategy for fine segmentation.
[0092] In one embodiment of the present invention, the integrated speech segmentation strategy includes:
[0093] Step 1a: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data;
[0094] Step 2a: Obtain the first segment coefficient using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data, and adjust the preset first segment speech rate threshold using the first segment coefficient to obtain the adjusted first segment speech rate threshold.
[0095] Step 3a: Compare the speech rate in the first speech data with the adjusted first segment speech rate threshold to obtain the first target speech data and the second target speech data;
[0096] Wherein, the first target speech data is the first speech data whose speech rate exceeds the adjusted first segment speech rate threshold; the second target speech data is the first speech data whose speech rate does not exceed the adjusted first segment speech rate threshold;
[0097] Furthermore, the first target speech data and the second target speech data are the speech data segments to be identified.
[0098] The working principle of the above technical solution is as follows: First, speech rate parameters, volume intensity, and pitch change rate are retrieved from the first speech data. These parameters reflect the dynamic characteristics and prosodic changes of the speech data. Using the retrieved speech rate parameters, volume intensity, and pitch change rate, a first segmentation coefficient is calculated through a certain algorithm or model. Then, this first segmentation coefficient is used to adjust the preset first segment speech rate threshold to obtain an adjusted first segment speech rate threshold that better matches the characteristics of the current speech data. The speech rate in the first speech data is compared with the adjusted first segment speech rate threshold. Based on the comparison result, the speech data is divided into two parts: the part with a speech rate exceeding the adjusted threshold is taken as the first target speech data, and the part with a speech rate not exceeding the adjusted threshold is taken as the second target speech data.
[0099] The above technical solution achieves the following results: by introducing multiple speech parameters (speech rate, volume intensity, and pitch change rate) and calculating segmentation coefficients, this strategy can more accurately reflect the characteristics of speech data, thereby improving the accuracy of segmentation. Since this strategy can adjust the segmentation speech rate threshold according to the actual situation of the speech data, it has strong adaptability to speech data with different speech rates, volumes, and pitch changes. The segmented speech data (first target speech data and second target speech data) can be more easily processed subsequently, such as speech recognition and speech synthesis. Especially in speech recognition, segmentation helps improve the accuracy and efficiency of recognition. Through segmentation, computational and storage resources can be allocated more rationally. For example, when processing large segments of continuous speech data, it can be divided into smaller segments for processing, thereby optimizing the utilization of computational resources.
[0100] In summary, this integrated speech segmentation strategy achieves accurate segmentation of speech data by introducing multiple speech parameters and calculating segmentation coefficients. It has the technical advantages of strong adaptability, convenience for subsequent processing, and optimized resource utilization.
[0101] In one embodiment of the present invention, the process of obtaining the first segmentation coefficient includes:
[0102] Step 201a: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data;
[0103] Step 202a: Obtain the first segmentation coefficient corresponding to the first speech data using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data;
[0104] The first piecewise coefficient is obtained by the following formula:
[0105] Among them, S 01 denoted as the first segment coefficient; n represents the number of time units contained in the first speech data, and the value of the time unit ranges from 5s to 3min; K i K represents the number of syllables contained in the first speech data in the i-th unit of time; c This indicates a preset reference value for the number of syllables; P si P represents the rate of change of the fundamental tone at the beginning of the i-th unit of time; zi T represents the rate of change of the fundamental tone at the end of the i-th unit of time; di P represents the time length corresponding to the i-th unit of time; maxi P represents the maximum rate of change of the fundamental tone in the i-th unit of time; mini T represents the minimum rate of change of the fundamental tone in the i-th unit of time; mi B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum in the i-th unit of time. maxi B represents the maximum volume intensity occurring in the i-th unit of time; mini B represents the minimum volume intensity occurring in the i-th unit of time; zpi K represents the average volume intensity of all syllables occurring in the i-th unit of time; di This represents the number of syllables that appear during the time interval between the time of maximum volume intensity and the time of minimum volume intensity in the i-th unit of time.
[0106] Step 203a: Adjust the first segment speech rate threshold using the first segment coefficient to obtain the adjusted first segment speech rate threshold;
[0107] The adjusted first segment speech rate threshold is obtained using the following formula:
[0108] Among them, K t01 K represents the adjusted first segment speech rate threshold; 01 S represents the speech rate threshold for the first segment; 01 X represents the first piecewise coefficient; p01This represents the average slope of the pitch change rate over n units of time corresponding to the first speech data.
[0109] The working principle of the above technical solution is as follows: Key information such as speech rate parameters, volume intensity, and pitch change rate are retrieved from the first speech data. For each unit of time (5s-3min), relevant speech features are calculated, including the number of syllables, pitch change rate, and volume intensity. Using these parameters, the first segmentation coefficient is calculated using a specific formula. This formula comprehensively considers multiple aspects of the number of syllables, the pitch change rate (start time, end time, maximum value, minimum value, time interval), and multiple features of the volume intensity (maximum value, minimum value, average value, number of syllables within a specific time period). This calculation method aims to comprehensively reflect the prosody and volume characteristics of the first speech data, providing an accurate basis for subsequent segmentation processing. Using the calculated first segmentation coefficient and the average slope of the pitch change rate of the first speech data, the first segmentation speech rate threshold is adjusted using a formula. The adjusted speech rate threshold better matches the actual situation of the first speech data, contributing to more accurate subsequent segmentation processing.
[0110] The above technical solution achieves the following results: By comprehensively considering multiple parameters such as speech rate, volume intensity, and pitch change rate, and calculating the first segmentation coefficient using a specific formula, this solution can more accurately reflect the characteristics of the first speech data. Furthermore, segmentation using an adjusted speech rate threshold improves segmentation accuracy, providing a more reliable data foundation for subsequent processing (such as speech recognition and speech analysis). This technical solution can be flexibly adjusted according to the characteristics of different speech data, making it suitable for speech data with varying speech rates, volumes, and pitch changes. This adaptability allows the solution to maintain high processing performance in various scenarios. Precise segmentation reduces the computational load and time cost of subsequent processing. Simultaneously, the solution utilizes concepts such as unit time, which helps to further improve processing efficiency while maintaining processing effectiveness. In applications such as speech recognition and speech synthesis, accurate segmentation helps improve the system's recognition accuracy and synthesis quality. This is of great significance for enhancing user experience and increasing user satisfaction.
[0111] On the other hand, by retrieving the speech rate parameters, volume intensity, and pitch change rate of the first speech data, and calculating the first segment coefficient accordingly, this scheme can dynamically reflect the characteristic changes of the speech data. This dynamic adaptability allows the scheme to more accurately capture subtle differences in the speech data, thereby improving the accuracy and reliability of subsequent processing. Adjusting the first segment speech rate threshold using the first segment coefficient makes the speech rate threshold more consistent with the characteristics of the actual speech data. This precise adjustment helps to more accurately determine the speech rate in speech recognition, speech analysis, and other fields, improving the performance of related applications. The scheme considers multiple dimensions of features such as speech rate parameters, volume intensity, and pitch change rate, and calculates the first segment coefficient by integrating these features. This fusion of multi-dimensional features allows the scheme to more comprehensively reflect the characteristics of the speech data, improving the accuracy and robustness of processing. The parameters in the formula (such as the range of values per unit time, reference values for the number of syllables, etc.) can be adjusted according to the actual situation. This parameterized design makes the scheme more flexible. By adjusting these parameters, it can adapt to the speech data processing needs in different scenarios. Because the solution considers multiple dimensions of features and uses parametric design, it maintains relatively stable performance under different speech data. This stability is particularly important for long-term operation and large-scale application scenarios.
[0112] In summary, the technical effects of the above-mentioned solution in terms of performance indicators are mainly reflected in enhanced dynamic adaptability, more precise speech rate threshold adjustment, multi-dimensional feature fusion, improved flexibility through parametric design, enhanced computational efficiency, and improved performance stability. These technical effects give the solution higher application value in fields such as speech recognition and speech analysis. Furthermore, by comprehensively considering multiple speech parameters, calculating the first segmentation coefficient using a specific formula, and adjusting the speech rate threshold accordingly for segmentation processing, this solution achieves improved segmentation accuracy, enhanced adaptability, optimized processing efficiency, and improved user experience.
[0113] In one embodiment of the present invention, the integrated speech segmentation strategy further includes:
[0114] Step 1b: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data;
[0115] Step 2b: Obtain the second segment coefficient using the speech rate parameter, volume intensity and pitch change rate corresponding to the second speech data, and adjust the preset second segment speech rate threshold using the second segment coefficient to obtain the adjusted second segment speech rate threshold;
[0116] Step 3b: Compare the speech rate in the second speech data with the adjusted second segment speech rate threshold to obtain the third target speech data and the fourth target speech data;
[0117] The third target speech data is the second speech data whose speech rate exceeds the adjusted second segment speech rate threshold; the fourth target speech data is the second speech data whose speech rate does not exceed the adjusted second segment speech rate threshold.
[0118] Furthermore, the third target speech data and the fourth target speech data are the speech data segments to be identified.
[0119] The working principle of the above technical solution is as follows: Speech rate parameters, volume intensity, and pitch change rate are retrieved from the second speech data. These parameters are fundamental characteristics of the speech signal, reflecting the characteristics and changes of speech. Using these parameters, a second segmentation coefficient is calculated through a certain algorithm or model. This coefficient is a factor used to adjust the preset second segment speech rate threshold. The calculated second segmentation coefficient is used to adjust the preset second segment speech rate threshold to obtain a speech rate threshold that better matches the characteristics of the current speech data. The adjusted second segment speech rate threshold is compared with the actual speech rate in the second speech data. Based on the comparison result, the second speech data is divided into two parts: the part where the speech rate exceeds the adjusted threshold (third target speech data) and the part where the speech rate does not exceed the adjusted threshold (fourth target speech data).
[0120] The above technical solution achieves the following results: By considering multiple factors such as speech rate, volume intensity, and pitch change rate, this solution can more accurately reflect the characteristics of speech data, thereby improving the accuracy of speech segmentation. This solution can dynamically adjust the segmented speech rate threshold according to different speech data characteristics, thus exhibiting strong adaptability. This helps in processing speech data with varying speech rates, volumes, and pitch changes, improving the flexibility of speech processing. After dividing the speech data into third-target speech data and fourth-target speech data, different processing strategies can be adopted for these two parts. For example, for faster-paced segments, more complex speech recognition algorithms or additional preprocessing may be required; while for slower-paced segments, relatively simpler processing strategies can be used. This helps optimize the speech processing flow and improve overall processing efficiency. Accurate speech segmentation provides more accurate and representative speech data for subsequent speech recognition tasks. This helps reduce the false recognition rate in speech recognition and improve recognition accuracy.
[0121] In summary, this comprehensive speech segmentation strategy achieves accurate segmentation of speech data by considering multiple speech feature parameters and dynamically adjusting the segmentation speech rate threshold. This not only improves the flexibility and efficiency of speech processing but also provides more accurate data support for subsequent speech recognition tasks.
[0122] In one embodiment of the present invention, the process of obtaining the second piecewise coefficient includes:
[0123] Step 201b: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data;
[0124] Step 202b: Obtain the second segmentation coefficients corresponding to the second speech data using the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data;
[0125] The second piecewise coefficient is obtained using the following formula:
[0126] Among them, S 02 represents the second segment coefficient; m represents the number of time units contained in the second speech data, and the value of each time unit is 20s-3mjn; K j K represents the number of syllables contained in the second speech data in the j-th unit of time; c This indicates a preset reference value for the number of syllables; T dj P represents the time length corresponding to the j-th unit of time; maxj P represents the maximum rate of change of the fundamental tone in the j-th unit of time; mjnj T represents the minimum rate of change of the fundamental tone in the j-th unit of time; mj B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum at the j-th unit of time. maxj B represents the maximum volume intensity occurring in the j-th unit of time; mjnj B represents the minimum volume intensity occurring in the j-th unit of time; zpj K represents the average volume intensity of all syllables occurring in the j-th unit of time; dj This represents the number of syllables that occur during the time interval between the maximum and minimum volume intensity of the j-th unit of time.
[0127] Step 203b: Adjust the second segment speech rate threshold using the second segment coefficient to obtain the adjusted second segment speech rate threshold;
[0128] The adjusted second segment speech rate threshold is obtained using the following formula:
[0129] Among them, K t02 K represents the adjusted second-segment speech rate threshold; 02 S represents the speech rate threshold for the second segment; 02 X represents the second piecewise coefficient; p02 This represents the average slope of the pitch change rate over m units of time corresponding to the second speech data.
[0130] The working principle of the above technical solution is as follows: Key parameters, including speech rate, volume intensity, and pitch change rate, are extracted from the second speech data. These parameters reflect the prosodic features, volume variations, and dynamic characteristics of the fundamental frequency. Using the extracted parameters, the second segmentation coefficients are calculated using a specific formula. This formula comprehensively considers multiple factors, including the number of syllables per unit time, the maximum and minimum values of the pitch change rate, the maximum and minimum values of the volume intensity, and the temporal distribution of these features (such as the time interval between the maximum and minimum values of the pitch change rate, and the number of syllables occurring within the time interval between the maximum and minimum values of the volume intensity). The combination of these factors aims to comprehensively evaluate the prosodic and volume characteristics of the second speech data, thereby providing an accurate basis for calculating the segmentation coefficients. Using the calculated second segmentation coefficients and the average slope of the pitch change rate of the second speech data, the second segmentation speech rate threshold is adjusted using a formula. The adjusted speech rate threshold better reflects the actual characteristics of the second speech data, contributing to the accuracy and effectiveness of subsequent speech segmentation processing.
[0131] The above technical solution achieves the following results: By comprehensively considering multiple speech feature parameters and calculating the second segmentation coefficient using a specific formula, this solution can more accurately reflect the prosody and volume characteristics of the second speech data. Furthermore, segmentation using an adjusted speech rate threshold significantly improves segmentation accuracy, providing a more reliable data foundation for subsequent processing (such as speech recognition and speech analysis). This technical solution can be flexibly adjusted according to the characteristics of different speech data, making it suitable for speech data with varying speech rates, volumes, and pitch variations. This adaptability allows the solution to maintain high processing performance in various scenarios, improving the versatility and practicality of speech processing. Precise segmentation helps reduce the computational load and time cost of subsequent processing, improving overall processing efficiency. Simultaneously, this solution utilizes concepts such as unit time, making the processing process clearer and more orderly, which helps optimize the processing flow. In applications such as speech recognition and speech synthesis, accurate segmentation helps improve the system's recognition accuracy and synthesis quality. This is of great significance for improving user experience and enhancing user satisfaction, and helps promote the development and application of speech technology.
[0132] On the other hand, by retrieving detailed parameters such as speech rate, volume intensity, and pitch change rate of the second speech data, this scheme can process and analyze speech data more precisely. This refined processing helps capture subtle changes in speech, improving the accuracy and reliability of subsequent processing. The second segment speech rate threshold is dynamically adjusted using the second segment coefficient, making the speech rate threshold more consistent with the characteristics of actual speech data. This dynamic adjustment mechanism helps to more accurately determine the speech rate in speech recognition, speech analysis, and other fields, improving the performance of related applications. The scheme integrates features from multiple dimensions, such as speech rate parameters, volume intensity, and pitch change rate, and calculates the second segment coefficient by combining these features. This fusion of multi-dimensional features allows the scheme to more comprehensively reflect the characteristics of speech data, improving processing accuracy and robustness. Parameters in the formula (such as the range of values per unit time, reference values for the number of syllables, etc.) can be adjusted according to actual conditions; this parameterized design makes the scheme more flexible. By adjusting these parameters, it can adapt to the speech data processing needs of different scenarios, improving the applicability of the scheme. Meanwhile, because the solution considers multiple dimensions of features and parametric design, it maintains relatively stable performance under different speech data, improving the reliability for long-term operation and large-scale applications. This technical solution can process speech data of varying lengths, speech rates, volume intensities, and pitch variability, demonstrating strong adaptability. This adaptability enables the solution to be widely applied in various fields of speech recognition and speech analysis.
[0133] In summary, the technical advantages of the above-mentioned solution in terms of performance indicators are mainly reflected in refined processing of speech data, dynamic adjustment of speech rate thresholds, multi-dimensional feature fusion to improve accuracy, parametric design to enhance flexibility, computational efficiency and performance stability, and strong adaptability. These technical advantages give the solution higher application value in fields such as speech recognition and speech analysis, and can meet the speech data processing needs in different scenarios. Furthermore, by comprehensively considering multiple speech feature parameters, calculating the second segmentation coefficient using a specific formula, and adjusting the speech rate threshold accordingly for segmentation processing, this solution achieves improved segmentation accuracy, enhanced adaptability, optimized processing flow, and improved user experience. These effects collectively promote the advancement of speech processing technology and the expansion of its application scope.
[0134] In one embodiment of the present invention, an ASR (Automatic Speech Recognition) system is used to perform text conversion on a speech data segment to be recognized, and the text data corresponding to the speech data segment to be recognized is obtained, including:
[0135] S201. Extract multiple speech data segments to be recognized;
[0136] S202. Adjust the unit time, obtain the adjusted unit time length, and use the adjusted unit time as the unit time period; wherein, the value range of the unit time period is 10s-1min;
[0137] S203. Perform speech preprocessing on each speech data segment to be recognized to obtain multiple preprocessed speech data segments to be recognized; wherein, the preprocessing includes, but is not limited to, noise suppression processing and echo cancellation processing;
[0138] S204. Retrieve the speech rate parameters corresponding to each preprocessed speech data segment to be recognized.
[0139] S205. Extract the time frame length corresponding to the preset frame segmentation process;
[0140] S206. Obtain the overlap ratio using the speech rate parameter corresponding to each speech data segment to be recognized;
[0141] The overlap ratio is obtained by the following formula:
[0142] Among them, P bI Indicates the overlap ratio; P bl0 This represents the preset overlap ratio baseline value; k represents the number of time segments contained in each segment of speech data to be recognized; K xi K represents the number of syllables corresponding to the i-th time unit; wp K represents the average number of syllables per unit time corresponding to the parent speech data segment for each speech data segment to be recognized; c This indicates a preset reference value for the number of syllables;
[0143] S207. Perform frame processing on each preprocessed speech data segment to be recognized according to the preset frame processing time frame length and overlap ratio, and obtain the audio data frame corresponding to each speech data segment to be recognized.
[0144] S208. Input the audio data frame corresponding to each speech data segment to be recognized into the ASR speech automatic recognition system to generate text data corresponding to each speech data segment to be recognized.
[0145] The working principle of the above technical solution is as follows: Multiple speech data segments to be recognized are extracted from the raw speech data. These segments are the basic units for subsequent processing. The unit time is adjusted to determine a suitable unit time interval length for subsequent framing and processing. The value of this unit time interval ranges from 10 seconds to 1 minute and can be flexibly set according to specific circumstances. Preprocessing is performed on each speech data segment to be recognized, including noise suppression and echo cancellation, to improve speech quality and reduce recognition errors. The speech rate parameters corresponding to each preprocessed speech data segment to be recognized are obtained. These parameters reflect the speed of speech and have a significant impact on subsequent framing and processing. A preset time frame length for framing processing is established, which determines the duration of each audio data frame. Using the speech rate parameters corresponding to each speech data segment to be recognized, combined with a preset overlap ratio benchmark value, the number of syllables within a unit time interval, the average number of syllables per unit time of the superior speech data segment, and a preset syllable count reference value, an overlap ratio is calculated. This ratio guides subsequent frame segmentation to ensure sufficient overlap between adjacent frames, thereby improving recognition accuracy. Based on the preset time frame length and the calculated overlap ratio, each speech data segment to be recognized is segmented into multiple audio data frames. These frames are the basic units for text conversion in the ASR speech recognition system. The audio data frames corresponding to each speech data segment to be recognized are input into the ASR speech recognition system to generate corresponding text data. This system utilizes advanced speech recognition algorithms to convert audio data into text data.
[0146] The above technical solution achieves the following results: Through meticulous frame segmentation and overlap ratio calculation, it ensures a certain overlap between adjacent frames, thereby improving the accuracy of speech recognition. This overlap processing significantly reduces the error rate, especially in cases of fast speech or poor speech quality. The solution can be flexibly adjusted according to the characteristics of different speech data, such as the length of the unit time period, the length of the time frame, and the overlap ratio. This adaptability makes it applicable to various speech data types, enhancing its versatility and practicality. Through reasonable preprocessing and frame segmentation steps, the solution significantly reduces the computational load and time cost of subsequent processing. Simultaneously, utilizing an ASR (Automatic Speech Recognition) system for text conversion achieves an automated and efficient processing flow. In speech recognition and other applications, accurate text conversion results are crucial for user experience. By improving recognition accuracy and optimizing the processing flow, this solution provides users with more accurate and efficient speech recognition services, thereby enhancing user satisfaction and experience.
[0147] On the other hand, by performing detailed preprocessing on the speech data segments to be recognized, such as noise suppression and echo cancellation, the impact of background noise and echo on speech recognition can be effectively reduced, thereby improving recognition accuracy. The scheme considers speech rate parameters and calculates the overlap ratio based on these parameters, then performs frame-based processing on the speech data. This speech rate-based adaptive frame-based processing can better capture key information in the speech, helping to improve recognition accuracy. This technical solution can handle speech data under different speech rates and noise environments, demonstrating strong adaptability. This adaptability enables the ASR (Automatic Speech Recognition) system to maintain stable performance in complex and changing speech environments. By adjusting parameters such as unit time length and overlap ratio, this scheme can flexibly handle speech data segments of different lengths, further enhancing the system's robustness. The scheme employs a frame-based processing method, cutting the speech data into smaller data frames for processing. This method can process multiple data frames in parallel, thereby improving processing efficiency. Furthermore, through reasonable parameter settings and algorithm optimization, this scheme can improve computational speed while ensuring accuracy, meeting the needs of applications with high real-time requirements. This technical solution reduces unnecessary computation through refined preprocessing and frame segmentation, thereby optimizing resource utilization. In practical applications, this helps reduce system energy consumption and costs, improving overall economic efficiency. Accurate speech recognition and efficient processing speed can significantly enhance the user experience. Whether in smart homes, in-vehicle navigation, or other voice interaction scenarios, users can enjoy a smoother and more natural voice interaction experience.
[0148] In summary, the technical effects of the above-mentioned solution in terms of performance indicators are mainly reflected in improving the accuracy of speech recognition, enhancing system robustness, increasing processing efficiency, optimizing resource utilization, and improving user experience. These technical effects make this solution have broad application prospects and significant practical value in the field of speech recognition. Furthermore, through refined frame segmentation processing, overlap ratio calculation, and the application of an ASR (Automatic Speech Recognition) system, this solution achieves improved recognition accuracy, enhanced adaptability, optimized processing flow, and improved user experience. These effects collectively promote the development of speech recognition technology and the expansion of its application scope.
[0149] In one embodiment of the present invention, the text data is preprocessed in conjunction with the speech data segment to be recognized to obtain preprocessed text data, including:
[0150] S301. Retrieve the text data corresponding to each speech data segment to be recognized;
[0151] S302. Remove leading and trailing whitespace characters, extra spaces, and special symbols that are not Chinese characters, numbers, or English characters from the text data to obtain the first sample data;
[0152] S303. Use a word segmentation tool to segment the first sample data to obtain the segmented second sample data.
[0153] S304. Perform error correction and modification processing on the second sample data to obtain the corrected and modified text data.
[0154] The corrected and revised text data refers to the preprocessed text data.
[0155] The working principle of the above technical solution is as follows: Text data, obtained after conversion by the ASR (Automatic Speech Recognition) system, is retrieved from each speech data segment to be recognized from previous processing steps or storage. This text data forms the basis for subsequent preprocessing. The text data undergoes initial cleaning, removing leading and trailing whitespace characters (such as spaces, tabs, etc.) and redundant spaces. Simultaneously, special symbols other than Chinese characters, numbers, and English words are removed, as these may be misidentification markers or irrelevant marks that could interfere with subsequent processing. After this step, the first sample data is obtained. The first sample data is then segmented using a word segmentation tool. Word segmentation is a crucial step in natural language processing, dividing continuous text into individual words or phrases. The result of word segmentation directly affects the accuracy of subsequent text analysis, understanding, and processing. After word segmentation, the second sample data is obtained. The second sample data undergoes error correction and refinement. This step may include grammatical error correction, spell checking, and semantic understanding, aiming to further improve the accuracy and readability of the text data. The corrected and revised text data is the preprocessed text data, which can be used for subsequent text analysis, information extraction, sentiment analysis and other applications.
[0156] The above technical solution significantly improves the quality of text data by removing whitespace, redundant spaces, and irrelevant special symbols, as well as performing word segmentation and error correction. This helps reduce noise and interference in subsequent processing, improving accuracy and efficiency. Preprocessed text data is cleaner, more standardized, and easier to read and understand. This is particularly important for scenarios requiring human intervention or review, reducing staff workload and improving efficiency. Preprocessed text data provides a reliable foundation for subsequent natural language processing tasks (such as text classification, sentiment analysis, and information extraction). These tasks rely on high-quality text data, and the preprocessing steps ensure data accuracy and consistency. In applications such as speech recognition and text analysis, high-quality text data enhances user experience. For example, in intelligent customer service systems, accurate text conversion and preprocessing enable users to have a smoother and more accurate interactive experience.
[0157] In summary, this technical solution improves the quality, readability, and accuracy of text data through meticulous text preprocessing steps, providing a reliable foundation for natural language processing tasks and contributing to enhanced user experience.
[0158] In one embodiment of the present invention, the deep learning model is a convolutional neural network, and the structure of the deep learning model is as follows:
[0159] The input layer is used to receive preprocessed text data.
[0160] An embedding layer is used to map the index of each word in the text data to a dense vector space of fixed dimension;
[0161] Convolutional layers are used to extract features from the text data and obtain feature maps corresponding to the text data;
[0162] Pooling layers are used to reduce the dimensionality of the feature map;
[0163] The concatenation layer is used to concatenate the dimensionality-reduced feature maps to form feature vectors;
[0164] Fully connected layers are used to integrate features from feature vectors;
[0165] The output layer is used to output the emotion recognition results.
[0166] The working principle of the above technical solution is as follows: Input layer: This layer receives preprocessed text data. Preprocessing typically includes steps such as word segmentation, stop word removal, lexical reconstruction, or stemming to ensure that the text data is suitable for model processing. The preprocessed text data is usually presented in the form of words or character sequences.
[0167] Embedding Layer: The embedding layer maps the index of each word in the text data to a dense vector space of fixed dimensions. This is typically achieved by looking up a pre-trained word embedding matrix (such as Word2Vec, GloVe, etc.), where each word is represented as a fixed-length vector. These vectors capture the semantic and syntactic relationships between words, laying the foundation for subsequent feature extraction.
[0168] Convolutional layers: Convolutional layers use multiple convolutional kernels of varying sizes to perform convolution operations on the word vector matrix output by the embedding layers, extracting local features of the text. The size of the convolutional kernel determines the range of features captured, and multiple kernels can extract a variety of different features. After the convolution operation, an activation function (such as ReLU) is typically used to increase the non-linearity of the model.
[0169] Pooling layers: Pooling layers reduce the dimensionality of the feature maps output by convolutional layers, thereby reducing the feature dimension while retaining key information. Common pooling operations include max pooling and average pooling. Pooling layers help reduce computational cost while improving the robustness of the model.
[0170] The concatenation layer combines the multiple feature maps output by the pooling layer to form a single feature vector. This feature vector contains global features of the text data, providing strong support for subsequent classification tasks.
[0171] Fully connected layer: The fully connected layer integrates and classifies the feature vectors output by the concatenation layer. This layer performs a linear transformation on the feature vectors through a weight matrix and a bias term, and outputs the probability distribution of each class through an activation function (such as softmax).
[0172] Output layer: Based on the output of the fully connected layers, the output layer provides the final emotion recognition result. Typically, the output layer selects the category with the highest probability as the recognition result.
[0173] The above technical solution achieves the following results: by combining convolutional and pooling layers, the model can efficiently extract local and global features from text data, providing strong support for emotion recognition. Because convolutional neural networks have powerful feature extraction capabilities, this model can handle text data from different domains and has good generalization ability. Compared to some complex deep learning models, the structure of convolutional neural networks is relatively simple and clear, making its output easier to interpret and understand. The model structure can be flexibly adjusted, such as changing the size and number of convolutional kernels and adjusting the pooling method, to adapt to different text classification tasks and datasets.
[0174] In summary, the described deep learning model structure achieves efficient feature extraction and sentiment recognition of text data through the collaborative work of the input layer, embedding layer, convolutional layer, pooling layer, concatenation layer, fully connected layer, and output layer. This model possesses advantages such as strong generalization ability, good interpretability, and strong adaptability, and has broad application prospects in text classification and sentiment analysis.
[0175] This invention proposes a real-time user call emotion recognition system based on ASR, as shown in Figure 2. The real-time user call emotion recognition system based on ASR includes:
[0176] The voice segmentation module is used to collect voice data during user calls in real time, and to segment the voice data using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to obtain multiple voice data segments to be recognized.
[0177] The text data acquisition module is used to convert the speech data segment to be recognized into text using the ASR speech automatic recognition system, and to acquire the text data corresponding to the speech data segment to be recognized.
[0178] The text data preprocessing module is used to preprocess the text data in conjunction with the speech data segment to be recognized, and obtain the preprocessed text data.
[0179] The emotion recognition module is used to input the preprocessed text data into a trained deep learning model for emotion analysis and to obtain the emotion recognition result corresponding to the preprocessed text data.
[0180] The working principle of the above technical solution is as follows: First, the method collects voice data during user calls in real time to ensure the timeliness and accuracy of the data. Next, a comprehensive voice segmentation strategy is used, combining first and second segmentation coefficients to segment the voice data. These coefficients may be set based on characteristics such as voice volume, speech rate, and pauses to ensure that the segmented voice data segments more accurately reflect the user's emotional changes. Through segmentation, multiple voice data segments to be recognized are obtained, providing basic data for subsequent steps. An Automatic Speech Recognition (ASR) system is used to convert the voice data segments to be recognized into text. The ASR system can convert voice signals into text information, which is a fundamental step in achieving emotion recognition. In this process, the ASR system needs to overcome various challenges, such as noise interference and accent differences, to ensure the accuracy and reliability of text conversion. The text data is preprocessed in conjunction with the voice data segments to be recognized. Preprocessing steps may include removing irrelevant characters, standardizing text format, and correcting spelling errors. The preprocessed text data is clearer and more accurate, which is beneficial for subsequent emotion analysis steps. The preprocessed text data is then input into a trained deep learning model for emotion analysis. Deep learning models can learn emotional features in text and classify and identify them. Through the computation and analysis of deep learning models, emotion recognition results corresponding to preprocessed text data are obtained. These results may include various emotion categories such as anger, sadness, and happiness.
[0181] The above technical solution achieves the following results: This method can collect voice data during user calls in real time, perform segmentation processing and text conversion, thereby enabling real-time emotion recognition. This helps to promptly capture changes in user emotions, providing a basis for subsequent decision-making and services. By integrating a voice segmentation strategy with an ASR (Automatic Speech Recognition) system, this method can more accurately identify user voice content and convert it into text information. Simultaneously, the application of a deep learning model also improves the accuracy of emotion recognition. This method can be applied to various scenarios, such as customer service, smart homes, and healthcare. By adjusting the segmentation coefficients and parameters of the deep learning model, it can adapt to the emotion recognition needs of different scenarios. For the needs of different user groups and application scenarios, this method can provide more personalized customized services. For example, in customer service, service strategies and processes can be adjusted and optimized based on the user's emotional state; in smart homes, the operating modes and settings of home devices can be adjusted based on changes in the user's emotions.
[0182] In summary, the ASR-based real-time emotion recognition method for user calls offers advantages such as real-time performance, accuracy, flexibility, and personalization, providing users with a more intelligent, convenient, and efficient service experience.
[0183] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for real-time user call emotion recognition based on ASR, characterized in that, The ASR-based real-time user call emotion recognition method includes: The system collects voice data during user calls in real time, and uses a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to segment the voice data to obtain multiple voice data segments to be recognized. The ASR speech recognition system is used to convert the speech data segment to be recognized into text, thereby obtaining the text data corresponding to the speech data segment to be recognized. The text data is preprocessed by combining the speech data segment to be recognized to obtain the preprocessed text data. The preprocessed text data is input into a trained deep learning model for sentiment analysis to obtain the sentiment recognition results corresponding to the preprocessed text data.
2. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, Real-time acquisition of voice data during user calls, and segmentation of the voice data using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient, to obtain multiple voice data segments to be recognized, including: Real-time collection of voice data during user calls; The speech data is initially segmented using a preset initial speech rate threshold to obtain the first speech data and the second speech data. Wherein, the first speech data is speech data whose speed exceeds a preset initial speech speed threshold; the second speech data is speech data whose speed does not exceed the preset initial speech speed threshold. The comprehensive speech segmentation strategy is used in conjunction with the first segmentation coefficient and the second segmentation coefficient to segment the first speech data and the second speech data respectively, so as to obtain the speech data segments to be recognized corresponding to the first speech data and the second speech data.
3. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, The comprehensive speech segmentation strategy includes: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data; The first segment coefficient is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data, and the first segment coefficient is used to adjust the preset first segment speech rate threshold to obtain the adjusted first segment speech rate threshold. The speech rate in the first speech data is compared with the adjusted first segment speech rate threshold to obtain the first target speech data and the second target speech data. Wherein, the first target speech data is the first speech data whose speech rate exceeds the adjusted first segment speech rate threshold; the second target speech data is the first speech data whose speech rate does not exceed the adjusted first segment speech rate threshold; Furthermore, the first target speech data and the second target speech data are the speech data segments to be identified.
4. The real-time user call emotion recognition method based on ASR according to claim 1 or 3, characterized in that, The process of obtaining the first piecewise coefficient includes: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the first speech data; The first segmentation coefficient corresponding to the first speech data is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the first speech data; The first piecewise coefficient is obtained by the following formula: Among them, S 01 denoted as the first segment coefficient; n represents the number of time units contained in the first speech data, and the value of the time unit ranges from 5s to 3min; K i K represents the number of syllables contained in the first speech data in the i-th unit of time; c This indicates a preset reference value for the number of syllables; P si P represents the rate of change of the fundamental tone at the beginning of the i-th unit of time; zi T represents the rate of change of the fundamental tone at the end of the i-th unit of time; di P represents the time length corresponding to the i-th unit of time; maxi P represents the maximum rate of change of the fundamental tone in the i-th unit of time; mini T represents the minimum rate of change of the fundamental tone in the i-th unit of time; mi B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum in the i-th unit of time. maxi B represents the maximum volume intensity occurring in the i-th unit of time; mini B represents the minimum volume intensity occurring in the i-th unit of time; zpi K represents the average volume intensity of all syllables occurring in the i-th unit of time; di This represents the number of syllables that appear during the time interval between the time of maximum volume intensity and the time of minimum volume intensity in the i-th unit of time. The first segment speech rate threshold is adjusted using the first segment coefficient to obtain the adjusted first segment speech rate threshold. The adjusted first segment speech rate threshold is obtained using the following formula: Among them, K t01 K represents the adjusted first segment speech rate threshold; 01 S represents the speech rate threshold for the first segment; 01 X represents the first piecewise coefficient; p01 This represents the average slope of the pitch change rate over n units of time corresponding to the first speech data.
5. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, The comprehensive speech segmentation strategy also includes: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data; The second segment coefficient is obtained by using the speech rate parameter, volume intensity and pitch change rate corresponding to the second speech data, and the preset second segment speech rate threshold is adjusted by using the second segment coefficient to obtain the adjusted second segment speech rate threshold. The speech rate in the second speech data is compared with the adjusted second segment speech rate threshold to obtain the third target speech data and the fourth target speech data. The third target speech data is the second speech data whose speech rate exceeds the adjusted second segment speech rate threshold; the fourth target speech data is the second speech data whose speech rate does not exceed the adjusted second segment speech rate threshold. Furthermore, the third target speech data and the fourth target speech data are the speech data segments to be identified.
6. The real-time user call emotion recognition method based on ASR according to claim 1 or 5, characterized in that, The process of obtaining the second piecewise coefficient includes: Retrieve the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data; The second segmentation coefficients corresponding to the second speech data are obtained by using the speech rate parameters, volume intensity, and pitch change rate corresponding to the second speech data. The second piecewise coefficient is obtained using the following formula: Among them, S 02 represents the second segment coefficient; m represents the number of time units contained in the second speech data, and the value of each time unit is 20s-3mjn; K j K represents the number of syllables contained in the second speech data in the j-th unit of time; c This indicates a preset reference value for the number of syllables; T dj P represents the time length corresponding to the j-th unit of time; maxj P represents the maximum rate of change of the fundamental tone in the j-th unit of time; mjnj T represents the minimum rate of change of the fundamental tone in the j-th unit of time; mj B represents the time interval between the moment when the fundamental rate of change reaches its maximum and the moment when the fundamental rate of change reaches its minimum at the j-th unit of time. maxj B represents the maximum volume intensity occurring in the j-th unit of time; mjnj B represents the minimum volume intensity occurring in the j-th unit of time; zpj K represents the average volume intensity of all syllables occurring in the j-th unit of time; dj This represents the number of syllables that occur during the time interval between the maximum and minimum volume intensity of the j-th unit of time. The second segment speech rate threshold is adjusted using the second segment coefficient to obtain the adjusted second segment speech rate threshold; The adjusted second segment speech rate threshold is obtained using the following formula: Among them, K t02 K represents the adjusted speech rate threshold for the second segment; 02 S represents the speech rate threshold for the second segment; 02 X represents the second piecewise coefficient; p02 This represents the average slope of the pitch change rate over m units of time corresponding to the second speech data.
7. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, The ASR (Automatic Speech Recognition) system is used to convert the speech data segment to be recognized into text, thereby obtaining the text data corresponding to the speech data segment to be recognized, including: Extract multiple speech data segments to be recognized; The unit time is adjusted to obtain the adjusted unit time length, and the adjusted unit time is used as the unit time period; wherein the value range of the unit time period is 10s-1min; Each speech data segment to be recognized is subjected to speech preprocessing to obtain multiple preprocessed speech data segments to be recognized; wherein, the preprocessing includes noise suppression processing and echo cancellation processing; Retrieve the speech rate parameters corresponding to each preprocessed speech data segment to be recognized; Extract the time frame length corresponding to the preset frame segmentation process; The overlap ratio is obtained using the speech rate parameter corresponding to each speech data segment to be recognized. The overlap ratio is obtained by the following formula: Among them, P bI Indicates the overlap ratio; P bl0 This represents the preset overlap ratio baseline value; k represents the number of time segments contained in each segment of speech data to be recognized; K xi K represents the number of syllables corresponding to the i-th time unit; wp K represents the average number of syllables per unit time corresponding to the parent speech data segment for each speech data segment to be recognized; c This indicates a preset reference value for the number of syllables; According to the preset frame processing time frame length and overlap ratio, each preprocessed speech data segment to be recognized is processed into frames to obtain the audio data frame corresponding to each speech data segment to be recognized. The audio data frame corresponding to each speech data segment to be recognized is input into the ASR speech automatic recognition system to generate text data corresponding to each speech data segment to be recognized.
8. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, The text data is preprocessed by combining the speech data segment to be recognized to obtain preprocessed text data, including: Retrieve the text data corresponding to each speech data segment to be recognized; Remove leading and several whitespace characters, extra spaces, and special symbols that are not Chinese characters, numbers, or English characters from the text data to obtain the first sample data; The first sample data is segmented using a word segmentation tool to obtain the second sample data after word segmentation. The second sample data is subjected to error correction and modification processing to obtain the corrected text data. The corrected and revised text data refers to the preprocessed text data.
9. The real-time user call emotion recognition method based on ASR according to claim 1, characterized in that, The deep learning model is a convolutional neural network, and the structure of the deep learning model is as follows: The input layer is used to receive preprocessed text data. An embedding layer is used to map the index of each word in the text data to a dense vector space of fixed dimension; Convolutional layers are used to extract features from the text data and obtain feature maps corresponding to the text data; Pooling layers are used to reduce the dimensionality of the feature map; The concatenation layer is used to concatenate the dimensionality-reduced feature maps to form feature vectors; Fully connected layers are used to integrate features from feature vectors; The output layer is used to output the emotion recognition results.
10. A real-time user call emotion recognition system based on ASR, characterized in that, The ASR-based real-time user call emotion recognition system includes: The voice segmentation module is used to collect voice data during user calls in real time, and to segment the voice data using a comprehensive voice segmentation strategy combined with a first segmentation coefficient and a second segmentation coefficient to obtain multiple voice data segments to be recognized. The text data acquisition module is used to convert the speech data segment to be recognized into text using the ASR speech automatic recognition system, and to acquire the text data corresponding to the speech data segment to be recognized. The text data preprocessing module is used to preprocess the text data in conjunction with the speech data segment to be recognized, and obtain the preprocessed text data. The emotion recognition module is used to input the preprocessed text data into a trained deep learning model for emotion analysis and to obtain the emotion recognition result corresponding to the preprocessed text data.