Barrier-free call system, method and equipment based on emotion recognition feedback and medium
By introducing emotion recognition modules and feedback modules into the barrier-free call system, real-time analysis and feedback call emotions have been solved, and the problem of difficulty in identifying and feedback emotional changes in existing systems is significantly improved.
Patent Information
- Application Number
- CN202510184708.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
AI Technical Summary
The existing accessible call system is difficult to analyze and feedback emotional changes in the call in real time, making it difficult for hearing-impaired people to accurately understand the other party's emotional state, affecting the interactivity and effectiveness of the call.
The barrier-free call system based on emotional recognition feedback is adopted, and the voice signals of both parties to the call are analyzed in real time through the emotion recognition module. The emotional recognition algorithm of the cross-modal fusion method is used to identify the emotional state with the highest degree of emotion in the voice, and the corresponding emotional feedback information is generated through the emotional feedback module, which is converted into text description display.
Real-time recognition and feedback on emotional changes in the call is realized, accurate emotional feedback is provided to hearing-impaired people, enhance the interactivity and authenticity of the call, and avoid possible misunderstandings due to the lack of emotional information.
Smart Images

Figure CN120050355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to a barrier-free call system, method, device, and medium based on emotion recognition feedback. Background Art
[0002] A barrier-free call system is a technology and service that helps people with hearing, speech, or other communication disabilities use telephone communication more conveniently. Existing barrier-free call systems can already help hearing-impaired people make basic calls by implementing speech-to-text conversion and text-to-speech synthesis. However, these systems often ignore the emotional expressions in calls, making it difficult for hearing-impaired people to accurately understand the emotional state of the other party and affecting the interactivity and effectiveness of the calls.
[0003] For this reason, a voice call content display method (patent application number: 202310898438.5) is disclosed in the prior art, which includes: receiving a first input; in response to the first input, performing speech recognition processing on the audio data in the voice call to obtain the content information corresponding to the audio data, as well as the object information and emotion information corresponding to the content information; and outputting the content information, the object information corresponding to the content information, and the emotion information.
[0004] However, in the actual use process, this method can only completely express all the language meanings of a certain party in the call through the text content, and cannot analyze the emotional changes in the call in real time to provide accurate emotional feedback for hearing-impaired people, thereby enhancing the interactivity and authenticity of the call.
[0005] Therefore, the present application particularly proposes a barrier-free call system based on emotion recognition feedback to solve the above technical problems. Summary of the Invention
[0006] The main purpose of the present invention is to provide a barrier-free call system based on emotion recognition feedback to solve the technical problems raised in the background art.
[0007] The present invention adopts the following technical solutions to solve the above technical problems:
[0008] A barrier-free call system based on emotion recognition feedback includes an emotion recognition module and an emotion feedback module, where:
[0009] The emotion recognition module is used to receive and analyze the voice signals of both parties in the call in real time, and use an emotion recognition algorithm with a cross-modal fusion method to identify the emotional state with the highest corresponding emotion degree in the voice.
[0010] The emotion feedback module is used to generate corresponding emotion feedback information according to the output of the emotion recognition module, convert it into a text description, and display it on the screen for hearing-impaired people to read.
[0011] Preferably, the emotion recognition algorithm is used to analyze and process speech signals to identify the emotional state with the highest corresponding emotional degree, and the following operations are performed by a computer:
[0012] S1. Speech signal preprocessing: The speech signal is successively pre-emphasized, framed, windowed, denoised, speech endpoint detected, and speech segmented to eliminate aliasing, high-order harmonic distortion, and high-frequency problems in the human voice of the speech signal;
[0013] S2. Feature extraction: Mel-frequency cepstral coefficients are used to capture the spectral features of the speech signal, TF-IDF is used to extract features from the text data, and at the same time, an emotion word library with specified special characters is created, and the specified special characters in the extracted text data are classified;
[0014] S3. Feature fusion: The features extracted from different modalities are fused using a weighted average algorithm according to the TF-IDF values;
[0015] S4. Emotion classification model training: The fused multi-modal features are trained using a support vector machine model to construct an emotion classification model;
[0016] S5. Model tuning: The model parameters of the emotion classification model are optimized using the labeled speech data set;
[0017] S6. Emotion recognition: Feature extraction is performed on the new language signal, and it is judged whether it contains the specified special characters stored in the emotion word library. If so, the trained model is used for emotion classification, the emotion category of the speech signal is identified, and the specified special characters are used as labels for output;
[0018] S7. Result optimization: The text content in the language signal is recognized by ASR speech recognition technology, and the specified special characters stored in the emotion word library are retrieved for matching, and the emotion recognition result in step S6 is optimized according to the audio size of the specified special character part.
[0019] Preferably, the specific operations for further optimizing the emotion recognition result according to the audio size of the specified special character part in step S7 include:
[0020] S71. Set and use a speech recognition model to recognize and match the specified special characters in the language signal and record the audio size of the characters;
[0021] S72. Calculate the emotion value of the specified special character according to the overall audio size of the language signal segment and the audio size of the specified special character;
[0022] S73. Optimize the emotion recognition result according to the emotion values of multiple specified special characters.
[0023] Preferably, the specific operation method for calculating the emotion value of the specified special text in step S72 includes:
[0024] First step, calculate the average audio value E according to the overall audio size of the language signal segment;
[0025] Second step, calculate the emotion value P of the specified special text in the language signal segment according to the average audio value E and the audio size U of the specified special text, where
[0026] Preferably, the specific operation method for optimizing the emotion recognition result according to the emotion weight value in step S73 includes:
[0027] L1. According to the emotion category of the speech signal recognized in step S6, calculate the optimized emotion result K through the emotion value P:
[0028]
[0029] L2. Further optimize according to the emotion category of the optimized emotion value K:
[0030] When K = strong emotion, a strong emotion label can be added before the emotion recognition result;
[0031] When K = normal emotion expression, output the emotion recognition result normally without optimization description;
[0032] When K = opposite emotion expression, modify the emotion recognition result to the opposite emotion.
[0033] A barrier-free call method based on emotion recognition feedback includes the following specific steps:
[0034] 1) Use the above-mentioned system to record the speech signals of both parties of the call user;
[0035] 2) The emotion recognition algorithm analyzes and processes the speech signal, and determines whether the speech signal segment contains specified special text. If so, the emotion recognition algorithm recognizes, obtains and optimizes the corresponding emotion state information;
[0036] 3) Generate corresponding emotion feedback information in the emotion feedback module according to the emotion state information, and display it in text description on the screen to realize a barrier-free call capable of emotion recognition feedback.
[0037] On the other hand, the present invention also discloses a computer device, including a connector, a memory and a processor. The connector is used to connect with a communication device and obtain a language signal. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above method.
[0038] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above-mentioned method.
[0039] As can be seen from the above technical solutions, the present invention provides a barrier-free call system based on emotion recognition feedback. Compared with the prior art, the present invention has the following advantages:
[0040] 1. The present invention analyzes and recognizes the emotional change states of both parties in a call in real time through an emotion recognition module and an emotion feedback module, provides accurate emotion feedback for the hearing-impaired, makes up for the deficiency that the traditional barrier-free call system can only transmit basic information, enables the hearing-impaired to also perceive the emotional changes of the other party, enhances the authenticity and interactivity of the call during the communication process, avoids misunderstandings that may occur due to the lack of emotional information, and makes the feedback content closer to the actual conversation situation.
[0041] 2. The present invention extracts features from text data by using TF-IDF, and fuses the features extracted from different modalities according to the TF-IDF values by using a weighted average algorithm, which can integrate the information of the two modalities, form a more comprehensive and accurate emotional feature representation, obtain more accurate emotional state information, and improve the robustness and accuracy of emotion recognition.
[0042] 3. The present invention optimizes the emotion recognition result by analyzing the audio characteristics of specified special texts (specific emotional words) before outputting the emotion recognition tags. By integrating the important emotional expression dimension of the high and low and size of the voice audio, the emotion recognition ability of the barrier-free call system is significantly enhanced, making it closer to the real interpersonal communication experience, improving the interactivity and understanding depth, and further improving the accuracy of the recognized emotion results.
[0043] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0045] Figure 1 is a schematic diagram of the system module of the present invention;
[0046] Figure 2Schematic diagram of the overall process for the emotion recognition algorithm of the present invention to recognize emotion states;
[0047] Figure 3 Schematic diagram of the overall process for the present invention to optimize emotion recognition results;
[0048] Figure 4 Schematic diagram of the process of the method of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] In the embodiment, refer in detail to Figures 1 to 4 .
[0051] As Figure 1 shown. An accessible call system based on emotion recognition feedback proposed in the embodiment of the present invention includes an emotion recognition module and an emotion feedback module:
[0052] In a specific embodiment, the emotion recognition module is used to receive and analyze the voice signals of both parties in the call in real time, and use an emotion recognition algorithm with a cross-modal fusion method to recognize the emotion state with the highest corresponding emotion degree in the voice, such as happiness, anger, sadness, etc.
[0053] Among them, as Figure 2 shown, the emotion recognition algorithm is used to analyze and process the voice signal to recognize the emotion state with the highest corresponding emotion degree, and the following operations are performed by a computer:
[0054] S1. Voice signal preprocessing: The voice signal is sequentially subjected to pre-emphasis, framing, windowing, denoising, voice endpoint detection, and voice segmentation operations to eliminate aliasing, high-order harmonic distortion, and high-frequency problems in the human voice of the voice signal, improve the quality of the voice signal, and ensure as much as possible that the signals obtained by subsequent voice processing are more uniform and smooth, provide high-quality parameters for signal parameter extraction, and improve the quality of voice processing;
[0055] S2. Feature Extraction: Use Mel Frequency Cepstral Coefficients (MFCC) to capture the spectral features of the speech signal, extract the discriminative components in the audio signal, then filter out other worthless information to reduce the interference of noise on emotion extraction and improve the extraction accuracy. After that, use TF-IDF to extract features from the text data, and at the same time create an emotion word library with specified special words and classify the specified special words in the extracted text data;
[0056] In the specific usage process, set the extracted MFCC feature vector as X = [x 1 , x 2 , x 3 , …, x n , where n is the dimension of the feature vector. At this time, each feature vector X is represented as a row vector, and a feature matrix X is constructed, with its size being mn, where m is the number of speech samples;
[0057] It should also be noted that before the TF-IDF process, the speech signal needs to be pre-emphasized: The average power spectrum of the speech signal s(n) is affected by glottal excitation and oral-nasal radiation. The high-frequency end attenuates by about 6 dB / oct (octave) above approximately 800 Hz, and the higher the frequency, the smaller the corresponding component. Therefore, before performing TF-IDF processing and analysis on the speech signal s(n), its high-frequency part needs to be enhanced. Usually, the measure is to use a digital filter to achieve pre-emphasis. The relationship between the output of the pre-emphasis network and the input speech signal s(n) is:
[0058]
[0059] where a is the pre-emphasis coefficient, generally taken between 0.9 and 1.0;
[0060] In the TF-IDF process, TF is expressed as "term frequency" and is expressed as:
[0061] Furthermore, the specific extraction method of using TF-IDF to extract features from the text data in step S2 includes:
[0062] S21. Calculate the frequency of occurrence of each word in the text by simple counting or normalization method, denoted as TF;
[0063] S22. Calculate the inverse document frequency of each word to measure the importance of the word, denoted as IDF:
[0064]
[0065] where N is the total number of documents, and df(t) is the number of documents containing the word t;
[0066] S23. Calculate the TF-IDF value of each word as the final feature representation. The calculation formula for the TF-IDF value is expressed as:
[0067] TF-IDF t, d = TF t, d × IDF(t)
[0068] Among them, TF t, d is the term frequency of word t in document d, and IDF(t) is the inverse document frequency of word t;
[0069] That is, when there are TF (term frequency) and IDF (inverse document frequency), multiplying these two terms can obtain the TF-IDF value of a word. The larger the TF-IDF of a certain word in the article, generally speaking, the more important this word is in this article. Therefore, by calculating the TF-IDF of each word in the article and sorting them from largest to smallest, the first few words are the keywords of the article;
[0070] S24. Extract features from the text according to the TF-IDF value.
[0071] S3. Feature fusion: The features extracted from different modalities are fused using the weighted average algorithm according to the TF-IDF value. The specific fusion method includes:
[0072] S31. Set M different features, which come from M different modalities respectively. The feature representation extracted from each modality is X 1 , X 2 , …, X M , where X i is an N i ×D i feature matrix, where N i represents the number of samples, and D i represents the feature dimension.
[0073] S32. Use the weighted average algorithm for fusion: Define the weight vector W = [w 1 , w 2 , …, w M , where w i represents the weight of the i-th modality feature and satisfies the condition Then the fused feature can be expressed as a weighted average:
[0074] Among them: X fused represents the fused feature matrix, with a size of N×D, N = N 1 = 2 = … = M expressed as the number of samples, D = D 1 = 2 = … = MExpressed as the fused feature dimension;
[0075] S4. Sentiment Classification Model Training: For the fused multi-modal features, use the Support Vector Machine (SVM) model to train and construct a sentiment classification model;
[0076] Before constructing the sentiment classification model, set multiple corresponding sentiment category labels Y = y 1 , y 2 , y 3 , …, y m , where y i is the sentiment category label of the i-th speech sample;
[0077] According to the above, the sentiment classification problem can be transformed into a binary classification or multi-classification problem. The goal of the SVM model is to find a hyperplane that can separate the projected sample points in the feature space to the greatest extent;
[0078] Specifically, the optimization objective function of the support vector machine can be expressed as:
[0079]
[0080] ξ i ≥0, i = 1, 2, …, m;
[0081] where W is the normal vector of the hyperplane, b is the bias term of the hyperplane, C is the regularization parameter, and ξ i is the slack variable used to handle the linearly inseparable situation;
[0082] S5. Model Tuning: Use the labeled speech dataset to optimize the model parameters of the sentiment classification model;
[0083] S6. Sentiment Recognition: Extract features from the new language signal, determine whether it contains the specified special characters stored in the emotion word library. If so, use the trained model for sentiment classification, identify the sentiment category of the speech signal, and output it with the specified special characters as labels. The specific labels are "happy", "sad", "angry", "afraid", "surprised", "disgusted";
[0084] S7. Result Optimization: Identify the text content in the language signal through ASR speech recognition technology, retrieve the specified special characters stored in the emotion word library for matching, and optimize the sentiment recognition result in step S6 according to the audio size of the matched specified special character part;
[0085] Specifically, as Figure 3 shown, the specific operations for further optimizing the sentiment recognition result according to the audio size of the specified special character part include:
[0086] S71. Set up and use a speech recognition model to recognize and match specified special characters in the language signal and record the audio size of the characters, where the speech recognition model can be directly obtained by training a convolutional neural network combined with multiple speech signals;
[0087] S72. Calculate the emotion value of the specified special character according to the overall audio size of the language signal segment and the audio size of the specified special character. The specific operation method for calculating the emotion value of the specified special character includes:
[0088] The first step is to calculate the average audio value E according to the overall audio size of the language signal segment;
[0089] The second step is to calculate the emotion value P of the specified special character in the language signal segment according to the average audio value E and the audio size U of the specified special character, where
[0090] S73. Optimize the emotion recognition result according to the emotion values of multiple specified special characters. The specific operation method includes:
[0091] L1. According to the emotion category of the speech signal recognized in step S6, calculate the optimized emotion result K through the emotion value P:
[0092]
[0093] L2. Further optimize according to the emotion category of the optimized emotion value K:
[0094] When K = strong emotion, an emotion strong label can be added before the emotion recognition result;
[0095] When K = normal emotion expression, the emotion recognition result is output normally without optimization description;
[0096] When K = opposite emotion expression, the emotion recognition result is modified to the opposite emotion.
[0097] In a specific embodiment, the emotion feedback module is used to generate corresponding emotion feedback information according to the output of the emotion recognition module and convert it into a text description, which is displayed on the screen for the hearing-impaired to read (the emotion label can be directly placed behind the text after speech recognition)
[0098] In addition, emoji, color change or vibration prompt and other methods can also be used to transmit emotion information through visual or tactile channels. These feedback methods can be selected and adjusted according to the user's settings and preferences.
[0099] On the other hand, the present invention also discloses a barrier-free call method based on emotion recognition feedback, refer to Figure 4 including the following specific steps:
[0100] 1) Use the barrier-free call system based on emotion recognition feedback as described above to record the voice signals of both parties of the call users.
[0101] 2) The emotion recognition algorithm analyzes and processes the voice signal, and determines whether the voice signal contains specified special characters. If so, the emotion recognition algorithm identifies, obtains, and optimizes the corresponding emotion state information.
[0102] 3) Generate corresponding emotion feedback information in the emotion feedback module according to the emotion state information, and display it in text description on the screen to achieve barrier-free calls capable of emotion recognition feedback.
[0103] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0104] On yet another aspect, the present invention also discloses a computer device including a connector, a memory, and a processor. The connector is used to connect to a communication device to obtain language signals. The memory stores a computer program, which when executed by the processor causes the processor to execute the steps of the above method.
[0105] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer causes the computer to execute any of the system methods in the above embodiments.
[0106] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. The explanations, examples, and beneficial effects of related content can refer to the corresponding parts in the above method.
[0107] The embodiments of the present application also provide an electronic device including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus.
[0108] The memory is used to store a computer program.
[0109] The processor is used to implement the above method when executing the program stored on the memory.
[0110] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, and the like.
[0111] The communication interface is used for communication between the above electronic device and other devices.
[0112] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0113] The above processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0114] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The specific technologies and device forms adopted by the terminal device in the embodiments of the present application are not limited.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
[0117] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0118] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, or scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
Claims
1. A barrier-free communication system based on emotion recognition feedback, characterized in that: It includes emotion recognition module and emotion feedback module, where: The emotion recognition module is used to receive and analyze the voice signals of both parties in real time, and use the emotion recognition algorithm with cross-modal fusion method to identify the emotional state with the highest emotional level in the speech; The emotion feedback module is used to generate corresponding emotion feedback information according to the output of the emotion recognition module and convert it into text description, which is displayed on the screen for the hearing-impaired to read.
2. The barrier-free communication system based on emotion recognition feedback as claimed in claim 1, characterized in that: The emotion recognition algorithm is used to analyze and process the speech signal to identify the emotional state with the highest corresponding emotion level, and the following operations are performed by a computer: S1. Speech signal preprocessing: pre-emphasize, frame, window, denoise, speech endpoint detection, and speech segmentation are performed on the speech signal in sequence to eliminate aliasing, high-order harmonic distortion, and high-frequency problems in the human voice of the speech signal; S2. Feature extraction: Use Mel frequency cepstral coefficients to capture the spectral features of speech signals, use TF-IDF to extract features from text data, create an emotional vocabulary with specified special words, and classify the specified special words in the extracted text data; S3. Feature fusion: The features extracted from different modalities are fused using a weighted average algorithm based on the TF-IDF value; S4. Sentiment classification model training: Use the support vector machine model to train and build a sentiment classification model for the fused multimodal features; S5. Model tuning: Use the labeled speech dataset to optimize the model parameters of the sentiment classification model; S6. Emotion recognition: Extract features from new language signals to determine whether they contain the specified special characters stored in the emotional vocabulary. If so, use the trained model to classify the emotions and identify the emotion category of the speech signal, and use the specified special characters as label output; S7. Result optimization: Use ASR speech recognition technology to identify the text content in the language signal, and retrieve the specified special text stored in the emotional vocabulary for matching, and optimize the emotion recognition result in step S6 according to the audio size of the matched specified special text part.
3. The barrier-free communication system based on emotion recognition feedback as claimed in claim 2, characterized in that: The specific operation of further optimizing the emotion recognition result according to the audio size of the specified special text part in step S7 includes: S71. Setting and using a speech recognition model to identify and match a specified special text in a language signal and record the audio size of the text for recognition; S72. Calculate the emotion value of the specified special text according to the overall audio size of the language signal segment and the audio size of the specified special text; S73. Optimize the emotion recognition result according to the emotion values of multiple specified special characters.
4. The barrier-free communication system based on emotion recognition feedback as claimed in claim 3, characterized in that: The specific operation method of calculating the emotion value of the specified special character in step S72 includes: The first step is to calculate the average audio value E according to the overall audio size of the language signal segment; The second step is to calculate the emotion value P of the specified special word in the language signal segment according to the average audio value E and the audio size U of the specified special word, where 5. The barrier-free communication system based on emotion recognition feedback as claimed in claim 4, characterized in that: The specific operation method of optimizing the emotion recognition result according to the emotion weight value in step S73 includes: L1. According to the emotion category of the speech signal identified in step S6, the optimized emotion result K is calculated by the emotion value P: L2. Further optimization based on the optimized sentiment value K sentiment category: When K = strong emotion, a strong emotion label can be added before the emotion recognition result; When K = normal emotional expression, the emotion recognition results are output normally without any optimization description; When K = opposite emotion expression, the emotion recognition result is modified to the opposite emotion.
6. A barrier-free communication method based on emotion recognition feedback, characterized in that: The specific steps include: 1) Using the barrier-free call system based on emotion recognition feedback as described in any one of claims 1 to 5 to record the voice signals of both call users; 2) The emotion recognition algorithm analyzes and processes the speech signal and determines whether the speech signal contains the specified special text. If so, the emotion recognition algorithm identifies, obtains and optimizes the corresponding emotional state information; 3) Generate corresponding emotional feedback information in the emotional feedback module according to the emotional state information, and display it in text description on the screen, so as to realize barrier-free communication with emotional recognition feedback.
7. A computer device, characterized in that: The method comprises a connector, a memory and a processor, wherein the connector is used to connect to a communication device for acquiring a language signal, and the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method as claimed in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice call content display method and device
CN116939091A