A method for diagnosing emotions of spoken english

By extracting and processing the frequency, visual, and waveform features of spoken English, and combining various emotional feature calculation methods, the reliability and validity of existing English spoken emotion diagnosis methods are addressed, achieving more accurate emotional state identification and analysis.

CN119360901BActive Publication Date: 2025-11-21GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411523874.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-11-21
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

Existing methods for diagnosing English spoken emotion are insufficient in terms of reliability and validity, and cannot effectively identify and analyze the phonological emotional features in spoken English.

Method used

An English speaking preprocessing module and an emotion diagnosis module are adopted. By extracting and processing frequency features, visual features and waveform features, and combining a bidirectional flow control unit, a convolutional feature learning network and a speech representation learner, a bidirectional-convolutional-representation emotion feature vector of English speaking is generated. Finally, the emotion diagnosis is performed using the maximum emotion probability calculation formula.

Benefits of technology

It improves the accuracy and reliability of English speaking emotion diagnosis, and can effectively identify and analyze the emotional state in English speaking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360901B_ABST
    Figure CN119360901B_ABST
Patent Text Reader

Abstract

The application provides an English oral emotion diagnosis method, which is a diagnosis method composed of sequentially connected English oral pre-processing modules and English oral emotion diagnosis modules. After the English oral is processed by the diagnosis method, the emotion diagnosis result of the English oral can be finally obtained. The application can solve the problem of poor reliability and validity of the existing English oral emotion diagnosis method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application is particularly a method for diagnosing the emotion of English oral language, and the method of the present application is only suitable for the emotion diagnosis of English oral language, and is not suitable for other emotion diagnosis. BACKGROUND

[0002] In various application scenarios such as intelligent customer service and oral language teaching, speech emotion diagnosis plays a significant role. English oral language emotion diagnosis aims to analyze the emotional state of the speaker in English oral language by recognizing and analyzing the speech emotion features in English oral language. The existing oral language emotion diagnosis has the problem of poor reliability and validity of emotion diagnosis. In order to solve the above problem, the present application provides an English oral language emotion diagnosis method including English oral language frequency features, English oral language visual features and English oral language waveform features. SUMMARY

[0003] The English oral language emotion diagnosis method of the present application comprises an English oral language preprocessing module and an English oral language emotion diagnosis module, and the overall processing flow chart is as shown in Figure 1 .

[0004] The processing flow of the English oral language preprocessing module of the present application is as follows: first, reading the English oral language to be diagnosed and enhancing the English oral language signal; second, dividing the enhanced English oral language signal according to a time interval of fifteen milliseconds to form a plurality of shorter English oral language segments, obtaining English oral language waveform features; third, converting the signal of each English oral language segment from time domain to frequency domain, analyzing the frequency components of each English oral language segment, obtaining the frequency spectrum of each English oral language signal and performing rotation and mapping, and then splicing the plurality of speech frequency spectrums after rotation and mapping to obtain English oral language visual features; fourth, performing analog human speech processing on the frequency spectrum of each English oral language segment after rotation and mapping, and performing frequency spectrum feature screening on the English oral language signal frequency spectrum after analog human speech processing; fifth, performing dimension reduction and numerical processing on the screened English oral language signal frequency spectrum features, and converting the English oral language frequency spectrum features after dimension reduction and numerical processing; sixth, analyzing the changes of English oral language frequency spectrum features at different time points, capturing the dynamic characteristics of English oral language frequency spectrum features, and outputting English oral language frequency features.

[0005] The processing flow of the English oral emotion diagnosis module of the application is as follows: first, reading the English oral frequency feature, the English oral visual feature and the English oral waveform feature obtained by the English oral preprocessing module; second, inputting the English oral frequency feature into the bidirectional flow control unit, and obtaining the English oral bidirectional emotion feature vector by using the English oral bidirectional emotion feature vector calculation formula (1); third, inputting the English oral visual feature into the convolution feature learning network unit, and obtaining the English oral convolution emotion feature vector by using the English oral convolution emotion feature vector calculation formula (2); fourth, inputting the English oral waveform feature into the speech representation learner, and obtaining the English oral representation emotion feature vector by using the English oral representation emotion feature vector calculation formula (3); fifth, performing splicing operation on the English oral bidirectional emotion feature vector, the English oral convolution emotion feature vector and the English oral representation emotion feature vector to generate the English oral bidirectional-convolution-representation emotion feature vector; sixth, performing English oral emotion probability maximum value calculation on the English oral bidirectional-convolution-representation emotion feature vector by using the English oral emotion probability maximum value calculation formula (4) to obtain the English oral emotion probability maximum value; seventh, reading the English oral emotion probability maximum value obtained by the English oral emotion extraction module; eighth, judging the English oral emotion diagnosis result by using the English oral emotion diagnosis result calculation formula (5) and outputting the English oral emotion diagnosis result.

[0006] The calculation formula of the diagnosis method of the application is defined as follows:

[0007] (1) English oral bidirectional emotion feature vector calculation formula

[0008] English oral bidirectional emotion feature vector = bidirectional flow control unit English oral frequency feature (1)

[0009] In formula (1), the English oral frequency feature is a feature for describing the frequency value of English oral speech, and the bidirectional flow control unit is a deep processing of the English oral frequency feature by using the mel scale.

[0010] (2) English oral convolution emotion feature vector calculation formula

[0011] English oral convolution emotion feature vector = convolution feature learning network English oral visual feature (2)

[0012] In formula (2), the English oral visual feature is a speech frequency graph feature reflecting the change of speech frequency in a period of English oral speech, and the convolution feature learning network is a convolutional coding processing of the English oral visual feature.

[0013] (3) English oral representation emotion feature vector calculation formula

[0014] English oral speech representation emotion feature vector = speech representation learner English oral speech waveform feature (3)

[0015] In formula (3), the English oral speech waveform feature is a waveform value reflecting the emotion of English oral speech, and the speech representation learner is an emotion waveform learning process for the English oral speech waveform feature.

[0016] (4) English oral speech emotion probability maximum calculation formula

[0017]

[0018] In formula (4), e represents the base number of the natural logarithm function, i represents the i-th emotion category, and there are four emotion categories in total: anger, sadness, natural, and happiness. The English oral speech bidirectional emotion feature vector, the English oral speech convolution emotion feature vector, and the English oral speech representation emotion feature vector are calculated by formulas (1), (2), and (3) respectively.

[0019] (5) English oral speech emotion diagnosis result calculation formula

[0020]

[0021] In formula (5), the English oral speech emotion probability maximum obtained by formula (4) is matched with the emotion category interval to obtain the English oral speech emotion diagnosis result. The emotion category is: anger, sadness, natural, and happiness, and the emotion category interval value is [0, 1].

[0022] Specific processing steps of the method of the application

[0023] The English oral speech preprocessing module, the English oral speech emotion diagnosis module, and the processing method steps of the English oral speech emotion diagnosis module of the analysis method of the application are described as follows.

[0024] As shown in Figure 2 The steps of the English oral speech pronunciation preprocessing module processing flow are as follows:

[0025] P201 starts;

[0026] P202 reads the English oral speech to be diagnosed and enhances the English oral speech signal;

[0027] P203 divides the enhanced English oral speech signal into multiple shorter English oral speech segments according to a fifteen-millisecond time interval, and obtains the English oral speech waveform feature;

[0028] P204 converts the signal of each English oral speech segment from the time domain to the frequency domain, analyzes the frequency components of each English oral speech segment, obtains the frequency spectrum of each English oral speech signal and performs rotation and mapping, and then splices the multiple frequency spectrums after rotation and mapping to obtain the English oral speech visual feature.

[0029] P205 simulates human speech processing on each segment of the rotated and mapped English spoken language signal spectrum, and performs spectrum feature screening on the English spoken language signal spectrum subjected to the simulated human speech processing;

[0030] P206 performs dimension reduction and numerical processing on the screened English spoken language signal spectrum features, and converts the English spoken language spectrum features subjected to the dimension reduction and numerical processing;

[0031] P207 analyzes the changes of the English spoken language spectrum features at different time points, captures the dynamic characteristics of the English spoken language spectrum features, and outputs the English spoken language frequency features;

[0032] P208 ends.

[0033] As shown in Figure 3 , the steps of the English spoken language emotion diagnosis module processing flow are as follows:

[0034] P301 starts;

[0035] P302 reads the English spoken language frequency features, English spoken language visual features and English spoken language waveform features obtained by the English spoken language preprocessing module;

[0036] P303 inputs the English spoken language frequency features into the bidirectional flow control unit, and obtains the English spoken language bidirectional emotion feature vector by using the English spoken language bidirectional emotion feature vector calculation formula (1);

[0037] P304 inputs the English spoken language visual features into the convolution feature learning network unit, and obtains the English spoken language convolution emotion feature vector by using the English spoken language convolution emotion feature vector calculation formula (2);

[0038] P305 inputs the English spoken language waveform features into the speech representation learner, and obtains the English spoken language representation emotion feature vector by using the English spoken language representation emotion feature vector calculation formula (3);

[0039] P306 performs splicing operation on the English spoken language bidirectional emotion feature vector, the English spoken language convolution emotion feature vector and the English spoken language representation emotion feature vector, to generate the English spoken language bidirectional-convolution-representation emotion feature vector;

[0040] P307 performs English spoken language emotion probability maximum value calculation on the English spoken language bidirectional-convolution-representation emotion feature vector by using the English spoken language emotion probability maximum value calculation formula (4), to obtain the English spoken language emotion probability maximum value;

[0041] P308 reads the obtained English spoken language emotion probability maximum value;

[0042] P309 calculates the English oral emotion diagnosis result by using the English oral emotion diagnosis result calculation formula (5), judges the English oral emotion diagnosis result, and outputs the English oral emotion diagnosis result.

[0043] P310 ends. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a general process flowchart of the method of the present application;

[0045] Figure 2 is an English oral pre-processing module processing flowchart of the method of the present application;

[0046] Figure 3 is an English oral emotion diagnosis module processing flowchart of the method of the present application. DETAILED DESCRIPTION

[0047] The present application will be further described below in conjunction with examples, but is not limited to the present application. The specific embodiment of the English oral emotion diagnosis method of the present application includes the following three steps.

[0048] Step 1: execute the "English oral pre-processing module"

[0049] The English oral data to be diagnosed is pre-processed, wherein the English oral data to be diagnosed contains 750 samples, covering four emotion types of happy, joyful, sad and natural, and the reference texts corresponding to the samples are as follows. Since the number of samples is large, only some of the reference texts corresponding to the samples are listed below, and the rest of the reference texts are replaced by ellipses:

[0050] Will we ever forget it.

[0051] Gad, your letter came just in time.

[0052] Not at this particular case, Tom, apologized Whittemore.

[0053] Lord, but I'm glad to see you again, Phil.

[0054] God bless them, I hope I'll go on seeing them forever.

[0055]

[0056] I'm playing a single hand in what looks like a losing game.

[0057] Gregson shoved back his chair and rose to his feet.

[0058] There was a change now.

[0059] Clubs and balls and cities grew to be only memories.

[0060] Hardly were our plans made public before we were met by powerful opposition.

[0061] Through the diagnosis of English oral English oral pretreatment module for pretreatment, get each English oral three kinds of acoustic features, respectively, English oral waveform features, English oral visual features and English oral frequency features, the first five sentences in the sample are shown, as follows, because the length of each sentence is inconsistent, three kinds of acoustic features length take the longest data in the standard, the data sequence of insufficient time is 0 processing:

[0062] (1) After step one processing of English oral frequency features

[0063] The first sentence of English oral English oral frequency features

[0064] [[-599.9907, 47.5119, -24.2490,..., 0.9349, -6.8700, -4.0621],

[0065] [-597.2722, 47.3519, -23.0500,..., 1.3695, -5.1366, -2.8849],

[0066] [-594.2827, 46.2890, -21.9260,..., 1.3969, -2.8165, -2.0617],

[0067] ...,

[0068] [0.0000, 0.0000, 0.0000,..., 0.0000, 0.0000, 0.0000],

[0069] [0.0000, 0.0000, 0.0000,..., 0.0000, 0.0000, 0.0000],

[0070] [0.0000, 0.0000, 0.0000,..., 0.0000, 0.0000, 0.0000]]

[0071] Frequency characteristics of spoken English in the second sentence

[0072] [[-5.8287e+02,1.4695e+01,-4.6094e+01,...,1.0172e+00,-8.1405e+00,-9.0633e+00],

[0073] [-5.6990e+02,1.6128e+01,-4.7454e+01,...,6.5461e-01,-7.4032e+00,-8.0288e+00],

[0074] [-5.6439e+02,1.7641e+01,-4.8666e+01,...,1.0434e-01,-5.9737e+00,-6.3430e+00],

[0075] ...,

[0076] [0.0000e+00,0.0000e+00,0.0000e+00,...,0.0000e+00,0.0000e+00,0.0000e+00],

[0077] [0.0000e+00,0.0000e+00,0.0000e+00,...,0.0000e+00,0.0000e+00,0.0000e+00],

[0078] [0.0000e+00,0.0000e+00,0.0000e+00,...,0.0000e+00,0.0000e+00,0.0000e+00]]

[0079] The frequency characteristics of spoken English in the third sentence

[0080] [[-6.5459e+02,3.5926e+01,-3.6259e+01,...,3.6956e+00,-3.1263e+00,1.1421e+00],[-6.3922e+02,3.6581e+01,-3.7816e+01,.. .,3.5393e+00,-3.4491e+00,-4.8856e-01],[-6.3073e+02,3.7314e+01,-3.9016e+01,...,2.9281e+00,-4.3269e+00,-7.8635e-03],

[0081] ...,

[0082] [-4.2343e+02,1.0743e+02,8.7176e+01,...,1.4282e+00,4.3656e+00,8.2641e+00],

[0083] [-4.3230e+02,1.1750e+02,-9.1742e+01,...,3.9004e+00,4.5268e+00,9.6107e+00],

[0084] [-4.3870e+02,1.2017e+02,-9.1920e+01,...,5.4412e+00,4.2130e+00,1.0909e+01]]

[0085] The frequency characteristics of spoken English in the fourth sentence

[0086] [[-592.3619,25.6732,-5.7658,...,-1.7858,-6.5272,-1.0995],

[0087] [-578.4620,28.2392,-10.6007,...,-1.3733,-5.9978,-1.0553],

[0088] [-570.9243,31.0595,-15.9202,...,-0.6141,-4.9857,-0.7592],

[0089] ...,

[0090] [0.0000,0.0000,0.0000,...,0.0000,0.0000,0.0000],

[0091] [0.0000,0.0000,0.0000,...,0.0000,0.0000,0.0000],

[0092] [0.0000,0.0000,0.0000,...,0.0000,0.0000,0.0000]]

[0093] The frequency characteristics of spoken English in the fifth sentence

[0094] [[4.3930e+02,6.7014e+01,8.6587e+01,...,1.7602e+01,1.0316e+01,4.0333e+00],

[0095] [-4.3264e+02, 7.2837e+01, 9.2337e+01,..., 1.8312e+01, 1.1684e+01, 5.2721e+00],

[0096] [-4.2887e+02, 7.8432e+01, -9.5508e+01,..., 1.8373e+01, 1.2399e+01, 6.5419e+00],

[0097] ...,

[0098] [-4.2128e+02, -1.0337e+02, -1.6662e+01,..., 2.0994e+00, 7.5810e+00, 5.7043e+00], [-4.2657e+02, -1.0196e+02, -1.5503e+01,..., 9.9634e-01, 5.1721e+00, 3.6053e+00],

[0099] [-4.2630e+02, -9.3132e+01, -1.2104e+01,..., -2.7240e+00, 1.8417e-01, 1.4195e+00]]

[0100] (2) English spoken language visual features of the English spoken language processed in step one

[0101] English spoken language visual features of the first sentence of English spoken language

[0102] [[[-0.5082, -0.5596, -0.5767,..., 2.2489, 2.2489, 2.2489],

[0103] [-0.7650, -0.8678, -0.8164,..., 2.2489, 2.2489, 2.2489],

[0104] [-1.1247, -1.1075, -0.9534,..., 2.2489, 2.2489, 2.2489],

[0105] ...,

[0106] [-0.7650, -0.5253, -0.3198,..., 2.2489, 2.2489, 2.2489],

[0107] [-1.2445, -0.7308, -0.4397,..., 2.2489, 2.2489, 2.2489],

[0108] [-1.7754,-0.9877,-0.6281,...,2.2489,2.2489,2.2489]],[[-0.3901,-0.4426,-0.4601,...,2.4286,2.4286,2.4286],

[0109] [-0.6527,-0.7577,-0.7052,...,2.4286,2.4286,2.4286],

[0110] [-1.0203,-1.0028,-0.8452,...,2.4286,2.4286,2.4286],

[0111] ...,

[0112] [-0.6527,-0.4076,-0.1975,...,2.4286,2.4286,2.4286],

[0113] [-1.1429,-0.6176,-0.3200,...,2.4286,2.4286,2.4286],

[0114] [-1.6856,-0.8803,-0.5126,...,2.4286,2.4286,2.4286]],[[-0.1661,-0.2184,-0.2358,...,2.6400,2.6400,2.6400],

[0115] [-0.4275,-0.5321,-0.4798,...,2.6400,2.6400,2.6400],

[0116] [-0.7936,-0.7761,-0.6193,...,2.6400,2.6400,2.6400],

[0117] ...,

[0118] [-0.4275,-0.1835,0.0256,...,2.6400,2.6400,2.6400],

[0119] [-0.9156,-0.3927,-0.0964,...,2.6400,2.6400,2.6400],

[0120] [-1.4559, -0.6541, -0.2881,..., 2.6400, 2.6400, 2.6400]]

[0121] English spoken language visual features of the second sentence of English spoken language

[0122] [[-1.5699, -1.3815, -1.3130,..., 2.2489, 2.2489, 2.2489],

[0123] [-1.6555, -1.4329, -1.3130,..., 2.2489, 2.2489, 2.2489],

[0124] [-1.8953, -1.5699, -1.3302,..., 2.2489, 2.2489, 2.2489],

[0125] ...,

[0126] [-1.2959, -1.2445, -1.2788,..., 2.2489, 2.2489, 2.2489],

[0127] [-1.3130, -1.4672, -1.5014,..., 2.2489, 2.2489, 2.2489],

[0128] [-1.3644, -1.8782, -1.9124,..., 2.2489, 2.2489, 2.2489]],[[-1.4755, -1.2829, -1.2129,..., 2.4286, 2.4286, 2.4286],

[0129] [-1.5630, -1.3354, -1.2129,..., 2.4286, 2.4286, 2.4286],

[0130] [-1.8081, -1.4755, -1.2304,..., 2.4286, 2.4286, 2.4286],

[0131] ...,

[0132] [-1.1954, -1.1429, -1.1779,..., 2.4286, 2.4286, 2.4286],

[0133] [-1.2129, -1.3704, -1.4055,..., 2.4286, 2.4286, 2.4286],

[0134] [-1.2654,-1.7906,-1.8256,...,2.4286,2.4286,2.4286]],[[-1.2467,-1.0550,-0.9853,...,2.6400,2.6400,2.6400],

[0135] [-1.3339,-1.1073,-0.9853,...,2.6400,2.6400,2.6400],

[0136] [-1.5779,-1.2467,-1.0027,...,2.6400,2.6400,2.6400],

[0137] ...,

[0138] [-0.9678,-0.9156,-0.9504,...,2.6400,2.6400,2.6400],

[0139] [-0.9853,-1.1421,-1.1770,...,2.6400,2.6400,2.6400],

[0140] [-1.0376,-1.5604,-1.5953,...,2.6400,2.6400,2.6400]]]

[0141] Third sentence of English spoken English spoken visual features [[[-1.6042,-1.3130,-1.1075,...,-0.7993,-0.9192,-1.0219],

[0142] [-1.7925,-1.4158,-1.2959,...,-0.9020,-0.9877,-1.1075],

[0143] [-1.7240,-1.3815,-1.2445,...,-1.0219,-0.9705,-1.0219],

[0144] ...,

[0145] [-1.6213,-1.3130,-1.0904,...,-1.3473,-1.2274,-1.1075],

[0146] [-1.7754,-1.4843,-1.3130,...,-1.6042,-1.3987,-1.1760],

[0147] [-2.0837,-1.7240,-1.5870,...,-1.8097,-1.5185,-1.3987]],[[-1.5105,-1.2129,-1.0028,...,-0.6877,-0.8102,-0.9153],

[0148] [-1.7031,-1.3179,-1.1954,...,-0.7927,-0.8803,-1.0028],

[0149] [-1.6331,-1.2829,-1.1429,...,-0.9153,-0.8627,-0.9153],

[0150] ...,

[0151] [-1.5280,-1.2129,-0.9853,...,-1.2479,-1.1253,-1.0028],

[0152] [-1.6856,-1.3880,-1.2129,...,-1.5105,-1.3004,-1.0728],

[0153] [-2.0007,-1.6331,-1.4930,...,-1.7206,-1.4230,-1.3004]],[[-1.2816,-0.9853,-0.7761,...,-0.4624,-0.5844,-0.6890],

[0154] [-1.4733,-1.0898,-0.9678,...,-0.5670,-0.6541,-0.7761],

[0155] [-1.4036,-1.0550,-0.9156,...,-0.6890,-0.6367,-0.6890],

[0156] ...,

[0157] [-1.2990,-0.9853,-0.7587,...,-1.0201,-0.8981,-0.7761],

[0158] [-1.4559,-1.1596,-0.9853,...,-1.2816,-1.0724,-0.8458],

[0159] [-1.7696, -1.4036, -1.2641,..., -1.4907, -1.1944, -1.0724]]

[0160] Fourth sentence of English spoken English spoken visual features of the English [[[-0.9534, -0.7993, -0.5767,..., 2.2489, 2.2489, 2.2489],

[0161] [-0.7650, -0.6452, -0.4911,..., 2.2489, 2.2489, 2.2489],

[0162] [-0.6109, -0.5596, -0.5424,..., 2.2489, 2.2489, 2.2489],

[0163] ...,

[0164] [-0.3883, -0.3541, -0.3541,..., 2.2489, 2.2489, 2.2489],

[0165] [-0.5424, -0.5938, -0.6281,..., 2.2489, 2.2489, 2.2489],

[0166] [-0.6452, -0.8335, -0.9705,..., 2.2489, 2.2489, 2.2489]],[[-0.8452, -0.6877, -0.4601,..., 2.4286, 2.4286, 2.4286],

[0167] [-0.6527, -0.5301, -0.3725,..., 2.4286, 2.4286, 2.4286],

[0168] [-0.4951, -0.4426, -0.4251,..., 2.4286, 2.4286, 2.4286],

[0169] ...,

[0170] [-0.2675, -0.2325, -0.2325,..., 2.4286, 2.4286, 2.4286],

[0171] [-0.4251, -0.4776, -0.5126,..., 2.4286, 2.4286, 2.4286],

[0172] [-0.5301, -0.7227, -0.8627,..., 2.4286, 2.4286, 2.4286], [[-0.6193, -0.4624, -0.2358,..., 2.6400, 2.6400, 2.6400],

[0173] [-0.4275, -0.3055, -0.1487,..., 2.6400, 2.6400, 2.6400],

[0174] [-0.2707, -0.2184, -0.2010,..., 2.6400, 2.6400, 2.6400],

[0175] ...,

[0176] [-0.0441, -0.0092, -0.0092,..., 2.6400, 2.6400, 2.6400],

[0177] [-0.2010, -0.2532, -0.2881,..., 2.6400, 2.6400, 2.6400],

[0178] [-0.3055, -0.4973, -0.6367,..., 2.6400, 2.6400, 2.6400]]]

[0179] Fifth sentence English spoken English spoken visual features [[ [0.0912, 0.1939, 0.0741,..., 0.7248, 0.9303, 1.1187],

[0180] [-0.4226, -0.0629, 0.0227,..., 0.8104, 0.9817, 1.1358],

[0181] [-0.5253, -0.1486, 0.0227,..., 0.9988, 0.9817, 0.8618],

[0182] ...,

[0183] [-0.5938, -0.7993, -0.8849,..., -0.4568, -0.3541, -0.3027],

[0184] [-0.9020, -1.1075, -1.1760,..., -0.6794, -0.6281, -0.5938],

[0185] [-1.3473,-1.3987,-1.5014,...,-0.9534,-1.0733,-1.0562]],[[0.2227,0.3277,0.2052,...,0.8704,1.0805,1.2731],

[0186] [-0.3025,0.0651,0.1527,...,0.9580,1.1331,1.2906],

[0187] [-0.4076,-0.0224,0.1527,...,1.1506,1.1331,1.0105],

[0188] ...,

[0189] [-0.4776,-0.6877,-0.7752,...,-0.3375,-0.2325,-0.1800],

[0190] [-0.7927,-1.0028,-1.0728,...,-0.5651,-0.5126,-0.4776],

[0191] [-1.2479,-1.3004,-1.4055,...,-0.8452,-0.9678,-0.9503]],

[0192] [[0.4439,0.5485,0.4265,...,1.0888,1.2980,1.4897],

[0193] [-0.0790,0.2871,0.3742,...,1.1759,1.3502,1.5071],

[0194] [-0.1835,0.1999,0.3742,...,1.3677,1.3502,1.2282],

[0195] ...,

[0196] [-0.2532,-0.4624,-0.5495,...,-0.1138,-0.0092,0.0431],

[0197] [-0.5670,-0.7761,-0.8458,...,-0.3404,-0.2881,-0.2532],

[0198] [-1.0201, -1.0724, -1.1770,..., -0.6193, -0.7413, -0.7238]]

[0199] (3) The English spoken language waveform features of the processed English spoken language in step one

[0200] The English spoken language waveform features of the first sentence of English spoken language

[0201] [5.5601e-04, 5.5601e-04, 5.5601e-04,..., -5.0515e-01, -7.1300e-01, 1.4834e-03]

[0202] The English spoken language waveform features of the second sentence of English spoken language

[0203] [4.1946e-05, 4.1946e-05, 4.1946e-05,..., -2.8711e-02, -2.2843e-03, 1.9071e-03]

[0204] The English spoken language waveform features of the third sentence of English spoken language

[0205] [-0.1186, -0.0628, -0.0290,..., 0.3281, 0.4915, 0.5587]

[0206] The English spoken language waveform features of the fourth sentence of English spoken language

[0207] [0.0002, 0.0002, 0.0002,..., -0.0297, -0.0981, -0.0755]

[0208] The English spoken language waveform features of the fifth sentence of English spoken language

[0209] [-0.6376, -0.5402, -0.3749,..., 0.2712, 3.9806, -2.0851]

[0210] Step two: execute the "English spoken language emotion diagnosis module"

[0211] The English spoken language emotion diagnosis module is to input the three acoustic features of the above-mentioned step one of the English spoken language to be diagnosed, to extract the English spoken language waveform features, English spoken language visual features and English spoken language features respectively, to input to the corresponding encoder, to learn the deeper spoken language emotion features, and to splice the English spoken language bidirectional emotion feature vector, English spoken language convolution emotion feature vector and English spoken language representation emotion feature vector extracted by different encoders to obtain the English spoken language bidirectional-convolution-representation emotion feature vector.

[0212] (1) The English oral two-way emotion feature vectors of the first sentence of English oral are as follows:

[0213] The English oral two-way emotion feature vectors of the first sentence of English oral are as follows:

[0214] [-0.0124, -0.0170, -0.0025,..., 0.0019, 0.0269, 0.0084]

[0215] The English oral two-way emotion feature vectors of the second sentence of English oral are as follows:

[0216] [-0.0368, -0.0059, 0.0127,..., 0.0036, -0.0117, 0.0048]

[0217] The English oral two-way emotion feature vectors of the third sentence of English oral are as follows:

[0218] [-0.0282, -0.0129, -0.0295,..., 0.0043, -0.0085, 0.0231]

[0219] The English oral two-way emotion feature vectors of the fourth sentence of English oral are as follows:

[0220] [-0.0104, -0.0425, -0.0263,..., -0.0056, 0.0213, 0.0263]

[0221] The English oral two-way emotion feature vectors of the fifth sentence of English oral are as follows:

[0222] [0.0032, 0.0253, -0.0490,..., -0.0265, 0.0583, 0.0563](2) The English oral convolution emotion feature vectors obtained by formula (2) are as follows:

[0223] The English oral convolution emotion feature vectors of the first sentence of English oral are as follows:

[0224] [0., 0., 0.,..., 0., 0., 0.]

[0225] The English oral convolution emotion feature vectors of the second sentence of English oral are as follows:

[0226] [0.0000, 0.0000, 0.0000,..., 0.1756, 0.0000, 0.0000]

[0227] The English oral convolution emotion feature vectors of the third sentence of English oral are as follows:

[0228] [0.0000, 0.0000, 0.0000,..., 0.0917, 1.3347, 2.9555]

[0229] English conversational representation sentiment feature vector of the fourth sentence of English conversational

[0230] [0.0000, 0.0000, 0.0000,..., 0.4597, 0.0000, 0.0000]

[0231] English conversational representation sentiment feature vector of the fifth sentence of English conversational

[0232] [0.0000, 0.0000, 0.0000,..., 1.0251, 0.2978, 0.1680]

[0233] (3) The English conversational representation sentiment feature vector of the English conversational obtained through formula (3) is as follows:

[0234] English conversational representation sentiment feature vector of the first sentence of English conversational

[0235] [-0.3337, -0.6586, 0.2126,..., -0.2429, -0.3902, 0.0193]

[0236] English conversational representation sentiment feature vector of the second sentence of English conversational

[0237] [-0.2351, -0.3815, -0.0053,..., 0.2747, 0.0991, -0.0712]

[0238] English conversational representation sentiment feature vector of the third sentence of English conversational

[0239] [-0.0337, -0.5608, 0.2722,..., 0.1183, 0.1333, 0.0953]

[0240] English conversational representation sentiment feature vector of the fourth sentence of English conversational

[0241] [-0.3347, -0.2667, 0.0440,..., 0.6145, 0.0385, -0.3312]

[0242] English conversational representation sentiment feature vector of the fifth sentence of English conversational

[0243] [0.4655,-1.1238,0.0838,...,0.1842,-0.0842,-0.3310](4)After splicing, the English oral two-way convolutional representation emotion feature vector of the English oral is as follows: The English oral two-way convolutional representation emotion feature vector of the first sentence of English oral

[0244] [0.4621,0.2205,0.0000,...,0.0000,0.0000,0.0215]

[0245] The English oral two-way convolutional representation emotion feature vector of the second sentence of English oral

[0246] [0.7225,0.5247,0.0327,...,0.3052,0.1102,0.0000]

[0247] The English oral two-way convolutional representation emotion feature vector of the third sentence of English oral

[0248] [0.3338,0.0000,0.1461,...,0.1315,0.1481,0.1058]

[0249] The English oral two-way convolutional representation emotion feature vector of the fourth sentence of English oral

[0250] [0.4502,0.2229,0.0000,...,0.6828,0.0427,0.0000]

[0251] The English oral two-way convolutional representation emotion feature vector of the fifth sentence of English oral

[0252] [0.2383,0.1560,0.0000,...,0.2046,0.0000,0.0000]

[0253] According to formula (4), the English oral emotion probability maximum value of each English oral is obtained, and formula (5) is used to find the emotion category corresponding to the English oral emotion probability maximum value, so as to obtain the English oral emotion diagnosis result.

[0254] The English oral emotion probability maximum value of the first sentence [0.6537]

[0256] The English oral emotion probability maximum value of the second sentence [0.7129]

[0258] The English oral emotion probability maximum value of the third sentence [0.7341]

[0260] The English oral emotion probability maximum value of the fourth sentence [0.6754]

[0262] Fifth sentence English spoken language sentiment probability maximum [0.7943]

[0264] Since the English spoken language sentiment probability maximum of the first five sentences are all greater than 0.6 and less than 0.8, the sentiment diagnosis results of the first five sentences are all natural, and the final English spoken language sentiment diagnosis result is shown as follows:

[0265] Will we ever forget it. Natural

[0266] Gad, your letter came just in time. Natural

[0267] Not at this particular case, Tom, apologized Whittemore. Natural

[0268] Lord, but I'm glad to see you again, Phil. Natural

[0269] God bless them, I hope I'll go on seeing them forever. Natural

Claims

1. A method for diagnosing English spoken emotion, characterized by including: The following is the processing flow: (1) English spoken language preprocessing process: First, read the English spoken language to be diagnosed and enhance the English spoken language signal; Second, divide the enhanced English spoken language signal into multiple shorter English spoken language segments according to a 15-millisecond time interval to obtain the English spoken language waveform features; Third, convert the signal of each English spoken language segment from the time domain to the frequency domain, analyze the frequency components of each English spoken language segment, obtain the spectrum of each English spoken language signal and perform rotation and mapping, and then splice the multiple speech spectrum segments after rotation and mapping to obtain the visual features of English spoken language; Fourth, perform simulated human speech processing on the spectrum of each English spoken language signal after rotation and mapping, and perform spectral feature screening on the spectrum of the English spoken language signal after simulated human speech processing; Fifth, perform dimensionality reduction and numerical processing on the selected English spoken language signal spectral features, and convert the dimensionality reduction and numerical processing of the English spoken language spectral features. Sixth, analyze the changes in the spectral characteristics of spoken English at different points in time, capture the dynamic characteristics of the spectral characteristics of spoken English, and output the frequency characteristics of spoken English. (2) English spoken emotion diagnosis process: First, input the English spoken frequency features into the two-way flow control unit to calculate the English spoken two-way emotion feature vector; Second, the visual features of spoken English are input into the convolutional feature learning network unit to calculate the spoken English convolutional sentiment feature vector. Third, the waveform features of spoken English are input into the speech representation learner to calculate the spoken English representation sentiment feature vector. Fourth, the spoken English bidirectional sentiment feature vector, the spoken English convolutional sentiment feature vector, and the spoken English representation sentiment feature vector are concatenated to generate the spoken English bidirectional-convolutional-representation sentiment feature vector. Fifth, using the formula for calculating the maximum value of spoken English sentiment probability, the maximum value of spoken English sentiment probability is calculated on the spoken English bidirectional-convolutional-representation sentiment feature vector to obtain the maximum value of spoken English sentiment probability. Sixth, read the maximum probability of English spoken emotion obtained from the English spoken emotion extraction module; Seventh, use the English spoken emotion diagnosis result calculation formula to determine the English spoken emotion diagnosis result and output the English spoken emotion diagnosis result. The formula for calculating the maximum probability of English spoken emotion is as follows: Where e represents the base of the natural logarithm function, and i represents the i-th emotion category. There are four emotion categories: anger, sadness, nature, and happiness. The formula for calculating the English spoken emotion diagnosis result is as follows:

Citation Information

Patent Citations

  • Voice emotion recognition method and system

    CN109767790A

  • Comprehensive evaluation method for oral English test

    CN118471233A